Synthetic colleagues, with a job description and a cost centre.
These are our own, running in our own estate. Each one has a charter written in business terms, a place it runs, a measured latency and a daily cost — because a model without those four things is a demo.
Who works here
Seven roles, each scheduled onto whichever target its trade-off points at. The placement is a configuration change, not a migration — which is the whole reason the fabric abstraction exists.
| Colleague | What it does | Runs on | Requests · 24h |
|---|---|---|---|
Aria · Voice Conciergellama-3.3-70b · Groq LPU | Answers the company line, qualifies inbound work, books discovery calls and hands a written brief to a human partner. | Groq | 1,284 |
Scribe · Transcriptionwhisper-large-v3-turbo | Transcribes and diarises every client call, then files a searchable summary against the right engagement. | Helios | 412 |
Echo · Speech Synthesisxtts-v2 · cloned brand voice | Gives every synthetic colleague a consistent, brand-owned voice across phone, video and the client portal. | Helios | 967 |
Vega · Video Presenterdiffusion pipeline · avatar-2 | Renders briefing videos and client walkthroughs on demand — one avatar, any language, published in minutes. | Modal | 88 |
Atlas · Analystllama-3.3-70b · vLLM | Reads the warehouse, drafts the weekly client report and flags the three numbers a partner should look at. | Helios | 3,140 |
Lint · Code Reviewerqwen2.5-coder-32b | Reviews every pull request against the engagement's architecture decisions before a human ever opens it. | Helios | 622 |
Index · Retrievalbge-m3 + qdrant | Keeps every engagement's documents, decisions and telemetry in one vector index the whole workforce reads from. | Helios | 18,400 |
Where a role runs is a trade-off, and it changes
Latency, privacy and cost decide it. Because those three move — a model gets cheaper, a regulation tightens, a load doubles — we build the routing layer first and treat placement as reversible.
Owned hardware
Cheapest per GPU-hour and completely private, because the machines are already paid for. This is where steady load and regulated data belong.
Steady load · regulated dataCloud accelerators
Elastic and regional, with data residency you can point at on a map. This is where client-facing work with a residency requirement belongs.
Elastic · residencyServerless inference
Instant, and you pay nothing for idle. This is where spiky work belongs, and where latency a human would notice has to disappear.
Spiky · latency-criticalWe will build you one of these in six weeks.
The AI Workforce Pilot takes one role to production with its evaluation harness, its cost model and its human checkpoints — chosen with you on the single criterion that the outcome is measurable.
Start a pilot