DISPATCH
30–50% LOWER INFERENCE COST
The Managed Model Router
One API key for your coding tools, agents and apps. Dispatch sends each request to the model with the best price for the work, attributes every dollar, and keeps your data out of everyone’s training set.
WORKS WITH YOUR CODING TOOLS, AGENT FRAMEWORKS AND APPS
SAVINGS
30–50%
less spent on inference, for the same work
Most of what a team sends doesn’t need a frontier model. Auto keeps that work on Swift, efficiency recommendations trim what every request carries, and attribution finds the spend nobody owns.
Everything between your team and the models
Five jobs you would otherwise staff, script or skip. Dispatch does them on every request, from every coding tool, agent and app.
01 · FULLY MANAGED ROUTER
We find the best price for performance, so you don’t have to
Our team benchmarks and prices the models behind every tier, and re-tunes them as the market moves. Auto reads each request and sends it to the tier that fits: frontier models for the hard turns, fast ones for everything else.
Tiers, not model names, in every tool
Re-benchmarked as new models ship
Auto decides per request, not per project
QUALITY
Cheaper where it’s safe. Frontier where it counts.
“Will cheaper models make our output worse?” It’s the first thing an engineer asks. Here’s how Dispatch makes sure they don’t.
Every new model is benchmarked first
When a model ships, our team scores it on independent intelligence, coding and agentic benchmarks and tests it on real agent workloads. It joins a tier only if it matches what that tier serves today, at a better price.
Auto escalates when the work gets hard
Every request is judged on its own. A failed tool call, an error to diagnose or a plan to make goes to Frontier, while settled, mechanical steps stay on Swift. One hard turn never makes the whole session expensive.
Pin a tier whenever you need certainty
Send darcy-frontier or darcy-swift instead of darcy-auto and every request stays on that tier, from any tool, agent or API. Admins can also turn tiers on or off for each team.
ONE AGENT SESSION, TURN BY TURN
02 · AUTOMATIC COST ATTRIBUTION
Every request tagged the moment it arrives
Dispatch inspects requests as they come in and files them under the team, client or project they were for. Nobody tags anything by hand, and nothing piles up as “unallocated”.
Tags from the key, its guardrail and the request itself
Classification for client and project questions
Teams inherit tags from their guardrail
03 · FULL COST REPORTING
Showback, chargeback and client invoices from one ledger
Break spend down by person, team, tier, tag or client, over any window. Share it with finance, charge it back to cost centers, or turn it into the line items on a client’s invoice.
Live report links for finance, no login needed
CSV export for your ERP or billing system
Per-client totals ready to invoice
04 · EFFICIENCY OPTIMIZATION
Fewer tokens for the same work
Dispatch reads what every agent turn spends its tokens on. When command output, file reads or long histories are doing the spending, it tells you which fix to install, like RTK or Headroom, and for whom.
Token anatomy for every agent turn
Recommendations ranked by what they save
Savings measured after you install them
05 · BUDGETS, GUARDRAILS AND ZDR
Control that holds on every request
Budgets per person or team that warn before they stop. Guardrails that catch secrets and personal data before any model sees them. And Zero Data Retention, enforced on every tier by default.
Live in an afternoon
STEP 01
Point your tools at Dispatch
Swap the base URL and key in your coding tools, agent frameworks or any OpenAI- or Anthropic-compatible app. Nothing to install.
STEP 02
Choose Auto, or pin a tier
Auto picks the tier for each request. Pin Frontier or Swift when a workload needs one.
model: “darcy-auto”
STEP 03
Watch the ledger fill
Every request lands tagged, priced and reported within seconds, with budgets and guardrails already in force.
Stop guessing which model to pay for
Bring your tools, agents and apps. We’ll bring the router, the ledger and the guardrails, and 30–50% off your inference bill.
