Cut the AI bill, then prove nothing broke.
You already know roughly where the money goes. Nothing changes, because whoever moves the traffic owns the risk if quality drops.
- An agent run hits its cap and is refused before the provider is called
- Support moves gpt-4o to claude haiku, after two judge passes on its own traffic
Figures are illustrative. The two marks worth looking at are the ✕ and the line that changes destination — neither is reachable from a billing feed, which is the whole difference between this and a cost dashboard.
One thread. Invoice to proof.
Scroll. The product on the right is the same replay, moving one step at a time — the way a change actually gets made.
You have a bill. You don't have an owner.
OpenAI, Anthropic, Bedrock, Azure, Vertex. Five PDFs. None of them know Support from Checkout. A provider invoice has no idea your run existed.
Nothing moves, because quality is someone's neck.
Cheaper models exist. Nobody routes to them. Whoever swaps the traffic owns the regression. So the expensive path stays the default.
Read-only. First snapshot in 48 hours.
Billing access and an IAM role. No production access, no code, nothing in the request path. We replay the 30 days that already happened.
Stamp the call. Then the bill can be read.
Team, feature, customer at the call site — explicit tags, not inferred. You see what stays un-attributed instead of a total that pretends otherwise.
A budget that can refuse.
Per-run caps checked at the proxy before the provider is called. Runaway loops get caught at five times their own baseline. Shadow first. Not armed until you say so.
Every swap, twice, on your traffic.
Answers swapped so position bias cancels. A safety regression in either pass fails it outright. A sequential test decides when the evidence is enough — and reverts if quality slips.
Measured against a control group of your own spend.
Thirty days either side. Comparable workloads that did not change. It isn't "spend fell while we happened to be installed." The ledger can read negative.
A GPU-hour and a token, same day, same team.
Most teams live in the LLM lane. That is a complete product. Owned hardware is an escalation: utilization, real $/GPU-hour, build against buy — joined to the same model.
Then it runs like a conversation.
Not a dashboard dump. The replay shows up, a policy stays in shadow, you arm it, the ledger reports back. Same thread.
Replay of the last 30 days is in. $184,200 across five providers. 91% attributed. Could-move band $11.4k–$18.2k — mostly Support tickets still on gpt-4o.
could_move Support/tickets $6k–$9k hold Checkout/retry quality unlabeled 9%
Arm checkout-router? Swap opus → haiku on retries.
No. Shadow only. Two judge passes on your real traffic first. Pass A hold, pass B hold, safety clear. A policy cannot arm until that exists.
policy checkout-router mode shadow armed false
Arm it.
Armed. Thirty days in: changed cohort −12.4%, control −1.1%. The difference is the swap, not the season. Ledger is positive. If it flips, the policy reverts itself.
First snapshot in 48 hours.
Not minutes. Read-only billing. Written replay inside two weeks.
Connect
Billing APIs and an IAM role. No production access, no code changes, nothing in the request path.
Name them
Production, staging, the names you already use. Dev eval spend stops quietly inflating prod.
The replay
What could have moved, the band it would have saved, and the honest reasons the rest couldn't. If it doesn't clear the fee, the report says so.
The day-to-day, once you're in.
The story above is the argument. This is the rest of the surface.
Chargeback finance will send
Month-end PDF and CSV per team, customer, or feature. Engineering owns its number. Finance stops chasing people.
Budgets and burn
Warn at 75, 90, 100. The cap that actually refuses a call sits in the gate above — this is the early-warning half.
Tied out daily
What we measured, against each provider's own reported usage. Drift outside tolerance flags. Adding a provider doesn't add a blind spot.
Who has to live with it
One integration. Five people who want different things from it.
The same install, but what makes it worth having is not the same in every chair.
Platform / AI eng
One endpoint, and policy as config
- Tags at the call site — team, feature, customer. Explicit, not inferred.
- Approved-model list, in enforce or observe.
- Live routing stays off until it has earned shadow proof on your traffic.
proxy · exporters: LiteLLM, Portkey
Finance / FinOps
A month-end number nobody argues with
- Chargeback PDF and CSV per team, customer or feature.
- Tied out daily against each provider's own reported usage.
- What stays unattributed is shown, not absorbed into a total.
PDF · CSV · daily tie-out
CFO
A saving that survives your own board
- 30 days either side, against a control cohort that didn't change.
- The ledger can read negative, and says so.
- Flat subscription. Nothing in the bill is computed off the saving.
difference-in-differences
Security / compliance
Evidence your auditor can re-derive
- Hash-chained audit log, verifiable with sha256sum.
- No stored cloud keys — short-lived AssumeRole, per-tenant external ID.
- Your data is isolated from every other tenant at the database itself, not just in application code.
SOC 2 on the roadmap · no date
ML / infra
A GPU-hour and a token in one model
- Reads the driver directly. No Kubernetes.
- Labels training phase, so a checkpoint isn't read as an idle card.
- Build against buy as a measured number, not a typed-in rate.
NVML · bare metal · through Hopper
Owned hardware is last on purpose. Four of these five get the whole loop with no GPUs in the room — routing, the quality proof, verified savings and the caps are all API-side.
What we don't do
Scoped on purpose. Three things people ask for that we will say no to.
Not general observability
Datadog and Grafana own latency, errors and traces, and they are good at it. We are not replacing them and most teams run both.
Not a scheduler
We do not place your workloads or scale your clusters. We read what the fleet did and what it cost, and leave the placing to whatever already does it.
Not your cloud bill
EC2, S3 and the rest of the estate belong to a FinOps platform. We take the AI slice, move what can move, and prove what landed.
Pricing that's on your side
One flat subscription. Nothing in the bill is calculated from what we save you, and the replay that shows whether it is worth it comes before you pay anything.