TensorCost TensorCost
The last 30 days, in writing

Cut the AI bill, then prove nothing broke.

You already know roughly where the money goes. Nothing changes, because whoever moves the traffic owns the risk if quality drops.

app.tensorcost.com · replay
WHAT SPENDSSupportCheckoutAgent runsEval jobsNotebooksOpenAIAnthropicBedrockAzure OpenAIVertexYour H100sWHERE IT LANDSNO CONTROL POINTnothing between the call and the card$184,200 last 30 days46% attributedno call can be refusedtagged at the call siteCONTROL · BEFORE THE CALLEVIDENCE · AFTER THE FACTAdmissionRouterJudgeAttributionLedgerAudit chain$184,200 last 30 days91% attributed · 9% still unlabeledmeasured against a control cohortWHAT SPENDSSupportCheckoutAgent runs+2 moreOpenAIAnthropicBedrock+3 moreWHERE IT LANDSNO CONTROL POINT46% attributedCONTROL · BEFORE THE CALLAdmissionRouterJudgeEVIDENCE · AFTER THE FACTAttributionLedgerAudit chain91% attributed · 9% still unlabeled
  • An agent run hits its cap and is refused before the provider is called
  • Support moves gpt-4o to claude haiku, after two judge passes on its own traffic

Figures are illustrative. The two marks worth looking at are the ✕ and the line that changes destination — neither is reachable from a billing feed, which is the whole difference between this and a cost dashboard.

Read-only billing · nothing in the request path
Read-only billing · nothing in the request path
Already on TensorCost
Works withOpenAIAnthropicBedrockAzureVertex

One thread. Invoice to proof.

Scroll. The product on the right is the same replay, moving one step at a time — the way a change actually gets made.

01 The invoice

You have a bill. You don't have an owner.

OpenAI, Anthropic, Bedrock, Azure, Vertex. Five PDFs. None of them know Support from Checkout. A provider invoice has no idea your run existed.

02 The freeze

Nothing moves, because quality is someone's neck.

Cheaper models exist. Nobody routes to them. Whoever swaps the traffic owns the regression. So the expensive path stays the default.

03 Connect

Read-only. First snapshot in 48 hours.

Billing access and an IAM role. No production access, no code, nothing in the request path. We replay the 30 days that already happened.

04 See

Stamp the call. Then the bill can be read.

Team, feature, customer at the call site — explicit tags, not inferred. You see what stays un-attributed instead of a total that pretends otherwise.

05 Gate

A budget that can refuse.

Per-run caps checked at the proxy before the provider is called. Runaway loops get caught at five times their own baseline. Shadow first. Not armed until you say so.

06 Judge

Every swap, twice, on your traffic.

Answers swapped so position bias cancels. A safety regression in either pass fails it outright. A sequential test decides when the evidence is enough — and reverts if quality slips.

07 Prove

Measured against a control group of your own spend.

Thirty days either side. Comparable workloads that did not change. It isn't "spend fell while we happened to be installed." The ledger can read negative.

08 If you own the cards

A GPU-hour and a token, same day, same team.

Most teams live in the LLM lane. That is a complete product. Owned hardware is an escalation: utilization, real $/GPU-hour, build against buy — joined to the same model.

Then it runs like a conversation.

Not a dashboard dump. The replay shows up, a policy stays in shadow, you arm it, the ledger reports back. Same thread.

#cost · TensorCost
TensorCost

Replay of the last 30 days is in. $184,200 across five providers. 91% attributed. Could-move band $11.4k–$18.2k — mostly Support tickets still on gpt-4o.

could_move  Support/tickets   $6k–$9k
hold        Checkout/retry    quality
unlabeled   9%
You

Arm checkout-router? Swap opus → haiku on retries.

TensorCost

No. Shadow only. Two judge passes on your real traffic first. Pass A hold, pass B hold, safety clear. A policy cannot arm until that exists.

policy   checkout-router
mode     shadow
armed    false
You

Arm it.

TensorCost

Armed. Thirty days in: changed cohort −12.4%, control −1.1%. The difference is the swap, not the season. Ledger is positive. If it flips, the policy reverts itself.

First snapshot in 48 hours.

Not minutes. Read-only billing. Written replay inside two weeks.

power

Connect

Billing APIs and an IAM role. No production access, no code changes, nothing in the request path.

sell

Name them

Production, staging, the names you already use. Dev eval spend stops quietly inflating prod.

search

The replay

What could have moved, the band it would have saved, and the honest reasons the rest couldn't. If it doesn't clear the fee, the report says so.

The day-to-day, once you're in.

The story above is the argument. This is the rest of the surface.

description

Chargeback finance will send

Month-end PDF and CSV per team, customer, or feature. Engineering owns its number. Finance stops chasing people.

notifications_active

Budgets and burn

Warn at 75, 90, 100. The cap that actually refuses a call sits in the gate above — this is the early-warning half.

cloud_queue

Tied out daily

What we measured, against each provider's own reported usage. Drift outside tolerance flags. Adding a provider doesn't add a blind spot.

Who has to live with it

One integration. Five people who want different things from it.

The same install, but what makes it worth having is not the same in every chair.

Platform / AI eng

One endpoint, and policy as config

  • Tags at the call site — team, feature, customer. Explicit, not inferred.
  • Approved-model list, in enforce or observe.
  • Live routing stays off until it has earned shadow proof on your traffic.

proxy · exporters: LiteLLM, Portkey

Finance / FinOps

A month-end number nobody argues with

  • Chargeback PDF and CSV per team, customer or feature.
  • Tied out daily against each provider's own reported usage.
  • What stays unattributed is shown, not absorbed into a total.

PDF · CSV · daily tie-out

CFO

A saving that survives your own board

  • 30 days either side, against a control cohort that didn't change.
  • The ledger can read negative, and says so.
  • Flat subscription. Nothing in the bill is computed off the saving.

difference-in-differences

Security / compliance

Evidence your auditor can re-derive

  • Hash-chained audit log, verifiable with sha256sum.
  • No stored cloud keys — short-lived AssumeRole, per-tenant external ID.
  • Your data is isolated from every other tenant at the database itself, not just in application code.

SOC 2 on the roadmap · no date

ML / infra

A GPU-hour and a token in one model

  • Reads the driver directly. No Kubernetes.
  • Labels training phase, so a checkpoint isn't read as an idle card.
  • Build against buy as a measured number, not a typed-in rate.

NVML · bare metal · through Hopper

Owned hardware is last on purpose. Four of these five get the whole loop with no GPUs in the room — routing, the quality proof, verified savings and the caps are all API-side.

What we don't do

Scoped on purpose. Three things people ask for that we will say no to.

monitoring

Not general observability

Datadog and Grafana own latency, errors and traces, and they are good at it. We are not replacing them and most teams run both.

dns

Not a scheduler

We do not place your workloads or scale your clusters. We read what the fleet did and what it cost, and leave the placing to whatever already does it.

receipt_long

Not your cloud bill

EC2, S3 and the rest of the estate belong to a FinOps platform. We take the AI slice, move what can move, and prove what landed.

Pricing that's on your side

One flat subscription. Nothing in the bill is calculated from what we save you, and the replay that shows whether it is worth it comes before you pay anything.

Developer
Free
Free
checkLLM spend observe by team, model & feature
check2 provider connections
checkNo GPU agents (upgrade to Scale)
check7-day metric retention
check2 budgets with burn-rate alerts
checkSavings preview (upgrade to apply)
checkMCP access
checkCommunity support
Start free
Most popular
Scale
$499/month
Scale
checkEverything in Developer, plus:
check1–2 customer GPUs monitored
checkLive applied routing + semantic prompt cache
checkPlatform-paid judging
checkTracing, governance, chargeback, guardrails
checkAll 5 providers + invoice reconciliation
checkSAML SSO · API/SDK
checkPriority support · 99.9% uptime target
Enterprise
Contact us
Enterprise
checkEverything in Scale, plus:
checkUnlimited / contracted GPUs
checkJudge + embedding BYOK
checkSelf-hosted term license or SaaS
checkSCIM · white-label · custom adapters
checkDedicated CSM · 99.95% SLA · custom DPA/BAA
Contact us

See what the last 30 days would have saved.

If the dollars don't clearly clear our fee, the report says so.