TensorCost TensorCost
See · Gate · Prove

Cut the AI bill, then prove nothing broke.

You already know roughly where the money goes. Nothing changes, because whoever moves the traffic owns the risk if quality drops. TensorCost shows the last 30 days, gates the next change, and proves it held.

Most teams live in the LLM lane. That is a complete product. Owned hardware is a later chapter, if you also own the cards.

Three chapters. One story.

Training
Inference
Fine-tune
Dev
Idle

01 See

Stamp the call. Then the bill can be read.

Team, feature, customer at the call site. Explicit tags, not inferred. OpenAI, Anthropic, Bedrock, Azure, Vertex in one explorer. You see what stays un-attributed instead of a total that pretends otherwise. Illustrative: Support spent $42K on GPT-4o last month, and $11K of it answered tickets a smaller model would've handled fine.

02 Gate

A budget that can refuse.

Per-run caps checked at the proxy before the provider is called. Runaway loops get caught at five times their own baseline. Shadow first, and not armed until you say so. A short YAML policy turns a pattern into a guardrail, and you preview what it would have done against last week's real numbers before anything goes live.

0%
control

03 Prove

Measured against a control group of your own spend.

Thirty days either side. Comparable workloads that did not change. It isn't "spend fell while we happened to be installed." The ledger can read negative. If the dollars don't clearly clear the fee, the report says so.

What we read, and from where

Six sources, one schema. Where a source is thinner than the others, it says so.

OpenAI

Usage and cost per model, normalized per call. Input, output and reasoning tokens kept separate, because reasoning is priced differently and folding it in hides the expensive calls.

Anthropic

Same schema, including cached input tokens, which change the arithmetic of a swap enough that ignoring them makes a cheaper model look worse than it is.

Amazon Bedrock

Per-model usage read alongside the AWS bill, so a Bedrock line and an EC2 line are not being compared to each other by hand in a spreadsheet.

Azure OpenAI

Deployment-level usage, onboarded without you handing over a client secret. The secretless path is the only path; there is no fallback that stores one.

Google Vertex

Per-model usage into the same schema as the other four, so a team that spans three providers gets one number rather than three exports.

Your own GPUs

A host-level agent reading NVML directly on bare metal, no Kubernetes required, labeling training phase so a checkpoint write is not read as an idle card. Read-only MIG inspection. Through Hopper, with no Blackwell support yet.

See what the last 30 days would have saved.

Read-only billing access. First snapshot in 48 hours. Written report inside two weeks: what could have moved, the band it would have saved, and the honest reasons the rest couldn't. If you also own the cards, the same ledger joins GPU-hours to tokens. Later chapter, not the opener.