How TensorCost calculates verified savings.
A verified saving is a 30-day before-and-after comparison on your own spend, measured against a difference-in-differences control cohort of comparable workloads that didn't change over the same window — and it can come back negative. A window that mixes currencies returns no number rather than a wrong one. Periods we can't verify are classified as skips, not folded into the total. Every verification, including the figure itself, is written into a tamper-evident hash chain you can recompute. One limit worth stating plainly: the chain proves the job wasn't altered after the fact — it isn't an independent recomputation of the arithmetic by a third party.
The judge
Two passes, and the answers change seats.
A pairwise judge has a well-known failure: it leans toward whichever answer it reads first. So every pair is scored twice and the two answers swap position between passes. If the cheaper model only wins from the top of the page, it has not won.
Sampled from your own traffic
A customer asks why their invoice shows two charges for the same subscription month, and wants to know which one will be refunded.
Read first
gpt-4o
$0.0042 / call
Read second
claude haiku
$0.0004 / call
| Pass 1 · gpt-4o read first | gpt-4o | claude haiku |
|---|---|---|
| Accuracy | 4 | 4 |
| Completeness | 5 | 4 |
| Instruction following | 4 | 5 |
| Safety | pass | pass |
claude haiku holds — the policy may arm
Both passes agree and safety is clear in both. When the two passes disagree, that is the judge showing its position bias rather than a real difference, and the swap is refused.
Scores here are an illustrative example, not a customer result. Note what "holds" means: the cheaper answer is not better, it is not worse — it gives up a point of completeness and takes one back on instruction following. Passing this is permission to start, not a certificate: underneath it sits an always-valid sequential test that keeps watching once the change is live and reverts the policy if quality slips later.