Gateways move the traffic. Something still has to account for it.
LiteLLM and Portkey are good inference gateways, and Portkey open-sourced its entire gateway under Apache 2.0 in March 2026, so either will get you to one endpoint across many models in an afternoon. We would no longer group Helicone with them: Mintlify acquired it in March 2026 and feature development has stopped. The limit of a gateway is structural rather than a missing feature. It sees the calls that pass through it, and nothing else: not whether a cheaper model would have held quality, not the invoice that lands at month end, not a control-group proof that a change paid for itself. TensorCost sits beside your gateway and accounts for all of it.
At a glance
Concrete capabilities, not adjectives. Where the proxies are stronger, this table says so.
Where TensorCost is the wrong fit (today)
Three customer shapes where a proxy serves you better than we do.
Pre-production / dev-stage AI apps
If your AI workload is in build phase — no production traffic, no scaled bill — a proxy gets you to one endpoint across many models with hours of integration. TensorCost's pilot pays back at production scale, not in development.
You need 250 providers
Portkey covers more model surfaces than we do today. If you need Replicate + Together + Anyscale + Mistral La Plateforme + ten specialty providers in one gateway, a proxy is the right tool. We go deep on five.
AI spend under $500K/yr
See /icp for the full self-qualification. Below this threshold, the engineering time to integrate TensorCost usually exceeds the recoverable savings. A proxy at the API boundary covers your routing needs at lower overhead.