TensorCost TensorCost
TensorCost vs LiteLLM / Portkey / Helicone

Gateways move the traffic. Something still has to account for it.

LiteLLM and Portkey are good inference gateways, and Portkey open-sourced its entire gateway under Apache 2.0 in March 2026, so either will get you to one endpoint across many models in an afternoon. We would no longer group Helicone with them: Mintlify acquired it in March 2026 and feature development has stopped. The limit of a gateway is structural rather than a missing feature. It sees the calls that pass through it, and nothing else: not whether a cheaper model would have held quality, not the invoice that lands at month end, not a control-group proof that a change paid for itself. TensorCost sits beside your gateway and accounts for all of it.

At a glance

Concrete capabilities, not adjectives. Where the proxies are stronger, this table says so.

LiteLLM / Portkey / Helicone
TensorCost
Primary buyer: developer or platform engineer building agentic apps.
Primary buyer: platform engineering + finance / FinOps + procurement.
Sits in the request path. Mandatory hop on every inference call.
Post-decision attribution today; inline routing shipped in August 2026 and stays default-off on every policy until it has earned shadow proof on your own traffic. You can also apply recommendations through your own gateway. We ship config exporters for LiteLLM and Portkey; the Helicone exporter still works, though Helicone itself stopped active development after the March 2026 Mintlify acquisition.
Number of providers: LiteLLM ~100, Portkey ~250, Helicone ~25.
Connections to the major AI providers with deep cost attribution (Bedrock, OpenAI, Anthropic, Vertex, Azure OpenAI). Provider breadth is intentionally narrower so each connection goes deep on cost, not just connectivity.
Routing policy: weighted fallback, retries, A/B splits. Quality verification is the developer's problem.
Routing recommendation includes a quality verification pass — evaluating output quality before we suggest the swap. Quality regression is our problem, not the developer's.
Observability: per-request traces, dashboards, alerts.
Spend visibility is one surface among several. Adds: per-team / per-feature attribution rolled up against the cloud invoice, with daily reconciliation against the actual bill.
GPU fleet: not in scope. These are API-layer tools.
Hardware telemetry per node — utilization, partition health, spot eligibility, agent connectivity. The hardware surface lives next to the AI spend surface in one record.
Pricing: per-request, per-seat, or self-hosted (compute-only).
Flat platform fee — $99 Growth, $499 Scale, negotiated for Enterprise. You keep 100% of what we save you. Every saving is still logged against a before/after window you can audit — see [/methodology](/methodology).
Savings ledger: none. Cost shown is the bill, not the savings.
Tamper-evident, append-only savings record. Customer-verifiable. Auditor-exportable. Public methodology page.
Compliance: Portkey SOC 2 Type II; LiteLLM/Helicone self-host or hosted-tier compliance varies.
No SOC 2 report yet — no auditor engaged and no date we would stand behind. The public trust portal at app.tensorcost.com/trust documents the controls we actually run today, and we will complete your security questionnaire.
Recommendations: minimal — gateway-level fallback or cache only.
4 shipped recommenders: model routing, repeat-cost reduction, capacity right-sizing, runaway-job catch. Each writes its verified savings to the record when accepted.

Where TensorCost is the wrong fit (today)

Three customer shapes where a proxy serves you better than we do.

Pre-production / dev-stage AI apps

If your AI workload is in build phase — no production traffic, no scaled bill — a proxy gets you to one endpoint across many models with hours of integration. TensorCost's pilot pays back at production scale, not in development.

You need 250 providers

Portkey covers more model surfaces than we do today. If you need Replicate + Together + Anyscale + Mistral La Plateforme + ten specialty providers in one gateway, a proxy is the right tool. We go deep on five.

AI spend under $500K/yr

See /icp for the full self-qualification. Below this threshold, the engineering time to integrate TensorCost usually exceeds the recoverable savings. A proxy at the API boundary covers your routing needs at lower overhead.

See where the proxy ends and the cost lens starts.

See the last 30 days alongside whichever proxy you already run. Read-only, written report inside two weeks. First snapshot in 48 hours. No card up front.