price of compute

Rent the GPU, or pay per token?

The two sides of this site, bridged: today’s GPU rental medians against the labs’ token list prices. We supply both price sides live; the throughput is yours to supply, because it varies enormously with model, quantization, and batching — and we don’t print performance numbers we can’t stand behind. The arithmetic is shown under every result.

Your self-hosted cost

$2.933 per million tokens

$7.39/hr × 1,000,000 ÷ (1,000 tok/s × 3600s × 70%)

API modelBlended $/MTokSelf-host vs APIBreak-even tok/s
deepseek-v4-flash deepseek$0.1816.76×16,757
mistral-small-4 mistral$0.2611.17×11,172
gpt-5.6-luna openai$0.456.52×6,517
deepseek-v4-pro deepseek$0.545.39×5,393
mistral-large-3 mistral$0.753.91×3,910
gemini-3.5-flash-lite google$0.853.45×3,450
claude-haiku-4-5 anthropic$2.001.47×1,466
gemini-3.6-flash google$3.000.98×978
mistral-medium-3.5 mistral$3.000.98×978
grok-4.5 xai$3.000.98×978
gemini-3.5-flash google$3.380.87×869
claude-sonnet-5 anthropic$4.000.73×733
gemini-3.1-pro-preview google$4.500.65×652
gpt-5.6-terra openai$4.500.65×652
gpt-5.4 openai$5.630.52×521
claude-opus-4-8 anthropic$10.000.29×293
claude-opus-5 anthropic$10.000.29×293
gpt-5.6-sol openai$11.250.26×261
claude-fable-5 anthropic$20.000.15×147
gpt-5.5-pro openai$67.500.04×43

Below 1.00× self-hosting is cheaper per token than that API at your inputs; the break-even column is the throughput where they match. This compares raw token economics only — it ignores your engineering time, the API’s quality/latency, and that a rented GPU bills whether or not you keep it busy (that’s the utilization slider).