Rent the GPU, or pay per token?
The two sides of this site, bridged: today’s GPU rental medians against the labs’ token list prices. We supply both price sides live; the throughput is yours to supply, because it varies enormously with model, quantization, and batching — and we don’t print performance numbers we can’t stand behind. The arithmetic is shown under every result.
Your self-hosted cost
$2.933 per million tokens
$7.39/hr × 1,000,000 ÷ (1,000 tok/s × 3600s × 70%)
| API model | Blended $/MTok | Self-host vs API | Break-even tok/s |
|---|---|---|---|
| deepseek-v4-flash deepseek | $0.18 | 16.76× | 16,757 |
| mistral-small-4 mistral | $0.26 | 11.17× | 11,172 |
| gpt-5.6-luna openai | $0.45 | 6.52× | 6,517 |
| deepseek-v4-pro deepseek | $0.54 | 5.39× | 5,393 |
| mistral-large-3 mistral | $0.75 | 3.91× | 3,910 |
| gemini-3.5-flash-lite google | $0.85 | 3.45× | 3,450 |
| claude-haiku-4-5 anthropic | $2.00 | 1.47× | 1,466 |
| gemini-3.6-flash google | $3.00 | 0.98× | 978 |
| mistral-medium-3.5 mistral | $3.00 | 0.98× | 978 |
| grok-4.5 xai | $3.00 | 0.98× | 978 |
| gemini-3.5-flash google | $3.38 | 0.87× | 869 |
| claude-sonnet-5 anthropic | $4.00 | 0.73× | 733 |
| gemini-3.1-pro-preview google | $4.50 | 0.65× | 652 |
| gpt-5.6-terra openai | $4.50 | 0.65× | 652 |
| gpt-5.4 openai | $5.63 | 0.52× | 521 |
| claude-opus-4-8 anthropic | $10.00 | 0.29× | 293 |
| claude-opus-5 anthropic | $10.00 | 0.29× | 293 |
| gpt-5.6-sol openai | $11.25 | 0.26× | 261 |
| claude-fable-5 anthropic | $20.00 | 0.15× | 147 |
| gpt-5.5-pro openai | $67.50 | 0.04× | 43 |
Below 1.00× self-hosting is cheaper per token than that API at your inputs; the break-even column is the throughput where they match. This compares raw token economics only — it ignores your engineering time, the API’s quality/latency, and that a rented GPU bills whether or not you keep it busy (that’s the utilization slider).