Self-hosting on a GPU vs paying per API call

Open formula. Browser-based calculations. Source-dated rates. No API keys. No provider login. No tracing SDK. No account required. Input payloads are not stored.
Pricing source date: 2026-05-28 · Build verification: 2026-05-28 · Stale threshold: 30 days. Estimates depend on hourly GPU and hosting-rate assumptions, not public LLM API pricing. GPU market rates vary by provider, region, and commitment term; verify current instance pricing with your hosting provider before sizing a deployment.

Worked example

Example inputs: 500M output tokens/month, API at $15 per 1M output tokens, self-hosted on a RunPod H100

Pricing last updated 2026-05-28 — verify with provider before committing to commitments or contracts. At 500,000,000 output tokens/mo, API cost ~$7500/mo vs self-host ~$2851/mo (incl. 40% overhead). Self-hosting is favorable — approximate savings $4649/mo.

Example inputs — your own numbers will differ. This is a computed planning estimate, not a quote.

Diagnostic output: Each calculation returns Cost classification, Dominant cost driver, Decision threshold, and Sensitivity. Focus for this tool: API-vs-GPU break-even volume and utilization sensitivity.

About this calculator

Enter your monthly output-token volume, current API rate, GPU type, and expected utilization; Model Ruler compares API spend against GPU self-hosting (with operational overhead) and reports the breakeven volume. Utilization and overhead dominate the answer — low utilization usually keeps the API cheaper.

Methodology summary

Compares API spend against single-GPU self-hosting spend at the requested token volume and expected utilization.

Formula: gpu_monthly_with_overhead = gpu_hourly * 730 * (1 + operational_overhead_pct); breakeven_tokens = gpu_monthly_with_overhead / api_cost_per_token

Audit coverage: Self-host break-even

Open the full methodology for this calculator.

Frequently asked

What does this calculator compute?
It compares monthly API spend at your token volume against the monthly cost of self-hosting on a GPU, including operational overhead, and reports the breakeven volume.

What should I check before committing to a workload shape?
Breakeven depends on sustained utilization, GPU hourly rate, and your request volume. Idle GPU time is the hidden cost — check whether your traffic keeps the hardware busy enough to beat per-token API pricing, since low utilization usually favors staying on an API.

Is this a quote or a benchmark?
Neither — it is a computed estimate, not a vendor quote. The worked example above shows the numbers this calculator produces for a representative self host breakeven workload; your own inputs move the result.

Related tools

Estimates depend on hourly GPU and hosting-rate assumptions, not public LLM API pricing. GPU market rates vary by provider, region, and commitment term; verify current instance pricing with your hosting provider before sizing a deployment. View methodology.