Open formula. Browser-based calculations. Source-dated rates. No API keys. No provider login. No tracing SDK. No account required. Input payloads are not stored.
Pricing source date: 2026-05-28 · Build verification: 2026-05-28 · Stale threshold: 30 days. Estimates depend on hourly GPU and hosting-rate assumptions, not public LLM API pricing. GPU market rates vary by provider, region, and commitment term; verify current instance pricing with your hosting provider before sizing a deployment.
Compares API spend against single-GPU self-hosting spend at the requested token volume and expected utilization.
Formula
gpu_monthly_with_overhead = gpu_hourly * 730 * (1 + operational_overhead_pct); breakeven_tokens = gpu_monthly_with_overhead / api_cost_per_token
Primary source register
GPU rates are planning rates from src/lib/pricing_table.js. Actual commitments, spot availability, throughput, and operational staffing can materially change the answer.
- https://openai.com/pricing
- https://www.anthropic.com/pricing
- https://platform.claude.com/docs/en/about-claude/pricing
- https://platform.claude.com/docs/en/about-claude/models/overview
- https://cloud.google.com/vertex-ai/generative-ai/pricing
- https://www.together.ai/pricing
- https://replicate.com/pricing
- https://fireworks.ai/pricing
- https://groq.com/pricing
- https://deepinfra.com/pricing
- https://aws.amazon.com/bedrock/pricing/
- https://azure.microsoft.com/pricing/details/cognitive-services/openai-service/
Included assumptions
- GPU hourly rate
- utilization assumption
- operational overhead
- API cost per 1M output tokens
Excluded assumptions
- multi-GPU networking overhead
- reserved/committed-use discounts
- engineering time outside the overhead assumption
- model-quality risk
Architecture-cost audit
Self-host break-evenGPU hourly cost, utilization, throughput, operational overhead, and API comparison.
Diagnostic output
This tool returns Cost classification, Dominant cost driver, Decision threshold, and Sensitivity. Diagnostic focus: API-vs-GPU break-even volume and utilization sensitivity..
Affiliate link. Model Ruler may earn a commission if you sign up. This does not affect what you pay.