What an LLM workload costs, per provider

Open formula. Browser-based calculations. Source-dated rates. No API keys. No provider login. No tracing SDK. No account required. Input payloads are not stored.
Pricing source date: 2026-05-28 · Build verification: 2026-05-28 · Stale threshold: 30 days. Estimates use published per-token API pricing for the model/provider path shown by the calculator, reviewed on the source date shown. Provider rates change frequently; verify current pricing on the provider's own page before committing to traffic volume, eval cadence, or agent-loop workload shape.

Worked example

Example inputs: 10,000 input + 2,000 output tokens per call, 100,000 calls/month, on anthropic/claude-sonnet-4-6

Pricing last updated 2026-05-28 — verify with provider before committing to commitments or contracts. $6000.00/mo at 100,000 calls on anthropic/claude-sonnet-4-6 (10,000 in / 2,000 out per call). deepinfra/llama-3.3-70b offers comparable tier at $310.00/mo — potential 95% savings.

OptionResult
deepinfra/llama-3.3-70b$310/mo
together/llama-3.3-70b-instruct-turbo$1,056/mo
fireworks/llama-3.3-70b$1,080/mo
google-vertex/gemini-3.5-flash$3,300/mo
openai/gpt-5.4$5,500/mo
anthropic/claude-sonnet-4-6$6,000/mo

Example inputs — your own numbers will differ. This is a computed planning estimate, not a quote.

Diagnostic output: Each calculation returns Cost classification, Dominant cost driver, Decision threshold, and Sensitivity. Focus for this tool: Dominant spend driver and routing threshold between budget/mid/premium models.

About this calculator

Enter your per-call token shape and monthly volume, and Model Ruler computes spend on your chosen provider, then compares budget, mid, and premium alternatives from a source-dated pricing table. Same inputs always produce the same output — it is a planning estimate, not a quote.

Methodology summary

Converts per-call token shape and monthly volume into provider/model spend, then compares alternatives from the committed pricing table.

Formula: monthly_cost = ((tokens_in * calls_per_month) / 1,000,000 * input_rate) + ((tokens_out * calls_per_month) / 1,000,000 * output_rate)

Audit coverage: Context growth

Open the full methodology for this calculator.

Frequently asked

What does this calculator compute?
It converts your per-call input and output tokens and monthly call volume into monthly spend on a chosen provider and model, then compares that against alternative models from a source-dated pricing table.

What should I check before committing to a workload shape?
Input and output token volume per request, request volume, and model choice dominate. Compare two or three providers on the specific request shape you actually run — a model that's cheaper per input token can be more expensive overall once output length and request frequency are included.

Is this a quote or a benchmark?
Neither — it is a computed estimate, not a vendor quote. The worked example above shows the numbers this calculator produces for a representative provider cost workload; your own inputs move the result.

Related tools

Estimates use published per-token API pricing for the model/provider path shown by the calculator, reviewed on the source date shown. Provider rates change frequently; verify current pricing on the provider's own page before committing to traffic volume, eval cadence, or agent-loop workload shape. View methodology.