Example inputs: 10,000 input + 2,000 output tokens per call, 100,000 calls/month, on anthropic/claude-sonnet-4-6
Pricing last updated 2026-05-28 — verify with provider before committing to commitments or contracts. $6000.00/mo at 100,000 calls on anthropic/claude-sonnet-4-6 (10,000 in / 2,000 out per call). deepinfra/llama-3.3-70b offers comparable tier at $310.00/mo — potential 95% savings.
| Option | Result |
|---|---|
| deepinfra/llama-3.3-70b | $310/mo |
| together/llama-3.3-70b-instruct-turbo | $1,056/mo |
| fireworks/llama-3.3-70b | $1,080/mo |
| google-vertex/gemini-3.5-flash | $3,300/mo |
| openai/gpt-5.4 | $5,500/mo |
| anthropic/claude-sonnet-4-6 | $6,000/mo |
Example inputs — your own numbers will differ. This is a computed planning estimate, not a quote.
Enter your per-call token shape and monthly volume, and Model Ruler computes spend on your chosen provider, then compares budget, mid, and premium alternatives from a source-dated pricing table. Same inputs always produce the same output — it is a planning estimate, not a quote.
Converts per-call token shape and monthly volume into provider/model spend, then compares alternatives from the committed pricing table.
Formula: monthly_cost = ((tokens_in * calls_per_month) / 1,000,000 * input_rate) + ((tokens_out * calls_per_month) / 1,000,000 * output_rate)
Audit coverage: Context growth
What does this calculator compute?
It converts your per-call input and output tokens and monthly call volume into monthly spend on a chosen provider and model, then compares that against alternative models from a source-dated pricing table.
What should I check before committing to a workload shape?
Input and output token volume per request, request volume, and model choice dominate. Compare two or three providers on the specific request shape you actually run — a model that's cheaper per input token can be more expensive overall once output length and request frequency are included.
Is this a quote or a benchmark?
Neither — it is a computed estimate, not a vendor quote. The worked example above shows the numbers this calculator produces for a representative provider cost workload; your own inputs move the result.
Estimates use published per-token API pricing for the model/provider path shown by the calculator, reviewed on the source date shown. Provider rates change frequently; verify current pricing on the provider's own page before committing to traffic volume, eval cadence, or agent-loop workload shape. View methodology.