Token price is not
task cost.
Every pricing page compares dollars per million tokens. That number ignores how many tokens a model burns to finish a job, and how often it fails and has to start again. Solvency measures cost per solved task.
Per output token
19x
cheaper
Per solved task
100x
cheaper
Understatement
5x
factor
Spread widens
38→146x
token price → cost
DeepSeek V4 Flash against Claude Opus 5, measured on agentic coding tasks.
Measured
The benchmark ran the model and observed this cost. No Solvency assumption is inside these figures.
| Model | Pass | $ / solved | $ / month |
|---|---|---|---|
| DeepSeek V4 Flash measured | 50% | $0.12 | $24.00 |
| Gemini 3.7 Flash measured | 60% | $2.12 | $423 |
| Grok 4.5 measured | 64% | $3.81 | $763 |
| GPT-5.6 Sol measured | 65% | $9.88 | $2.0k |
| Claude Opus 5 measured | 68% | $12.01 | $2.4k |
| Claude Fable 5 measured | 67% | $17.46 | $3.5k |
Modelled
Pass rate published, cost estimated by Solvency's loop model — an assumption.
| Model | Pass | $ / solved | $ / month |
|---|---|---|---|
| GPT-5.4 | 59% | $1.86 | $372 |
| Gemini 3.1 Pro (preview) | 46% | $1.91 | $382 |
| Claude Opus 4.6 | 52% | $3.85 | $771 |
| Claude Opus 4.5 | 46% | $4.36 | $872 |
Modelled from stale pass rates
Pass rates published before 2026. Cost recomputed at current prices; the pass rate is old.
| Model | Pass | $ / solved | $ / month |
|---|---|---|---|
| GPT-5 | 88% | $0.74 | $148 |
| Gemini 2.5 Pro | 83% | $0.78 | $156 |
| o3 | 81% | $0.89 | $177 |
| GPT-4.1 | 52% | $1.37 | $275 |
| Claude Sonnet 4 | 61% | $1.96 | $392 |
| Claude Opus 4 | 72% | $8.33 | $1.7k |
| o3-pro | 85% | $8.48 | $1.7k |
Measured rows carry a cost the benchmark observed, so the tier, cache and efficiency controls cannot move them. Modelled rows are priced by an assumed loop model.
Research Note 01 — latest
Cost Per Solved Task
Per-token pricing does not predict what a coding task actually costs. Measured across six current models, the gap widens from 38x to 146x.
Read the note →