SOLVENCY

Token price is not
task cost.

Every pricing page compares dollars per million tokens. That number ignores how many tokens a model burns to finish a job, and how often it fails and has to start again. Solvency measures cost per solved task.

Per output token

19x

cheaper

Per solved task

100x

cheaper

Understatement

5x

factor

Spread widens

38→146x

token price → cost

DeepSeek V4 Flash against Claude Opus 5, measured on agentic coding tasks.

DeepSeek V4 Flash costs $0.12 per solved task against Claude Opus 5 at $12.01 — 100.1x more for 18 more points of pass rate. Over 200 tasks that difference is $2.4k a month.

Measured

The benchmark ran the model and observed this cost. No Solvency assumption is inside these figures.

ModelPass $ / solved$ / month
DeepSeek V4 Flash measured 50% $0.12 $24.00
Gemini 3.7 Flash measured 60% $2.12 $423
Grok 4.5 measured 64% $3.81 $763
GPT-5.6 Sol measured 65% $9.88 $2.0k
Claude Opus 5 measured 68% $12.01 $2.4k
Claude Fable 5 measured 67% $17.46 $3.5k

Modelled

Pass rate published, cost estimated by Solvency's loop model — an assumption.

ModelPass $ / solved$ / month
GPT-5.4 59% $1.86 $372
Gemini 3.1 Pro (preview) 46% $1.91 $382
Claude Opus 4.6 52% $3.85 $771
Claude Opus 4.5 46% $4.36 $872

Modelled from stale pass rates

Pass rates published before 2026. Cost recomputed at current prices; the pass rate is old.

ModelPass $ / solved$ / month
GPT-5 88% $0.74 $148
Gemini 2.5 Pro 83% $0.78 $156
o3 81% $0.89 $177
GPT-4.1 52% $1.37 $275
Claude Sonnet 4 61% $1.96 $392
Claude Opus 4 72% $8.33 $1.7k
o3-pro 85% $8.48 $1.7k
Not shown — no published pass rate, reported as missing rather than estimated: Claude Sonnet 5, Claude Haiku 4.5, GPT-5.6 Terra, GPT-5.6 Luna, GPT-5.3 Codex, DeepSeek V4 Pro, Grok 4.6, Mistral Medium 3.5.

Measured rows carry a cost the benchmark observed, so the tier, cache and efficiency controls cannot move them. Modelled rows are priced by an assumed loop model.

Research Note 01 — latest

Cost Per Solved Task

Per-token pricing does not predict what a coding task actually costs. Measured across six current models, the gap widens from 38x to 146x.

Read the note →