GPT-6.1 Sol is OpenAI's best deal, but it disappoints in finance, ranking 30th of 72
OpenAI's newest model is the best value on Vals's general index, 91% of the top score for 15% of the cost. On Vals's finance benchmark it ranks 30th of 72, and 17 cheaper systems beat it.


OpenAI released GPT-6.1 Sol at DevDay on 29 September, at $2 per million input tokens and $10 per million output tokens. Vals AI updated two of its leaderboards the same day, and they tell two different stories about it.
Across finance, coding, legal and tax work, Sol is the best value OpenAI sells: 61.15% on the Vals Index at $3.24 per test, 91% of the top score for 15% of its cost. On Finance Agent v2, Vals's benchmark of core financial analyst tasks, it ranks 30th of 72 at $1.62 per test. Twenty-nine systems score higher, and 17 of them cost less.
Latest video: How AI companies tricked you with Claude skills for finance
Where Sol sits on the cost curve

Each point is one model: its accuracy on the Vals Index against what Vals paid to run it on one test. A model sits on the Pareto frontier when no other model is both cheaper and more accurate. Seven of the 30 make it, and GPT-6.1 Sol is where the curve bends. Below Sol, accuracy is cheap. Above it, every extra point costs far more:
| Step along the frontier | Points gained | Extra cost per test | Cost per extra point |
|---|---|---|---|
| Xiaomi MiMo V2.6 Flash to MiMo V2.6 Pro | +1.97 | +$0.21 | $0.11 |
| MiMo V2.6 Pro to GPT-6.1 Sol | +5.95 | +$2.83 | $0.48 |
| GPT-6.1 Sol to GPT-6 Astra | +1.98 | +$15.22 | $7.69 |
| GPT-6 Astra to Claude Opus 5 | +0.54 | +$0.85 | $1.57 |
| Claude Opus 5 to Claude Sonnet 5.5 | +3.37 | +$2.03 | $0.60 |
Going from MiMo V2.6 Pro to Sol costs $0.48 per extra point. Going from Sol to OpenAI's own GPT-6 Astra costs $7.69 per point, 16 times more, for 1.98 points. Claude Sonnet 5.5, the top model at 67.04%, sits another 3.91 points higher at $21.34. Artificial Analysis, scoring a different set of evaluations, found the same gap inside OpenAI's line: Sol one point below Astra on its Intelligence Index, at $0.72 per task against $3.26.

What a test costs: list price times tokens

Cost per test is list price multiplied by the tokens a model consumes, so two models with the same price list can produce very different bills. Vals publishes both for the index run:
| Model | List price, $ per M tokens in / out | Input tokens | Output tokens | Cost per test | Accuracy |
|---|---|---|---|---|---|
| Claude Opus 5.5 | $4 / $20 | 28.76B | 491.9M | $32.14 | 66.97% |
| Claude Sonnet 5.5 | $2 / $10 | 40.78B | 608.4M | $21.34 | 67.04% |
| GPT-6 Astra | $10 / $50 | 3.93B | 94.5M | $18.46 | 63.13% |
| Kimi K3 | $3 / $15 | 7.55B | 98.7M | $7.26 | 50.70% |
| GLM 5.3 | $1.40 / $4.40 | 12.09B | 211.6M | $7.25 | 53.51% |
| GPT-6.1 Sol | $2 / $10 | 4.50B | 104.8M | $3.24 | 61.15% |
| MiMo V2.6 Pro | $0.435 / $0.87 | 8.35B | 118.0M | $0.41 | 55.20% |
| DeepSeek V4.1 Flash | $0.30 / $1.20 | 10.05B | 138.8M | $0.33 | 51.32% |
Sol and Claude Sonnet 5.5 share a list price, $2 in and $10 out. Sonnet read 9.1 times as many input tokens across the index, 40.78 billion against 4.50 billion, and costs 6.6 times as much per test. GPT-6 Astra goes the other way: it read fewer tokens than Sol, 3.93 billion, at five times the price. Sol is cheap because of both numbers.
A low price isn't enough on its own. Moonshot's Kimi K3 and Z.ai's GLM 5.3 have cheaper or similar list prices, yet cost more per test than Sol, $7.26 and $7.25, and score lower. Cheap models pair low prices with moderate token use: Xiaomi's MiMo V2.6 Pro at $0.41 per test and DeepSeek V4.1 Flash at $0.33.
On finance, Sol ranks 30th

Finance Agent v2 gives each model six tools (SEC EDGAR search, web search, a page parser, retrieval over fetched pages, a calculator and price history) and questions written by financial experts, from disclosure analysis to DCF models and precedent transactions. Google's Gemini 3.8 Flash leads at 61.44% for $2.00 per test. Sol scores 52.03% at $1.62, and GPT-6 Astra 53.54% at $6.82. No OpenAI or Anthropic model sits on the finance frontier, which runs through NVIDIA, Ant Group, Z.ai, Meta and Google.

Of the 29 systems above Sol, 17 cost less per test. Some of them cost a small fraction of it:
| Model | Accuracy | Cost per test | Points above Sol | Cheaper than Sol |
|---|---|---|---|---|
| Ant Group Ling 3.0 Flash Fin | 54.93% | $0.04 | +2.90 | 40x |
| Z.ai GLM 5.3 Flash | 57.85% | $0.05 | +5.82 | 32x |
| Xiaomi MiMo V2.6 Flash | 56.28% | $0.07 | +4.25 | 23x |
| Xiaomi MiMo V2.6 Pro | 57.34% | $0.20 | +5.31 | 8x |
| DeepSeek V4.1 Flash | 53.48% | $0.21 | +1.45 | 8x |
| OpenAI GPT-5.6 Luna | 55.04% | $0.28 | +3.01 | 6x |
| Meta Muse Spark 1.2 | 60.60% | $0.77 | +8.57 | 2x |
OpenAI's own GPT-5.6 Luna beats Sol on finance at a sixth of its cost. Ant Group's Ling 3.0 Flash Fin, a model tuned for finance, beats it at four cents a test.

Finance is far from solved, though. Vals reports that only two models clear 60%, that none reaches 51% under its stricter all-or-nothing scoring, and that category leaders in financial modelling and precedent transactions reach only 34.52% and 36.37%.
One model, two ranks

Sol falls from 7th on the index to 30th on finance, and Astra from 5th to 27th. Claude Sonnet 5.5 goes from 1st to 9th. Gemini 3.8 Flash climbs from 13th to 1st and Meta's Muse Spark 1.3 Max from 8th to 3rd. Xiaomi's MiMo V2.6 Pro barely moves, from 11th to 12th. A general leaderboard is a weak guide to finance performance.
What it means for finance teams
- Test on finance work. Sol's index rank pointed to a top-ten finance result. It came 30th.
- Route by task. On Finance Agent v2, models at $0.05 to $0.77 per test score within 3.6 points of the leader. Frontier models are worth testing where the work stays hard, such as financial modelling and precedent transactions.
- Price the whole run. Token appetite moves cost per test as much as list price does. Cost per correct answer on a sample of real work is a better guide than a price sheet.
- Check where the model runs. API calls send data to the provider. Models with published weights can run on a firm's own infrastructure instead.
Limits
- One harness and one snapshot, taken on 29 September 2026. Scores carry run-to-run error: Vals puts it at about ±0.9 points for the index leaders and ±2.06 for Muse Spark 1.3 Max on finance.
- Costs come from Vals's list-price runs. Caching, batch discounts and the choice of provider change real bills.
- Vals's test sets are private, and ranks on the index (30 models) and on finance (72 systems) aren't on the same scale.
- OpenAI shelved GPT-6.1 Astra before release, so Sol is the only 6.1 model on either board.
- A benchmark task isn't a production workflow.
Sources: Vals Index v2.1 and Finance Agent v2 on vals.ai, both updated 29 September 2026 and read on 30 September 2026; Finance Agent Benchmark paper, arXiv 2508.00828; Artificial Analysis, "GPT-6.1 Sol replaces GPT-6 Sol after just 7 days" (29 September 2026); OpenAI DevDay coverage, The Next Web (29 September 2026). Charts drawn from Vals's published tables; screenshots of Vals's charts with arrows added. Nothing here is investment advice.


