• GPT-6.1
  • LLM costs
  • Finance agents

GPT-6.1 Sol is OpenAI's best deal, but it disappoints in finance, ranking 30th of 72

OpenAI's newest model is the best value on Vals's general index, 91% of the top score for 15% of the cost. On Vals's finance benchmark it ranks 30th of 72, and 17 cheaper systems beat it.

David Arias, CFA5 min read

GPT-6.1 Sol: best deal, 30th in finance. All tasks: 91% of the top score for 15% of the cost, rank 7 of 30. Finance: rank 30 of 72, 17 cheaper models score higher
Two charts: on the Vals Index GPT-6.1 Sol sits at the bend of the cost curve with 91% of the top score for 15% of the cost; on Finance Agent v2 it ranks 30th of 72 and a shaded box holds the 17 systems that beat it on both accuracy and cost
Both halves of the title in one figure. Left: all tasks. Right: finance, where the shaded box holds every system that beats Sol on both accuracy and cost.

OpenAI released GPT-6.1 Sol at DevDay on 29 September, at $2 per million input tokens and $10 per million output tokens. Vals AI updated two of its leaderboards the same day, and they tell two different stories about it.

Across finance, coding, legal and tax work, Sol is the best value OpenAI sells: 61.15% on the Vals Index at $3.24 per test, 91% of the top score for 15% of its cost. On Finance Agent v2, Vals's benchmark of core financial analyst tasks, it ranks 30th of 72 at $1.62 per test. Twenty-nine systems score higher, and 17 of them cost less.

Latest video: How AI companies tricked you with Claude skills for finance

Get the next research article by email

One short email when something new is published: the test, the numbers and the code. No spam.

Prefer X? Follow @Davidariasfin

Where Sol sits on the cost curve

Vals Index accuracy against cost per test on a log scale, with the Pareto frontier bending at GPT-6.1 Sol
Vals Index v2.1, 30 models. Blue: OpenAI. Cost on a log scale so the cheap models don't pile up on the left.

Each point is one model: its accuracy on the Vals Index against what Vals paid to run it on one test. A model sits on the Pareto frontier when no other model is both cheaper and more accurate. Seven of the 30 make it, and GPT-6.1 Sol is where the curve bends. Below Sol, accuracy is cheap. Above it, every extra point costs far more:

Step along the frontierPoints gainedExtra cost per testCost per extra point
Xiaomi MiMo V2.6 Flash to MiMo V2.6 Pro+1.97+$0.21$0.11
MiMo V2.6 Pro to GPT-6.1 Sol+5.95+$2.83$0.48
GPT-6.1 Sol to GPT-6 Astra+1.98+$15.22$7.69
GPT-6 Astra to Claude Opus 5+0.54+$0.85$1.57
Claude Opus 5 to Claude Sonnet 5.5+3.37+$2.03$0.60

Going from MiMo V2.6 Pro to Sol costs $0.48 per extra point. Going from Sol to OpenAI's own GPT-6 Astra costs $7.69 per point, 16 times more, for 1.98 points. Claude Sonnet 5.5, the top model at 67.04%, sits another 3.91 points higher at $21.34. Artificial Analysis, scoring a different set of evaluations, found the same gap inside OpenAI's line: Sol one point below Astra on its Intelligence Index, at $0.72 per task against $3.26.

Screenshot of Vals's own Vals Index chart with arrows pointing at GPT-6.1 Sol and GPT-6 Astra
Vals's own chart of the Vals Index, on a linear cost axis, with arrows added.

What a test costs: list price times tokens

Three bar charts side by side: list price per million input tokens, input tokens used and cost per test for eight models, with GPT-6.1 Sol highlighted
Cost per test broken into its two drivers. Token totals cover Vals's whole index run for each model.

Cost per test is list price multiplied by the tokens a model consumes, so two models with the same price list can produce very different bills. Vals publishes both for the index run:

ModelList price, $ per M tokens in / outInput tokensOutput tokensCost per testAccuracy
Claude Opus 5.5$4 / $2028.76B491.9M$32.1466.97%
Claude Sonnet 5.5$2 / $1040.78B608.4M$21.3467.04%
GPT-6 Astra$10 / $503.93B94.5M$18.4663.13%
Kimi K3$3 / $157.55B98.7M$7.2650.70%
GLM 5.3$1.40 / $4.4012.09B211.6M$7.2553.51%
GPT-6.1 Sol$2 / $104.50B104.8M$3.2461.15%
MiMo V2.6 Pro$0.435 / $0.878.35B118.0M$0.4155.20%
DeepSeek V4.1 Flash$0.30 / $1.2010.05B138.8M$0.3351.32%

Sol and Claude Sonnet 5.5 share a list price, $2 in and $10 out. Sonnet read 9.1 times as many input tokens across the index, 40.78 billion against 4.50 billion, and costs 6.6 times as much per test. GPT-6 Astra goes the other way: it read fewer tokens than Sol, 3.93 billion, at five times the price. Sol is cheap because of both numbers.

A low price isn't enough on its own. Moonshot's Kimi K3 and Z.ai's GLM 5.3 have cheaper or similar list prices, yet cost more per test than Sol, $7.26 and $7.25, and score lower. Cheap models pair low prices with moderate token use: Xiaomi's MiMo V2.6 Pro at $0.41 per test and DeepSeek V4.1 Flash at $0.33.

On finance, Sol ranks 30th

Finance Agent v2 accuracy against cost per test on a log scale, with GPT-6.1 Sol in 30th place and Z.ai GLM 5.3 Flash near the top at five cents
Finance Agent v2, 70 systems with a published cost. Blue: OpenAI.

Finance Agent v2 gives each model six tools (SEC EDGAR search, web search, a page parser, retrieval over fetched pages, a calculator and price history) and questions written by financial experts, from disclosure analysis to DCF models and precedent transactions. Google's Gemini 3.8 Flash leads at 61.44% for $2.00 per test. Sol scores 52.03% at $1.62, and GPT-6 Astra 53.54% at $6.82. No OpenAI or Anthropic model sits on the finance frontier, which runs through NVIDIA, Ant Group, Z.ai, Meta and Google.

Horizontal bars of cost per test for the 29 systems that score above GPT-6.1 Sol on Finance Agent v2, with Sol's cost as a dashed line
Every system above GPT-6.1 Sol on Finance Agent v2. Orange bars cost less per test than Sol.

Of the 29 systems above Sol, 17 cost less per test. Some of them cost a small fraction of it:

ModelAccuracyCost per testPoints above SolCheaper than Sol
Ant Group Ling 3.0 Flash Fin54.93%$0.04+2.9040x
Z.ai GLM 5.3 Flash57.85%$0.05+5.8232x
Xiaomi MiMo V2.6 Flash56.28%$0.07+4.2523x
Xiaomi MiMo V2.6 Pro57.34%$0.20+5.318x
DeepSeek V4.1 Flash53.48%$0.21+1.458x
OpenAI GPT-5.6 Luna55.04%$0.28+3.016x
Meta Muse Spark 1.260.60%$0.77+8.572x

OpenAI's own GPT-5.6 Luna beats Sol on finance at a sixth of its cost. Ant Group's Ling 3.0 Flash Fin, a model tuned for finance, beats it at four cents a test.

Screenshot of Vals's own Finance Agent v2 chart with arrows pointing at GPT-6.1 Sol and Z.ai GLM 5.3 Flash
Vals's own chart of Finance Agent v2, on a linear cost axis, with arrows added.

Finance is far from solved, though. Vals reports that only two models clear 60%, that none reaches 51% under its stricter all-or-nothing scoring, and that category leaders in financial modelling and precedent transactions reach only 34.52% and 36.37%.

One model, two ranks

Slope chart of each model's rank on the Vals Index against its rank on Finance Agent v2, with GPT-6.1 Sol falling from 7th to 30th
Rank on the Vals Index (30 models) against rank on Finance Agent v2 (72 systems).

Sol falls from 7th on the index to 30th on finance, and Astra from 5th to 27th. Claude Sonnet 5.5 goes from 1st to 9th. Gemini 3.8 Flash climbs from 13th to 1st and Meta's Muse Spark 1.3 Max from 8th to 3rd. Xiaomi's MiMo V2.6 Pro barely moves, from 11th to 12th. A general leaderboard is a weak guide to finance performance.

What it means for finance teams

  • Test on finance work. Sol's index rank pointed to a top-ten finance result. It came 30th.
  • Route by task. On Finance Agent v2, models at $0.05 to $0.77 per test score within 3.6 points of the leader. Frontier models are worth testing where the work stays hard, such as financial modelling and precedent transactions.
  • Price the whole run. Token appetite moves cost per test as much as list price does. Cost per correct answer on a sample of real work is a better guide than a price sheet.
  • Check where the model runs. API calls send data to the provider. Models with published weights can run on a firm's own infrastructure instead.

Limits

  • One harness and one snapshot, taken on 29 September 2026. Scores carry run-to-run error: Vals puts it at about ±0.9 points for the index leaders and ±2.06 for Muse Spark 1.3 Max on finance.
  • Costs come from Vals's list-price runs. Caching, batch discounts and the choice of provider change real bills.
  • Vals's test sets are private, and ranks on the index (30 models) and on finance (72 systems) aren't on the same scale.
  • OpenAI shelved GPT-6.1 Astra before release, so Sol is the only 6.1 model on either board.
  • A benchmark task isn't a production workflow.

Sources: Vals Index v2.1 and Finance Agent v2 on vals.ai, both updated 29 September 2026 and read on 30 September 2026; Finance Agent Benchmark paper, arXiv 2508.00828; Artificial Analysis, "GPT-6.1 Sol replaces GPT-6 Sol after just 7 days" (29 September 2026); OpenAI DevDay coverage, The Next Web (29 September 2026). Charts drawn from Vals's published tables; screenshots of Vals's charts with arrows added. Nothing here is investment advice.

David Arias, CFA

Written by David Arias, CFA

CFA charterholder and licensed portfolio manager working on emerging markets private debt, derivatives and quantitative finance. Builds open-source AI tools for finance, like Glassbench, and publishes the code behind every model on this site.

See more from me