A 12-agent AI investment committee for $0.09. Here's what it decided.
10 AI agents build the case, 2 decide, and neither judge ever reads the data. Inside GitHub's most starred trading project (108,000 stars), tested over 68 runs with every log published.

On 19 September 2026, at 22:07, twelve AI agents met as an investment committee and rated Microsoft Underweight. Two minutes later the same committee met again, with the same filings, prices, models and settings, and rated it Overweight. Each meeting lasted about eight minutes and cost about $0.09.
This committee is TradingAgents, an open-source framework running here on DeepSeek's models. To see how it can disagree with itself, it helps to know where the idea came from and how the meeting is organised.
Three years, from one model to a whole firm

Finance started with models. BloombergGPT, published in March 2023, was a 50 billion parameter model trained on a corpus heavy with financial data. FinGPT followed in June with open-source financial models anyone could fine-tune. You could ask either one about a filing or a headline. Neither could do anything with the answer.
Agents changed that, though at first they worked alone. FinMem, in November 2023, was a single trading agent with a layered memory and a profile you could configure. FinAgent, in February 2024, read price charts, news and tool results together before it traded. Nobody inside either system checked the agent's reasoning.
A fix came from general AI research. In May 2023, Yilun Du and colleagues at MIT published multiagent debate: several copies of a model propose answers and criticise each other for a few rounds, and together they reason better and hallucinate less than one copy alone.
Finance took that idea and gave the agents job titles. FinRobot, in May 2024, put specialised financial agents on one open-source platform. FinCon, in July, organised its agents as a manager with analysts, a setup its authors describe as "inspired by effective real-world investment firm organizational structures," plus a risk step that criticises past decisions. AI Hedge Fund, a GitHub repository started that November, gives each agent the style of a famous investor and has more than 63,000 stars.
Committees have one more appeal for finance. Every report, argument and ruling is written down, so a person can audit exactly how a decision was reached. In December 2024, one project took the firm idea further than anyone else.
TradingAgents

By GitHub stars, TradingAgents is the biggest project on that timeline: more than 108,000 stars and 20,800 forks, 22 contributors and eleven releases, published by Tauric Research under the Apache 2.0 licence. Search GitHub for trading projects and it comes first, with about twice the stars of freqtrade, the runner-up.
It began as the code for a paper by Yijia Xiao, Edward Sun, Di Luo and Wei Wang. Their abstract notes that earlier work had "largely focused on single-agent systems handling specific tasks," and proposes a whole trading firm instead: analysts, researchers who argue, a trader, a risk team and a manager who signs off. On Apple, over the first quarter of 2024, the paper reported a 26.62% cumulative return and a Sharpe ratio of 8.21, while simply holding Apple lost 5.23%.
Those results came from three stocks over three months, one backtest each, and later tests were less kind. FINSABER, a KDD 2026 paper, tested LLM trading strategies over twenty years and more than 100 symbols and found the earlier advantages "deteriorate significantly." Its strategies were too cautious in bull markets and too aggressive in bear markets. Agent Market Arena, launched in October 2025, ran four agent designs live on stocks and crypto, each on five models from OpenAI, Anthropic and Google. How a system was designed changed its behaviour far more than which model it ran on. In the authors' words, "model backbones contribute less to outcome variation."
The code kept moving after the paper. Releases in 2026 widened model support to most major providers, from OpenAI, Anthropic and Google to DeepSeek and Qwen, then added structured outputs for the managers and the trader, a grounded sentiment analyst, and FRED and Polymarket data. Version 0.5.0, released on 18 September 2026, brought point-in-time data on every dated path, SEC filings as filed and grid backtests, and 0.5.1 followed six days later. Its README now calls the project "a research scaffold for studying multi-agent analysis," and forks adapt it to Chinese A-shares and to brokers like Alpaca.
Setup, a full run and a paper order at a broker are shown step by step in this video:

How the committee works

A run takes one stock and one date. Five stages follow in a fixed order, built with LangGraph, a library for wiring LLM calls into a graph.
First, four analysts work one after another, each with its own data tools: a market analyst for prices and indicators, a sentiment analyst for headlines and social posts, a news analyst for headlines and macro data, and a fundamentals analyst for statements and filings. Each writes a report, and only these four ever touch data. Version 0.5.0 also made that data safer: the company is identified from the ticker before anyone starts, the market analyst checks exact prices against a verified snapshot, and on past dates no source goes beyond what was known that day.
Second comes a debate with fixed sides. A bull researcher argues for buying and a bear argues against, whatever the reports say, for one round by default. Sides are assigned so that every argument meets the strongest opposing case.
Third, a research manager reads the debate and writes an investment plan with one of five ratings: Buy, Overweight, Hold, Underweight or Sell. Overweight means holding more than a normal position, Underweight less.
Fourth, a trader turns that plan into an order idea with an entry price, a stop and a size.
Fifth, three risk voices, one aggressive, one neutral and one conservative, argue about the trader's proposal. A portfolio manager then makes the final call: a rating, a price target and a time horizon.
Once a decision's result is known, the framework writes a short lesson into a memory log. Later runs load those lessons, limited to what was known on their own date, and in version 0.5.0 only the portfolio manager reads them.
At the end you get a rating for one stock on one date, with sizes described relative to "a standard allocation." Turning those ratings into a portfolio is up to you.
Two design choices matter most. Only two agents decide, and nobody votes: the research manager and the portfolio manager run on DeepSeek's larger reasoning model, Pro, and the other ten run on the faster Flash model. And neither judge reads the data.

In these runs, the four analysts read about 109,000 tokens of raw material between them. Each judge read fewer than 8,000. Of the two, the research manager sees only the bull and bear transcript, and the portfolio manager sees the plans and the risk debate. Neither sees a single analyst report. That's intended, since the debate is supposed to boil the evidence down for the judges. It also means a ruling can depend on how an argument was worded.
Here's what it decided
Across eight full runs with all twelve agents, the median cost was $0.09 at DeepSeek's list prices on 24 September 2026, and a run took about eight minutes. More than half of everything the models wrote was hidden reasoning, produced before the visible answer.

Now Microsoft. Both runs started from the same filings and the same closing price of $493.78, and both fundamentals analysts spotted the same $5.9 billion of gains on securities Microsoft had sold. They framed it differently. One called a $15.6 billion swing in other income "not purely operational." Its counterpart in the second run called normalized income growth of 26.4% "the cleaner read, still exceptional."

Everything downstream followed that framing. In the first run, the bear pointed to deferred revenue growing 13% while revenue grew 17.8%. In the second, the bull read that same deferred revenue as a sign of demand. One research manager concluded that "the bear won the risk/reward argument," and the other that "the bull's structural case is the stronger one, but not decisive." One trader sold and the other bought.
Position sizes narrow the gap. Underweight meant cutting to half or three quarters of a normal position, "not an exit." Overweight meant buying a small first slice and staying at or below a normal position until the move confirmed. So the two labels point in opposite directions, while the money involved barely differs.
TradingAgents' own README explains why this happens: model sampling is non-deterministic, and reasoning models "vary the most" and "largely ignore temperature," the setting normally used to make output repeatable. Two runs are a single example, so they can't show how often ratings flip. Until someone publishes reruns of the same stock on the same date, no single rating should size a position.
Watch it run

Every run in this article was recorded in Glassbench, an open-source workbench built around TradingAgents. It shows the twelve agents working live, stores every run in a searchable database, compares LLMs on the same setup, and sends paper orders only, never real money. All 68 runs behind this article come with the repository, including reports, tool results and reasoning, so every number here can be checked without spending anything.
One meeting costs $0.09. Trusting its verdict takes many meetings: the same stock on the same date, run again and again until you know how far the answers spread. At today's prices, a hundred of them cost about $9.
Figures come from the Glassbench runs database (68 runs, read 24 September 2026), TradingAgents 0.5.0 source and the TradingAgents repository page on 25 September 2026. Costs use DeepSeek's off-peak list prices on 24 September. Written by a TradingAgents contributor (merged fix #1370) who builds Glassbench, with no affiliation to Tauric Research. Site: davidariasfinance.com/glassbench. Nothing here is investment advice.
Watch my latest video
Open-source AI trading agents on DeepSeek
Twelve AI agents take a ticker from research to a paper order at a broker, set up step by step.


