Run TradingAgents, the 108,000 star multi agent LLM trading framework, on a real stock. Watch the twelve agents live, keep every run in a database, benchmark LLMs on the same harness, backtest honestly, and send the decision to an Interactive Brokers paper account.
AI trading agents are everywhere, and the quant community argues about them. The research got serious with FinMem, FinAgent and FinCon, and the most famous framework, TradingAgents, has a paper reporting 26% on Apple in a quarter where buy and hold lost 5%. Nobody outside the authors had a way to watch it work.
Glassbench is a local workbench on top of the official engine, which runs untouched. It records what the twelve agents read, argued and decided, one card per agent, live and replayable. Every run lands in a SQLite database with its framework version, models, rating, entry, stop, target, cost and flags. A simulator replays the saved decisions against buy and hold and a placebo. And one script turns a finished run into a bracket order at a paper broker account. The point is to judge the method yourself instead of trusting a rating.
The video explains the method, installs the framework in the terminal, runs it on DeepSeek, then follows every agent in Glassbench and sends the NVIDIA decision to Interactive Brokers.
I post the models as I build them, the changes I make to them, and the results that did not work. Shorter than these write ups, and there are a lot more of them.
Follow me on XThe app, the tests, two evidence documents and a database of 68 recorded runs are in the repository, under the Apache 2.0 license. The engine is cloned separately and never modified.
A run is twelve agents in five stages: four data analysts (market, sentiment, news, fundamentals) with tools, a bull and a bear researcher, a research manager, a trader, three risk debaters and a portfolio manager. Cheap fast models do the reading and the debating, and the expensive model is called only by the two managers. Glassbench shows one card per agent, with what it is reading, what it is writing and its one line conclusion. Click a card for the full output: every LLM call, every tool call with its result, the reasoning stream.

Below the board, a stage strip shows where the run is and a timeline shows who ran when and for how long. A four analyst run takes about eight minutes and costs about six cents on DeepSeek. Every run can be replayed step by step afterwards.
The terminal version of the framework works for one run. After ten you cannot compare anything, filter, see what each run cost, or send a decision anywhere. In Glassbench each run lands in SQLite with its framework and version, provider and models, ticker, date, analysts, debate depth, rating, entry, stop, target, horizon, tokens, tool calls, cost and flags.

The Database page facets the same runs by framework version, variant, quick and deep model, the model actually served, ticker, rating, horizon and purpose, and searches the full text of every report and every reasoning trace. Any view exports to CSV. The repository ships the author's database as a courtesy: 68 runs on 8 tickers, including two Microsoft runs from the video that disagree on the same day.

The framework is fixed; the models are yours to choose. Pick the provider and type any model id for the quick role (reads and debates) and the deep role (the two managers decide). DeepSeek is the default; OpenAI, Anthropic, Google, xAI, Qwen, GLM, MiniMax, Mistral, Kimi, Groq, OpenRouter and a local Ollama are one key away. Every run records which model was asked for and which was actually served, so two models on the same stock and date sit side by side in the same table with their cost.

One to four analysts, one to five debate rounds, and named variants that change how the agents are fed (prompt wording, SEC EDGAR statements as filed, point in time valuation) without touching the engine. On 24 matched runs, serving valuation as of the run date cut the agents' "valuation is missing" statements from 55 to 1 and raised correctly cited multiples from 25 to 397, at the same cost. The write up is in the repository.
Weekly grids of runs with a budget cap, replayed at zero LLM cost against buy and hold, a 50/200 day moving average rule and a shuffled rating placebo, with next open fills and trading costs. The pilot on AAPL and NVDA had the agents' trader levels lose money while buy and hold gained, and the ratings alone never traded. Four decisions per stock is a method check, not evidence, and the page says so.

Entry, stop and target are what a broker needs for an order. One script lets the committee decide, then sends the decision to an Interactive Brokers paper account as one bracket order per stock, with the trader's stop and the portfolio manager's target attached. Long only, one decision, fixed size, paper ports only, and a y/N confirmation before anything is sent. The order stays linked to the run that produced it, so every trade traces back to the argument behind it.
trade.py
py -3 trade.py NVDA MSFT # run the twelve agents on each stock, then trade
py -3 trade.py NVDA MSFT --reuse # replay each stock's latest finished run, then trade
py -3 trade.py NVDA --dry-run # everything except sending the ordersPython 3.11 or newer, Node 22 or newer, Git, and an API key for one LLM provider. The engine is cloned at its released tag, next to the app.
Terminal
git clone https://github.com/davidalmeida90/glassbench.git
cd glassbench
git clone --branch v0.5.0 --depth 1 https://github.com/TauricResearch/TradingAgents.git
py -3.11 -m venv .venv # Mac or Linux: python3 -m venv .venv
.\.venv\Scripts\Activate.ps1 # source .venv/bin/activate
pip install -r requirements.txt
pip install -e .\TradingAgents
copy .env.example .env # paste your key after the equals sign
.\glassbench.ps1 # Mac or Linux: ./glassbench.shThe first start builds the frontend, then the app is at http://127.0.0.1:8765. Press New run, pick a ticker, a date, the analysts, the provider and the two models, and watch the committee work.