Glassbench, an open-source UI for AI trading agents

Run TradingAgents, the 108,000 star multi agent LLM trading framework, on a real stock. Watch the twelve agents live, keep every run in a database, benchmark LLMs on the same harness, backtest honestly, and send the decision to an Interactive Brokers paper account.

What it is

AI trading agents are everywhere, and the quant community argues about them. The research got serious with FinMem, FinAgent and FinCon, and the most famous framework, TradingAgents, has a paper reporting 26% on Apple in a quarter where buy and hold lost 5%. Nobody outside the authors had a way to watch it work.

Glassbench is a local workbench on top of the official engine, which runs untouched. It records what the twelve agents read, argued and decided, one card per agent, live and replayable. Every run lands in a SQLite database with its framework version, models, rating, entry, stop, target, cost and flags. A simulator replays the saved decisions against buy and hold and a placebo. And one script turns a finished run into a bracket order at a paper broker account. The point is to judge the method yourself instead of trusting a rating.

Watch the video

The video explains the method, installs the framework in the terminal, runs it on DeepSeek, then follows every agent in Glassbench and sends the NVIDIA decision to Interactive Brokers.

Open source AI trading agents on DeepSeek, from the setup to a broker order. Subtitles in English, Portuguese and Spanish.

On X

Follow me on X, my models are there

I post the models as I build them, the changes I make to them, and the results that did not work. Shorter than these write ups, and there are a lot more of them.

Follow me on X
Newsletter

Get exclusive models, prompts and code

One short email when there's something new. finance models, code walkthroughs, agentic prompts. No spam.

Built with
Anthropic Claude Python R ChatGPT GitHub Jupyter Excel PowerPoint VS Code SQL MATLAB

The app, the tests, two evidence documents and a database of 68 recorded runs are in the repository, under the Apache 2.0 license. The engine is cloned separately and never modified.

Open Glassbench on GitHub

1. The committee, live

A run is twelve agents in five stages: four data analysts (market, sentiment, news, fundamentals) with tools, a bull and a bear researcher, a research manager, a trader, three risk debaters and a portfolio manager. Cheap fast models do the reading and the debating, and the expensive model is called only by the two managers. Glassbench shows one card per agent, with what it is reading, what it is writing and its one line conclusion. Click a card for the full output: every LLM call, every tool call with its result, the reasoning stream.

The committee board of a finished NVIDIA run in Glassbench

Below the board, a stage strip shows where the run is and a timeline shows who ran when and for how long. A four analyst run takes about eight minutes and costs about six cents on DeepSeek. Every run can be replayed step by step afterwards.

2. A database of every run

The terminal version of the framework works for one run. After ten you cannot compare anything, filter, see what each run cost, or send a decision anywhere. In Glassbench each run lands in SQLite with its framework and version, provider and models, ticker, date, analysts, debate depth, rating, entry, stop, target, horizon, tokens, tool calls, cost and flags.

The Runs page: every recorded analysis with its rating, levels and cost

The Database page facets the same runs by framework version, variant, quick and deep model, the model actually served, ticker, rating, horizon and purpose, and searches the full text of every report and every reasoning trace. Any view exports to CSV. The repository ships the author's database as a courtesy: 68 runs on 8 tickers, including two Microsoft runs from the video that disagree on the same day.

The Database page: facets and a full text search over every run

3. Benchmark LLMs on the same harness, choose the agents

The framework is fixed; the models are yours to choose. Pick the provider and type any model id for the quick role (reads and debates) and the deep role (the two managers decide). DeepSeek is the default; OpenAI, Anthropic, Google, xAI, Qwen, GLM, MiniMax, Mistral, Kimi, Groq, OpenRouter and a local Ollama are one key away. Every run records which model was asked for and which was actually served, so two models on the same stock and date sit side by side in the same table with their cost.

The New run dialog: provider, model ids, analysts and debate rounds

One to four analysts, one to five debate rounds, and named variants that change how the agents are fed (prompt wording, SEC EDGAR statements as filed, point in time valuation) without touching the engine. On 24 matched runs, serving valuation as of the run date cut the agents' "valuation is missing" statements from 55 to 1 and raised correctly cited multiples from 25 to 397, at the same cost. The write up is in the repository.

4. Backtests with an honest simulator

Weekly grids of runs with a budget cap, replayed at zero LLM cost against buy and hold, a 50/200 day moving average rule and a shuffled rating placebo, with next open fills and trading costs. The pilot on AAPL and NVDA had the agents' trader levels lose money while buy and hold gained, and the ratings alone never traded. Four decisions per stock is a method check, not evidence, and the page says so.

Backtest results: the agents against buy and hold and a moving average rule

5. Then trade it, on paper

Entry, stop and target are what a broker needs for an order. One script lets the committee decide, then sends the decision to an Interactive Brokers paper account as one bracket order per stock, with the trader's stop and the portfolio manager's target attached. Long only, one decision, fixed size, paper ports only, and a y/N confirmation before anything is sent. The order stays linked to the run that produced it, so every trade traces back to the argument behind it.

trade.py

py -3 trade.py NVDA MSFT              # run the twelve agents on each stock, then trade
py -3 trade.py NVDA MSFT --reuse      # replay each stock's latest finished run, then trade
py -3 trade.py NVDA --dry-run         # everything except sending the orders

6. Install in four steps

Python 3.11 or newer, Node 22 or newer, Git, and an API key for one LLM provider. The engine is cloned at its released tag, next to the app.

Terminal

git clone https://github.com/davidalmeida90/glassbench.git
cd glassbench
git clone --branch v0.5.0 --depth 1 https://github.com/TauricResearch/TradingAgents.git

py -3.11 -m venv .venv            # Mac or Linux: python3 -m venv .venv
.\.venv\Scripts\Activate.ps1      #                source .venv/bin/activate
pip install -r requirements.txt
pip install -e .\TradingAgents

copy .env.example .env            # paste your key after the equals sign
.\glassbench.ps1                  # Mac or Linux: ./glassbench.sh

The first start builds the frontend, then the app is at http://127.0.0.1:8765. Press New run, pick a ticker, a date, the analysts, the provider and the two models, and watch the committee work.

Honest limits