Glassbench · Apache 2.0

Watch AI trading agents think.

An open-source UI for TradingAgents, AI Hedge Fund and other open-source trading frameworks. Watch every agent work live, benchmark different LLMs on the same framework and backtest their decisions. Ships with 77 recorded runs and their full logs, so everything can be checked before spending a cent.

Research tool · paper accounts only · not investment advice

Agents per run
12in 5 stages
Runs published, full logs
77
Cost per run on DeepSeek
$0.06
LLM providers
13+ any model id
RATING=FRAMEWORK+MODEL+DATA+GLASS

Every step on the record

A rating is the last line of a long conversation between twelve agents. Glassbench keeps the whole conversation, so any rating can be checked against what produced it.

Framework, untouched

TradingAgents 0.5.1 and AI Hedge Fund 2.4.0 run as released. Glassbench observes each one through its own adapter with its own tests, so an engine update breaks in one known place and nowhere else.

backend/deskapp/adapter.py

Model, your choice

Any provider the engine supports and any model id you type. Cheap fast models read and debate; only the two managers call the expensive one, twice per run.

quick model · deep model

Record, complete

Every prompt, reasoning trace, tool call and result lands in SQLite as a numbered event. Searchable, exportable, and replayable step by step long after the run ended.

81,186 events in the published database
Inside Glassbench

Built for people who want to check

Most AI trading tools hand over a signal. Glassbench hands over the committee that produced it, with the receipts.

Watch the committee work, card by card

Twelve agents in five stages, each on its own card: what it is reading, what it is writing, what it concluded. Click one for every LLM call, every tool result and the reasoning as it streams, then replay the whole run later.

127.0.0.1:8765/runs · MSFT · running

A database, never a log folder

Every run lands in SQLite with its framework version, models, ticker, date, rating, levels, tokens and cost. Filter on any field, search the full text of every report and reasoning trace, and export any view to CSV.

127.0.0.1:8765/database
Database page: facets on the left, a breakdown table and every run with its framework, variant, models and rating

Benchmark LLMs on one framework

Keep the framework fixed and swap the models. Pick any of 13 providers, type any model id for the quick and deep roles, and compare the runs side by side with their cost. DeepSeek, the default, costs about six cents a run.

New run · provider and models
New run dialog with OpenRouter selected and two custom model ids typed in

Backtests that are allowed to say no

Weekly grids of past dates with a budget cap, replayed against buy and hold, a moving average rule and a shuffled rating placebo, with next open fills and trading costs. On the first pilot, buy and hold won on both stocks.

Backtests · Pilot: AAPL + NVDA · 10 bps
Backtest results: variants on the same dates, and growth of 100 for the agents against buy and hold

Then trade it, on paper

Entry, stop and target are what a broker needs. One script sends the committee's decision to an Interactive Brokers paper account as a bracket order, and the order stays linked to the run that argued for it.

Glassbench trade
$ py -3 trade.py NVDA MSFT --reuse 3 DECISIONS NVDA Overweight buy 220.50 stop 214.35 target 235.00 MSFT Underweight no order: none held, long only 4 ORDERS one bracket order per stock NVDA BUY 44 LMT 225.62 stop 214.35 target 235.00 Send 1 order to the PAPER account? [y/N] y 5 EXECUTION NVDA order BUY 44 sent NVDA fill 44 at 225.11 NVDA target 235.00 attached, GTC NVDA stop 214.35 attached, GTC 6 RESULT 1 order, linked to its run run 20260921-172349-NVDA-4761
Glassbench trade · terminal and Interactive Brokers TWS · sped up

Committee decision to filled bracket order, target and stop attached, on an Interactive Brokers paper account.

Already running TradingAgents?

What Glassbench adds to the terminal

Same engine, same agents, same prompts. Glassbench changes what you can see, keep and compare.

TradingAgents terminalGlassbench
Follow a runA status table and a scrolling feed of messagesOne card per agent with what it is doing and what it concluded, a stage strip and a timeline of every call
Open one agentScroll back through the terminalEvery LLM call with its tokens, every tool call with its full result, the reasoning stream, one click away
ReplayNot availableAny run, step by step, long after it finished
Compare runsA report folder per runEvery run in one table, filtered on any field, exported to CSV
Search the logsGrep the filesFull text search over every report and every reasoning trace
CostToken counts at the bottom of the screenTokens and dollar cost stored per run and per agent
ModelsAny supported provider, picked at the promptAny provider and model id, recorded per run next to the model actually served
BacktestA grid of runs scored on realised alphaWeekly grids with a budget cap, replayed with fills and costs against buy and hold, a moving average rule and a placebo
Change what agents readEdit the engineNamed variants applied per run, with the engine untouched
Send to a brokerNot availableOne bracket order at an Interactive Brokers paper account, linked to its run
Compared with the TradingAgents 0.5.1 command line interface. Glassbench runs that same release, unmodified.
Published with the code

Check the logs database before spending a cent

Not investment advice. Ratings, price levels and reports below are language model output from a research tool, published so the method can be judged. None of it is a recommendation to buy, sell or hold any security.

Glassbench ships with the database of every run made while building it, complete logs included: every report, every tool output, every reasoning trace, and what each run cost. Open it in the app, query the SQLite file, or read any run on GitHub from the table below. No API key needed.

Runs
77
Tickers
9
Tokens
12.7M
LLM calls
974
Tool calls
1,422
Reasoning events
54,821
Total cost
$2.85
Ratings by ticker67 rated runs, both frameworks
OverweightHoldUnderweightSell
Same stock, same day, two minutes apartMSFT · 2026-09-19
Underweight22:07
Entry498.98
Stop513.53
Target464.20
Trim MSFT toward 0.5 to 0.75x a standard allocation, a 25 to 50% reduction, not an exit. Use strength into the 498.98 to 505.06 supply zone as the preferred trim window. Open run 2e8c on GitHub →
Overweight22:09
Entry493.78
Stop473.00
Target590.00
Build MSFT toward an Overweight using a scaled entry, but start smaller than the original plan: take a quarter to a third of the intended overweight target near 493.78. Open run 2938 on GitHub →
Same models, same data, same prompts. One rating says little until you know how often a rerun flips it, which is why every run stays in the database.
Every run
DateTickerRatingEntryStopTargetHorizonModelsEngineCostLogs

Every row links to its folder on GitHub: analyst reports, both debates, trader proposal, risk debate, final decision, every tool output and the final state. All of it also lives in data/desk.db, which opens in the app with no key at all.

Measured, not claimed

What the runs show so far

Two results from the published database, with the method written down and the limits stated next to the numbers.

Serving valuation as of the run date24 runs · 12 cells
Statements that valuation data is missing
Statements
55
+ valuation
1
Multiples cited that match the true figure on the date
Statements
25
+ valuation
397
LLM cost of the 12 runs
Statements
$0.548
+ valuation
$0.523
12 of 12 cells moved the same way · sign test p = 0.0005 · ratings not claimedMethod and data →
Backtest pilot, 256 trading days4 decisions per stock · 10 bps
StrategyAAPLNVDAEqual weight
Agents, trader levels-6.0%-7.4%-6.7%
Agents, rating only0.0%0.0%0.0%
Buy and hold+43.5%+29.6%+36.6%
50/200 day MA rule+31.0%+29.6%+30.3%
Replayed 23 Sep 2026, Sep 2025 to Sep 2026. Every level based entry was stopped out; four decisions per stock is a method check, not evidence.Backtest method →

Watch it end to end

13 min · setup, method, live run, Glassbench, a paper order · subtitles EN, PT, ES
Frameworks

One workbench, many committees

Each framework plugs in through its own adapter and keeps its own name and version on every run, so results from different frameworks never blur together.

TradingAgents

Tauric Research · paper arXiv 2412.20138
Connected

A trading firm in twelve agents: four analysts pull data, a bull and a bear argue, a research manager writes a plan, a trader proposes levels, three risk debaters push back and a portfolio manager rates the stock on a five point scale.

v0.5.1Apache 2.0Python · LangGraph108k stars

AI Hedge Fund

virattt · open source
Connected

A fund in agents: investor personas and an earnings drift model each signal the stock, their blended conviction passes risk limits and a simulated fill. Glassbench runs it untouched in its own environment, shows every analyst as a lane and maps the blend to the same five ratings, so both frameworks land in one table.

v2.4.0MIT8 runs published63k stars
Get started

Running in four steps

Python 3.11 or newer, Node 22 or newer, Git, and a key for one LLM provider. Reading the published runs needs no key at all.

Install

Clone Glassbench, clone the engine next to it at its released tag, install both, add a key.

> git clone https://github.com/davidalmeida90/glassbench.git
> cd glassbench
> git clone --branch v0.5.1 --depth 1 https://github.com/TauricResearch/TradingAgents.git
> py -3.11 -m venv .venv
> .\.venv\Scripts\Activate.ps1
> pip install -r requirements.txt
> pip install -e .\TradingAgents
> copy .env.example .env     # paste your key
> .\glassbench.ps1

First run

Open http://127.0.0.1:8765 once the frontend finishes its first build.

  1. Press New run and type a ticker, for example NVDA.
  2. Pick the analysts, the debate rounds, the provider and the two models.
  3. Watch the twelve cards fill in. A full run takes about eight minutes.
  4. Open the run later from the Runs page, or compare it with the 77 that ship with the repository.
# no browser: one run from the command line
$ cd backend
$ python -m deskapp run NVDA --analysts market,news,fundamentals

Judge the agents yourself

Open source under Apache 2.0. Clone it, run your own committee, or start with the logs database of the 77 that already ran.

David Arias, CFA
Built by

David Arias, CFA

Licensed portfolio manager and CFA charterholder, working on emerging markets private debt, derivatives and quantitative finance. Every model on davidariasfinance.com ships with its code and its write up, and Glassbench follows the same rule: the method, the runs and the results are all public.

New runs, new frameworks, the results

One short email when something ships: a new framework adapter, a new batch of runs with their logs, or a backtest that says something. No spam.