Agent4CT — Live dashboard

Every run, every iteration, every observation.

← Back to project GitHub
Home Setup Pentathlon Agents Performance Live dashboard GitHub

Pick a dataset

Each dataset has its own dashboard — runs, progress curves, and comparison figures load only for the dataset you choose, so the page stays fast.

loading datasets…

← All datasets

Dataset

Best headroom per run vs iteration (running max). 1.0 = matched the oracle on the in-loop val subset.

loading…

Runs

Each card is one autoresearch run (<solver>-YYYYMMDD-NN) with its best-iter validation and held-out test figures. Click to expand the per-iteration timeline.

loading runs…


Run

Progress curve

Iterations

Stage checks (every 30 iter)


Shared scratch pad

Cross-run observations for this dataset (most recent first, capped).

Data fetched directly from docs/runs/. Last refreshed —.
↗