Leaderboards

Best-of-best per solver per dataset under the canonical calibrated-SSIM-headroom scoring convention (evaluate_calibrated). All 3 leaderboards have full 19-solver inventory coverage. Mayo-LDCT is being re-run live (search-20260619-01) under the corrected metric — the Mayo champion row below auto-updates each wave from the run data; the breast/demo boards are stable. The cross-dataset counts further down are a 2026-06-09 snapshot.

Champions — rendered from the registry

The per-dataset champion (single canonical ranking = headroom, SSIM tiebreak) is rendered live from the registry below, so it can never drift from the dataset boards. Each panel is that dataset’s full all-solver board (below-baseline solvers dimmed, never dropped). For the prose write-ups and the baseline tables, open the dataset board linked under “Full standings”.

Mayo-LDCT

loading leaderboard…

Breast-CT

loading leaderboard…

BreastCT-Noise (high-dose robustness re-eval, I0 = 100k photons)

loading leaderboard…

BreastCT-Noise-Retrained (retrained on noisy train data, I0 = 100k photons)

loading leaderboard…

Demo-DL

loading leaderboard…

Full standings — every solver

The complete per-dataset rankings — all 19 solvers, with every column (params, SSIM, hr, PSNR, RMSE, time) — are on the dataset boards below. These are the single source of truth and list every solver (no truncated summary):

Methodology

See solver_plan.md for the canonical methodology — calibrated metric, per-solver hyperparameter spaces, autoresearch+TPE protocol.

Per-solver design docs and cross-dataset transfer records: pentathlon/demo_dl_reference/. Every canonical-19 solver has a dedicated .md design doc with cross-dataset hr record, CONFIG defaults, and “hints for the next autoresearch agent”.

Cross-cutting findings (DDPM training quality NOT predictive of DPS performance, FBP-init 1st-step-no-op mechanism, etc.) live in docs/findings.md.