Explainable-AI methods are usually impossible to check: nobody knows the true inner workings of a deep network, so an explanation can be convincing and wrong with no way to tell. We swap in a 1977 Atari 2600 game console instead — a real, complex computer whose every wire and byte is known. For any game we can compute the true cause of any pixel by literally changing a byte and re-running the machine bit-for-bit. That true answer is the ground truth every method is graded against. This page lists, game by game, what ground truth we have and where it comes from.
Not all ground truth is the same. We separate what the machine can tell us by itself from what only a human can supply.
An explanation is not “right” just because it looks convincing. We ask three separate things, and an explanation only counts as right when it clears all three.
On this machine all three are computed against the oracle (T1/T2), never against another method — which is what keeps the grade honest. The per-method F/S/M scores live on the P2 Methods page; this page shows, per game, what those measures can be computed against.
The T3 names are external, so we are explicit about their source. We import from AtariARI (Anand et al., 2019 — RAM annotations for 22 games) and OCAtari (Delfosse et al., 2024 — 54 games, with the render-time offset correction that aligns a RAM coordinate to where the sprite is actually drawn). Every label is then verified on our bit-exact machine before use, so a name is only kept if changing that byte demonstrably moves the named object. A method that “matches” a label is therefore aligning to a meaning we supplied — it did not discover it.
All 64 bit-exact ROMs (54 carry a T3 label, 10 are T1/T2 only), ordered by how much ground truth the shared analysis frame supports (position-regime games first). Screenshots are the actual analysis frame the ground truth is computed on. Colour: green = available, amber = present but limited at this frame, red = not available. The 10 games with no external label carry exact T1/T2 ground truth but are held out of the label-dependent (T3) study.
Of the 54 labelled ROMs, 42 make up the scored battery; the other 12 are excluded (dimmed and tagged below). We exclude a game for one of two honest reasons. Eight fail the cause-density gate: at the shared frame the chosen output has almost no true causes, so there is nothing for a method to be right or wrong about. Four have no sprite that moves under intervention — their verified labels are static things like a score counter, so there is no position for a position method to recover; three games fail both tests. Running interpretability methods on such a frame cannot measure faithfulness — it would only add noise to the leaderboard — so we leave those games out of scoring while keeping their exact T1/T2 ground truth on record. Every game in the scored battery, by contrast, both passes the gate and carries a moving sprite, so it contributes to the all-regime and the position regime alike.
b1.xyplayer.xyenemy_x_part_1player.xyplayer.xyenemy_robots_x[0]ball_xenemy_xplayer_xmissile.xyplayer.xyplayer.xyenemy_projectile_yp1s.hook_positionc2.xyfourth_row_iceflow_xgopher.xydestructible_wall.xyball.xychild.xyplayer.xyenemy_skull_xplayer.xyenemy_logs_xball_xcar.xyplayer_yplayer_xenemy.xydiver_or_enemy_missile_xenemies_xball_xplayer_xball_x




















