BenchMendRead the essay

ROBOTWIN · EVALUATION GALLERY

RoboTwin

Explore GPT-6-Astra rollouts with original and revised RoboTwin instructions. Compare available versions, inspect the instruction edits, and watch each evaluation.

Explore the episodes

WATCH, COMPARE, EXPLORE

The episode collection

Browse all files

Original task instructions, evaluated across all 50 tasks.

Loading episodes…

Outcomes come from the original benchmark checker. Select a numbered episode to watch and compare instructions.

ABOUT THE DATA

A repair you can
inspect, episode by episode.

BenchMend uses an agent to help identify gaps between task instructions and benchmark verification. This gallery makes the RoboTwin evaluations available for direct inspection.

Open the data provenance
What is in each version?

Original contains 500 episodes: 10 slots for each of 50 tasks, including 482 original NVIDIA-run recordings and 18 local Codex recoveries. Original has 131 episodes where the agent reported completion but the benchmark judged failure. Revision contains exactly one latest confirmed original-linked rerun for each of those episodes, across 29 tasks. The six latest dual-shoe revisions are included, with failures and timeout retained.

What does success mean here?

Success and failure are the native benchmark checker's outcomes. They are not a human judgment of whether the visible behavior satisfies the instruction. Instructions were revised; the benchmark checker was retained.

How do I compare a rollout?

Open an episode and switch between recordings for the same task, episode slot and reported seed. Revision keeps one latest confirmed rerun per original disagreement episode. Added or replaced words appear in bold. Matching seeds across different source runs do not by themselves guarantee identical simulator state.

RoboTwin

Episode

ORIGINAL INSTRUCTION

REVISED INSTRUCTION · EDITS IN BOLD