BenchMendRead the essay

LIBERO · EVALUATION GALLERY

LIBERO

Explore GPT-6-Astra rollouts with original and revised LIBERO instructions. Compare the same initial state, inspect the instruction edits, and watch each evaluation.

Explore the episodes

WATCH, COMPARE, EXPLORE

The episode collection

Browse all files

Original task instructions, evaluated across all 40 tasks.

Loading episodes…

Outcomes come from the original benchmark checker. Select a numbered episode to watch and compare instructions.

ABOUT THE DATA

A repair you can
inspect, episode by episode.

BenchMend uses an agent to help identify gaps between task instructions and benchmark verification. This gallery makes the LIBERO evaluations available for direct inspection.

Open the dataset card
What is in each version?

Original contains 400 episodes: 10 initial states for each of 40 tasks. Revision contains 90 episodes using the latest revised instructions for nine tasks, evaluated on the same 10 initial states per task.

What does success mean here?

Success and failure are the native benchmark checker's outcomes. They are not a human judgment of whether the visible behavior satisfies the instruction. Instructions were revised; the benchmark checker was retained.

How do I compare a rollout?

Open an episode and switch between the available versions for the same task and initial state. Added or replaced words appear in bold in the revised instruction. Removed words can be seen in the original instruction.

LIBERO

Episode

ORIGINAL INSTRUCTION

REVISED INSTRUCTION · EDITS IN BOLD