LIBERO · EVALUATION GALLERY
LIBERO
Explore GPT-6-Astra rollouts with original and revised LIBERO instructions. Compare the same initial state, inspect the instruction edits, and watch each evaluation.
Explore the episodesWATCH, COMPARE, EXPLORE
The episode collection
Original task instructions, evaluated across all 40 tasks.
Loading episodes…
Outcomes come from the original benchmark checker. Select a numbered episode to watch and compare instructions.
The collection could not load.
Please reload the page to try again.
No matching episodes.
Try a different task, suite, or outcome.
ABOUT THE DATA
A repair you can
inspect, episode by episode.
BenchMend uses an agent to help identify gaps between task instructions and benchmark verification. This gallery makes the LIBERO evaluations available for direct inspection.
Open the dataset cardWhat is in each version?
Original contains 400 episodes: 10 initial states for each of 40 tasks. Revision contains 90 episodes using the latest revised instructions for nine tasks, evaluated on the same 10 initial states per task.
What does success mean here?
Success and failure are the native benchmark checker's outcomes. They are not a human judgment of whether the visible behavior satisfies the instruction. Instructions were revised; the benchmark checker was retained.
How do I compare a rollout?
Open an episode and switch between the available versions for the same task and initial state. Added or replaced words appear in bold in the revised instruction. Removed words can be seen in the original instruction.