Fiatlux, a new simulation benchmark, asks a humanoid to climb a ladder, swap a lightbulb, and dispose of the old one without breaking it. Policies with no task specific training clear none of its twelve subtasks.
Roboticists have published Fiatlux, a new simulation benchmark that asks a humanoid to do one long, multi-step task end to end: position a step ladder under a ceiling or wall fixture, climb it, swap a spent bulb for a fresh one, then drop the old bulb in a disposal crate. The bulb must survive the whole run (preprint, project page).
Fiatlux bundles three weaknesses into one episode: vertical mobility, fragile-payload handling, and a dexterous swap. According to the authors, prior benchmarks have scored those as separate groups. The full replacement episode is split into twelve registered subtasks, scored on both completion and force or drop violations (paper HTML).
Zero-shot GR00T N1.7, a zero-action policy, and a random-action policy each completed zero of the twelve subtasks across four layout seeds in an Isaac Lab simulation of a Unitree G1. The reported difficulty-weighted partial-progress scores sit near 0.087, which measures gate progress, not completion rates (project page).
A teleoperated policy scored 0.8419 on the eight non-climbing subtasks. The four climbing subtasks remain outstanding; the on-ladder stance demonstrations only initialize the robot on the top tread, not how it got there. The scorer also leans on a privileged simulator mode that exposes exact internal state, so the ceiling reflects what the simulator can verify, not what a real humanoid could sense (project page).