Our client is an applied R&D lab building an open-source humanoid robot platform and the full software stack behind it, working across hardware, locomotion, autonomy, simulation and infrastructure. This role owns the policies that take the robot from walking to running, recovering and moving robustly through the real world — not just performing well in simulation.
What You'll Do
Design, train and ship reinforcement learning policies for bipedal and whole-body locomotion on real hardware
Own the sim-to-real pipeline end to end, from training environment to first hardware run
Push balance, gait and recovery behavior past demo stage into something that survives pushes, slips and untrained terrain
Build and refine reward design, domain randomization and training environments in a physics simulator (e.g. MuJoCo)
Close the loop between simulation and hardware using real telemetry, feeding failures back into the next policy iteration
Collaborate closely with hardware, controls and manipulation researchers, since whole-body control spans team boundaries
Publish or open-source findings and tooling so the broader community can build on the work
Core Key Skills
Deep hands-on experience with reinforcement learning for continuous control, legged locomotion, or whole-body control
Proven track record of deploying learned policies onto real robots, not only into papers or simulators
Strong command of a physics simulator such as MuJoCo or Isaac, including reward shaping and domain randomization
Fluency in Python and modern RL tooling, with comfort working in a ROS2-based control stack
A bias for shipping — prioritizing real hardware iteration over polishing results in simulation
Strong diagnostic thinking about why a policy fails, not just whether it does
Nice to Have
Published work in locomotion, legged robotics, or sim-to-real transfer
Experience with model predictive control or classical locomotion methods alongside learning-based approaches
Contributions to open-source robotics or RL projects
Experience bringing up new hardware and debugging the gap between model and motor