Use when training RL policies for robots (locomotion, manipulation, navigation), debugging reward hacking or sim-to-real transfer failures, choosing between scripted and learned control, or setting up Isaac Gym/Lab massively-parallel training. Provides PPO configuration, reward shaping patterns, dom
Use when training RL policies for robots (locomotion, manipulation, navigation), debugging reward hacking or sim-to-real transfer failures, choosing between scripted and learned control, or setting up Isaac Gym/Lab massively-parallel training. Provides PPO configuration, reward shaping patterns, domain randomization ranges, and safe-RL constraints that work on real hardware.