isaaclab-debugging-rl-training
Installation
SKILL.md
Debugging RL Training
When To Use
Use this skill when a user needs to debug learned behavior, reward hacking, checkpoint compatibility, unstable training, or training-result trustworthiness.
Do not use this skill for first-time training commands. Use isaaclab-training-rl-agents for launch commands and agent config wiring.
Workflow
- Identify task name, workflow type, RL library, agent config, seed, backend, and exact launch command.
- Confirm the environment contract: action space, observation space, reward terms, termination terms, reset logic, and success metric.
- Run the smallest reproduction: import, reset/step, one-iteration training, or deterministic playback depending on where the failure appears.
- Change one variable per training experiment. Mark multi-variable runs as exploratory.
- Compare reward curves against task metrics. Reward increases are not proof that the task behavior improved.
- For reward issues, map every reward term to a named task phase and check that success reward, termination, and evaluation metric use consistent geometry.
- For checkpoint issues, compare current observation/action dimensions with the saved training configuration before editing policy code.
- For contact-rich tasks, collect state traces for controlled-frame pose, object pose, contacts, gripper state, per-term rewards, and termination flags.
- Select checkpoints by task metrics, rollout behavior, and stability, not reward alone.