isaaclab-debugging-rl-training

Installation
SKILL.md

Debugging RL Training

When To Use

Use this skill when a user needs to debug learned behavior, reward hacking, checkpoint compatibility, unstable training, or training-result trustworthiness.

Do not use this skill for first-time training commands. Use isaaclab-training-rl-agents for launch commands and agent config wiring.

Workflow

  1. Identify task name, workflow type, RL library, agent config, seed, backend, and exact launch command.
  2. Confirm the environment contract: action space, observation space, reward terms, termination terms, reset logic, and success metric.
  3. Run the smallest reproduction: import, reset/step, one-iteration training, or deterministic playback depending on where the failure appears.
  4. Change one variable per training experiment. Mark multi-variable runs as exploratory.
  5. Compare reward curves against task metrics. Reward increases are not proof that the task behavior improved.
  6. For reward issues, map every reward term to a named task phase and check that success reward, termination, and evaluation metric use consistent geometry.
  7. For checkpoint issues, compare current observation/action dimensions with the saved training configuration before editing policy code.
  8. For contact-rich tasks, collect state traces for controlled-frame pose, object pose, contacts, gripper state, per-term rewards, and termination flags.
  9. Select checkpoints by task metrics, rollout behavior, and stability, not reward alone.
Installs
1
First Seen
Aug 13, 2026
isaaclab-debugging-rl-training — kaweees/isaaclab-skill