isaaclab-transferring-policies-sim-to-sim
Installation
SKILL.md
Transfer Policies Between PhysX And Newton
When To Use
Read the sim-to-sim how-to first. This skill follows that page in the same order. Before transfer, make the asset and task MJWarp-ready with isaaclab-preparing-assets-for-newton.
Workflow
- Task readiness and checkpoint compatibility. Keep one MDP for the same registered task. Resolve the explicit
isaacsim_physxandnewton_mjwarpbackend alternatives from eachPresetCfg, using them for intentional physics, asset, and control differences without silently changing policy-facing MDP terms. If an MDP-term preset differs, restore one checkpoint contract or treat it as a different task and retrain. Require successful training in both engines and exact action, observation, policy-state, timing, mechanism, and episode contracts. - Mimic-joint action nuance. Newton MJWarp preserves both Franka finger coordinates but creates one active drive: the leader is driven and an equality moves the follower. PhysX preserves the mimic coupling while leaving the follower driveable. If one logical command targets both fingers and both have nonzero gains, PhysX applies two PD-drive contributions. Drive
panda_finger_joint1and setpanda_finger_joint2stiffness and damping to zero. - Transferring control behavior. Match nominal actuator response before tuning the policy. Distinguish rated and solver velocity limits, use per-joint effort, gains, friction, and armature, preserve
dt * decimation, keep targets away from hard stops, and monitor saturation and action sign changes. Increase damping to prevent bang-bang control and retune it after increasing armature. - Introducing domain randomization. Randomize plausible robot/object friction, object mass/inertia, joint gains/friction, joint armature, gravity, actuator response, reset pose/geometry, and observation noise. Keep inertia valid and coupled mechanisms coherent. If transfer needs extreme ranges, revisit the nominal model. Use curriculum when the final distribution blocks learning, promote to final deployment difficulty, and keep a separate deterministic nominal evaluation.
- Validate the full matrix. Evaluate PP, PN, NN, and NP. For each source policy, reproduce the same-backend baseline and deploy the exact checkpoint in the other backend.
- Run the Franka lift transfer. Use
Isaac-Lift-Frankafor both training and inference, selectingisaacsim_physxfor the PhysX runs. The play entry point applies the environment'splay_modeoverrides automatically. Follow the PhysX-to-MJWarp and MJWarp-to-PhysX commands in the how-to.
Validation
Require a task trainable in both backends, exact environment-contract equality, one active Franka finger drive in each backend, matched nominal control behavior, plausible randomization, and all four training/deployment combinations.