PyTorch
Installation
SKILL.md
Train vs Eval Mode
model.train()enables dropout, BatchNorm updates — default after initmodel.eval()disables dropout, uses running stats — MUST call for inference- Mode is sticky — train/eval persists until explicitly changed
model.eval()doesn't disable gradients — still needtorch.no_grad()
Gradient Control
torch.no_grad()for inference — reduces memory, speeds up computationloss.backward()accumulates gradients — calloptimizer.zero_grad()before backwardzero_grad()placement matters — before forward pass, not after backward.detach()to stop gradient flow — prevents memory leak in logging
Device Management
- Model AND data must be on same device —
model.to(device)andtensor.to(device) .cuda()vs.to('cuda')— both work,.to(device)more flexible- CUDA tensors can't convert to numpy directly —
.cpu().numpy()required torch.device('cuda' if torch.cuda.is_available() else 'cpu')— portable code