AI Agent Hub
Back to skills
PyTorch Development Patterns icon

PyTorch Development Patterns

Development Updated 2026.08.30

Paste the following prompt into your AI chat to install this skill:

Please install @user_e514343f/xrqtest according to https://skillhub.cn/install/skillhub.md.

About this skill

Problem

PyTorch projects often fail in training details rather than model architecture: hard-coded devices, missing random seeds, unchecked tensor shapes, forgotten train()/eval() transitions, growing GPU memory, and slow data loading. This skill turns those recurring issues into reviewable patterns for new models, training scripts, code review, training-loop debugging, and data pipelines.

How It Works

It organizes guidance around device independence, reproducibility, and explicit shape management. Common practices include using .to(device) instead of hard-coding GPUs, fixing random sources with torch.manual_seed, documenting output shapes in nn.Module, and using torch.no_grad(), optimizer.zero_grad(set_to_none=True), and model.train()/model.eval() for standard training and validation loops. Data coverage includes Dataset, DataLoader, pin_memory=True, and variable-length collate_fn logic. Experiment handling covers checkpoint save/load, with a reminder to prefer weights_only=True for safer loading. Performance guidance covers torch.amp.autocast mixed precision, gradient_checkpointing, and torch.compile, with validation through torch.profiler and torch.cuda.memory_summary().

Boundaries

It fits engineering checks for PyTorch training code and inference services, not full distributed training, data-science experimentation, or model compression. Behavior of torch.compile, mixed precision, and memory strategies depends on version, operators, and input shapes, so results should be profiled on real data.

Use Cases

  • Use it when writing a new PyTorch model to check device placement, random seeds, and explicit tensor shape documentation before training starts.
  • Use it while reviewing training loops to catch missing train/eval toggles, no_grad inference, and confirm optimizer gradient clearing is efficient.
  • Use it when debugging data pipelines to configure DataLoader, pin_memory, and collate_fn for variable-length samples while keeping CPU to GPU transfer efficient.
  • Use it for GPU memory or speed issues to plan autocast, gradient checkpointing, and torch.compile validation before rolling into training.

Best For

  • ML engineers maintaining training scripts who need stable seeds, train/eval toggles, and lower memory pressure.
  • Engineers reviewing PyTorch code who need device, shape, checkpoint, and secure-load checks.
  • ML engineers maintaining data pipelines who need DataLoader, pin_memory, and variable-length collate tuning.
  • Engineers deploying inference services who need no_grad, autocast, and torch.compile performance checks.