PyTorch Development Patterns
Paste the following prompt into your AI chat to install this skill:
Please install @user_e514343f/xrqtest according to https://skillhub.cn/install/skillhub.md.
About this skill
Problem
PyTorch projects often fail in training details rather than model architecture: hard-coded devices, missing random seeds, unchecked tensor shapes, forgotten train()/eval() transitions, growing GPU memory, and slow data loading. This skill turns those recurring issues into reviewable patterns for new models, training scripts, code review, training-loop debugging, and data pipelines.
How It Works
It organizes guidance around device independence, reproducibility, and explicit shape management. Common practices include using .to(device) instead of hard-coding GPUs, fixing random sources with torch.manual_seed, documenting output shapes in nn.Module, and using torch.no_grad(), optimizer.zero_grad(set_to_none=True), and model.train()/model.eval() for standard training and validation loops. Data coverage includes Dataset, DataLoader, pin_memory=True, and variable-length collate_fn logic. Experiment handling covers checkpoint save/load, with a reminder to prefer weights_only=True for safer loading. Performance guidance covers torch.amp.autocast mixed precision, gradient_checkpointing, and torch.compile, with validation through torch.profiler and torch.cuda.memory_summary().
Boundaries
It fits engineering checks for PyTorch training code and inference services, not full distributed training, data-science experimentation, or model compression. Behavior of torch.compile, mixed precision, and memory strategies depends on version, operators, and input shapes, so results should be profiled on real data.
Use Cases
- Use it when writing a new PyTorch model to check device placement, random seeds, and explicit tensor shape documentation before training starts.
- Use it while reviewing training loops to catch missing train/eval toggles, no_grad inference, and confirm optimizer gradient clearing is efficient.
- Use it when debugging data pipelines to configure DataLoader, pin_memory, and collate_fn for variable-length samples while keeping CPU to GPU transfer efficient.
- Use it for GPU memory or speed issues to plan autocast, gradient checkpointing, and torch.compile validation before rolling into training.
Best For
- ML engineers maintaining training scripts who need stable seeds, train/eval toggles, and lower memory pressure.
- Engineers reviewing PyTorch code who need device, shape, checkpoint, and secure-load checks.
- ML engineers maintaining data pipelines who need DataLoader, pin_memory, and variable-length collate tuning.
- Engineers deploying inference services who need no_grad, autocast, and torch.compile performance checks.
Related Skills
A one-shot coding agent built on Claude Code CLI that runs non-interactively, supports a specified workdir, and can be monitored in the foreground or background.
Preview and confirm file sorting by extension, with recursive cleanup, ignore rules, and transactional rollback.
An engineering assistant for static HTML/CSS/JS pages, design-token extraction, IE8-compatible review, and structured delivery.
An engineering workflow for requirement analysis, scenario modeling, risk planning, quality gates, testing, and knowledge capture, with lightweight, standard, and full modes.