General Algorithm Training
Paste the following prompt into your AI chat to install this skill:
Please install @user_db1d96f9/music-algo-train according to https://skillhub.cn/install/skillhub.md.
About this skill
Problem
In amt_ai and YourMT3-based melody transcription projects, the hard part is rarely discovering a training framework. It is repeatedly reconciling repository-specific conventions: when to use scratch, resume, or warm-start; how to size batch size and num_workers for DDP; when to choose bf16-mixed over 16-mixed; and where to start when NaN, stalled loss, DDP hangs, or empty W&B runs appear. This skill packages the training conventions already used in the repository into reusable context for editing training scripts, tuning hyperparameters, and debugging runs.
How It Works
- Training identity comes first: the skill distinguishes
scratch,resume, andwarm-startbased on--init-ckpt, the presence oflast.ckpt, andAMT_FORCE_SCRATCH, and reminds users thatmax_stepsmust exceed the currentglobal_stepbefore resume. - Script shape stays consistent: new training scripts should follow the “environment variable overrides +
argsarray + explicitinit-ckpt/resume” pattern, with# 修改:style comments where defaults are changed. - Hyperparameters have defaults: examples include
AdamWas the more stable optimizer,bf16-mixedas preferred precision,warmup_steps=500, checkpoint monitoring onvalidation/macro_onset_f,save_top_k=7, andlast.ckptstored as a real copy. - Debugging path is concrete: move from
bf16-mixedto a short32run, enableAMT_DEBUG_GRAD_FINITEandAMT_DETECT_ANOMALY, inspect DDPnum_workers, and check whether W&B credentials fall back toCSVLogger.
Boundaries
This skill is specific to AMT / YourMT3 training and should not be treated as a general PyTorch template. Capabilities that are not implemented or are disabled by default, such as certain deepspeed offload paths, additional RL branches, or dynamic dataset weighting, should be enabled only after checking repository defaults. If the project does not use the ~/code/amt_ai layout, confirm that src/train.py, config.py, and data_presets.py match before relying on the conventions.
Use Cases
- Create a new training script in amt_ai and set scratch, resume, or warm-start correctly.
- Tune DDP training by setting batch size, num_workers, bf16-mixed, and checking last.ckpt.
- Debug NaN or stalled loss by enabling gradient and anomaly detection to find the first bad op.
- Configure W&B and checkpoints, monitor the main metric, and keep a recoverable last.ckpt.
Best For
- Algorithm engineers maintaining AMT melody transcription training who need consistent scratch, resume, and warm-start rules.
- ML engineers running DDP pipelines who need to set batch size, workers, precision, and checkpoints.
- Researchers debugging NaN, DDP hangs, or missing W&B runs who need repository-level safeguards and first-error traces.
- Applied researchers writing new AMT training scripts who want the same optimizer, scheduler, and logging conventions.
Related Skills
Automatically indexes Gradle-cached AAR/JAR dependency classes and returns library coordinates, versions, and public APIs by fully qualified name, using only the Python standard library.
Retrieve relevant chunks from a customer-managed PKM dataset by dataset_id and return concise, source-annotated answers.
Convert PRDs, user stories, or functional specs into prioritized test-point checklists covering functional, business-rule, boundary, exception, and non-functional dimensions.
Scan frontend repos and monorepos to generate executable Python or Node.js CLIs with API, type, enum, and auth support.