AI Agent Hub
Back to skills
General Algorithm Training icon

General Algorithm Training

Development Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Please install @user_db1d96f9/music-algo-train according to https://skillhub.cn/install/skillhub.md.

About this skill

Problem

In amt_ai and YourMT3-based melody transcription projects, the hard part is rarely discovering a training framework. It is repeatedly reconciling repository-specific conventions: when to use scratch, resume, or warm-start; how to size batch size and num_workers for DDP; when to choose bf16-mixed over 16-mixed; and where to start when NaN, stalled loss, DDP hangs, or empty W&B runs appear. This skill packages the training conventions already used in the repository into reusable context for editing training scripts, tuning hyperparameters, and debugging runs.

How It Works

  • Training identity comes first: the skill distinguishes scratch, resume, and warm-start based on --init-ckpt, the presence of last.ckpt, and AMT_FORCE_SCRATCH, and reminds users that max_steps must exceed the current global_step before resume.
  • Script shape stays consistent: new training scripts should follow the “environment variable overrides + args array + explicit init-ckpt / resume” pattern, with # 修改: style comments where defaults are changed.
  • Hyperparameters have defaults: examples include AdamW as the more stable optimizer, bf16-mixed as preferred precision, warmup_steps=500, checkpoint monitoring on validation/macro_onset_f, save_top_k=7, and last.ckpt stored as a real copy.
  • Debugging path is concrete: move from bf16-mixed to a short 32 run, enable AMT_DEBUG_GRAD_FINITE and AMT_DETECT_ANOMALY, inspect DDP num_workers, and check whether W&B credentials fall back to CSVLogger.

Boundaries

This skill is specific to AMT / YourMT3 training and should not be treated as a general PyTorch template. Capabilities that are not implemented or are disabled by default, such as certain deepspeed offload paths, additional RL branches, or dynamic dataset weighting, should be enabled only after checking repository defaults. If the project does not use the ~/code/amt_ai layout, confirm that src/train.py, config.py, and data_presets.py match before relying on the conventions.

Use Cases

  • Create a new training script in amt_ai and set scratch, resume, or warm-start correctly.
  • Tune DDP training by setting batch size, num_workers, bf16-mixed, and checking last.ckpt.
  • Debug NaN or stalled loss by enabling gradient and anomaly detection to find the first bad op.
  • Configure W&B and checkpoints, monitor the main metric, and keep a recoverable last.ckpt.

Best For

  • Algorithm engineers maintaining AMT melody transcription training who need consistent scratch, resume, and warm-start rules.
  • ML engineers running DDP pipelines who need to set batch size, workers, precision, and checkpoints.
  • Researchers debugging NaN, DDP hangs, or missing W&B runs who need repository-level safeguards and first-error traces.
  • Applied researchers writing new AMT training scripts who want the same optimizer, scheduler, and logging conventions.