AI Agent Hub
Back to skills
AutoML Automated Machine Learning icon

AutoML Automated Machine Learning

Data Analysis Updated 2026.08.29

Paste the following prompt into your AI chat to install this skill:

Follow https://skillhub.cn/install/skillhub.md to install @user_b09806db/automl in your AI assistant.

About this skill

Problem

Tabular model tuning often depends on manual trial and error: learning_rate, num_leaves, cleaning, model switching, and ensemble strategies all require repeated runs and comparisons. This skill wraps CSV modeling into an autonomous experiment loop, letting an agent iterate on CPU toward a target metric and keep effective configurations.

How It Works

The flow starts from parameters: data_path, target_col, task_type, metric, time_budget, and max_experiments. It then checks the data, initializes a workspace, and sets a baseline. In the loop, it reads results.tsv, designs experiments by priority: hyperparameter tuning > model switching > feature engineering > ensembling; edits train.py, runs training, parses validation metric, memory, and runtime, then decides whether to commit or revert. It supports classification, regression, and multiclass tasks, with metrics such as accuracy, f1_macro, rmse, mae, and r2, and reports feature importance and overfitting diagnostics when available.

Boundaries

It is best suited to medium-scale structured data, limited compute budgets, and interpretable classical ML configurations. If the data has severe missing values, the target column is unclear, or the task requires deep learning or long-running model training, clean or curate the data first or use a different pipeline. The experiment loop runs autonomously until max_experiments is reached or the user stops it.

Use Cases

  • Given a CSV file with a target column, run CPU-based AutoML for classification or regression and output the best classical ML configuration.
  • When a baseline model stalls, continue experiments in order of tuning, model switching, feature engineering, and ensembling, then keep only effective changes.
  • For datasets around 100K rows, compare LGB and XGB configurations quickly using metrics like `f1_macro` or `rmse` instead of writing training scripts manually.
  • Review train-validation gap, feature importance, and kept/discarded experiment counts to diagnose overfitting.

Best For

  • ML engineers turning sales, support, or user-behavior tables into prediction models and reducing manual hyperparameter tuning time.
  • Business analysts building churn, pricing, or risk classification models who need interpretable classical ML and evaluation reports.
  • Engineers using AI agents for data experiments who want autonomous training runs, result logging, and rollback of ineffective configs.
  • Platform engineers maintaining data-science pipelines who need CSV modeling, baselines, and experiment records in one workflow.