Data Scientist Workflow Assistant
Paste the following prompt into your AI chat to install this skill:
Please follow https://skillhub.cn/install/skillhub.md to install @org-02qudk26/data-scientist-zh into your AI assistant.
About this skill
Problem to Solve
Data science work often stalls at concrete points: EDA stays at summary statistics without clarifying metric definitions, A/B tests focus on p-values while ignoring statistical power and business impact, and deployed models lack drift monitoring and reproducible documentation.
data-scientist-zh turns these steps into an engineering workflow: clarify the goal, constraints, and inputs, select statistical or machine-learning methods, then provide validation checks and actionable recommendations. It fits analysis and modeling tasks spanning pandas, scikit-learn, XGBoost, and PyTorch, as well as production-related work with SQL, PySpark, MLflow, and FastAPI.
How It Works
- Statistics and experiments: hypothesis testing,
A/B testing, causal inference, time-series analysis, Bayesian modeling - Machine learning: feature engineering, cross-validation,
SHAPexplainability, hyperparameter tuning - Data engineering: data profiling, missing-value handling, data quality, feature stores, pipeline orchestration
- Delivery and communication: visualization, dashboards, executive reporting, reproducible code and method notes
A typical run defines the business objective, explores the data, selects the right method, validates results, recommends next actions, and plans monitoring and maintenance.
Scope and Caveats
It focuses on data analysis and model workflows rather than replacing permissions, data access, or deployment environments. For compliance, production incidents, or financial risk work, regulatory, security, and governance constraints should be included in the task definition.
Use Cases
- Analyze churn data and train an interpretable churn-risk model.
- Evaluate a website A/B test and report statistical power and business impact.
- Build a time-series demand forecast with prediction intervals for inventory planning.
- Audit offline recommendation metrics and design drift monitoring after launch.
Best For
- Algorithm engineers building churn models who need interpretable metrics.
- Growth analysts running A/B tests who need validation and causal checks.
- Data scientists planning supply-chain demand forecasts with time series.
- Data engineers building analysis pipelines who need quality and monitoring norms.
Related Skills
Analyzes smart customer-service chat logs with semantic clustering to generate word clouds, Top K frequent questions, and personalized guess-you-ask recommendations.
Maps natural-language Reddit requests to KeyAPI REST workflows, validates endpoint contracts against official docs, and executes search, detail, comment, ranking, and report tasks.
Schedules fetching from NEP and custom academic sites, filters and ranks papers by keywords, generates Chinese summaries, and pushes Feishu cards with local download archiving.
Browser-based CSV/TXT visualizer with multi-Y-axis curves, smoothing, and PNG/JSON/CSV export.