| name | ml-experiment |
| description | Design and run machine learning experiments with proper evaluation using jupyter_execute, including training, benchmarking, and ablation studies |
ML Experiment Skill
Description
Design, implement, and evaluate machine learning experiments with reproducible workflows, proper baselines, and statistical analysis.
Tools Used
jupyter_execute - Execute ML code in Python (auto-switches to Jupyter)
jupyter_notebook - Manage experiment notebooks
update_notebook - Set up experiment cells
update_latex - Write experiment results to papers
latex_compile - Compile CS conference papers (auto-switches to LaTeX)
arxiv_to_prompt - Read related work from arXiv papers
update_notes - Write experiment logs and analysis summaries
Capabilities
Experiment Design
- Proper train/validation/test splits
- Cross-validation and bootstrap confidence intervals
- Ablation study design
- Hyperparameter search (grid, random, Bayesian)
Implementation
- PyTorch and TensorFlow model building
- Data loading and augmentation pipelines
- Training loops with logging and checkpointing
- Distributed training setup
Evaluation
- Standard metrics per task (accuracy, F1, BLEU, mAP, etc.)
- Statistical significance testing (paired t-test, bootstrap)
- Comparison with baselines
- Error analysis and visualization
Usage Patterns
Run an Experiment
When user says: "Train a model for [task]"
- Clarify dataset, metrics, and baselines
- Implement data loading and preprocessing
- Build model architecture
- Train with proper logging
- Evaluate and compare to baselines
- Report results with confidence intervals
Reproduce a Paper
When user says: "Reproduce [paper title/arXiv ID]"
- Fetch paper using arxiv_to_prompt
- Extract key method details
- Implement core algorithm
- Run experiments matching paper setup
- Compare results to reported numbers