Analyze Nixtla baseline forecasting results (sMAPE/MASE on M4 or other benchmark datasets). Use when the user asks about baseline performance, model comparisons, or metric interpretation for Nixtla time-series experiments. Trigger with "baseline review", "interpret sMAPE/MASE", or "compare AutoETS vs AutoTheta".
Analyze Nixtla baseline forecasting results (sMAPE/MASE on M4 or other benchmark datasets). Use when the user asks about baseline performance, model comparisons, or metric interpretation for Nixtla time-series experiments. Trigger with "baseline review", "interpret sMAPE/MASE", or "compare AutoETS vs AutoTheta".
Claude Code 1.0+; Python 3.10+; reads outputs from the nixtla-baseline-m4 workflow (statsforecast 1.7+).
Nixtla Baseline Review Skill
Overview
Analyze baseline forecasting results from the nixtla-baseline-m4 workflow. Interpret metrics, compare models, surface patterns, and recommend next steps.
When to Use This Skill
Activate this skill when the user:
Asks "Which baseline model performed best?"
Requests interpretation of sMAPE or MASE metrics
Wants to compare AutoETS vs AutoTheta vs SeasonalNaive
Says "Explain these baseline results"
Needs guidance on model selection based on baseline performance
For StatsForecast Power Users
This baseline lab is built on Nixtla's statsforecast library. What this plugin provides:
Real statsforecast models (SeasonalNaive, AutoETS, AutoTheta)
M4 dataset integration via datasetsforecast
Standard train/test evaluation with sMAPE and MASE metrics
MASE: < 1.0 means better than seasonal naive baseline
Step 5: Generate Structured Explanation
Provide analysis in this format:
## Baseline Performance Analysis
Based on {N} series from M4-Daily with horizon={H}:
**Overall Winner**: {ModelName}
- Mean sMAPE: {X.XX}% (vs {Y.YY}% for second place)
- Wins on {N} out of {Total} series ({Percentage}%)
- Consistent performance (std dev {X.X}%)
**Key Insights**:
1. {Insight about winner's strengths}
2. {Insight about other models or patterns}
3. {Insight about failure cases or opportunities}
**Recommendation**: Use {ModelName} as production baseline. {Additional context}
**Next Steps**:
- {Actionable recommendation 1}
- {Actionable recommendation 2}
Examples
Example 1: Simple Query
User: "Which model was best?"
Response:
Based on the baseline results (50 series, horizon=14):
**AutoTheta won** with average sMAPE of 12.3%
- Outperformed AutoETS (13.5%) and SeasonalNaive (15.2%)
- Most consistent with std dev of 4.2%
- Recommended for production use
Example 2: Detailed Comparison
User: "Compare AutoETS vs AutoTheta performance"
Response:
## AutoETS vs AutoTheta Comparison
| Metric | AutoETS | AutoTheta | Winner |
|---------------|---------|-----------|----------|
| Mean sMAPE | 13.5% | 12.3% | Theta ✓ |
| Median sMAPE | 12.9% | 11.8% | Theta ✓ |
| Std Dev | 5.1% | 4.2% | Theta ✓ |
| Series Won | 18/50 | 32/50 | Theta ✓ |
**Key Differences**:
- Theta handles trends better (exponential smoothing with drift)
- ETS excels on strong seasonality (state space framework)
- Theta more consistent across diverse patterns
**Recommendation**: Use AutoTheta as default, fallback to AutoETS for highly seasonal data.
Example 3: Metric Interpretation
User: "What does sMAPE of 12.3% mean?"
Response:
sMAPE of 12.3% is **good performance** for forecasting:
- This means on average, predictions are off by about 12% from actual values
- For M4 Daily benchmark, anything under 15% is considered good
- This is in the "acceptable" range (10-20%)
Context: If forecasting daily sales of 100 units, a 12.3% sMAPE means your forecast will typically be within ±12 units of the true value.
For more details, see resources/METRIC_INTERPRETATION.md
Advanced Features
Benchmark Reports
If a benchmark report exists (benchmark_report_*.md), use Read tool to view formatted summaries suitable for GitHub issues or documentation.
For details on benchmark reports, see the MCP server documentation.
TimeGPT Showdown
If TimeGPT comparison data is present, incorporate it into your analysis. See resources/TIMEGPT_COMPARISON.md for detailed guidance.
Help users create GitHub issue drafts to share results with Nixtla maintainers. See resources/GITHUB_ISSUES.md for complete guidance.
When to suggest:
User wants to ask Nixtla maintainers a question
User suspects a bug or unexpected behavior
User wants to share benchmark results with the community
Output
A structured markdown analysis with an overall winner, key insights, and recommended next steps.
Optional: a short “power user” section highlighting series-level anomalies and failure cases.
Error Handling
If results files are missing:
I don't see baseline results in nixtla_baseline_m4/.
Please run the baseline command first:
Run: `/nixtla-baseline-m4 horizon=14 series_limit=50`
This will generate the metrics files I need to analyze.
If CSV is malformed:
The results file exists but appears malformed. Expected columns:
- series_id, model, sMAPE, MASE
Please re-run /nixtla-baseline-m4 to regenerate clean results.