| name | result-analyzer |
| description | Compare reproduced results against paper-reported values. Generate Markdown/JSON/Beamer reports. |
Result-Analyzer Skill
Purpose
The final piece of the reproduction pipeline. After running experiments, this skill compares your results against the paper's reported metrics, figures, and training curves — then generates a structured report.
When to Use
- After
code-reproducer has finished training on the remote GPU
- User says "compare my results" or "how close are we to the paper?"
- When writing a reproduction report or paper
- To decide if reproduction was successful
Architecture
Inputs Comparators Outputs
─────────────────── ────────────── ────────────────
paper-parser JSON ──┐
├──→ table_extractor ───┐
paper metrics ──┘ │
├──→ metric_comparator ──→ Markdown Report
repro metrics ────────────────────────────┘ ──→ JSON Report
──→ Beamer Data
repro train.csv ────→ curve_comparator ──→ comparison plots ──→ latex-paper-skills
──→ beamer-skill PPT
repro images ────→ image_comparator ──→ SSIM/PSNR/FID
paper figures ────┘ ──→ side-by-side figures
Quick Start
Simple Metric Comparison
python result-analyzer/analyze.py \
--paper-metrics '{"accuracy": 95.3, "FID": 3.17}' \
--repro-metrics '{"accuracy": 94.8, "FID": 3.45}' \
--title "DDPM Reproduction" \
-o report/
Full Pipeline (paper JSON + training log + images)
python result-analyzer/analyze.py \
--paper-json workspace/<paper>/paper_content.json \
--repro-log workspace/<paper>/train_log.csv \
--repro-images workspace/<paper>/generated/ \
--paper-figures workspace/<paper>/figures/ \
--fid \
--beamer \
-o workspace/<paper>/report/
From Paper-Parser JSON Only
python result-analyzer/analyze.py \
--paper-json paper_content.json \
--repro-metrics '{"accuracy": 94.8}' \
--method-name "Ours" \
-o report/
Comparators
1. Metric Comparator (Core)
Compares reproduced values against paper values with tolerance-based judgment.
Built-in Knowledge Base: 50+ metrics with direction info:
higher_better: accuracy, BLEU, PSNR, SSIM, IS, mAP, AUC, F1...
lower_better: FID, loss, perplexity, WER, MAE, RMSE, LPIPS...
Judgment Logic:
| Status | Condition |
|---|
| ✅ PASS | Within tolerance (default: ±1 abs, ±5% rel) |
| 🟢 PASS | Better than paper! |
| ⚠️ WARN | Within 2× tolerance |
| ❌ FAIL | Outside tolerance |
2. Curve Comparator
- Loads CSV/JSON training logs
- Compares final values against paper
- Pearson correlation between curves
- Convergence speed analysis
- Generates matplotlib comparison plots
3. Image Comparator
- SSIM (Structural Similarity) — via scikit-image
- PSNR (Peak Signal-to-Noise Ratio)
- FID (Fréchet Inception Distance) — optional, requires
pytorch-fid
- Side-by-side comparison figures
4. Table Extractor
- Reads paper-parser JSON output → finds result tables
- Auto-detects "Ours" / "Proposed" / last row
- Handles
±, bold markers, % signs
- Free-text metric extraction (e.g., "We achieve 95.3% accuracy")
Output Formats
Markdown Report (reproduction_report.md)
# Reproduction Report
## 🟢 Overall: PASS
| Status | Metric | Paper | Reproduced | Diff | Note |
|--------|--------|-------|-----------|------|------|
| ✅ PASS | Accuracy | 95.3 | 94.8 | -0.5 ↓ | Within tolerance |
| ⚠️ WARN | FID | 3.17 | 3.45 | +0.28 ↑ | Close but slightly off |
JSON Report (reproduction_report.json)
Structured data for downstream consumption:
latex-paper-skills → results-backfill skill
beamer-skill → reproduction PPT
Beamer Data (beamer_report_data.json)
Slide-by-slide data for generating a Beamer PPT:
- Title slide (paper name + overall status)
- Metric comparison table
- Training curves figure
- Sample comparison figure
- Conclusion slide
Integration with Other Skills
→ beamer-skill (Generate PPT)
python result-analyzer/analyze.py ... --beamer -o report/
python beamer-skill/generate_beamer_report.py report/beamer_report_data.json -o report/reproduction_slides.tex
python beamer-skill/generate_beamer_report.py report/beamer_report_data.json -o slides.tex --compile
→ latex-paper-skills (Write LaTeX Paper via latex_bridge.py)
python result-analyzer/analyze.py ... -o report/
python result-analyzer/latex_bridge.py report/reproduction_report.json -o paper/results/
← paper-parser (Input)
paper-parser outputs paper_content.json with tables
→ result-analyzer extracts paper metrics from tables
→ compares against reproduced values
← code-reproducer (Input)
code-reproducer outputs train_log.csv + generated images
→ result-analyzer loads CSV for curve comparison
→ compares images via SSIM/PSNR
Dependencies
pip install -r result-analyzer/requirements.txt
📚 Reference URLs (for agent self-help)
| Topic | URL |
|---|
| scikit-image SSIM docs | https://scikit-image.org/docs/stable/api/skimage.metrics.html |
| scikit-image PSNR docs | https://scikit-image.org/docs/stable/api/skimage.metrics.html |
| pytorch-fid | https://github.com/mseitzer/pytorch-fid |
| matplotlib savefig | https://matplotlib.org/stable/api/_as_gen/matplotlib.pyplot.savefig.html |
| pandas read_csv | https://pandas.pydata.org/docs/reference/api/pandas.read_csv.html |
| latex-paper-skills | https://github.com/yunshenwuchuxun/latex-paper-skills |
| results-backfill skill | https://github.com/yunshenwuchuxun/latex-paper-skills/tree/main/.codex/skills/results-backfill |
| ML metrics overview | https://paperswithcode.com/task/image-generation |
Limitations
- SSIM is sensitive to image alignment — ensure consistent cropping
- FID requires ≥2048 images per directory for reliable scores
- Table extraction depends on paper-parser JSON quality
- Free-text metric extraction uses regex patterns — may miss complex phrasing