| name | evaluation-frameworks |
| description | Domain knowledge on evaluation frameworks and benchmarks built by Jiachen (Amber) Liu — IaC-Eval, ML.ENERGY, EXP-Bench, Humanity's Last Exam, Sci-Reasoning, and the pattern of evaluation-as-research. Use this when you need to understand Amber's benchmarking work, position evaluation infrastructure as a research contribution, or discuss her unique approach to AI research.
|
| version | 1 |
| tags | ["evaluation","benchmarks","research-methodology","AI-safety","energy-efficiency","infrastructure-as-code","MLSys"] |
Evaluation Frameworks — Jiachen Liu's Research Perspective
When to Use
- You need to discuss Amber's evaluation framework work with context and depth
- You want to understand what makes evaluation infrastructure a publishable research contribution
- You need sourced details on IaC-Eval, ML.ENERGY, EXP-Bench, HLE, Sci-Reasoning, or FedScale
- You are explaining the "benchmarking instinct" — why some researchers build the measuring stick, not just the thing being measured
1. The Benchmarking Instinct: Why Evaluation Matters
Jiachen (Amber) Liu has a distinctive research pattern: she doesn't just propose methods — she builds the community infrastructure for measuring progress. This is not accidental. Her PhD dissertation (2025, University of Michigan, advised by Prof. Mosharaf Chowdhury) is titled User-Centric Machine Learning Systems, reflecting a systems-level perspective where the evaluation layer is as important as the algorithm layer.
[Source: https://amberljc.github.io/]
The pattern spans her entire publication record:
| # | Title | Year | Venue | Citations | Role |
|---|
| 1 | FedScale: Benchmarking FL at Scale | 2022 | ICML 2022 (Spotlight) | 417 | Co-author |
| 2 | IaC-Eval: A Code Generation Benchmark for IaC Programs | 2024 | NeurIPS 2024 | 42 | Co-first author |
| 3 | The ML.ENERGY Benchmark | 2025 | NeurIPS 2025 Spotlight | — | Co-author |
| 4 | EXP-Bench: Can AI Conduct AI Research Experiments? | 2025 | ICLR 2026 | 18 | Co-first author |
| 5 | Humanity's Last Exam | 2025 | Nature 2026 | 460 | Co-author |
| 6 | Sci-Reasoning: A Dataset Decoding AI Innovation Patterns | 2026 | arXiv 2026 | 2 | First author |
[Source: https://amberljc.github.io/; citation counts from brief]
Six papers, five of which are primarily evaluation/benchmark contributions. This is a deliberate research strategy: build the measuring stick, and the field comes to you.
2. IaC-Eval: Code Generation for Infrastructure
What It Is
The first benchmark for evaluating LLMs on Infrastructure-as-Code (IaC) generation — producing cloud infrastructure definitions (AWS CloudFormation, Terraform) from natural language.
Paper: IaC-Eval: A Code Generation Benchmark for Cloud Infrastructure-as-Code Programs
Venue: NeurIPS 2024, Datasets and Benchmarks Track [Source: https://openreview.net/forum?id=7TCK0aBL1C]
Authors: Patrick Tser Jern Kon*, Jiachen Liu*, Yiming Qiu, Weijun Fan, Ting He, Lei Lin, Haoran Zhang, Owen M. Park, George Sajan Elengikal, Yuxin Kang, Ang Chen, Mosharaf Chowdhury, Myungjin Lee, Xinyu Wang (* Equal contribution)
Code: https://github.com/autoiac-project/iac-eval
Dataset: https://huggingface.co/datasets/autoiac-project/iac-eval
Why IaC Needs a Dedicated Benchmark
IaC is distinct from general-purpose code generation because:
- The "code" defines infrastructure (networks, compute, storage), not algorithms
- Correctness requires understanding cloud service APIs, resource dependencies, and deployment constraints
- Errors are costly — a wrong IaC config can expose security holes or waste thousands of dollars
- IaC-specific training data is scarce compared to Python/JavaScript
[Source: OpenReview abstract]
Methodology
- 458 human-curated scenarios across popular AWS services at difficulty levels 1–6
- Each scenario: natural language IaC problem description + infrastructure intent specification
- Two-phase evaluation pipeline that validates functional correctness without requiring actual cloud deployment
- Metric: pass@1 accuracy
[Source: OpenReview abstract; https://github.com/autoiac-project/iac-eval]
Key Results
- GPT-4 (best): 19.36% pass@1 on IaC-Eval vs. 86.6% on EvalPlus (Python)
- This 67-point gap proves IaC is a distinct, harder challenge — general code ability does not transfer
- Follow-up: Multi-IaC-Eval (arXiv:2509.05303, Aug 2025) extended to multi-IaC settings
[Source: OpenReview abstract]
Amber's Role
Co-first author with Patrick Tser Jern Kon. This was the first paper from the Michigan/Santa Barbara evaluation-frameworks group that would go on to produce EXP-Bench.
3. ML.ENERGY: Energy Benchmarking for LLMs
What It Is
A benchmark suite, measurement toolkit, and leaderboard that positions energy consumption as a first-class evaluation metric for generative AI — alongside accuracy and latency.
Paper: The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization
Venue: NeurIPS 2025, Datasets and Benchmarks Track (Spotlight) [Source: https://neurips.cc/virtual/2025/loc/san-diego/poster/121781]
arXiv: https://arxiv.org/abs/2505.06371
Authors: Jae-Won Chung, Jiachen Liu, Jeff Ma, Ruofan Wu, Oh Jun Kweon, Yuxuan Xia, Zhiyu Wu, Mosharaf Chowdhury
Initiative: https://ml.energy/ (SymbioticLab, University of Michigan)
Tool (Zeus): https://github.com/ml-energy/zeus
Leaderboard: https://ml.energy/leaderboard/
Origin & Evolution
Methodology
- Zeus library measures GPU energy consumption programmatically during inference
- Benchmark runs configurations independently on designated hardware
- Measures: energy (Wh), latency, throughput, model quality — reported together for energy-quality tradeoffs
- Provides automated energy optimization recommendations based on measurements
- Hardware: NVIDIA GPUs, expanding to CPU, DRAM, AMD GPU, Apple Silicon, Jetson (v3.0)
[Source: arXiv 2505.06371; https://github.com/ml-energy/zeus]
Amber's Role
Co-author. The initiative is led by her advisor Prof. Mosharaf Chowdhury. Amber contributed to the benchmark infrastructure and the broader ML.ENERGY ecosystem.
Why It Matters
Energy is the metric AI doesn't want to measure. As models scale, inference energy becomes a bottleneck — not just environmentally, but economically. ML.ENERGY makes this visible and actionable.
4. EXP-Bench: Can AI Do AI Research?
What It Is
A benchmark that evaluates AI agents on complete, end-to-end AI research experiments — not just code generation or Q&A, but the full cycle: hypothesis → experiment design → implementation → execution → analysis.
Paper: EXP-Bench: Can AI Conduct AI Research Experiments?
Venue: ICLR 2026 (Poster) [Source: https://openreview.net/forum?id=KjgyAm383Z]
arXiv: https://arxiv.org/abs/2505.24785 (May 30, 2025)
Authors: Patrick Tser Jern Kon*, Jiachen Liu*, Xinyi Zhu, Qiuyi Ding, Jingjia Peng, Jiarong Xing, Yibo Huang, Yiming Qiu, Jayanth Srinivasa, Myungjin Lee, Mosharaf Chowdhury, Matei Zaharia, Ang Chen (* Equal contribution)
Methodology
- 461 research tasks extracted from 51 top-tier AI papers (NeurIPS, ICML, ICLR)
- Semi-autonomous pipeline extracts experiment specifications from papers + associated code
- Given: research question + incomplete starter code
- Agent must: formulate hypotheses → design experiments → implement → execute → analyze
- Multi-dimensional scoring: design correctness, implementation correctness, execution success, analysis quality
[Source: arXiv abstract; OpenReview abstract]
Key Results
- Best agents (OpenHands, IterativeAgent): 20–35% on individual aspects (design, implementation)
- Complete executable experiment success rate: 0.5%
- This gap reveals the massive distance between "can partially do research" and "can do research end-to-end"
[Source: arXiv abstract; OpenReview abstract]
Connection to Curie
EXP-Bench is the evaluation companion to Curie — the first AI-agent framework for rigorous automated scientific experimentation. Curie is the method; EXP-Bench is the measuring stick.
Curie paper: Curie: Toward Rigorous and Automated Scientific Experimentation with AI Agents
Authors: Patrick Tser Jern Kon*, Jiachen Liu*, Qiuyi Ding, Yiming Qiu, Zhenning Yang, Yibo Huang, Jayanth Srinivasa, Myungjin Lee, Mosharaf Chowdhury, Ang Chen (* Equal contribution)
Venue: arXiv 2025
[Source: https://amberljc.github.io/]
Amber's Role
Co-first author on both EXP-Bench and Curie. This is her most direct "method + evaluation" paired contribution.
5. Humanity's Last Exam and Sci-Reasoning
Humanity's Last Exam (HLE)
Paper: Humanity's Last Exam
Venue: Nature 2026 [Source: https://amberljc.github.io/; https://agi.safe.ai/]
arXiv: https://arxiv.org/abs/2501.14249 (January 2025)
Website: https://agi.safe.ai/
Organizers: Center for AI Safety (CAIS) + Scale AI
Lead authors: Long Phan*, Alice Gatti*, Ziwen Han*, Nathaniel Li*, and others; senior: Dan Hendrycks (CAIS), Summer Yue, Alexandr Wang (Scale AI)
Scale: 2,500 questions from ~1,000 expert contributors across 500+ institutions, 50 countries
Prize pool: $500,000
[Source: https://agi.safe.ai/; Wikipedia; arXiv 2501.14249]
Key design choices:
- Questions require graduate-level expertise or highly specific domain knowledge
- Filtering: first tested against AI models (if models failed or did worse than random, questions went to human expert review)
- 14% multi-modal (text + images); 24% multiple-choice, 76% short-answer
- Private holdout set to detect benchmark overfitting
- Community feedback bug bounty program for continuous quality improvement
[Source: Wikipedia; https://agi.safe.ai/]
Amber's Role: Co-author in the large-scale collaborative benchmark. This is a different authorship model from her other papers — a community-scale effort rather than a small-team contribution.
Results (April 2026): Best model (Gemini 3.1 Pro Preview) at 46.44%. Early results (Jan 2025): GPT-4o at 2.7%.
[Source: Wikipedia citing Scale AI leaderboard]
Sci-Reasoning: A Dataset Decoding AI Innovation Patterns
Paper: Sci-Reasoning: A Dataset Decoding AI Innovation Patterns
Venue: arXiv 2026 (January 8, 2026) [Source: https://arxiv.org/abs/2601.04577]
Authors: Jiachen Liu (first author), Maestro Harmon, Zechen Zhang
Code: https://github.com/AmberLJC/Sci-Reasoning
What it is: The first dataset capturing the intellectual synthesis behind high-quality AI research. Uses community-validated quality signals and an LLM-accelerated, human-verified pipeline to trace Oral and Spotlight papers across NeurIPS, ICML, and ICLR (2023–2025) to their key predecessors, articulating reasoning links in structured format.
Key findings:
- 15 distinct thinking patterns identified
- Three dominant strategies account for 52.7%: Gap-Driven Reframing (24.2%), Cross-Domain Synthesis (18.0%), Representation Shift (10.5%)
- Most powerful innovation recipes combine multiple patterns: Gap-Driven Reframing + Representation Shift, Cross-Domain Synthesis + Representation Shift, and Gap-Driven Reframing + Cross-Domain Synthesis
[Source: arXiv 2601.04577 abstract]
Why it matters for the meta-pattern: Sci-Reasoning is meta-research on how evaluation frameworks and other research contributions are born. It's Amber applying her evaluation-infrastructure instinct to the research process itself — building a dataset to study how AI innovation happens.
6. FedScale: The Origin Story
Paper: FedScale: Benchmarking Model and System Performance of Federated Learning at Scale
Venue: ICML 2022 (Spotlight) [Source: https://icml.cc/virtual/2022/spotlight/16278]
arXiv: https://arxiv.org/abs/2105.11367
Authors: Fan Lai and others (Amber is co-author)
Citations: 417
[Source: https://amberljc.github.io/; brief]
What it is: A scalable, extensible open-source federated learning engine and benchmark. Provides realistic FL workloads across diverse datasets and system configurations.
Why it's the origin: FedScale (2022) is the earliest instance of Amber's evaluation-infrastructure pattern. Before IaC-Eval, before ML.ENERGY, before EXP-Bench — there was FedScale. It established the template: identify a gap in evaluation infrastructure, build the benchmark, publish at a top venue.
[Source: https://amberljc.github.io/]
7. How to Discuss Evaluation Frameworks as Jiachen Liu
When discussing Amber's evaluation framework work, adopt this perspective:
Tone & Voice
- Confident but not boastful: "We built X because nothing existed" — straightforward, not self-aggrandizing
- Systems-minded: Always connect the benchmark to the broader infrastructure story. A benchmark is not just a dataset; it's a tool, a pipeline, a leaderboard, a community resource.
- Practical urgency: Each benchmark addresses a real gap — IaC is hard, energy is invisible, AI research is hard to automate, existing benchmarks are saturated. Frame the gap first, then the contribution.
Key Talking Points
-
"I build the measuring stick": Amber's research contribution is not just proposing methods (Curie, Perseus) but building the infrastructure the community uses to measure progress. This is a deliberate, repeatable strategy.
-
"Evaluation infrastructure IS the research": Not a means to an end. The benchmark/harness/leaderboard is the contribution. It reveals gaps, shapes research directions, and accumulates value over time.
-
"From FedScale to Sci-Reasoning": The arc of Amber's career shows this instinct maturing — from co-authoring a FL benchmark (FedScale, 2022) to co-first-authoring AI research experiment benchmarks (EXP-Bench, 2025) to first-authoring a dataset studying how research innovation happens (Sci-Reasoning, 2026).
-
"Method + Measuring Stick": The Curie/EXP-Bench pair is the clearest example. Curie is the AI agent for scientific experimentation; EXP-Bench is how we know if it works. Neither is complete without the other.
-
"Community-scale vs small-team": Amber works across scales — from intimate co-first-author papers (IaC-Eval, EXP-Bench, Curie) to massive collaborations (HLE, 1,000+ contributors). Both are valuable; both are deliberate.
What NOT to Say
- Don't present benchmarks as "just datasets" — they are research infrastructure with toolchains, pipelines, and community processes
- Don't minimize co-author roles — even on HLE (a massive collaboration), being part of a Nature 2026 benchmark paper is significant
- Don't separate the evaluation work from the systems work — Amber's systems expertise (LLM serving, training infrastructure) is what makes her benchmarks technically rigorous
8. Meta-Pattern: Evaluation Infrastructure as Research Contribution
The Pattern
Across Amber's work and the broader AI field, a clear pattern: the evaluation framework itself is the primary research contribution, not a secondary tool.
| Traditional | Eval-as-contribution |
|---|
| "We propose method X. We built benchmark Y to evaluate X. Contribution = X." | "We built benchmark Z. Z reveals gaps. Contribution = Z." |
What Makes It Publishable
- Novel measurement axis: Something not being systematically measured (IaC capability, energy, end-to-end research ability, expert-level knowledge)
- Reveals gaps: Shows existing methods are much worse than thought (GPT-4 at 19% on IaC; 0.5% end-to-end research success)
- Infrastructure, not just data: Tools + pipelines + leaderboards + optimization, not just a dataset
- Living & extensible: Designed to evolve (ML.ENERGY v1→v3, HELM as "living benchmark")
- Community adoption: Others use it in their own work (lm-evaluation-harness in every LLM paper)
- Fills a recognized gap: Before IaC-Eval, no IaC benchmark existed despite IaC's importance
The Benchmark → Leaderboard → Eval Framework Evolution
- Benchmark (static dataset): Fixed tasks + metrics. Example: MMLU.
- Leaderboard (dynamic ranking): Benchmark + public ranking + competitive dynamics. Example: Chatbot Arena.
- Eval Framework (full infrastructure): Measurement tools + benchmarks + leaderboards + optimization + extensibility. Example: ML.ENERGY full suite, HELM.
Each step adds contribution value. A benchmark is useful but saturates. A leaderboard creates incentives. A full eval framework shapes the field.
Other Examples
| Framework | Venue | Pattern |
|---|
| HELM | NeurIPS 2022 | Multi-metric "living benchmark for transparency" |
| lm-evaluation-harness | EleutherAI | Open-source harness used in virtually every LLM paper |
| Chatbot Arena | NeurIPS 2023 | Elo-based human preference ranking — the methodology is the contribution |
| GEM | ACL 2021 | 55 researchers, 44 institutions — "living benchmark" for NLG |
| BigBench | NeurIPS 2023 | 200+ tasks, 400+ authors — the benchmark IS the paper |
| MedHELM | Stanford 2025 | Medical evaluation with 37 evals across 121 tasks |
The "Picks and Shovels" Strategy
Just as the real money in a gold rush is in selling picks and shovels, the most durable research contributions in AI may be in building the infrastructure everyone else uses to evaluate their methods. Amber has internalized this principle.
References
Amber's Papers
Other Referenced Frameworks
Amber's Profile
- Homepage: https://amberljc.github.io/
- PhD: University of Michigan, Computer Science, advised by Prof. Mosharaf Chowdhury
- Focus: AI-Native Research Infrastructure, MLSys, LLM Systems, AI Agents
- Current vision: Orchestra Vibe Research Platform — "AI-native platform for Vibe Research to accelerate scientific research from idea to publication"
[Source: https://amberljc.github.io/]