Skip to main content

horizon-facts-eval-sweep

Stars8
Forks6
UpdatedJune 24, 2026 at 00:51

Run and analyze the horizon-facts cross-model evaluation sweep (graph-grounded vs parametric+web closed-QA), sweeping harvester × query × judge models into a 3×3×3 score tensor with an LLM-generated bias-aware report. Use when collecting fresh sweep results, re-running a subset of cells, changing the model set, or interpreting judge bias.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly