Audit a gold manifest row end-to-end — verify answer correctness, grounding pages, and free-text quality. Reads from manifest, searches corpus independently, and updates the row in-place (review_status → audited or flagged). Use this skill for post-labeling…
Score a pipeline's free-text answer against the gold answer and gold pages on the 5 competition criteria (correctness, completeness, grounding, confidence calibration, clarity). Use this skill to evaluate answer quality for free_text questions before…
Label a deterministic legal RAG question (boolean, number, date, name, names) by independently searching the DIFC legal corpus and producing a gold manifest row with the correct answer and minimal grounding pages. Use this skill whenever you need to create or…
Label a free-text legal RAG question by searching the DIFC legal corpus and producing a gold manifest row with a grounded answer and minimal citation pages. Use this skill whenever you need to create gold labels for free_text questions — warm-up calibration,…
Score a free-text gold answer against the 5 LLM judge criteria (correctness, completeness, grounding, confidence calibration, clarity/relevance) and suggest improvements. Use this skill after labeling free-text questions to catch quality issues before they…
Audit grounding pages for a labeled gold manifest row. Uses adversarial methodology to find missing pages, unnecessary pages, and off-by-one errors. Works for both deterministic and free-text questions. Use this skill after labeling to verify that the cited…