Skip to main content

scix-eval-and-evidence

스타6
포크1
업데이트2026년 7월 11일 02:49

How retrieval and extraction quality is measured and what counts as evidence in SciX: the gold sets (50q curated, 1200q recall gold, claim extraction, lexical stress), the fusion-calibration sweep and its honest verdict (dense_only < bm25_only), nDCG@10 / Recall@K / MRR, Wilson 95% CIs, the OAuth persona/UMBRELA judges, the claim_blame gold-set plan (bead 6ajy), and the reporting rules (read-only harness, lane provenance, per-bucket numbers, null results stated plainly). Load when running or interpreting an eval, adding a gold set, judging relevance, choosing an acceptance threshold, or writing up a result. NOT for RRF fusion internals (scix-retrieval-architecture), Qdrant mechanics (scix-vector-serving-qdrant), CI/pytest (scix-build-test-ci), whether a change needs an ADR (scix-change-control), or query_log telemetry (scix-db-safety-and-telemetry).

설치

Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.

SKILL.md
readonly