Skip to main content
hamelsmu
Profil créateur GitHub

hamelsmu

Vue par dépôt de 20 skills collectés dans 6 dépôts GitHub.

skills collectés
20
dépôts
6
mis à jour
12 juil. 2026
explorateur de dépôts

Dépôts et skills représentatifs

validate-evaluator
Développeurs de logiciels

Calibrate an LLM judge against human labels using data splits, TPR/TNR, and bias correction. Use after writing a judge prompt (write-judge-prompt) when you need to verify alignment before trusting its outputs. Do NOT use for code-based evaluators (those are…

10 juin 2026
build-review-interface
Analystes en assurance qualité des logiciels et testeurs

Build a custom browser-based annotation interface tailored to your data for reviewing LLM traces and collecting structured feedback. Use when you need to build an annotation tool, review traces, or collect human labels.

3 mars 2026
error-analysis
Analystes en assurance qualité des logiciels et testeurs

Help the user systematically identify and categorize failure modes in an LLM pipeline by reading traces. Use when starting a new eval project, after significant pipeline changes (new features, model switches, prompt rewrites), when production metrics drop, or…

3 mars 2026
eval-audit
Analystes en assurance qualité des logiciels et testeurs

Audit an LLM eval pipeline and surface problems: missing error analysis, unvalidated judges, vanity metrics, etc. Use when inheriting an eval system, when unsure whether evals are trustworthy, or as a starting point when no eval infrastructure exists. Do NOT…

3 mars 2026
evaluate-rag
Analystes en assurance qualité des logiciels et testeurs

Guides evaluation of RAG pipeline retrieval and generation quality. Use when evaluating a retrieval-augmented generation system, measuring retrieval quality, assessing generation faithfulness or relevance, generating synthetic QA pairs for retrieval testing,…

3 mars 2026
generate-synthetic-data
Analystes en assurance qualité des logiciels et testeurs

Create diverse synthetic test inputs for LLM pipeline evaluation using dimension-based tuple generation. Use when bootstrapping an eval dataset, when real user data is sparse, or when stress-testing specific failure hypotheses. Do NOT use when you already…

3 mars 2026
write-judge-prompt
Analystes en assurance qualité des logiciels et testeurs

Design LLM-as-Judge evaluators for subjective criteria that code-based checks cannot handle. Use when a failure mode requires interpretation (tone, faithfulness, relevance, completeness). Do NOT use when the failure mode can be checked with code (regex,…

3 mars 2026
6 dépôts affichés sur 6
Tous les dépôts sont affichés