Skip to main content
hamelsmu
ملف منشئ GitHub

hamelsmu

عرض على مستوى المستودعات لـ ٢٠ skills مجمعة عبر ٦ مستودعات GitHub.

skills مجمعة
٢٠
مستودعات
٦
محدث
١٢ يوليو ٢٠٢٦
مستكشف المستودعات

المستودعات و skills الممثلة

validate-evaluator
مطوّرو البرمجيات

Calibrate an LLM judge against human labels using data splits, TPR/TNR, and bias correction. Use after writing a judge prompt (write-judge-prompt) when you need to verify alignment before trusting its outputs. Do NOT use for code-based evaluators (those are…

١٠ يونيو ٢٠٢٦
build-review-interface
محللو ضمان جودة البرمجيات والمختبرون

Build a custom browser-based annotation interface tailored to your data for reviewing LLM traces and collecting structured feedback. Use when you need to build an annotation tool, review traces, or collect human labels.

٣ مارس ٢٠٢٦
error-analysis
محللو ضمان جودة البرمجيات والمختبرون

Help the user systematically identify and categorize failure modes in an LLM pipeline by reading traces. Use when starting a new eval project, after significant pipeline changes (new features, model switches, prompt rewrites), when production metrics drop, or…

٣ مارس ٢٠٢٦
eval-audit
محللو ضمان جودة البرمجيات والمختبرون

Audit an LLM eval pipeline and surface problems: missing error analysis, unvalidated judges, vanity metrics, etc. Use when inheriting an eval system, when unsure whether evals are trustworthy, or as a starting point when no eval infrastructure exists. Do NOT…

٣ مارس ٢٠٢٦
evaluate-rag
محللو ضمان جودة البرمجيات والمختبرون

Guides evaluation of RAG pipeline retrieval and generation quality. Use when evaluating a retrieval-augmented generation system, measuring retrieval quality, assessing generation faithfulness or relevance, generating synthetic QA pairs for retrieval testing,…

٣ مارس ٢٠٢٦
generate-synthetic-data
محللو ضمان جودة البرمجيات والمختبرون

Create diverse synthetic test inputs for LLM pipeline evaluation using dimension-based tuple generation. Use when bootstrapping an eval dataset, when real user data is sparse, or when stress-testing specific failure hypotheses. Do NOT use when you already…

٣ مارس ٢٠٢٦
write-judge-prompt
محللو ضمان جودة البرمجيات والمختبرون

Design LLM-as-Judge evaluators for subjective criteria that code-based checks cannot handle. Use when a failure mode requires interpretation (tone, faithfulness, relevance, completeness). Do NOT use when the failure mode can be checked with code (regex,…

٣ مارس ٢٠٢٦
عرض ٦ من أصل ٦ مستودعات
تم تحميل كل المستودعات