Skip to main content

maragudk/evals-skills

SkillsMP 已收集 maragudk/evals-skills 中的 4 个 Skill。打开任一 Skill 可查看来源和详情。

最近记录的来源活动
SkillsMP 收录数据更新
已收集 skills
4
GitHub 星标
12
GitHub Forks
0

这个仓库中的 skills

已展示 4 / 4 个已收集 Skill。

职业分类
软件开发工程师
描述

Generate a custom trace annotation web app for open coding during LLM error analysis. Use when the user wants to review LLM traces, annotate failures with freeform comments, and do first-pass qualitative labeling (open coding). Also use when the user mentions…

原文语言:英语

更新
职业分类
数据科学家
描述

Build a structured taxonomy of failure modes from open-coded trace annotations. Use this skill whenever the user has freeform annotations from reviewing LLM traces and wants to cluster them into a coherent, non-overlapping set of binary failure categories…

原文语言:英语

更新
职业分类
软件开发工程师
描述

Use this skill when crafting, reviewing, or improving prompts for LLM pipelines — including task prompts, system prompts, and LLM-as-Judge prompts. Triggers include: requests to write or refine a prompt, diagnose why an LLM produces inconsistent or incorrect…

原文语言:英语

更新
职业分类
软件质量保证分析师与测试员
描述

Build, validate, and deploy LLM-as-Judge evaluators for automated quality assessment of LLM pipeline outputs. Use this skill whenever the user wants to: create an automated evaluator for subjective or nuanced failure modes, write a judge prompt for Pass/Fail…

原文语言:英语

更新
已展示 4 / 4 个已收集 Skill。