Design, implement, validate, and calibrate a new eval for the convex-evals suite. Use when the user wants to add a new eval, create an eval, test a new Convex concept, or expand eval coverage.
原文语言:英语
菜单
SkillsMP 已收集 get-convex/convex-evals 中的 6 个 Skill。打开任一 Skill 可查看来源和详情。
复制这段 Prompt,发给你正在使用的 AI 助手。
请帮我从这个仓库安装 Agent Skills:https://github.com/get-convex/convex-evals
先检查来源、SKILL.md 和配套文件,列出可安装的 Skill 让我选择。得到选择后,将选中的完整 Skill 目录安装到当前项目,并确认所需文件已就位。已展示 6 / 6 个已收集 Skill。
Design, implement, validate, and calibrate a new eval for the convex-evals suite. Use when the user wants to add a new eval, create an eval, test a new Convex concept, or expand eval coverage.
原文语言:英语
Add a new AI model to the eval runner, update the manual eval workflow, push changes, and trigger baseline eval runs. Use when the user wants to add a new model, onboard a model, or mentions a new model name/link to add to the leaderboard.
原文语言:英语
Investigate a single failing eval from the convex-evals system. Use when the user shares a visualizer URL pointing to a specific eval, asks about a specific failing eval, or references a specific eval ID.
原文语言:英语
Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations. Use when the user asks to analyze an entire run, review all failures in a run, or wants to…
原文语言:英语
Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions. Use when proposing or reviewing changes to runner/models/guidelines.ts, or when the user asks to validate guidelines.
原文语言:英语
Analyze guideline ablation experiment results to determine which guideline sections are essential, marginal, or dispensable. Use when the user asks to analyze ablation results, interpret guideline compaction data, or wants to know which guidelines to keep for…
原文语言:英语