Skip to main content

get-convex/convex-evals

SkillsMP 已收集 get-convex/convex-evals 中的 6 个 Skill。打开任一 Skill 可查看来源和详情。

最近记录的来源活动
SkillsMP 收录数据更新
已收集 skills
6
GitHub 星标
128
GitHub Forks
10

用 AI 助手安装

复制这段 Prompt,发给你正在使用的 AI 助手。

请帮我从这个仓库安装 Agent Skills:https://github.com/get-convex/convex-evals
先检查来源、SKILL.md 和配套文件,列出可安装的 Skill 让我选择。得到选择后,将选中的完整 Skill 目录安装到当前项目,并确认所需文件已就位。

这个仓库中的 skills

已展示 6 / 6 个已收集 Skill。

职业分类
软件开发工程师
描述

Design, implement, validate, and calibrate a new eval for the convex-evals suite. Use when the user wants to add a new eval, create an eval, test a new Convex concept, or expand eval coverage.

原文语言:英语

更新
职业分类
软件开发工程师
描述

Add a new AI model to the eval runner, update the manual eval workflow, push changes, and trigger baseline eval runs. Use when the user wants to add a new model, onboard a model, or mentions a new model name/link to add to the leaderboard.

原文语言:英语

更新
职业分类
软件质量保证分析师与测试员
描述

Investigate a single failing eval from the convex-evals system. Use when the user shares a visualizer URL pointing to a specific eval, asks about a specific failing eval, or references a specific eval ID.

原文语言:英语

更新
职业分类
软件质量保证分析师与测试员
描述

Analyze all failures in a convex-evals run, spawning parallel sub-agents to investigate each failure and producing a report with classifications and recommendations. Use when the user asks to analyze an entire run, review all failures in a run, or wants to…

原文语言:英语

更新
职业分类
软件质量保证分析师与测试员
描述

Empirically verify guideline changes by running before/after eval runs across multiple models and ensuring no regressions. Use when proposing or reviewing changes to runner/models/guidelines.ts, or when the user asks to validate guidelines.

原文语言:英语

更新
职业分类
软件质量保证分析师与测试员
描述

Analyze guideline ablation experiment results to determine which guideline sections are essential, marginal, or dispensable. Use when the user asks to analyze ablation results, interpret guideline compaction data, or wants to know which guidelines to keep for…

原文语言:英语

更新
已展示 6 / 6 个已收集 Skill。