用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/mofa-org/mofa-skills --skill mofa-eval命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
基于 SOC 职业分类
正在显示 SKILL.md
| name | mofa-eval |
| description | An LLM-as-a-Judge agent evaluation skill built in Rust for MoFA IDE Testing. |
mofa-eval (Agent Testing Skill)This skill brings a robust agentic testing platform to MoFA. Utilizing OpenAI (via async-openai) as an LLM Judge, this skill can autonomously grade the outputs of other MoFA agents against standard rubrics and track regressions in a SQLite database.
run_id into a SQLite database.compare_runs tool to explicitly notify if Agent Prompts have regressed performance between run iterations.Trigger a single evaluation:
echo '{"run_id":"test-run-01", "expected":"The capital is Paris.", "actual":"Paris is the capital city of France."}' | mofa-eval evaluate_response
Get a summary of a test run:
echo '{"run_id":"test-run-01"}' | mofa-eval score_summary