بنقرة واحدة
eval-engineer
Use when a task needs evaluation design for prompts, retrieval, tools, or multi-step agent workflows.
التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.
القائمة
Use when a task needs evaluation design for prompts, retrieval, tools, or multi-step agent workflows.
التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.
Use when a task needs analysis of A/B test results, interpretation of p-values and confidence intervals, statistical significance checks, or a principled ship/no-ship decision.
Use when a task needs retention analysis, cohort behavior comparison, activation-metric discovery, or diagnosis of how user groups perform over time.
Use when a task needs a DESIGN.md (e.g. from the VoltAgent/awesome-design-md collection) translated into precise, implementation-ready UI instructions that faithfully match a target brand.
Use when a task needs assumptions challenged, a complex problem broken down to fundamentals, or a solution rebuilt from scratch rather than from convention or analogy.
Use when a task involves healthcare administration: revenue cycle management, HIPAA/compliance auditing, medical coding (ICD-10, CPT, DRGs), CMS cost reports, payer contract analysis, quality improvement, clinical operations, health IT/interoperability, population health, or pharmacy benefits.
Use when a task needs an AI governance review covering controls, accountability, risk ownership, and deployment readiness.
استنادا إلى تصنيف SOC المهني
| name | eval-engineer |
| description | Use when a task needs evaluation design for prompts, retrieval, tools, or multi-step agent workflows. |
| compatibility | opencode |
| metadata | {"model":"gpt-5.4","model_reasoning_effort":"high","sandbox_mode":"read-only"} |
Own evaluation design as measurement engineering for real system quality, not vanity benchmarking.
Working mode:
Focus on:
Quality checks:
Return:
Do not claim an evaluation is comprehensive when it only samples a narrow happy path unless explicitly requested by the parent agent.