원클릭으로
ai-evaluation-evals
Create AI evaluation plans with benchmarks, rubrics, and error analysis workflows.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
Create AI evaluation plans with benchmarks, rubrics, and error analysis workflows.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
Plan and run Codex-native dynamic workflows for complex tasks that benefit from explicit orchestration, goal mode, subagents or simulated work packets, approval gates, integration, verification, and reusable workflow artifacts. Use when the user invokes this skill, asks for dynamic workflows, subagents, parallel agents, swarm-like work, Goal Maker orchestration, large audits, migrations, multi-track research plus implementation, or Claude Code-style fan-out/fan-in workflows.
Direct browser control via CDP. Use when the user wants to automate, scrape, test, or interact with web pages. Connects to the user's already-running Chrome.
This skill should be used for browser automation tasks using Chrome DevTools Protocol (CDP). Triggers when users need to launch Chrome with remote debugging, navigate pages, execute JavaScript in browser context, capture screenshots, or interactively select DOM elements. No MCP server required.
Use Orca's computer-use CLI to inspect and control local desktop apps through accessibility trees, screenshots, and safe UI actions. Use when an agent needs to list desktop apps, get an app state, read visible UI, click, type, press keys, scroll, drag, set values, or perform app accessibility actions. Triggers include "computer use", "orca computer", "list apps", "get app state", "read Spotify", "read Slack", "click app", "type text", "press key", "set value", "scroll app", "drag app", and desktop app interaction tasks.
Use the `orca` CLI to drive a running Orca editor — manage Orca worktrees; create and manage scheduled automations; create, read, and run shell commands in Orca-managed terminals; and automate Orca's built-in browser (snapshot/click/fill/screenshot/tabs). Use this instead of raw `git worktree`, ad hoc shell PTYs, or Playwright whenever the task touches Orca state. Coding agents inside an Orca worktree should also use it to keep the worktree comment fresh at meaningful checkpoints. Boundary with `orchestration`: if the recipient of a terminal write is another AI agent (Claude Code, Gemini, Codex, a worker), use `orchestration` — it is the only correct way to send messages, nudges, replies, or task hand-offs to agents. orca-cli writes are for non-agent terminals (shells, build/test commands); reading or `wait`ing on any terminal — including agent terminals — stays in orca-cli.
Use Orca orchestration for structured multi-agent coordination: threaded messages, blocking ask/reply flows, task dispatch, worker_done/escalation waits, task DAGs, decision gates, coordinator loops, or decomposing work across agents. Use `orca-cli` instead for ordinary terminal control, lightweight terminal prompts, shell commands, Orca worktree management, reading or waiting on terminals, and automation of the browser embedded inside Orca. Use Computer Use for browser windows, webviews, Orca app UI, or desktop UI outside Orca's embedded browser.
| name | ai-evaluation-evals |
| description | Create AI evaluation plans with benchmarks, rubrics, and error analysis workflows. |
Category: AI & Technology
Source: https://refoundai.com/lenny-skills/s/ai-evals
AI Evaluation (Evals) | Refound AI
Lenny Skills Database SKILLS PLAYBOOKS GUESTS ABOUT SKILLS PLAYBOOKS GUESTS ABOUT AI & Technology 2 guests | 2 insights
AI Evaluation (Evals) AI evaluation (evals) is the emerging skill of systematically testing and measuring AI model performance. As models become products, evals become the product requirements document. This involves error analysis, creating rubrics, building benchmarks, and developing systematic tests - a critical bottleneck for AI labs and a new core competency for product builders.
Download Claude Skill
Read Guide
The Guide 3 key steps synthesized from 2 experts.
1 Treat evals as your product requirements In AI products, the eval suite defines what the product should do. If you can't measure it, you can't improve it. Before building features, define how you'll evaluate success. The eval is the spec - it tells the model (and your team) exactly what 'good' looks like.
Featured guest perspectives
"If the model is the product, then the eval is the product requirement document."
— Brendan Foody 2 Build systematic evaluation workflows Develop a multi-step process: start with error analysis to understand where the model fails, use open coding to categorize failure modes, create rubrics based on those categories, and build automated tests. This systematic approach replaces gut-feel assessments with rigorous measurement.
Featured guest perspectives
"Both the chief product officers of Anthropic and OpenAI shared that evals are becoming the most important new skill for product builders."
— Hamel Husain & Shreya Shankar 3 Invest in this as a core skill The heads of product at major AI labs consider evals one of the most important emerging skills. This isn't traditional QA or software testing - it's a new discipline that product builders need to develop. Treat it as a first-class competency worth investing significant time in learning.
Featured guest perspectives
"Both the chief product officers of Anthropic and OpenAI shared that evals are becoming the most important new skill for product builders."
— Hamel Husain & Shreya Shankar
✗ Common Mistakes
Treating AI testing like traditional software testingRelying on vibes instead of systematic measurementNot investing in eval infrastructure earlyEvaluating only accuracy without considering other dimensions like safety, helpfulness, or style ✓ Signs You're Doing It Well
You can quantify model performance across multiple dimensionsYou have automated eval suites that run on every model changeYour product decisions are informed by eval results, not intuitionYou can explain exactly why one model version is better than another
All Guest Perspectives
Deep dive into what all 2 guests shared about ai evaluation (evals).
Hamel Husain & Shreya Shankar 1 quote
Listen to episode →
"Both the chief product officers of Anthropic and OpenAI shared that evals are becoming the most important new skill for product builders."
View all skills from Hamel Husain & Shreya Shankar →
Brendan Foody 1 quote
Listen to episode →
"If the model is the product, then the eval is the product requirement document."
View all skills from Brendan Foody →
Install This Skill
Add this skill to Claude Code, Cursor, or any AI coding assistant that supports Agent Skills.
1 Download the skill
Download SKILL.md
2 Add to your project
Create a folder in your project root and add the skill file:
.claude/skills/ai-evals/SKILL.md 3 Start using it
Claude will automatically detect and use the skill when relevant. You can also invoke it directly:
Help me with ai evaluation (evals) Related Skills Other AI & Technology skills you might find useful. 94 guests AI Product Strategy AI strategy should focus on using algorithms to scale human expertise and judgment rather than just... View Skill → → 60 guests Building with LLMs Using LLMs for text-to-SQL can democratize data access and reduce the burden on data analysts for ad... View Skill → → 24 guests Platform Strategy Platform and ecosystem success comes from identifying 'gardening' opportunities—projects with inhere... View Skill → → 22 guests Evaluating New Technology Be skeptical of 'out-of-the-box' AI solutions for enterprises; real ROI requires a pipeline that acc... View Skill → →
AI Transformation Partner
Start Your Journey
SERVICES AI Audit AI Automation AI Training COMPANY About Case Studies Book a Call
© 2026 Refound. All rights reserved.