- name
- assess
- license
- MIT
- compatibility
- Claude Code 2.1.277+. Requires memory MCP server.
- description
- Assesses and rates quality 0-10 across multiple dimensions (correctness, maintainability, security, performance, testability, simplicity) with pros/cons analysis. Compares against project conventions and prior decisions from memory. Produces structured evaluation reports with actionable improvement suggestions. Use when evaluating code, designs, architectures, or comparing alternative approaches.
- context
- fork
- background
- false
- version
- 1.8.0
- author
- OrchestKit
- tags
- ["assessment","evaluation","quality","comparison","pros-cons","rating"]
- user-invocable
- true
- allowed-tools
- ["AskUserQuestion","Read","Write","Grep","Glob","Agent","TaskCreate","TaskUpdate","TaskList","ToolSearch","mcp__memory__search_nodes","Bash"]
- skills
- ["code-review-playbook","quality-gates","architecture-decision-record","memory","chain-patterns"]
- argument-hint
- [code-path-or-topic] [--render=markdown|json-render|both] [--effort=low|medium|high|xhigh]
- complexity
- high
- persuasion-type
- guidance
- effort
- high
- model
- sonnet
- hooks
- {"PreToolUse":[{"matcher":"Read","command":"${CLAUDE_PLUGIN_ROOT}/hooks/bin/run-hook.mjs skill/assessment-baseline-loader","once":true}]}
- metadata
- {"category":"document-asset-creation","mcp-server":"memory"}
# Assess
Host-neutral workflow. Invoke by skill name (`assess`). Claude Code slash routing, YAML hook loaders, and `.claude/chain` live in `references/claude-code.md`.
Comprehensive assessment skill for answering "is this good?" with structured evaluation, scoring, and actionable recommendations.
## 🎯 Quick Start
```bash
assess backend/app/services/auth.py
assess our caching strategy
assess --model=opus the current database schema
assess frontend/src/components/Dashboard
```
### Effort levels (CC 2.1.111+ adds `xhigh`)
| Effort | Behavior |
|---|---|
| `low` / `medium` | Subset of dimensions, faster turnaround |
| `high` (default) | All six dimensions with pros/cons |
| `xhigh` | All six dimensions + one additional assessor pass focused on uncertainty/caveats; emits `confidence` per dimension |
> `xhigh` silently falls back to `high` on a model that does not implement it: no error, no log line. `doctor` Category 14 reports this, and only when it can positively prove the active model lacks the tier.
---
## Argument Resolution
### Step 0: resolve a conversational reference first
`$ARGUMENTS` is often not a path. For a bare pronoun or deictic (`them`, `this`, `that`,
`these`, `they`, `same`, `the above`, `the last one`, `what we just did`) or an empty target
after flags are stripped, the subject is in the conversation. Read back for the NEAREST
concrete one (a file just discussed, a diff or PR just opened, a component just investigated)
and announce the resolution in one line, so a wrong guess costs a correction rather than a
turn: *"Reading 'them' as the 3 pretool guards we just probed; say otherwise and I'll switch."*
**Refusing is the bug, not the safe option.** Asking "what does this refer to?" when the
previous turn named the subject burns a round-trip re-deriving what is already on screen.
Measured 2026-08-28: the operator sent `assess them throguhly` one message after "bug in
orchestkit hooks", mid-investigation of `pretool/bash/dangerous-command-blocker`, and this
skill replied that "them" had "no antecedent anywhere in this conversation". It had two.
Ask only when the conversation is genuinely empty (a fresh session opening with a bare
pronoun). Every other case: resolve and announce.
> Not unique to this skill: `verify`, `cover`, `fix-issue`, `review-pr` and `implement` all
> read `$ARGUMENTS` as a literal path or topic, and no skill mentions resolving a reference.
> Tracked separately; this one fixes its own door.
```python
TARGET = "$ARGUMENTS" # Full argument string, e.g., "backend/app/services/auth.py"
# $ARGUMENTS[0] is the first token (CC 2.1.59 indexed access)
# Model override detection (CC 2.1.72)
MODEL_OVERRIDE = None
for token in "$ARGUMENTS".split():
if token.startswith("--model="):
MODEL_OVERRIDE = token.split("=", 1)[1] # "opus", "sonnet", "haiku", "fable"
TARGET = TARGET.replace(token, "").strip()
```
Pass `MODEL_OVERRIDE` to all Agent() calls via `model=MODEL_OVERRIDE` when set. Accepts symbolic names (`opus`, `sonnet`, `haiku`, `fable` on harnesses whose Agent tool lists it; note fable is premium API spend after 2026-07-12) or full IDs (`claude-opus-4-8`) per CC 2.1.74.
> **Switching to Opus via `/model` (CC 2.1.144+):** `/model` now changes the model for the current session only, so picking Opus for an assess run no longer persists past it. Press `d` in the picker only to set a default for new sessions.
### Effort detection (CC 2.1.120+)
`$CLAUDE_EFFORT` is the primary signal. CC 2.1.120 sets this env var from `/effort` or the model picker. `--effort=` token in `$ARGUMENTS` is the explicit override fallback (also covers older CC).
```python
# Read env first (CC 2.1.120+), then check explicit override
EFFORT = os.environ.get("CLAUDE_EFFORT") # "low" | "medium" | "high" | "xhigh" | None
for token in "$ARGUMENTS".split():
if token.startswith("--effort="):
EFFORT = token.split("=", 1)[1] # explicit override wins
TARGET = TARGET.replace(token, "").strip()
EFFORT = EFFORT or "high" # default when CC < 2.1.120 and no flag
```
Use `EFFORT` to gate dimension count, agent count, and the optional `xhigh` uncertainty pass — see "Effort levels" table above. On CC < 2.1.120 the env var is unset; the explicit `--effort=` override is the only path. `doctor` Category 14 reports a provably unsupported `xhigh` request.
---
## STEP -1: MCP Probe + Resume Check
> Load: `Read("../chain-patterns/references/mcp-detection.md")`
```python
# 1. Probe MCP servers (once at skill start)
# memory is alwaysLoad in .mcp.json (CC 2.1.121+, #1541) — probe below kept as fallback for older CC:
ToolSearch(query="select:mcp__memory__search_nodes")
# 2. Store capabilities
Write(".claude/chain/capabilities.json", {
"memory": probe_memory.found,
"skill": "assess",
"timestamp": now()
})
# 3. Check for resume
state = Read(".claude/chain/state.json") # may not exist
if state.skill == "assess" and state.status == "in_progress":
last_handoff = Read(f".claude/chain/{state.last_handoff}")
```
### Phase Handoffs
| Phase | Handoff File | Contents |
|-------|-------------|----------|
| 0 | `00-intent.json` | Dimensions, target, mode |
| 1 | `01-baseline.json` | Initial codebase scan results |
| 2 | `02-evaluation.json` | Per-dimension scores + evidence |
| 3 | `03-report.json` | Final report, grade, recommendations |
---
## STEP 0: Verify User Intent with AskUserQuestion
**BEFORE creating tasks**, clarify assessment dimensions:
```python
AskUserQuestion(
questions=[{
"question": "What dimensions to assess?",
"header": "Dimensions",
"options": [
{"label": "Full assessment (Recommended)", "description": "All dimensions: quality, maintainability, security, performance"},
{"label": "Code quality only", "description": "Readability, complexity, best practices"},
{"label": "Security focus", "description": "Vulnerabilities, attack surface, compliance"},
{"label": "Quick score", "description": "Just give me a 0-10 score with brief notes"}
],
"multiSelect": false
}]
)
```
**Based on answer, adjust workflow:**
- **Full assessment**: All 7 phases, parallel agents
- **Code quality only**: Skip security and performance phases
- **Security focus**: Prioritize security-auditor agent
- **Quick score**: Single pass, brief output
---
## STEP 0b: Select Orchestration Mode
Load details: `Read("references/orchestration-mode.md")` for env var check logic, Agent Teams vs Task Tool comparison, and mode selection rules.
---
## 🚨 Task Management (CC 2.1.16)
```python
# 1. Create main task IMMEDIATELY
TaskCreate(
subject="Assess: {target}",
description="Comprehensive evaluation with quality scores and recommendations",
activeForm="Assessing {target}"
)
# 2. Create subtasks for each assessment phase
TaskCreate(subject="Understand target and gather context", activeForm="Understanding target") # id=2
TaskCreate(subject="Discover scope and build file list", activeForm="Discovering scope") # id=3
TaskCreate(subject="Rate quality across 6 dimensions", activeForm="Rating quality") # id=4
TaskCreate(subject="Analyze pros and cons", activeForm="Analyzing pros/cons") # id=5
TaskCreate(subject="Compare alternatives", activeForm="Comparing alternatives") # id=6
TaskCreate(subject="Generate improvement suggestions", activeForm="Generating suggestions") # id=7
TaskCreate(subject="Compile assessment report", activeForm="Compiling report") # id=8
# 3. Set dependencies for sequential phases
TaskUpdate(taskId="3", addBlockedBy=["2"]) # Scope needs target understanding
TaskUpdate(taskId="4", addBlockedBy=["3"]) # Rating needs scoped file list
TaskUpdate(taskId="5", addBlockedBy=["4"]) # Pros/cons needs quality scores
TaskUpdate(taskId="6", addBlockedBy=["4"]) # Alternatives need quality scores
TaskUpdate(taskId="7", addBlockedBy=["5", "6"]) # Suggestions need analysis
TaskUpdate(taskId="8", addBlockedBy=["7"]) # Report needs suggestions
# 4. Update status as you progress
TaskUpdate(taskId="2", status="in_progress") # When starting
TaskUpdate(taskId="2", status="completed") # When done — repeat for each subtask
```
---
## What This Skill Answers
| Question | How It's Answered |
|----------|-------------------|
| "Is this good?" | Quality score 0-10 with reasoning |
| "What are the trade-offs?" | Structured pros/cons list |
| "Should we change this?" | Improvement suggestions with effort |
| "What are the alternatives?" | Comparison with scores |
| "Where should we focus?" | Prioritized recommendations |
---
## 🔄 Workflow Overview
| Phase | Activities | Output |
|-------|------------|--------|
| **1. Target Understanding** | Read code/design, identify scope | Context summary |
| **1.5. Scope Discovery** | Build bounded file list | Scoped file list |
| **2. Quality Rating** | 6-dimension scoring (0-10) | Scores with reasoning |
| **3. Pros/Cons Analysis** | Strengths and weaknesses | Balanced evaluation |
| **4. Alternative Comparison** | Score alternatives | Comparison matrix |
| **5. Improvement Suggestions** | Actionable recommendations | Prioritized list |
| **6. Effort Estimation** | Time and complexity estimates | Effort breakdown |
| **7. Assessment Report** | Compile findings | Final report |
---
## Phase 1: Target Understanding
Identify what's being assessed and gather context. `TARGET` here is the value Step 0 already
resolved, which is not necessarily what the user typed.
```python
# PARALLEL - Gather context
Read(file_path=TARGET) # only when TARGET is a path
Grep(pattern=TARGET, output_mode="files_with_matches") # topic or symbol
mcp__memory__search_nodes(query=TARGET) # past decisions
```
`Read` failing is NOT a reason to stop. A target resolved from the conversation is usually a
subject rather than a filename ("the three pretool guards", "today's hook fixes"), so the Read
misses and the Grep plus the conversation carry the context. Treat a failed Read as "this is a
topic, not a path" and continue to Phase 1.5, which discovers the real file list anyway.
---
## Phase 1.5: Scope Discovery
Load `Read("references/scope-discovery.md")` for the full file discovery, limit application (MAX 30 files), and sampling priority logic. **Always include the scoped file list** in every agent prompt.
### Progressive Output (CC 2.1.76)
Output results **incrementally** as each evaluation phase completes:
| After Phase | Show User |
|-------------|-----------|
| 1. Target Understanding | Scope summary, file list, context |
| 1.5. Scope Discovery | Bounded file list (max 30 files) |
| 2. Quality Rating | Each dimension's score as the evaluating agent returns |
| 3. Pros/Cons | Balanced evaluation summary |
For Phase 2 parallel agents, show each dimension's score **as soon as the evaluating agent returns** — don't wait for all 4 agents. If any dimension scores below 4/10, flag it immediately as a priority concern requiring user attention.
---
## Phase 2: Quality Rating (6 Dimensions)
Rate each dimension 0-10 with weighted composite score. Load `Read("../quality-gates/references/unified-scoring-framework.md")` for dimensions, weights, grade interpretation, and per-dimension criteria. Load `Read("references/quality-model.md")` for assess-specific overrides.
Load `Read("references/agent-spawn-definitions.md")` for Task Tool mode spawn patterns and Agent Teams alternative.
**Composite Score:** Weighted average of all 6 dimensions (see quality-model.md).
---
## Phase 2.5: Adversarial Refutation (effort-gated)
The assessor that scores a dimension is also its only judge — self-preferential bias.
A separate **blind refuter** verifies decision-bearing scores before they reach the
composite. **Effort gate:** `low`/`medium` skip this phase entirely; `high` runs up-to-4
single refuters (advisory, no auto-swing); `xhigh` runs 3-refuter majority with auto-revise.
Load the protocol + assess bindings: `Read("references/adversarial-refutation.md")`
(which loads the shared engine `../../shared/rules/adversarial-refutation.md`).
Producer findings must first pass the evidence-replay gate before entering any score or verdict: `Read("../../shared/rules/evidence-replay.md")`.
### Cross-model refuter (optional, provenance-labeled, cost-gated)
When `ORK_ALT_MODEL_CMD` is configured and effort is `high`/`xhigh`, one quorum slot per high-weight or boundary-adjacent dimension score can route to a non-Claude model (Codex/GPT) for diverse failure modes. Off by default; substitutes one same-model slot, stamps `refuter_model` for provenance, cannot silently raise the grade (engine §7), owns no credentials/egress (shells out via `ORK_ALT_MODEL_CMD`, matches the egress guard #2533), and degrades to same-model on an absent command. Shares the review-pr operational doc: `Read("../review-pr/references/cross-model-refuter.md")`.
Runs after Phase 2 returns, before the composite/grade and Phases 3-7. Refuters are ALWAYS
isolated `Agent(...)` Task spawns (never team members, even in Agent Teams mode) fed only the
serialized claim — no producer score, identity, or prose. Revised scores recompute the
composite; the refutation ledger (`02b-refutation.json`) records survived/killed/downgraded
so wrong scores are auditable. Keep the producer-basis score AND a labeled post-refutation
score — refutation never silently raises the grade.
---
## Phases 3-7: Analysis, Comparison & Report
View on GitHub