Skip to main content

assess

Assesses and rates quality 0-10 across multiple dimensions (correctness, maintainability, security, performance, testability, simplicity) with pros/cons analysis. Compares against project conventions and prior decisions from memory. Produces structured evaluation reports with actionable improvement suggestions. Use when evaluating code, designs, architectures, or comparing alternative approaches.

Datos de origen

Repositorio
yonatangross/orchestkit
Última actividad en el origen
29 de septiembre de 2026 a las 15:03
Idioma detectado de SKILL.md
inglés
Estrellas
285
Forks
35

Opciones de instalación

De forma predeterminada está seleccionado el prompt que primero revisa el origen. Puedes cambiar a un comando directo o descargar una copia local.

Revisa los archivos de origen

Lee SKILL.md y los archivos complementarios que muestra SkillsMP antes de decidir si quieres instalarlo.

Explorador de archivos
26 archivos

Mostrando SKILL.md

SKILL.md
Instrucciones de origen · Vista previa de solo lectura
name
assess
license
MIT
compatibility
Claude Code 2.1.277+. Requires memory MCP server.
description
Assesses and rates quality 0-10 across multiple dimensions (correctness, maintainability, security, performance, testability, simplicity) with pros/cons analysis. Compares against project conventions and prior decisions from memory. Produces structured evaluation reports with actionable improvement suggestions. Use when evaluating code, designs, architectures, or comparing alternative approaches.
context
fork
background
false
user-invocable
true
allowed-tools
AskUserQuestion Read Write Grep Glob Agent Workflow TaskCreate TaskUpdate TaskList ToolSearch mcp__memory__search_nodes Bash
skills
["code-review-playbook","quality-gates","architecture-decision-record","memory","chain-patterns"]
argument-hint
[code-path-or-topic] [--render=markdown|json-render|both] [--effort=low|medium|high|xhigh]
effort
high
model
sonnet
hooks
{"PreToolUse":[{"matcher":"Read","command":"${CLAUDE_PLUGIN_ROOT}/hooks/bin/run-hook.mjs skill/assessment-baseline-loader","once":true}]}
metadata
{"category":"document-asset-creation","mcp-server":"memory","version":"1.9.0","author":"OrchestKit","complexity":"high","tags":"assessment, evaluation, quality, comparison, pros-cons, rating"}
# Assess Host-neutral workflow. Invoke by skill name (`assess`). Claude Code slash routing, YAML hook loaders, and `.claude/chain` live in `references/claude-code.md`. Comprehensive assessment skill for answering "is this good?" with structured evaluation, scoring, and actionable recommendations. ## 🎯 Quick Start ```bash assess backend/app/services/auth.py assess our caching strategy assess --model=opus the current database schema assess frontend/src/components/Dashboard ``` ### Effort levels (CC 2.1.111+ adds `xhigh`) | Effort | Behavior | |---|---| | `low` / `medium` | Subset of dimensions, faster turnaround | | `high` (default) | All six dimensions with pros/cons | | `xhigh` | All six dimensions + one additional assessor pass focused on uncertainty/caveats; emits `confidence` per dimension | > `xhigh` silently falls back to `high` on a model that does not implement it: no error, no log line. `doctor` Category 14 reports this, and only when it can positively prove the active model lacks the tier. --- ## Argument Resolution ### Step 0: resolve a conversational reference first `$ARGUMENTS` is often not a path. For a bare pronoun or deictic (`them`, `this`, `that`, `these`, `they`, `same`, `the above`, `the last one`, `what we just did`) or an empty target after flags are stripped, the subject is in the conversation. Read back for the NEAREST concrete one (a file just discussed, a diff or PR just opened, a component just investigated) and announce the resolution in one line, so a wrong guess costs a correction rather than a turn: *"Reading 'them' as the 3 pretool guards we just probed; say otherwise and I'll switch."* **Refusing is the bug, not the safe option.** Asking "what does this refer to?" when the previous turn named the subject burns a round-trip re-deriving what is already on screen. Measured 2026-08-28: the operator sent `assess them throguhly` one message after "bug in orchestkit hooks", mid-investigation of `pretool/bash/dangerous-command-blocker`, and this skill replied that "them" had "no antecedent anywhere in this conversation". It had two. Ask only when the conversation is genuinely empty (a fresh session opening with a bare pronoun). Every other case: resolve and announce. > Not unique to this skill: `verify`, `cover`, `fix-issue`, `review-pr` and `implement` all > read `$ARGUMENTS` as a literal path or topic, and no skill mentions resolving a reference. > Tracked separately; this one fixes its own door. ```python TARGET = "$ARGUMENTS" # Full argument string, e.g., "backend/app/services/auth.py" # $ARGUMENTS[0] is the first token (CC 2.1.59 indexed access) # Model override detection (CC 2.1.72) MODEL_OVERRIDE = None for token in "$ARGUMENTS".split(): if token.startswith("--model="): MODEL_OVERRIDE = token.split("=", 1)[1] # "opus", "sonnet", "haiku", "fable" TARGET = TARGET.replace(token, "").strip() ``` Pass `MODEL_OVERRIDE` to all Agent() calls via `model=MODEL_OVERRIDE` when set. Accepts symbolic names (`opus`, `sonnet`, `haiku`, `fable` on harnesses whose Agent tool lists it; note fable is premium API spend after 2026-07-12) or full IDs (`claude-opus-5-5`) per CC 2.1.74. > **Switching to Opus via `/model` (CC 2.1.144+):** `/model` now changes the model for the current session only, so picking Opus for an assess run no longer persists past it. Press `d` in the picker only to set a default for new sessions. ### Effort detection (CC 2.1.120+) `$CLAUDE_EFFORT` is the primary signal. CC 2.1.120 sets this env var from `/effort` or the model picker. `--effort=` token in `$ARGUMENTS` is the explicit override fallback (also covers older CC). ```python # Read env first (CC 2.1.120+), then check explicit override EFFORT = os.environ.get("CLAUDE_EFFORT") # "low" | "medium" | "high" | "xhigh" | None for token in "$ARGUMENTS".split(): if token.startswith("--effort="): EFFORT = token.split("=", 1)[1] # explicit override wins TARGET = TARGET.replace(token, "").strip() EFFORT = EFFORT or "high" # default when CC < 2.1.120 and no flag ``` Use `EFFORT` to gate dimension count, agent count, and the optional `xhigh` uncertainty pass — see "Effort levels" table above. On CC < 2.1.120 the env var is unset; the explicit `--effort=` override is the only path. `doctor` Category 14 reports a provably unsupported `xhigh` request. --- ## STEP -1: MCP Probe + Resume Check > Load: `Read("../chain-patterns/references/mcp-detection.md")` ```python # 1. Probe MCP servers (once at skill start) # memory is alwaysLoad in .mcp.json (CC 2.1.121+, #1541) — probe below kept as fallback for older CC: ToolSearch(query="select:mcp__memory__search_nodes") # 2. Store capabilities Write(".claude/chain/capabilities.json", { "memory": probe_memory.found, "skill": "assess", "timestamp": now() }) # 3. Check for resume state = Read(".claude/chain/state.json") # may not exist if state.skill == "assess" and state.status == "in_progress": last_handoff = Read(f".claude/chain/{state.last_handoff}") ``` ### Phase Handoffs | Phase | Handoff File | Contents | |-------|-------------|----------| | 0 | `00-intent.json` | Dimensions, target, mode | | 1 | `01-baseline.json` | Initial codebase scan results | | 2 | `02-evaluation.json` | Per-dimension scores + evidence | | 3 | `03-report.json` | Final report, grade, recommendations | --- ## STEP 0: Verify User Intent with AskUserQuestion **BEFORE creating tasks**, clarify assessment dimensions: ```python AskUserQuestion( questions=[{ "question": "What dimensions to assess?", "header": "Dimensions", "options": [ {"label": "Full assessment (Recommended)", "description": "All dimensions: quality, maintainability, security, performance"}, {"label": "Code quality only", "description": "Readability, complexity, best practices"}, {"label": "Security focus", "description": "Vulnerabilities, attack surface, compliance"}, {"label": "Quick score", "description": "Just give me a 0-10 score with brief notes"} ], "multiSelect": false }] ) ``` **Based on answer, adjust workflow** (passed to Phase 2 as `focus`): - **Full assessment** (`full`): All 7 phases, parallel agents, dimension subset scaled by effort - **Code quality only** (`quality`): Skip security and performance phases - **Security focus** (`security`): Prioritize security-auditor agent - **Quick score** (`quick`): Single pass, brief output --- ## STEP 0b: Select Orchestration Mode Load details: `Read("references/orchestration-mode.md")` for env var check logic, Agent Teams vs Agent Tool comparison, and mode selection rules. --- ## 🚨 Task Management (CC 2.1.16) ```python # 1. Create main task IMMEDIATELY TaskCreate( subject="Assess: {target}", description="Comprehensive evaluation with quality scores and recommendations", activeForm="Assessing {target}" ) # 2. Create subtasks for each assessment phase TaskCreate(subject="Understand target and gather context", activeForm="Understanding target") # id=2 TaskCreate(subject="Discover scope and build file list", activeForm="Discovering scope") # id=3 TaskCreate(subject="Rate quality across 6 dimensions", activeForm="Rating quality") # id=4 TaskCreate(subject="Analyze pros and cons", activeForm="Analyzing pros/cons") # id=5 TaskCreate(subject="Compare alternatives", activeForm="Comparing alternatives") # id=6 TaskCreate(subject="Generate improvement suggestions", activeForm="Generating suggestions") # id=7 TaskCreate(subject="Compile assessment report", activeForm="Compiling report") # id=8 # 3. Set dependencies for sequential phases TaskUpdate(taskId="3", addBlockedBy=["2"]) # Scope needs target understanding TaskUpdate(taskId="4", addBlockedBy=["3"]) # Rating needs scoped file list TaskUpdate(taskId="5", addBlockedBy=["4"]) # Pros/cons needs quality scores TaskUpdate(taskId="6", addBlockedBy=["4"]) # Alternatives need quality scores TaskUpdate(taskId="7", addBlockedBy=["5", "6"]) # Suggestions need analysis TaskUpdate(taskId="8", addBlockedBy=["7"]) # Report needs suggestions # 4. Update status as you progress TaskUpdate(taskId="2", status="in_progress") # When starting TaskUpdate(taskId="2", status="completed") # When done — repeat for each subtask ``` --- ## 🔄 Workflow Overview | Phase | Activities | Output | |-------|------------|--------| | **1. Target Understanding** | Read code/design, identify scope | Context summary | | **1.5. Scope Discovery** | Build bounded file list | Scoped file list | | **2. Quality Rating** | 6-dimension scoring (0-10) | Scores with reasoning | | **3. Pros/Cons Analysis** | Strengths and weaknesses | Balanced evaluation | | **4. Alternative Comparison** | Score alternatives | Comparison matrix | | **5. Improvement Suggestions** | Actionable recommendations | Prioritized list | | **6. Effort Estimation** | Time and complexity estimates | Effort breakdown | | **7. Assessment Report** | Compile findings | Final report | --- ## Phase 1: Target Understanding Identify what's being assessed and gather context. `TARGET` here is the value Step 0 already resolved, which is not necessarily what the user typed. ```python # PARALLEL - Gather context Read(file_path=TARGET) # only when TARGET is a path Grep(pattern=TARGET, output_mode="files_with_matches") # topic or symbol mcp__memory__search_nodes(query=TARGET) # past decisions ``` `Read` failing is NOT a reason to stop. A target resolved from the conversation is usually a subject rather than a filename ("the three pretool guards", "today's hook fixes"), so the Read misses and the Grep plus the conversation carry the context. Treat a failed Read as "this is a topic, not a path" and continue to Phase 1.5, which discovers the real file list anyway. --- ## Phase 1.5: Scope Discovery Load `Read("references/scope-discovery.md")` for the full file discovery, limit application (MAX 30 files), and sampling priority logic. **Always include the scoped file list** in every agent prompt. ### Progressive Output (CC 2.1.76) Output results **incrementally** as each evaluation phase completes: | After Phase | Show User | |-------------|-----------| | 1. Target Understanding | Scope summary, file list, context | | 1.5. Scope Discovery | Bounded file list (max 30 files) | | 2. Quality Rating | Each dimension's score as the evaluating agent returns | | 3. Pros/Cons | Balanced evaluation summary | The Phase 2 workflow returns once, so show every dimension's score from its result and lead with `priorityConcerns` (any dimension below 4/10) as a concern needing user attention. Per-agent streaming applies only on the Agent tool fallback. --- ## Phase 2: Quality Rating (6 Dimensions) Rate each dimension 0-10 with weighted composite score. Load `Read("../quality-gates/references/unified-scoring-framework.md")` for dimensions, weights, grade interpretation, and per-dimension criteria. Load `Read("references/quality-model.md")` for assess-specific overrides. Do NOT hand-roll the assessors. Run the executor, which owns Phases 2 and 2.5: ```python result = Workflow( scriptPath="${CLAUDE_SKILL_DIR}/workflows/assess-fanout.js", args={"target": TARGET, "effort": EFFORT, "focus": FOCUS, # FOCUS from STEP 0 "mode": "comparison" if COMPARING else "default", # quality-model.md "domain": "frontend" or "backend", # picks the performance engineer "scopeFiles": SCOPE_FILES, # Phase 1.5 list "projectContext": MEMORY_CONTEXT, # Phase 1 memory search "rubric": Read("rubric.json"), "modelOverride": MODEL_OVERRIDE, "feature": FEATURE}) Write(".claude/chain/02-evaluation.json", result) ``` **The script owns the mechanics, not the prose.** It picks the assessors from `focus` and `effort` (security first), gives each a score schema that demands `file:line` evidence, sends every decision-bearing score to blind refuters (Phase 2.5), and computes the weighted composite, grade, rubric verdict and blockers. A score with no `file:line` evidence counts as unscored, a repeated dimension keeps only its first entry, and a selected dimension with a `min_blocker` that nobody scored is a blocker. It returns `composite`, `grade`, `verdict`, `blockers` (producer basis), `postRefutation`, `chainVerdict`, `chainVerdictIfConfirmed`, `revisions`, `confirmationNeeded`, `manualReview`, `advisory`, `priorityConcerns`, `quickWins`, `unscored`, `rejectedDimensions`, `unassessed`, `dimensions`, `ledger` and `reasons`. **It never asks and never writes**: those stay in this shell. Fallback (no Workflow tool, `ORCHESTKIT_FORCE_TASK_TOOL=1`, or the cross-model lane below): `Read("references/agent-spawn-definitions.md")` for Agent tool and Agent Teams spawns, then run Phase 2.5 by hand. **Composite Score:** Weighted average of the scored dimensions (see quality-model.md). --- ## Phase 2.5: Adversarial Refutation (effort-gated) The assessor that scores a dimension is also its only judge, a self-preferential bias. A separate **blind refuter** forms its own band for each decision-bearing score. **Effort gate:** `low`/`medium` skip it; `high` runs up to 4 single advisory refuters (no auto-swing); `xhigh` runs a 3-refuter majority that revises to the near band edge. On the Workflow path the script already ran it; the shell finishes it: 1. Write the returned `ledger` to `.claude/chain/02b-refutation.json` (engine section 10).
Ver en GitHub
Este SKILL.md es muy grande, por eso SkillsMP muestra aqui solo la primera seccion. Ver en GitHub