| name | omk-debugging |
| description | Systematic debugging: reproduce → hypothesize → verify → fix. Trigger when encountering bugs, test failures, errors, crashes, unexpected behavior, or when user says 'debug', 'not working', 'broken', 'fails', 'error', 'exception', 'why does this happen', 'investigate'. Also trigger on stack traces, error logs, or non-zero exit codes. |
Trigger Examples
- "这个测试跑不过"
- "why is this returning null?"
- "报错了,帮我看看"
- "build fails with exit code 1"
- "investigate why the hook isn't firing"
Systematic Debugging
Overview
Random fixes waste time and create new bugs. Quick patches mask underlying issues.
Core principle: ALWAYS find root cause before attempting fixes. Symptom fixes are failure.
Violating the letter of this process is violating the spirit of debugging.
The Iron Law
NO FIXES WITHOUT ROOT CAUSE INVESTIGATION FIRST
If you haven't completed Phase 1, you cannot propose fixes.
Three Iron Laws of Code Debugging
- No goto_definition = No modify — Don't change code you haven't navigated to its definition
- No find_references = No refactor — Don't refactor without knowing all usage sites
- No get_diagnostics = No claim fixed — Don't claim a fix without verifying zero new diagnostics
Investigation Document Protocol
The investigation document (docs/investigations/{date}-{topic}.md) is the persistent, cross-session record of a debugging effort. Create one at the start of every non-trivial debug session using skills/omk-debugging/investigation-template.md.
Three-Level Evidence System
| Level | Label | Meaning | Mutability |
|---|
| 🔒 L0 | Machine Facts | Command output, code analysis, API responses, experiment results | Immutable — cannot be overridden except by stronger L0 evidence |
| 👤 L1 | Human Observations | User-reported operations and phenomena | Protected — need L0 evidence + documented rationale to challenge |
| 🤖 L2 | Agent Inferences | AI deductions, hypotheses, analysis conclusions | Mutable — can be revised freely; mark old entry struck, keep original |
Section Update Rules
| Section | Update Mode | Rule |
|---|
| Status Overview | Overwrite | Refresh after every Stage completion — must reflect current state at all times |
| Evidence Table | Append-only | Never delete or edit existing entries; add new rows only |
| Investigation Tree | Overwrite | Update branch statuses (⬜→✅/❌) as investigation progresses |
| Decision Log | Append-only | Superseded decisions stay; mark status as "❌ 被 D{n} 取代" |
| Experiment Log | Append-only | Each experiment gets a unique EXP-{n} ID for cross-reference |
| Ruled Out | Append-only | Eliminated directions stay permanently to prevent re-investigation |
| Timeline | Append-only | Chronological milestones, never rewritten |
Anti-regression Rules
These rules prevent new sessions from undoing prior investigation progress:
- Overturning 🔒 L0 facts: Requires stronger L0 evidence (new command output or experiment). Must record both old and new evidence in the Evidence Table with explicit comparison.
- Challenging 👤 L1 observations: Requires L0 evidence that contradicts the observation. Must document the challenge rationale in the Timeline.
- Revising 🤖 L2 inferences: Free to revise. Mark the old inference
struck (do not delete). New inference must note "取代 I{n}".
- Decision Log: Rejected decisions must include rejection rationale. New sessions must NOT re-attempt a rejected approach unless new L0 evidence invalidates the rejection reason.
- Ruled Out: Eliminated investigation directions must NOT be re-investigated unless new L0 evidence emerges that was unavailable when the direction was ruled out.
Tool Decision Matrix
| Bug Type | Tool Sequence | Why |
|---|
| Compile/type error | get_diagnostics → goto_definition → get_hover | Diagnostics pinpoint error; definition shows context; hover reveals types |
| Wrong behavior | search_symbols → find_references → goto_definition | Find the symbol, trace all callers, read implementation |
| Unknown codebase | get_document_symbols → goto_definition → get_hover | Map file structure, navigate to definitions, understand types |
| Refactor broke something | find_references → get_diagnostics → search_symbols | Find all usage sites, check for new errors, locate related symbols |
| Test failure | get_diagnostics → search_symbols → find_references | Check compiler errors first, find test subject, trace dependencies |
When to use grep instead: Only for searching comments, string literals, config values, or non-code files.
When to Use
Use for ANY technical issue:
- Test failures
- Bugs in production
- Unexpected behavior
- Performance problems
- Build failures
- Integration issues
Use this ESPECIALLY when:
- Under time pressure (emergencies make guessing tempting)
- "Just one quick fix" seems obvious
- You've already tried multiple fixes
- Previous fix didn't work
- You don't fully understand the issue
Don't skip when:
- Issue seems simple (simple bugs have root causes too)
- You're in a hurry (rushing guarantees rework)
- Manager wants it fixed NOW (systematic is faster than thrashing)
The Four Phases
You MUST complete each phase before proceeding to the next.
Phase 1: Root Cause Investigation
BEFORE attempting ANY fix:
Step -1: Build Architectural Context (MANDATORY — do NOT skip)
Before looking at any specific code, build a map of the system around the bug. This step solves the "structural blindness" problem — LLMs process code as text and cannot see dependency graphs, call chains, or architectural boundaries without explicit navigation.
- Run
generate_codebase_overview to get the project's high-level module structure
- For the bug's core symbol(s), run
find_references to discover all callers and usage sites
- For the bug's file(s), run
get_document_symbols to understand internal structure
- Produce an Architectural Context summary:
Architectural Context:
- Module: [where this code sits in the system architecture]
- Upstream callers: [who calls the buggy code — list file:function]
- Downstream dependencies: [what the buggy code depends on]
- Cross-module impact: [which other modules could be affected by a fix]
- Hidden dependencies: [files with no semantic overlap but architectural connection]
Why this is mandatory: CodeCompass research (258 trials) showed that dependency graph navigation achieves 99.4% architectural coverage on hidden dependencies vs 76.2% without (+23.2pp). Agents only explore dependencies voluntarily 42% of the time — when skipped, performance equals the no-tool baseline. This step must be forced, not optional.
Step 0: Check Past Episodes
- Read
knowledge/episodes.md for similar past bugs
- Past mistakes often repeat — check before investigating from scratch
Step 0.5: Classify Failure Type (Failure Classification)
Before diving into investigation, classify the failure to guide your approach. Pick the most likely category:
| Category | Description | Typical Signal |
|---|
| Logic/Semantic Error | Code logic is wrong | Test fails, wrong output |
| Environment/Config Error | Environment, config, or dependency issue | Works locally, fails in CI |
| Concurrency/Timing Error | Race condition, timing dependency | Intermittent failure |
| Invalid Invocation | Tool/API call malformed or missing args | Schema error, 400 response |
| Misinterpretation of Output | Acted on wrong assumption about a return value | Downstream logic wrong |
| Intent–Plan Misalignment | Solving the wrong problem | Fix doesn't address user's actual issue |
| Plan Adherence Failure | Skipped required steps or did unplanned actions | Expected step not executed |
| Invention of New Information | Hallucinated data not in trace/tool output | References nonexistent variable/file |
| Under-specified Intent | Not enough information to proceed | Need more context from user |
Write your classification into the Diagnostic Evidence (Step 6). If unsure, pick the top 2 candidates and investigate both.
Step 1: Run get_diagnostics First
- Run
get_diagnostics on the failing file(s) to get compiler errors, warnings, hints
- This gives you the exact error locations and types — far more precise than reading logs
Step 2: Use search_symbols to Find Relevant Code
- Use
search_symbols to locate the function/class/variable involved in the error
- Follow up with
goto_definition to read the actual implementation
- Use
find_references to understand all callers and usage sites
Step 3: Read Error Messages Carefully
- Don't skip past errors or warnings
- They often contain the exact solution
- Read stack traces completely
- Note line numbers, file paths, error codes
Step 4: Reproduce Consistently
- Can you trigger it reliably?
- What are the exact steps?
- Does it happen every time?
- If not reproducible → gather more data, don't guess
Step 5: Check Recent Changes
- What changed that could cause this?
- Git diff, recent commits
- New dependencies, config changes
- Environmental differences
- If this is a regression (it used to work, now it doesn't): use the Git Bisect flow in
reference.md to pinpoint the exact commit that introduced the bug
Step 6: Gather Diagnostic Evidence
You MUST produce a Diagnostic Evidence summary before moving to Phase 2:
Diagnostic Evidence:
- failure_type: [category from Step 0.5 classification]
- get_diagnostics: [what errors/warnings were found]
- search_symbols: [what symbols were located]
- find_references: [what callers/usage sites were found]
- get_hover: [what type information was revealed]
- key_variables:
- var_name: [variable name]
expected: [expected value at this point]
actual: [actual value observed]
location: [file:line where divergence occurs]
- ...
- Root cause hypothesis: [your conclusion based on above]
Without this evidence, you cannot proceed.
Step 7: Gather Evidence in Multi-Component Systems
WHEN system has multiple components (CI → build → signing, API → service → database):
BEFORE proposing fixes, add diagnostic instrumentation:
For EACH component boundary:
- Log what data enters component
- Log what data exits component
- Verify environment/config propagation
- Check state at each layer
Run once to gather evidence showing WHERE it breaks
THEN analyze evidence to identify failing component
THEN investigate that specific component
Step 8: Trace Data Flow
WHEN error is deep in call stack:
- Use
goto_definition to navigate to the error source
- Use
find_references to trace callers backward
- Where does bad value originate?
- Keep tracing up until you find the source
- Fix at source, not at symptom
Phase 2: Pattern Analysis
Find the pattern before fixing:
-
Find Working Examples
- Locate similar working code in same codebase
- Use
search_symbols to find similar patterns
- What works that's similar to what's broken?
-
Compare Against References
- If implementing pattern, read reference implementation COMPLETELY
- Don't skim - read every line
- Understand the pattern fully before applying
-
Identify Differences
- What's different between working and broken?
- List every difference, however small
- Don't assume "that can't matter"
-
Understand Dependencies
- What other components does this need?
- What settings, config, environment?
- What assumptions does it make?
Phase 3: Hypothesis and Testing
Scientific method:
-
Form Single Hypothesis
- State clearly: "I think X is the root cause because Y"
- Write it down
- Be specific, not vague
-
Test Minimally
- Make the SMALLEST possible change to test hypothesis
- One variable at a time
- Don't fix multiple things at once
-
Verify Before Continuing
- Did it work? Yes → Phase 4
- Didn't work? Form NEW hypothesis
- DON'T add more fixes on top
-
When You Don't Know
- Say "I don't understand X"
- Don't pretend to know
- Ask for help
- Research more
Phase 4: Implementation
Fix the root cause, not the symptom:
-
Run get_diagnostics Before Fix (baseline)
- Record current diagnostics count and details
- This is your "before" snapshot
-
Create Failing Test Case
- Simplest possible reproduction
- Automated test if possible
- MUST have before fixing
-
Implement Single Fix
- Address the root cause identified
- ONE change at a time
- No "while I'm here" improvements
- No bundled refactoring
-
Run get_diagnostics After Fix (verify)
- Compare with baseline: new diagnostics must be 0
- All original diagnostics should be resolved or unchanged
- If new diagnostics appeared, your fix introduced problems — revert
-
Verify Fix
- Test passes now?
- No other tests broken?
- Issue actually resolved?
5.5. Self-Explanation (Rubber Duck Verification)
After verifying the fix works, explain in natural language:
- Root cause: What exactly was wrong and why?
- Fix logic: Why does this fix resolve the root cause?
- Side effects: Could this fix introduce new problems? Check against the Architectural Context from Step -1.
If you discover a logical contradiction while explaining → STOP, return to Phase 3 and re-verify your hypothesis. The act of explaining often reveals flawed reasoning that testing alone misses.
-
If Fix Doesn't Work
- STOP
- Count: How many fixes have you tried?
- If < 3: Return to Phase 1, re-analyze with new information
- If ≥ 3: STOP and question the architecture (step 7 below)
- DON'T attempt Fix #4 without architectural discussion
-
If 3+ Fixes Failed: Question Architecture
Pattern indicating architectural problem:
- Each fix reveals new shared state/coupling/problem in different place
- Fixes require "massive refactoring" to implement
- Each fix creates new symptoms elsewhere
STOP and question fundamentals:
- Is this pattern fundamentally sound?
- Are we "sticking with it through sheer inertia"?
- Should we refactor architecture vs. continue fixing symptoms?
Discuss with your human partner before attempting more fixes
-
Post-Fix Review
After the fix is verified and working:
- Compare the original buggy code vs the fixed code — is the fix minimal?
- Check: did the fix introduce any code smells (duplication, magic numbers, missing error handling)?
- Check: does the fix need defensive programming at other call sites? (Refer to Architectural Context from Step -1)
- Write a one-line summary for
knowledge/episodes.md if this is a new bug pattern
- Update the investigation document: set Status Overview to 🟢 Resolved, record final root cause and fix in Decision Log
Red Flags - STOP and Follow Process
If you catch yourself thinking:
- "Quick fix for now, investigate later"
- "Just try changing X and see if it works"
- "Add multiple changes, run tests"
- "Skip the test, I'll manually verify"
- "It's probably X, let me fix that"
- "I don't fully understand but this might work"
- "Pattern says X but I'll adapt it differently"
- "Here are the main problems: [lists fixes without investigation]"
- Proposing solutions before tracing data flow
- "One more fix attempt" (when already tried 2+)
- Each fix reveals new problem in different place
- Using grep to find code instead of search_symbols/find_references
ALL of these mean: STOP. Return to Phase 1.
If 3+ fixes failed: Question the architecture (see Phase 4.7)
Your Human Partner's Signals You're Doing It Wrong
Watch for these redirections:
- "Is that not happening?" - You assumed without verifying
- "Will it show us...?" - You should have added evidence gathering
- "Stop guessing" - You're proposing fixes without understanding
- "Ultrathink this" - Question fundamentals, not just symptoms
- "We're stuck?" (frustrated) - Your approach isn't working
When you see these: STOP. Return to Phase 1.
Common Rationalizations
| Excuse | Reality |
|---|
| "Issue is simple, don't need process" | Simple issues have root causes too. Process is fast for simple bugs. |
| "Emergency, no time for process" | Systematic debugging is FASTER than guess-and-check thrashing. |
| "Just try this first, then investigate" | First fix sets the pattern. Do it right from the start. |
| "I'll write test after confirming fix works" | Untested fixes don't stick. Test first proves it. |
| "Multiple fixes at once saves time" | Can't isolate what worked. Causes new bugs. |
| "Reference too long, I'll adapt the pattern" | Partial understanding guarantees bugs. Read it completely. |
| "I see the problem, let me fix it" | Seeing symptoms ≠ understanding root cause. |
| "One more fix attempt" (after 2+ failures) | 3+ failures = architectural problem. Question pattern, don't fix again. |
| "I'll just grep for it" | grep is text matching. Use LSP tools for semantic code analysis. |
Quick Reference
| Phase | Key Activities | Success Criteria |
|---|
| 1. Root Cause | get_diagnostics, search_symbols, find_references, reproduce, gather evidence | Diagnostic Evidence produced |
| 2. Pattern | Find working examples, compare | Identify differences |
| 3. Hypothesis | Form theory, test minimally | Confirmed or new hypothesis |
| 4. Implementation | Pre/post get_diagnostics, create test, fix, verify | Bug resolved, zero new diagnostics |
When Process Reveals "No Root Cause"
If systematic investigation reveals issue is truly environmental, timing-dependent, or external:
- You've completed the process
- Document what you investigated
- Implement appropriate handling (retry, timeout, error message)
- Add monitoring/logging for future investigation
But: 95% of "no root cause" cases are incomplete investigation.
Supporting Techniques
These techniques are part of systematic debugging and available in this directory:
root-cause-tracing.md - Trace bugs backward through call stack to find original trigger
defense-in-depth.md - Add validation at multiple layers after finding root cause
condition-based-waiting.md - Replace arbitrary timeouts with condition polling
Related skills:
- superpowers:test-driven-development - For creating failing test case (Phase 4, Step 2)
- superpowers:verification-before-completion - Verify fix worked before claiming success