Skip to main content

systematic-debugging

4-phase root cause debugging: understand bugs before fixing.

الانتقال إلى التثبيت

معلومات المصدر

المستودع
Ntizar/NtizarBrainMasterMind
آخر نشاط في المصدر
٢٩ يونيو ٢٠٢٦ في ١٨:٣٧
لغة SKILL.md المكتشفة
لغات متعددة
النجوم
٢
التفرعات
٠

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.

مستكشف الملفات
6 ملفات

عرض SKILL.md

SKILL.md
تعليمات المصدر · معاينة للقراءة فقط
name
systematic-debugging
description
4-phase root cause debugging: understand bugs before fixing.
version
1.2.0
author
Hermes Agent (adapted from obra/superpowers)
license
MIT
platforms
["linux","macos","windows"]
metadata
{"hermes":{"tags":["debugging","troubleshooting","problem-solving","root-cause","investigation"],"related_skills":["test-driven-development","writing-plans","subagent-driven-development"]}}
# Systematic Debugging ## Overview Random fixes waste time and create new bugs. Quick patches mask underlying issues. **Core principle:** ALWAYS find root cause before attempting fixes. Symptom fixes are failure. **Violating the letter of this process is violating the spirit of debugging.** ## The Iron Law ``` NO FIXES WITHOUT ROOT CAUSE INVESTIGATION FIRST ``` If you haven't completed Phase 1, you cannot propose fixes. ## When to Use Use for ANY technical issue: - Test failures - Bugs in production - Unexpected behavior - Performance problems - Build failures - Integration issues **Use this ESPECIALLY when:** - Under time pressure (emergencies make guessing tempting) - "Just one quick fix" seems obvious - You've already tried multiple fixes - Previous fix didn't work - You don't fully understand the issue **Don't skip when:** - Issue seems simple (simple bugs have root causes too) - You're in a hurry (rushing guarantees rework) - Someone wants it fixed NOW (systematic is faster than thrashing) ## The Four Phases You MUST complete each phase before proceeding to the next. --- ## Phase 0: Pre-Implementation Audit **BEFORE any changes, run automated checks to validate the current state.** This prevents introducing new bugs and identifies real issues vs. perceived ones. ### 1. Syntax Validation Check all modified files for syntax errors before making changes: ```bash # Node.js node -c server.js # Python python3 -m py_compile module.py # JS/HTML (basic brace check) python3 -c "c=open('dashboard.html').read(); print('OK' if c.count('{')==c.count('}') else 'MISMATCH')" ``` ### 2. Variable Scope Analysis For JavaScript/TypeScript files, check for: - Undefined variables used in expressions (e.g., `perfil && perfil.altura_cm` when `perfil` is never declared in scope) - `const charts` vs `var charts = window.charts = {}` (must use `var` for Chart.js instances) - Mismatched function parameter names (e.g., function takes `objHidr` but caller passes `objetivoHidr`) - Missing variable declarations in closures/lambdas - **JavaScript class constructor crash — missing method call:** If a constructor calls `this.something()` but that method is NOT defined in the class, the constructor throws `TypeError: this.something is not a function` and **the entire class instantiation fails silently** (no method after the failing call executes). Symptom: app loads HTML but nothing works — no render loop, no event handlers, overlays stuck. Diagnosis: check that every `this.method()` called in the constructor is actually defined as a class method. - **Global function vs class method mismatch:** HTML `onclick="foo()"` calls a **global** function `foo()`, NOT `app.foo()`. If `foo` only exists as a class method, the call fails with `ReferenceError: foo is not defined`. Symptom: button clicks do nothing. Fix: either add a global wrapper (`function foo() { if (app) app.foo(); }`) or change the onclick to `app.foo()`. ```python # Check for undefined variable usage in a function scope import re with open('file.js', 'r') as f: content = f.read() # Find function scope and check for used-but-undefined vars # Look for patterns like `var && var.` where `var` is not declared with `var` in scope ``` ### 3. Field Name Consistency Check that frontend field names match backend expectations: ```bash # Find all references to a field name grep -rn "objetivo_peso_kg" dashboard.html server.js grep -rn "peso_objetivo_kg" dashboard.html server.js # If frontend uses X and backend expects Y → BUG ``` ### 4. API Endpoint Verification Verify all API endpoints called from frontend exist in backend: ```bash # Extract all API calls from frontend grep -oP "api\(['\"](/api/\S+)" dashboard.html | sort -u # Compare with backend routes grep -oP "app\.(get|post|put|delete)\(['\"](/api/\S+)" server.js | sort -u ``` ### 5. Chart.js Memory Management Check that charts are properly destroyed before recreation: ```bash # Count chart destroy calls vs. new Chart calls grep -c "\.destroy()" dashboard.html grep -c "new Chart(" dashboard.html # Should be equal or destroy > new (for safety) ``` ### 6. XSS Risk Assessment For files using `innerHTML`, check that user input is escaped: ```bash # Count innerHTML assignments grep -c "innerHTML =" dashboard.html # Verify escape functions exist and are used grep -c "escapeHtml" dashboard.html grep -c "formatMarkdown" dashboard.html ``` ### 7. CDN MIME Type Verification **WHEN frontend loads third-party libraries via CDN and they don't work:** Check the actual MIME type served by the CDN — this is the #1 silent failure in vanilla JS projects: ```bash curl -sI "https://cdn.jsdelivr.net/npm/docx@9.7.1/dist/index.umd.cjs" | grep -i content-type # application/node → BLOCKED by browser curl -sI "https://unpkg.com/docx@9.7.1/dist/index.umd.cjs" | grep -i content-type # text/javascript → OK ``` **Common pitfall:** jsdelivr serves `.cjs` files as `application/node` which modern browsers reject silently. Switch to unpkg which serves `text/javascript`. The symptom: network shows HTTP 200, but `window.libName` is `undefined` with NO console error. **Check list:** - [ ] Verify CDN URL returns the expected file (not a 404/redirect) - [ ] Verify Content-Type is `text/javascript` or `application/javascript` - [ ] Add defensive check: `if (typeof window.libName === 'undefined') { alert('No cargada'); return; }` - [ ] Check that SPA fallback (serving index.html for everything) isn't intercepting CDN paths ### 8. DOM Element Existence Against JS References **WHEN UI elements were recently removed from HTML:** Check that ALL JavaScript references to removed DOM elements are also removed: ```bash # Find all JS references to a removed element grep -rn "shpModosContainer" js/ grep -rn "btnDescargarSHP" js/ grep -rn "getElementById.*shp" js/ ``` **Pattern:** Remove `<div id="shpModosContainer">` from HTML but JS still references `document.getElementById('shpModosContainer')` or adds event listeners to it → JS silently fails. **Fix:** 1. Remove the dead JS code entirely (preferred) 2. Guard with `if (document.getElementById('elID')) { ... }` as fallback ### Phase 0 Completion Checklist - [ ] Syntax validation passed for all files - [ ] Variable scope checked (no undefined vars in scope) - [ ] Field names match between frontend and backend - [ ] All API endpoints verified - [ ] Chart.js memory management checked - [ ] XSS risk assessed (user input escaped) - [ ] CDN MIME types verified (if third-party libraries used) - [ ] DOM references match HTML elements - [ ] Known bugs documented before fixing **STOP:** Do not proceed to Phase 1 until Phase 0 is complete. This phase takes 2-5 minutes but prevents 80% of "fixes that introduce new bugs." --- ## Phase 1: Root Cause Investigation **BEFORE attempting ANY fix:** ### 1. Read Error Messages Carefully - Don't skip past errors or warnings - They often contain the exact solution - Read stack traces completely - Note line numbers, file paths, error codes **Action:** Use `read_file` on the relevant source files. Use `search_files` to find the error string in the codebase. ### 2. Reproduce Consistently - Can you trigger it reliably? - What are the exact steps? - Does it happen every time? - If not reproducible → gather more data, don't guess **Action:** Use the `terminal` tool to run the failing test or trigger the bug: ```bash # Run specific failing test pytest tests/test_module.py::test_name -v # Run with verbose output pytest tests/test_module.py -v --tb=long ``` ### 3. Check Recent Changes - What changed that could cause this? - Git diff, recent commits - New dependencies, config changes **Action:** ```bash # Recent commits git log --oneline -10 # Uncommitted changes git diff # Changes in specific file git log -p --follow src/problematic_file.py | head -100 ``` ### 4. Gather Evidence in Multi-Component Systems **WHEN system has multiple components (API → service → database, CI → build → deploy):** **BEFORE proposing fixes, add diagnostic instrumentation:** For EACH component boundary: - Log what data enters the component - Log what data exits the component - Verify environment/config propagation - Check state at each layer Run once to gather evidence showing WHERE it breaks. THEN analyze evidence to identify the failing component. THEN investigate that specific component. ### 5. Trace Data Flow **WHEN error is deep in the call stack:** - Where does the bad value originate? - What called this function with the bad value? - Keep tracing upstream until you find the source - Fix at the source, not at the symptom **Action:** Use `search_files` to trace references: ```python # Find where the function is called search_files("function_name(", path="src/", file_glob="*.py") # Find where the variable is set search_files("variable_name\\s*=", path="src/", file_glob="*.py") ``` ### Phase 1 Completion Checklist - [ ] Error messages fully read and understood - [ ] Issue reproduced consistently - [ ] Recent changes identified and reviewed - [ ] Evidence gathered (logs, state, data flow) - [ ] Problem isolated to specific component/code - [ ] Root cause hypothesis formed **STOP:** Do not proceed to Phase 2 until you understand WHY it's happening. --- ## Phase 2: Pattern Analysis **Find the pattern before fixing:** ### 1. Find Working Examples - Locate similar working code in the same codebase - What works that's similar to what's broken? **Action:** Use `search_files` to find comparable patterns: ```python search_files("similar_pattern", path="src/", file_glob="*.py") ``` ### 2. Compare Against References - If implementing a pattern, read the reference implementation COMPLETELY - Don't skim — read every line - Understand the pattern fully before applying ### 3. Identify Differences - What's different between working and broken? - List every difference, however small - Don't assume "that can't matter" ### 4. Understand Dependencies - What other components does this need? - What settings, config, environment? - What assumptions does it make? --- ## Phase 3: Hypothesis and Testing **Scientific method:** ### 1. Form a Single Hypothesis - State clearly: "I think X is the root cause because Y" - Write it down - Be specific, not vague ### 2. Test Minimally - Make the SMALLEST possible change to test the hypothesis - One variable at a time - Don't fix multiple things at once ### 3. Verify Before Continuing - Did it work? → Phase 4 - Didn't work? → Form NEW hypothesis - DON'T add more fixes on top ### 4. When You Don't Know - Say "I don't understand X" - Don't pretend to know - Ask the user for help - Research more --- ## Phase 4: Implementation **Fix the root cause, not the symptom:** ### 1. Create Failing Test Case - Simplest possible reproduction - Automated test if possible - MUST have before fixing - Use the `test-driven-development` skill ### 2. Implement Single Fix - Address the root cause identified - ONE change at a time - No "while I'm here" improvements - No bundled refactoring ### 3. Verify Fix ```bash # Run the specific regression test pytest tests/test_module.py::test_regression -v # Run full suite — no regressions pytest tests/ -q ``` ### 4. If Fix Doesn't Work — The Rule of Three - **STOP.** - Count: How many fixes have you tried? - If < 3: Return to Phase 1, re-analyze with new information - **If ≥ 3: STOP and question the architecture (step 5 below)** - DON'T attempt Fix #4 without architectural discussion ### 5. If 3+ Fixes Failed: Question Architecture **Pattern indicating an architectural problem:** - Each fix reveals new shared state/coupling in a different place - Fixes require "massive refactoring" to implement - Each fix creates new symptoms elsewhere **STOP and question fundamentals:** - Is this pattern fundamentally sound? - Are we "sticking with it through sheer inertia"? - Should we refactor the architecture vs. continue fixing symptoms? **Discuss with the user before attempting more fixes.** This is NOT a failed hypothesis — this is a wrong architecture. --- ## Red Flags — STOP and Follow Process If you catch yourself thinking: - "Quick fix for now, investigate later" - "Just try changing X and see if it works" - "Add multiple changes, run tests" - "Skip the test, I'll manually verify" - "It's probably X, let me fix that" - "I don't fully understand but this might work" - "Pattern says X but I'll adapt it differently" - "Here are the main problems: [lists fixes without investigation]" - Proposing solutions before tracing data flow - **"One more fix attempt" (when already tried 2+)** - **Each fix reveals a new problem in a different place** **ALL of these mean: STOP. Return to Phase 1.** **If 3+ fixes failed:** Question the architecture (Phase 4 step 5). ## Common Rationalizations | Excuse | Reality | |--------|---------| | "Issue is simple, don't need process" | Simple issues have root causes too. Process is fast for simple bugs. | | "Emergency, no time for process" | Systematic debugging is FASTER than guess-and-check thrashing. | | "Just try this first, then investigate" | First fix sets the pattern. Do it right from the start. | | "I'll write test after confirming fix works" | Untested fixes don't stick. Test first proves it. | | "Multiple fixes at once saves time" | Can't isolate what worked. Causes new bugs. | | "Reference too long, I'll adapt the pattern" | Partial understanding guarantees bugs. Read it completely. | | "I see the problem, let me fix it" | Seeing symptoms ≠ understanding root cause. | | "One more fix attempt" (after 2+ failures) | 3+ failures = architectural problem. Question the pattern, not fix again. | ## Async-Specific Pitfalls ### Async infinite recursion ≠ stack overflow (harder to catch) Async recursion doesn't stack overflow the same way sync does — the event loop unwinds between awaits. This makes it INVISIBLE in local testing yet deadly in production: - **Symptom**: Container restarts silently (OOM kill), "no available server", not reproducible locally - **Root cause**: Function A calls A internally with a shifted parameter (e.g., `buildSummary(today)` → `buildSummary(yesterday)` → `buildSummary(tomorrow)`...) without a base case. Each "recursive" call leaks memory (pending HTTP requests, unresolved promises, response buffers) until the process hits the RAM limit. - **Why local testing misses it**: Local dev has abundant RAM + open file handles. The recursion eventually hits a date with no API data (returns null → error → terminates). On production with cached data, every date succeeds → infinite chain. - **Fix**: Never let function A call function A (directly or indirectly) unless you have a hard depth limit. Replace with targeted single-call fetches. **Detection trick**: If you see `const data = await buildSummary(ayer, token)` inside function `buildSummary(...)` — that's the bug. The call chain is `buildSummary(d)` → `buildSummary(d-1)` → `buildSummary(d-2)` → ... with no termination. ### Event loop starvation por concurrencia HTTP en servidores de 1 vCPU Cuando un servidor tiene 1 vCPU (NaN free, contenedores pequeños, VPS baratos), **múltiples llamadas HTTP simultáneas pueden colapsar el event loop de Node.js**:
عرض على GitHub
ملف SKILL.md هذا كبير جدا، لذلك يعرض SkillsMP القسم الاول فقط هنا. عرض على GitHub