Triple-blind AI testing protocol for scientifically rigorous experiments that test AI behavior changes (e.g. CLAUDE.md variants, prompt strategies) while eliminating coordinator and evaluator bias. Use when the user wants to run /experiment, design a…
Skills in this repository
jleechanorg/claude-commands - Page 2
SkillsMP has collected 447 skills from jleechanorg/claude-commands. Open a skill to review its source and details.
jleechanorg/claude-commandsShowing 40 of 447 collected skills.
Read-only fleet-size consumer for the 6 mac + 16 linux ezgha runner invariant. Restart-only remediation is UNAVAILABLE (fail-closed) pending jleechanorg/ez-gh-actions PR 67 (dual-Lima convergence) and PR 70/issue 60 (recovery controller) merging AND being…
Analyze conversation and git history to find gaps where cold reviews (codex, Bugbot, CodeRabbit, /reviewdeep) caught issues that dark-factory in-pipeline reviewer nodes missed. Fans out subagents, opens PRs, drives each through /green, and merges (with…
Display the Dark Factory pipeline node graphs, gates, node types, edge conditions, and handler mappings.
Use when a persistent repository, tool, configuration, automation, launcher, wrapper, or installed CLI fix could remain only in a working tree, topic branch, or one machine's live state.
Analyze and fix PR blockers (CI failures, merge conflicts, bot/review feedback) to get a PR into a mergeable state — never merges it. GitHub status is the sole source of truth over local assumptions. Slash command /fixpr. Also embodied as the copilot-fixpr…
Evidence-based E2E test generator for your-project.com (Real Mode Only). Use for the /generatetest slash command — generates self-contained MCPTestBase test files under testing_mcp/ with built-in evidence bundle generation (git provenance, checksums,…
Goal-driven harness loop — keep working on a goal until /es, /er, /code_standards, and Codex plugin review all pass. All gates use adversarial subagents. Alias: /h
Use when auditing or optimizing CLAUDE.md, AGENTS.md, GEMINI.md, slash commands, skills, or other coding-agent harness surfaces, or when recurring agent failures suggest instruction, policy, memory, test, or automation gaps.
Harness-fix meta-skill (slash: /meta) — analyze an agent behavior failure (autonomy violation, refusal, premature stopping) and run /harness to fix the agent, NOT the underlying task. Input: Slack thread URL, pasted conversation, or freeform description.
Rotate or repair Hermes Slack credentials, provision the complete Hermes Slack app permission baseline from the beginning, update macOS shell exports safely, and verify the live gateway path. Use for Hermes Slack token rotation, Slack app reinstall, Socket…
Alert when the Mac runner host's free disk drops below 50GB, auto-clean safe targets (evidence bundles, scratchpads, merged-PR worktrees) below 20GB. Closes the gap where mac-runner-di[REDACTED_OPENAI_KEY] (retired) and ezgha's own docker-daemon-view disk…
Harden a stated goal into ironclad exit criteria (stronger than asked, binary, executable, externally anchored, anti-gaming, iterate-until), set them durably (bead/goal/STATE/memory), then execute toward them. Use when the user invokes /ironclad, asks for…
Terminate macOS SecurityAgent and dismiss all stacked keychain modal prompts. Programmatically dismiss active macOS credential popups and debug headless/background keyring access issues (launchd, cron, isolated tmux test runners). Slash command:…
Capture durable learnings from failures, corrections, repeated mistakes, successful recovery patterns, or direct /learn requests. Use whenever the user invokes /learn, asks to save/remember/capture a lesson, or when a workflow failure should become reusable…
Validate that level-up flows are complete, auto-selected, editable, and evidenced with real LevelUpAgent output.
Trigger and monitor the level-up organic test (testing_mcp/core/test_level_up_organic.py) against the PR's deployed preview server.
Steering work on $USER's Ubuntu machine (jeff-ubuntu) via SSH. Use when asked to run commands, install software, check services, or manage files on the Ubuntu box.
Detect Gemini SDK / BYOK leaks when local AGY-default provider is expected.
Use when supervising the level-up ZFC migration loop in this repo, especially when the cleanup-first roadmap is drifting, AO workers need steering, or PR sequencing must stay aligned with the canonical roadmap.
Search across all memory systems — ~/roadmap, beads, claude memories, hermes sqlite, hermes briefings, hermes index, openclaw memories, wiki, history, and slack. Use whenever the user asks to search memories, find something in memories, or looks for anything…
Use for mobile browser investigations in this repo, especially Firebase/Google auth, iOS Safari or Chrome iOS, Incognito/Private Browsing, storage partitioning, redirect flows, and user-visible mobile-only failures.
Situational assessment, beads + roadmap sync after a work block; writes a self-contained nextsteps markdown doc (TOC, executive summary, full detail, bead links), updates learnings + README, Claude auto-memory, mem0, and beads. Prefers editing existing…
OpenClaw agent model configs — which work, which are broken/quota-limited, and how to switch
Autonomous convergence via orchestration — runs /converge repeatedly inside a tmux agent (managed by the existing /orch system) until success criteria are met, time limit reached, or max attempts reached, then executes a final /pushl → /reviewdeep → /copilot…
Run one benchmark script that compares pairv2, pair via Claude Teams, and pair via direct Python
Pairv2 implementation philosophy: LLM-decided outcomes, fail-soft file handling, and evidence-first promotion
Use when designing, debugging, reviewing, or scripting any work that has independent items (rows, files, tests, migrations, jobs, agent lanes, API sweeps, builds). Slash command `/parallel`. Enforces a single rule — the speed ceiling is the workload's real…
Run deterministic Playwright UI tests for local or test environments, including isolated browser profiles, traces, video, and multi-browser coverage. Use when the user asks for /playwright, end-to-end tests, reproducible UI testing, visual regression, or…
Lazy senior-dev mode — the seven-rung ladder that decides whether to write code at all, whether to reuse code that already exists, and whether to prefer stdlib/platform/installed deps over a new dependency. Use before writing any code, every PR diff, every…
Use this skill any time a .pptx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx file (even if the extracted content will be used…
Drive all open PRs toward /green (CI green + no merge conflicts) by fixing CI failures, and toward draft-phase quality readiness by resolving comments and running the smoke gate. CodeRabbit/Bugbot are optional advisory reviewers — surfaced for information,…
Use for production-code PRs ($PROJECT_ROOT/**, gates, ZFC). Requires full GitHub URL to governing roadmap/design doc and a `br` bead ID in the body.
Canonical /green definition — CI green + no merge conflicts, both verified at current PR HEAD SHA. Quality gates (evidence, review, comment-resolution) live in the draft-first-pr skill, not here.
This skill should be used when the user asks for a PR report, PR audit, PR delta analysis, /pr-report, or wants per-PR summary with delta files/lines. Generates a structured report for one or more open PRs covering purpose, files changed, +/− line counts,…
Interpret user-requested read-only work as protecting Git-tracked files while allowing commands and writes outside the Git index.
Canonical /repro workflow for copy+targeted bug repro, twin clones, evidence exports, same-symptom verdicts, and red/green provenance.
Run matched A/B reviewer calibration for /f: compare factory reviewer, raw codex exec, and delegated subagent reviews on the same frozen PR/work-item envelope.
Use this skill whenever a bug tempts you to add backend protection, fallback, clamp, sanitizer, retry, suppression, guardrail, or workaround logic. It enforces root-cause analysis first: inspect raw model prompts/responses and fix upstream prompt/schema/agent…
Operational health snapshot for the jleechanorg self-hosted runner fleet. Use when user says "check runners", "are runners up", "is jeff-ubuntu down", "diagnose runner health", or before/after runner operations. Produces a structured Markdown report at…