en un clic
claude_harness_forge
claude_harness_forge contient 47 skills collectées depuis rlpatrao, avec une couverture métier par dépôt et des pages de détail sur le site.
Skills dans ce dépôt
Switch model mid-session preserving thinking blocks. Used by /model command and by BRD §6.2 failover when the primary provider rate-limits or fails.
Six-phase ReAct loop (pre-check, thinking, self-critique, action, tool, post). Activates per-workflow via the thinking_level knob in config/workflows.yaml. Replaces standard ReAct for workflows where independent verification of the reasoning improves outcomes.
Mines completed sessions for repeating {tool sequence → outcome} tuples. Scores by frequency × success_rate × novelty. High-scoring tuples become "instincts" — candidate skill seeds — in instincts/pending/.
Walk-back algorithm for the Spec-Auditor subagent. Given a failure at phase N, identifies the earliest phase whose spec, if tightened, would have prevented the failure. Proposes a surgical amendment.
Sessions stored as trees (not lists). /fork creates a branch from any point; /tree navigates; /branch labels a path; /export produces HTML for review. All branches live in one session file under sessions/<project>/<session_id>.json.
Progressive context refinement for subagents. Don't dump everything; retrieve in passes. Start with a list of file paths and one-line abstracts; load full content only when the abstract proves insufficient.
Create a Business Requirements Document through Socratic five-dimension dialogue with the human. First step of the SDLC pipeline, before spec/design/build. Supports greenfield projects or single-feature additions.
Autonomous build loop with Karpathy ratcheting, GAN evaluator, browser console capture, UI standards review, 8-gate ratchet, session chaining, and cross-project learnings. Iterates story groups until all features pass or stopping criteria met.
Evaluation patterns — sprint contract format, three-layer verification, scoring rubric references.
Run the application and verify sprint contract criteria via API tests, Playwright interaction, and schema validation.
Verify that long-running operations with UI report incremental progress. Flags blocking all-at-once calls that leave users staring at a spinner with no feedback.
Reference patterns for building resilient AI-native applications — retry, fallback, circuit breaker, graceful degradation, checkpoint/resume, and LLM-specific error handling.
Pull latest forge from GitHub and upgrade scaffolded project files in place. Preserves project state, merges config, reports what changed.
Full 12-phase SDLC pipeline. BRD → Architect (up to 11 rounds) → Spec → Design → Observe → Comply → Initialize → Auto (11 gates) → Post-build. Human gates on phases 1-4. Conditional phases 6-7 for AI-native projects.
Autonomous self-testing of the forge. Creates a test project, runs the full 12-phase pipeline, self-heals on failures, fixes forge bugs when found, and produces a dogfooding report.
Test planning, test case design, test data generation, and Playwright E2E automation. Use when creating test plans, writing test cases, generating test data, setting up Playwright, or automating end-to-end tests.
Generate and display a terminal dashboard showing project progress — stories spec'd, coded, unit-tested, and E2E-verified per group. Writes to specs/status.md.
Log a requirement change to the BRD changelog, run impact analysis, and cascade updates through affected specs, design, and implementation.
Design the system architecture including layered dependencies, API contracts (endpoints, schemas, errors), data models, folder structure, and deployment topology. Output a detailed design document to `specs/design/` with all decisions justified.
Reference patterns for AI compliance — OWASP Agentic Top 10, bias detection, fairness testing, PII scanning, GDPR/HIPAA/SOC2 checklists, model cards, and content filtering.
Run compliance review — PII handling, audit trails, data retention, and ML fairness checks with model card generation.
Analyze and optimize token usage — cost summary, cost per agent, per story, per gate, cache hit rates, and recommendations for context reduction. Use --summary for a quick cost overview.
Review and submit anonymized harness findings as GitHub issues to the forge repo. Opt-in, user-confirmed, no secrets or PII.
Decompose a Business Requirements Document into epics and user stories with acceptance criteria, estimate effort, and identify hard dependencies for parallel execution. Output stories in `specs/stories/` with dependency-graph.md and story files.
Decompose a BRD into epics, user stories, acceptance criteria, and a dependency graph with parallel groups for agent team execution.
Generate test plan, test cases, test data fixtures, and Playwright E2E tests mapped to acceptance criteria.
Interactive stack interrogation, design artifact generation, decision verification, and learnings persistence. Runs after BRD approval, before spec decomposition.
UX patterns for agentic AI applications — intent preview, autonomy dial, confidence signals, audit trails, escalation, streaming, multi-agent dashboards, and error recovery.
Reference patterns for managing LLM context windows — token budgets, prompt caching, progressive disclosure, context compression, cost optimization, and anti-patterns.
Generate a model card by extracting model metadata, metrics, dataset info, and bias analysis from training and evaluation code.
Scaffold observability for a project — OpenTelemetry tracing, structured logging, Grafana dashboards, alerting rules, and Docker Compose wiring.
Reference patterns for Retrieval-Augmented Generation — chunking strategies, embedding models, vector databases, retrieval patterns, reranking, evaluation, and agentic RAG.
Scaffold a RAG pipeline — embedding service, vector store, retrieval service, chunking, and evaluation tests.
Add resilience patterns to existing code — retry with backoff, circuit breakers, timeouts, fallbacks, and graceful degradation.
Scaffold multi-tenancy — tenant middleware, row-level security, per-tenant rate limiting, feature flags, and tenant admin API.
Scaffold durable workflow orchestration — workflow definitions, activities, HITL signal handlers, checkpoint/resume, and saga compensations.
Generate production code and tests for a story group using agent teams for parallel execution.
Generate Docker Compose stack, Dockerfiles, environment config, init.sh bootstrap script, and verify local deployment with health checks.
Generate system architecture and UI mockups. Spawns architect (if not already run) and ui-designer concurrently.
Standard workflow for fixing a GitHub issue. Fetches issue details, creates a branch, implements the fix with tests, and prepares a PR.