with one click
autonomous-fleet
autonomous-fleet contains 30 collected skills from ravidsrk, with repository-level occupation coverage and site-owned skill detail pages.
Skills in this repository
The portable, tool-agnostic ENGINE for running fully-autonomous multi-agent engineering jobs. Shipped mission skills (doc-sync, test-coverage, adversarial-review-and-fix) invoke THIS engine plus exactly one ADAPTER (orca, claude-code, grok, or another runtime). Eighteen additional missions are documented under docs/exploratory/missions/ and re-promote on real-run evidence. This core holds everything that does NOT depend on orchestration tool: self-orientation, fully-autonomous coordinator behaviour with file-ledger boolean gates, context-handoff to survive compaction, the worker-placement DECISION LOGIC (dependent vs independent), the PR-per-task pipeline with commits-preserved + conflict-aware merge + worktree cleanup, the empirical risk tiers, safety rails, secret hygiene, and commit/authorship policy. It speaks in PRIMITIVES; the ACTIVE ADAPTER maps each primitive to its tool's real commands. Load with a mission skill and one runtime adapter — do not run alone.
[Tier 2 · moderate autonomy · full review gate · developer-experience] Run a live developer experience audit with frozen DX scorecard and ranked documentation gaps. Maps gstack /devex-review and /plan-devex-review into fleet ledgers with optional document-generate follow-ups. Use when onboarding friction, CLI help, or docs TTHW need evidence before ship. Trigger on: "DX audit", "developer experience review", "devex mission", "audit onboarding", "score the getting started flow", "plan devex review".
[Tier 2 · moderate autonomy · full review gate] Force a production landing page that has DIVERGED from an approved design into full fidelity with that design - section by section, with a named divergence checklist as the forcing function so it converges instead of drifting. Use when a live page no longer matches its design export and needs to be brought back into line (a single page/site, not a whole app - that's design-integration). Extracts the design system, diffs production against it, and closes each divergence as its own PR. Runs via the autonomous-fleet-core engine. Trigger on: "make the landing page match the design", "the production page diverged from the design", "bring the landing page into parity", "fix the landing page to match the mockup".
[Tier 3 · high blast radius · expect rework · run deliberately, review the floor + architecture] Adversarially review a legacy app, research the best current architecture, and rebuild it end to end on a modern foundation while PRESERVING everything it currently does. Use when an app uses outdated everything (old deps, poor architecture, e.g. JS-inlined-in-HTML, no build/module system) and needs a real modernization, not a patch. Incremental and shippable per PR — NOT a big-bang rewrite — rebuilt against a captured behaviour floor so nothing is silently lost. Runs via the autonomous-fleet-core engine. Trigger on: "rebuild this legacy app", "modernize the whole codebase", "this uses all old versions, rebuild it properly", "re-architect end to end".
[Tier 2 · moderate autonomy · full review gate · pre-build] Turn vague product intent into a frozen, review-approved product spec before any implementation. Runs gstack-style office-hours forcing questions plus CEO/design/eng plan reviews (via optional community skills), then emits docs/product-spec.md as the single source of truth for downstream build missions. Use when the user has an idea, wedge, or feature direction but no frozen spec yet. Trigger on: "frame this product", "office hours on this idea", "is this worth building", "freeze the product spec", "run plan reviews before we code", "product framing mission".
[Tier 2 · moderate autonomy · full review gate · post-ship docs] Post-ship documentation sweep and deploy verification checklist after a version lands. Maps gstack /document-release, /ship, /land-and-deploy, and /canary into fleet frozen artifacts. Use after merge or deploy when changelog, user docs, and canary notes need closure. Trigger on: "document this release", "post-ship docs", "release documentation mission", "update docs after ship", "canary doc sweep".
[Tier 2 · moderate autonomy · full review gate · security-specialist] Run a gstack-style CSO security audit (secrets archaeology, supply chain, CI/CD, OWASP, STRIDE, LLM/skill supply chain) with daily or comprehensive modes, freeze findings, then remediate confirmed issues one PR at a time with skeptic narrowing. Distinct from general adversarial-review-and-fix by infrastructure- first CSO lenses. Trigger on: "CSO audit", "security audit mission", "OWASP review", "threat model this repo", "run /cso as a fleet mission", "supply chain security pass".
[Tier 2 · moderate autonomy · full review gate · browser-grounded] Systematically QA a web application via real browser sessions, freeze findings with screenshot EVID, then fix and re-verify in a loop until health score meets threshold. Maps gstack /qa + /browse + /qa-only into fleet worktrees with build-blind review. Use when a staging or local URL exists and the user wants "test the site and fix bugs", "browser QA", or "dogfood this flow". Trigger on: "browser qa fix", "test the site and fix", "QA this app", "find UI bugs and fix them", "run QA on staging".
[Tier 1 · merge-friendly chore-category work · lower-risk when scoped] Targeted code-health pass — remove dead code, kill duplication, fix a named anti-pattern, tidy structure — WITHOUT re-architecting. The light counterpart to the archived docs/exploratory/missions/archive/legacy-rebuild/ design: improves the code as it is, preserves all behaviour, no big structural change. Use for "clean up this mess", tech-debt paydown, removing cruft after a feature, or eliminating a specific smell. Behaviour-preserving by definition; every change covered by existing or added tests proving nothing broke. Runs fully autonomously via the autonomous-fleet-core engine. Trigger on: "clean up the code", "remove dead code", "reduce duplication", "tidy this up", "pay down tech debt", "fix this anti-pattern" (when a full rebuild is NOT wanted).
[Tier 2 · moderate autonomy · full review gate] Adopt a fresh design across an existing product to full parity - VISUAL and FEATURE-WISE. Import the design (Claude Design export/URL or the claude_design MCP connector), reskin every existing screen to it, and where the design implies a feature the product lacks, design-and-build it; nothing the old product did is lost. Use for a whole-app redesign adoption, not a single page (the parked single-page variant lives under docs/exploratory/missions/archive/landing-page-convergence/). Reuses the existing backend; rewires the UI to the new design and fills gaps to full depth. Runs via the autonomous-fleet-core engine. Trigger on: "adopt this design across the product", "integrate the new design end to end", "make the app match this design fully", "redesign the whole app to this".
[Tier 2 · moderate autonomy · full review gate · incident-response] Root-cause investigation with frozen RCA document and mandatory regression test per confirmed cause. Maps gstack /investigate, /retro, and /learn into fleet worktrees. Use for production incidents, SEV postmortems, or recurring failures needing durable learnings. Trigger on: "investigate this incident", "root cause analysis", "incident mission", "write the RCA", "retro this outage", "regression test for this bug".
[Tier 3 · high blast radius · expect rework · run deliberately, review the scope artifact] Take a STALLED product to a full-fledged, complete, shippable state — adversarial review of what exists, research the market, freeze a COHERENT PRODUCT BOUNDARY, then build everything inside it to full depth (no stubs, no half-built screens) and rebuild the landing page. Use when a product has been started but never finished, when "done" keeps expanding so nothing ships, or when you need a side project driven to a credible v1. NOT for a thin MVP — the user wants real depth — but the boundary is frozen once so the run converges instead of expanding forever. Highest-risk mission: feature/cross-module work has no category in arXiv 2601.15195, so expect rework; the scope artifact (the IN/ROADMAP/FIX boundary) is the thing to eyeball. Runs via the autonomous-fleet-core engine. Trigger on: "take this product to the finish line", "finish this stalled project", "make this shippable end to end", "complete the whole product".
[Tier 2 · moderate autonomy · full review gate] Migrate ONE axis of a codebase — a framework version, a library swap, a language/runtime bump, a database/ORM change, an API-version move — while preserving everything else and keeping the suite green. Use for a single deliberate migration that's too big for dependency-update but is NOT a full rebuild: React class→hooks, a major framework major-version, swapping one library for another, a DB engine/ORM change, a REST→GraphQL of one surface. Changes one axis; preserves architecture and behaviour on every other axis. Runs via the autonomous-fleet-core engine. Trigger on: "migrate to X", "swap library A for B", "upgrade framework to v-next", "move from REST to GraphQL", "migrate the database/ORM".
[Tier 2 · moderate autonomy · full review gate · the proven two-phase workhorse] Run a rigorous CODE-GROUNDED adversarial architecture review of a repo, FREEZE it as the source of truth, then close every confirmed finding one at a time until done. Use for a security/architecture/ reliability hardening pass, a pre-production audit-and-remediate, or "review the whole app and fix everything." Phase 0 reviews the actual source (not existing docs) and a skeptic narrows out false findings; Phase 1 fixes the confirmed set with full safety rails. Runs via the autonomous-fleet-core engine. Trigger on: "adversarial review and fix", "audit and remediate", "review the whole app and fix the issues", "harden this before production", "find and fix the architecture problems".
The CLAUDE CODE adapter for autonomous-fleet-core. Maps each engine PRIMITIVE to Claude Code's native mechanics — subagents via the Task tool, git worktrees for isolation, the Bash tool for git/gh, and TodoWrite as the live task mirror. Load this alongside autonomous-fleet-core when running a mission in Claude Code instead of Orca. Because Claude Code has no separate orchestration runtime, the coordinator IS the main Claude Code session and workers are subagents or worktree-scoped sub-sessions; the file ledger is the durable source of truth and TodoWrite mirrors it.
The CODEX adapter for autonomous-fleet-core. Maps each engine PRIMITIVE to OpenAI Codex mechanics — subagents, git worktrees, shell for git/gh, and the file ledger as durable truth. Load alongside autonomous-fleet-core when running a mission in Codex (app, IDE, or CLI). The coordinator IS the main Codex thread; workers are subagents or worktree-scoped sessions; the file ledger survives compaction.
The GROK adapter for autonomous-fleet-core. Maps each engine PRIMITIVE to Grok Build mechanics — subagents via the Task tool, git worktrees for isolation, the Shell tool for git/gh, and the file ledger as the durable source of truth. Load this alongside autonomous-fleet-core when running a mission in Grok instead of Orca. Because Grok has no separate orchestration daemon, the coordinator IS the main Grok session and workers are subagents (Task tool) or worktree-scoped shell-driven sessions; the file ledger is the authority.
TEMPLATE for writing a new autonomous-fleet adapter (e.g. codex, gemini-cli, a custom CLI fleet, or a raw tmux+worktrees setup). Copy this, rename to autonomous-fleet-adapter-YOUR-TOOL, and fill in how YOUR runtime implements each PRIMITIVE the core calls. The missions and the core never change — only this mapping does. Use when adding a new orchestration runtime to autonomous-fleet. Not a runnable mission skill.
Entry point for the autonomous-fleet multi-agent engineering framework. Use whenever the user wants fully-autonomous coding runs, multi-agent orchestration, PR-per-task pipelines, or mentions autonomous-fleet — even if they have not named a specific mission yet. Routes to one mission or to fleet-program for sequential chains, loads autonomous-fleet-core plus a runtime adapter, and runs unattended on the current repo. Install from github.com/ravidsrk/autonomous-fleet. Trigger on: "use autonomous-fleet", "run autonomous fleet", "autonomous multi-agent run", "fleet mission", "which fleet mission should I use".
Orchestrate autonomous-fleet missions on one repo — sequential chains and conditional campaign DAGs with if-outcome edges. Reads fleet-outcome YAML from each mission's readiness doc to branch (e.g. audit then tests if no P0s, else dependency-update). One mission active at a time per repo; cross-repo parallel via separate sessions. Does not run parallel missions on the same repo. Use for "repo health program", "audit then test", "docs then bugs if needed", mission chains, fleet campaign, ship with proof, align then ship, quality gate. Install from github.com/ravidsrk/autonomous-fleet. Trigger on: "fleet program", "fleet campaign", "mission chain", "if P0 then", "repo health", "conditional fleet run", "ship safely", "ship with proof", "finish stalled product", "align then ship", "production ready", "quality gate".
User-invoked setup only — do not auto-activate. Configure a repo for autonomous-fleet: runtime adapter, branch prefix, default campaign bundle, optional community installs. Run when the user says setup autonomous fleet, configure fleet, or first fleet run on a repo.
First-class Orca adapter for autonomous-fleet-core — the reference runtime for supervised fleet missions. Maps engine primitives to Orca orchestration CLI commands (task-create, dispatch --inject, check --wait, worktree/terminal placement, worker_done/ask/reply, gh PR pipeline). Includes routing to orca-cli for full handoffs vs supervised fleet runs. Load with autonomous-fleet-core and one mission on Orca. Default roles: interactive @codex or @grok builds, fresh build-blind @claude reviews, @claude integrates. Trigger on Orca fleet runs, multi-agent PR pipelines on Orca, or when the host is Orca and autonomous-fleet is active.
[Tier 3 · one-axis stub→live cutover · full review gate] Fill the live implementation of the agent/service seam recorded in docs/build-plan.md, swapping the frozen stub for the configured agents framework (default eve) one seam-member-group at a time, behind the exact contract the stub froze, then remove the stub. Use after the build plan is frozen and the typed depth is done; acceptance is contract-tested against the stub fixtures plus local evals, with live deploy left to OPS. NOT for deriving the seam (scaffold-align owns that) and NOT for building typed API/data/UI depth (contract-first-build owns that); this mission only wires the live impl and cuts over the selector. Runs via the autonomous-fleet-core engine. Trigger on: "wire the live agents layer", "swap the agent stub for the real framework", "cut the agent seam over to live", "implement the agents framework behind the stub".
[Tier 2 · moderate autonomy · full review gate · reproduce-first required] Close a batch of bugs — from a list, an issue tracker, or a described set — one PR per bug, each gated by a FAILING TEST written first that reproduces the bug, then a fix that turns it green. Use for a backlog of bugs, a tracker label, or a known cluster of defects. Bug-fixes are among the categories agents are WEAKEST at (they need exact, not approximate, changes), so this mission forces exactness: prove the bug with a red test before fixing. Does not add features. Runs via the autonomous-fleet-core engine. Trigger on: "fix these bugs", "work through the bug backlog", "close the bugs labelled X", "fix this list of issues", "batch bug fixing".
[Tier 3 · high blast radius · expect rework · review the frozen build plan] Build the typed product depth (API/router, server, data/ORM, auth, payments, integrations, read-path persistence) to full depth against a PRE-FROZEN boundary, consuming docs/build-plan.md (the seam contract, invariants, and work-areas frozen by scaffold-align) without re-deriving any of it. Use after the build plan is frozen and the agents-live row can stay stubbed, to fill every typed work-area to real depth against the discovered seam plus its stub fixtures. NOT for deriving the boundary (scaffold-align owns that) and NOT for wiring the live agents impl (agents-layer owns that): never touch the live impl file or widen the seam interface. Expect rework; the artifact to eyeball is the pre-frozen docs/build-plan.md, not one this mission invents. Runs via the autonomous-fleet-core engine. Trigger on: "build the contract depth", "implement the build plan", "build the API/data/auth/payments layers", "fill in the typed product".
[Tier 1 · among the highest cross-agent merge-success categories (build/chore) · safe to run unattended] Update a repo's dependencies to current versions, fix the breakages each bump causes, and keep the suite green — one PR per logical group. Use when deps are stale, for routine maintenance, to clear security advisories, or before building on an old base. Handles version bumps, lockfile updates, and the code/config changes a bump requires (deprecations, renamed APIs, breaking changes); does NOT add features. Runs fully autonomously via the autonomous-fleet-core engine. Trigger on: "update dependencies", "bump packages", "our deps are out of date", "upgrade to latest", "fix security advisories", "dependency maintenance".
[Tier 2 - measurement-first cost optimization - full review gate] Reduce AI inference spend while holding output quality constant. Use when a repo has high LLM/API spend, needs cost controls, or wants model routing, prompt caching, batch/flex-tier pricing, provider abstraction, or token-hygiene cleanup. First gate is a baseline cost+quality harness; ship only sanctioned levers and block output-quality regressions. Hard-refuse subscription-token-as-backend/token-pool-proxy hacks; use billed provider API keys from env only. Runs via autonomous-fleet-core. Trigger on: "reduce inference cost", "optimize LLM spend", "lower token usage", "route cheaper models".
[Tier 1 · verification-led · safe to run unattended] Self-orient to a design+spec handoff, certify the scaffold builds green, certify the agent/service seam (its typed stub/live boundary) and HARDEN the stub so it exercises the WHOLE contract, then FREEZE docs/build-plan.md as the single handoff artifact the downstream build missions consume. This is verification and freeze work only: it confirms the scaffold typechecks, lints, and builds via the repo's own scripts, enumerates the seam, and records discovered per-product facts in the artifact. NOT for building product depth (that is contract-first-build) and NOT for wiring the live agents impl (that is agents-layer) — it leaves the seam stubbed and the live impl untouched. Runs via the autonomous-fleet-core engine. Trigger on: "align the scaffold", "freeze the build plan", "certify the handoff", "prep this handoff for the fleet", "orient to the design+spec handoff".
[Tier 1 · highest cross-agent merge-success category (documentation) · safe to run unattended] Bring a repo's documentation back into alignment with its actual code: README, docs/, AGENTS.md/CLAUDE.md, API references, setup/usage instructions, code comments that have drifted, and inline examples that no longer run. Use when docs are stale, after a refactor or dependency change, when onboarding docs are wrong, or for a periodic documentation-truth pass. This is a documentation mission ONLY — it does not change application behaviour or logic; it makes the docs match the code as it actually is. Runs fully autonomously via the autonomous-fleet-core engine. Trigger on: "sync the docs", "our README is out of date", "docs don't match the code", "update documentation", "fix onboarding/setup instructions", "documentation audit".
[Tier 1 · lower-than-doc/build cross-agent merge-success (test) · guard hollow-test risk] Raise real test coverage on a repo or a target area with behaviour-exercising tests — not coverage-padding stubs. Use when a module is undertested, before a refactor to lock current behaviour, after a feature shipped without tests, or for a periodic coverage pass. Adds/strengthens unit, integration, and where relevant UI tests that genuinely assert behaviour; does NOT change application logic. Runs fully autonomously via the autonomous-fleet-core engine. Trigger on: "add tests", "raise coverage", "this module has no tests", "write tests for X", "improve test coverage", "lock current behaviour with tests".