Skip to main content
Ejecuta cualquier Skill en Manus
con un clic

solo-founder-nodes

Estrellas1
Forks0
Actualizado10 de julio de 2026 a las 19:21

Benchmark-driven development for AI agents — turns "I have an idea / prototype / half-built app and an agent that demos but does not hold up" into "an agent that completes real benchmark tasks IN the live app, browser-verified, without cheating." Runs the loop discover → benchmark → setup → build → adapter → verify → iterate under four non-negotiables (held-out · no answer-keys · in-app transfer · honest provenance), with the honest-lane clean-probe rule, a local-first memory substrate, and a Design Bridge for UI. Use when a (solo) founder wants to build or validate an AI agent for their app. Triggers: "build the agent layer for my app", "benchmark my agent", "prove my agent works in production", "which benchmark fits my agent", "make my agent pass SpreadsheetBench/BankerToolBench/SWE-bench in my app", "my agent demos but fails on real tasks", "eval my agent honestly". The user's coding agent (Claude Code, Codex, OpenClaw, Hermes, Trae) drives; the user steers by comment. This single skill IS the suite — it r

Instalación

Instalar con Codex o Claude Copia este prompt, pégalo en Codex, Claude u otro asistente, y deja que revise la página de la skill y la instale por ti.

Explorador de archivos
100 archivos
SKILL.md
readonly