Skip to main content
Run any Skill in Manus
with one click

solo-founder-nodes

Stars1
Forks0
UpdatedJuly 10, 2026 at 19:21

Benchmark-driven development for AI agents — turns "I have an idea / prototype / half-built app and an agent that demos but does not hold up" into "an agent that completes real benchmark tasks IN the live app, browser-verified, without cheating." Runs the loop discover → benchmark → setup → build → adapter → verify → iterate under four non-negotiables (held-out · no answer-keys · in-app transfer · honest provenance), with the honest-lane clean-probe rule, a local-first memory substrate, and a Design Bridge for UI. Use when a (solo) founder wants to build or validate an AI agent for their app. Triggers: "build the agent layer for my app", "benchmark my agent", "prove my agent works in production", "which benchmark fits my agent", "make my agent pass SpreadsheetBench/BankerToolBench/SWE-bench in my app", "my agent demos but fails on real tasks", "eval my agent honestly". The user's coding agent (Claude Code, Codex, OpenClaw, Hermes, Trae) drives; the user steers by comment. This single skill IS the suite — it r

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

File Explorer
100 files
SKILL.md
readonly