원클릭으로
spike
Throwaway experiments to validate an idea before build.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
Throwaway experiments to validate an idea before build.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
| name | spike |
| description | Throwaway experiments to validate an idea before build. |
| version | 1.1.0 |
| author | Adapted from Hermes Agent (gsd-build/get-shit-done) |
| license | MIT |
| tags | ["spike","prototype","experiment","feasibility","throwaway","exploration","research","planning","mvp","proof-of-concept"] |
Use this skill when you want to feel out an idea before committing to a real build — validating feasibility, comparing approaches, or surfacing unknowns that no amount of research will answer. Spikes are disposable by design. Throw them away once they've paid their debt.
Load this when you hear things like "let me try this", "I want to see if X works", "spike this out", "before I commit to Y", "quick prototype of Z", "is this even possible?", or "compare A vs B".
Every spike follows this loop:
decompose → research → build → verdict
↑__________________________________________↓
iterate on findings
Break the user's idea into 2-5 independent feasibility questions. Each question is one spike. Present them as a table with Given/When/Then framing:
| # | Spike | Validates (Given/When/Then) | Risk |
|---|---|---|---|
| 001 | websocket-streaming | Given a WS connection, when LLM streams tokens, then client receives chunks < 100ms | High |
| 002a | pdf-parse-pdfjs | Given a multi-page PDF, when parsed with pdfjs, then structured text is extractable | Medium |
| 002b | pdf-parse-camelot | Given a multi-page PDF, when parsed with camelot, then structured text is extractable | Medium |
Spike types:
a/b/c)Good spike questions: specific feasibility with observable output. Bad spike questions: too broad, no observable output, or just "read the docs about X".
Order by risk. The spike most likely to kill the idea runs first. No point prototyping the easy parts if the hard part doesn't work.
Skip decomposition only if the exact spike question is already known. Then take it as a single spike.
Present the spike table. Ask: "Build all in this order, or adjust?" Let adjustments happen before writing any code.
Spikes are not research-free — you research enough to pick the right approach, then you build. Per spike:
Brief it. 2-3 sentences: what this spike is, why it matters, key risk.
Surface competing approaches if there's real choice:
| Approach | Tool/Library | Pros | Cons | Status |
|---|---|---|---|---|
| ... | ... | ... | ... | maintained / abandoned / beta |
Pick one. State why. If 2+ are credible, build quick variants within the spike.
Skip research for pure logic with no external dependencies.
Use web search and docs for the research step:
# Find candidates
# Search: "python websocket streaming libraries 2025"
# Read the actual docs
# Extract: https://websockets.readthedocs.io/...
# Check what's installed
pip show websockets | grep Version
One directory per spike. Keep it standalone.
spikes/
├── 001-websocket-streaming/
│ ├── README.md
│ └── main.py
├── 002a-pdf-parse-pdfjs/
│ ├── README.md
│ └── parse.js
└── 002b-pdf-parse-camelot/
├── README.md
└── parse.py
Bias toward something interactive. Spikes fail when the only output is a log line that says "it works." You want to feel the spike working. Default choices, in order of preference:
Depth over speed. Never declare "it works" after one happy-path run. Test edge cases. Follow surprising findings. The verdict is only trustworthy when the investigation was honest.
Avoid unless the spike specifically requires it: complex package management, build tools/bundlers, Docker, env files, config systems. Hardcode everything — it's a spike.
Building one spike — a typical sequence:
mkdir -p spikes/001-websocket-streaming
# Write README.md with question, approach, research
# Write main.py with the experiment
cd spikes/001-websocket-streaming && python3 main.py
# Observe output, iterate.
Parallel comparison spikes (002a / 002b) — when two approaches need real engineering, fan out with subagents (one per approach). Each returns its own verdict; you write the head-to-head.
Each spike's README.md closes with:
## Verdict: VALIDATED | PARTIAL | INVALIDATED
### What worked
- ...
### What didn't
- ...
### Surprises
- ...
### Recommendation for the real build
- ...
VALIDATED = the core question was answered yes, with evidence. PARTIAL = it works under constraints X, Y, Z — document them. INVALIDATED = doesn't work, for this reason. This is a successful spike.
When two approaches answer the same question (002a / 002b), build them back to back, then do a head-to-head comparison at the end:
## Head-to-head: pdfjs vs camelot
| Dimension | pdfjs (002a) | camelot (002b) |
|-----------|--------------|----------------|
| Extraction quality | 9/10 structured | 7/10 table-only |
| Setup complexity | npm install, 1 line | pip + ghostscript |
| Perf on 100-page PDF | 3s | 18s |
| Handles rotated text | no | yes |
**Winner:** pdfjs for our use case. Camelot if we need table-first extraction later.
If spikes already exist and you're asked "what should I spike next?", walk the existing directories and look for:
Propose 2-4 candidates as Given/When/Then.
spikes/ in the repo rootNNN-descriptive-name/README.md per spike captures question, approach, results, verdictAdapted from the GSD (Get Shit Done) project's /gsd-spike workflow — MIT © 2025 Lex Christopherson (gsd-build/get-shit-done).
Browser automation CLI for AI agents. Use when the user needs to interact with websites, including navigating pages, filling forms, clicking buttons, taking screenshots, extracting data, testing web apps, or automating browser actions.
Reviews and removes unsupported, non-idiomatic code and UI/design changes introduced on a branch. Use when implementing or reviewing a feature or fix before handoff.
Design principles for REST and GraphQL APIs in 2025. Enforces OpenAPI-first, versioning strategies, and standardized error responses.
Use the `fm` CLI (Apple Foundation Models) for on-device and Private Cloud Compute LLM inference on macOS. Use when the user wants to prompt, chat, serve, or count tokens with Apple's built-in foundation models. Also use for structured output generation (`fm schema`), checking model availability (`fm available`), or checking quota usage (`fm quota-usage`). Trigger on phrases like 'run fm', 'Apple Foundation Model', 'on-device LLM', 'Apple AI', 'fm respond', 'fm chat', 'fm serve', 'fm token-count', 'fm schema', or any task involving Apple's local AI models.
This skill should be used when the user asks to "trigger a build", "check build status", "watch a build", "view build logs", "retry a build", "cancel a build", "list builds", "download artifacts", "upload artifacts", "manage secrets", "create a pipeline", "list pipelines", or "interact with Buildkite from the command line". Also use when the user mentions bk commands, bk build, bk job, bk pipeline, bk secret, bk artifact, bk cluster, bk package, bk auth, bk configure, bk use, bk init, bk api, or asks about Buildkite CLI installation, terminal-based Buildkite workflows, or command-line CI/CD operations.
Buildkite CI/CD pipeline setup, migration, validation, hardening, and live pipeline management. Use when Codex needs to create, review, debug, or update Buildkite pipelines, .buildkite/pipeline.yml files, steps, queues, triggers, artifacts, annotations, secrets, OIDC, or deployment stages.