원클릭으로
ai-starter-kit
ai-starter-kit에는 Andersonlimahw에서 수집한 skills 42개가 있으며, 저장소 수준 직업 범위와 사이트 내 skill 상세 페이지를 제공합니다.
이 저장소의 skills
Visualize an agent's perceive→think→act loop. Inspects tool calls, latency per step, decision branches, retry patterns. Surfaces bottlenecks. Triggers on "trace agent", "why is agent slow", "agent loop analysis", "what did the agent do".
Compute context window budget for a prompt/agent — sums system prompt, tool descriptions, examples, conversation history; flags overruns; suggests cuts. Triggers on "context budget", "how much context", "will this fit", "budget my prompt".
Treat dataset / eval JSONL as code — lint, validate schema, detect duplicates, leakage between train/eval, label imbalance, PII. Karpathy "Software 2.0" lens. Triggers on "audit dataset", "lint jsonl", "check for leakage", "validate eval data".
Enforce "evals over vibes" — before any prompt, agent, or model change, require an eval set with success criteria and baseline measurement. Blocks vibe-only changes. Triggers on "tune prompt", "improve agent", "change model" without prior eval set.
Build a 10-30 case eval dataset from a PRD, skill, or prompt — covering golden path, edge cases, and adversarial inputs. Output as JSONL with input/expected/category. Triggers on "build eval set", "create evals", "generate test cases for prompt".
Break down any ML/AI concept to first principles — strip jargon, derive from basics, build physical/mathematical analogy. Karpathy intuition-building style. Triggers on "first principles", "explain like I'm 5", "derive from scratch", "why does X actually work".
Generate minimal, runnable from-scratch implementation of LLM concepts (attention, BPE tokenizer, backprop, RoPE, KV-cache, MoE, LoRA, RLHF). Karpathy-style pedagogy — code first, theory in comments. Triggers on "explain X from scratch", "minimal X implementation", "show me how X works in code", "build X from zero".
Add a new entry to the existing llm-wiki — following the project's template (concept → intuition → math → code → references). Maintains consistency across entries. Triggers on "add to llm-wiki", "new wiki entry", "document concept X".
Generate a progressive learning notebook for any domain — starts at bigram baseline, escalates to MLP, RNN, transformer. Each stage adds one concept. Karpathy makemore pedagogy. Triggers on "build progressive notebook", "makemore-style tutorial", "teach me X step by step".
Draw the autograd computation graph for any Python function. Outputs ASCII graph or graphviz DOT. Use to visualize backprop, debug gradient flow, teach chain rule. Triggers on "trace gradient", "show autograd graph", "visualize backprop", "compute graph for".
Re-explain any ML/AI Python file line-by-line in Karpathy lecture style. Walks reader through the code with intuition, shape annotations, and why-decisions. Triggers on "explain this file like Karpathy", "walk me through this code", "lecture-style explain", "explain line by line".
Karpathy-style TL;DR of an ML paper — intuition first, math second, with a runnable code sketch and "why this matters". Triggers on "tldr paper", "explain this paper", "summarize arxiv", "paper summary".
Compress a system prompt or skill body while preserving semantic content. More aggressive than caveman:compress — targets system prompts and tool descriptions, validates with eval set. Triggers on "compress this prompt", "shrink system prompt", "reduce prompt tokens".
Version prompts/skills like code — append entry to CHANGELOG, run eval, attach eval delta to version. Forces "every prompt change = versioned + measured". Triggers on "version this prompt", "release prompt v2", "log prompt change".
Run current eval suite vs stored baseline. Reports per-case delta, aggregate score, and flags regressions. Triggers on "run regression eval", "did this regress", "compare to baseline".
Audit an eval, reward function, or agent for reward hacking / Goodhart's law / spec gaming. Karpathy "outcome supervision" lens. Triggers on "reward hacking", "goodhart check", "is this metric gameable", "spec gaming".
Estimate cost / compute / data tradeoff before fine-tuning or scaling. Applies Chinchilla scaling, FLOP budgets, and rough cloud cost. Triggers on "should I fine-tune", "estimate training cost", "scaling laws", "chinchilla optimal".
Inspect how text gets tokenized by BPE — count tokens, highlight edge cases (unicode, whitespace, digits, code), suggest compaction. Karpathy "Let's build the GPT Tokenizer" lens. Triggers on "count tokens", "tokenize this", "tokenizer audit", "how many tokens".
Scan a repo or diff for "vibe-coded" sections — code that lacks tests, evals, logging, error handling, or has hardcoded prompts/magic numbers. Karpathy warning lens. Triggers on "vibe check", "audit for vibes", "is this vibe-coded".
Semantic diff between two model checkpoints, configs, or training runs. Reports param count delta, layer changes, hyperparam delta, eval score delta. Triggers on "diff checkpoints", "compare models", "what changed in this checkpoint".
Generate a Karpathy Zero-to-Hero style learning path for any AI/ML topic. Outputs ordered: prerequisites, videos, papers, hands-on exercises, evaluation milestones. Triggers on "learning path for", "how do I learn X", "study plan", "roadmap for".
Pipeline stage-0 prompt refiner. Runs FIRST — before skills-selector and smart-dispatch — turning a raw idea, pasted draft, or the current chat into a definitive, production-grade prompt (Anthropic prompt-engineering best practices), then hands the refined prompt + an Execution Map (agents, skills, models, effort, time & token estimate) to the routers so they pick the best skills/models with sharpened context. When no file path is given, the input IS the chat/pasted text and the output is returned inline (processed naturally), not forced to a file. Use when the user invokes /senior-prompt-engineer, says "refina meu prompt", "melhora esse prompt", "transforma essa ideia em prompt", "reescreve idea.md em prompt.md", "prompt definitivo", "prompt profissional", or pastes a draft prompt/idea to be hardened. Cross-CLI: Claude, Codex, Gemini, OpenCode, Antigravity (agy), lemon-code (lemon).
Meta-skill gatekeeper that runs FIRST on every task to decide which other skill(s) — if any — should activate. Analyzes the user request, matches it against a ranked catalog (frontend-design, ui-ux-pro-max, superpowers/planning, git-*, mcp-builder, smart-dispatch, claude-api, etc.), and emits a tight SELECTION plan so other skills are loaded on-demand only. Minimizes token/context consumption by refusing to activate heavy skills unless clearly justified. Triggers on ANY user instruction that could plausibly benefit from a specialized skill — i.e. essentially every turn that starts new work. Keywords: "plan", "design", "ui", "ux", "implement", "build", "create", "fix", "refactor", "commit", "pr", "review", "test", "debug", "deploy", "mobile", "video", "mcp", "skill", "architecture", "api", "ci", "docs", "copy", "social", "setup", "doctor".
Automatically routes tasks to the optimal AI agent, model, or provider based on complexity, cost, and capability. Use when implementing features, fixing bugs, or any multi-step development work. Triggers on "implement", "build", "create", "fix", "add feature", "develop", or when the user asks to do any coding task.
Multi-skill workflow composer. Activates ONLY for tasks that require combining multiple skills in sequence or resolving skill conflicts. Single-skill tasks: return [] — individual skills self-trigger via their own descriptions.
Debugging phase. Guided loop for hard bugs: reproduce, minimize, hypothesize, instrument, fix. Trigger: "/diagnose", "diagnose this bug", "debug this".
Alignment phase. Use before writing any code to ensure every design decision is resolved through an intensive interview process. Trigger: "/grill-me", "grill me", "interview me about this plan".
Refactoring phase. Identifies opportunities to make the codebase more modular and AI-friendly. Trigger: "/improve-codebase-architecture", "improve architecture", "refactor for better structure".
Initialization phase. Scaffolds the per-repo config for engineering skills. Trigger: "/setup-matt-pocock-skills", "setup engineering skills".
Quality phase. Enforces a strict Red-Green-Refactor loop. Trigger: "/tdd", "build this using TDD", "write tests first".
Execution planning phase. Breaks a PRD or complex plan into "vertical slices" (tasks). Trigger: "/to-issues", "create issues", "break this down into tasks".
Documentation phase. Synthesizes the current conversation into a formal Product Requirements Document (PRD). Trigger: "/to-prd", "create PRD", "write a PRD for this".
Maintenance phase. Organizes messy issues into actionable tasks using a state machine. Trigger: "/triage", "triage these issues", "organize the backlog".
Contextual awareness phase. Explains the system-wide impact of a change or architectural pattern. Trigger: "/zoom-out", "zoom out", "give me the big picture".
Jest Expert. Mock configuration, coverage, and asynchronous testing.
Automatically routes tasks to the optimal AI agent, model, or provider based on complexity, cost, and capability. Use when implementing features, fixing bugs, or any multi-step development work. Triggers on "implement", "build", "create", "fix", "add feature", "develop", or when the user asks to do any coding task.
Commit semântico. Usar quando: "commit", "salvar", "cn".
Checklist genérico de revisão de código. Usar quando: "review", "revisar código", "verificar qualidade", "code review", "está bom esse código?".
Metodologia científica de debugging. Usar quando: "debug", "bug", "erro", "não funciona", "investigar problema", "por que isso quebrou".
Conceitos fundamentais de LLMs como contexto para decisões de arquitetura de agentes. Ativar quando discutir: prompts, tokens, contexto, temperatura, embeddings, RAG, fine-tuning, agentes, escolha de modelo, latência de inferência.