ワンクリックで
dotfiles
dotfiles には drewstone から収集した 53 個の skills があり、リポジトリ単位の職業カバレッジとサイト内 skill 詳細ページを表示します。
このリポジトリの skills
Raise the ceiling, don't climb toward it. When a target is nearly hit or a metric has plateaued, question the target itself: name the binding constraint, separate the physics floor from the assumed floor, design the regime change that makes 10x reachable, commit through the valley. Triggers: 'breakout', '10x this', 'raise the ceiling', 'why are we capped', 'think way bigger', 'this metric is a cage'.
Prove that a merged change is actually live and production behavior is measured against the deployed artifact. Use after merge, before claiming deployed/live/validated, for Cloudflare Pages/Workers deploys, cache claims, perf claims, and release closeout.
Diagnose and improve eval harnesses built on @tangle-network/agent-eval or similar trace-first systems. Use when success/failure rates, scorecard deltas, judge outputs, promotion gates, benchmark rows, or improvement campaigns may be contaminated by harness bugs, evaluator drift, routing/auth failures, missing traces, or invalid baselines.
Goal-pursuit loop: measure, diagnose, experiment, verify, compare, and iterate against a measurable target. Triggers: evolve, optimize, make this better, push to target.
Split one messy experiment branch into N clean, independent branches — one logical change each, built from the merge-base, disjoint change-sets, each an isolated reviewable PR. Triggers: 'finalize', 'split this branch', 'atomize this', 'make reviewable PRs', 'decompose this mess'.
Read repo/session state and choose the single next skill: exploit (evolve/polish), explore (pursue/meta-harness/breakout), bootstrap (eval-agent), diagnose (diagnose/eval-harness-diagnose), or step back (reflect). One dispatch, then exit.
Adopt @tangle-network/hub-sdk in any product that needs to talk to the Tangle hub (OAuth connections, tools/search/describe/invoke, capability tokens, policies). Use whenever a product imports @tangle-network/agent-integrations, maintains a local hub-client.ts, or hand-rolls fetch() to /v1/hub/*. Forces kill-and-replace, not additive code.
The THINK phase before you spend compute. Research the solution space (prior art, competitors, the real ceiling), generate a diverse field of candidate mechanisms, rank them by expected value, sequence by information gain, and hand a ranked portfolio to /evolve or /pursue. Turns greedy poke-and-measure into science. Triggers: 'hypothesize', 'what should we try', 'research the space first', 'generate options', 'before we optimize', 'what are the bets'.
Automated architecture evolution after metric progress plateaus. Discover the code loop, create missing evals, run parallel proposers, benchmark variants, and keep the best patches.
Dynamic multi-agent workflow composition: decompose a goal no single skill covers, pick a dependency structure, layer coordination policies, wire skills as stages, compile to a Workflow script, run, synthesize. Triggers: 'orchestrate', 'compose a workflow', 'too big for one skill', 'fan this out', 'pipeline of agents'.
Design and build a generational improvement when the current approach is wrong or plateaued. Audit, choose a coherent architecture, implement, test, and hand off.
Merged into /evolve's structured-hypothesis mode. Use /evolve for hypothesis-driven experimentation, competitive landscape, and the bootstrap-CI promotion gate.
Before optimizing, debugging, or speeding up any LIVE system, stand up the FULL measured harness of the real production path FIRST — instrument every hop, benchmark the real (not local) path, build a reversible test loop, trace your own run, baseline + decompose — in one parallel fan-out. Skip it and you burn days optimizing a system you can't see, acting on a number true only in a narrower context than you present it.
Audit technical docs for weak claims, AI slop, unclear product boundaries, false promises, and prose that misleads builders or buyers.
Design or revise visible product UI with reference-first judgment, real mode-specific controls, low-copy surfaces, and screenshot-based verification.
Capability-preserving simplification for active code work. Use when the user asks to simplify, modularize, remove god objects, reduce duplication, make code reusable, clean up abstractions, deep clean without losing capability, or asks "anything else to simplify?" after a feature/refactor/PR.
Design + run a rigorous, equal-compute, executable-graded comparison of agent ARCHITECTURES (topologies, coordination policies, profiles) across a controlled difficulty axis — to find WHERE one approach beats another, not just whether it solves a task. Use before any "does smart multi-agent / topology X beat dumb loop / baseline Y" experiment. Reuse the substrate; never rebuild the harness.
Adopt @tangle-network/agent-app — the shared application-shell framework for agent products — either greenfield (new product) or by migrating an existing app (from ANY stack). Starts with a discovery interview (product surface, agent surface, eval surface, features, sandbox-or-not, billing, integrations), then routes to the right module set + path. Covers the engine/shell/domain layering rule, per-module seams, sandbox AND non-sandbox (browser/edge copilot) wiring, the migration lift-loop, and anti-patterns. Use when standing up a new agent product, deciding what belongs in the app vs the framework, or porting an existing app onto agent-app.
Before running any eval, A/B, benchmark, or experiment, prove the metric can actually see what you claim to measure — and prove the task is hard enough to need the capability. Skip this and you measure the wrong thing for three experiments straight.
When tempted to collapse an ambitious-but-unproven architecture (multi-agent topology, context-lifecycle management, a recursive loop system) into a dumb/old pattern because an early A/B looked marginal — don't. Marginal-early almost always means the regime that makes the architecture pay off wasn't active. Find that regime, build the missing competency, then judge.
When you catch yourself doing the safe/easy version of a task or experiment, force the harder one that could actually fail. Timidity disguises itself as "let's start simple" and as the flattering result you didn't try to kill.
Audit whether an autonomous agent actually observes state, uses tools, improves from outcomes, and stays aligned to user intent using traces, logs, and artifacts.
Extend @tangle-network/agent-eval internals: campaigns, scorecards, trace capture, backend integrity, held-out gates, analysts, auto-PR loops, RL bridge, and eval release checks.
Root-cause one null, surprising, or suspicious run result. Verify raw data, classify the cause, and decide whether to fix code, metric, design, or belief.
Browser Agent Driver CLI operator for browser automation, UI/design audits, auth state, showcases, and benchmark runs. Triggers: "run bad", "browser agent", "design audit", "webbench", "automate this site".
Drive failing CI to green by reading remote failures, reproducing locally when possible, fixing root causes, pushing, waiting, and repeating without shortcuts.
Staff-engineer review of diffs, branches, docs, APIs, SDKs, and customer-facing surfaces. Findings first, severity ranked, file:line grounded, with a concrete fix plan.
Measured codebase cleanup: dead code, dependency cycles, weak types, duplicate logic, deprecated paths, test debt, and complexity, using real tools and before/after proof.
Analyze test, CI, benchmark, or eval failures; cluster by root cause; rank by impact and fix effort; produce concrete fix hypotheses.
Build LLM-as-judge components from real references: gather examples, generate rubrics, score outputs, return findings, and wire improvement loops.
Produce a session-to-session brief with current state, git/PR status, decisions, blockers, verification, and exact next actions for a fresh agent.
Security adversarial validation: derive invariants, attack surface, fuzz targets, credential risks, race conditions, and coverage gaps; extend existing tests to prove fixes.
Run multiple independent /pursue-grade architecture tracks in parallel, each with its own brief, build, verification, and central synthesis.
Apply a fixed quality rubric to existing work and fix every gap: correctness, design, robustness, tests, and public interface. Triggers: polish, tighten, production-grade.
Product UI audit/redesign for apps, dashboards, workflows, marketing pages, fake components, visual slop, information architecture, copy hierarchy, and rendered proof.
First-principles product innovation audit for product bets, workflows, AI agents, marketplaces, developer tools, enterprise apps, differentiation, kill/ship decisions, and 10/10 marketability.
Analyze sessions or projects for patterns, misses, product signals, process improvements, automation opportunities, and skill effectiveness. Modes: session, project, portfolio.
Run opaque or custom releases to production with a ledger, artifact decision, deploy proof, smoke checks, rollback path, ETA updates, and handoff.
Integrate @tangle-network/sandbox SDK without rebuilding stream durability, session replay, browser-safe clients, or idempotent dispatch already provided by the platform.
Run Semgrep static analysis for security findings. Supports important-only or full scans, Semgrep Pro when available, merged SARIF, triage, and remediation plans.