| name | pony-software-design |
| description | Disciplines for software design work. Load when designing APIs, type systems, features, or system boundaries. Counters the tendency to retrieve familiar patterns instead of discovering what the problem actually needs. Has full (8-persona) and lightweight (5-persona) modes. |
| disable-model-invocation | false |
Software Design
Load this skill when doing design work — APIs, type systems, features, system
boundaries. The core problem it addresses: LLMs default to retrieving
familiar patterns from training data rather than discovering what a specific
problem needs. The result is designs that look right (they have the right nouns)
but weren't derived from the problem.
Design is the act of discovering what is needed. It's about finding surprising
affordances and avoiding candy-machine interfaces riddled with footguns.
Design is not a phase. It's a continuous loop: observe what you have, orient
against what you know, decide, act, then observe again. Every decision you make
reveals new information. That information might confirm prior decisions or
invalidate them. Either way, you have to look. "I already decided that" is never
a reason to skip re-evaluation — it was decided with less information than you
have now.
Mode selection
The skill has two modes: full and lightweight. The orchestrator
selects the appropriate mode based on the criteria below and proceeds.
Report the mode choice when presenting results — the human can request
full mode if lightweight was used and they want deeper coverage.
Full mode is the default. Use it when:
- Defining new boundaries — a new subsystem, API surface, or type hierarchy
- Multiple ownership or abstraction decisions need to be made simultaneously
- There's genuine design space to explore — multiple viable approaches
- The design must be described from first principles, not as an extension of
something that already exists
Lightweight mode is for bounded design work within established patterns:
- Adding a method to an existing API, a new variant in a type family, a new
handler following established conventions
- The boundaries are already decided — you're filling in, not redrawing
- Consumer code patterns are already established by adjacent features
- The task can be described as "another X that does Y" where X already exists
in the codebase
When in doubt, use full mode. Lightweight is appropriate when there's a
clear existing pattern being extended and the boundaries are already
decided.
Process: full mode
Design work is where pattern-matching failures are most costly and hardest to
self-detect. A single agent applying design disciplines will still
pattern-match — the disciplines become post-hoc rationalizations for a
retrieved design rather than actual constraints on the exploration.
The full design process runs in two stages with a feedback loop. Load
pony-ensemble for the mechanical process; the personas defined in this skill
replace the generic attention focuses.
Relationship to Ensemble Workflow
This skill uses the ensemble workflow with domain-specific customizations.
Stage 1 (design) runs as a standard ensemble with 3 personas. Stage 2
(evaluation) runs as a second ensemble with 5 personas, using the Stage 1
synthesis output as its input. The two-stage loop and finding categorization
(Rejection/Adjustment/Tension) are additions specific to this skill — the
base ensemble protocol handles agent spawning, triage, and synthesis
mechanics.
Orchestrator pre-spawn: understand the problem
Before spawning design personas, the orchestrator must decompose the problem
statement. Ask: "What are we trying to accomplish? What pain point does this
address?"
If the problem statement implies a solution rather than stating a problem —
"add convenience wrapper methods for X" rather than "users struggle with X
because Y" — peel it back to the underlying pain point. If you can't
confidently identify the pain point, ask the human. This is not a failure;
it's the most valuable question the orchestrator can ask, because every
persona will anchor on whatever framing they receive.
Brief personas with the actual problem. The original problem statement can be
included as context, but the briefing should lead with the pain point: what the
user is struggling with and why. This gives personas room to explore different
solutions rather than anchoring on the solution implied by the problem
statement.
Stage 1: Design
Three design personas explore the problem space in parallel. Each applies all
the disciplines below but enters the problem from a different direction. The
decorrelation comes from where they start, not what they know. Persona
definitions are in personas/design/.
| File | Focus |
|---|
consumer-first.md | Starts from usage code, derives API from what makes call sites clean |
skeptic.md | Questions every abstraction, tries to subtract, proposes the smallest design |
principle-checker.md | Hard verification of each design principle with evidence |
Orchestrator post-return: evaluate exploration quality
When personas return, the orchestrator evaluates the quality of exploration
before forwarding outputs to synthesis. Check:
- Did the personas explore meaningfully different approaches? Read each
persona's "Key decisions / alternatives considered" section. If a persona
only considered one interpretation of the problem and designed for it, the
exploration was insufficient.
- Did they all anchor on the same narrow interpretation? If all three
personas produced designs based on the same framing of the problem — e.g.,
all three designed wrapper methods — the ensemble failed to decorrelate.
The different entry points weren't enough to overcome the anchoring effect
of the problem statement.
- Is the reasoning for rejecting alternatives sound? A persona that
considered a wild idea and rejected it with good reasoning explored
genuinely. A persona that listed an alternative and dismissed it in one
sentence didn't.
If exploration quality is poor — narrow interpretation, insufficient
alternatives, weak rejection reasoning — don't forward to synthesis. Refine
the briefing: restate the problem more broadly, point out the anchoring you
observed, and ask the human to clarify if needed. Then re-spawn.
Stage 1 synthesis
Stage 1 synthesis produces a candidate design using the standard ensemble
synthesis process. The synthesis should pay special attention to:
- Where the consumer-first designer's sketches conflict with the skeptic's
subtractions — the tension usually reveals the right boundary
- Where the principle checker found a violation that the others missed —
this is the highest-value finding
- Whether all three converged on the same abstraction — convergence from
different starting points is strong signal
- When the skeptic says "no value" and redirects: The skeptic is the only
persona whose job includes saying "this shouldn't exist." When the skeptic
concludes the proposed design doesn't earn its keep and redirects toward a
different problem, this is not a minority position to be outvoted because
the other two produced designs. The other two will always produce a
design — that's their job. Evaluate the skeptic's case on its merits. If
the skeptic is right that the design is a thin wrapper, the synthesis
output should adopt the skeptic's redirect as the foundation: "the
proposed design doesn't earn its keep — here's the actual problem worth
solving" becomes the candidate, not a blended version of the thin wrapper.
Stage 2: Evaluation
Five evaluation personas stress-test the candidate design in parallel. Their
input is the Integrated Result from Stage 1 synthesis — the candidate design
with its consumer sketches, type definitions, and boundary decisions. They
evaluate these design artifacts, not implementation code. Persona definitions
are in personas/evaluation/.
| File | Focus |
|---|
security.md | Trust boundaries, attack surfaces, resource bounds in the design |
performance.md | Architectural bottlenecks, coordination points, data structure choices |
adversarial.md | Concrete usage scenarios that lead to bad outcomes |
testability.md | Whether the design is verifiable — observable effects, isolatable components |
wildcard.md | What all the other personas missed |
For the wildcard persona specifically: include the identity statement (first
paragraph) from each of the other personas so the wildcard knows what
territory is already covered.
Before spawning evaluation personas, create a temporary directory for evidence
files (~/tmp/design-eval-<timestamp>/). Each persona writes its detailed
analysis to a file in this directory and returns a structured summary to the
orchestrator. The synthesizer works from summaries and digs into evidence
files only when it needs to examine a finding more closely. This prevents
context overload during synthesis.
Evaluation personas identify problems and assess impact — they do not
categorize their own findings as Rejection/Adjustment/Tension. Categorization
is the synthesis step's responsibility.
Evaluation persona output format
Each evaluation persona produces two artifacts:
Evidence file — written to the path provided by the orchestrator. Contains
the full detailed analysis: every finding with complete evidence, full design
element excerpts, detailed reasoning, and complete pass/fail evaluations. This
is the authoritative record.
Summary (returned to orchestrator) — a structured summary for the
synthesizer to work from:
Findings — ordered by impact (Structural > Significant > Minor). Each:
- Design element: The type, boundary, API, or interaction being evaluated
- Concern: What the problem is
- Impact: Structural (requires rethinking the approach), significant
(requires notable changes to the candidate), or minor (a local
change; the rest of the candidate is unaffected)
- Evidence: Brief — full evidence is in the file
- Suggested change: Write it whenever you can. A concern you can name but
cannot resolve is still a finding — say so, and say what blocks you.
Impact is what the finding costs the candidate: how much of it has to
change. It is not permission to leave the finding alone — report a minor
finding exactly as you report a structural one, and stage 2 synthesis categorizes
every finding, at every impact level.
The impact assessment helps the synthesizer with categorization without
pre-empting it. A persona's "structural" assessment is a strong signal toward
Rejection, but the synthesizer may disagree if it sees the concern addressed
by another persona's suggestion.
Passes — things checked that look correct. Brief.
Uncertainties — things the persona couldn't determine, and why.
Stage 2 synthesis works from the persona summaries and categorizes each
finding. Provide the paths to each persona's evidence file so the synthesizer
can dig in when it needs more context — when impact assessments conflict, when
a finding's summary is ambiguous, or when it needs to verify the evidence
supports the concern.
Stage 2 synthesis categorizes each finding:
- Rejection: A structural problem that invalidates the design direction.
The candidate cannot be fixed by adjustment — the design personas need to
rethink the approach. The rejection includes why the direction fails and
what constraint it violates.
- Adjustment: A specific aspect that needs to change, but the overall
direction is sound. Becomes a constraint for the next design iteration.
- Tension: A fundamental conflict that the personas cannot resolve — it
requires human judgment. Collected and presented at the end.
The loop
After stage 2 synthesis:
- If there are only tensions (no rejections or adjustments), the loop
terminates. Present the design with the tensions for human review.
- If there are adjustments and/or rejections, feed them back to the design
personas. Each design persona receives: the prior candidate design, the
original problem statement, and the categorized findings. Rejections include
the rationale for why the direction failed — design personas should explore
a different approach, not patch the rejected one. Adjustments include the
specific change needed and why — design personas should revise the candidate
to incorporate them as constraints.
- The design personas run again with this context, producing a revised
candidate.
- Evaluation runs again on the revised candidate. Evaluation personas run with
fresh context (no knowledge of prior evaluations). The synthesis step
receives the full history so it can track convergence.
- Repeat until clean or until convergence failure is detected.
Convergence failure
The orchestrator monitors the loop for signs that it isn't converging:
- The same evaluation concern keeps appearing across iterations, even after
design revisions attempt to address it
- Rejections and adjustments are contradicting each other (fixing one
evaluation concern breaks another)
- The design is growing more complex with each iteration rather than settling
When a convergence failure is detected, the orchestrator stops the loop and
escalates to the human: "Here's the fundamental tension — these concerns pull
in opposite directions, and we need you to decide which matters more." This is
not a failure of the process. Surfacing genuine tensions is one of its primary
outputs.
Output
The final output includes:
- Accepted design (if one emerged): The candidate that passed evaluation,
with consumer sketches, type definitions, and boundary decisions.
- Rejected designs: Each candidate that was rejected during the loop, with
the rejection rationale. These are valuable — they document explored
territory and why it didn't work.
- Unresolved tensions: Findings categorized as tensions that require human
judgment.
If the loop terminated via convergence failure rather than a clean evaluation,
there is no accepted design — only the history of attempts, the rejections, and
the tensions.
Process: lightweight mode
Lightweight mode uses fewer personas and a single pass. It keeps all three
design personas but reduces evaluation to two personas and drops the
feedback loop. Load pony-ensemble for the mechanical process.
Orchestrator pre-spawn: understand the problem
Same as full mode. Before spawning personas, decompose the problem statement
to the underlying pain point. If the problem statement implies a solution,
peel it back or ask the human. Brief personas with the actual problem.
Stage 1: Design
Three design personas explore the problem in parallel — the same as full
mode. The same disciplines apply; the decorrelation still comes from
different entry points.
| File | Focus |
|---|
consumer-first.md | Starts from usage code, derives API from what makes call sites clean |
skeptic.md | Questions every abstraction, tries to subtract, proposes the smallest design |
principle-checker.md | Hard verification of each design principle with evidence |
Orchestrator post-return: evaluate exploration quality
Same as full mode. Evaluate whether personas explored meaningfully different
approaches before forwarding to synthesis. If all three anchored on the same
narrow interpretation, re-brief and re-spawn.
Stage 1 synthesis produces a candidate design using the standard ensemble
synthesis process. The same synthesis guidance as full mode applies:
consumer-first vs skeptic tensions reveal boundaries, principle-checker
violations the others missed are the highest-value findings, and
convergence from different starting points is strong signal.
Stage 2: Evaluation
Two evaluation personas stress-test the candidate. The adversarial evaluator
always runs. The second slot is context-dependent — if the human specifies
which evaluator to use, use that. Otherwise the orchestrator picks
whichever lens is most relevant to the task:
| When | Pick |
|---|
| Design touches trust boundaries or external input | security.md |
| Design is on a hot path or introduces coordination points | performance.md |
| Design has complex state or will be hard to test in isolation | testability.md |
Pick whichever is closest — every design has some risk profile. If
multiple conditions apply, pick the most relevant one. If the reason
for picking a particular evaluator is a characteristic that also
appears in the full-mode selection criteria, that's a signal the task
warrants full mode — don't use the persona pick to compensate for a
wrong mode selection.
Before spawning evaluation personas, create a temporary directory for
evidence files (~/tmp/design-eval-<timestamp>/), same as full mode.
Each persona writes its detailed analysis to a file in this directory and
returns a structured summary. Evaluation personas use the same output
format as full mode (evidence file + summary with findings ordered by
impact, passes, uncertainties).
Stage 2 synthesis categorizes each finding as Rejection, Adjustment, or
Tension using the same scheme as full mode.
No loop — single pass
Lightweight mode does not iterate. After Stage 2 synthesis:
- Adjustments and tensions: Present the design with findings to the human.
Adjustments are expected to be small enough that the orchestrator or human
can apply them directly. If adjustments collectively amount to redesigning
rather than tweaking, that's the same escalation signal as high finding
density — present it to the human.
- Rejection: The design direction is wrong. Present the rejected
candidate, the rejection rationale, and all other findings to the human.
The human decides what to do — escalate to full mode, fix it directly,
rethink the problem statement, or something else. Lightweight doesn't
prescribe the response; it presents the information.
If the review produces an unexpectedly high density of findings relative to
the change size, if a finding reveals the approach is fundamentally wrong,
or if a finding reveals the change touches more subsystems or has more
complex interactions than the mode selection assumed, the orchestrator
presents this to the human. The human decides what to do — the same options
apply.
Output
The final output includes:
- Candidate design: With consumer sketches and type definitions.
- Adjustments: Specific changes needed, small enough to apply directly.
- Tensions: Conflicts requiring human judgment.
- Rejection rationale: If applicable — the structural finding and why
the design direction was rejected.
Design values
When principles conflict, these values set the priority. We value the left side over the right — but the right side still matters when the left isn't at stake.
API safety over API minimality — an error-prone API should be fixed even if the fix adds surface. Prefer solutions that don't expand the API, but never leave a footgun to preserve minimality.
Correctness over performance — never sacrifice correctness for speed. Get it right first, then optimize. A faster wrong answer is still wrong.
Correctness over concision — correct but verbose beats concise but wrong. Don't simplify code or APIs at the cost of correct behavior.
Security over performance — never skip validation at trust boundaries for speed. Optimize how you validate, not whether you validate. Security is correctness.
Interface simplicity over implementation simplicity — accept a harder implementation to give users a clean interface. The consumer's experience matters more than the implementer's convenience.
Performance over interface simplicity — runtime speed matters more than programmer convenience. It's acceptable to make things harder on the user to improve performance, but never at the cost of correctness.
Simplicity over consistency — don't force artificial consistency when it makes things harder to use. If two similar things genuinely need different interfaces, let them be different.
Explicitness over implicitness — when the language allows something to work by magic (implicit conversions, convention-based wiring, unnamed dependencies), prefer the version that states what's happening. The cost of a few extra characters is less than the cost of reconstructing hidden knowledge.
Type safety over convenience — use the type system to encode constraints even when it's more work. Distinct types for distinct semantics, validated wrappers over raw primitives, explicit error vocabularies over generic errors. "We could just use a String here" is almost always wrong.
Changeability over predictive design — make designs modular and replaceable so future needs can be accommodated, but don't add abstractions, extension points, or features for changes that haven't happened yet. Easy to modify beats designed for a specific predicted modification.
The disciplines
These are the foundation each persona builds on. Every agent applies all of
them.
Start from the problem, not the solution
State what problem the user has before proposing any types, traits, or APIs.
"The user needs X" comes before "here's a SessionStore trait." If you can't
articulate the problem without referencing your solution, you don't understand
the problem yet.
Explore before committing
If you only have one idea, you don't have any. The first interpretation of a
problem statement is pattern retrieval — the LLM equivalent of "this looks like
a thing I've seen before." Design starts when you generate a second
interpretation that's genuinely different, not a variation on the first.
Before committing to a direction, explore the design space. Generate multiple
approaches internally — different framings of what the problem is asking for,
not just different implementations of the same framing. Include wild ideas.
They're valuable not because you'll use them, but because they reveal what
matters about the problem. An idea you reject teaches you something about why
you rejected it — that "why" is design knowledge.
Present your best idea, not your first idea. The output is the winner of an
internal competition. The ensemble output format has a "Key decisions:
alternatives considered" section — use it to document what you explored, what
you picked, why, and why you rejected the alternatives. This isn't bookkeeping;
it gives the orchestrator visibility into whether real exploration happened.
The problem statement is where exploration starts, not where it ends. "Add
convenience wrapper methods" is a hint about a pain point, not a design
specification. What pain? Why is the current approach painful? What would
"convenience" actually mean to the user? Different answers to those questions
lead to genuinely different designs — one of which might be "thin wrappers"
and another might be "a unified API that abstracts away the underlying
complexity." Both are valid interpretations of "convenience"; only exploration
reveals which one earns its keep.
Sketch consumer code first
Before designing any API, write the code that uses it. The handler, the
call site, the configuration — the actual application code a user would write.
Not as an afterthought example, but as the first artifact. The consumer sketch
is the specification. It reveals:
- What the API actually needs to provide (and what it doesn't)
- Where type safety breaks down (runtime casts, stringly-typed maps)
- Whether the abstraction can actually serve its purpose (can middleware do
async work? can the handler access typed data?)
- What the error paths look like from the consumer's perspective
If the consumer code is awkward, the API is wrong. Fix the API, not the
consumer code.
When claiming consistency between two APIs (e.g., "guards use the same API
as handlers"), write both consumer sketches side by side. If the method names,
signatures, or interaction patterns differ, the claim is false — address the
discrepancy before proceeding.
Inventory before inventing
Before proposing a new type, trait, or abstraction, write down what already
exists that addresses the same need: in the codebase, in the language's stdlib,
in the ecosystem. If nothing exists, say so explicitly. If something exists,
start from it — extend, adapt, or compose it rather than building a parallel
structure.
On a greenfield project, "what already exists" means the language's built-in
types, stdlib, and idioms. A new type that duplicates what the language provides
is a smell.
This is not "reuse for reuse's sake." It's a forcing function against the
pattern-matching tendency to invent new abstractions when the problem doesn't
require them.
Build up incrementally
Don't design the whole thing at once. Start with the smallest coherent piece.
Validate it with a consumer sketch. Then add the next piece and see if it fits.
At each step, ask:
- Does the new piece fit naturally with what's already there?
- Or is it fighting the existing design?
- If it's fighting, is the problem upstream? Would a different foundation make
this piece fit naturally?
This is how you discover the shape of the problem. A big-bang design papers
over these tensions. Incremental exploration surfaces them while they're cheap
to fix.
Every step changes what you can see
Every design decision is made with incomplete information. As you explore
further, the territory expands. At step B you could see one option and picked
it. At step D you can see two options that weren't visible from B. Step C might
provide evidence that the option you didn't pick is actually better. None of
this means B was wrong at the time — it means the landscape changed and you need
to look again.
The way you discover this is by constantly pushing on the design: "what if we
did it this other way?" "Does this conform to our principles?" "This doesn't
feel right — why?" These aren't idle questions. They're how you explore more of
the map. Each question might reveal new options, new evidence, or new
connections between decisions you thought were independent.
When the landscape changes, trace back through prior decisions. Not to check
whether they were "disproved" — that's too binary. Check whether the option
space has expanded. Maybe at step B you picked X because it was the only option
you could see. Now at step D you can see X and Y. Which is better given
everything you've learned? Maybe X is still right. Maybe Y is clearly better.
Maybe you need to explore further to tell. All three of those outcomes require
you to actually go back and look rather than assuming B is settled.
This is the core of the design loop. Skipping it is how you end up with designs
that look coherent on the surface but have quiet contradictions baked in — an
app-level-only registration model sitting next to a radix tree that already
knows the route hierarchy, because nobody went back to check whether "app-level
only" still made sense after discovering what the tree could do.
The cost of revisiting decisions is real but bounded. The cost of building on a
decision that should have been revisited compounds with every step forward.
Question every abstraction
For each type, trait, or interface in your design, ask: is this here because the
problem requires it, or because other systems have it? "Sessions usually have a
SessionStore" is not a reason. "The framework needs to persist session data" is
a reason — but only if the framework actually needs to own that responsibility.
The strongest signal that you're importing rather than discovering: your design
has the same nouns as Rails/Phoenix/Express/Django and you're working in a
language with fundamentally different idioms.
Name things precisely
Names are the primary interface of a design — they're how it communicates
intent to every future reader. A design that's structurally sound can still
fail in practice because a name suggested something different from what the
thing actually does, and the user built a wrong mental model from it.
This discipline earns its keep at boundaries: public API types, method names,
parameter names, module names — anywhere a user encounters a name and forms an
expectation about what it means. Internal names matter too, but the cost of a
misleading public name is much higher because it shapes how every consumer
understands the system.
At each design step, ask:
- Does the name come from the problem domain or the solution domain?
Problem-domain names ("invoice," "route," "subscription") connect the code
to what users already understand. Solution-domain names ("manager,"
"handler," "processor") describe implementation roles that tell the user
nothing about what the thing actually does. Prefer problem-domain names for
types the user interacts with; reserve solution-domain names for internal
machinery where they're the clearest description of the role. Note that some
of these words become problem-domain vocabulary in specific contexts —
"handler" is the natural term in web frameworks and event systems. The
concern is with using them as generic suffixes that avoid naming what the
thing actually handles or manages.
- Does the name describe what this thing does, or just what category it
belongs to? "Validator" says it validates — but what?
"InvoiceAmountValidator" says what it validates. "Store" says it stores —
but a
UserStore that also sends email notifications lies about its
responsibilities. A name that categorizes without specifying is a name that
lets scope creep in without anyone noticing.
- Could the name mislead a user about what this thing does? If someone
reading only the name would form a wrong expectation about the behavior, the
name is wrong. This includes names that are accurate but incomplete — a
Cache that also writes through to the database is misleading because the
name suggests read-only caching.
- Are there names that sound similar but mean different things?
Session
and SessionState, Token and TokenData, Route and Router — when
names differ by a suffix or prefix, users will assume the relationship is
systematic. If SessionState is not the state of a Session, the naming
implies a relationship that doesn't exist.
- Are there names that sound different but mean the same thing? If the
design uses
user in one place and account in another for the same
concept, readers will assume these are different things and spend effort
trying to understand the distinction. One concept, one name — everywhere.
This is about vocabulary consistency, not boundary-qualified variants like
UserInput and UserRecord — those are distinct types that need distinct
names per "Distinguish values with distinct semantics."
"Distinguish values with distinct semantics" ensures different concepts get
different types. This discipline ensures those types get names that communicate
what they actually are. A type can be correctly distinct and still misleading
if its name suggests the wrong thing.
When this discipline pushes for more precise names and the skeptic questions
whether the named concepts need to exist at all, surface the tension — both
are needed. Good names make abstractions easier to evaluate: a precisely named
type reveals its purpose, which either justifies or undermines its existence.
Reason about ownership boundaries
For every capability in the design, ask: does the framework/library own this, or
does the user own this? The answer should come from analysis of the consumer
sketch, not from "frameworks usually own this."
The test: if you removed this from the framework and the user did it themselves,
would anything break? Would anything get worse? If the user can do it better