Skip to main content

design-mcp-server

Design the tool surface, resources, and service layer for a new MCP server. Use when starting a new server, planning a major feature expansion, or when the user describes a domain/API they want to expose via MCP. Produces a design doc at docs/design.md that drives implementation.

설치로 이동

소스 정보

저장소
cyanheads/obsidian-mcp-server
최근 소스 활동
2026년 9월 13일 18:16
감지된 SKILL.md 언어
영어
스타
679
포크
101

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
design-mcp-server
description
Design the tool surface, resources, and service layer for a new MCP server. Use when starting a new server, planning a major feature expansion, or when the user describes a domain/API they want to expose via MCP. Produces a design doc at docs/design.md that drives implementation.
metadata
{"author":"cyanheads","version":"2.25","audience":"external","type":"workflow"}
## When to Use - User says "I want to build a ___ MCP server" - User has an API, database, or system they want to expose to LLMs - User wants to plan tools before scaffolding - Existing server needs a new capability area — two or more tools sharing a new noun or service, or any new upstream source (design the addition, not just a single tool) Do NOT use for a single tool on an existing noun — use `add-tool` directly. ## Inputs Gather before designing. Ask the user if not obvious from context: 1. **Domain** — what system, API, or capability is this server wrapping? Or is the server providing internal capability with no external dependency (computation, text/code utilities, in-memory state)? 2. **Data sources / source of truth** — APIs, databases, file systems, external services? Or is the server itself the source (in-memory state, pure computation, local-only utility, embedded model)? 3. **Target users** — what will the LLM (and its human) be trying to accomplish? 4. **Scope constraints** — read-only? write access? admin operations? what's off-limits? If the domain has a public API, read its docs before designing. For internal-only servers, skip API research and go straight to user goals. Don't design from vibes either way. ### Server scope and audience Before committing to a server boundary, answer: **what workflow does this server serve, and who is the audience?** The unit of a server is a *user workflow*, not an API. A single rich API can earn its own server when the audience is large and the API surface supports a full workflow (PubMed for literature research, SEC EDGAR for financial analysis, Shodan for internet-wide device intelligence). Multiple APIs should collapse into one server when they serve the same workflow from different angles — a "threat intelligence" server that aggregates VirusTotal, AbuseIPDB, and GreyNoise is more useful than three separate servers because the user's goal is "assess this indicator," not "query VirusTotal." **Don't default to one-API-one-server.** That's the right call when the API is deep enough and the audience is large enough, but it's not the starting point. The starting point is the workflow: | Signal | Server boundary | |:-------|:----------------| | Single API with rich surface, large audience | Standalone server named for the platform (`pubmed-mcp-server`, `secedgar-mcp-server`) | | Multiple APIs serving the same workflow | One server named for the workflow (`threat-intel-mcp-server`), APIs are internal sources | | Domain with distinct sub-audiences | Consider splitting — a pentester and a SOC analyst have different workflows even in the same domain | | Pure computation, no external deps | Standalone server named for the capability (`calculator-mcp-server`, `pentest-mcp-server`) | When multiple APIs collapse into one server, the tool surface is organized around what the user is doing, not which API gets called. The agent says "investigate this domain" and the server routes to the best available source internally. Individual APIs become service-layer implementation details, not tool-surface identities. ## Server Naming Usually settled before this skill runs; confirm it passes one test before designing against it. A name is earned when someone reading only the name — an npm result, a marketplace grid — forms no wrong expectation about scope. Name a wrapper for its source (`pubmed-mcp-server`, `secedgar-mcp-server`); a generic domain name is earned only by aggregating independent sources (`threat-intel-mcp-server`), and a single-jurisdiction scope goes in the name (`uk-legislation-mcp-server`). Spell out an acronym that reads as something else (`libofcongress`, not `loc`), or pair it with its domain (`eia-energy`, `bls-labor`). The tool prefix is a separate, stable identifier (see the Design table) — it need not move when the package is renamed, and renaming it is a breaking change for every client. ## Steps ### 1. Research External Dependencies **Applies when:** the server wraps an external API or service. Skip for internal-only servers (computation, local file ops, in-memory state, code analysis utilities) and jump to Step 2. Before designing, verify the APIs and services the server will wrap. Read the docs, then **hit the API** — real requests reveal what docs omit. Research inline by default — fetch docs, read SDK readmes, confirm assumptions before committing them to the design. For each external dependency: - Fetch API docs, confirm endpoint availability, auth methods, rate limits - Check for official SDKs or client libraries (npm packages) - Note any API quirks, pagination patterns, or data format considerations When research is genuinely parallelizable (multiple independent APIs, several SDKs to evaluate), spawn background agents for the independent legs while you proceed with domain mapping. Skip the overhead for a single API — just read it yourself. **Live API probing.** After reading docs, make real requests against the API to verify assumptions: - **Response shapes** — confirm actual field names, nesting, and types. Docs frequently lag or omit fields. - **Batch/filter endpoints** — look for `filter.ids`, bulk GET, or query-by-multiple-IDs patterns. A single batch request replaces N individual fetches and eliminates serial-request bottlenecks and rate-limit accumulation. - **Field selection** — check if the API supports `fields` or `select` parameters to request only the data you need. This reduces payload size dramatically for large objects. - **Pagination behavior** — verify token format, page size limits, and what happens when results exceed one page. - **Error shapes** — trigger real 400/404/429 responses to see the actual error format, not just what docs claim. - **Unknown-param behavior** — send one deliberately misspelled parameter. If the API silently ignores it (plausible but unfiltered results instead of an error), every typo'd or unverified param name becomes a silent-wrongness bug — the service layer then needs a strict allowlist of confirmed spellings, and new filters require a probe before they ship. - **Omission semantics** — for each major optional parameter, check what omitting it actually returns. Some APIs default to the intuitive scope; others silently widen (all historical versions, all statuses, global instead of regional). A default that changes result *meaning* becomes a server-side default plus an echoed output field, not something left to the agent. **Stopping condition:** at minimum, probe one list/search endpoint, one single-item GET, one error case (force a 404 or 400), and one unknown-param request. For large APIs with many resource types, add one probe per major noun. Stop when the response shapes and error envelope are confirmed. This step prevents building a service layer against assumed response shapes that don't match reality. ### 2. Map User Goals, Then Domain Operations Start with **user goals**, not endpoints. Enumerate the outcomes an agent (and its human) will actually try to accomplish with this server — usually 3–10, scaled to domain size. These drive the workflow tools that form the spine of the surface. Endpoint-inventory-first design produces 1:1 API mirrors; goal-first design produces tools agents reach for. For internal-only servers, goals map to capabilities rather than endpoints — e.g., "format markdown to GFM," "tokenize text by model," "compute file hash." Example user goals for a project management server: - Find tasks I'm assigned to that are due soon - Create a task in a project, assign it, and notify the owner - Mark a task complete and log the outcome - Audit a project's overdue work Then enumerate the underlying **domain operations** the system supports, grouped by noun. These are the raw material workflow tools compose and single-action tools back-fill where workflows don't cover an edge case. | Noun | Operations | |:-----|:-----------| | Project | list, get, create, archive | | Task | list (by project), get, create, update status, assign, comment | | User | list, get current | The user-goal list shapes the tool surface; the operation list fills in the gaps. Not every operation becomes a tool — an operation stays as raw material (not its own tool) when it's already fully covered by an existing tool's output, or when the only agents who'd use it are in scenarios outside this server's stated purpose. ### 3. Classify into MCP Primitives **Tools are the primary interface.** Not all MCP clients expose resources, and those that do rarely surface one to the model without a human selecting it. Design the tool surface to be self-sufficient: an agent with only tool access should be able to do everything the server is built for. Resources add convenience for clients that support them (injectable context, stable URIs), but are not a reliable access path. | Primitive | Use when | Examples | |:----------|:---------|:--------| | **Tool** | The default. Any operation or data access an agent needs to accomplish the server's purpose. | Search, create, update, analyze, fetch-by-ID, list reference data | | **App Tool** | **Rare — default to a standard tool.** Only when a human will actively interact with the result in real time *and* the target client supports MCP Apps. Most clients are tool-only and most agent workflows are read-by-LLM, not viewed-by-human. App tools add an iframe + CSP, `app.ontoolresult`/`callServerTool` plumbing, host-context wiring, and a `format()` text twin that still has to be content-complete (since most clients only see that). Two surfaces to keep in sync, two failure modes per change. | Dense tabular state a human scrubs through; form-based human approval in an MCP Apps-capable client | | **Resource** | *Additionally* expose as a resource when the data is addressable by stable URI, read-only, and useful as injectable context. | Config, schemas, status, entity-by-ID lookups | | **Prompt** | Reusable message template that structures how the LLM approaches a task | Analysis framework, report template, review checklist | | **Client round-trip** | Not a registered primitive — a handler *returns* `ctx.requestInput(...)` to ask the client for what only it has, and is re-entered with the answer: a confirmation or form (`inputRequired.elicit`), an authorization or hosted-form URL (`inputRequired.elicitUrl`), the client model's judgment (`inputRequired.createMessage` — borrow the caller's model rather than bundling one), filesystem roots (`inputRequired.listRoots`). Design it into the tool that needs it; see Workflow tool safety and `api-context`. | Destructive-arm confirmation, OAuth consent, summarize-with-the-client's-model | | **Neither** | Internal detail, admin-only, not useful to an LLM | Token refresh, webhook setup, migrations | What the tool surface needs to cover depends on the server: a read-only research server has different economics than a CRUD project management server. Consider the domain, the expected agent workflows, whether it wraps one API or many, and what data relationships exist. **Common traps:** - **Data locked behind resources**: If something an agent needs is only accessible via a resource, it's invisible to tool-only clients. That data might warrant its own tool, or it might already be covered by an existing tool's output — but it needs a tool path somewhere. - **CRUD explosion**: Don't map every REST endpoint to a tool. Related operations on the same noun often belong in one tool with an `operation`/`mode` parameter (see Step 4). - **1:1 endpoint mirroring**: API endpoints are designed for programmatic consumers. LLM tools should be designed for workflows — what an agent is *trying to accomplish*, not what HTTP calls happen under the hood. **Irreversible operations stay in the UI.** The "Neither" bucket above covers operations that aren't useful to an LLM. There's a second, sharper reason to exclude something from the tool surface: operations whose failure mode is catastrophic and unrecoverable. Examples span domains — dropping a production database table (data loss across every row), force-emptying a versioned cloud-storage bucket (no recovery once the lifecycle policy fires), revoking the workspace's last admin role (locks everyone out, recovery requires vendor support), GDPR permanent-delete on a customer profile (un-restorable by design), purging an analytics warehouse partition older than the retention window (auditable history gone), or deleting the single audience on a free-plan email platform (nukes every subscriber and historical report in one call). These are useful to an LLM *in principle*, but the blast radius of a mis-call is disproportionate to any agent workflow. Humans do these in the vendor UI, where confirmation dialogs and undo paths exist. Agents shouldn't have the tool at all. This is distinct from `destructiveHint` — that annotation is for operations that are destructive but recoverable (deleting a task, reverting a commit) and agents should still have them. The "stays in the UI" line applies only to operations whose failure is both catastrophic *and* irreversible. ### 4. Design Tools This is the highest-leverage step. Tool definitions — names, descriptions, parameters, output schemas — are the **entire interface contract** the LLM reads to decide whether and how to call a tool. Every field is context. Design accordingly. #### Tool shapes you'll encounter Most tools follow the `{server}_{verb}_{noun}` default — one focused responsibility, one clear verb, often (but not always) one upstream call. API-wrapping examples: `pubmed_search_articles`, `pubmed_fetch_articles`. Internal-only examples: `markdown_format_text`, `regex_test_pattern`, `tokens_count_text` — same naming convention, no external dep. Two variants warrant explicit design pressures of their own: | Shape | Purpose | Typical form | Examples | |:------|:--------|:-------------|:---------| | **Workflow** | Multi-step orchestration that replaces a common agent chain | N upstream calls (often parallelized); may request confirmation; may need mid-flow cleanup | `clinicaltrials_find_eligible` (search → filter → rank) | | **Instruction** | State-aware procedural guidance — advice, not action | Static markdown + a few live-state fetches, `readOnlyHint: true`, outputs `nextToolSuggestions` pre-filling the recommended follow-up. No writes. | `git_wrapup_instructions` | | **Reference** | Decode opaque domain vocabulary — codes, enums, identifier formats, coverage windows — so agents can build valid inputs for the rest of the surface | Static tables or one cached fetch (often zero upstream calls); consolidate N lists under one `topic` enum; `readOnlyHint: true`, `openWorldHint: false` when offline | `medcode_list_systems`, `osv_list_ecosystems` | These aren't boxes every tool must fit into — some blend shapes — but the design pressures differ enough that naming them helps avoid re-discovering the patterns per server. #### Think in workflows, not endpoints The unit of a tool is a *useful action*, not an API call. Ask: "What is the agent trying to accomplish?" — not "What endpoints does the API have?" A single tool can call multiple APIs internally, apply local filtering, reshape data, and return enriched results. The LLM doesn't know or care about the underlying calls. ```ts // Workflow tool — search + local filter pipeline, not a raw API proxy const findEligible = tool('clinicaltrials_find_eligible', { description: 'Match a patient profile to eligible clinical trials, filtering by age, sex, conditions, location, and healthy volunteer status. Results are ranked and carry a per-study eligibility explanation.', // handler: listStudies() → filter by eligibility → rank by location proximity → slice }); ``` > **Tip — mode consolidation.** When a tool has several related operations on the same noun, you can consolidate them under one tool with a `mode`/`operation` enum. This affects both naming (noun-led, e.g., `github_pull_request`) and handler design (dispatch by mode). Use when it tightens the surface; skip when ops diverge enough to warrant separate tools. When the arms need *different* required fields (look up by ID vs. search by name), declare the input as `z.discriminatedUnion('mode', [...])` rather than making every field optional and checking the combination by hand — each arm advertises its own `required`, the handler narrows on the discriminator, and mixed arguments are rejected. Two constraints: `output` stays a flat `z.object`, and a union root rules out `headerParam`. See `add-tool` § *Multi-mode tools*. #### Multi-source tools and fallback chains **Applies when:** a server aggregates multiple data sources for the same workflow, and the "best" source varies by input type, availability, or coverage. Skip for single-API servers. When a tool's goal can be served by multiple sources, design it as a **multi-source tool** — the agent calls one tool, the handler routes to the best source (or fans out to several) internally. This is the difference between a "PubMed wrapper" and a "literature research server": a hypothetical `literature_search_articles` tries PubMed first, falls back to EuropePMC for broader coverage, then Unpaywall for open access. The agent doesn't choose which API to hit — the server makes that decision based on what works. Two patterns: **Source fallback chains** — try sources in priority order, fall through on failure or empty results. Best when sources cover the *same corpus* with different depth or availability. The output should indicate which source provided the data so the agent (and human) can assess provenance. When the fallback changes what is being searched — a different corpus, different identifiers, different licensing — don't chain: expose the second source as a sibling tool so the agent chooses the corpus knowingly (the shipped `pubmed-mcp-server` keeps `pubmed_europepmc_search` separate for exactly this reason). **Multi-source fan-out** — query multiple sources in parallel, merge results. Best when sources provide complementary data about the same entity. Use `Promise.allSettled` so one failing source doesn't tank the whole call. ```ts // Handler pseudocode — indicator enrichment across threat intel sources async handler(input, ctx) { const [vt, abuse, greynoise] = await Promise.allSettled([ vtService.lookup(input.indicator), abuseIpService.check(input.indicator), greynoiseService.query(input.indicator), ]); return { indicator: input.indicator, sources: { virustotal: vt.status === 'fulfilled' ? vt.value : { error: vt.reason.message }, abuseipdb: abuse.status === 'fulfilled' ? abuse.value : { error: abuse.reason.message }, greynoise: greynoise.status === 'fulfilled' ? greynoise.value : { error: greynoise.reason.message }, }, // Server synthesizes a verdict from available data — the agent gets a conclusion, not raw API dumps assessment: synthesizeVerdict(vt, abuse, greynoise), }; } ``` In both patterns, the tool surface is organized around what the user is doing. Sources are service-layer details — the agent sees `threat_enrich_indicator`, not `virustotal_lookup` + `abuseipdb_check` + `greynoise_query`. Mode-based dispatch by input type (e.g., `indicator_type: 'ip' | 'domain' | 'hash'`) naturally routes to different source chains per mode, since different sources cover different indicator types. #### Cut the surface There is no fixed ceiling on tool count and no target either. The ceiling is workflow coverage — if the domain genuinely has 20 distinct workflows, expose 20 tools; the cut is per-tool overlap and reach. After mapping tools, review the full list critically. A tool that covers a niche use case, serves a tiny fraction of agents, or duplicates what another tool already handles is a candidate for deferral. Drop it from the design and note it as a future addition if demand warrants. Every tool in the surface is cognitive load for tool selection — a tight surface outperforms a comprehensive one. #### Instruction tools **Applies when:** the domain has recurring "how do I do X well given my current state" questions worth merging with static procedural content. Skip otherwise. Some domains benefit from a tool whose output is **guidance, not data** — a markdown playbook tailored by live account state, with pre-filled next-step tool calls. These sit between Prompts (static templates, client-invokable) and action tools (do work, return data): they return advice, but the advice is worth more than static text because it merges procedural content with the agent's actual situation. Characteristics: - **Output is markdown guidance**, not structured data (though the output schema still has fields — typically `guidance`, `diagnostics`, and `nextToolSuggestions`) - **Merges static procedural content with live state** — the value is the tailoring. "You have 12 staged files spanning 4 unrelated changes — split them into separate commits before pushing" beats a generic best-practices article. The same shape works in other domains: "Your slowest query is 2.3s on `orders.customer_id` — add the index before tuning the planner" (database advisor), "Error rate spiked 4× at 14:32 UTC, 4 minutes after the `web@a3f9c2` deploy — roll back before chasing the upstream provider" (incident triage). - **`readOnlyHint: true`; `openWorldHint` follows where the diagnostics come from** — `false` when the live state is local (a repo on disk), `true` when it is fetched from an external API. No writes either way. - **Outputs `nextToolSuggestions`** — an array of recommended follow-up tool calls with arguments **pre-filled** from the diagnostics, not just tool names. The agent consumes the playbook, then executes steps with other tools. - **Consolidate by `topic` enum** — what could be N separate per-topic tools collapses into one ```ts const wrapupInstructions = tool('git_wrapup_instructions', { description: 'Get procedural guidance tailored to the current repo state: best-practice markdown merged with live diagnostics (staged/unstaged files, branch info, recent commits) and pre-filled follow-up tool calls. Read-only; execute the steps with other tools.', annotations: { readOnlyHint: true, openWorldHint: false }, input: z.object({ topic: z.enum(['review-changes', 'stage-and-commit', 'push-to-remote']) .describe('Playbook topic. Determines which static guidance is returned and which live state is fetched for tailoring.'), }), output: z.object({ guidance: z.string() .describe('Markdown playbook content, tailored to current account state.'), diagnostics: z.record(z.unknown()) .describe('Live state used to tailor the guidance (e.g., staged file count, branch divergence, recent commit cadence).'), nextToolSuggestions: z.array(z.object({ toolName: z.string().describe('Tool to call next.'), reason: z.string().describe('Why this step is recommended given current state.'), args: z.record(z.unknown()).describe('Arguments pre-filled from diagnostics.'), })).describe('Recommended follow-up calls with arguments already populated.'), }), }); ```
GitHub에서 보기
이 SKILL.md는 매우 커서 SkillsMP가 여기에는 첫 섹션만 미리 보여줍니다. GitHub에서 보기