Skip to main content

api-canvas

DataCanvas primitive reference — a Tier 3 SQL/analytical workspace for tabular MCP servers, backed by DuckDB. Use when registering tables from upstream APIs, running ad-hoc SQL across them, and exporting results. Covers the acquire → register → query → export flow, per-table TTL, the token-sharing pattern for multi-agent collaboration, env config, and Cloudflare Workers fail-closed behavior.

Aller à l'installation

Informations de source

Dépôt
cyanheads/obsidian-mcp-server
Dernière activité de la source
19 septembre 2026 à 15:47
Langue détectée de SKILL.md
anglais
Étoiles
682
Forks
103

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.

Affichage de SKILL.md

SKILL.md
Instructions source · Aperçu en lecture seule
name
api-canvas
description
DataCanvas primitive reference — a Tier 3 SQL/analytical workspace for tabular MCP servers, backed by DuckDB. Use when registering tables from upstream APIs, running ad-hoc SQL across them, and exporting results. Covers the acquire → register → query → export flow, per-table TTL, the token-sharing pattern for multi-agent collaboration, env config, and Cloudflare Workers fail-closed behavior.
metadata
{"author":"cyanheads","version":"2.3","audience":"external","type":"reference"}
## Overview `DataCanvas` is a primitive for **storage stashes, canvas computes**. The existing `IStorageProvider` is a key/value abstraction — it can stash blobs but exposes no analytical surface. `DataCanvas` is the analytical surface: register tabular data from upstream APIs, run SQL across multiple registered tables, and export results as CSV/Parquet/JSON. **Tier 3** — `@duckdb/node-api` is an optional peer dependency (`bun add @duckdb/node-api`). Servers that don't enable canvas pay zero install cost. Lazy-loaded on first use. **Disabled by default.** Set `CANVAS_PROVIDER_TYPE=duckdb` to enable. Otherwise `core.canvas` is `undefined`. **Cloudflare Workers:** unsupported. DuckDB has no V8-isolate build. Setting `CANVAS_PROVIDER_TYPE=duckdb` on a Worker fails closed with a `ConfigurationError` at init time. --- ## When canvas earns its keep Two gates before wiring canvas in — **both** must be yes. Canvas that fails either is a SQL surface nobody queries. 1. **Is the data analytical, not just large?** Canvas is for tabular/numeric result sets an agent runs SQL over — aggregate, group, join, time-series filter. A **discovery/search surface** returning categorical metadata (titles, IDs, types, dates) where the workflow is *find the record, then drill into it* does **not** qualify, regardless of row count. A 5,000-row search result is still discovery. The gate is **shape, not size**: the right question is "would an agent write `SELECT … GROUP BY` against this?", not "does it have many rows?" For name→ID resolution over a bounded list, reach for MCP-side list filtering (see the `design-mcp-server` skill) instead. 2. **Is it too big to inline?** A result that fits the response (≤ ~100 rows of compact data) just gets inlined — no canvas. Canvas is the third option only when shape *and* size both call for it. If canvas earns its keep, it carries an obligation: **a tool that emits a `canvas_id` MUST ship a `dataframe_query` tool in the same server's surface** (see the [simple-shape Tools row](#simple-shape-defaults) and the [Checklist](#checklist)). A `canvas_id` with no query tool is dead output — the agent literally cannot reach the staged data. --- ## Imports ```ts import type { DataCanvas, CanvasInstance, ColumnSchema } from '@cyanheads/mcp-ts-core/canvas'; ``` The framework wires the optional service onto `CoreServices`, accessible in the `setup()` callback — **not on `Context`**. Handlers access canvas via a module-level accessor: ```ts // src/services/canvas-accessor.ts import type { DataCanvas } from '@cyanheads/mcp-ts-core/canvas'; let _canvas: DataCanvas | undefined; export const setCanvas = (c: DataCanvas | undefined) => { _canvas = c; }; export const getCanvas = () => _canvas; ``` ```ts // src/index.ts — wire in setup() import { setCanvas } from './services/canvas-accessor.js'; await createApp({ setup(core) { setCanvas(core.canvas); }, }); ``` ```ts interface CoreServices { canvas?: DataCanvas; // present when CANVAS_PROVIDER_TYPE !== 'none' // ... other services } ``` --- ## The token-sharing model A canvas is identified by an opaque 10-character URL-safe `canvasId` (~10¹⁸ keyspace). Tools that touch canvas state accept an optional `canvas_id` input parameter: | Caller passes | Result | |:--------------|:-------| | **Omitted** | Framework mints a fresh canvasId, returns it in the tool output. Caller surfaces it to the user / next tool call / another agent. | | **Existing id (own tenant)** | Resolves to that canvas, slides TTL forward, returns `isNew: false`. | | **Existing id (other tenant)** | Throws `NotFound` — uniform with unknown to avoid leaking existence across tenants. | | **Unknown id** | Throws `NotFound` (`data.reason: 'canvas_not_found'`) with a recovery hint to re-run the producing tool or re-check the id. | | **Malformed id** | Throws `ValidationError` (`data.reason: 'canvas_id_malformed'`) before any lookup, with a hint naming the format. A value that cannot be an id is an input error; only a well-formed id that is absent is a lookup miss. | | **Omitted, tenant at its cap** | Throws `RateLimited` (`data.reason: 'canvas_capacity_exhausted'`, `retryable: true`) carrying `tenantId`, `activeCount`, and `cap`. The hint leads with reusing an id the caller already holds — the one reclaim path present in every configuration. | When auth is enabled, the effective scope is the composite `(tenantId, canvasId)`. In `MCP_AUTH_MODE=none`, `tenantId` collapses to `'default'` and the canvasId is the only differentiator — entropy + TTL + the framework's rate limiter make brute-force discovery operationally infeasible. **Designed for public-data servers (BrAPI, OpenFEC, etc.). Don't put PII on a no-auth canvas.** That collapse is also why the capacity hint reads the way it does: under `default` the occupied slots may belong to other callers, and a consumer's dataframe-drop tool is off by default, so "drop an unused canvas" is advice nobody can follow. The cap is reached only on the mint path, when `canvas_id` was omitted. The refusal keeps `-32003` and its HTTP 429 mapping; `data.reason` is what separates it from upstream throttling, including in the `mcp.tool.error_category` metric, where it files under `server` rather than `upstream`. ### Advertising the id shape `CanvasIdSchema` is exported from `@cyanheads/mcp-ts-core/canvas` — `z.string().regex(/^[A-Za-z0-9_-]{10}$/)` with a `.describe()` naming where an id comes from. A tool that declares its `canvas_id` field with it advertises the constraint in `inputSchema`, so a model sees the shape before it calls and an impossible value is rejected at argument validation rather than inside the handler: ```ts import { CanvasIdSchema } from '@cyanheads/mcp-ts-core/canvas'; input: z.object({ canvas_id: CanvasIdSchema.optional().describe( 'Optional canvas ID from a prior call. Omit on first call to start a fresh canvas.', ), }), ``` The two halves are independent. On a tool that adopts the shape, `"x"` fails as `InvalidParams` (-32602) with the framework's own `reason: 'invalid_arguments'` and a schema-derived hint, and the handler never runs — so `canvas_id_malformed` never fires there. It covers tools that have not adopted it and ids the registry receives from somewhere other than a validated argument, `importFrom`'s source id in particular. Adopting the shape does not change any existing server's advertised schema until that server adopts it. --- ## Lifecycle | Behavior | Default | Override | |:---------|:--------|:---------| | Sliding TTL | 24 h, extended on every operation | `CANVAS_TTL_MS` | | Absolute cap from creation | 7 days | `CANVAS_ABSOLUTE_CAP_MS` | | Per-tenant active cap | 100 canvases | `CANVAS_MAX_CANVASES_PER_TENANT` | | Sweeper interval | 60 s | `CANVAS_SWEEPER_INTERVAL_MS` (0 to disable) | | Persistence | In-memory only | — (v1; restart drops all canvases) | The sweeper runs as an `unref`'d `setInterval` — does not keep the event loop alive on its own. Shutdown via `core.canvas.shutdown(ctx)` (called automatically from `ServerHandle.shutdown()`) stops the sweeper and tears down every active DuckDB instance. --- ## API ### `canvas.acquire(maybeId, ctx, options?) → CanvasInstance` Resolves an existing canvas or creates a new one. Returns a {@link CanvasInstance} bound to `(canvasId, tenantId)`. Subsequent operations don't repeat them. ```ts import { getCanvas } from '@/services/canvas-accessor.js'; const canvas = getCanvas(); if (!canvas) throw new Error('DataCanvas is not enabled. Set CANVAS_PROVIDER_TYPE=duckdb.'); const instance = await canvas.acquire(input.canvas_id, ctx); // instance.canvasId — surface to the agent // instance.isNew — true on first call // instance.expiresAt — ISO 8601 after sliding extension ``` ### `instance.registerTable(name, rows, options?)` Register an in-memory or async-iterable rowset as a canvas table. ```ts await instance.registerTable('germplasm', rows); // Explicit schema for AsyncIterable (required — sniffer can't peek). await instance.registerTable('big_dataset', asyncRows, { schema: [ { name: 'id', type: 'BIGINT' }, { name: 'label', type: 'VARCHAR', nullable: true }, ], }); // Per-table TTL — this table ages on its own clock (30 min sliding window). // The canvas itself is unaffected; other tables on the same canvas are not touched. await instance.registerTable('recent_fetch', rows, { ttlMs: 30 * 60 * 1000 }); ``` **Schema inference** when `schema` is omitted: sniffer materializes the first 100 rows, unions JS-side types per column, and maps to DuckDB types. All inferred columns are **always nullable** — a sample can prove a column is nullable, but can never prove NOT NULL (a null may appear past the sniff window). Pass an explicit `schema` when `NOT NULL` enforcement is required. Fall-backs to `VARCHAR` for ambiguous unions (string mixed with numerics). Numeric widening: `INTEGER + DOUBLE → DOUBLE`, `INTEGER + BIGINT → BIGINT`. Column ordering follows first-appearance. **Per-table TTL (`ttlMs`)** — optional sliding TTL for this table specifically. When set: - The sweep loop drops the table (and clears its bookkeeping) when its window expires. - The TTL slides on any read or write against this table: on `registerTable` (initial set), on `query()` (both when the table appears in the SQL text and when it is the `registerAs` target). - The canvas itself is unaffected — canvas-level expiry is independent. - Tables registered without `ttlMs` inherit the canvas lifecycle exactly as before (no change to default behavior). - `instance.describe()` surfaces `TableInfo.expiresAt` (ISO 8601) for tables that have a per-table TTL; absent otherwise. ### `instance.query(sql, options?)` Run SQL across registered tables. Returns at most `rowLimit` rows (default 10 000). When the result exceeds `rowLimit`, the response carries `truncated: true` and `rowCount` reflects the number of materialized rows (not the full result set). For full result sets and exact counts, pass `registerAs` — the result is materialized as a new canvas table; the response carries a `preview` slice and the exact `rowCount`. Querying a table that does not exist throws `NotFound` (`data.reason: 'missing_table'`) with a recovery hint to re-run the tool that staged the table or list what is currently staged. This happens when a table has expired (per-table TTL), been dropped, or the name is mistyped. The error is `NotFound`, not `ValidationError` — agents should re-stage, not fix the SQL shape. A well-formed but unknown or expired `canvas_id` fails the same way (`data.reason: 'canvas_not_found'`, with its own recovery hint) — thrown by `acquire()` and every canvas operation. An id that fails the format check is a different failure: `ValidationError` with `data.reason: 'canvas_id_malformed'`, raised before the lookup on each of the three entry points that take a caller-supplied id — `acquire`, `drop` (which previously reported it as a silent `false`), and `importFrom`'s source id. A `SELECT` that parses but fails to prepare for any other reason — a mistyped column, an unknown function, an invalid expression — throws `ValidationError` (`data.reason: 'invalid_sql'`) and preserves the DuckDB binder detail in `data.binderMessage` (e.g. `Referenced column "x" not found...`, often with a candidate suggestion). This is distinct from `non_select_statement`, reserved for statements that genuinely aren't `SELECT`s — here the shape is fine, so the agent should fix the named column or function. A `SELECT` that prepares and then fails on the staged data throws `ValidationError` (`data.reason: 'sql_execution_error'`) with the engine message preserved and a hint pointing at `TRY_CAST` or filtering the offending rows. The split follows DuckDB's own execution-error classes — `Conversion Error`, `Invalid Input Error`, `Out of Range Error` — matched on the message prefix. Engine faults (`IO Error`, `INTERNAL Error`, `Out of Memory Error`, and anything unmatched) stay `DatabaseError`, so an export or import failing on I/O is never reported to the caller as bad SQL. `DUCKDB_ERROR_REASONS` exports these alongside `SQL_GATE_REASONS`. **Every gate and engine rejection carries `data.recovery.hint`**, which the framework mirrors into `content[]` as a `Recovery:` line — so the guidance reaches `structuredContent`-only and `content[]`-only clients alike. The hints name a capability, never a framework method: an MCP client sees only the consuming server's tool names, so `registerTable()` or `describe()` in a hint is guidance it cannot follow. Write your own hints the same way (see `api-errors`). ```ts const result = await instance.query(` SELECT germplasmName, COUNT(*) AS n FROM germplasm GROUP BY germplasmName ORDER BY n DESC `); // Materialize a join result for follow-up queries. const joined = await instance.query(` SELECT g.germplasmName, o.value FROM germplasm g JOIN observations o ON g.germplasmDbId = o.germplasmDbId `, { registerAs: 'g_with_obs', preview: 10 }); // joined.tableName === 'g_with_obs'; joined.rows.length === 10; joined.rowCount === <full count> // Materialize with a per-table TTL so the chained result ages independently. const chained = await instance.query( 'SELECT * FROM recent_fetch WHERE score > 0.8', { registerAs: 'high_score', ttlMs: 15 * 60 * 1000 }, ); ``` `registerAs` rejects with `ValidationError` (`data.reason: 'register_as_clash'`) if the target name already exists — drop it first. `ttlMs` on `query({ registerAs })` assigns a per-table TTL to the materialized table — the same sliding semantics as `registerTable({ ttlMs })`. The SQL text is also scanned for referenced table names; any tracked per-table TTL entry found is slid on each `query()` call. `denySystemCatalogs?: boolean` (default `false`) — when `true`, the gate rejects any reference to system catalog namespaces (`information_schema`, `pg_catalog`, `sqlite_master`, `duckdb_<name>()` calls) at the text-scan layer before the query executes. Use on shared canvases where handle possession is the access boundary — catalog namespaces let callers enumerate every staged handle. Rejection throws `ValidationError` with `data.reason: 'system_catalog_access'`. Canvas-token servers that explicitly expose `describe()` to agents do not need this; only servers that intentionally hide the full catalog should opt in. **Read-only enforcement** (four layers + optional catalog layer): 1. Text-level deny-list — pre-parse scan for file/HTTP-reading table functions (`read_csv*`, `read_json*`, `read_parquet*`, `read_text`, `read_blob`, `glob`, `iceberg_scan`, `delta_scan`, `postgres_scan`, `mysql_scan`, `sqlite_scan`, plus pre-staged spatial ones). 2. Statement count (must be 1) via `extractStatements`. 3. Statement type (must be `SELECT`) via `prepared.statementType`. 4. EXPLAIN-plan walk against an allowlisted set of physical operators + a denied-function rescan over plan metadata strings. Any layer's rejection throws `ValidationError` with a structured `data.reason`. File-reading scans (`READ_CSV`, `READ_PARQUET`, `READ_JSON`), DDL (`CREATE_*`, `DROP_*`, `ALTER_*`), DML (`INSERT`, `UPDATE`, `DELETE`), exports (`COPY_TO_FILE`), and utility statements (`PRAGMA`, `ATTACH`, `LOAD`, `SET`) are all rejected. ### `instance.registerView(name, selectSql, options?)` Register a SQL view on the canvas. The `SELECT` runs through the same gate `query()` enforces (four layers), so a malicious definition fails at registration time, not later when the view is referenced. Pass `{ denySystemCatalogs: true }` to also block catalog namespace references in the view definition — same semantics as the `query()` flag. ```ts await instance.registerView( 'sales_by_region', 'SELECT region, SUM(amount) AS total FROM sales GROUP BY region', ); // { viewName: 'sales_by_region', columns: ['region', 'total'] } // Subsequent queries against the view inherit normal gate enforcement at execution time. const result = await instance.query("SELECT total FROM sales_by_region WHERE region = 'a'"); ``` `CREATE OR REPLACE VIEW` semantics: re-registering the same name succeeds. Conflict with an existing base table throws `validationError({ reason: 'view_table_clash' })`. ### `instance.importFrom(sourceCanvasId, sourceTableName, options?)`
Voir sur GitHub
Ce SKILL.md est tres volumineux, SkillsMP affiche donc ici seulement la premiere section. Voir sur GitHub