Skip to main content

threadlight-demo-data-factory

Generate per-domain Faker-style synthetic data + Cosmos seed/reset scripts for a threadlight process. Reads spec § 11d Demo Data and industry realism rules from threadlight-design, then produces scripts/seed_data.py and scripts/reset_data.py plus the seed JSON files in specs/sample-data/. USE FOR: generate demo data, synthetic data, Faker generators, Cosmos seed script, demo reset, mock data factory, threadlight demo data, populate sample-data, golden cases, idempotent reset. DO NOT USE FOR: indexing real customer data (use foundry-iq for that), live data ingestion (use threadlight-event-triggers), MCP server scaffold (use foundry-mcp-aca).

Aller à l'installation

Informations de source

Dépôt
aiappsgbb/threadlight-skills
Dernière activité de la source
25 septembre 2026 à 14:02
Langue détectée de SKILL.md
anglais
Étoiles
1
Forks
5

Options d'installation

Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.

Vérifiez les fichiers source

Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.

Explorateur de fichiers
2 fichiers

Affichage de SKILL.md

SKILL.md
Instructions source · Aperçu en lecture seule
name
threadlight-demo-data-factory
description
Generate per-domain Faker-style synthetic data + Cosmos seed/reset scripts for a threadlight process. Reads spec § 11d Demo Data and industry realism rules from threadlight-design, then produces scripts/seed_data.py and scripts/reset_data.py plus the seed JSON files in specs/sample-data/. USE FOR: generate demo data, synthetic data, Faker generators, Cosmos seed script, demo reset, mock data factory, threadlight demo data, populate sample-data, golden cases, idempotent reset. DO NOT USE FOR: indexing real customer data (use foundry-iq for that), live data ingestion (use threadlight-event-triggers), MCP server scaffold (use foundry-mcp-aca).
metadata
{"version":"1.0.1"}
# Threadlight Demo Data Factory ## Presenter-ready incumbent data For an explicitly selected presenter-ready profile, consume the process-owned `specs/presenter-contract.json` and [incumbent adoption map](../threadlight-deploy/references/presenter-adoption.md). Retain existing producers and data; adoption is not permission to run seed/reset. This selected-profile rule takes precedence over the generic reset examples below: never wipe retained receipts, reset consumed approvals/operation IDs, or clear UNKNOWN effects. Ordinary authorized corrections remain in scope; expired source or authorization does not grant a new write. Preparation requires the designated producer and a distinct synthetic revision with new lawful operation intent. Reads must not refresh data. Exercise the day-after case: expiry denies new work while the shell and historical reads remain available only under their own contract, authorization and retention. Treat malformed configuration separately from valid-but-expired data; do not force full regeneration or extend leases. Generate per-domain synthetic demo data + idempotent reset/seed scripts for a threadlight process. The output drives both the mock MCP server (via `foundry-mcp-aca` Option D) and the workspace UI (via `threadlight-workspace-ui`) — every demo surface reads the same seed. > **Why a separate skill?** `foundry-mcp-aca` knows how to *serve* mock > data via FastMCP; `threadlight-design` knows how to *declare* what > entities and shapes are needed. This skill bridges them: it generates > the actual JSON files (with realistic distributions, golden cases, and > reset semantics) that the MCP server returns and the workspace renders. ## When to Use - Process spec has at least one system marked `availability: mock` in § 5 - Process spec has § 11d Demo Data populated - Demo needs reset-to-pristine for live recovery (every customer-facing demo does) - Pilot is the canonization opportunity for an industry's realism rules ## When NOT to Use - Customer's real backend is already accessible (no mocks needed) - Demo data is so tiny it's hand-authored faster (e.g. 3 records total) - Process talks only to public APIs (no synthetic data needed) --- ## Industry canons — anchor pilot pattern Per-industry realism rules live in `threadlight-design/references/data-realism/{industry}.md` (one of: `fsi.md` · `retail.md` · `telco.md` · `mfg.md`). A canon is **only as credible as the pilot it's been anchored against** — generic vocabulary lists are aspirational; rules referenced from a *shipped* pilot's seed JSON are testable. | Industry | Canon file | Anchor pilot status | |----------|-----------|---------------------| | FSI | `data-realism/fsi.md` | **Anchored** to reusable financial-services realism patterns (PCI-DSS PAN masking · MCC anchors · Reg E/Z timer math · Fed holiday list · CNP fraud signals · worked examples). Compliance golden cases are aspirational pending future pilots. | | Retail | `data-realism/retail.md` | Aspirational — awaiting PIM / returns-triage pilot. | | Telco | `data-realism/telco.md` | Aspirational — awaiting a future telco operations pilot. | | Mfg | `data-realism/mfg.md` | Aspirational — awaiting supplier-risk / shift-handover pilot. | **The anchor-pilot rule.** When you generate seed data for a pilot, mine the produced JSON for new realism rules and fold them back into the canon as an "Anchor pilot" subsection with a worked-example table (canon section → pilot file/field → worked value). The FSI canon's `### Anchor-pilot worked example` table is the template — copy that shape for sibling pilots. **Honesty rule.** When the shipped pilot violated the canon (e.g. real merchant names slipping through), document it in a `### Traps` subsection with a "Don't / Do" table, NOT silently. The next pilot's generator should reject the v3-style anti-pattern up front. --- ## Input contract / Output artifacts **Input contract**: - `specs/SPEC.md` § 4 **Data Models** — entity schemas - `specs/SPEC.md` § 5 **System Integrations** — which systems are `mock` - `specs/SPEC.md` § 11d **Demo Data (Realism rules)** — required: - Per-entity volumes - Distribution rules - Named golden cases - Reset semantics (`idempotent` / `append-only` / `none`) - Industry realism reference (e.g. `industry: fsi-banking` → loads `threadlight-design/references/data-realism/fsi.md`) - `threadlight-design/references/data-realism/{industry}.md` — the per-industry rule book (universal rules + industry-specific overrides) **Output**: ``` specs/sample-data/ ├── {entity1}.json # Generated, includes _meta block ├── {entity2}.json └── README.md # Generation date + golden case list scripts/ ├── seed_data.py # Faker-driven generator (regenerates JSON files) ├── reset_data.py # Wipes Cosmos + reloads from JSON ├── pyproject.toml # uv-managed (faker, azure-cosmos, etc.) └── README.md # How to run + golden case scripts src/mcp/data/ # COPIED from specs/sample-data/ at deploy time ├── {entity1}.json # (handled by threadlight-deploy / foundry-mcp-aca) └── {entity2}.json ``` --- ## The realism stack ``` Universal rules ↓ (always applied) Industry rules (e.g. fsi.md) ↓ (industry-specific overrides) Spec § 11d (per-process tweaks) ↓ (process-specific overrides) Generated data ``` Each layer can override the layer above it. The factory composes them deterministically (with a seeded RNG) so two runs of the same spec produce byte-identical output. --- ## Generation procedure ### Step 1: Read inputs ```python industry = spec["demo_data"]["industry"] # "fsi-banking", "retail-cpg", "telco", "mfg" realism_rules = load_realism(f"data-realism/{industry}.md") volumes = spec["demo_data"]["volumes"] # {"customers": 50, "orders": 200, ...} distributions = spec["demo_data"]["distributions"] # per-entity skew rules golden_cases = spec["demo_data"]["golden_cases"] # named hand-curated records reset_mode = spec["demo_data"]["reset_semantics"] # "idempotent" / "append-only" / "none" entities = spec["data_models"] # field schemas from § 4 ``` ### Step 2: Generate `scripts/seed_data.py` ```python """Synthetic data generator — driven by spec § 11d. Run: uv run scripts/seed_data.py [--seed 42] Output: specs/sample-data/*.json (deterministic with --seed) """ import json, random from datetime import datetime, timezone from pathlib import Path from faker import Faker # Deterministic seeding so two runs produce identical output random.seed(42) fake = Faker(["en_US"]) fake.seed_instance(42) DATA_DIR = Path(__file__).parent.parent / "specs" / "sample-data" DATA_DIR.mkdir(parents=True, exist_ok=True) def gen_customers(n: int) -> list: """Generates n customer records per universal + FSI rules.""" out = [] for i in range(n): out.append({ "customer_id": f"DEMO-CUST-{i:05d}", "name": _shifted_company_name(), # never a real bank "incorporation_date": fake.date_between(start_date="-30y", end_date="-1y").isoformat(), # ... fields from spec § 4 Customer entity ... }) # Splice in golden cases by ID for golden in [g for g in GOLDEN_CASES if g["entity"] == "customers"]: out = [g for g in out if g["customer_id"] != golden["data"]["customer_id"]] out.append(golden["data"]) return out # ... one gen_* function per entity ... if __name__ == "__main__": for entity in ENTITIES: records = ENTITY_GENERATORS[entity](VOLUMES[entity]) path = DATA_DIR / f"{entity}.json" # Wrap shape: top-level object with _meta + records — NOT a list # of [meta, record, record, ...]. Loaders MUST do # `json.load(f)["records"]` and `json.load(f)["_meta"]` rather than # filter "is the first element a meta?". The list-with-meta-prepended # form silently breaks any consumer that iterates the JSON as a flat # array (eval scenarios, the workspace UI dataset loader, etc.). out = { "_meta": { "generated_at": datetime.now(timezone.utc).isoformat(), "seed": 42, "version": "1.0", "entity": entity, "count": len(records), }, "records": records, } path.write_text(json.dumps(out, indent=2)) print(f" wrote {path} ({len(records)} records)") ``` ### Cosmos data-plane RBAC — required before running seed_data.py / reset_data.py > **Critical first-time gotcha.** Cosmos uses **separate** control-plane > and data-plane RBAC. Control-plane Owner / Contributor lets you create > containers but NOT write items. Data-plane writes need the **`Cosmos > DB Built-in Data Contributor`** role (definition id > `00000000-0000-0000-0000-000000000002`). Without it, `seed_data.py` > fails with `Forbidden ... required RBAC permissions to perform > action [Microsoft.DocumentDB/databaseAccounts/sqlDatabases/containers/items/upsert]` > on EVERY upsert. > > Origin: recent pilot retrospective — first run of `seed_data.py` > failed every upsert with Forbidden; control-plane Owner doesn't > grant data-plane writes. Grant once per pilot, both for the deployer (so they can run seed_data locally) AND for the agent/bot UAMI (so the runtime can read/write): ```bash RG=<resource-group> COSMOS=<cosmos-account-name> DEPLOYER_OID=$(az ad signed-in-user show --query id -o tsv) UAMI_OID=$(az identity show -g $RG -n <uami-name> --query principalId -o tsv) SCOPE=$(az cosmosdb show -g $RG -n $COSMOS --query id -o tsv) for OID in $DEPLOYER_OID $UAMI_OID; do az cosmosdb sql role assignment create -g $RG -a $COSMOS \ --role-definition-id "00000000-0000-0000-0000-000000000002" \ --principal-id "$OID" \ --scope "$SCOPE" done sleep 30 # RBAC propagation; 25-30s is the empirically observed minimum ``` The Bicep in `infra/modules/cosmos-db.bicep` should declare the UAMI assignment so it survives `azd provision`; the deployer assignment is typically a one-time bootstrap step the SE runs by hand or that lives in `infra/scripts/postprovision.py`. ### Step 3: Generate `scripts/reset_data.py` For `idempotent` reset: ```python """Reset Cosmos containers from specs/sample-data/. Safe to run while the agent is up — completes in <30s for live demo recovery. Usage: uv run scripts/reset_data.py [--container <name>] """ import asyncio, json, os from pathlib import Path from azure.identity.aio import DefaultAzureCredential from azure.cosmos.aio import CosmosClient DATA_DIR = Path(__file__).parent.parent / "specs" / "sample-data" COSMOS_ENDPOINT = os.environ["COSMOS_ENDPOINT"] DB_NAME = os.environ["COSMOS_DATABASE"] # Cap concurrent Cosmos requests so the SDK doesn't exhaust connections # under "wipe + reload 5 containers x ~1000 docs each" parallel fan-out. SEM = asyncio.Semaphore(8) async def _bounded(coro): async with SEM: return await coro async def reset_container(client, db_name: str, container_name: str, entity_file: str): db = client.get_database_client(db_name) container = db.get_container_client(container_name) # Read the partition-key path from container metadata so we use the # CORRECT field, not a guessed `partition_key` attribute. Docs can # legitimately omit the partition_key value (it's derived from the path). container_props = await container.read() pk_path = container_props["partitionKey"]["paths"][0].lstrip("/") def _pk_value(item): # nested-path support — pk_path can be `tenant_id` or `org/tenant_id` cursor = item for part in pk_path.split("/"): cursor = cursor.get(part) if isinstance(cursor, dict) else None return cursor if cursor is not None else item["id"] # Load fresh records FIRST — if loading fails, abort BEFORE deleting, # so a malformed JSON doesn't leave the container empty mid-demo. payload = json.loads((DATA_DIR / entity_file).read_text()) records = payload["records"] if isinstance(payload, dict) else [ r for r in payload if not r.get("_meta") # back-compat for old shape ] # Wipe (bounded concurrency) await asyncio.gather(*[ _bounded(container.delete_item(item["id"], partition_key=_pk_value(item))) async for item in container.read_all_items() ]) # Reload (bounded concurrency) await asyncio.gather(*[_bounded(container.upsert_item(r)) for r in records]) print(f" reset {container_name}: {len(records)} records") async def main(): async with DefaultAzureCredential() as cred: async with CosmosClient(COSMOS_ENDPOINT, cred) as client: await asyncio.gather(*[ reset_container(client, DB_NAME, name, file) for name, file in CONTAINER_MAP.items() ]) if __name__ == "__main__": asyncio.run(main()) ``` For `append-only`: delete only items with `_meta.demo: true` — preserve human-introduced records. Use the same partition-key-discovery pattern; do NOT assume a field name. For `none`: skip generating reset_data.py (read-only datasets). ### Step 4: Generate `specs/sample-data/README.md` ```markdown # Sample data > Generated by `threadlight-demo-data-factory` from spec § 11d. > Industry realism: `{industry}` (see `threadlight-design/references/data-realism/{industry}.md`) ## Files | Entity | Records | Notes | |--------|---------|-------| | customers.json | 50 | log-normal distribution; 3 golden cases | | orders.json | 200 | mean $1,200, p99 $50K | ## Golden cases Hand-curated records the demo script narrates around. **Do not delete or rename.**
Voir sur GitHub
Ce SKILL.md est tres volumineux, SkillsMP affiche donc ici seulement la premiere section. Voir sur GitHub