Skip to main content

threadlight-demo-data-factory

Generate per-domain Faker-style synthetic data + Cosmos seed/reset scripts for a threadlight process. Reads spec § 11d Demo Data and industry realism rules from threadlight-design, then produces scripts/seed_data.py and scripts/reset_data.py plus the seed JSON files in specs/sample-data/. USE FOR: generate demo data, synthetic data, Faker generators, Cosmos seed script, demo reset, mock data factory, threadlight demo data, populate sample-data, golden cases, idempotent reset. DO NOT USE FOR: indexing real customer data (use foundry-iq for that), live data ingestion (use threadlight-event-triggers), MCP server scaffold (use foundry-mcp-aca).

Zur Installation springen

Quellinformationen

Repository
aiappsgbb/threadlight-skills
Letzte Quellaktivität
25. September 2026 um 14:02
Erkannte Sprache von SKILL.md
Englisch
Sterne
1
Forks
5

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.

Datei-Explorer
2 Dateien

SKILL.md wird angezeigt

SKILL.md
Quellanweisungen · Schreibgeschützte Vorschau
name
threadlight-demo-data-factory
description
Generate per-domain Faker-style synthetic data + Cosmos seed/reset scripts for a threadlight process. Reads spec § 11d Demo Data and industry realism rules from threadlight-design, then produces scripts/seed_data.py and scripts/reset_data.py plus the seed JSON files in specs/sample-data/. USE FOR: generate demo data, synthetic data, Faker generators, Cosmos seed script, demo reset, mock data factory, threadlight demo data, populate sample-data, golden cases, idempotent reset. DO NOT USE FOR: indexing real customer data (use foundry-iq for that), live data ingestion (use threadlight-event-triggers), MCP server scaffold (use foundry-mcp-aca).
metadata
{"version":"1.0.1"}
# Threadlight Demo Data Factory ## Presenter-ready incumbent data For an explicitly selected presenter-ready profile, consume the process-owned `specs/presenter-contract.json` and [incumbent adoption map](../threadlight-deploy/references/presenter-adoption.md). Retain existing producers and data; adoption is not permission to run seed/reset. This selected-profile rule takes precedence over the generic reset examples below: never wipe retained receipts, reset consumed approvals/operation IDs, or clear UNKNOWN effects. Ordinary authorized corrections remain in scope; expired source or authorization does not grant a new write. Preparation requires the designated producer and a distinct synthetic revision with new lawful operation intent. Reads must not refresh data. Exercise the day-after case: expiry denies new work while the shell and historical reads remain available only under their own contract, authorization and retention. Treat malformed configuration separately from valid-but-expired data; do not force full regeneration or extend leases. Generate per-domain synthetic demo data + idempotent reset/seed scripts for a threadlight process. The output drives both the mock MCP server (via `foundry-mcp-aca` Option D) and the workspace UI (via `threadlight-workspace-ui`) — every demo surface reads the same seed. > **Why a separate skill?** `foundry-mcp-aca` knows how to *serve* mock > data via FastMCP; `threadlight-design` knows how to *declare* what > entities and shapes are needed. This skill bridges them: it generates > the actual JSON files (with realistic distributions, golden cases, and > reset semantics) that the MCP server returns and the workspace renders. ## When to Use - Process spec has at least one system marked `availability: mock` in § 5 - Process spec has § 11d Demo Data populated - Demo needs reset-to-pristine for live recovery (every customer-facing demo does) - Pilot is the canonization opportunity for an industry's realism rules ## When NOT to Use - Customer's real backend is already accessible (no mocks needed) - Demo data is so tiny it's hand-authored faster (e.g. 3 records total) - Process talks only to public APIs (no synthetic data needed) --- ## Industry canons — anchor pilot pattern Per-industry realism rules live in `threadlight-design/references/data-realism/{industry}.md` (one of: `fsi.md` · `retail.md` · `telco.md` · `mfg.md`). A canon is **only as credible as the pilot it's been anchored against** — generic vocabulary lists are aspirational; rules referenced from a *shipped* pilot's seed JSON are testable. | Industry | Canon file | Anchor pilot status | |----------|-----------|---------------------| | FSI | `data-realism/fsi.md` | **Anchored** to reusable financial-services realism patterns (PCI-DSS PAN masking · MCC anchors · Reg E/Z timer math · Fed holiday list · CNP fraud signals · worked examples). Compliance golden cases are aspirational pending future pilots. | | Retail | `data-realism/retail.md` | Aspirational — awaiting PIM / returns-triage pilot. | | Telco | `data-realism/telco.md` | Aspirational — awaiting a future telco operations pilot. | | Mfg | `data-realism/mfg.md` | Aspirational — awaiting supplier-risk / shift-handover pilot. | **The anchor-pilot rule.** When you generate seed data for a pilot, mine the produced JSON for new realism rules and fold them back into the canon as an "Anchor pilot" subsection with a worked-example table (canon section → pilot file/field → worked value). The FSI canon's `### Anchor-pilot worked example` table is the template — copy that shape for sibling pilots. **Honesty rule.** When the shipped pilot violated the canon (e.g. real merchant names slipping through), document it in a `### Traps` subsection with a "Don't / Do" table, NOT silently. The next pilot's generator should reject the v3-style anti-pattern up front. --- ## Input contract / Output artifacts **Input contract**: - `specs/SPEC.md` § 4 **Data Models** — entity schemas - `specs/SPEC.md` § 5 **System Integrations** — which systems are `mock` - `specs/SPEC.md` § 11d **Demo Data (Realism rules)** — required: - Per-entity volumes - Distribution rules - Named golden cases - Reset semantics (`idempotent` / `append-only` / `none`) - Industry realism reference (e.g. `industry: fsi-banking` → loads `threadlight-design/references/data-realism/fsi.md`) - `threadlight-design/references/data-realism/{industry}.md` — the per-industry rule book (universal rules + industry-specific overrides) **Output**: ``` specs/sample-data/ ├── {entity1}.json # Generated, includes _meta block ├── {entity2}.json └── README.md # Generation date + golden case list scripts/ ├── seed_data.py # Faker-driven generator (regenerates JSON files) ├── reset_data.py # Wipes Cosmos + reloads from JSON ├── pyproject.toml # uv-managed (faker, azure-cosmos, etc.) └── README.md # How to run + golden case scripts src/mcp/data/ # COPIED from specs/sample-data/ at deploy time ├── {entity1}.json # (handled by threadlight-deploy / foundry-mcp-aca) └── {entity2}.json ``` --- ## The realism stack ``` Universal rules ↓ (always applied) Industry rules (e.g. fsi.md) ↓ (industry-specific overrides) Spec § 11d (per-process tweaks) ↓ (process-specific overrides) Generated data ``` Each layer can override the layer above it. The factory composes them deterministically (with a seeded RNG) so two runs of the same spec produce byte-identical output. --- ## Generation procedure ### Step 1: Read inputs ```python industry = spec["demo_data"]["industry"] # "fsi-banking", "retail-cpg", "telco", "mfg" realism_rules = load_realism(f"data-realism/{industry}.md") volumes = spec["demo_data"]["volumes"] # {"customers": 50, "orders": 200, ...} distributions = spec["demo_data"]["distributions"] # per-entity skew rules golden_cases = spec["demo_data"]["golden_cases"] # named hand-curated records reset_mode = spec["demo_data"]["reset_semantics"] # "idempotent" / "append-only" / "none" entities = spec["data_models"] # field schemas from § 4 ``` ### Step 2: Generate `scripts/seed_data.py` ```python """Synthetic data generator — driven by spec § 11d. Run: uv run scripts/seed_data.py [--seed 42] Output: specs/sample-data/*.json (deterministic with --seed) """ import json, random from datetime import datetime, timezone from pathlib import Path from faker import Faker # Deterministic seeding so two runs produce identical output random.seed(42) fake = Faker(["en_US"]) fake.seed_instance(42) DATA_DIR = Path(__file__).parent.parent / "specs" / "sample-data" DATA_DIR.mkdir(parents=True, exist_ok=True) def gen_customers(n: int) -> list: """Generates n customer records per universal + FSI rules.""" out = [] for i in range(n): out.append({ "customer_id": f"DEMO-CUST-{i:05d}", "name": _shifted_company_name(), # never a real bank "incorporation_date": fake.date_between(start_date="-30y", end_date="-1y").isoformat(), # ... fields from spec § 4 Customer entity ... }) # Splice in golden cases by ID for golden in [g for g in GOLDEN_CASES if g["entity"] == "customers"]: out = [g for g in out if g["customer_id"] != golden["data"]["customer_id"]] out.append(golden["data"]) return out # ... one gen_* function per entity ... if __name__ == "__main__": for entity in ENTITIES: records = ENTITY_GENERATORS[entity](VOLUMES[entity]) path = DATA_DIR / f"{entity}.json" # Wrap shape: top-level object with _meta + records — NOT a list # of [meta, record, record, ...]. Loaders MUST do # `json.load(f)["records"]` and `json.load(f)["_meta"]` rather than # filter "is the first element a meta?". The list-with-meta-prepended # form silently breaks any consumer that iterates the JSON as a flat # array (eval scenarios, the workspace UI dataset loader, etc.). out = { "_meta": { "generated_at": datetime.now(timezone.utc).isoformat(), "seed": 42, "version": "1.0", "entity": entity, "count": len(records), }, "records": records, } path.write_text(json.dumps(out, indent=2)) print(f" wrote {path} ({len(records)} records)") ``` ### Cosmos data-plane RBAC — required before running seed_data.py / reset_data.py > **Critical first-time gotcha.** Cosmos uses **separate** control-plane > and data-plane RBAC. Control-plane Owner / Contributor lets you create > containers but NOT write items. Data-plane writes need the **`Cosmos > DB Built-in Data Contributor`** role (definition id > `00000000-0000-0000-0000-000000000002`). Without it, `seed_data.py` > fails with `Forbidden ... required RBAC permissions to perform > action [Microsoft.DocumentDB/databaseAccounts/sqlDatabases/containers/items/upsert]` > on EVERY upsert. > > Origin: recent pilot retrospective — first run of `seed_data.py` > failed every upsert with Forbidden; control-plane Owner doesn't > grant data-plane writes. Grant once per pilot, both for the deployer (so they can run seed_data locally) AND for the agent/bot UAMI (so the runtime can read/write): ```bash RG=<resource-group> COSMOS=<cosmos-account-name> DEPLOYER_OID=$(az ad signed-in-user show --query id -o tsv) UAMI_OID=$(az identity show -g $RG -n <uami-name> --query principalId -o tsv) SCOPE=$(az cosmosdb show -g $RG -n $COSMOS --query id -o tsv) for OID in $DEPLOYER_OID $UAMI_OID; do az cosmosdb sql role assignment create -g $RG -a $COSMOS \ --role-definition-id "00000000-0000-0000-0000-000000000002" \ --principal-id "$OID" \ --scope "$SCOPE" done sleep 30 # RBAC propagation; 25-30s is the empirically observed minimum ``` The Bicep in `infra/modules/cosmos-db.bicep` should declare the UAMI assignment so it survives `azd provision`; the deployer assignment is typically a one-time bootstrap step the SE runs by hand or that lives in `infra/scripts/postprovision.py`. ### Step 3: Generate `scripts/reset_data.py` For `idempotent` reset: ```python """Reset Cosmos containers from specs/sample-data/. Safe to run while the agent is up — completes in <30s for live demo recovery. Usage: uv run scripts/reset_data.py [--container <name>] """ import asyncio, json, os from pathlib import Path from azure.identity.aio import DefaultAzureCredential from azure.cosmos.aio import CosmosClient DATA_DIR = Path(__file__).parent.parent / "specs" / "sample-data" COSMOS_ENDPOINT = os.environ["COSMOS_ENDPOINT"] DB_NAME = os.environ["COSMOS_DATABASE"] # Cap concurrent Cosmos requests so the SDK doesn't exhaust connections # under "wipe + reload 5 containers x ~1000 docs each" parallel fan-out. SEM = asyncio.Semaphore(8) async def _bounded(coro): async with SEM: return await coro async def reset_container(client, db_name: str, container_name: str, entity_file: str): db = client.get_database_client(db_name) container = db.get_container_client(container_name) # Read the partition-key path from container metadata so we use the # CORRECT field, not a guessed `partition_key` attribute. Docs can # legitimately omit the partition_key value (it's derived from the path). container_props = await container.read() pk_path = container_props["partitionKey"]["paths"][0].lstrip("/") def _pk_value(item): # nested-path support — pk_path can be `tenant_id` or `org/tenant_id` cursor = item for part in pk_path.split("/"): cursor = cursor.get(part) if isinstance(cursor, dict) else None return cursor if cursor is not None else item["id"] # Load fresh records FIRST — if loading fails, abort BEFORE deleting, # so a malformed JSON doesn't leave the container empty mid-demo. payload = json.loads((DATA_DIR / entity_file).read_text()) records = payload["records"] if isinstance(payload, dict) else [ r for r in payload if not r.get("_meta") # back-compat for old shape ] # Wipe (bounded concurrency) await asyncio.gather(*[ _bounded(container.delete_item(item["id"], partition_key=_pk_value(item))) async for item in container.read_all_items() ]) # Reload (bounded concurrency) await asyncio.gather(*[_bounded(container.upsert_item(r)) for r in records]) print(f" reset {container_name}: {len(records)} records") async def main(): async with DefaultAzureCredential() as cred: async with CosmosClient(COSMOS_ENDPOINT, cred) as client: await asyncio.gather(*[ reset_container(client, DB_NAME, name, file) for name, file in CONTAINER_MAP.items() ]) if __name__ == "__main__": asyncio.run(main()) ``` For `append-only`: delete only items with `_meta.demo: true` — preserve human-introduced records. Use the same partition-key-discovery pattern; do NOT assume a field name. For `none`: skip generating reset_data.py (read-only datasets). ### Step 4: Generate `specs/sample-data/README.md` ```markdown # Sample data > Generated by `threadlight-demo-data-factory` from spec § 11d. > Industry realism: `{industry}` (see `threadlight-design/references/data-realism/{industry}.md`) ## Files | Entity | Records | Notes | |--------|---------|-------| | customers.json | 50 | log-normal distribution; 3 golden cases | | orders.json | 200 | mean $1,200, p99 $50K | ## Golden cases Hand-curated records the demo script narrates around. **Do not delete or rename.**
Auf GitHub ansehen
Diese SKILL.md ist sehr gross, daher zeigt SkillsMP hier nur den ersten Abschnitt. Auf GitHub ansehen