- name
- threadlight-demo-data-factory
- description
- Generate per-domain Faker-style synthetic data + Cosmos seed/reset scripts for a threadlight process. Reads spec § 11d Demo Data and industry realism rules from threadlight-design, then produces scripts/seed_data.py and scripts/reset_data.py plus the seed JSON files in specs/sample-data/. USE FOR: generate demo data, synthetic data, Faker generators, Cosmos seed script, demo reset, mock data factory, threadlight demo data, populate sample-data, golden cases, idempotent reset. DO NOT USE FOR: indexing real customer data (use foundry-iq for that), live data ingestion (use threadlight-event-triggers), MCP server scaffold (use foundry-mcp-aca).
- metadata
- {"version":"1.0.1"}
# Threadlight Demo Data Factory
## Presenter-ready incumbent data
For an explicitly selected presenter-ready profile, consume the process-owned
`specs/presenter-contract.json` and
[incumbent adoption map](../threadlight-deploy/references/presenter-adoption.md).
Retain existing producers and data; adoption is not permission to run seed/reset.
This selected-profile rule takes precedence over the generic reset examples below:
never wipe retained receipts, reset consumed approvals/operation IDs, or clear
UNKNOWN effects. Ordinary authorized corrections remain in scope; expired source
or authorization does not grant a new write. Preparation requires the designated
producer and a distinct synthetic revision with new lawful operation intent.
Reads must not refresh data. Exercise the day-after case: expiry denies new work
while the shell and historical reads remain available only under their own
contract, authorization and retention. Treat malformed configuration separately
from valid-but-expired data; do not force full regeneration or extend leases.
Generate per-domain synthetic demo data + idempotent reset/seed scripts for
a threadlight process. The output drives both the mock MCP server (via
`foundry-mcp-aca` Option D) and the workspace UI (via
`threadlight-workspace-ui`) — every demo surface reads the same seed.
> **Why a separate skill?** `foundry-mcp-aca` knows how to *serve* mock
> data via FastMCP; `threadlight-design` knows how to *declare* what
> entities and shapes are needed. This skill bridges them: it generates
> the actual JSON files (with realistic distributions, golden cases, and
> reset semantics) that the MCP server returns and the workspace renders.
## When to Use
- Process spec has at least one system marked `availability: mock` in § 5
- Process spec has § 11d Demo Data populated
- Demo needs reset-to-pristine for live recovery (every customer-facing
demo does)
- Pilot is the canonization opportunity for an industry's realism rules
## When NOT to Use
- Customer's real backend is already accessible (no mocks needed)
- Demo data is so tiny it's hand-authored faster (e.g. 3 records total)
- Process talks only to public APIs (no synthetic data needed)
---
## Industry canons — anchor pilot pattern
Per-industry realism rules live in
`threadlight-design/references/data-realism/{industry}.md` (one of:
`fsi.md` · `retail.md` · `telco.md` · `mfg.md`). A canon is **only as
credible as the pilot it's been anchored against** — generic vocabulary
lists are aspirational; rules referenced from a *shipped* pilot's seed
JSON are testable.
| Industry | Canon file | Anchor pilot status |
|----------|-----------|---------------------|
| FSI | `data-realism/fsi.md` | **Anchored** to reusable financial-services realism patterns (PCI-DSS PAN masking · MCC anchors · Reg E/Z timer math · Fed holiday list · CNP fraud signals · worked examples). Compliance golden cases are aspirational pending future pilots. |
| Retail | `data-realism/retail.md` | Aspirational — awaiting PIM / returns-triage pilot. |
| Telco | `data-realism/telco.md` | Aspirational — awaiting a future telco operations pilot. |
| Mfg | `data-realism/mfg.md` | Aspirational — awaiting supplier-risk / shift-handover pilot. |
**The anchor-pilot rule.** When you generate seed data for a pilot,
mine the produced JSON for new realism rules and fold them back into
the canon as an "Anchor pilot" subsection with a worked-example table
(canon section → pilot file/field → worked value). The FSI canon's
`### Anchor-pilot worked example` table is
the template — copy that shape for sibling pilots.
**Honesty rule.** When the shipped pilot violated the canon (e.g.
real merchant names slipping through), document it in a `### Traps`
subsection with a "Don't / Do" table, NOT silently. The next pilot's
generator should reject the v3-style anti-pattern up front.
---
## Input contract / Output artifacts
**Input contract**:
- `specs/SPEC.md` § 4 **Data Models** — entity schemas
- `specs/SPEC.md` § 5 **System Integrations** — which systems are `mock`
- `specs/SPEC.md` § 11d **Demo Data (Realism rules)** — required:
- Per-entity volumes
- Distribution rules
- Named golden cases
- Reset semantics (`idempotent` / `append-only` / `none`)
- Industry realism reference (e.g. `industry: fsi-banking` → loads
`threadlight-design/references/data-realism/fsi.md`)
- `threadlight-design/references/data-realism/{industry}.md` — the
per-industry rule book (universal rules + industry-specific overrides)
**Output**:
```
specs/sample-data/
├── {entity1}.json # Generated, includes _meta block
├── {entity2}.json
└── README.md # Generation date + golden case list
scripts/
├── seed_data.py # Faker-driven generator (regenerates JSON files)
├── reset_data.py # Wipes Cosmos + reloads from JSON
├── pyproject.toml # uv-managed (faker, azure-cosmos, etc.)
└── README.md # How to run + golden case scripts
src/mcp/data/ # COPIED from specs/sample-data/ at deploy time
├── {entity1}.json # (handled by threadlight-deploy / foundry-mcp-aca)
└── {entity2}.json
```
---
## The realism stack
```
Universal rules
↓ (always applied)
Industry rules (e.g. fsi.md)
↓ (industry-specific overrides)
Spec § 11d (per-process tweaks)
↓ (process-specific overrides)
Generated data
```
Each layer can override the layer above it. The factory composes them
deterministically (with a seeded RNG) so two runs of the same spec
produce byte-identical output.
---
## Generation procedure
### Step 1: Read inputs
```python
industry = spec["demo_data"]["industry"] # "fsi-banking", "retail-cpg", "telco", "mfg"
realism_rules = load_realism(f"data-realism/{industry}.md")
volumes = spec["demo_data"]["volumes"] # {"customers": 50, "orders": 200, ...}
distributions = spec["demo_data"]["distributions"] # per-entity skew rules
golden_cases = spec["demo_data"]["golden_cases"] # named hand-curated records
reset_mode = spec["demo_data"]["reset_semantics"] # "idempotent" / "append-only" / "none"
entities = spec["data_models"] # field schemas from § 4
```
### Step 2: Generate `scripts/seed_data.py`
```python
"""Synthetic data generator — driven by spec § 11d.
Run: uv run scripts/seed_data.py [--seed 42]
Output: specs/sample-data/*.json (deterministic with --seed)
"""
import json, random
from datetime import datetime, timezone
from pathlib import Path
from faker import Faker
# Deterministic seeding so two runs produce identical output
random.seed(42)
fake = Faker(["en_US"])
fake.seed_instance(42)
DATA_DIR = Path(__file__).parent.parent / "specs" / "sample-data"
DATA_DIR.mkdir(parents=True, exist_ok=True)
def gen_customers(n: int) -> list:
"""Generates n customer records per universal + FSI rules."""
out = []
for i in range(n):
out.append({
"customer_id": f"DEMO-CUST-{i:05d}",
"name": _shifted_company_name(), # never a real bank
"incorporation_date": fake.date_between(start_date="-30y", end_date="-1y").isoformat(),
# ... fields from spec § 4 Customer entity ...
})
# Splice in golden cases by ID
for golden in [g for g in GOLDEN_CASES if g["entity"] == "customers"]:
out = [g for g in out if g["customer_id"] != golden["data"]["customer_id"]]
out.append(golden["data"])
return out
# ... one gen_* function per entity ...
if __name__ == "__main__":
for entity in ENTITIES:
records = ENTITY_GENERATORS[entity](VOLUMES[entity])
path = DATA_DIR / f"{entity}.json"
# Wrap shape: top-level object with _meta + records — NOT a list
# of [meta, record, record, ...]. Loaders MUST do
# `json.load(f)["records"]` and `json.load(f)["_meta"]` rather than
# filter "is the first element a meta?". The list-with-meta-prepended
# form silently breaks any consumer that iterates the JSON as a flat
# array (eval scenarios, the workspace UI dataset loader, etc.).
out = {
"_meta": {
"generated_at": datetime.now(timezone.utc).isoformat(),
"seed": 42,
"version": "1.0",
"entity": entity,
"count": len(records),
},
"records": records,
}
path.write_text(json.dumps(out, indent=2))
print(f" wrote {path} ({len(records)} records)")
```
### Cosmos data-plane RBAC — required before running seed_data.py / reset_data.py
> **Critical first-time gotcha.** Cosmos uses **separate** control-plane
> and data-plane RBAC. Control-plane Owner / Contributor lets you create
> containers but NOT write items. Data-plane writes need the **`Cosmos
> DB Built-in Data Contributor`** role (definition id
> `00000000-0000-0000-0000-000000000002`). Without it, `seed_data.py`
> fails with `Forbidden ... required RBAC permissions to perform
> action [Microsoft.DocumentDB/databaseAccounts/sqlDatabases/containers/items/upsert]`
> on EVERY upsert.
>
> Origin: recent pilot retrospective — first run of `seed_data.py`
> failed every upsert with Forbidden; control-plane Owner doesn't
> grant data-plane writes.
Grant once per pilot, both for the deployer (so they can run seed_data
locally) AND for the agent/bot UAMI (so the runtime can read/write):
```bash
RG=<resource-group>
COSMOS=<cosmos-account-name>
DEPLOYER_OID=$(az ad signed-in-user show --query id -o tsv)
UAMI_OID=$(az identity show -g $RG -n <uami-name> --query principalId -o tsv)
SCOPE=$(az cosmosdb show -g $RG -n $COSMOS --query id -o tsv)
for OID in $DEPLOYER_OID $UAMI_OID; do
az cosmosdb sql role assignment create -g $RG -a $COSMOS \
--role-definition-id "00000000-0000-0000-0000-000000000002" \
--principal-id "$OID" \
--scope "$SCOPE"
done
sleep 30 # RBAC propagation; 25-30s is the empirically observed minimum
```
The Bicep in `infra/modules/cosmos-db.bicep` should declare the UAMI
assignment so it survives `azd provision`; the deployer assignment is
typically a one-time bootstrap step the SE runs by hand or that lives
in `infra/scripts/postprovision.py`.
### Step 3: Generate `scripts/reset_data.py`
For `idempotent` reset:
```python
"""Reset Cosmos containers from specs/sample-data/.
Safe to run while the agent is up — completes in <30s for live demo recovery.
Usage: uv run scripts/reset_data.py [--container <name>]
"""
import asyncio, json, os
from pathlib import Path
from azure.identity.aio import DefaultAzureCredential
from azure.cosmos.aio import CosmosClient
DATA_DIR = Path(__file__).parent.parent / "specs" / "sample-data"
COSMOS_ENDPOINT = os.environ["COSMOS_ENDPOINT"]
DB_NAME = os.environ["COSMOS_DATABASE"]
# Cap concurrent Cosmos requests so the SDK doesn't exhaust connections
# under "wipe + reload 5 containers x ~1000 docs each" parallel fan-out.
SEM = asyncio.Semaphore(8)
async def _bounded(coro):
async with SEM:
return await coro
async def reset_container(client, db_name: str, container_name: str, entity_file: str):
db = client.get_database_client(db_name)
container = db.get_container_client(container_name)
# Read the partition-key path from container metadata so we use the
# CORRECT field, not a guessed `partition_key` attribute. Docs can
# legitimately omit the partition_key value (it's derived from the path).
container_props = await container.read()
pk_path = container_props["partitionKey"]["paths"][0].lstrip("/")
def _pk_value(item):
# nested-path support — pk_path can be `tenant_id` or `org/tenant_id`
cursor = item
for part in pk_path.split("/"):
cursor = cursor.get(part) if isinstance(cursor, dict) else None
return cursor if cursor is not None else item["id"]
# Load fresh records FIRST — if loading fails, abort BEFORE deleting,
# so a malformed JSON doesn't leave the container empty mid-demo.
payload = json.loads((DATA_DIR / entity_file).read_text())
records = payload["records"] if isinstance(payload, dict) else [
r for r in payload if not r.get("_meta") # back-compat for old shape
]
# Wipe (bounded concurrency)
await asyncio.gather(*[
_bounded(container.delete_item(item["id"], partition_key=_pk_value(item)))
async for item in container.read_all_items()
])
# Reload (bounded concurrency)
await asyncio.gather(*[_bounded(container.upsert_item(r)) for r in records])
print(f" reset {container_name}: {len(records)} records")
async def main():
async with DefaultAzureCredential() as cred:
async with CosmosClient(COSMOS_ENDPOINT, cred) as client:
await asyncio.gather(*[
reset_container(client, DB_NAME, name, file)
for name, file in CONTAINER_MAP.items()
])
if __name__ == "__main__":
asyncio.run(main())
```
For `append-only`: delete only items with `_meta.demo: true` — preserve
human-introduced records. Use the same partition-key-discovery pattern;
do NOT assume a field name.
For `none`: skip generating reset_data.py (read-only datasets).
### Step 4: Generate `specs/sample-data/README.md`
```markdown
# Sample data
> Generated by `threadlight-demo-data-factory` from spec § 11d.
> Industry realism: `{industry}` (see `threadlight-design/references/data-realism/{industry}.md`)
## Files
| Entity | Records | Notes |
|--------|---------|-------|
| customers.json | 50 | log-normal distribution; 3 golden cases |
| orders.json | 200 | mean $1,200, p99 $50K |
## Golden cases
Hand-curated records the demo script narrates around. **Do not delete or rename.**
Ver en GitHub