Skip to main content

velth-doc-verify

Automated L5 - verify correctness of generated VELTH documents (all 6 doc types, any vertical) without opening a PDF by hand. Deterministic defect checks plus an adversarial DIFFERENT-FAMILY LLM judge plus a localhost/UI propagation smoke. Load whenever a task touches a renderer, extractor, the export route, the gate, or any vertical's output. Scope each run to the doc types and verticals of the current session.

Source facts

Repository
Ansh2508/assistance-system
Last source activity
September 26, 2026 at 11:37
Detected SKILL.md language
English
Stars
0
Forks
0

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
velth-doc-verify
description
Automated L5 - verify correctness of generated VELTH documents (all 6 doc types, any vertical) without opening a PDF by hand. Deterministic defect checks plus an adversarial DIFFERENT-FAMILY LLM judge plus a localhost/UI propagation smoke. Load whenever a task touches a renderer, extractor, the export route, the gate, or any vertical's output. Scope each run to the doc types and verticals of the current session.
when_to_use
Before reporting any change to document_renderers.py, an extractor, the export route, or a vertical config as done — runs after velth-test- strategy's L1-L4, never as a replacement for them.
# VELTH Document Correctness Skill (automated L5) ## WHAT THIS REPLACES velth-test-strategy L5 is currently a HUMAN rung: open `smoke_t16_combined_*.pdf` and eyeball it. This skill automates that rung so Cowork verifies document correctness itself before reporting DONE. It runs AFTER L1-L4 (pre_commit_gate.py), never instead of them. ## WHY TWO LAYERS (research-grounded - do not collapse to one) - Deterministic layer = mechanical defects. Exact, zero false negatives, free. Catches the known P0 blockers: mojibake, mid-word truncation, umlaut ASCII-fallback, US-locale leak, non-German risk_class, RPZ != Schwere x Wahrscheinlichkeit, the #504 class/band mismatch, UUID / (needs_review) leaks, empty Feststellungen while hazards exist, duplicate Massnahmenplan text. - LLM judge = semantic / legal correctness regex cannot see. Graded by a SEPARATE model. Huang (ICLR 2024) + Kamoi (TACL 2024): a model grading its own reasoning with no external signal can DEGRADE output. The 2026 self-repair scaling study: logical errors are the hardest to self-detect. So: maker != checker, adversarial, reference-guided. ## WHY THE JUDGE IS A DIFFERENT FAMILY (not Claude) Cowork's implementer is Claude. LLM judges systematically favor their OWN family ("Play Favorites" arXiv 2508.06709: Claude/GPT over-score their own and same-family outputs) and that bias is AMPLIFIED inside self-refinement loops (Xu 2024; arXiv 2509.00462). Documented mitigation: judge with a DIFFERENT provider, reference-guided, calibrated to human spot-checks (Adaline 2026; Zheng 2023). => Default judge = Gemini (already trusted for GBU prose), keyed from the dev shell env, NOT the Cowork session. Set: `VELTH_JUDGE_PROVIDER=google` and `GEMINI_API_KEY=...` in `.env` `VELTH_JUDGE_MODEL=gemini-2.5-flash` (or your current Gemini id; flash is fine, use a pro tier for ambiguous legal calls - detection capability is the bottleneck). Fallbacks (selectable, lower priority): `anthropic` (same family - higher self-bias, only sensible if the maker is NOT Claude or as one voice in a jury) and `groq` (fast/cheap, weaker detection - never the sole judge on a legal doc). For high-stakes docs, a 2-family jury (Gemini + one other) cuts correlated blind spots; single different-family judge is the 80/20. **Checked against `CLAUDE.md`, 2026-08-26 — not previously verified, worth flagging rather than silently assuming stale.** Production AI generation for velth runs "Claude via AWS Bedrock (eu-central-1) primary · Mistral EU fallback" — not Gemini. That does not make the Gemini-judge default wrong: Cowork's implementer is Claude, so Gemini still satisfies "different family from the maker" regardless of what production generation uses. But it does mean **Mistral is a stronger fallback candidate than `groq`** for the `--fallback` slot if Gemini access is ever unavailable — Mistral is already a confirmed, currently-used production dependency (credentials almost certainly already exist), where `groq`'s wiring status here was never confirmed one way or the other. This has not been implemented or tested; it's a candidate worth checking before reaching for `groq`, not a verified replacement. The deterministic layer is the PRIMARY gate and has zero model bias - correctness never rides on the judge alone. ## INVOKE (scope to the session's doc types / verticals) First use: materialize the two scripts at the bottom of this skill into `apps/backend/scripts/` as no-BOM UTF-8, then run from `apps/backend/`: ```bash # deterministic + judge on a doc type for a vertical you are working on. # --fixture is a JSON content dict (live/exported); omit for the smoke fixture. uv run --no-sync python scripts/doc_verify.py \ --doc-type gbu --vertical safetransport-svg --fixture ./_fixtures/svg_gbu.json # repeat per doc type touched this session (begehungsbericht, ba, gfv, ...) ``` ```bash # UI / localhost propagation smoke (start the stack first). uv run --no-sync uvicorn api.main:app --reload --port 8000 ( cd ../frontend && bun run dev ) uv run --no-sync python scripts/ui_smoke.py --expect-class Mittel \ --export-url http://localhost:8000/<export route> --token-env VELTH_SMOKE_JWT ``` ## REPO WIRING (one-time, inside Cowork - per velth-preflight: read, don't guess) `doc_verify.py` already calls the canonical `render_project_document_sections( doc_type, title, content)` and imports the repo's own `_risk_class_from_rpz`, so it checks against the SAME mapping the renderer uses (catalog owns the number). Two hooks need wiring (they need .env + Supabase admin, so they run in Cowork): 1. `load_live_project(project_id)` -> pull `project_risk_sets.risk_set_json` (active=True) via `get_supabase_admin_client()` to verify real exports. 2. `ui_smoke.py --export-url` -> point at `export_vault_doc_version` (+ the `documents.py` theater path) with a smoke JWT in `VELTH_SMOKE_JWT`. Confirm the rpz->class band matches `document_renderers.py` on first run (the import is tried first; the hardcoded band is only a fallback). `doc_verify.py` calls the judge API over raw HTTP (no `anthropic` SDK) so it does not trip L1 import-safety; keep it where `test_no_direct_anthropic.py` does not scan (next to smoke_e2e_t16.py). ## ABSOLUTE RULES 1. Run AFTER L1-L4 (pre_commit_gate.py), never instead. 2. Any P0/P1 from either layer => NOT done. Fix, re-run. Do not commit. 3. Judge error / missing key => treated as NOT PASS (never silently green). 4. Scope to the session's doc types + verticals; do not sweep all 43 unless asked. 5. Surgical fixes (velth-preflight); re-run after each fix; cap iterations per velth-loop (do not snowball a wrong fix across rounds). 6. Report the issue list and PASS/FAIL IN CHAT - never a .log sidecar. 7. COWORK NEVER COMMITS - Anshu does. ## REFERENCES Design decisions in this skill are grounded in the following. Inline cites above map to these entries (author-date / arXiv). Self-correction & verification loops: - Madaan et al. (2023). Self-Refine: Iterative Refinement with Self-Feedback. NeurIPS 2023. arXiv:2303.17651. - Shinn et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. NeurIPS 2023. arXiv:2303.11366. - Huang et al. (2024). Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024. arXiv:2310.01798. [no external signal -> cannot reliably self-correct] - Kamoi et al. (2024). When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey. TACL 2024. [intrinsic self-correction can DEGRADE output] - Iterative Self-Repair in LLM Code Generation Across Model Scales (2026). arXiv:2604.10508. [pass-rate gains saturate after ~2 rounds; logical errors hardest to self-detect -> cap=3 + a non-LLM layer] - From Hallucination to Structure Snowballing (2026). arXiv:2604.06066. [reflection can snowball errors; detection is the bottleneck, citing Tyen et al. 2024 -> checkpoint/restore] LLM-as-judge self-preference & family bias (why a different-family judge): - Zheng et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023. arXiv:2306.05685. [~80% human agreement; position/verbosity/self-enhancement bias] - Wataoka et al. (2024). Self-Preference Bias in LLM-as-a-Judge. arXiv:2410.21819. - Xu et al. (2024). Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement. ACL 2024. [self-bias widespread; AMPLIFIED in self-refinement] - Panickssery et al. (2024). LLM Evaluators Recognize and Favor Their Own Generations. NeurIPS 2024. arXiv:2404.13076. - Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge (2025). arXiv:2508.06709. [Claude 3.5 Sonnet / GPT-4o self- and same-family bias] - AI Self-preferencing in Algorithmic Hiring (2025). arXiv:2509.00462. [reviews self-preference; amplified in self-refinement pipelines] Engineering practice (NOT peer-reviewed - workflow shape, not load-bearing claims): - Anthropic. Building Effective Agents (2024); Effective context engineering for AI agents; How we built our multi-agent research system; Claude Code best practices (code.claude.com/docs). - Cherny, B. Claude Code workflow thread (X, Jan 2026). [verification loop 2-3x quality; plan mode; one git checkout per parallel session] - Adaline (2026), LLM-as-a-Judge reliability & bias [secondary synthesis of the above]. VELTH-specific checks, gates, entrypoints, and German class invariants are grounded in the repo and in velth-preflight / velth-test-strategy - codebase facts, not papers. ## EMBEDDED SCRIPTS - materialize into apps/backend/scripts/ (no-BOM UTF-8) ### scripts/doc_verify.py ```python #!/usr/bin/env python3 """velth-doc-verify - automated L5 document-correctness verifier. Replaces the manual "open the PDF and eyeball it" L5 rung with two layers: 1. Deterministic checks - mechanical defects (mojibake, mid-word truncation, umlaut ASCII-fallback, US-locale leak, RPZ math, German risk_class, UUID / (needs_review) leaks, empty Feststellungen, duplicate Massnahmen). Cheap, exact, zero false negatives. 2. Adversarial LLM judge - semantic / legal correctness the regex layer cannot see, graded by a SEPARATE, DIFFERENT-FAMILY model, never by the agent that wrote the code. Research basis (why two layers, why a separate judge, why bounded): - Madaan 2023 (Self-Refine) / Shinn 2023 (Reflexion): iterative critique improves code -- but only with a feedback signal. - Huang ICLR 2024 + Kamoi TACL 2024: INTRINSIC self-correction (a model grading its own reasoning, no external signal) often fails to improve and can DEGRADE output. => verification must be external. maker != checker. - Self-repair scaling study (arXiv 2604.10508, 2026): logical / "assertion" errors are the hardest for a model to self-detect; gains saturate after ~2 rounds. => deterministic layer catches mechanics; LLM judge catches semantics; the loop that calls this caps iterations (see velth-loop skill). Run from apps/backend/ : uv run --no-sync python <skills_dir>/velth-doc-verify/verify_docs.py \ --doc-type gbu --vertical safetransport-svg [--fixture path.json] Exit code 0 = clean, 1 = P0/P1 issues found, 2 = harness error. """ from __future__ import annotations import argparse import json import os import re import sys import urllib.request from pathlib import Path # --- VELTH invariants (kept in sync with velth-test-strategy L5) --------------- GERMAN_CLASSES = {"Gering", "Mittel", "Hoch", "Sehr Hoch"} ENGLISH_LEAKS = {"low", "medium", "high", "critical", "very high"} UUID_RE = re.compile(r"[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}", re.IGNORECASE) NEEDS_REVIEW_RE = re.compile(r"\(?\s*needs[_ ]review\s*\)?", re.IGNORECASE) # mojibake signatures seen in Nachweisprotokoll (UTF-8 read as cp1252, etc.) MOJIBAKE = ("ä", "ö", "ü", "ß", "â€", "Ä", "Ö", "Ü", "\ufffd") # ASCII-fallback umlauts that should be real umlauts in a German legal doc ASCII_FALLBACK = re.compile(r"\b(Massnahme|Gefaehrdung|fuer|Pruef|Begehungsber)", re.IGNORECASE) US_DATE_RE = re.compile(r"\b(0?[1-9]|1[0-2])/(0?[1-9]|[12]\d|3[01])/(\d{4})\b") # MM/DD/YYYY # Per-doc-type ground-truth expectations fed to the judge (from velth-test-strategy L5). DOC_EXPECTATIONS = { "gbu": "Must contain a 'Psychische Belastungen' section (GDA category 10). " "Risk classes must be German. RPZ must equal Schwere x Wahrscheinlichkeit.", "ba": "Section 4 STOP order must be Technisch (T:) before Organisatorisch (O:) " "before Personenbezogen (P:). German only.", "betriebsanweisung": "STOP order T: -> O: -> P:. Correct GefStoffV paragraph for the " "actual substance (do NOT assume a hardcoded paragraph).", "gfv": "Gefahrstoffverzeichnis must have an AGW/Grenzwert column with units " "(e.g. 0.05 mg/m3 for Quarzfeinstaub).", "gefahrstoffverzeichnis": "AGW/Grenzwert column with correct units per substance.", "begehungsbericht": "If hazards were found, the Feststellungen table must be " "populated (never all placeholder dashes).", "begehungsprotokoll": "Feststellungen must reflect the risk_set; no '§4 Keine " "Feststellungen dokumentiert' false negative when hazards exist.", "unterweisung": "Topics must match the vertical's hazards. German only.", "pruefprotokoll": "Pruef results present; no mid-word truncation in the protocol body.", } # --- env (matches the established VELTH .env idiom) ----------------------------- def load_env() -> None: """Load apps/backend/.env into os.environ if GROQ_API_KEY is not already set.""" if os.environ.get("GROQ_API_KEY"): return for candidate in (Path(".env"), Path("apps/backend/.env"), Path("../.env")): if candidate.exists(): for line in candidate.read_text(encoding="utf-8", errors="ignore").splitlines(): m = re.match(r"^\s*([^#][^=]*)=(.*)$", line) if m: os.environ.setdefault(m.group(1).strip(), m.group(2).strip().strip('"')) return # --- LLM judge (provider-agnostic; DEFAULT a DIFFERENT family from the maker) ---- # Cowork's implementer is Claude. LLM judges favor their OWN family and the bias is # AMPLIFIED in self-refinement loops (Xu 2024; Panickssery 2024; "Play Favorites" # arXiv 2508.06709; arXiv 2509.00462). Mitigation: judge with a DIFFERENT family, # reference-guided, keyed from the dev shell env (NOT the Cowork session). So default # = Gemini (already trusted for GBU prose). anthropic/groq selectable as fallbacks; # do NOT use the anthropic SDK here (L1 import-safety) - raw HTTP only, and keep this # script where test_no_direct_anthropic.py does not scan (next to smoke_e2e_t16.py). JUDGE_PROVIDER = os.environ.get("VELTH_JUDGE_PROVIDER", "google") # google|anthropic|groq def _judge_system(doc_type: str, vertical: str) -> str: expect = DOC_EXPECTATIONS.get(doc_type.lower(), "Document must be legally correct and complete.") return ( "You are a senior DACH Fachkraft fuer Arbeitssicherheit (SiFa) auditing a German " "workplace-safety document before it is signed. Be ADVERSARIAL: assume errors exist " "and find them. Check: German-only risk classes (Gering/Mittel/Hoch/Sehr Hoch), " "RPZ = Schwere x Wahrscheinlichkeit, in-force legal references (DGUV/TRGS/ArbSchG/" "GefStoffV/BetrSichV), completeness, no contradictions, no English leakage, no " "truncation, correct paragraph citations. " f"Document type: {doc_type}. Vertical: {vertical}. Ground truth (reference): {expect} " 'Output ONLY JSON, no prose: ' '{"pass": bool, "issues": [{"severity":"P0|P1|P2","category":str,"detail":str}]}' ) def _post(url: str, data: bytes, headers: dict, timeout: int = 60) -> str: req = urllib.request.Request(url, data=data, headers=headers, method="POST") with urllib.request.urlopen(req, timeout=timeout) as r: return r.read().decode("utf-8", "ignore") def _extract_json(raw: str) -> dict: fence = "`" * 3 # built, not literal, so this is safe to embed inside markdown raw = raw.strip().strip(fence).strip() if raw[:4].lower() == "json": raw = raw[4:].strip() return json.loads(raw)
View on GitHub
This SKILL.md is very large, so SkillsMP previews the first section here. View on GitHub