ワンクリックで
extract-computations
Extract Per-File Section/Computation Data Into policy_facets/computations/
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
メニュー
Extract Per-File Section/Computation Data Into policy_facets/computations/
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
SOC 職業分類に基づく
Emit Catala
Create Sample Tests
Draft or Update Test Cases
Expand Test Coverage
Extract Test Cases from Policy Documents
Generate a Demo App (Catala-Python Backend)
| name | extract-computations |
| description | Extract Per-File Section/Computation Data Into policy_facets/computations/ |
Read one policy doc under <domain>/input/policy_docs/, load the naming manifest, parse its H1–H3 sections, and write a YAML map {sections} to the mirrored destination at <domain>/policy_facets/computations/<rel>.md.yaml.
This skill is invoked per file by /index-inputs (in a loop over REINDEX entries) and may also be invoked standalone by the analyst against a single source file. The non-AI half (file enumeration, manifest sync, mirror-deletes, plan/finalize handoff) lives in xlator extract-computations <domain> --plan and --finalize; this skill is the AI half that does the per-file content generation.
/extract-computations <path_to_policy_file>
<path_to_policy_file> must resolve to a path under <DOMAINS_DIR>/<domain>/input/policy_docs/. If not, pre-flight emits :::error and stops.
Read ../../core/output-fencing.md now.
Run these checks before doing anything else:
Argument provided?
Path resolves under <DOMAINS_DIR>/<domain>/input/policy_docs/?
<path_to_policy_file> to an absolute path. Accept absolute paths, paths relative to the project root, or paths relative to $DOMAINS_DIR.input/policy_docs/. The directory three levels up from input/policy_docs/<rel> (i.e., the ancestor whose child input/policy_docs/ exists) is <domain>. The grandparent is $DOMAINS_DIR.<DOMAINS_DIR>/<domain>/input/policy_docs/<rel> for any valid <domain> →
:::error
Path must be under /input/policy_docs/. Got:
:::
Then stop.Source file exists and is readable?
Source has .md extension?
.md →
:::error
Only .md sources are supported in v1; got:
:::
Then stop.md_quality gate.
<DOMAINS_DIR>/<domain>/policy_facets/input-index.yaml exists, look up input/policy_docs/<rel>.md in its files: block. If found and md_quality.score < 40 (the project's REJECTED threshold) →
:::error
Source rejected by index (md_quality.score=). Fix or remove from input/policy_docs/.
:::
Then stop.input-index.yaml exists yet (analyst running this skill before /index-inputs), proceed without the gate. The gate only fails on a known-low-quality source.This skill has 5 steps:
policy_facets/computations/<rel>.md.yaml via xlator emit-per-file-yamlRead the full file content at <DOMAINS_DIR>/<domain>/input/policy_docs/<rel>.md.
Extract all H1 (#), H2 (##), and H3 (###) headings in document order. For each heading, collect the section body — the text between this heading and the next heading of equal or higher level.
If the file has no H1–H3 headings, treat the whole file as a single section. The heading for that single section is the filename stem (without .md) prefixed with #.
Read ../../core/naming_guide.md now — the static plugin-wide style guide consulted on every run.
Then resolve <domain_dir> from the source path: walk ancestors of <path_to_policy_file> to find input/policy_docs/; the directory three levels up is <domain>; its parent is $DOMAINS_DIR. (The pre-flight already does this lookup; reuse the resolved <domain_dir>.)
If <domain_dir>/specs/naming-manifest.yaml exists, build a {normalize(policy_phrase) → {name, source_doc, section}} lookup map by walking each entry's observations: list across inputs.*.*, computed.*, and outputs.* (post-3.0 schema):
observations: list (when present). For every observation with a policy_phrase field, add one map key. The map value is {name: <leaf_key>, source_doc, section} taken from the per-observation triple. One manifest entry may contribute multiple map keys when it carries multi-source observations.policy_phrase (phrase-absent observation — variable was seen in a section but no verbatim phrase was recorded) contribute no map key; they're not useful for phrase-based reconciliation.observations: list (synthesized outputs that never appeared in any source) contribute no map keys — same effect as today's "entries with no policy_phrase are skipped."inputs.Household.gross_income and inputs.Applicant.gross_income both contributing an observation with phrase "gross monthly income"), prefer the observation whose source_doc: matches the file currently being processed; deterministic tiebreak: alphabetical by entity name. The collision rule operates at observation granularity, not entry granularity.The normalizer used here: lowercase, strip leading articles (a, an, the), strip ASCII punctuation, collapse whitespace.
If the manifest does not exist (first run on a domain), the lookup map is empty — Step 3 derives fresh names from the static guide. This is normal and expected on the first /index-inputs run.
Schema-version tolerance. This skill reads observations: lists only. Pre-3.0 manifests with scalar policy_phrase / source_doc / section fields are not supported — /index-inputs will hit merge_naming_manifest.py's version-rejection gate (v3.0 required) before this skill runs. If a stale manifest somehow reaches this skill (e.g., a hand-edited domain), the absence of observations: produces an empty lookup map and Step 3 derives fresh names. No silent shape coercion.
Section filter: Only sections that contain identifiable rule logic (i.e., would carry a non-empty computations: field) are emitted. Sections with no rule logic — narrative, definitions-only prose, intro/overview text, table of contents, etc. — are dropped from the output entirely. Parse all H1–H3 sections internally so you can attribute stage: ancestors and surface variables, but the output sections: list contains only the surviving sections.
For each surviving section produce:
heading: — verbatim heading text including the # / ## / ### prefix. The prefix encodes the level; do NOT strip it.
summary: — one sentence describing what this section covers, in the policy's own terminology.
tags: — 3–5 short noun-phrase tags (lowercase, hyphenated or single-word). These are downstream filtering signals.
stage: — optional snake_case identifier naming the stage of analysis the section belongs to. Populate ONLY when the source doc surfaces an explicit phase or stage signal — examples that justify a stage::
# Phase 1 — Initial Screening (the heading itself is the signal).Phase 2: Detailed Eligibility — the stage label is attributable to the ancestor).Convert the surfaced label to a snake_case identifier (Phase 1 — Initial Screening → initial_screening). Omit the field entirely when no such signal exists in or above the section. Inventing a stage: when the source has no signal degrades downstream defaults — an absent field is stronger than a hallucinated one.
stage_source: — required when stage: is present, omitted when stage: is omitted. Value is the verbatim source-text phrase that justified the stage: identifier — copied character-for-character from the source .md, no paraphrasing, no truncation that breaks substring matching. Downstream consumers run grep -F "<stage_source>" <input/policy_docs/<rel>.md> to verify the AI honored the explicit-signal rule. If you cannot find a verbatim quote in the source, the signal is not explicit — omit stage: entirely rather than invent or paraphrase.
computations: — required, non-empty list. Every emitted section has at least one entry; sections that would carry an empty list are filtered out per the section-filter rule above. Each entry has:
description: — one sentence describing the computation in plain language.
preconditions: — optional boolean expression describing when the expr_hint: applies, derived from the section's own heading, its parent headings, and the surrounding text. The value is a list of terms joined by implicit AND at the top level. Each term is one of:
"household contains a working adult", "gross_income > 0").{all_of: [<term>, ...]} — an explicit AND group; useful for nesting inside any_of.{any_of: [<term>, ...]} — an OR group; terms inside may themselves be string clauses or further all_of / any_of groups, so arbitrary nesting is allowed.Example: a parent H2 "Working Adults" plus a phrase "if the applicant is over 65, or is married to an employed spouse" yields
preconditions:
- "household contains a working adult"
- any_of:
- "applicant is over 65"
- all_of:
- "applicant is married"
- "spouse is employed"
Omit the field when the computation applies unconditionally within the section.
expr_hint: — optional assignment of the form output_name = <expression> capturing the formula stated or clearly implied by the source. Include when the section states a formula or condition; omit when the logic is descriptive only.
Shape: single = separator. The LHS is a snake_case identifier naming the computed output. The RHS is the prior bare-expression form — short formula or expression fragment referencing input variables by name. The emitter rejects bare-expression expr_hint: payloads (no =); both sides must be present and non-empty when the field is present.
Examples:
gross_income = earned_income + unearned_incomeeligibility = gross_income < poverty_linemonthly_payment = base_rate * monthsFor descriptive-only computations (e.g., "households receiving SSI are categorically eligible"), omit expr_hint: entirely. Downstream consumers fall back to scanning description: prose for variable names.
Drop the section entirely when no rule logic is present — do not emit it in sections:. Never emit computations: [] and never emit a section without a computations: field.
variables: — map persisting each variable's source-anchored phrase, keyed by the snake_case variable name. Required on every section that has a computations: field; emit an entry for every variable referenced by any computation in the section (LHS of expr_hint:, RHS tokens of expr_hint:, or variable names surfaced from description: prose for descriptive-only computations). Each entry carries exactly one field:
policy_phrase: — required, non-empty string. Apply the verbatim rule from core/naming_guide.md: copy a verbatim noun phrase from the source .md body when one exists (grep -F "<policy_phrase>" <input/policy_docs/<rel>.md> must match). When no verbatim noun phrase for the variable exists in the source body, fall back to the deterministic anchor chain — a contiguous noun phrase from the variable's own computations[].description: text first, then the section heading: text, then the parent heading, then the first sentence of the section. Emit the fallback string verbatim; never paraphrase, never omit. The fallback is stable across re-runs because it is derived from source structure (or from the description text this skill itself emits, which is deterministic given the same source).Dedup is by variable name within a section: emit each variable at most once even if it appears in multiple computations. Order: alphabetical by variable name (deterministic across re-runs).
For each variable a section's computations: references (whether on the LHS or RHS of expr_hint:, or named in description: for descriptive-only computations):
policy_phrase: per the verbatim rule in core/naming_guide.md — a verbatim noun phrase from the source body (or the most specific deterministic anchor when no noun phrase exists). Never paraphrase.Use the resolved snake_case name on both sides of expr_hint: (LHS for the computed output, RHS for inputs) and in any description: mentions.
Persist the phrase in the section's variables: block on every variable referenced by the section's computations. The phrase emitted is the verbatim noun phrase from the source body when one exists; otherwise it is the deterministic-fallback string (description noun phrase → section heading → parent heading → first sentence). The variable's variables: entry always carries a non-empty policy_phrase: — phrase absence at the section level is not a first-class state for variables that are observed in that section. (Variables that never appear in any source section produce no variables: entry and surface downstream via the manifest's observations:-less synthesized-output path.)
xlator emit-per-file-yamlCompute the destination: <DOMAINS_DIR>/<domain>/policy_facets/computations/<rel>.md.yaml (where <rel>.md mirrors the source filename verbatim, with .yaml appended so the file's content type is unambiguous to editors and tooling).
Do NOT hand-format YAML. Build a JSON payload and pipe it to xlator emit-per-file-yaml via stdin. The tool validates per-computation expr_hint: shape, omits absent optional fields cleanly, and writes the destination atomically (tmp + os.replace) with the standard preamble.
JSON payload shape:
{
"destination": "<absolute path to .md.yaml file>",
"source_rel": "input/policy_docs/<rel>.md",
"sections": [
{
"heading": "# Section Title",
"summary": "...",
"tags": ["tag1", "tag2"],
"stage": "initial_screening",
"stage_source": "Phase 1 — Initial Screening",
"computations": [
{
"description": "...",
"preconditions": [...],
"expr_hint": "output_var = var1 * 0.20"
}
],
"variables": {
"output_var": {"policy_phrase": "output variable phrasing from source"},
"var1": {"policy_phrase": "var1 phrasing from source"}
}
}
]
}
Invocation pattern:
echo "$JSON_PAYLOAD" | xlator emit-per-file-yaml
The tool emits the YAML map as:
# Auto-generated by /extract-computations — do not edit manually
# Source: input/policy_docs/<rel>.md
sections:
- heading: "# Section Title"
summary: "..."
tags: [tag1, tag2, tag3]
stage: initial_screening
stage_source: "Phase 1 — Initial Screening"
computations:
- description: "..."
expr_hint: "net_income = gross_income - deductions"
# ...
variables:
net_income:
policy_phrase: "net monthly income"
gross_income:
policy_phrase: "gross monthly income"
deductions: {} # phrase not verbatim in source body — omitted
Conventions enforced by the emitter:
sections. Consumers read data["sections"] for section blocks.stage:, stage_source:, preconditions:, expr_hint:) are omitted entirely when absent from the JSON payload — never written as null or []. Per-variable policy_phrase: is required (non-empty); the payload supplies the verbatim phrase or the deterministic fallback string per core/naming_guide.md.computations: is required on every emitted section. Sections lacking rule logic are filtered out upstream (see Step 3) — they never appear in the JSON payload at all.expr_hint: shape: when present, must match <snake_case_identifier> = <non-empty expression>. The emitter rejects payloads with no =, empty LHS, empty RHS, or non-snake_case LHS.sections[*].computations: reflects source order. Downstream consumers (notably /create-ruleset-modules's sequential_chain heuristic) rely on this — within a section, the first computation in the list is the first in document order, the second is next, and so on. Build the JSON computations: array in source order.variables: shape: when present, must be a map keyed by snake_case variable name. Each value is a map with exactly one key: policy_phrase: (required, non-empty string). The emitter rejects non-dict values, missing-or-empty policy_phrase:, and non-string values.Always rewrite the destination file in full; this skill is idempotent at the file level. Per-file caching is the manifest's job (handled by xlator extract-computations --finalize), not the skill's.
:::important ✓ Wrote policy_facets/computations/.md.yaml ( section(s), computation(s)). :::
Do NOT emit :::next_step from this skill — it is per-file and is normally invoked from a parent loop. The parent (e.g., /index-inputs) emits the workflow's next-step suggestion.
path: field — the destination filename encodes the source path; path: is redundant and was removed.heading: "# Title" not heading: "Title"; the # characters encode the level.computations: block are dropped from the output — narrative, definitions-only prose, intros, and TOC sections are excluded from sections:. Never emit computations: [], and never emit a section without a computations: field.xlator emit-per-file-yaml. The tool handles quoting, optional-field omission, and the expr_hint: shape check.expr_hint: as a bare expression — it must be output_name = <expression>. The emitter rejects payloads with no =. For descriptive-only computations, omit expr_hint: entirely; downstream consumers fall back to description: prose.policy_phrase: — verbatim from the source body when one exists, otherwise the deterministic-fallback string verbatim (description noun phrase → section heading → parent heading → first sentence per core/naming_guide.md). The fallback applies uniformly across the manifest-lookup phrase AND the section-level variables: block. Paraphrase drift across re-runs silently breaks alignment with confirmed specs entries.variables[<var>].policy_phrase — every variable in the variables: block must carry a non-empty policy_phrase:. When no verbatim noun phrase exists in the source body, emit the deterministic-fallback string (per the chain above) rather than <var>: {}. The empty-entry shape is rejected by the emitter.variables: block when it's referenced by any computation — every variable named in any expr_hint: (LHS or RHS) or in description: prose must have an entry. Missing entries break /suggest-target-ruleset's observation aggregation.xlator extract-computations --finalize. When invoked standalone (outside /index-inputs), the per-file file is written but the manifest is not updated; the next --plan will simply re-extract this file (matching destination + missing manifest entry → to_extract). This is acceptable best-effort behavior for the standalone path.md_quality.score < 40. If the gate fires, fix the source or remove it from input/policy_docs/.stage: value when the source has no explicit signal — stages must be anchored to a heading, body sentence, or attributable ancestor heading. An absent stage: is the safe default; hallucinated stages flow through /create-ruleset-groups into guidance/ruleset-groups.yaml and ultimately produce clerk typecheck errors at the /extract-ruleset stage.stage_source: — it must be a verbatim substring of the source .md so grep -F "<stage_source>" <source> matches. If you cannot find a verbatim quote in the source, the signal is not explicit — omit stage: entirely rather than invent or paraphrase.stage: without stage_source: (or vice versa) — the two fields ship together. The quote is the proof that the AI honored the explicit-signal rule.