- name
- ai4s-task-harness
- description
- Use when starting any AI4S analysis task (bioinformatics pipeline, model training, benchmark, data integration) to build a task-specific harness BEFORE writing analysis code. Symptoms - about to run a multi-step experiment, need reproducible results, want to prevent metric drift or story drift, setting up a new analysis phase.
# AI4S Task Harness Builder
## Overview
Build a **task-specific harness** before writing analysis code, so Claude orchestrates the task with contracts, sanity checks, and evidence generation — not just free-form coding. The harness is the scaffolding that prevents silent metric drift, story drift, and irreproducibility.
Core idea from Anthropic's dynamic workflows: Claude's default coding harness works for general tasks, but AI4S tasks need **domain-specific guardrails** — data contracts, statistical sanity, evidence locking, and negative controls. This skill builds those guardrails per-task.
## When to Use
- Starting a new analysis phase (e.g. "run differential expression", "build a classifier", "benchmark against SILVA")
- Any task that will produce numbers cited in a manuscript
- Multi-step pipelines where intermediate results feed downstream
- Tasks involving HPC submission, large-scale benchmarking, or cross-dataset validation
## When NOT to Use
- Pure writing tasks (use `ml-paper-writing`)
- One-off exploratory plots that won't appear in the paper
- Project-level scaffolding from scratch (use `/init_ai4s` instead)
## The Harness Stack
```
┌─────────────────────────────────────────┐
│ STORY_LOCK.md — which sub-question? │ ← narrative anchor
├─────────────────────────────────────────┤
│ HYPOTHESIS.md — falsifiable claim │ ← scientific rigor
├─────────────────────────────────────────┤
│ harness/contracts/ — input schemas │ ← data integrity
├─────────────────────────────────────────┤
│ harness/sanity/ — negative controls │ ← anti-artifact
├─────────────────────────────────────────┤
│ harness/benchmarks/ — metric capture │ ← evidence lock
├─────────────────────────────────────────┤
│ harness/registry.py — phase object │ ← orchestration
├─────────────────────────────────────────┤
│ tests/test_contracts/ — locked numbers │ ← drift guard
├─────────────────────────────────────────┤
│ Workflow script (optional) — parallel │ ← scale & verify
└─────────────────────────────────────────┘
```
## Build Process (follow in order)
### Step 1: Anchor to the story
Before any code, answer:
- Which sub-question in `plan/STORY_LOCK.md` does this task serve?
- What evidence tier will it produce? (Causal / Inferential / Correlative / Descriptive)
- Main text or extended data?
If no sub-question fits, **stop** — the task may be scope creep.
### Step 2: Write HYPOTHESIS.md
```markdown
## Hypothesis
<one falsifiable sentence>
## Null hypothesis
<what we expect if the effect is absent>
## Variables
- Independent: ...
- Dependent: ...
- Controlled: ...
## Statistical test (pre-registered)
<test name, correction method, threshold>
## Sample size / power
<n per group, expected effect size, power>
## Seed
<from configs/, never hardcoded>
```
Run `bash check_science.sh` — it must pass before proceeding.
### Step 3: Write the data contract
Create `harness/contracts/<task_name>.py`:
```python
def check_<task_name>_inputs() -> None:
"""Assert input data matches expected schema."""
# 1. File existence
# 2. Column schema + dtypes
# 3. Expected dimensions (rows, features)
# 4. Value ranges (no negative counts, no NaN in key cols)
# 5. Hash check against manifest (if data is locked)
```
### Step 4: Write the sanity check
Create `harness/sanity/<task_name>.py` with at least ONE negative control:
| Task type | Negative control |
|---|---|
| Classification | Shuffled labels → accuracy ~ chance |
| Clustering | Random features → no structure |
| Differential expression | Permuted groups → no significant genes |
| Dimensionality reduction | Uniform random → no clusters |
| Benchmark comparison | Null classifier → floor metric |
### Step 5: Register the phase
```python
@register_phase("phase_<task_name>")
class Phase<TaskName>(Phase):
inputs = [PROJECT_ROOT / "data" / "..."]
outputs = [PROJECT_ROOT / "results" / "phase_<task_name>" / "..."]
seed = cfg.seed # from config, never literal
def run(self):
from src.<module> import run_<task_name>
run_<task_name>(self.inputs, self.outputs, seed=self.seed)
def benchmark(self):
# Return flat dict of metrics for evidence table
return {"accuracy": ..., "f1": ..., "n_samples": ...}
def sanity(self):
from harness.sanity.<task_name> import check_shuffled
return [lambda: check_shuffled(seed=self.seed)]
```
### Step 6: Lock the evidence (after first successful run)
Add to `tests/test_contracts/test_evidence_constants.py`:
```python
def test_<task_name>_accuracy():
assert 0.85 <= result["accuracy"] <= 0.95, "accuracy drifted"
```
### Step 7: (Optional) Build a workflow for scale
For tasks needing parallelism, adversarial verification, or multi-dataset benchmarking, write a Claude Code workflow script. Pick from these AI4S patterns:
## AI4S Workflow Patterns
### Pattern A: Fan-Out Benchmark
Run the same pipeline across N datasets in parallel, then synthesize.
```
datasets ──┬── agent(dataset_1) ──┐
├── agent(dataset_2) ──┤── synthesize metrics
└── agent(dataset_N) ──┘
```
Use when: benchmarking a method across multiple cohorts (e.g. 276 16S projects).
### Pattern B: Adversarial Method Verification
One agent runs the analysis; separate agents try to break it.
```
run_analysis ──┬── verify: shuffled labels
├── verify: held-out samples
└── verify: known-positive recovery
```
Use when: any result that will appear in main text. Prevents self-preferential bias.
### Pattern C: Hypothesis Tournament
Generate competing explanations, judge each against evidence.
```
evidence ──┬── hypothesis_1 ──┐
├── hypothesis_2 ──┤── judge panel ── winner
└── hypothesis_3 ──┘
```
Use when: root-cause investigation, model selection, method comparison.
### Pattern D: Pipeline-Then-Audit
Run the full pipeline, then audit every intermediate for biological sense.
```
pipeline(steps) ── audit_agent(step_1_output)
audit_agent(step_2_output)
...
```
Use when: multi-step bioinformatics pipelines where each step can silently produce garbage.
## Failure Modes This Prevents
| Failure mode | Without harness | With harness |
|---|---|---|
| **Metric drift** | Paper says 92%, code produces 89% | `test_evidence_constants` fails |
| **Story drift** | Manuscript claims X, code does Y | `evidence_table.py` auto-generates from code |
| **Silent fallback** | Method A fails, silently uses method B | Contract rejects; no silent fallback rule |
| **Artifact survival** | Shuffled labels still show signal | `sanity/shuffled_labels.py` catches it |
| **Agentic laziness** | Claude stops after step 2 of 5 | Workflow script enforces all steps |
| **Self-preferential bias** | Claude judges own output as "good" | Adversarial verification agents |
| **Irreproducibility** | Results differ on re-run | Seed locked, env pinned, manifest hashed |
## Quick Reference: Harness Completeness Check
Before calling the task done, verify:
```
[ ] HYPOTHESIS.md written and check_science.sh passes
[ ] Data contract in harness/contracts/
[ ] At least one negative control in harness/sanity/
[ ] Phase registered in harness/registry.py
[ ] Benchmark metrics captured (flat dict, no nested)
[ ] Evidence constants locked in tests/test_contracts/
[ ] Seed from config, not hardcoded
[ ] `python -m harness.run_all` exits 0
[ ] `python -m harness.evidence_table` regenerates without diff
```
## Common Mistakes
- **Writing analysis code before the contract**: the contract defines what "correct input" means — without it, garbage-in is invisible.
- **Skipping the negative control**: "it works" means nothing without "it fails when it should fail."
- **Hardcoding thresholds in tests**: use tolerances that reflect actual pipeline variance, not wishes.
- **Registering the phase before outputs exist**: register only after the first successful run; until then the phase stays commented out.
- **Using workflows for simple tasks**: a single-dataset, single-method analysis doesn't need fan-out. Reserve workflows for genuine parallelism or adversarial verification.
View on GitHub