| name | slang-test-feature |
| description | Orchestrator that researches a language feature, produces a test plan for user review, and fans out parallel agents that each deliver test branches with commits. Only invoke when explicitly called via /slang-test-feature. |
| argument-hint | <feature-name> [--dry-run | --live] [--max-agents N] [--reference-url URL] [--wsl] |
| license | Apache-2.0 |
Feature Test Flow
For: End-to-end test coverage for a Slang language feature — from research through parallel test implementation to bug triage.
Usage: /slang-test-feature <feature-name> [--dry-run | --live] [--max-agents N] [--reference-url URL] [--wsl]
--dry-run (default): Local branches + local bug files, no PRs or GitHub issues filed
--live: Push branches, create PRs, file GitHub issues
--max-agents N: Maximum parallel agents (default 5)
--reference-url URL: Additional documentation URL to fetch
--wsl: Force native WSL git/gh in orchestrator and subagents. Without
it, WSL requires Windows-native git.exe/gh.exe and stops if either is
missing.
Tool Selection For GitHub Steps
Before the orchestrator or any subagent runs a git or gh command, initialize
selected tools:
ARGS="${ARGUMENTS:-}"
USE_WSL_TOOLS=false
if printf '%s\n' "$ARGS" | grep -Eq '(^|[[:space:]])--wsl([[:space:]]|$)'; then
USE_WSL_TOOLS=true
ARGS="$(printf '%s\n' "$ARGS" | sed -E 's/(^|[[:space:]])--wsl([[:space:]]|$)/ /; s/^[[:space:]]+//; s/[[:space:]]+$//')"
fi
is_wsl() {
[ -n "${WSL_DISTRO_NAME:-}" ] || grep -qi microsoft /proc/version 2>/dev/null
}
choose_tool() {
tool="$1"
if is_wsl && [ "$USE_WSL_TOOLS" = false ]; then
if command -v "${tool}.exe" >/dev/null 2>&1; then
printf '%s.exe\n' "$tool"
return 0
fi
printf 'Missing Windows-hosted tool: %s.exe\n' "$tool" >&2
printf 'Install it on Windows or rerun with --wsl to use native WSL %s.\n' "$tool" >&2
return 1
fi
if command -v >/dev/null 2>&1;
0
>&2
1
}
GIT= || 1
GH= || 1
() { -d ; }
Use $GIT and $GH for all subsequent git and gh commands. Include this
tool-selection block in every subagent prompt that can commit, push, create PRs,
or file GitHub issues.
Phase 1: RESEARCH
Gather all available knowledge about the feature. This runs in the orchestrator (main conversation).
Information Sources (priority order)
| Source | What to extract |
|---|
--reference-url (if provided) | Fetch and extract feature documentation |
external/spec/specification/ | Grammar, semantics, constraints, edge cases |
external/spec/proposals/ | Motivation, design decisions, known limitations |
docs/user-guide/ | User-facing behavior, examples, restrictions |
DeepWiki (mcp__deepwiki__ask_question) | Implementation details, edge cases, cross-feature interactions |
source/slang/slang-diagnostics.lua | Error codes related to the feature |
tests/ | Existing test inventory + coverage assessment |
source/slang/ | Key implementation files, TODO/FIXME gaps |
Research Steps
- Check for spec repo: If
external/spec/ doesn't exist, ask user whether to clone it
- Fetch reference URL if provided (use WebFetch)
- Search spec and proposals for the feature keyword
- Search user guide chapters
- Query DeepWiki for implementation details and edge cases
- Enumerate ALL diagnostics (see below)
- Inventory existing tests: find all tests related to the feature, categorize what they cover
- Search compiler source for TODO/FIXME related to the feature
Exhaustive Diagnostic Enumeration (mandatory)
Do NOT rely on keyword search of documentation alone. Extract ALL
diagnostic codes related to the feature directly from compiler source:
rg -i "generic|constraint|speciali|conform|type.param|pack|where.clause" \
source/slang/slang-diagnostic*.h --context 2
For each code found:
- Record: code number, diagnostic name, message text
- Search
tests/ to check if any test triggers this code
- Mark as COVERED or UNCOVERED
This produces the ground truth for error-path coverage. Keyword search
of docs misses codes that use different terminology -- the generics
feature had 22 diagnostic codes (out of 57 total) that were missed by
doc-based keyword search.
Line/Branch Coverage Baseline (optional)
Check the nightly coverage report at
https://shader-slang.org/slang-coverage-reports/reports/latest/linux/index.html
for key source files related to the feature. Record baseline coverage
numbers in research.md and use low branch coverage as a signal for
untested error paths. Do not use coverage % as a test-writing target.
Output
Write tmp/<feature>-<issue-id>/research.md:
- Feature overview (from spec + docs)
- Complete list of behaviors/constraints/error conditions
- Complete diagnostic code table with COVERED/UNCOVERED status
- Coverage baseline for key source files (if checked)
- Existing test inventory with coverage assessment
- Known limitations or open issues
Phase 2: PLAN
Decompose the feature into semantic dimensions — each becomes an independent sub-plan that one agent can implement as one focused branch.
Decomposition Strategy
Split by semantic area, NOT by test type:
Feature: generics
├── Sub-plan A: Type parameters on structs
│ ├── Positive: basic usage, multiple params, nested
│ ├── Negative: constraint violations, recursive types
│ └── Targets: cpu (+ GPU targets for CI)
├── Sub-plan B: Type parameters on functions
├── Sub-plan C: Where clauses and constraints
├── Sub-plan D: Generic interfaces and conformance
└── Sub-plan E: Specialization and cross-feature interactions
Why semantic dimensions?
- Naturally independent → parallelizable with no conflicts
- Each covers positive + negative + edge within its scope
- Each becomes one focused, reviewable branch/PR — one sub-plan = one agent = one branch = one PR
- Keep sub-plans small enough for a single reviewable PR (aim for 5-8 test files, ~100-200 lines)
- Related sub-plans can be combined into a single PR if the total stays under ~500 lines (e.g., all diagnostic tests for a feature, or all functional tests). Use judgement to group by theme.
Negative testing requirement: Every sub-plan that includes positive
functional tests for constrained features (interface conformance, where
clauses, generic constraints, typealias constraints) MUST also include
companion negative diagnostic tests that verify the constraints are
enforced. A positive-only test for a constrained feature will be flagged
in code review. Plan the negative companion at sub-plan time, not as an
afterthought.
Sub-plan Format
Write each to tmp/<feature>-<issue-id>/sub-plans/sub-plan-<letter>.md:
# Sub-plan [Letter]: [Name]
## Branch Name
<feature>-<dimension>
## Title
Add tests for [feature]: [dimension]
## Context
[2-3 sentences from spec about this dimension]
## Spec References
- Section X.Y: [relevant quote/summary]
## Existing Coverage
- [existing tests and what they cover]
- [identified gaps]
## Tests to Write
### 1. feature-scenario.slang
- **Path**: tests/language-feature/<feature>/<name>.slang
- **Type**: compute | diagnostic | interpret
- **Validates**: [specific behavior from spec]
- **Key assertion**: [what the CHECK verifies]
- **Negative companion**: [yes/no — if yes, list the -negative.slang file]
- **Priority**: HIGH | MEDIUM | LOW
- **Score**: N/10
### 2. ...
## Anti-overlap
- Do NOT test [X] — covered by sub-plan [other]
## Overlap with Existing Tests
- Checked against: [list files searched]
- No significant overlap / Extending [file] instead of creating new
Gap Traceability (mandatory)
Every gap from research.md must map to either a sub-plan or an explicit
SKIP with reason. No gap may be silently dropped.
Write tmp/<feature>-<issue-id>/gap-traceability.md:
| Gap | Action | Sub-plan | Reason |
|-----|--------|----------|--------|
| 30400 generic-type-needs-args | WRITE | A | No test exists |
| 30404 invalid-equality-constraint | SKIP | — | Already tested in existing file |
| Coercion constraints | SKIP | — | Low priority, score 3/10 |
| Constructor type inference | SKIP | — | Feature not implemented |
Present this table alongside the plan for user review. If any research
gap is not accounted for, the plan is incomplete.
Plan Summary
Write tmp/<feature>-<issue-id>/plan.md with overview table:
| Sub-plan | Dimension | Tests | Priority |
|---|
| A | Type params on structs | 5 | HIGH |
| B | Type params on functions | 4 | HIGH |
| ... | | | |
User Review Gate
STOP HERE and present the plan to the user. The user can:
- Approve all sub-plans
- Drop specific sub-plans
- Adjust scope (add/remove tests)
- Merge sub-plans that are too small
Share the plan (if GitHub issue exists)
STOP and ask the user before posting. Show a preview of the comment.
If approved, post an [Agent]-prefixed test plan summary (sub-plan overview table +
gap traceability) to the linked GitHub issue. This makes the coverage strategy visible
to other contributors and documents what's planned vs gaps.
Phase 3: EXECUTE
Launch parallel agents (one per approved sub-plan) in worktree isolation using
subagent_type="best-of-n-runner". Each agent gets its own git worktree and branch,
so agents cannot conflict with each other even when creating files in the same
directory hierarchy.
Constructing the Agent Prompt
Agent prompts must be self-contained. Agents cannot read other skills.
The orchestrator must read the following skills and include their content
in each agent's prompt at the marked {placeholder} locations:
slang-write-test → {slang-write-test content} (test syntax reference)
slang-build → {slang-build content} (build commands, preset selection)
slang-run-tests → {slang-run-tests content} (test commands, skip detection)
slang-create-issue "Commit Rules" section → {commit-rules}
Also include the Tool Selection For GitHub Steps block above in every prompt.
Launch agents in parallel by sending multiple Task tool calls in a single message,
each with subagent_type="best-of-n-runner".
Agent Prompt Template
You are implementing tests for the Slang compiler.
You are running in an isolated git worktree with your own branch.
## Your Sub-plan
{sub-plan content}
## Test Syntax Reference
{slang-write-test content}
## Build Reference
{slang-build content}
## Test Runner Reference
{slang-run-tests content}
## Feature Context
{research.md overview section}
## Project Context
This is the Slang shading language compiler (C++). Additional conventions:
- Format: ./extras/formatting.sh
- Single-dash CLI options: -target spirv (not --target)
- No trailing whitespace or blank lines with only spaces/tabs
## Instructions
### Step 0: Build (if needed)
Build slangc and slang-test using the build reference above, then select
`$SLANGC` and `$SLANG_TEST` using the test runner reference. Under WSL with a
Windows-hosted build, these must be `slangc.exe` and `slang-test.exe`; stop if
they are missing.
### Step 1: Write and validate tests
For each test in the sub-plan:
a. Create the .slang test file at the specified path
b. Follow the test templates from the syntax reference
c. Write natural comments explaining semantic behavior
d. Run the test: "$SLANG_TEST" tests/path/to/test.slang
e. If it fails:
- Wrong expected value → fix the test
- Compiler bug → record bug in structured format (see below)
For DISABLED tests: rename the //TEST directive to //DISABLE_TEST
### Step 2: Pre-commit quality checks
Before formatting or committing, verify EVERY test file:
a. Filename matches what the test actually verifies
b. All comments reference the correct interfaces, error codes, and
behavior (cross-check against actual code in the file)
c. No dead code: every declared function/struct/variable is used
d. No duplicates: search existing tests for the same scenario
(rg "keyword" tests/language-feature/<feature>/)
e. Feature support: for functional tests, confirm the feature compiles
with a quick "$SLANGC" invocation before writing the full test
f. DIAGNOSTIC_TEST uses exhaustive mode unless there is a documented
reason for non-exhaustive; never use non-exhaustive "just in case"
g. Negative companion: if any functional test exercises a constrained
feature (interface conformance, where clause, generic constraint),
verify a companion -negative.slang diagnostic test exists that
proves the constraint is enforced by rejecting invalid types
### Step 3: Format
Run ./extras/formatting.sh on changed files.
### Step 4: Create branch & commit
{commit-rules}
Create a dedicated branch from master for this sub-plan:
"$GIT" checkout -b <branch-name>
Stage all new test files. Create a commit with message:
"Add tests for <feature>: <dimension>"
**IMPORTANT**: Each agent creates its own branch and PR independently.
The orchestrator does NOT merge agents' work into a single branch/PR.
One sub-plan = one branch = one PR. This keeps PRs focused and reviewable.
### Step 5: Mode-dependent actions
**Dry-run mode (default)**:
- Do NOT push or create PRs
- The branch stays local in the worktree
- The orchestrator will collect results and can copy files to the main repo later
**Live mode**:
- Push branch with `"$GIT" push -u origin HEAD`
- Create PR using `"$GH" pr create` with:
- Title: {title from sub-plan}
- Label: pr: non-breaking
- Assignee: @me
- Body: Summary of tests added + what they cover
- Each agent is responsible for creating its own PR — do not defer to the orchestrator
### Step 6: Report back
Return a structured result:
BRANCH: <branch-name>
COMMIT: <commit-hash>
PR_URL: <url or "dry-run: not created">
TESTS:
- name: <test-file-path>
status: pass | fail | disabled
notes: <any relevant detail>
BUGS:
- id: <sub-plan-letter>-<number> (e.g. B-1)
test: <test-file-path that triggered it>
symptom: <what happened — crash, wrong output, missing error>
error_output: <exact compiler output, truncated to key lines>
reproducer: |
<minimal .slang code that triggers the bug>
target: <which target(s) affected>
severity: ICE | wrong-codegen | missing-diagnostic | validation-error
Constraints
- Max 5 parallel agents (configurable via --max-agents)
- Each agent uses
subagent_type="best-of-n-runner" for git worktree isolation
- Plan phase assigns unique branch names and file paths
- Agents run fully autonomously: write tests, build, run, format, commit
- In dry-run mode agents stop at commit; in live mode they push and create PRs
Phase 4: BUG TRIAGE
After all agents complete, collect and triage all reported bugs.
Step 1: Collect
Gather all bug reports from all agents into tmp/<feature>-<issue-id>/bugs/.
Step 2: Deduplicate
Group bugs by:
- Same error code → likely same root cause
- Same crash location / ICE message
- Same symptom on same target
For each group, keep one canonical bug report (best reproducer, clearest symptom).
Step 3: Write bug files
For each unique bug, write tmp/<feature>-<issue-id>/bugs/bug-<id>.md:
# Bug <ID>: [Short title]
## Severity
ICE | wrong-codegen | missing-diagnostic | validation-error
## Reproducer
\`\`\`hlsl
[minimal code]
\`\`\`
**Command:**
\`\`\`bash
"$SLANGC" -target [target] test.slang
\`\`\`
## Expected Behavior
[What the spec/docs say should happen]
## Actual Behavior
[Error output]
## Affected Tests
- `tests/path/to/disabled-test.slang` (disabled in branch `<feature>-<dim>`)
## Duplicates
- Sub-plan A, bug A-1: [same/different]
- Existing GitHub issue: [#NNNN or "not found"]
Step 4: Mode-dependent actions
Dry-run mode: Stop here. Bug files in tmp/<feature>-<issue-id>/bugs/ are the deliverable.
Live mode:
- Search existing GitHub issues for each bug
- File new issues for confirmed new bugs using
slang-create-issue bug report format with --assignee @me
- Comment on existing issues with new reproducers
- Update PR descriptions to link to filed issues
Phase 5: REPORT
Write tmp/<feature>-<issue-id>/SUMMARY.md and present to user:
## Feature Test Flow: [feature] — Results
### Mode
dry-run | live
### Branches Created
| # | Sub-plan | Branch | Tests | Pass | Disabled | Bugs |
|---|----------|--------|-------|------|----------|------|
| A | Type params on structs | generics-struct-params | 5 | 5 | 0 | 0 |
| B | Type params on functions | generics-func-params | 4 | 3 | 1 | 1 |
| ...
### PRs Created (live mode only)
| Sub-plan | PR |
|----------|-----|
| A | #1234 |
| ...
### Bugs Found
| Bug ID | Severity | File | Description |
|--------|----------|------|-------------|
| B-1 | ICE | tmp/generics/bugs/bug-B-1.md | Crash when inferring... |
### Bugs Skipped (duplicates)
| Bug ID | Reason |
|--------|--------|
| D-2 | Same as B-1 (same error code) |
### Remaining Gaps
- [anything from the plan that couldn't be tested]
### Next Steps
- [ ] Review branch generics-struct-params
- [ ] Review branch generics-func-params
- [ ] Investigate bug B-1
Output Structure
tmp/<feature>-<issue-id>/
├── research.md # Phase 1 output
├── plan.md # Phase 2 overview
├── sub-plans/
│ ├── sub-plan-a.md
│ ├── sub-plan-b.md
│ └── ...
├── bugs/
│ ├── bug-A-1.md
│ ├── bug-B-1.md
│ └── ...
├── agent-results/
│ ├── result-a.md # Raw agent output
│ ├── result-b.md
│ └── ...
└── SUMMARY.md # Phase 5 final report