cross-vendor-judging
Routes selective local Codex plan, code, verification, and visual judgment with evidence-backed blocking and strict cost caps.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Routes selective local Codex plan, code, verification, and visual judgment with evidence-backed blocking and strict cost caps.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
| name | cross-vendor-judging |
| description | Routes selective local Codex plan, code, verification, and visual judgment with evidence-backed blocking and strict cost caps. |
Read docs/JUDGE-CONTRACT.md and .claude/judge-policy.json. Codex is a critic, not an implementer or deterministic verifier.
PLAN_DUCK: after Scribe's plan and applicable design pass, before approval/engineering.CODE_REVIEW: after JOB_DONE, before Verifier.VERIFICATION_CHALLENGE: after VERIFICATION_PASS, before Browser Validator, only when coverage/evidence risk remains.VISUAL_REVIEW: after browser evidence and before final acceptance, only for Studio, brand-critical, accessibility-critical, or explicitly high-value UI.Skip trivial deterministic single-file work without public-contract, security, data-flow, or user-visible impact. T1 normally skips unless a trigger applies. T2 uses Luna/Terra at relevant boundaries. T3/high-consequence work uses Terra. Standard work has at most two calls; high-risk/Studio has at most four. Put adversarial criteria into the one applicable Terra rubric rather than making a duplicate call.
gpt-5.6-luna, medium: clear repeatable plan completeness and bounded evidence checks.gpt-5.6-terra, medium: everyday multi-file, cross-layer, contract, state, and integration review; also all high-risk, adversarial, verification, and brand-critical visual review.Start at the routed tier. Escalate once only after ABSTAIN, insufficient evidence, or a substantiated risk-tier conflict. Never use Claude as a cross-vendor fallback.
The local adapter, not the judge, computes the policy action. Only an unsuppressed major/critical finding with per-finding high confidence, criterion, evidence, and location can require remediation. Minor, advisory, style-only, and lower-confidence findings remain visible. Verifier/Browser failures always dominate a judge pass. One judge-driven remediation loop is permitted.
In shadow mode, all valid results are advisory regardless of policy action. Never claim a skipped or unavailable required high-risk check passed.
Screenshot-first visual end-to-end validation policy for browser-rendered UI and webview jobs after verifier approval.
Default engineering workflow. Classify a technical request, route it through the native T1/T2/T3 engineering fleet, and return only independently verified completion to the orchestrator.
Safely operate Claude Code's opt-in advanced capabilities—sessions, worktrees, agent teams, Skills, MCP, plugins, hooks, Channels, schedules, goals, and Agent SDK integrations—when the user explicitly asks for one. Do not use for an ordinary build request.
Shared evidence, accessibility, context, and handoff rules for the local design-agent fleet.
Produces contextual creative theses and bounded visual directions without generic model defaults.
Creates one disposable selected-direction prototype under the isolated design prototype root.