cross-vendor-judging
Routes selective local Codex plan, code, verification, and visual judgment with evidence-backed blocking and strict cost caps.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
Routes selective local Codex plan, code, verification, and visual judgment with evidence-backed blocking and strict cost caps.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
Screenshot-first visual end-to-end validation policy for browser-rendered UI and webview jobs after verifier approval.
Default engineering workflow. Classify a technical request, route it through the native T1/T2/T3 engineering fleet, and return only independently verified completion to the orchestrator.
Safely operate Claude Code's opt-in advanced capabilities—sessions, worktrees, agent teams, Skills, MCP, plugins, hooks, Channels, schedules, goals, and Agent SDK integrations—when the user explicitly asks for one. Do not use for an ordinary build request.
Shared evidence, accessibility, context, and handoff rules for the local design-agent fleet.
Produces contextual creative theses and bounded visual directions without generic model defaults.
Creates one disposable selected-direction prototype under the isolated design prototype root.
| name | cross-vendor-judging |
| description | Routes selective local Codex plan, code, verification, and visual judgment with evidence-backed blocking and strict cost caps. |
Read docs/JUDGE-CONTRACT.md and .claude/judge-policy.json. Codex is a critic, not an implementer or deterministic verifier.
PLAN_DUCK: after Scribe's plan and applicable design pass, before approval/engineering.CODE_REVIEW: after JOB_DONE, before Verifier.VERIFICATION_CHALLENGE: after VERIFICATION_PASS, before Browser Validator, only when coverage/evidence risk remains.VISUAL_REVIEW: after browser evidence and before final acceptance, only for Studio, brand-critical, accessibility-critical, or explicitly high-value UI.Skip trivial deterministic single-file work without public-contract, security, data-flow, or user-visible impact. T1 normally skips unless a trigger applies. T2 uses Luna/Terra at relevant boundaries. T3/high-consequence work uses Terra. Standard work has at most two calls; high-risk/Studio has at most four. Put adversarial criteria into the one applicable Terra rubric rather than making a duplicate call.
gpt-5.6-luna, medium: clear repeatable plan completeness and bounded evidence checks.gpt-5.6-terra, medium: everyday multi-file, cross-layer, contract, state, and integration review; also all high-risk, adversarial, verification, and brand-critical visual review.Start at the routed tier. Escalate once only after ABSTAIN, insufficient evidence, or a substantiated risk-tier conflict. Never use Claude as a cross-vendor fallback.
The local adapter, not the judge, computes the policy action. Only an unsuppressed major/critical finding with per-finding high confidence, criterion, evidence, and location can require remediation. Minor, advisory, style-only, and lower-confidence findings remain visible. Verifier/Browser failures always dominate a judge pass. One judge-driven remediation loop is permitted.
In shadow mode, all valid results are advisory regardless of policy action. Never claim a skipped or unavailable required high-risk check passed.