원클릭으로
testdata-quality
Verify final tests with hard quality gates: integrity, consistency, validator, limit semantics, and wrong-solution kill.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
메뉴
Verify final tests with hard quality gates: integrity, consistency, validator, limit semantics, and wrong-solution kill.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
SOC 직업 분류 기준
Use when creating competitive programming problems with AutoCode MCP tools. Enforces the plugin workflow for problem statements, std/brute solutions, validators, generators, stress tests, final data verification, and Polygon packaging.
Validate statement samples and sample files for competitive programming problems. Ensures the expected outputs in problem statements match the actual solution output.
Turn deterministic problem_audit difficulty signals into a CF-style rating with reasons and confidence.
Audit statement, tutorial, and samples for consistency and publication readiness before packaging.
Define and enforce project-wide quality standards for agent and skill documents. Use when creating, reviewing, or refactoring files under agents/ and skills/ to keep structure, terminology, and output contracts consistent.
Audit std/brute assumptions with MCP evidence, including worst/average complexity risk and stress readiness.
| name | testdata-quality |
| description | Verify final tests with hard quality gates: integrity, consistency, validator, limit semantics, and wrong-solution kill. |
| disable-model-invocation | false |
Final test data must pass:
problem_verify_tests checks: file_count / answer_consistency / validator / no_empty. When manifest.json has special_judge: true and stress_comparison: "checker", answer_consistency and wrong_solution_kill use files/checker (testlib); with stress_comparison: "exact" they still compare strings to .ans.limit_ratio and limit_semantics (type=3 and type=4 must not be semantically overlapping).wrong_solution_kill: for each role=wrong solution in manifest.json, behavior depends on expected (default fail). expected=fail: at least one final test must be non-AC under checker mode, or not match .ans under exact mode (wrong solution must be "killed"). expected=pass: every final test must be AC / match .ans (e.g. alternate valid output). The tool reports killed, expected, and a hint per wrong solution; interpret them together.decision: go / no_gofailed_checks: list of failed checks (with reasons)coverage_report: scale distribution, limit-case ratio, and wrong-solution kill statisticsrepair_priority: fix priority (P0 / P1 / P2)reverify_plan: re-validation order after fixesgo: all required checks pass, limit semantics are valid, and wrong-solution checks match each entry's expected semantics (not only the legacy "always killed by one test" reading).no_go: any required check fails or coverage quality is insufficient.