Skip to main content

AurganicSubstance/Harness-Pair-Benchmarks

SkillsMP は AurganicSubstance/Harness-Pair-Benchmarks から 8 件の skill を収集しています。skill を開くとソースと詳細を確認できます。

記録された最新のソース活動
SkillsMP カタログ更新
収集済み skills
8
GitHub スター
4
GitHub フォーク
0

このリポジトリの skills

収集済み skill 8 件中 8 件を表示しています。

職業分類
ソフトウェア品質保証アナリスト・テスター
説明

Grade a backend corrective (debugging) benchmark submission under submissions/be-corrective/ by running each question's Node check.mjs (binary 10/0). Use when asked to evaluate, score, or grade a be-corrective submission.

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

Grade a data-analyst constructive (build-from-scratch) benchmark submission under submissions/da-constructive/ against the hidden reference answers and rubrics, producing per-question and overall scores out of 10. Use when asked to evaluate, score, or grade a…

原文の言語: 英語

更新
職業分類
ソフトウェア品質保証アナリスト・テスター
説明

Grade a data-analyst corrective (find-the-mistakes) benchmark submission under submissions/da-corrective/ - run the scripted Q1 HTML checker (10/0), the scripted Q2 verify.py (net counting), and LLM-match the Q3 mistakes.md entries to the MISTAKES keys. Use…

原文の言語: 英語

更新
職業分類
ソフトウェア品質保証アナリスト・テスター
説明

Grade a frontend corrective (debugging) benchmark submission under submissions/fe-corrective/ by running each question's Playwright check.mjs (binary 10/0). Use when asked to evaluate, score, or grade a fe-corrective submission.

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

Take the backend corrective (debugging) capability benchmark - create a submission folder named <llm>+<harness> under submissions/be-corrective/ and fix all nine broken Node backends (api/db/service × easy/mid/hard). Use when asked to take the be-corrective…

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

Take the data-analyst constructive (build-from-scratch) capability benchmark - create a submission folder named <llm>+<harness> under submissions/da-constructive/ and produce all three deliverables (offline dashboard, CSV analysis, architecture design) from…

原文の言語: 英語

更新
職業分類
ソフトウェア開発者
説明

Take the data-analyst corrective (find-the-mistakes) capability benchmark - create a submission folder named <llm>+<harness> under submissions/da-corrective/ and answer all nine corrective questions across three types (fix a broken HTML dashboard, find the…

原文の言語: 英語

更新
職業分類
ソフトウェア品質保証アナリスト・テスター
説明

Take the frontend corrective (debugging) capability benchmark - create a submission folder named <llm>+<harness> under submissions/fe-corrective/ and fix all nine broken frontend apps (web/react/app-shell × easy/mid/hard). Use when asked to take the…

原文の言語: 英語

更新
収集済み skill 8 件中 8 件を表示しています。