Skip to main content

sourcegraph/CodeScaleBench

جمع SkillsMP عدد ٣٠ من skills من sourcegraph/CodeScaleBench. افتح أي skill لمراجعة مصدره وتفاصيله.

آخر نشاط مصدر مسجل
آخر تحديث لفهرس SkillsMP
skills مجمعة
٣٠
نجوم GitHub
٣٢
تفرعات GitHub
٥

عرض ٣٠ من أصل ٣٠ skills مجمعة.

المهنة
مديرو الشبكات وأنظمة الحاسوب
الوصف

Archive old completed benchmark runs to save disk space and speed up scans. Triggers on archive runs, clean up runs, disk space, old runs.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
محللو ضمان جودة البرمجيات والمختبرون
الوصف

Audit benchmark suites against ABC framework (Task/Outcome/Reporting validity). Checks instruction quality, verifier correctness, reproducibility. Triggers on benchmark audit, audit benchmark, abc audit, task validity.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مديرو الشبكات وأنظمة الحاسوب
الوصف

Verify infrastructure readiness before launching benchmark runs — tokens, Docker, disk, credentials. Triggers on check infra, infrastructure check, ready to run, pre-run check.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
محللو ضمان جودة البرمجيات والمختبرون
الوصف

Compare benchmark results across agent configurations (baseline, SG_full). Show where configs diverge. Triggers on compare configs, config comparison, which config wins, MCP impact.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجياتعلماء البيانات
الوصف

Token and cost analysis per run, suite, and config. Shows most expensive tasks and config cost comparison. Triggers on cost report, how much did it cost, token usage, spending.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Generate the aggregate CSB evaluation report from completed Harbor runs. Triggers on generate report, eval report, ccb report, benchmark report.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
علماء البيانات
الوصف

Compute information retrieval quality metrics (precision, recall, MRR, nDCG, MAP) comparing file retrieval across baseline and MCP configs against ground truth. Triggers on ir analysis, retrieval metrics, file recall, ground truth, search quality.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
علماء البياناتمطوّرو البرمجيات
الوصف

Analyze MCP tool usage patterns, reward/time deltas conditioned on MCP adoption, and zero-MCP investigation. Triggers on mcp audit, mcp analysis, mcp impact, tool usage analysis, did mcp help.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Point at any repo (GitHub, GitLab, Bitbucket, Azure DevOps, self-hosted, or local) and get eval tasks that compare baseline coding agents vs MCP-augmented agents. Mines merged PRs/MRs for real code-change tasks, auto-generates ground truth from patches,…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
محللو ضمان جودة البرمجيات والمختبرون
الوصف

Run a single benchmark task locally to verify a fix. Uses haiku for speed. Triggers on quick rerun, rerun task, verify fix, test task.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
علماء البيانات
الوصف

Batch re-extract task_metrics.json for all runs after extraction bug fixes or schema changes. Triggers on reextract metrics, refresh metrics, update task metrics, fix extraction.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Run the repo health gate before commit or push to reduce doc drift and keep branches clean. Triggers on commit, push, before push, repo health, reduce drift, doc drift, clean branch.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
محللو ضمان جودة البرمجيات والمختبرون
الوصف

Configure and launch CodeScaleBench runs with current paired-run and curation guardrails.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
محللو ضمان جودة البرمجيات والمختبرون
الوصف

Quick status check on the currently active benchmark run. Lighter than watch-benchmarks, scoped to recent activity. Triggers on run status, how's it going, are tasks done, active run.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Scaffold a new Harbor-compatible benchmark task (SDLC or org-scale) and optionally a new benchmark suite. Generates task.toml, instruction.md, Dockerfile, test.sh, and registers the task. Triggers on scaffold task, new task, create task, add task, new…

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
محللو ضمان جودة البرمجيات والمختبرون
الوصف

Score individual task quality on instruction clarity, verifier quality, and reproducibility. Triggers on score task, task quality, rate tasks, task score.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
محللو ضمان جودة البرمجيات والمختبرون
الوصف

Reconcile task metadata between selected_benchmark_tasks.json and task.toml files. Finds and fixes drift. Triggers on sync metadata, check metadata, metadata mismatch, reconcile tasks.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
محللو ضمان جودة البرمجيات والمختبرونمطوّرو البرمجيات
الوصف

Investigate a specific failed benchmark task — read logs, identify root cause, check if known pattern, suggest fix. Triggers on triage, investigate failure, debug task, diagnose failure.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
محللو ضمان جودة البرمجيات والمختبرون
الوصف

Pre-flight validation of benchmark tasks before launching runs. Catches truncated instructions, metadata mismatches, missing test scripts. Triggers on validate tasks, preflight, pre-flight check, check tasks.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
محللو ضمان جودة البرمجيات والمختبرون
الوصف

Live benchmark monitoring — scan run directories, classify task status, fingerprint errors, present structured summaries. Triggers on watch benchmarks, benchmark status, run status, monitor runs.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
محللو أنظمة الحاسوب
الوصف

Analyze current benchmark state and recommend what to work on next. Triggers on whats next, what should I do, next steps, prioritize work.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
محللو ضمان جودة البرمجيات والمختبرون
الوصف

Run repo health checks, validate benchmark tasks, and audit run integrity.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
محللو ضمان جودة البرمجيات والمختبرون
الوصف

Extract metrics, score traces, and evaluate benchmark task results.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مديرو الشبكات وأنظمة الحاسوب
الوصف

Check infrastructure readiness, manage MCP tools, and audit system dependencies.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Plan upcoming work, analyze coverage gaps, and recommend next steps for benchmarking.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مطوّرو البرمجيات
الوصف

Generate evaluation reports, analyze run costs, and compare configurations.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
محللو ضمان جودة البرمجيات والمختبرون
الوصف

Launch and manage CodeScaleBench benchmark runs with paired-run guardrails, quick reruns, and execution orchestration.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
محللو ضمان جودة البرمجيات والمختبرون
الوصف

Create, mine, and validate new benchmark tasks and task suites.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
مديرو الشبكات وأنظمة الحاسوب
الوصف

Monitor active runs, check task completion status, and watch benchmark execution progress.

لغة النص الأصلي: الإنجليزية

آخر تحديث
المهنة
محللو ضمان جودة البرمجيات والمختبرون
الوصف

Investigate and triage failed benchmark tasks, analyze root causes, and plan reruns.

لغة النص الأصلي: الإنجليزية

آخر تحديث
عرض ٣٠ من أصل ٣٠ skills مجمعة.