Skip to main content
Execute qualquer Skill no Manus
com um clique
Repositório GitHub

deep-swe-bench

deep-swe-bench contém 9 skills coletadas de Whamp, com cobertura ocupacional por repositório e páginas de detalhe dentro do site.

skills coletadas
9
Stars
2
atualizado
2026-07-08
Forks
0
Cobertura ocupacional
3 categorias ocupacionais · 100% classificado
explorador de repositórios

Skills neste repositório

codegraph
Desenvolvedores de software

CodeGraph scout before broad grep/read. Use for repo explanation, navigation, diagnosis, runtime/reconnect flow, contract/RPC/schema tracing, refactor/cycle seams, dead-code cleanup, test targeting, and code review.

2026-07-08
codegraph
Desenvolvedores de software

CodeGraph scout before broad grep/read. Use for repo explanation, navigation, diagnosis, runtime/reconnect flow, contract/RPC/schema tracing, refactor/cycle seams, dead-code cleanup, test targeting, and code review.

2026-07-08
paired-trajectory-analysis
Analistas de garantia de qualidade de software e testadores

Paired trajectory analysis for benchmark churn. Use when comparing two configs on matched task/rep cells, explaining solve flips, diagnosing why one config solved and another failed, separating net score from churn, or preparing evidence to improve a skill/prompt/tool from trajectory differences.

2026-07-08
prompt-embedding-analysis
Cientistas de dados

Prompt embedding analysis. Use when clustering benchmark prompt/config text, comparing semantic neighbors, or separating prompt-shaped effects from behavioral wrappers in deep-swe-bench results.

2026-07-08
benchmark-launch
Desenvolvedores de software

Use before launching harness/run_batch.py for a benchmark, especially when configs use advisor, observational-memory workers, subagents, local-vllm shims, or any model beyond the main executor; use before claiming a benchmark launch is working.

2026-07-03
benchmark-config-validation
Desenvolvedores de software

Use before adding or changing a deep-swe-bench config, model leaf, provider/model API path, usage parser, smoke contract, or extension/subagent worker usage accounting.

2026-07-02
runboard
Desenvolvedores de software

Open a Herdr tail tab for a harness run when the user explicitly asks for a runboard, tail, Herdr tab, or raw log view. For new harness/run_batch.py monitoring, prefer the structured dashboard at scripts/run_dashboard.py reading results/_runs/<run_id>; this skill is the compatibility/tail workflow.

2026-07-02
codegraph
Desenvolvedores de software

Local symbol-and-relationship map of the repo. Use to see who calls what (blast radius) before editing a function, class, or method. The binary is at /arm/bin/cg.

2026-07-01
benchmark-social-graphics
Desenvolvedores de software

Benchmark graphics. Use when creating social cards, README benchmark tables, chart images, X/Twitter graphics, or visual summaries from eval result artifacts where exact numbers, axes, labels, or datapoint placement matter.

2026-06-30