en un clic
AgentArena
AgentArena contient 14 skills collectées depuis aabbcdl, avec une couverture métier par dépôt et des pages de détail sur le site.
Skills dans ce dépôt
Plan and execute code changes using a spec-driven workflow with structured documents. Use when making non-trivial changes (multi-file, multi-step, or behavior-impacting) that need structured planning before implementation. Not for trivial single-file edits.
Create and manage changesets for versioning with @changesets/cli. Use when making changes that need a changelog entry, preparing a release, or managing package versions.
Diagnose and fix failed benchmark runs. Use when a benchmark run fails, scores are unexpected, judges show errors, or traces are incomplete.
Build the pnpm monorepo correctly respecting inter-package dependencies. Use when building, verifying changes, or troubleshooting build failures.
Write and run tests using Node.js built-in test runner (node:test) with .mjs files. Use when creating new tests, debugging test failures, or understanding test infrastructure.
Understand the vanilla JS web-report app structure, patterns, and contribution guidelines. Use when modifying the UI, adding new views, or fixing frontend bugs.
Add a new AI coding agent adapter so it can participate in benchmarks. Use when integrating any new CLI-based coding tool (Cursor-like, Copilot alternative, open-source agent, etc.) into the adapter system.
Add a new evaluation judge type so task packs can use it to verify agent correctness. Use when creating a new check, custom validator, or specialized assertion.
Run, monitor, and troubleshoot a benchmark execution. Use when executing benchmarks, checking results, debugging failed runs, or comparing agent performance.
Modify scoring logic, weight presets, or score modes. Use when adding/changing scoring modes, adjusting weights, fixing score calculation bugs, or adding new score components.
Author or modify a task pack YAML that defines benchmark tasks. Use when creating new tasks, editing existing ones, or configuring metadata.
Review git changes with project-specific risk checks layered on top of general senior code review. Use when asked for a code review, merge readiness assessment, or to flag likely regressions.
Run release-readiness checks and return a go/no-go verdict. Use when preparing a new version, asking whether the project can ship, or checking for blockers.
Modify the web report UI. Use when adding/changing dashboard views, scoring selectors, share buttons, code review panels, cost calculators, i18n strings, or any visual element.