Skip to main content

agent-evaluation

Tests and benchmarks LLM agents covering behavioral testing, capability assessment, reliability metrics, and production monitoring. Use when evaluating agent quality, designing eval suites, building regression tests, or measuring real-world reliability beyond benchmark scores.

Jump to install

Source facts

Repository
VKirill/antigravity-for-claude-code
Last source activity
May 25, 2026 at 13:43
Detected SKILL.md language
English
Stars
15
Forks
2

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.