Skip to main content
Run any Skill in Manus
with one click

inspect-ai-python

Stars17
Forks23
UpdatedMay 19, 2026 at 06:50

Build LLM evaluations with Inspect AI (inspect-ai Python package by UK AISI). Use this skill whenever the user mentions Inspect AI, inspect-ai, LLM evaluation frameworks, eval tasks, eval datasets, solvers, scorers, agent evals, model grading, agentic benchmarks, SWE-bench evals, or any task involving evaluating language models systematically. Also trigger for questions about sandboxing model code execution, tool-use evals, multi-agent evaluation, eval log analysis, or running benchmarks like MMLU, HumanEval, GSM8K, HellaSwag, ARC with Inspect. Even if the user just says "write an eval" or "benchmark this model", consider this skill.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

File Explorer
11 files
SKILL.md
readonly