Initialize a new agentic-usability benchmark pipeline project. Use when setting up a new SDK benchmark, creating a config.json, or starting a new evaluation project.
Langue du texte source : anglais
Menu
SkillsMP a collecté 10 skills depuis PSPDFKit-labs/agentic-usability. Ouvrez un skill pour examiner sa source et ses détails.
Affichage de 10 skills collectés sur 10.
Initialize a new agentic-usability benchmark pipeline project. Use when setting up a new SDK benchmark, creating a config.json, or starting a new evaluation project.
Langue du texte source : anglais
Launch an interactive shell inside a microsandbox for debugging. Supports bare mode, executor setup, or judge setup with optional test case scaffolding.
Langue du texte source : anglais
Run the full evaluation pipeline (execute, judge, report) for an SDK usability benchmark. Use when running a complete benchmark end-to-end, resuming an interrupted pipeline, or checking pipeline status.
Langue du texte source : anglais
Execute benchmark test cases in sandboxed environments with AI agents. Spins up microsandbox containers for each test case and extracts solutions.
Langue du texte source : anglais
Export a benchmark pipeline as a zip file for sharing or archiving. Excludes cache and large snapshots.
Langue du texte source : anglais
Generate SDK usability test cases by exploring source code. Use when creating benchmark test suites, generating test cases for an SDK, or when the user wants to create evaluation scenarios.
Langue du texte source : anglais
Analyze benchmark results and identify SDK improvement areas. Use when reviewing evaluation results, finding failure patterns, identifying documentation gaps, or understanding API design issues.
Langue du texte source : anglais
Open the web UI to visually inspect, edit, and run the benchmark pipeline. Use when the user wants a visual interface for their pipeline.
Langue du texte source : anglais
Have an LLM judge compare reference and generated solutions, scoring on API discovery, correctness, completeness, and functional correctness.
Langue du texte source : anglais
Display a terminal scorecard of benchmark results showing pass rates, scores by difficulty, and per-test breakdowns. Use when the user asks about benchmark results, scores, or wants to see how their SDK performed.
Langue du texte source : anglais