Skip to main content

PSPDFKit-labs/agentic-usability

SkillsMP has collected 10 skills from PSPDFKit-labs/agentic-usability. Open a skill to review its source and details.

Latest recorded source activity
SkillsMP catalog refreshed
skills collected
10
GitHub stars
19
GitHub forks
0

Showing 10 of 10 collected skills.

occupation
Software Developers
description

Initialize a new agentic-usability benchmark pipeline project. Use when setting up a new SDK benchmark, creating a config.json, or starting a new evaluation project.

updated
occupation
Network & Computer Systems Administrators
description

Launch an interactive shell inside a microsandbox for debugging. Supports bare mode, executor setup, or judge setup with optional test case scaffolding.

updated
occupation
Software Quality Assurance Analysts & Testers
description

Run the full evaluation pipeline (execute, judge, report) for an SDK usability benchmark. Use when running a complete benchmark end-to-end, resuming an interrupted pipeline, or checking pipeline status.

updated
occupation
Software Quality Assurance Analysts & Testers
description

Execute benchmark test cases in sandboxed environments with AI agents. Spins up microsandbox containers for each test case and extracts solutions.

updated
occupation
Software Developers
description

Export a benchmark pipeline as a zip file for sharing or archiving. Excludes cache and large snapshots.

updated
occupation
Software Quality Assurance Analysts & Testers
description

Generate SDK usability test cases by exploring source code. Use when creating benchmark test suites, generating test cases for an SDK, or when the user wants to create evaluation scenarios.

updated
occupation
Software Developers
description

Analyze benchmark results and identify SDK improvement areas. Use when reviewing evaluation results, finding failure patterns, identifying documentation gaps, or understanding API design issues.

updated
occupation
Software Developers
description

Open the web UI to visually inspect, edit, and run the benchmark pipeline. Use when the user wants a visual interface for their pipeline.

updated
occupation
Software Quality Assurance Analysts & Testers
description

Have an LLM judge compare reference and generated solutions, scoring on API discovery, correctness, completeness, and functional correctness.

updated
occupation
Software Quality Assurance Analysts & Testers
description

Display a terminal scorecard of benchmark results showing pass rates, scores by difficulty, and per-test breakdowns. Use when the user asks about benchmark results, scores, or wants to see how their SDK performed.

updated
Showing 10 of 10 collected skills.