Skip to main content

PSPDFKit-labs/agentic-usability

SkillsMP 已收集 PSPDFKit-labs/agentic-usability 中的 10 个 Skill。打开任一 Skill 可查看来源和详情。

最近记录的来源活动
SkillsMP 收录数据更新
已收集 skills
10
GitHub 星标
19
GitHub Forks
0

已展示 10 / 10 个已收集 Skill。

职业分类
软件开发工程师
描述

Initialize a new agentic-usability benchmark pipeline project. Use when setting up a new SDK benchmark, creating a config.json, or starting a new evaluation project.

原文语言:英语

更新
职业分类
网络与计算机系统管理员
描述

Launch an interactive shell inside a microsandbox for debugging. Supports bare mode, executor setup, or judge setup with optional test case scaffolding.

原文语言:英语

更新
职业分类
软件质量保证分析师与测试员
描述

Run the full evaluation pipeline (execute, judge, report) for an SDK usability benchmark. Use when running a complete benchmark end-to-end, resuming an interrupted pipeline, or checking pipeline status.

原文语言:英语

更新
职业分类
软件质量保证分析师与测试员
描述

Execute benchmark test cases in sandboxed environments with AI agents. Spins up microsandbox containers for each test case and extracts solutions.

原文语言:英语

更新
职业分类
软件开发工程师
描述

Export a benchmark pipeline as a zip file for sharing or archiving. Excludes cache and large snapshots.

原文语言:英语

更新
职业分类
软件质量保证分析师与测试员
描述

Generate SDK usability test cases by exploring source code. Use when creating benchmark test suites, generating test cases for an SDK, or when the user wants to create evaluation scenarios.

原文语言:英语

更新
职业分类
软件开发工程师
描述

Analyze benchmark results and identify SDK improvement areas. Use when reviewing evaluation results, finding failure patterns, identifying documentation gaps, or understanding API design issues.

原文语言:英语

更新
职业分类
软件开发工程师
描述

Open the web UI to visually inspect, edit, and run the benchmark pipeline. Use when the user wants a visual interface for their pipeline.

原文语言:英语

更新
职业分类
软件质量保证分析师与测试员
描述

Have an LLM judge compare reference and generated solutions, scoring on API discovery, correctness, completeness, and functional correctness.

原文语言:英语

更新
职业分类
软件质量保证分析师与测试员
描述

Display a terminal scorecard of benchmark results showing pass rates, scores by difficulty, and per-test breakdowns. Use when the user asks about benchmark results, scores, or wants to see how their SDK performed.

原文语言:英语

更新
已展示 10 / 10 个已收集 Skill。