Use when the user wants to re-evaluate a previous arksim simulation with different metrics, thresholds, or judge model without re-running the agent. Cheaper than re-simulating.
原文语言:英语
菜单
已展示 9 / 9 个已收集 Skill。
Use when the user wants to re-evaluate a previous arksim simulation with different metrics, thresholds, or judge model without re-running the agent. Cheaper than re-simulating.
原文语言:英语
Use when the user wants to inspect arksim evaluation results, debug specific failures turn by turn, or compare two runs to measure improvement.
原文语言:英语
Use when the user wants to generate, edit, or extend arksim test scenarios. Reads the agent's source code to derive realistic scenarios; can build regression scenarios from past failures.
原文语言:英语
Use when the user wants to simulate multi-turn conversations against an AI agent. Alias for the arksim-test skill; the canonical flow lives there.
原文语言:英语
Use when the user wants to test, simulate, or evaluate an AI agent against multi-turn scenarios (also exposed as the arksim-simulate alias). Discovers the agent, generates scenarios, runs simulation and evaluation, surfaces failures.
原文语言:英语
Use when the user wants to launch the arksim web dashboard to browse evaluation results visually rather than in CLI output.
原文语言:英语
Generate a PR title and description from your changes
原文语言:英语
Check your branch before opening a PR (lint, test, title validation)
原文语言:英语
Review a pull request against arksim project standards
原文语言:英语