Skip to main content

wicked-testing-ai-feature-test-engineer

星标0
分支0
更新时间2026年7月12日 22:44

Tier-2 specialist — testing LLM-backed features. Prompt-injection probes (direct / indirect / payload-smuggling / multi-turn), jailbreak library (DAN, grandma, token-smuggling, base64), refusal-rate regression, hallucination drift against a caller-provided golden set, output-drift monitors (JSON-schema, token length, citation fidelity). Use when: LLM feature under test, prompt-injection review, jailbreak sweep, refusal-rate check, hallucination regression, RAG citation audit, "does this AI feature still behave after the prompt change". NOT THIS WHEN: - Post-deploy LLM cost/latency monitoring — use `production-quality-engineer` - Classical model-accuracy metrics (precision/recall on labelled data) — use `data-quality-tester` - Security bugs in the surrounding app (authz, secrets) — use `security-test-engineer`; AI-specific attack surface stays here <example> Context: Reviewer wants to verify a support-bot didn't regress after a system-prompt change. user: "Run the prompt-injection + refusal-rate suite a

安装

用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。

SKILL.md
readonly