一键导入
eval-audit-review
Audit SoulMap AI evals so datasets, assertions, source markers, and golden responses stay source-backed, failure-oriented, and hard to game.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
菜单
Audit SoulMap AI evals so datasets, assertions, source markers, and golden responses stay source-backed, failure-oriented, and hard to game.
用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
基于 SOC 职业分类
SoulMap, a reflective companion that helps people stop abandoning themselves. Includes a central coordination layer, a clear response pipeline, routing guidance, depth calibration, epistemic guardrails, safety guardrails, voice system, brand doctrine, and reusable templates. Mirror, not guide.
Build and maintain Python detector modules that identify framework selection signals in conversation history.
SoulMap reflective response frameworks covering emotional de-escalation, grief, existential reflection, inner parts, life direction, shadow work, synthesis, and relational inquiry. Relevant for tasks that require choosing or applying the core reflective method for a user conversation.
SoulMap safety and boundary rules covering crisis handling, dependency prevention, trauma-informed language, prompt injection defense, and scope control. Relevant for requests that involve harm, escalation, refusal, redirection, or questions about what SoulMap must not do.
SoulMap symbolic spiritual materials covering brand-safe numerology, chakra policy, healing metaphors, and archetypal language. Relevant for tasks that involve spiritual framing within SoulMap's grounded, non-predictive, non-grandiose boundaries.
SoulMap brand doctrine, positioning, message hierarchy, surface-specific rules, and strategic direction. Relevant for tasks that concern what SoulMap is, what it is not, how it sounds in public, or how brand language stays aligned across product surfaces.
| name | eval-audit-review |
| description | Audit SoulMap AI evals so datasets, assertions, source markers, and golden responses stay source-backed, failure-oriented, and hard to game. |
Use this skill when the task is to inspect or improve the trustworthiness of SoulMap's eval system rather than only adding one more case.
evals/datasets/ as the main task, use
eval-suite-maintainerrelease-readiness-reviewpython-maintainerKeep evals honest, source-backed, and useful against real failure modes instead of optimizing for easy green runs.
evals/README.mdevals/datasets/tests/contract/tests/eval_regression/src/soulmap/devtools/evals/AGENTS.md, skills/, or templates/source_markers when confidence needs to be explicit.List the eval weaknesses first, especially loose assertions, stale source links, or blind spots.
Summarize the dataset, harness, or contract changes that improved audit quality.
State which eval and pytest commands were run.
The audited eval surface should be: