用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/Abd0r/porcupineai --skill benchmark-evals命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
Write code that does not look like AI output. Reject low-evidence and low-signal TypeScript patterns (chained assertions, unknown/object widening, Reflect, runtime typeof, module mocks) in favor of typed, boundary-checked code.
Control a Home Assistant smart home through its MCP server. Query sensor and entity states read-only, then act on devices (lights, switches, climate) via call_service with explicit user approval. Use for "turn off the living room lights", "what is the thermostat set to", "set bedroom temp", or any Home Assistant entity/device task.
Produce clear, accessible diagrams and charts (architecture, flow, sequence, ER, state, org, Gantt, timeline) as Mermaid or polished self-contained SVG. Covers picking the right diagram type, a consistent visual system, and accessibility. Use whenever you need to explain a system, a flow, data, or a process in a README, doc, benchmark report, plan, or launch graphic.
基于 SOC 职业分类
正在显示 SKILL.md
| name | benchmark-evals |
| description | Compare methods fairly - baselines, held-out data, leakage prevention, and honest significance. |
| stack | sci |
Use this skill whenever methods are compared on measured numbers — model evals, algorithm benchmarks, A/B tests, or reproduction studies. An unfair comparison is worse than no comparison.
reproducible-experiments).EVIDENCE.md.reproducible-experiments to pin the eval environment and run command.data-analysis for the statistics behind the comparison.literature-review to check prior evals and their reported baselines.research-writing when the comparison becomes a results section.project-hygiene to keep the eval protocol and results in a Project/ workspace.