Skip to main content
在 Manus 中运行任何 Skill
一键导入

llm-eval-robustness

星标0
分支0
更新时间2026年6月24日 06:01

Use when integrating or benchmarking an LLM API (especially reasoning models like GLM-5.2, DeepSeek V4, or any new provider/endpoint), or when a model "looks weak" and you need to rule out integration artifacts before concluding. Covers anti-hang call layers, reasoning-model handling (max_tokens, empty content), endpoint comparison, prefix caching, cheat-proof evaluation (faithfulness), and variance-aware measurement. Generalizable beyond fmhub.

安装

用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。

SKILL.md
readonly