Skip to main content
Manusで任意のスキルを実行
ワンクリックで

llm-eval-robustness

スター0
フォーク0
更新日2026年6月24日 06:01

Use when integrating or benchmarking an LLM API (especially reasoning models like GLM-5.2, DeepSeek V4, or any new provider/endpoint), or when a model "looks weak" and you need to rule out integration artifacts before concluding. Covers anti-hang call layers, reasoning-model handling (max_tokens, empty content), endpoint comparison, prefix caching, cheat-proof evaluation (faithfulness), and variance-aware measurement. Generalizable beyond fmhub.

インストール

Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。

SKILL.md
readonly