Skip to main content
Run any Skill in Manus
with one click

llm-eval-robustness

Stars0
Forks0
UpdatedJune 24, 2026 at 06:01

Use when integrating or benchmarking an LLM API (especially reasoning models like GLM-5.2, DeepSeek V4, or any new provider/endpoint), or when a model "looks weak" and you need to rule out integration artifacts before concluding. Covers anti-hang call layers, reasoning-model handling (max_tokens, empty content), endpoint comparison, prefix caching, cheat-proof evaluation (faithfulness), and variance-aware measurement. Generalizable beyond fmhub.

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly