Skip to main content
在 Manus 中运行任何 Skill
一键导入

uniclawbench-proactive-agents-real-world-tasks

星标2
分支0
更新时间2026年7月12日 14:22

Capability-driven benchmark for evaluating proactive agents in dynamic real-world settings. UniClawBench evaluates five foundational capabilities (Skill Usage, Exploration, Long-Context Reasoning, Multimodal Understanding, Cross-Platform Coordination) across 400 bilingual tasks in live Docker containers with closed-loop evaluation. Activation: proactive agents, real-world benchmark, agent evaluation, capability-driven, multimodal agents, closed-loop evaluation.

安装

用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。

SKILL.md
readonly