Skip to main content

tangle-network/agent-eval

SkillsMP 已收集 tangle-network/agent-eval 中的 5 个 Skill。打开任一 Skill 可查看来源和详情。

最近记录的来源活动
SkillsMP 收录数据更新
已收集 skills
5
GitHub 星标
2
GitHub Forks
0

这个仓库中的 skills

1 个职业分类 · 已分类 80%

已展示 5 / 5 个已收集 Skill。

职业分类
未分类
描述

Maintain agent-eval cases, judges, records, traces, campaigns, comparisons, and releases.

原文语言:英语

更新
职业分类
软件开发工程师
描述

Parse classic SWE-agent .traj trajectories: one step per entry of the trajectory array, in order, matching the CodeTraceBench annotation convention. No upstream CodeTracer commit ships a SWE-agent parser.

原文语言:英语

更新
职业分类
软件开发工程师
描述

Parse OpenHands trajectories published as per-call LiteLLM completion logs (<model>-<timestamp>.json files with messages/response/kwargs/cost), the layout of CodeTraceBench's swe_raw OpenHands trials. Reads the call file whose messages hold the most tool-call…

原文语言:英语

更新
职业分类
软件开发工程师
描述

Parse OpenHands session-event trajectories at the CodeTraceBench annotation granularity: one step per action event with a cause-paired observation (run, run_ipython, read, edit, think, recall). Outranks the seed openhands skill, which counts only…

原文语言:英语

更新
职业分类
软件开发工程师
描述

Parse Terminus2 trajectories at command granularity: one step per entry of each episode response's commands array, matching the CodeTraceBench annotation convention. Outranks the seed terminus2 skill, whose episode-level steps do not align with the…

原文语言:英语

更新
已展示 5 / 5 个已收集 Skill。