Verify a tested web agent's real behavior from its recorded Pi or Codex session and selected screenshots before scoring. The judge prompt provides absolute artifact paths; use targeted queries to compare actual outputs, errors, and the final answer with the…
citrolabs/ego-browser-benchmark-framework
SkillsMP has collected 9 skills from citrolabs/ego-browser-benchmark-framework. Open a skill to review its source and details.
- Latest recorded source activity
- SkillsMP catalog refreshed
- skills collected
- 9
- GitHub stars
- 6
- GitHub forks
- 0
Skills in this repository
Showing 9 of 9 collected skills.
ego-browser (ego-lite) is a Chromium-based browser that gives AI agents a CLI-accessible Node.js runtime for driving a real browser. Use this skill whenever the user needs to interact with a website opening pages, filling forms, clicking buttons, taking…
Author or expand browser-agent benchmark tasks in data/real_world_bench.json (Odysseys schema, per-rubric graded). Use when creating, adding, or expanding real-world browser tasks, writing rubrics, or picking bot-friendly sites — every task MUST be verified…
Analyze Codex provider rollout JSONL and ego-benchmark-harness run artifacts to identify slow or failed individual tasks, tool retries, token/cost anomalies, and differences in agent execution paths. Use when users ask to analyze a Codex session or rollout,…
Export Codex rollout/session JSONL as a self-contained HTML report that can be opened offline. Use when a user provides a complete session file path or Codex session ID, or asks to export a Codex session to HTML, generate a session report, or convert a…
Generate self-contained HTML/CSS benchmark and comparison charts in an ego-inspired Browser-Native Editorial style. Use for blue-white product metrics, model benchmarks, A/B comparisons, efficiency reports, difficulty breakdowns, and compact technical data…
Architecture design and governance protocol for system structure, boundaries, contracts, APIs, schemas, migrations, rewrites, and long-lived codebase health. Use when designing or reviewing architecture, planning structural refactors or migrations, governing…
Analyze pi (coding-agent) session JSONL and ego-benchmark-harness run artifacts to explain slow or failed individual tasks, compare execution-path differences between agents or browser tools, and count tool calls and error types. Use this skill whenever users…
Compare standardized metrics across multiple ego-benchmark-harness runs by recalculating them from tasks.jsonl and session JSONL, never from HTML reports. Covers overall aggregates (success rate/average score), Odysseys per-rubric pass rates, duration…