Skip to main content

appworld-leaderboard

Run, resume, retry and pack a CUGA AppWorld leaderboard submission (test_normal + test_challenge) batch-by-batch from eval_config.toml. Use when asked to run the full AppWorld test splits, continue an interrupted AppWorld run, retry failed/uncompleted AppWorld tasks, check AppWorld SGC/TGC, or produce leaderboard.bundle files.

Zur Installation springen

Quellinformationen

Repository
cuga-project/cuga-eval
Letzte Quellaktivität
3. September 2026 um 07:47
Erkannte Sprache von SKILL.md
Englisch
Sterne
1
Forks
1

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.

SKILL.md wird angezeigt

SKILL.md
Quellanweisungen · Schreibgeschützte Vorschau
name
appworld-leaderboard
description
Run, resume, retry and pack a CUGA AppWorld leaderboard submission (test_normal + test_challenge) batch-by-batch from eval_config.toml. Use when asked to run the full AppWorld test splits, continue an interrupted AppWorld run, retry failed/uncompleted AppWorld tasks, check AppWorld SGC/TGC, or produce leaderboard.bundle files.
trigger
/appworld-leaderboard
# AppWorld leaderboard flow Everything is driven by keys in `benchmarks/appworld/eval_config.toml`. One cuga-eval workspace (`benchmarks/appworld/evaluation_bundles/<name>`) and one AppWorld experiment directory (`benchmarks/appworld/appworld/experiments/outputs/<prefix>_<split>`) per split. Never create a second workspace for the same prefix+split. ## 0. Prepare batch keys (once per split) uv run python -m benchmarks.appworld.leaderboard split-key test_challenge_all --batch-size 100 uv run python -m benchmarks.appworld.leaderboard split-key test_normal_all --batch-size 100 Writes `test_challenge_all_b1..b5` (100/100/100/100/17) and `test_normal_all_b1..b2` (100/68); scenarios `_1/_2/_3` of a base always stay in the same batch. ## 1. First batch ./benchmarks/appworld/eval.sh --sdk --experiment cuga_v1_chal --leaderboard cuga_v1 \ --eval-key test_challenge_all_b1 --background Watch: `./benchmarks/appworld/eval.sh --status --resume-experiment cuga_v1_chal` → `cuga_v1_chal split=test_challenge completed 100/417 errored 0 score<1: 31 missing 317` Direct form: `uv run python -m benchmarks.appworld.leaderboard status --bundle-dir benchmarks/appworld/evaluation_bundles/cuga_v1_chal` The console/background.log also prints `cuga-viz experiment: <card name>`. ## 2. Inspect in cuga-viz (http://localhost:8988/) Open the card named in the log. **Uncompleted** = never finished (kill, crash). **Failed** in cuga-viz only lists score == 0.0 and misses AppWorld's fractional scores — prefer the harness list: uv run python -m benchmarks.appworld.leaderboard retry-key errored --bundle-dir benchmarks/appworld/evaluation_bundles/cuga_v1_chal --of-key test_challenge_all_b1 uv run python -m benchmarks.appworld.leaderboard retry-key uncompleted --bundle-dir benchmarks/appworld/evaluation_bundles/cuga_v1_chal --of-key test_challenge_all_b1 Either command appends a key like `cuga_v1_chal_errored = [...]` to eval_config.toml. Pasting the cuga-viz line (`<card>_uncompleted_tasks = [...]`) into the toml works too. Decide what to retry: open a failed task's trajectory; timeout / connection reset / 5xx / empty LLM reply → retry. A genuine agent mistake is NOT retried on a leaderboard run (one attempt per task). ## 3. Retry (same workspace, same AppWorld dir) ./benchmarks/appworld/eval.sh --resume-experiment cuga_v1_chal --eval-key cuga_v1_chal_errored A key that `retry-key` wrote for **this** workspace re-runs every id even if its partial is clean (the workspace records it under `retry_keys` in `metadata.json`). A key that merely ends in `_failed` but was not recorded is an ordinary key. For a hand-written key add `--force-retry`. ## 4. Next batches ./benchmarks/appworld/eval.sh --resume-experiment cuga_v1_chal --eval-key test_challenge_all_b2 --background # inspect / retry, then b3, b4, b5 Batch keys skip ids that already completed. Ids must belong to the workspace's split or the run aborts. ## 5. Validate + official numbers uv run python -m benchmarks.appworld.leaderboard validate cuga_v1 --split test_challenge uv run python -m benchmarks.appworld.leaderboard evaluate cuga_v1_test_challenge --split test_challenge \ --bundle-dir benchmarks/appworld/evaluation_bundles/cuga_v1_chal `validate` exits 1 on missing tasks/files/scenarios. SDK eval copies ToolCallTracker records into `environment_io.md` / `api_calls.jsonl` after invoke (HTTP still goes to port 9111; the APIs are not re-executed). Pass `--allow-low-interactions` only when the ≤1-interaction tasks are one of: - the task really made no AppWorld API call besides `complete_task` (crash / no-op), or - the known logging gap: the merge could not run or found nothing to copy, so only the harness `complete_task` interaction was recorded even though the agent did call APIs. It never silences a missing-task, missing-file or missing-scenario failure — those always exit 1. `evaluate` prints TGC + SGC by difficulty and writes them into the workspace `report.md` under "AppWorld official metrics". ## 6. Pack both splits ./benchmarks/appworld/pack_leaderboard.sh cuga_v1 "CUGA" "CUGA lite via SDK" "gpt-4.1" "gpt-4.1-2025-04-14" \ https://github.com/cuga-project/cuga-agent Refuses unless both splits validate — or just the one split, with `--only test_normal` / `--only test_challenge`. Then runs `appworld evaluate` (skipped when `evaluations/<split>.json` is already newer than every task output; `--re-evaluate` forces it) and `appworld pack`, unpacks the bundle into a temp dir and byte-compares every file; prints the two `leaderboard.bundle` paths and the `/add-to-leaderboard --python … --appworld … cuga_v1` comment for the PR. If you already ran the step-5 `evaluate` for a split, this reuses it rather than re-running it — a full `evaluate` on `test_challenge` loads the start/end DBs of all 417 tasks. ## Do not - Rename an AppWorld experiment dir after packing (the bundle then refuses to unpack). - Run `--task` for leaderboard retries; use a toml key so the attempt is recorded in `resume_history`. - Trust `appworld pack` output alone: it prints WARNINGs and still writes the bundle, and says nothing about absent task dirs. Only `pack_leaderboard.sh` / `leaderboard pack` verify.
Auf GitHub ansehen