Skip to main content

benchmark-vlm-qa

Benchmark video Q&A accuracy and latency of a deployed RT-VLM (Cosmos Reason 3) via vss vlm run, using questions and videos from the DSS vss-devx-base dataset. Replaces the deprecated nat eval / vss-agent QA path. Not for tool-calling or trajectory evaluation, and not for LVS summarization throughput.

설치로 이동

소스 정보

저장소
NVIDIA-AI-Blueprints/video-search-and-summarization
최근 소스 활동
2026년 9월 14일 12:55
감지된 SKILL.md 언어
영어
스타
1,865
포크
390

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

파일 탐색기
5 개 파일

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
benchmark-vlm-qa
description
Benchmark video Q&A accuracy and latency of a deployed RT-VLM (Cosmos Reason 3) via vss vlm run, using questions and videos from the DSS vss-devx-base dataset. Replaces the deprecated nat eval / vss-agent QA path. Not for tool-calling or trajectory evaluation, and not for LVS summarization throughput.
license
Apache-2.0
metadata
{"version":"3.3.0","author":"NVIDIA Video Search and Summarization Team","github-url":"https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization","tags":"nvidia blueprint performance benchmarking vlm qa"}
# Benchmark video Q&A via `vss vlm` Measure **accuracy** (LLM-as-judge vs ground truth) and **latency** of end-to-end video question answering by calling **`vss vlm run`** against a deployed Cosmos Reason 3 RT-VLM. Questions and clips come from DSS dataset **`vss-devx-base`** (`nvdataset`). This replaces `docker exec vss-agent nat eval` for the QA slice. It does **not** score tool-calling or trajectories. ## When to use - The user asks to benchmark / evaluate VLM video Q&A after vss-agent / NAT eval was removed. - The user wants latency and answer accuracy on `vss-devx-base`. ## When not to use - Tool-calling or trajectory evaluation — out of scope. - LVS summarization throughput — use `benchmark-video-summarization`. - Ad-hoc single questions — use `/vss-ask-video`. ## Prerequisites - A VSS stack with RT-VLM serving Cosmos Reason 3, and `vss configure` already run so `vss configure check` lists `rt_vlm` as `ok` and `vst` as `ok`. **Configure with a routable address, not `localhost`.** Clips are addressed as VIOS sensors so RT-VLM fetches them by URL; the URL VIOS mints is built from the configured origin. A loopback origin mints a loopback URL, which means nothing inside the RT-VLM container, so the CLI falls back to inlining the clip as base64 and the VLM rejects anything large with `HTTP 422 ... content ... valid string`. `vss configure --base-url http://<host-ip>:7777` avoids that — `--base-url` is a `vss configure` flag, not a benchmark one. `--inline-media` *is* a benchmark flag; it forces the old inline behaviour and is only safe for clips under ~10 MB. - `uv` and this checkout (CLI via `uv run --project libs/vss vss`). - The `nvdataset` CLI. It is **not** on PyPI, and the index used by the old deep-search eval (`urm.nvidia.com/.../sw-ngc-data-platform-pypi`) returns 403. Install from the documented read-only index instead — no credentials needed: ```bash uv tool install --index https://artifactory.pdx.nvidia.com/artifactory/api/pypi/sw-ngc-data-platform-pypi-local/simple nvdataset ``` - DSS access, one of: - `NVDATASET_API_KEY` — a **Personal Key** from [org.ngc.nvidia.com/setup/personal-keys](https://org.ngc.nvidia.com/setup/personal-keys) scoped to the service `NVIDIA Dataset Service`, with the NGC org switched to the one owning the dataset. This is *not* the global NGC key used by the NGC CLI; a global key returns 403. `NVDATASET_NGC_API_KEY` and `NGC_API_KEY` are also read, in that order, for backward compatibility only — the run prints the variable it picked as `dss credential: <name>`, so check that line if a 403 surprises you. - `nvdataset auth login` (Starfleet SSO), which needs no key. Add `--flow device` on a remote box with no browser. Group access requires membership in `ngc-datasetservice-viewer-<tenant>-<group>` (reader) or `...-user-...` (writer). Plus tenancy, which SSO does **not** supply — after `auth login`, `nvdataset auth status` still reports `"tenant_id": null` and every call fails with `Did not find tenant_id`. The script names no tenant, so set one yourself: export `NVDATASET_TENANTID` and `NVDATASET_GROUPID`, or save them once with `nvdataset auth context add`. Ask the dataset's owning team for its coordinates. Another dataset needs no change to the script. - An OpenAI-compatible judge LLM: `EVAL_LLM_JUDGE_BASE_URL` and `EVAL_LLM_JUDGE_NAME`, authenticated with `EVAL_LLM_JUDGE_API_KEY`. `NGC_API_KEY` is deliberately **not** sent to non-NVIDIA judge hosts — it is set for the dataset download and must not reach a third party. Any chat-completions endpoint will do; the judge moves absolute scores on its own, so hold it fixed across runs you mean to compare, and read `judge_model` in `summary.json` before comparing two numbers. `--skip-judge` gives latency only. Bootstrap is in the repo-root [AGENTS.md](../../../AGENTS.md). Do not construct RT-VLM URLs; `vss vlm run` reads the recorded config. ## Run ```bash export NVDATASET_API_KEY=<personal-key> # or: nvdataset auth login [--flow device] export NVDATASET_TENANTID=<tenant> # SSO does not set this; see Prerequisites export NVDATASET_GROUPID=<group> export EVAL_LLM_JUDGE_BASE_URL="${LLM_BASE_URL}" # OpenAI-compat origin, e.g. http://127.0.0.1:8000 export EVAL_LLM_JUDGE_NAME="${LLM_NAME}" # Optional: already-extracted dataset # export VSS_EVAL_DATASET=/path/to/vss-devx-base <repo>/skills/benchmarking/benchmark-vlm-qa/scripts/run_vlm_qa_benchmark.sh \ --dataset-name vss-devx-base \ --dataset-file dataset_single_turn.json ``` Both dataset flags are **required** — the script carries no default dataset, so it never assumes one team's DSS coordinates. Useful flags (forwarded to `benchmark_vlm_qa.py`): | Flag | Purpose | |---|---| | `--dry-run` | Resolve QA items and video files; no VLM calls | | `--limit N` | First N QA items (smoke) | | `--skip-judge` | Latency only | | `--skip-download` | Use an already-downloaded `vss-devx-base` | | `--timeout SEC` | Passed through as `vss vlm run --timeout` (default 300) | | `--num-frames N` | Frame budget (default 20, matching the old RT-VLM agent config) | | `--model ID` | Override the RT-VLM model `vss configure` recorded | Outputs under `<dataset>/../../results/vlm_qa/` (or `--output-dir`): - `summary.json` — mean accuracy, latency mean / p50 / p90 / p95 / p99, and the model the deployment reported serving, so a number is never left unattributable - `qa_evaluator_output.json` — per-item judge scores (same shape as NAT QA output) - `latency_summary.json` — per-item wall-clock around `vss vlm run` - `workflow_output.json` — raw answers - `summary.csv` ## Rules - Drive the VLM only through `vss vlm run`. Never `POST /generate` or hand-built `/v1/chat/completions`. - Do not wrap `vss` in retries. `--timeout` is the bound; the script adds only a hard kill 60 s past it, so a CLI that never returns cannot cost the whole run. A killed item is recorded as an error naming the watchdog, never as a low score. - Items must declare `evaluation_method` containing `qa` and carry a text `ground_truth`. Report, trajectory-only, and unmarked items are skipped. ## Failures Branch on the exit code; never scrape stdout for the word "error". | Exit | Meaning | What to do | |---|---|---| | 0 | Every item answered | Read `summary.json` | | 2 | Precondition wrong — a dataset flag missing, no DSS credential, no judge configured, dataset or videos not found, no QA items | Fix the setup. Re-running unchanged fails identically | | 3 | The download failed, or at least one item errored | Read each item's `error` in `summary.json` | A `vss` call that exits 4 (service missing from the recorded config) surfaces as an item error, so the run ends at exit 3 — the fix is `vss configure`, not a flag. Failures worth recognising by their message: - `HTTP 422 ... content ... valid string` on the big clips — the recorded origin is loopback, so clips are being inlined as base64. Reconfigure with a routable address. - `Did not find tenant_id` — SSO signed you in but selected no tenant. Export `NVDATASET_TENANTID`, or `nvdataset auth context use`. - `LLM judge HTTP 403 ... key_model_access_denied` or `400 Invalid model name` on every item — the judge id is not what that gateway calls the model. Gateways that front several providers usually want a fully-qualified id and reject the bare name. `GET <judge-base-url>/models` lists the ids the key may use; copy one verbatim into `EVAL_LLM_JUDGE_NAME`. The VLM answers are unaffected, so only scoring is lost. - An item error naming the watchdog — the CLI never returned and was killed at `--timeout` + 60 s. That is recorded as an error, never as a low score. Do not retry. - Accuracy far from the ~0.465 baseline is not a harness failure. The judge model and `--num-frames` both move it; check `judge_model` and `model_served` before filing. Implementation: [`scripts/benchmark_vlm_qa.py`](scripts/benchmark_vlm_qa.py), tested by [`scripts/tests/`](scripts/tests/). Dataset download contract: [`README_eval.md`](../../../deploy/docker/developer-profiles/dev-profile-base/eval/README_eval.md).
GitHub에서 보기