Port an existing third-party benchmark or eval into this repo as an Inspect AI task, faithfully. Use this whenever the task is "add benchmark X", "port this eval", "can we run <github repo> here", or adapting any external eval harness (aiewf-eval,…
awslabs/llm-evaluation-system
SkillsMP has collected 3 skills from awslabs/llm-evaluation-system. Open a skill to review its source and details.
- Latest recorded source activity
- SkillsMP catalog refreshed
- skills collected
- 3
- GitHub stars
- 22
- GitHub forks
- 3
Skills in this repository
1 occupation categories · 100% classified
Showing 3 of 3 collected skills.
skill
occupation
description
updated
occupation
Software Developers
description
updated
occupation
Software Developers
description
Use when the user asks to create, update, or render an AWS architecture diagram (or a cloud/infrastructure diagram using AWS service icons) — e.g. "diagram our AWS setup", "architecture diagram with CloudFront/S3/EKS/RDS", "update docs/image.png", "make an…
updated
occupation
Software Developers
description
Ship work in the llm-evaluation-system repo end-to-end — commit with conventional-commit messages, push to a feature branch (never directly to main), open a PR with the proper title format, and after merge either run `make release` to publish to PyPI or just…
updated
Showing 3 of 3 collected skills.