Skip to main content

run-360-eval

Run LLM evaluations with 360-eval. Use when the user wants to benchmark, evaluate, score, or compare one or more LLMs (Amazon Bedrock, OpenAI, Gemini, Azure) on a dataset of prompts using LLM-as-a-jury scoring — including quality evals (correctness/completeness/format/etc. scored 1-5), latency/throughput benchmarks, vision/multimodal evals, multi-turn evals, and Bedrock Advanced Prompt Optimization (APO). Covers preparing inputs, running the engine CLI, reading results, and generating an HTML report. This skill drives the engine programmatically (no web UI needed).

跳到安装

来源信息

仓库
aws-samples/sample-bedrock-migration-and-modernization-tools
最近来源活动
2026年7月3日 15:08
检测到的 SKILL.md 语言
英语
星标
36
分支
13

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。