Skip to main content

run-360-eval

Run LLM evaluations with 360-eval. Use when the user wants to benchmark, evaluate, score, or compare one or more LLMs (Amazon Bedrock, OpenAI, Gemini, Azure) on a dataset of prompts using LLM-as-a-jury scoring — including quality evals (correctness/completeness/format/etc. scored 1-5), latency/throughput benchmarks, vision/multimodal evals, multi-turn evals, and Bedrock Advanced Prompt Optimization (APO). Covers preparing inputs, running the engine CLI, reading results, and generating an HTML report. This skill drives the engine programmatically (no web UI needed).

インストールへ移動

ソース情報

リポジトリ
aws-samples/sample-bedrock-migration-and-modernization-tools
ソースの最終更新活動
2026年7月3日 15:08
検出された SKILL.md の言語
英語
スター
36
フォーク
13

インストール方法

デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。

ソースファイルを確認

インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。