Skip to main content

run-360-eval

Run LLM evaluations with 360-eval. Use when the user wants to benchmark, evaluate, score, or compare one or more LLMs (Amazon Bedrock, OpenAI, Gemini, Azure) on a dataset of prompts using LLM-as-a-jury scoring — including quality evals (correctness/completeness/format/etc. scored 1-5), latency/throughput benchmarks, vision/multimodal evals, multi-turn evals, and Bedrock Advanced Prompt Optimization (APO). Covers preparing inputs, running the engine CLI, reading results, and generating an HTML report. This skill drives the engine programmatically (no web UI needed).

الانتقال إلى التثبيت

معلومات المصدر

المستودع
aws-samples/sample-bedrock-migration-and-modernization-tools
آخر نشاط في المصدر
٣ يوليو ٢٠٢٦ في ١٥:٠٨
لغة SKILL.md المكتشفة
الإنجليزية
النجوم
٣٦
التفرعات
١٣

خيارات التثبيت

يُحدَّد Prompt الذي يراجع المصدر أولًا بشكل افتراضي. يمكنك التبديل إلى أمر مباشر أو تنزيل نسخة محلية.

مراجعة ملفات المصدر

اقرأ SKILL.md وأي ملفات مرافقة يعرضها SkillsMP قبل أن تقرر التثبيت.