Skip to main content

hebrew-llm-eval-suite

Stars11
Forks4
UpdatedJune 18, 2026 at 21:12

Benchmark and compare LLMs on Hebrew reasoning, comprehension, sentiment, translation, and Israeli cultural knowledge. Wraps the HuggingFace Open Hebrew LLM Leaderboard tasks (HeQ reading comprehension, HebrewSentiment, Hebrew Winograd, NeuLabs-TedTalks translation) plus DictaLM 3.0 benchmark tasks (Summarization, Nikud diacritization, Israeli Trivia) into a reproducible evaluation harness. Runs evals against Claude, GPT, Gemini, AI21 Jamba, DictaLM, Llama, and local HuggingFace models. Produces comparison scorecards in JSON and markdown with per-task breakdowns. Use when choosing an LLM for a Hebrew product, answering procurement questions about Hebrew performance, validating a fine-tuned Hebrew model, or tracking Hebrew regressions after a model upgrade. Do NOT use for Arabic NLP evaluation, speech recognition benchmarking (use ivrit.ai leaderboard for ASR), or general English LLM benchmarks. Activate for: ื‘ื ืฆ'ืžืจืง ืขื‘ืจื™ืช, ื”ืฉื•ื•ืืช ืžื•ื“ืœื™ื, ื”ืขืจื›ืช ืžื•ื“ืœ ืฉืคื”, ื‘ื™ืฆื•ืขื™ ืขื‘ืจื™ืช, ืžื‘ื—ืŸ ื”ื‘ื ืช ื”ื ืงืจื, ื ื™ืชื•ื— ืจื’ืฉื•ืช, ื‘ื—ื™ืจืช ืžื•ื“ืœ ืœ

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

File Explorer
11 files
SKILL.md
readonly