Skip to main content

r-vitals

Use when code loads or uses vitals (library(vitals), vitals::), evaluating LLM output quality, scoring AI responses, testing RAG retrieval accuracy, or benchmarking prompt changes in R

インストールへ移動

ソース情報

リポジトリ
arthurgailes/r-package-skills
ソースの最終更新活動
2026年4月24日 15:12
検出された SKILL.md の言語
英語
スター
21
フォーク
1

インストール方法

デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。

ソースファイルを確認

インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。

ファイルエクスプローラー
3 ファイル

SKILL.md を表示中

SKILL.md
ソースの指示 · 読み取り専用プレビュー
name
r-vitals
description
Use when code loads or uses vitals (library(vitals), vitals::), evaluating LLM output quality, scoring AI responses, testing RAG retrieval accuracy, or benchmarking prompt changes in R
# vitals: LLM Evaluation and Testing ## Overview **vitals tests LLM output quality.** Create test datasets, define solvers (LLM pipelines), score outputs. Benchmark RAG systems, prompt changes, model performance. **Install:** `install.packages("vitals")` ## References Read `references/API.md` before writing code. - `references/API.md` - Complete function reference - `references/package-docs.md` - Test suite creation and scoring patterns ## When to Use - Test LLM output quality - Evaluate RAG retrieval accuracy - Benchmark different prompts/models - Score AI-generated responses - Create LLM test suites ## When NOT to Use - Just running one test case manually - Traditional unit testing (use testthat) - Non-LLM code testing ## Quick Reference ```r library(vitals) # Create test dataset test_cases <- tibble::tibble( input = c("question 1", "question 2"), target = c("answer 1", "answer 2") ) # Define solver (your LLM pipeline) chat <- chat_openai() solver <- function(input) { chat$chat(input, echo = "none") } # Run evaluation task <- Task$new( dataset = test_cases, solver = solver, scorer = model_graded_qa() ) task$run() task$view() # Test RAG system ragnar_register_tool_retrieve(chat, store) task$run(chat) # Tests with RAG ``` ## Common Mistakes | Issue | Solution | |-------|----------| | No test dataset | Create tibble with input/target columns | | Solver not function | Wrap chat in function: `function(input) chat$chat(input)` | | Using for non-LLM tests | Use testthat for traditional testing | | Forgetting echo = "none" | Solver should return text, not print | ## Core Functions **Task Management:** - `Task$new()`: Create evaluation task - `task$run()`: Execute tests - `task$view()`: View results **Scorers:** - `model_graded_qa()`: LLM grades Q&A quality - Custom scorers for specific domains ## Advanced See `references/` for: - **API.md**: Complete function reference - **Package docs**: Full package documentation ## Integration **With ellmer:** Test chat quality **With ragnar:** Evaluate RAG accuracy **Cross-package patterns:** See r-ai meta-skill
GitHubで見る