Skip to main content 홈 크리에이터 adu2021 skillxiv ruscarl-rubric-scaffolded-rl
ruscarl-rubric-scaffolded-rl Guide LLM exploration through rubric-based scaffolding that gradually diminishes, enabling models to internalize reasoning patterns while maintaining exploration quality for robust RL training.
설치로 이동 Skills Marketplace 커뮤니티가 만든 AI 스킬을 발견하고 탐색하세요.
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
직접 명령은 검토 Prompt를 거치지 않습니다. 실행하기 전에 소스를 확인하세요.
npx skills add https://github.com/ADu2021/skillXiv --skill ruscarl-rubric-scaffolded-rl명령은 한 줄로 유지됩니다. 복사하기 전에 가로로 스크롤해 전체 내용을 확인하세요.
로컬 사본을 원하시나요? SkillsMP에서 현재 제공할 수 있는 파일을 다운로드하세요.
Zip 다운로드 다운로드 중... 이 저장소의 다른 Skills meaningful-kebab-case-name Convert arXiv papers into ready-to-use agent skills using category-aware extraction. First classifies the paper into one or more of 11 research categories, then applies a specialized extraction pipeline for each category — because different types of papers produce different types of usable knowledge. A single paper can yield multiple skills if it spans categories. Use this skill whenever the user wants to turn a paper into a skill, extract practical techniques from research, build a skill library from papers, convert arXiv papers into reusable agent instructions, or batch-process multiple papers into skills. Also trigger when someone asks about extracting actionable knowledge from papers, making research practical for LLM agents, or systematically converting academic contributions into structured agent capabilities.
action-quantization-behavior-cloning Establish regret bounds for behavior cloning with discretized actions combining statistical error and quantization error terms. Prove smoothness requirements for safe quantizer design, show that learning-based quantizers fail these requirements, and propose model-based augmentation to reduce error dependence from H² to H.
adaptive-lora-personalized-ranks Dynamically allocate LoRA ranks per-layer during fine-tuning instead of using fixed uniform ranks. Learn optimal rank for each layer and subject via variational framework with discretized exponential distribution, reducing memory footprint while maintaining fidelity and text-alignment.
name ruscarl-rubric-scaffolded-rl title RuscaRL: Rubric-Scaffolded RL for General LLM Reasoning version 0.0.2 engine skillxiv-v0.0.2-claude-opus-4.6 license MIT url https://arxiv.org/abs/2508.16949 keywords ["reinforcement-learning","rubric-guidance","exploration-guidance","reasoning-scaffolding","curriculum-learning"] description Guide LLM exploration through rubric-based scaffolding that gradually diminishes, enabling models to internalize reasoning patterns while maintaining exploration quality for robust RL training.
RuscaRL: Rubric-Scaffolded Reinforcement Learning
Core Concept
The core challenge in LLM RL is exploration: what cannot be explored cannot be learned. RuscaRL (Rubric-Scaffolded RL) addresses this by using checklist-style rubrics in two phases: first as explicit guidance during exploration (gradually fading), then as reference for scoring. This approach breaks the constraint that LLM limitations restrict sample quality, enabling models to learn general reasoning capabilities through guided exploration.
Architecture Overview
Exploration Rubrics : Checklist guidance for high-quality response generation
Gradual Scaffolding Fade : Curriculum learning removing explicit guidance
Rubric-Based Scoring : Reference for reward computation
LLM-as-Judge : Automated evaluation using rubrics
General Reasoning Transfer : Patterns learned with rubrics transfer without them
Implementation Steps
1. Design Task-Specific Rubrics
Create structured evaluation checklist templates:
from dataclasses import dataclass
from typing import List , Dict , Any
from enum import Enum
class RubricLevel (Enum ):
EXPLICIT = "explicit"
IMPLICIT = "implicit"
ABSENT = "absent"
@dataclass
class RubricCriterion :
"""Single rubric evaluation criterion."""
name: str
description: str
exemplar_good: str
exemplar_bad: str
weight: float =
:
task_name:
criteria: [RubricCriterion]
instructions:
( ) -> :
guidance =
guidance +=
criterion .criteria:
guidance +=
guidance +=
guidance +=
guidance
:
( ):
.rubrics: [ , TaskRubric] = {}
( ):
.rubrics[task_type] = rubric
( ) -> TaskRubric:
.rubrics.get(task_type)
( ) -> TaskRubric:
TaskRubric(
task_name= ,
criteria=[
RubricCriterion(
name= ,
description= ,
exemplar_good= ,
exemplar_bad= ,
weight=
),
RubricCriterion(
name= ,
description= ,
exemplar_good= ,
exemplar_bad= ,
weight=
),
RubricCriterion(
name= ,
description= ,
exemplar_good= ,
exemplar_bad= ,
weight=
),
RubricCriterion(
name= ,
description= ,
exemplar_good= ,
exemplar_bad= ,
weight=
)
],
instructions=
)
1.0
@dataclass
class
TaskRubric
"""Complete rubric for a task."""
str
List
str
def
render_as_guidance
self
str
"""Render rubric as exploration guidance in prompt."""
f"Task: {self.task_name} \n\n"
"Evaluation Rubric (aim to satisfy these):\n"
for
in
self
f"\n- {criterion.name} : {criterion.description} \n"
f" Good example: {criterion.exemplar_good} \n"
f" Avoid: {criterion.exemplar_bad} \n"
return
class
RubricLibrary
"""Library of task-specific rubrics."""
def
__init__
self
self
Dict
str
def
register_rubric
self, task_type: str , rubric: TaskRubric
"""Register rubric for task type."""
self
def
get_rubric
self, task_type: str
return
self
def
create_math_rubric
self
"""Example: rubric for mathematical reasoning."""
return
"Mathematical Problem Solving"
"Clear Problem Understanding"
"State what is being asked before solving"
"We need to find the value of x. Given: 2x + 5 = 15"
"Just give the answer: x=5"
1.0
"Step-by-Step Solution"
"Show each calculation step explicitly"
"2x + 5 = 15\nSubtract 5: 2x = 10\nDivide by 2: x = 5"
"x = 5"
2.0
"Verification"
"Check the answer by substitution"
"Check: 2(5) + 5 = 10 + 5 = 15 ✓"
"(no verification)"
1.0
"Correct Final Answer"
"Clearly state the answer"
"Therefore, x = 5"
"(answer buried in work)"
3.0
"Solve the problem following the rubric above."
2. Implement Exploration with Rubric Guidance Use rubrics to enhance exploration quality:
class RubricGuidedExploration :
"""Guide exploration with rubric scaffolding."""
def __init__ (
self,
model: "LLM" ,
rubric_library: RubricLibrary
):
self .model = model
self .rubric_library = rubric_library
def generate_with_rubric_guidance (
self,
task: str ,
task_type: str ,
guidance_level: RubricLevel = RubricLevel.EXPLICIT,
temperature: float = 0.9
) -> str :
"""
Generate response with rubric guidance.
Guidance level controls how much structure to provide.
"""
rubric = self .rubric_library.get_rubric(task_type)
if guidance_level == RubricLevel.EXPLICIT:
prompt = f"{rubric.render_as_guidance()} \n\nProblem: {task} \n\nSolution:"
elif guidance_level == RubricLevel.IMPLICIT:
criterion_names = ", " .join(c.name for c in rubric.criteria)
prompt = f"Solve this task considering: {criterion_names} .\n\nProblem: {task} \n\nSolution:"
elif guidance_level == RubricLevel.ABSENT:
prompt = f"Solve this problem:\n\n{task} \n\nSolution:"
response = self .model.generate(
prompt,
temperature=temperature,
max_tokens=500
)
return response
def collect_exploration_trajectories (
self,
tasks: List [Dict ],
task_type: str ,
num_trajectories_per_task: int = 5 ,
guidance_schedule: List [RubricLevel] = None
) -> List [Dict ]:
"""
Collect exploration trajectories with curriculum guidance schedule.
"""
if guidance_schedule is None :
guidance_schedule = [
RubricLevel.EXPLICIT,
RubricLevel.EXPLICIT,
RubricLevel.IMPLICIT,
RubricLevel.IMPLICIT,
RubricLevel.ABSENT
]
trajectories = []
for task_idx, task in enumerate (tasks):
for traj_idx in range (num_trajectories_per_task):
schedule_idx = min (traj_idx, len (guidance_schedule) - 1 )
guidance_level = guidance_schedule[schedule_idx]
response = self .generate_with_rubric_guidance(
task["prompt" ],
task_type,
guidance_level
)
trajectories.append({
"task" : task["prompt" ],
"response" : response,
"guidance_level" : guidance_level,
"task_id" : task.get("id" ),
"expected_answer" : task.get("answer" )
})
return trajectories
3. Implement Rubric-Based Scoring Create LLM judge using rubrics:
class RubricBasedJudge :
"""Evaluate responses using rubric criteria."""
def __init__ (
self,
judge_model: "LLM" ,
rubric_library: RubricLibrary
):
self .judge_model = judge_model
self .rubric_library = rubric_library
def score_response (
self,
task: str ,
response: str ,
task_type: str
) -> Dict [str , float ]:
"""
Score response against rubric criteria.
Returns: {criterion_name: score, "overall": score}
"""
rubric = self .rubric_library.get_rubric(task_type)
scores = {}
for criterion in rubric.criteria:
score = self ._score_criterion(
task, response, criterion
)
scores[criterion.name] = score
overall = sum (
scores[c.name] * c.weight
for c in rubric.criteria
) / sum (c.weight for c in rubric.criteria)
scores["overall" ] = overall
return scores
def _score_criterion (
self,
task: str ,
response: str ,
criterion: RubricCriterion
) -> float :
"""Score single criterion using LLM judge."""
prompt = f"""Evaluate the following response against this criterion:
Criterion: {criterion.name}
Description: {criterion.description}
Good example: {criterion.exemplar_good}
Bad example: {criterion.exemplar_bad}
Task: {task}
Response: {response}
Rate this response on criterion '{criterion.name} ' on a scale 0-1:
- 0: Does not satisfy criterion at all
- 0.5: Partially satisfies criterion
- 1: Fully satisfies criterion
Score (just the number, 0-1):"""
score_text = self .judge_model.generate(prompt, max_tokens=10 )
try :
score = float (score_text.strip())
return min (max (score, 0.0 ), 1.0 )
except :
return 0.5
def score_batch (
self,
trajectories: List [Dict ],
task_type: str
) -> List [Dict ]:
"""Score multiple trajectories."""
for trajectory in trajectories:
scores = self .score_response(
trajectory["task" ],
trajectory["response" ],
task_type
)
trajectory["scores" ] = scores
trajectory["reward" ] = scores["overall" ]
return trajectories
4. Implement RuscaRL Training Loop Train with rubric-guided RL:
class RuscaRLTrainer :
"""Training with rubric-scaffolded RL."""
def __init__ (
self,
model: "LLM" ,
rubric_library: RubricLibrary,
learning_rate: float = 1e-5
):
self .model = model
self .rubric_library = rubric_library
self .explorer = RubricGuidedExploration(model, rubric_library)
self .judge = RubricBasedJudge(model, rubric_library)
self .optimizer = torch.optim.Adam(model.parameters(), lr=learning_rate)
def train_with_rubric_curriculum (
self,
tasks: List [Dict ],
task_type: str ,
num_iterations: int = 5 ,
guidance_schedule: List [RubricLevel] = None
) -> Dict [str , float ]:
"""
Train model with rubric curriculum.
Gradually remove rubric guidance over iterations.
"""
metrics = {
"iteration" : [],
"avg_reward" : [],
"without_guidance_reward" : []
}
for iteration in range (num_iterations):
print (f"RuscaRL Iteration {iteration + 1 } /{num_iterations} " )
trajectories = self .explorer.collect_exploration_trajectories(
tasks,
task_type,
num_trajectories_per_task=3 ,
guidance_schedule=guidance_schedule
)
trajectories = self .judge.score_batch(trajectories, task_type)
rewards = [t["reward" ] for t in trajectories]
avg_reward = sum (rewards) / len (rewards)
metrics["iteration" ].append(iteration)
metrics["avg_reward" ].append(avg_reward)
print (f" Avg Reward: {avg_reward:.4 f} " )
for trajectory in trajectories:
response = trajectory["response" ]
reward = trajectory["reward" ]
loss = -reward * self .model.get_log_prob(response)
self .optimizer.zero_grad()
loss.backward()
torch.nn.utils.clip_grad_norm_(self .model.parameters(), 1.0 )
self .optimizer.step()
print (f" Evaluating without guidance..." )
guidance_free_trajectories = self .explorer.collect_exploration_trajectories(
tasks,
task_type,
num_trajectories_per_task=1 ,
guidance_schedule=[RubricLevel.ABSENT] * 5
)
guidance_free_trajectories = self .judge.score_batch(
guidance_free_trajectories, task_type
)
guidance_free_rewards = [t["reward" ] for t in guidance_free_trajectories]
avg_guidance_free = sum (guidance_free_rewards) / len (guidance_free_rewards)
metrics["without_guidance_reward" ].append(avg_guidance_free)
print (f" Without Guidance Reward: {avg_guidance_free:.4 f} " )
return metrics
Practical Guidance
When to Use RuscaRL
General reasoning tasks without fixed rubrics
Curriculum learning scenarios
Exploration-heavy RL problems
Multi-criterion optimization
Tasks where quality evaluation is clear
When NOT to Use
Tasks without clear evaluation criteria
Real-time systems (rubric scoring is slow)
Domains where guidance misleads exploration
Single-criterion optimization
Key Hyperparameters
guidance_schedule : Length equal to trajectory budget
temperature : 0.8-1.0 for exploration
rubric weights : Adjust by criterion importance
judge model : Can be same or separate from student
curriculum fade rate : Linear or exponential
Performance Expectations
Exploration Quality: 2-3x improvement with guidance
Guidance Transfer: Patterns learned with guidance apply without it
Sample Efficiency: Faster convergence than unguided RL
Final Performance: Approaches or exceeds baseline after guidance removal
Reference Researchers. (2024). Breaking the Exploration Bottleneck: Rubric-Scaffolded RL for General LLM Reasoning. arXiv preprint arXiv:2508.16949.