Skip to main content الرئيسية المنشئون adu2021 skillxiv code-a1-adversarial-rl-code
code-a1-adversarial-rl-code Train code and test generators through adversarial co-evolution where test LLM generates adversarial test cases to expose code defects. Prevent self-collusion by separating models and enabling white-box test generation.
الانتقال إلى التثبيت سوق المهارات اكتشف واستكشف مهارات الذكاء الاصطناعي التي بناها المجتمع.
التثبيت باستخدام Codex أو Claude انسخ هذا Prompt والصقه في Codex أو Claude أو مساعد آخر ليراجع صفحة Skill ويثبّتها لك.
نسخ Promptعرض تفاصيل Prompt يتجاوز الأمر المباشر Prompt المخصّص للمراجعة. افحص المصدر قبل تشغيله.
npx skills add https://github.com/ADu2021/skillXiv --skill code-a1-adversarial-rl-codeيبقى الأمر في سطر واحد. مرّر أفقيًا لمراجعته كاملًا قبل النسخ.
تفضّل نسخة محلية؟ نزّل الملفات المتاحة حاليًا لدى SkillsMP.
تحميل Zip جاري التحميل... المزيد من هذا المستودع meaningful-kebab-case-name Convert arXiv papers into ready-to-use agent skills using category-aware extraction. First classifies the paper into one or more of 11 research categories, then applies a specialized extraction pipeline for each category — because different types of papers produce different types of usable knowledge. A single paper can yield multiple skills if it spans categories. Use this skill whenever the user wants to turn a paper into a skill, extract practical techniques from research, build a skill library from papers, convert arXiv papers into reusable agent instructions, or batch-process multiple papers into skills. Also trigger when someone asks about extracting actionable knowledge from papers, making research practical for LLM agents, or systematically converting academic contributions into structured agent capabilities.
action-quantization-behavior-cloning Establish regret bounds for behavior cloning with discretized actions combining statistical error and quantization error terms. Prove smoothness requirements for safe quantizer design, show that learning-based quantizers fail these requirements, and propose model-based augmentation to reduce error dependence from H² to H.
adaptive-lora-personalized-ranks Dynamically allocate LoRA ranks per-layer during fine-tuning instead of using fixed uniform ranks. Learn optimal rank for each layer and subject via variational framework with discretized exponential distribution, reducing memory footprint while maintaining fidelity and text-alignment.
المهن ذات الصلة SOC
استنادا إلى تصنيف SOC المهني
name code-a1-adversarial-rl-code title Code-A1: Adversarial Evolving of Code LLM and Test LLM via RL version 0.0.2 engine skillxiv-v0.0.2-claude-opus-4.6 license MIT url https://arxiv.org/abs/2603.15611 keywords ["Code Generation","Reinforcement Learning","Adversarial Training","Test Generation","Self-Collusion Prevention"] description Train code and test generators through adversarial co-evolution where test LLM generates adversarial test cases to expose code defects. Prevent self-collusion by separating models and enabling white-box test generation.
Code-A1: Adversarial Co-Evolution for Code and Test Generation
Standard reinforcement learning for code generation suffers from self-collusion: a single model learns to generate trivial tests to easily pass, or creates generic tests that miss implementation-specific bugs. Code-A1 solves this through architectural separation—a Code LLM generates implementations while a Test LLM generates adversarial tests designed to expose defects. This adversarial co-evolution enables the Code LLM to improve robustly against high-quality tests without gaming the reward signal, matching or exceeding performance on human-annotated test benchmarks.
The technique combines architectural separation with two stabilizing mechanisms: a mistake book preserving high-quality test cases, and composite rewards balancing test validity against difficulty.
Core Concept
The adversarial co-evolution loop operates as:
Code Generation — Code LLM proposes solution implementations
Adversarial Test Creation — Test LLM inspects code to craft targeted tests exposing bugs
Execution & Reward — Evaluate code against tests; track both code and test quality
Co-Evolution — Update both models, Code LLM to pass tests, Test LLM to expose defects
Stability Mechanisms — Preserve high-quality tests and prevent reward gaming
The architectural separation ensures the Test LLM cannot trivially satisfy itself, forcing generation of genuinely challenging test cases.
Architecture Overview
Separate LLM Instances — Independent Code LLM and Test LLM models, each optimized for different objectives
White-Box Test Generation — Test LLM accesses code implementation to craft targeted adversarial tests
Mistake Book — Experience replay mechanism storing high-quality failing test cases
Composite Reward — Combines test validity (code actually passes) with adversarial difficulty (test exposes bugs)
Trial Execution Environment — Sandboxed runtime for safe code and test execution
Quality Filtering — Reject trivial tests and syntactically invalid code before inclusion in training
Implementation Steps
Start by setting up the adversarial reward structure that scores both code and test quality.
import ast
import subprocess
from typing import
:
( ):
.validity_weight = validity_weight
.difficulty_weight = difficulty_weight
( ) -> :
._is_valid_code(code):
-
passed =
failed =
test tests:
result = ._run_test(code, test)
result[ ]:
passed +=
:
failed +=
passed / (passed + failed, )
( ) -> :
._is_valid_test(test, code):
-
validity_result = ._run_test(reference_code code, test)
validity_score = validity_result[ ]
candidate_result = ._run_test(code, test)
difficulty_score = candidate_result[ ]
( .validity_weight * validity_score +
.difficulty_weight * difficulty_score)
( ) -> :
:
ast.parse(code)
SyntaxError:
( ) -> :
:
ast.parse(test)
test test
SyntaxError:
( ) -> :
full_script =
:
result = subprocess.run([ , , full_script],
capture_output= ,
timeout=timeout,
text= )
{
: result.returncode == ,
: result.stdout,
: result.stderr
}
subprocess.TimeoutExpired:
{ : , : , : }
Tuple
class
AdversarialReward
"""Compute rewards for code and test quality in adversarial setting."""
def
__init__
self, validity_weight=0.6 , difficulty_weight=0.4
self
self
def
score_code
self, code: str , tests: list
float
"""Score code by tests passed / failed."""
if
not
self
return
1.0
0
0
for
in
self
if
'success'
1
else
1
return
max
1
def
score_test
self, test: str , code: str , reference_code: str = None
float
"""Score test by validity and difficulty (finding bugs)."""
if
not
self
return
0.5
self
or
1.0
if
'success'
else
0.0
self
1.0
if
not
'success'
else
0.0
return
self
self
def
_is_valid_code
self, code: str
bool
"""Check if code is syntactically valid Python."""
try
return
True
except
return
False
def
_is_valid_test
self, test: str , code: str
bool
"""Check if test is syntactically valid and runs."""
try
return
'assert'
in
or
'self.assert'
in
except
return
False
def
_run_test
self, code: str , test: str , timeout=5
dict
"""Execute test against code in sandbox."""
f"{code} \n\n{test} "
try
'python'
'-c'
True
True
return
'success'
0
'output'
'error'
except
return
'success'
False
'output'
''
'error'
'Timeout'
Next, implement the Mistake Book—a memory buffer storing high-quality failing test cases to maintain training stability.
from collections import deque
import heapq
class MistakeBook :
"""Store high-quality test cases that expose bugs."""
def __init__ (self, max_size=1000 ):
self .tests = deque(maxlen=max_size)
self .quality_scores = []
def add_test (self, test: str , quality_score: float ):
"""Add test if it has sufficient quality."""
if quality_score > 0.3 :
self .tests.append(test)
self .quality_scores.append(quality_score)
def sample_batch (self, batch_size: int ) -> list :
"""Sample tests, biasing toward high-quality ones."""
if not self .tests:
return []
total_quality = sum (self .quality_scores)
if total_quality == 0 :
sampled_indices = np.random.choice(len (self .tests), batch_size,
replace=True )
else :
probabilities = [q / total_quality for q in self .quality_scores]
sampled_indices = np.random.choice(len (self .tests), batch_size,
p=probabilities, replace=True )
return [self .tests[i] for i in sampled_indices]
def get_top_k (self, k: int ) -> list :
"""Return top-k quality tests for evaluation."""
indexed = list (enumerate (zip (self .tests, self .quality_scores)))
top_k = heapq.nlargest(k, indexed, key=lambda x: x[1 ][1 ])
return [test for _, (test, _) in top_k]
Now implement the main adversarial training loop coordinating code and test generation.
import torch
from torch.optim import AdamW
class AdversarialCodeTrainer :
"""Co-train Code LLM and Test LLM with adversarial objectives."""
def __init__ (self, code_llm, test_llm, reward_fn, max_code_len=1024 ):
self .code_llm = code_llm
self .test_llm = test_llm
self .reward_fn = reward_fn
self .mistake_book = MistakeBook(max_size=2000 )
self .max_code_len = max_code_len
def step (self, problem_specs: list , num_samples=4 ):
"""One training step with adversarial loop."""
results = {
'code_rewards' : [],
'test_rewards' : [],
'code_loss' : 0 ,
'test_loss' : 0
}
for spec in problem_specs:
code_samples = self .code_llm.generate_batch(
spec['prompt' ],
num_samples=num_samples,
max_length=self .max_code_len
)
valid_code = [c for c in code_samples
if self .reward_fn._is_valid_code(c)]
if not valid_code:
continue
test_context = f"Code under test:\n{valid_code[0 ]} \n\nGenerate adversarial tests:"
test_samples = self .test_llm.generate_batch(
test_context,
num_samples=num_samples,
max_length=512
)
valid_tests = [t for t in test_samples
if self .reward_fn._is_valid_test(t, valid_code[0 ])]
code_rewards = []
for code in valid_code:
code_reward = self .reward_fn.score_code(code, valid_tests)
code_rewards.append(code_reward)
results['code_rewards' ].append(code_reward)
test_rewards = []
for test in valid_tests:
best_code = valid_code[np.argmax(code_rewards)]
test_reward = self .reward_fn.score_test(test, best_code)
test_rewards.append(test_reward)
results['test_rewards' ].append(test_reward)
if test_reward > 0.5 :
self .mistake_book.add_test(test, test_reward)
code_rewards_tensor = torch.tensor(code_rewards, dtype=torch.float32)
code_loss = self ._compute_policy_loss(code_samples, code_rewards_tensor)
self ._update_model(self .code_llm, code_loss)
results['code_loss' ] += code_loss.item()
test_rewards_tensor = torch.tensor(test_rewards, dtype=torch.float32)
test_loss = self ._compute_policy_loss(test_samples, test_rewards_tensor)
self ._update_model(self .test_llm, test_loss)
results['test_loss' ] += test_loss.item()
return results
def _compute_policy_loss (self, samples: list , rewards: torch.Tensor ):
"""Policy gradient loss: maximize reward * log_prob."""
rewards = (rewards - rewards.mean()) / (rewards.std() + 1e-8 )
logprobs = self .code_llm.compute_logprob_batch(samples)
loss = -(rewards * logprobs).mean()
return loss
def _update_model (self, model, loss ):
"""Gradient update step."""
optimizer = AdamW(model.parameters(), lr=1e-5 )
optimizer.zero_grad()
loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0 )
optimizer.step()
def train (self, problem_specs, num_steps=100 ):
"""Run full adversarial training."""
for step in range (num_steps):
results = self .step(problem_specs)
if (step + 1 ) % 10 == 0 :
avg_code_reward = np.mean(results['code_rewards' ])
avg_test_reward = np.mean(results['test_rewards' ])
print (f"Step {step+1 } : Code Reward={avg_code_reward:.3 f} , "
f"Test Reward={avg_test_reward:.3 f} " )
Practical Guidance Hyperparameters and When to Use:
Validity weight typically 0.6, difficulty weight 0.4; adjust based on how strict test quality requirements are
Mistake book size of 1000-2000 balances diversity with computational cost
Use separate model instances for code and test LLMs; same architecture works but keep parameters independent
Effective for programming problems with clear test cases (competitive programming, algorithmic problems, code completion)
For problems without clear oracles (creative coding, open-ended design tasks)
When high-quality human-written tests are already available; use supervised learning instead
For code generation in domains requiring safety verification (e.g., cryptography, safety-critical systems)
Test generation becoming too difficult; use curriculum learning starting with simpler problems
Code and test LLMs converging to trivial solutions; regularly refresh the mistake book with historical high-quality tests
Insufficient diversity in generated tests; apply dropout/temperature sampling during generation
Execution timeout issues from infinite loops; set strict time limits (typically 1-5 seconds per execution)
Test LLM seeing answers in code context leading to "teaching to the test"; use code anonymization if possible
Reference