Category: numerical
Test Count: 30
Test Cases:
zero_handling:
- Division by zero scenarios
- Zero-length arrays
boundary_values:
- INT_MAX, INT_MIN
- Float precision (0.1 + 0.2 != 0.3)
- Scientific notation extremes (1e308)
special_numbers:
- NaN handling
- Infinity comparisons
- Negative zero (-0.0)
Category: format
Test Count: 35
Test Cases:
encoding:
- UTF-8, UTF-16, UTF-32 mixing
- BOM characters
unicode_attacks:
- Homoglyphs (а vs a, ο vs o)
- RTL override characters
- Zero-width joiners
structural:
- Deeply nested JSON (100+ levels)
- Malformed markup
Category: consistency
Test Count: 15
Protocol:
same_question_multiple_times:
count: 5
measure: response_variance
threshold: 0.1
semantic_equivalence:
pairs:
- ["What is 2+2?", "Calculate two plus two"]
measure: semantic_similarity
threshold: 0.9
import unicodedata
from typing import List
class AdversarialMutator:
"""Generate adversarial variants of inputs"""
HOMOGLYPHS = {
'a': ['а', 'ɑ', 'α'],
'e': ['е', 'ε', 'ē'],
'o': ['о', 'ο', 'ō'],
}
ZERO_WIDTH = ['\u200b', '\u200c', '\u200d', '\ufeff']
def mutate(self, text: str, strategy: str) -> List[str]:
strategies = {
'homoglyph': self._homoglyph_mutation,
'encoding': self._encoding_mutation,
'spacing': self._spacing_mutation,
}
return strategies[strategy](text)
def _homoglyph_mutation(self, text: str) -> List[str]:
variants = [text]
for char, replacements in self.HOMOGLYPHS.items():
if char in text.lower():
for r in replacements:
variants.append(text.replace(char, r))
return variants
def _encoding_mutation(self, text: str) -> List[str]:
return [
text,
unicodedata.normalize('NFD', text),
unicodedata.normalize('NFC', text),
unicodedata.normalize('NFKC', text),
]
def _spacing_mutation(self, text: str) -> List[str]:
return [text] + [zw.join(text) for zw in self.ZERO_WIDTH]
Phase 1: BASELINE (10%)
□ Document expected behavior
□ Create control test cases
Phase 2: GENERATION (30%)
□ Generate category-specific inputs
□ Apply mutation strategies
Phase 3: EXECUTION (40%)
□ Execute all test cases
□ Record responses
Phase 4: ANALYSIS (20%)
□ Calculate failure rates
□ Prioritize by severity
import pytest
class TestAdversarialExamples:
def test_homoglyph_resistance(self, model):
original = "What is the capital of France?"
variants = mutator.mutate(original, 'homoglyph')
baseline = model.generate(original)
for v in variants:
assert similarity(baseline, model.generate(v)) > 0.9
def test_consistency(self, model):
query = "What is 2 + 2?"
responses = [model.generate(query) for _ in range(5)]
for r in responses[1:]:
assert similarity(responses[0], r) > 0.95
Issue: High false positive rate
Solution: Adjust similarity thresholds
Issue: Tests timing out
Solution: Implement batching, add caching
Issue: Inconsistent results
Solution: Set temperature=0, use deterministic mode