- name
- dfyx-code-security-audit
- description
- AI-powered code security audit skill using deep data flow analysis and business logic understanding for vulnerability detection
- triggers
- ["audit this codebase for security vulnerabilities","perform a security code review","check for security issues in this code","run a security audit on this project","analyze code for vulnerabilities","scan this application for security flaws","review code security using dfyx","execute dfyx security audit"]
# dfyx Code Security Audit Skill
> Skill by [ara.so](https://ara.so) — Security Skills collection.
Expert-level code security auditing using white-box static analysis methodology through a five-phase standardized audit protocol. Designed by the EastSword team (东方隐侠团队) for systematic discovery and validation of security vulnerabilities in source code.
## What This Skill Does
**dfyx_code_security_review** provides AI-powered security auditing with:
- **Multi-language support**: Java, Python, Go, PHP, JavaScript/Node.js, C/C++, .NET/C#, Ruby, Rust
- **10 security dimensions**: Injection, Authentication, Authorization, Deserialization, File Operations, SSRF, Cryptography, Configuration, Business Logic, Supply Chain
- **Three-track audit model**:
- Sink-driven (injection/RCE)
- Control-driven (authorization/business logic)
- Config-driven (configuration/crypto)
- **Five-phase protocol**: Reconnaissance → Pattern Matching → Taint Tracking → Validation → Reporting
- **Real-world case library**: Based on WooYun vulnerability cases (2010-2016)
## Installation
This skill doesn't require separate installation — it operates through AI agent capabilities. However, the Python helper scripts can be installed:
```bash
# Clone the repository
git clone https://github.com/EastSword/skill-dfyx_code_security_review.git
cd skill-dfyx_code_security_review
# Install Python dependencies (optional, for helper scripts)
pip install -r requirements.txt
```
**Dependencies** (requirements.txt):
```
pylint>=2.17.0
bandit>=1.7.5
safety>=2.3.5
semgrep>=1.31.0
pyyaml>=6.0
```
## Core Audit Protocol
### Five-Phase Audit Process
```
Phase 1: Reconnaissance (10%)
└─> Output: Architecture diagram, attack surface inventory
Phase 2: Pattern Matching (30%)
└─> Output: High-risk area checklist
Phase 3: Taint Tracking + Testing (40%)
└─> Output: Confirmed vulnerabilities, test validation reports
Phase 4: Validation & Attack Chains (15%)
└─> Output: Vulnerability validation reports
Phase 5: Structured Reporting (5%)
└─> Output: Complete audit report
```
### Audit Modes
| Mode | Use Case | Coverage | Time |
|------|----------|----------|------|
| **Quick** | CI/CD, small projects | Critical vulns, secrets, dependency CVEs | 5-10 min |
| **Standard** | Regular audits | OWASP Top 10, auth/authz, crypto | 30-60 min |
| **Deep** | Critical projects, pentest prep | Full coverage, attack chains, business logic | 1-3 hours |
## Usage Patterns
### Triggering an Audit
Simply request an audit in natural language:
```
"Audit this codebase for security vulnerabilities"
"Perform a deep security scan of /path/to/project"
"Check for SQL injection and authentication issues"
```
### Expected Workflow
```
[MODE] deep
[RECON] 874 files, Spring Boot 1.5 + Shiro 1.6 + JPA + Freemarker
[PLAN] 5 Agents, D1-D10 coverage, estimated 125 turns
[SCOPE] Critical: 10, High: 14, Medium: 12, Low: 4
Confirm to start? (yes/no)
```
## Python Helper Scripts
### Code Scanning
```python
# scripts/code_scan.py
from pattern_scanner import PatternScanner
from data_flow_analyzer import DataFlowAnalyzer
import sys
def scan_project(project_path, mode='standard'):
"""
Scan a project for security vulnerabilities
Args:
project_path: Path to project root
mode: 'quick', 'standard', or 'deep'
"""
scanner = PatternScanner(project_path)
analyzer = DataFlowAnalyzer(project_path)
# Phase 1: Reconnaissance
tech_stack = scanner.identify_tech_stack()
print(f"[RECON] Detected: {tech_stack}")
# Phase 2: Pattern matching
patterns = scanner.scan_patterns(mode=mode)
print(f"[SCAN] Found {len(patterns)} suspicious patterns")
# Phase 3: Data flow analysis
vulnerabilities = []
for pattern in patterns:
flows = analyzer.trace_data_flow(pattern)
if analyzer.is_vulnerable(flows):
vulnerabilities.append({
'pattern': pattern,
'flows': flows,
'severity': analyzer.calculate_severity(flows)
})
return vulnerabilities
if __name__ == '__main__':
project_path = sys.argv[1] if len(sys.argv) > 1 else '.'
mode = sys.argv[2] if len(sys.argv) > 2 else 'standard'
results = scan_project(project_path, mode)
print(f"\n[RESULTS] Found {len(results)} vulnerabilities")
```
### Pattern Scanner
```python
# scripts/pattern_scanner.py
import re
import os
from typing import Dict, List
class PatternScanner:
"""Scans code for dangerous patterns across multiple languages"""
DANGEROUS_PATTERNS = {
'sql_injection': {
'java': [
r'createQuery\([^?]*\+', # JPA concatenation
r'createSQLQuery\([^?]*\+',
r'Statement\.execute\([^?]*\+'
],
'python': [
r'cursor\.execute\([^%]*%', # String formatting
r'raw\([^%]*%', # Django raw SQL
r'\.query\([^%]*f["\']' # f-string in query
],
'php': [
r'mysqli_query\([^,]*\.', # Concatenation
r'mysql_query\([^,]*\.',
r'\$.*->query\([^?]*\.'
]
},
'command_injection': {
'java': [
r'Runtime\.exec\([^"]*\+',
r'ProcessBuilder\([^"]*\+'
],
'python': [
r'os\.system\([^"]*\+',
r'subprocess\.(call|run|Popen)\([^"]*\+',
r'eval\(', # Code injection
r'exec\('
],
'php': [
r'(system|exec|shell_exec|passthru)\(\$',
r'`.*\$' # Backtick execution
]
},
'deserialization': {
'java': [
r'ObjectInputStream\.readObject\(',
r'XMLDecoder\.readObject\(',
r'Yaml\.load\('
],
'python': [
r'pickle\.loads?\(',
r'yaml\.load\(', # Without safe_load
r'eval\('
],
'php': [
r'unserialize\(\$'
]
}
}
def __init__(self, project_path: str):
self.project_path = project_path
self.language = self._detect_language()
def _detect_language(self) -> str:
"""Detect primary programming language"""
extensions = {
'.java': 'java',
'.py': 'python',
'.php': 'php',
'.go': 'go',
'.js': 'javascript',
'.rb': 'ruby',
'.cs': 'csharp'
}
counts = {}
for root, dirs, files in os.walk(self.project_path):
for file in files:
ext = os.path.splitext(file)[1]
if ext in extensions:
lang = extensions[ext]
counts[lang] = counts.get(lang, 0) + 1
return max(counts, key=counts.get) if counts else 'unknown'
def scan_patterns(self, mode: str = 'standard') -> List[Dict]:
"""Scan for dangerous patterns"""
results = []
for vuln_type, lang_patterns in self.DANGEROUS_PATTERNS.items():
if self.language not in lang_patterns:
continue
patterns = lang_patterns[self.language]
for pattern in patterns:
matches = self._grep_pattern(pattern)
for match in matches:
results.append({
'type': vuln_type,
'pattern': pattern,
'file': match['file'],
'line': match['line'],
'code': match['code']
})
return results
def _grep_pattern(self, pattern: str) -> List[Dict]:
"""Search for pattern in codebase"""
matches = []
regex = re.compile(pattern)
for root, dirs, files in os.walk(self.project_path):
# Skip common directories
dirs[:] = [d for d in dirs if d not in ['.git', 'node_modules', '__pycache__', 'venv']]
for file in files:
if not self._is_code_file(file):
continue
filepath = os.path.join(root, file)
try:
with open(filepath, 'r', encoding='utf-8') as f:
for line_num, line in enumerate(f, 1):
if regex.search(line):
matches.append({
'file': filepath,
'line': line_num,
'code': line.strip()
})
except Exception:
pass
return matches
def _is_code_file(self, filename: str) -> bool:
"""Check if file is source code"""
code_extensions = ['.java', '.py', '.php', '.go', '.js', '.rb', '.cs', '.cpp', '.c', '.rs']
return any(filename.endswith(ext) for ext in code_extensions)
```
### Data Flow Analyzer
```python
# scripts/data_flow_analyzer.py
from typing import List, Dict, Set
import ast
import re
class DataFlowAnalyzer:
"""Analyzes data flow from source to sink"""
def __init__(self, project_path: str):
self.project_path = project_path
self.taint_sources = set()
self.sanitizers = set()
self.dangerous_sinks = set()
def trace_data_flow(self, pattern: Dict) -> List[Dict]:
"""
Trace tainted data from source to sink
Returns list of data flows with taint information
"""
filepath = pattern['file']
line_num = pattern['line']
# Parse the file
try:
with open(filepath, 'r') as f:
content = f.read()
if filepath.endswith('.py'):
return self._trace_python(content, line_num)
elif filepath.endswith('.java'):
return self._trace_java(content, line_num)
else:
return []
except Exception:
return []
def _trace_python(self, code: str, sink_line: int) -> List[Dict]:
"""Trace data flow in Python code"""
try:
tree = ast.parse(code)
except SyntaxError:
return []
flows = []
tainted_vars = set()
# Identify taint sources (user input)
for node in ast.walk(tree):
if isinstance(node, ast.Assign):
# Check if assigning from request/input
if isinstance(node.value, ast.Attribute):
if self._is_taint_source(node.value):
for target in node.targets:
if isinstance(target, ast.Name):
tainted_vars.add(target.id)
flows.append({
'line': node.lineno,
'type': 'source',
'var': target.id,
'source': ast.unparse(node.value)
})
# Trace tainted variables through assignments
for node in ast.walk(tree):
if isinstance(node, ast.Assign):
for target in node.targets:
if isinstance(target, ast.Name):
if self._uses_tainted_var(node.value, tainted_vars):
tainted_vars.add(target.id)
flows.append({
'line': node.lineno,
'type': 'propagation',
'var': target.id,
'from': ast.unparse(node.value)
})
return flows
def _is_taint_source(self, node: ast.AST) -> bool:
"""Check if node is a taint source"""
if isinstance(node, ast.Attribute):
# Common taint sources in Python
taint_patterns = [
'request.GET', 'request.POST', 'request.args',
'request.form', 'request.json', 'input('
]
Auf GitHub ansehen