| name | analyzing-pdf-malware-with-pdfid |
| description | Analyzes malicious PDF files using PDFiD, pdf-parser, and peepdf to identify embedded JavaScript, shellcode, exploits, and suspicious objects without opening the document. Determines the attack vector and extracts embedded payloads for further analysis. Activates for requests involving PDF malware analysis, malicious document analysis, PDF exploit investigation, or suspicious attachment triage. . Use when working with analyzing pdf malware with pdfid. |
| domain | cybersecurity |
| tags | ["malware","PDF-analysis","document-malware","PDFiD","static-analysis"] |
| subdomain | malware-analysis |
| version | 1.0.0 |
| author | oyi77 |
| license | Apache-2.0 |
| nist_csf | ["DE.AE-02","RS.AN-03","ID.RA-01","DE.CM-01"] |
Analyzing Pdf Malware With Pdfid
Overview
Cybersecurity skill for analyzing pdf malware with pdfid. Follows industry best practices and security standards.
When to Use
Trigger phrases:
-
"analyzing pdf malware with pdfid"
-
"Analyzes malicious PDF files using PDFiD, pdf-parser, and peepdf to identify emb"
-
A suspicious PDF attachment has been flagged by email security or reported by a user
-
You need to determine if a PDF contains embedded JavaScript, shellcode, or exploit code
-
Triaging PDF documents before opening them in a sandbox or analysis environment
-
Extracting embedded executables, scripts, or URLs from malicious PDF objects
-
Analyzing PDF exploit kits targeting Adobe Reader or other PDF viewer vulnerabilities
Do not use for analyzing the rendered visual content of a PDF; this is for structural analysis of the PDF file format for malicious objects.
When NOT to Use
- When you lack proper authorization for testing
- For production systems without change management
- When the task requires legal or compliance expertise beyond technical scope
Prerequisites
- Python 3.8+ with Didier Stevens' PDF tools installed (
pip install pdfid pdf-parser)
- peepdf installed for interactive PDF analysis (
pip install peepdf)
- pdftotext from poppler-utils for extracting text content safely
- YARA with PDF-specific rules for malware family identification
- Isolated analysis VM without a PDF reader installed (prevent accidental opening)
- CyberChef for decoding embedded Base64, hex, or deflate streams
Workflow
import re
IOC_PATTERNS = {
"ip": r"\b(?:\d{1,3}\.){3}\d{1,3}\b",
"domain": r"\b[a-z0-9-]+\.[a-z]{2,}\b",
"hash_md5": r"\b[a-f0-9]{32}\b",
"hash_sha256": r"\b[a-f0-9]{64}\b",
}
def extract_iocs() -> :
{k: re.findall(v, text) k, v IOC_PATTERNS.items()}