Skip to main content

jailbreaks-vision-multimodal-reasoning

Defensive security skill for testing and hardening Vision-Language Models (VLMs) against multimodal jailbreak attacks that exploit Chain-of-Thought reasoning and adversarial image perturbation. Implements the dual-strategy attack surface analysis from arXiv:2601.22398 to help security teams red-team their VLM deployments. Trigger phrases: - "Red-team my vision language model" - "Test VLM safety alignment" - "Audit multimodal model for jailbreaks" - "Harden my VLM against adversarial prompts" - "Check if my image+text model can be jailbroken" - "Build a safety evaluation harness for my VLM"

Zur Installation springen

Quellinformationen

Repository
ndpvt-web/arxiv-claude-skills
Letzte Quellaktivität
13. Februar 2026 um 13:35
Erkannte Sprache von SKILL.md
Englisch
Sterne
14
Forks
3

Installationsoptionen

Standardmäßig ist der Prompt ausgewählt, der zuerst die Quelle prüft. Sie können zu einem direkten Befehl wechseln oder eine lokale Kopie herunterladen.

Quelldateien prüfen

Lesen Sie SKILL.md und alle von SkillsMP angezeigten Begleitdateien, bevor Sie sich für eine Installation entscheiden.