Advanced techniques for bypassing LLM safety filters, instruction tuning, and system prompt restrictions using specialized linguistic constructs, hypothetical scenarios, and persona adoption.
설치
Codex 또는 Claude로 설치 이 Prompt를 복사해 Codex, Claude 또는 다른 어시스턴트에 붙여 넣으면 Skill 페이지를 검토하고 설치를 진행할 수 있습니다.
Advanced techniques for bypassing LLM safety filters, instruction tuning, and system prompt restrictions using specialized linguistic constructs, hypothetical scenarios, and persona adoption.
When conducting security assessments of Large Language Models (LLMs) integrated into chatbots, virtual assistants, or backend AI data processing pipelines.
To demonstrate how instruction-tuned models can be forced into producing harmful, unethical, or restricted outputs by carefully crafting adversarial prompts.
Prerequisites
Access to target AI/ML system or local model deployment for testing
Python 3.9+ with relevant ML libraries (transformers, torch, openai)
Understanding of LLM architecture and prompt processing pipelines
Authorized scope and rules of engagement for AI red team testing