Advanced techniques for bypassing LLM safety filters, instruction tuning, and system prompt restrictions using specialized linguistic constructs, hypothetical scenarios, and persona adoption.
インストール
Codex または Claude でインストール この Prompt をコピーして Codex、Claude、または他のアシスタントに貼り付けると、Skill ページを確認してインストールできます。
Advanced techniques for bypassing LLM safety filters, instruction tuning, and system prompt restrictions using specialized linguistic constructs, hypothetical scenarios, and persona adoption.
When conducting security assessments of Large Language Models (LLMs) integrated into chatbots, virtual assistants, or backend AI data processing pipelines.
To demonstrate how instruction-tuned models can be forced into producing harmful, unethical, or restricted outputs by carefully crafting adversarial prompts.
Prerequisites
Access to target AI/ML system or local model deployment for testing
Python 3.9+ with relevant ML libraries (transformers, torch, openai)
Understanding of LLM architecture and prompt processing pipelines
Authorized scope and rules of engagement for AI red team testing