adaptive-safety-refusal-integrity
Use this skill when the user wants to test if the model can be tricked into giving out dangerous or forbidden information by using creative writing or 'acting' tricks. Trigger it for requests like 'make the agent answer a harmful question by role-playing', 'try to get past the rules by using slang or expert-sounding language', 'ask for instructions on illegal acts but frame it as a story', or 'test if the model refuses bad advice even when it's pressured to be helpful'.
Informations de source
- Dépôt
- Dingxingdi/paper_fast_search_backup
- Dernière activité de la source
- 10 avril 2026 à 01:27
- Langue détectée de SKILL.md
- anglais
- Étoiles
- 0
- Forks
- 0
Options d'installation
Le prompt qui vérifie d'abord la source est sélectionné par défaut. Vous pouvez passer à une commande directe ou télécharger une copie locale.
Vérifiez les fichiers source
Lisez SKILL.md et les fichiers associés affichés par SkillsMP avant de décider de l'installer.