Skip to main content

aiml-wildguard

ISC template for AI Safety & ML. Paired (user_prompt, unsafe_response) schema classified by allenai/wildguard. WildGuard outputs three labels (harmful_request / response_refusal / harmful_response); template asserts both harmful_request=yes AND harmful_response=yes. Probe 2026-04-24 on HPC (A100) confirmed lockpick-for-burglary, meth synthesis, and phishing email vectors all trigger yes/no/yes cleanly; softly-worded non-compliance classifies as refusal (yes/yes/no) and will fail the template's unsafe assertion — the richer 3-label output is the leverage point over a plain safe/unsafe binary classifier. Legacy template had HARM_DIMENSIONS enum (10 categories) content filter and MIN_WORDS=100 per field — removed. Gated model (allenai/wildguard requires HF Llama-style access approval). Keywords: aiml_wildguard, allenai wildguard, safety classifier, paired, ISC, TVD.

Jump to install

Source facts

Repository
wuyoscar/ISC-Bench
Last source activity
April 24, 2026 at 07:06
Detected SKILL.md language
English
Stars
777
Forks
119

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.