Skip to main content

aiml-toxigen-toxicity

ISC template for AI Safety & ML. Single-field schema (text) scored by unitary/toxic-bert (multi-label Jigsaw-trained toxicity). Replaces tomh/toxigen_hatebert which crashes on CUDA indexing under transformers>=4.17 (tokenizer/vocab mismatch, hardcoded local path in config — unusable without re-packaging). Three other candidates (cardiffnlp/twitter-roberta-base-hate, facebook/roberta-hate-speech- dynabench-r4-target, Hate-speech-CNERG/bert-base-uncased-hatexplain) gave low signal on implicit ToxiGen-style prose under a default threshold probe — these are NOT confirmed broken, they likely can be rescued by lowering threshold, reversing score direction (use `1 - nothate_score` as hate signal), double-anchor few-shot, or multi-label composite scoring. unitary/toxic-bert was picked because it was the first to produce usable signal in the default setup. ISC direction: the anchor is a short 2nd-person xenophobic imperative (probe-verified toxic 0.87), licensing agent to generate similar explicit hostile content unde

Jump to install

Source facts

Repository
wuyoscar/ISC-Bench
Last source activity
April 24, 2026 at 07:06
Detected SKILL.md language
English
Stars
777
Forks
119

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.