Instalar com Codex ou Claude Copie este prompt, cole no Codex, Claude ou outro assistente e deixe que ele revise a página da skill e instale para você.
Um comando direto ignora o prompt de revisão. Verifique a origem antes de executá-lo.
Instruções da origem · Visualização somente leitura
name
openai-privacy-filter
description
OpenAI Privacy Filter — bidirectional token-classification model for PII detection and masking in text
triggers
["detect PII in text","redact personally identifiable information","mask private data with OpenAI privacy filter","run privacy filter on a file","finetune PII detection model","evaluate privacy filter on labeled data","filter sensitive information from text","anonymize text with privacy filter"]
OpenAI Privacy Filter is a bidirectional token-classification model (1.5B params, 50M active) for detecting and masking PII spans in text. It runs in a single forward pass with constrained Viterbi decoding, supports a 128k-token context window, and is licensed Apache 2.0.
Installation
pip install -e .
# or from a cloned repo:
git clone https://github.com/openai/privacy-filter
cd privacy-filter
pip install -e .
After install, the opf CLI is available. On first use it downloads the model checkpoint to ~/.opf/privacy_filter unless OPF_CHECKPOINT is set.
# Redact inline text
opf "Alice was born on 1990-01-02 and her email is alice@example.com."# Force CPU inference
opf --device cpu "Alice was born on 1990-01-02."# Use a specific checkpoint
opf --checkpoint /path/to/checkpoint_dir "Alice Johnson, SSN 123-45-6789"# Redact an entire file
opf -f /path/to/document.txt
# Pipe inputcat document.txt | grep "sensitive" | opf
# Interactive mode (no input provided)
opf
Evaluation
# Evaluate on a labeled JSONL dataset
opf eval examples/data/sample_eval_five_examples.jsonl
# See all eval options
opf eval --help
Finetuning
# Finetune on your labeled dataset
opf train /path/to/train.jsonl --output-dir /path/to/finetuned_checkpoint
# See all training options
opf train --help
Python API
from opf import PrivacyFilter
# Load with default checkpoint (~/.opf/privacy_filter or OPF_CHECKPOINT)
pf = PrivacyFilter()
# Or specify a checkpoint explicitly
pf = PrivacyFilter(checkpoint="/path/to/checkpoint_dir")
# Redact a single string
result = pf.redact("Alice Johnson called from +1-800-555-0199.")
print(result.redacted_text)
# "██████████████ called from ██████████████."# Access detected spansfor span in result.spans:
print(span.label, span.text, span.start, span.end)
Batch processing
from opf import PrivacyFilter
pf = PrivacyFilter(device="cuda") # or "cpu"
texts = [
"Contact Bob Smith at bob@example.com",
"Her SSN is 123-45-6789 and DOB is 1985-03-15",
"API key: sk-abc123xyz789",
]
results = pf.redact_batch(texts)
for r in results:
print(r.redacted_text)
print(r.spans)
Precision/Recall tuning via operating points
from opf import PrivacyFilter
# High recall (broader masking, more false positives)
pf_recall = PrivacyFilter(operating_point="high_recall")
# High precision (stricter masking, fewer false positives)
pf_precision = PrivacyFilter(operating_point="high_precision")
# Default balanced
pf_default = PrivacyFilter()
Data Format
Input for eval and training (JSONL)
Each line is a JSON object:
{"text": "Alice was born on 1990-01-02.", "spans": [{"start": 0, "end": 5, "label": "private_person"}, {"start": 18, "end": 28, "label": "private_date"}]}
{"text": "Email bob@corp.com for details.", "spans": [{"start": 6, "end": 18, "label": "private_email"}]}
JSON output schema
{"redacted_text":"██████ was born on ██████████.","spans":[{"label":"private_person","text":"Alice","start":0,"end":5,"score":0.987},{"label":"private_date","text":"1990-01-02","start":18,"end":28,"score":0.973}]}
See OUTPUT_SCHEMAS.md in the repo for full payload spec.
Finetuning Workflow
# Prepare labeled JSONL (see data format above)# Run finetuning
opf train train.jsonl \
--output-dir ./my_finetuned_model \
--eval-file eval.jsonl \
--epochs 3 \
--batch-size 8
# Use the finetuned model
opf --checkpoint ./my_finetuned_model "redact this text"
See FINETUNING.md and examples/scripts/finetuning/ for runnable demo harnesses.
Environment Variables
Variable
Purpose
OPF_CHECKPOINT
Path to model checkpoint directory (overrides default ~/.opf/privacy_filter)
Pipeline: sanitize files before uploading to an LLM
from opf import PrivacyFilter
import json
pf = PrivacyFilter()
defsanitize_for_llm(raw_text: str) -> str:
result = pf.redact(raw_text)
return result.redacted_text
withopen("raw_data.txt") as f:
clean = sanitize_for_llm(f.read())
print(clean)
Audit: log all detected PII spans without redacting
from opf import PrivacyFilter
pf = PrivacyFilter()
defaudit_pii(text: str) -> list[dict]:
result = pf.redact(text)
return [
{"label": s.label, "text": s.text, "start": s.start, "end": s.end}
for s in result.spans
]
findings = audit_pii("Bob Jones (DOB: 1978-06-15) owes $1,200.")
print(json.dumps(findings, indent=2))
Filter specific label types only
from opf import PrivacyFilter
pf = PrivacyFilter()
defredact_only(text: str, labels: list[str]) -> str:
result = pf.redact(text)
# Rebuild text redacting only chosen labels
chars = list(text)
for span in result.spans:
if span.label in labels:
for i inrange(span.start, span.end):
chars[i] = "█"return"".join(chars)
# Only redact emails and phones, keep names
output = redact_only(
"Call Alice at 555-1234 or alice@example.com",
labels=["private_phone", "private_email"]
)
print(output)
# "Call Alice at ████████ or █████████████████"
Troubleshooting
Model not found / auto-download fails
Set OPF_CHECKPOINT to a local checkpoint directory, or ensure internet access for the first run.