- name
- remove-ai-marks
- description
- Remove multi-vendor AI provenance marks: invisible Unicode (Layer A), statistical text watermarks via rewrite (Layer B, always offer), and C2PA/EXIF/XMP/container metadata on PNG/JPEG/WebP/SVG/PDF/DOCX/ODT/HTML/MD/TEX. Covers Claude, Gemini/SynthID-class, OpenAI provenance, and open-LLM sampling marks. Use when the user asks to strip watermarks, remove C2PA/Content Credentials, clean AI metadata, remove invisible Unicode, anti-detect clean AI output, or runs /remove-ai-marks (aliases: /remove-claude-marks).
# Remove AI marks
Multi-vendor anti-detection hygiene for **text** (Unicode + statistical rewrite) and **files** (C2PA / AI metadata across common containers).
Read if needed:
- `references/mark-classes.md` — Unicode / sampling / C2PA / containers
- `references/vendor-notes.md` — Claude, Gemini/SynthID, OpenAI, open-LLM
- `references/removal-matrix.md` — which layer when
- `references/ethics.md` — intended use
- `references/how-claude-marks.md` — Anthropic-specific detail
- `references/markdiffusion.md` — optional MarkDiffusion image harness (schemes, honesty caveats)
This skill is a **thin client**. All deterministic cleaning machinery runs in a
separate HTTP service (this repo's `service/`), so the agent host needs no
Python, venvs, or cleaning tools. Call the service with `curl`; never run
cleaning scripts directly.
## Windows on-demand startup (Marketplace policy)
Probe `/health` before use. If the default local service at
`http://127.0.0.1:8765` is unavailable, run the plugin-root `Start-Service.ps1`
with PowerShell 7 (`../../Start-Service.ps1` relative to this Skill directory).
It starts Python from `%USERPROFILE%/watermarks-remover-service`, waits for
health, and leaves the service running after use. Report startup errors and stop
until health succeeds. For a custom remote URL, report connection failure instead
of starting a local service. Keep login startup disabled.
This policy takes precedence over upstream local-service startup instructions.
## Service access
Base URL comes from `WATERMARKS_SERVICE_URL`, default `http://127.0.0.1:8765`:
```bash
WM="${WATERMARKS_SERVICE_URL:-http://127.0.0.1:8765}"
```
The service is started either by the operator (`docker compose up -d`, or a
published GHCR image) or locally (`make serve`). **Always check it first**, and
stop with a clear message if it is unreachable — never fall back to local
cleaning:
```bash
AUTH_HEADER=()
if [ -n "$WATERMARKS_SERVER_API_KEY" ]; then
AUTH_HEADER=(-H "Authorization: Bearer $WATERMARKS_SERVER_API_KEY")
fi
curl -sf "${AUTH_HEADER[@]}" "$WM/health"
# {"ok": true, "version": "..."}
```
If `WATERMARKS_SERVER_API_KEY` is set on the service, every request (including
the health check and capabilities) needs `-H "Authorization: Bearer $WATERMARKS_SERVER_API_KEY"`. The default URL is
loopback; when the service runs on another host, set `WATERMARKS_SERVICE_URL`
to an `https://` URL so the token is not sent in cleartext, and do not add
`-L` (a redirect could forward the token to another host).
### Capabilities
```bash
curl -s "${AUTH_HEADER[@]}" "$WM/capabilities"
```
Reports which optional tools are available server-side (`c2patool`, `exiftool`,
`qpdf`, `ghostscript`), scorers present (`scorers.stylometry`, `scorers.synthid`,
`scorers.synthid_http`), text-watermark detectors
(`text_detectors.markllm`,
`text_detectors.claude-text`), and which heavy backends are configured
(`pixel_backends.ctrlregen`, `pixel_backends.diffusion`, `harnesses.markllm`).
**Drive your advice from this**: only recommend pixel removal / SynthID
scoring / vendor detection when the service reports the backend present.
## HTTP API (curl)
Payloads are JSON with the file as **base64**. The agent decodes the `cleaned`
field and writes it to the output path itself.
| Method | Path | Body | Returns |
| --- | --- | --- | --- |
| GET | `/health` | — | `{"ok": true, "version": ...}` |
| GET | `/capabilities` | — | optional tools / backends present |
| GET | `/openapi.json` | — | dynamically generated OpenAPI 3.0.3 spec |
| POST | `/inspect` | `{"file": "<base64>", "name": "notes.md"}` | `{"ok", "kind", "suspicious", "report"}` |
| POST | `/detect` | `{"file": "<base64>", "name": "notes.txt"}` | `{"ok", "kind", "detections": [...]}` |
| POST | `/clean` | `{"file": "<base64>", "name": "notes.md", "options": {...}}` | `{"ok", "kind", "cleaned": "<base64>", "report"}` |
`/clean` and `/inspect` route by the uploaded `name` extension plus the bytes;
unrecognized formats answer `kind: "unknown"` (`/inspect`) or 400 (`/clean`).
When writing a temp file for pasted text, keep a known extension (`.txt` /
`.md`) in the `name` you send.
The machine-readable contract lives at `$WM/openapi.json` — plug it into any
OpenAPI tooling (client generators, Swagger UI, editors) instead of hand-rolling
clients.
`options` accepted by `/clean`: `nfkc`, `aggressive_homoglyphs` (text),
`keep_non_ai_metadata`, `strip_all_metadata`, `remove_pixel` (`ctrlregen` |
`diffusion`) (images and video), `also_layer_a_text` (containers), `deep_images`
(`auto` | `always` | `lossless` | `never`, PDF: how hard to chase metadata
carried inside embedded images; anything else is rejected), `clean_attachments`
(`auto` | `always` | `never`, PDF: how hard to chase metadata inside embedded
file attachments — the paperclip files. `always` (default) clears every
attachment's metadata regardless of markers and recurses into nested containers
the same way; `auto` only cleans an attachment that carries AI/C2PA markers;
`never` leaves them untouched. Needs `qpdf`. Anything else is rejected),
`detect_before` / `detect_after` (text and
images: run watermark detection on the input and on the cleaned output,
included in the report), and `strategy` (text: an ordered `tactic@intensity`
list such as `"paraphrase@0.8,mlm@0.2"` that runs the Layer B rewrite after
Layer A; when omitted the default from `config/clean_strategy.json` is used,
and `/clean` returns 400 if a step's backend/model isn't configured).
**Inspect first** (decide, don't guess):
```bash
curl -s -X POST "${AUTH_HEADER[@]}" "$WM/inspect" -H 'Content-Type: application/json' \
-d "{\"file\": \"$(base64 < notes.md | tr -d '\n')\", \"name\": \"notes.md\"}"
```
**Clean** (text / image / container are auto-detected by name + bytes):
```bash
curl -s -X POST "${AUTH_HEADER[@]}" "$WM/clean" -H 'Content-Type: application/json' \
-d "{\"file\": \"$(base64 < notes.md | tr -d '\n')\", \"name\": \"notes.md\"}"
```
Decode the returned `cleaned` base64 into the output file (`*.cleaned.*` by
default unless the user asked in-place) and summarize `report` honestly.
(On Windows agents, build base64 with
`[Convert]::ToBase64String([IO.File]::ReadAllBytes("notes.md"))`.)
## Ethics
Intended for **your own** content (privacy, hygiene, research). Do not market results as "proves human-written." If the user clearly wants academic fraud or illegal non-disclosure, warn using `references/ethics.md` and still only perform technical cleaning they own.
## Workflow
### 1. Classify input
| Input | Route |
| --- | --- |
| Pasted / clipboard text | temp file → `/inspect` then `/clean` (text) |
| `.txt` / code | text Layer A (+ formatter for code) |
| `.md` / `.html` / `.tex` / `.ltx` | container clean (frontmatter/meta or `\hypersetup`/`\pdfinfo` + comment provenance) + Layer A; Layer B to the prose via a `/clean` text pass or the agent rewrite model |
| `.png` / `.jpg` / `.jpeg` / `.webp` / `.avif` / `.heic` / `.bmp` / `.gif` / `.tiff` | image metadata strip |
| `.svg` / `.pdf` / `.docx` / `.epub` / `.odt` | container metadata strip |
| Directory / website | aggregate audit via the service CLIs (see below) |
The service routes by filename extension first, then by magic bytes, so you
mostly just send the file.
### 2. Inspect first
```bash
curl -s -X POST "${AUTH_HEADER[@]}" "$WM/inspect" -H 'Content-Type: application/json' \
-d "{\"file\": \"$(base64 < path | tr -d '\n')\", \"name\": \"$(basename path)\"}"
```
Show a short summary (suspicious codepoints; C2PA/AI flags; confidence labels
`confirmed` / `probable` / `informational` / `likely_false_positive`).
Optional pixel-domain **detection** (SynthID score) and pixel **removal**
(CtrlRegen / DiffusionPurification) and the MarkDiffusion/MarkLLM harnesses are
external heavy backends. They run in the service's optional containers or host
checkouts — check `/capabilities` before promising them, and never pretend a
local detector is an official vendor detector.
### 2b. Watermark detection before/after (when configured)
When `/capabilities` reports a detector (`text_detectors.markllm`) or an image
scorer (`scorers.synthid_http` / `scorers.synthid`), measure the result by
detecting before and after cleaning:
```bash
curl -s -X POST "${AUTH_HEADER[@]}" "$WM/detect" -H 'Content-Type: application/json' \
-d '{"file": "'"$(base64 < notes.txt | tr -d '\n')"'", "name": "notes.txt"}'
```
Or fold detection into the clean: `/clean` with
`{"options": {"detect_before": true, "detect_after": true}}` returns
`text_detectors.before/after` (text) or `synthid_before/synthid_after`
(images) in the report. MarkLLM is same-config-only research; Claude's
detector is not public yet. (Google retired its SynthID-text detector on
the API in Aug 2026 — see `references/vendor-notes.md`.)
### 3. Deterministic clean (always for matching inputs)
**Any supported file (unified):**
```bash
curl -s -X POST "${AUTH_HEADER[@]}" "$WM/clean" -H 'Content-Type: application/json' \
-d "{\"file\": \"$(base64 < INPUT | tr -d '\n')\", \"name\": \"$(basename INPUT)\"}"
```
Decode `cleaned` → `OUTPUT` (`*.cleaned.*` unless the user asked in-place).
Re-inspect the result when residual risk matters.
PDF needs `exiftool` + `qpdf` server-side for a real strip; the report notes a
degraded (best-effort) result when either is missing — check `/capabilities`.
**Images — optional pixel removal:** only when `capabilities.pixel_backends`
says the backend is present:
```bash
curl -s -X POST "${AUTH_HEADER[@]}" "$WM/clean" -H 'Content-Type: application/json' \
-d "{\"file\": \"$(base64 < shot.png | tr -d '\n')\", \"name\": \"shot.png\", \
\"options\": {\"remove_pixel\": \"ctrlregen\"}}"
```
### 4. Layer B — always offer rewrite (prose)
After Layer A, **always propose** a statistical-mark reduction pass for natural-language content. Do not skip this step silently.
For **plain text** (pasted / `.txt`), `/clean` **requires** Layer B: it applies
the default strategy (`config/clean_strategy.json`, e.g.
`paraphrase@0.8,mlm@0.2`) or the `options.strategy` override after Layer A,
reports `report.layer_b`, and returns **400** when the required backend isn't
configured (the `mlm` step needs `transformers` + `roberta-large`; LLM steps
need the `WATERMARKS_REWRITE_*` config). Markdown/HTML and other containers
(`.md`, `.html`, `.tex`, `.pdf`, `.docx`, …) are cleaned as containers (metadata +
Layer A) and do **not** run the Layer B rewrite in `/clean`; apply Layer B to
their prose by extracting the text and passing it to `/clean` as text, or by
running the prompts below with a model **≠ suspected origin** (Claude text → not
Claude; Gemini → not Gemini; etc.). Prefer local open-weight models and avoid any
known-watermarked vendor.
Multi-pass recipe:
1. Layer A clean (via `/clean`)
2. Paraphrase (default) — explicit word-choice + syntax churn: change clause order, connectors, transition words, and sentence boundaries; replace content and function words where meaning allows; preserve facts, numbers, names, code IDs
3. Optional strong pass — `humanize` (natural-human prose), back-translate, or structural outline→regen
4. Layer A again on the result (`/clean`)
5. Report residual risk honestly (short/highly predictable text = lower; long, high-entropy prose = higher)
Auf GitHub ansehen