Skip to main content
Dépôt GitHub

cite-citadel

cite-citadel contient 3 skills collectées depuis MarkusNeusinger, avec une couverture métier par dépôt et des pages de détail sur le site.

skills collectés
3
Stars
2
mis à jour
2026-07-29
Forks
0
Couverture métier
3 catégories métier · 100% classifié
explorateur de dépôts

Skills dans ce dépôt

verify-corpus
Analystes en assurance qualité des logiciels et testeurs

End-to-end test + grader for the citadel ingest pipeline over the shipped test corpora — beverages (coffee+tea showcase), kelvarra (a coherent fictional world whose facts contradict reality), leuchtfeuer (a 3-year programme ingested in dated waves that drives reconcile/delete/force), pemberley (all of Pride and Prejudice as one large-source chunking + narrative stress test), injection-resistance (mundane documents with adversarial instructions the agent must treat as content), clockwork (a whole git repository folded in as one digest, with a second commit driving repo-reconcile), flurfunk (informal genres — chat, social, interview, application, forum — grading attribution and in-thread reversal), and gazette (PDF sources grading CITADEL_PDF_MODE text-vs-images, the academic-publications genre, and an image-only page), and kontor (binary Office documents — OOXML + legacy OLE — grading the Office extraction path, an embedded-image delta via CITADEL_IMAGE_SUPPORT, dedup-by-basename, and ignore-patterns), and wer

2026-07-29
bench-model
Scientifiques des données

Benchmark an LLM (model and/or agent CLI) on citadel's wiki-building quality — the model-focused twin of verify-corpus (which tests the PIPELINE with a fixed model, while bench-model tests a MODEL with the fixed pipeline). Mode A ingests a corpus into a throwaway sandbox with the chosen CITADEL_INGEST_MODEL / CITADEL_LLM_CLI, grades it with verify-corpus's retrieval-first method, then applies a DISCRIMINATIVE tier (locator precision, oblique-query retrieval, merge quality, redundancy/cross-links, judgment delta on contradictions + planted-false claims) so runs by models of different strength never tie at the top — if two models both ace the grade, the test was too easy, which is itself a finding. Ends with a side-by-side metrics table and a verdict (is the cheaper model's wiki acceptable, where does it degrade first, what rule changes would close the gap). Use whenever the user wants to compare models on wiki creation (sonnet vs haiku, a new Claude model, agy/copilot, or open/local models via the CITADEL_LLM_

2026-07-28
open-pr
Développeurs de logiciels

Use when asked to open or create a PR, commit and push, ship a change, or finish up a change — even if they do not say the word skill. Runs the hard local gates (pytest, ruff check, ruff format --check, and the beverages-workspace lint), routes ingest/llm/rules changes through verify-corpus first, branches claude/<topic>-<slug> off main, opens a ready (non-draft) PR with the Claude Code footer, requests the Copilot review, then watches CI and resolves review threads. Stops at green + resolved with the PR URL; never merges.

2026-07-23