| name | deglacer |
| description | MANDATORY gate BEFORE running jq on any .jsonl under ~/.claude/ or reading past CC sessions. Invoke FIRST when introspecting conversations, searching session history, parsing transcripts, or building tools that read ~/.claude/projects/ data. Provides the CC JSONL schema reference and `deglacer` CLI tool, plus routing to `deja` ranked content search — prevents the 54-attempt fumble pattern where Claudes guess at field names. Triggers on 'what happened last session', 'find when we discussed', 'parse session', 'read conversation', 'session history', 'token usage', 'deglacer', 'deja'. Do NOT use for git history (use git log) or your own current-session context. (user)
|
| allowed-tools | ["Bash","Read","Grep","Glob"] |
Déglacer — CC Session JSONL Reference
Deglazing the pan to lift the fond — extracting the good bits from past sessions.
When to Use This Skill
You are working with Claude Code session data. This includes:
- Introspecting past conversations — "what did we discuss last session?", "when did we first talk about X?"
- Searching session history — finding sessions that mention a topic, tool, or file
- Parsing JSONL transcripts — extracting human messages, tool calls, thinking blocks, token usage
- Building tooling — anything that reads
~/.claude/projects/ session data
- Debugging session format — understanding why a jq query returns nothing
Do NOT guess at the schema. The CC JSONL format has multiple entry types, triple-duty user entries, streaming-duplicated message.ids, and inconsistent field presence across versions. This reference is the source of truth.
When NOT to Use
- Git history — use
git log / git blame for code change history
- Current conversation state — you already have context, no need to parse your own session
- Non-CC JSONL files — this schema is specific to Claude Code sessions
deglacer — The CLI Tool
deglacer is the CC JSONL extraction CLI, installed as a uv tool. Use it instead of raw jq for structured extraction.
uv tool install 'deglacer @ git+https://github.com/spm1001/deglacer'
deglacer SESSION.jsonl
Commands
deglacer SESSION.jsonl
deglacer --summary SESSION.jsonl
deglacer --with-tools SESSION.jsonl
deglacer --with-thinking SESSION.jsonl
deglacer --last 5 SESSION.jsonl
deglacer --json SESSION.jsonl
deglacer --stats SESSION.jsonl
deglacer --timeline SESSION.jsonl
deglacer --find "search term"
deglacer --recent
deglacer --recent 10
deglacer --today
deglacer --since 2026-03-25
Combining Flags
deglacer --with-tools --last 10 SESSION.jsonl
deglacer --with-tools --with-thinking --json ...
deglacer --summary --last 5 SESSION.jsonl
deja — Ranked Content Search (optional companion)
deja is a third-party single-binary search engine (MIT, Go — no LLM, no embeddings) that maintains a local BM25 index over all CC sessions. It is the right tool for "find when we discussed X" across months of history: deglacer --find substring-scans recent sessions only; deja ranks the whole estate.
It ships separately from deglacer — check before reaching for it:
command -v deja || echo "not installed"
deja distinctive terms
Measured routing (13-task bench, banc/session-search, 2026-08-09 — private estate eval):
- Default: deja with 2–3 distinctive terms — ranked, ~12 results, ~2s warm; 85% recall on the bench. Whole-question queries dropped that to 64% (AND-matching cliff).
- Backstop when deja returns nothing or absence must be proven:
rg -li -F 'fragment' ~/.claude/projects -g '*.jsonl' — 100% recall on the bench (it sees raw bytes deja's tokeniser can miss, e.g. dotted versions like 25.12.5), but ~120 unranked files per query.
deglacer --find fragment — only worth it inside its window (the 200 most-recently-modified sessions); short substrings only.
- hit@1 was 8–15% for every tool — treat any result list as a candidate set to skim, never as an oracle.
Quirks that matter
- Query with a few distinctive terms, not the whole question. Multi-word queries are AND-matched (filler words dropped; double-quoted phrases mean contiguous text), so a six-content-word question routinely matches zero sessions — the tool itself says "try fewer words". Two or three rare terms is the sweet spot;
--re PATTERN for regex.
deja --help is not help — it searches for the literal string --help. Real flags and subcommands do exist: --since 30d, --project, --limit N, --json, and version, show <id>, blame <path>, resume <id>, stats, doctor among others. The repo README is the full surface.
- It auto-indexes on every run. Warm runs answer in ~2 seconds; the first run after weeks of inactivity re-indexes the backlog and can take minutes. Don't pipe it through
head (SIGPIPE kills it mid-index) — redirect to a file and read that.
- Results are capped (default ~15 sessions) and recency-weighted, NOT exhaustive. A session containing your term can be absent when the term is common across your history (measured: a 3-week-old session with 3 matches lost every slot to fresher, denser hits). Absence from results is not absence from history: re-probe with a rarer term or raise
--limit before concluding something was never discussed.
- Echo hits. It indexes tool results as well as prose, so a session that quotes old content (reading a handoff, grepping a transcript) matches alongside the original. Use the date column to tell originals from echoes.
From deja result to deglacer
[claude] owner/repo · Jul 19 · 511191c5-458 — 12 matches
The third field is a session-id prefix. deja show <id> prints the conversation directly; for schema-aware extraction (tool calls, thinking, token usage, timelines), resolve to the file and hand it to deglacer:
deglacer --summary ~/.claude/projects/*/511191c5-458*.jsonl
File Discovery
Sessions live at:
~/.claude/projects/{encoded-cwd}/{session-uuid}.jsonl
Where {encoded-cwd} replaces / with - in the project path.
Subagent transcripts: {session-uuid}/subagents/agent-{id}.jsonl.
Find recent sessions:
ls -lt ~/.claude/projects/*/*.jsonl | head -20
Find sessions for a project:
ls -lt ~/.claude/projects/-home-modha-Repos-myproject/*.jsonl
Match session to slug/name:
head -1 SESSION.jsonl | jq '{sessionId, slug, version}'
The Schema
Each line in a .jsonl file is one JSON object. The .type field discriminates.
Entry Types
| Type | Purpose | Has timestamp? |
|---|
assistant | Claude's response | Yes |
user | Human msg / tool result / skill injection | Yes |
progress | Streaming bash/hook/agent output | Yes |
system | Turn timing, API errors, slash commands | Yes |
summary | Context compaction | No |
queue-operation | Input typed while Claude busy | Yes |
last-prompt | Records last user text | No |
custom-title | User-set session name | No |
agent-name | Session agent name | No |
file-history-snapshot | File backup state | No |
pr-link | Created PR reference | Yes |
saved_hook_context | Persisted hook output | Yes |
Common Fields (on user/assistant entries)
uuid string Unique entry ID
parentUuid string? Previous entry (linked list)
sessionId string Session UUID (matches filename)
timestamp string ISO 8601
cwd string Working directory
gitBranch string Current git branch
version string CC version (e.g. "2.1.85")
slug string Human-readable session name
userType string Always "external"
entrypoint string "cli" (absent in v2.0.x)
isSidechain boolean Side conversation flag
assistant entries
{
"type": "assistant",
"message": {
"id": "msg_...",
"type": "message",
"role": "assistant",
"model": "claude-opus-4-6",
"content": [],
"stop_reason": "end_turn" | "tool_use" | null,
"usage": {
"input_tokens": 119,
"cache_creation_input_tokens": 18531,
"cache_read_input_tokens": 36004,
"output_tokens": 500
}
},
Content block types:
{type: "text", text: "..."} — Claude's text
{type: "tool_use", id: "toolu_...", name: "Bash", input: {...}} — tool call
{type: "thinking", thinking: "...", signature: "..."} — extended thinking
DRAGON: Multiple entries share the same message.id. CC streams
incremental updates. Merge content blocks by message.id, dedup
tool_use blocks by their id field. deglacer handles this automatically.
DRAGON: stop_reason is null in older sessions (pre-v2.1.79).
user entries (TRIPLE DUTY)
The user type serves three purposes. Discriminate with:
| Subtype | How to detect | Content shape |
|---|
| Human message | typeof content === "string", has permissionMode | String |
| Tool result | Has toolUseResult | Array of {type: "tool_result"} |
| Skill/system injection | isMeta: true | Array of {type: "text"} |
Human message:
{
"type": "user",
"message": {"role": "user", "content": "the actual human text"},
"permissionMode": "default",
"promptId": "..."
}
Tool result:
{
"type": "user",
"message": {"role": "user", "content": [
{"type": "tool_result", "tool_use_id": "toolu_...", "content": "output text"}
]},
"toolUseResult": {},
"sourceToolAssistantUUID": "..."
}
toolUseResult shapes:
| Tool | Keys |
|---|
| Bash | stdout, stderr, interrupted, isImage, noOutputExpected |
| Bash (large) | + persistedOutputPath, persistedOutputSize |
| Write/Edit | content, filePath, originalFile, structuredPatch, type |
| Read | file, type |
| Agent | agentId, agentType, content, prompt, status, totalDurationMs, totalTokens, totalToolUseCount, usage |
| Error | Bare string — wording drifts by CC version: older sessions say "User rejected tool use", current ones "The user doesn't want to proceed with this tool use. The tool use was rejected…" (both live in the corpus, the new form now the majority). Count rejections by toolUseResult being a string, or match a stable substring (rejected / doesn't want to proceed) — never the full documented sentence. |
Token Counting
The input_tokens field is ONLY the non-cached portion.
Real input = input_tokens + cache_creation_input_tokens + cache_read_input_tokens.
summary entries (minimal)
{"type": "summary", "leafUuid": "...", "summary": "short text"}
No uuid, parentUuid, timestamp, version, or sessionId. Three fields only.
system entries
.subtype discriminates:
turn_duration: {durationMs, messageCount}
api_error: {error: {status, headers, requestID}, retryInMs, retryAttempt}
local_command: {content: "...", level: "info"} — slash commands
jq Recipes (when deglacer isn't enough)
Quick schema discovery (do this FIRST, not head | jq .):
jq -r '.type' FILE.jsonl | sort | uniq -c | sort -rn
Extract human messages:
jq -r 'select(.type == "user" and .permissionMode and (.isMeta | not))
| .message.content' FILE.jsonl
Extract assistant text (handles multi-block):
jq -r 'select(.type == "assistant")
| [.message.content[]? | select(.type == "text") | .text]
| select(length > 0) | join("\n")' FILE.jsonl
Extract tool calls:
jq -c 'select(.type == "assistant")
| [.message.content[]? | select(.type == "tool_use")
| {tool: .name, input_keys: (.input | keys)}]
| select(length > 0)' FILE.jsonl
Session timeline:
jq -c 'select(.type == "user" or .type == "assistant")
| {ts: .timestamp, type, model: .message.model?}' FILE.jsonl
Token usage per turn:
jq -c 'select(.type == "assistant") | .message.usage
| {in: (.input_tokens + .cache_creation_input_tokens + .cache_read_input_tokens),
out: .output_tokens}' FILE.jsonl
Find sessions mentioning a term:
for f in ~/.claude/projects/*/*.jsonl; do
if jq -e 'select(.type == "user" and (.message.content | type) == "string"
and (.message.content | test("term"; "i")))' "$f" >/dev/null 2>&1; then
echo "$f"
fi
done
Anti-Patterns (DON'T)
| Don't | Why | Do instead |
|---|
jq -s '.' on JSONL | Slurps entire file into memory as array | Stream line-by-line (default jq behaviour) |
jq '.[]' on JSONL | JSONL isn't an array | Each line is already a separate object |
.role at top level | Role is at .message.role, not top-level | Use .type for entry type |
.type == "message" | No such type | Types: user, assistant, progress, system, etc. |
.type == "human" | No such type | .type == "user" + check it's not a tool result |
head -1 | jq . for discovery | Wastes a turn, first line may be queue-operation | jq -r '.type' | sort | uniq -c |
| Assume content is string | Assistant content is always array; user content varies | Check type before accessing |
2>/dev/null on everything | Hides real errors | Understand the schema, don't hedge |
| Guess at field names | 39% of jq-on-.claude commands are schema discovery | Read this reference |
Grep with a space after the colon ("skill": "x") | CC serializes JSONL compact — "skill":"x" — so the spaced pattern is a false zero that reads like absence | jq on the parsed field, or match the compact form |
| Treat deja's list as exhaustive | Top-K, recency-weighted — common terms rank-cut older sessions | Re-probe with a rarer term before claiming absence |