Designed for Claude Code, also compatible with Codex
LangChain Content Blocks (Python)
Overview
On Claude, AIMessage.content is list[dict] even for pure text — so any
code from an OpenAI-first tutorial that calls message.content.lower() or
message.content.split() crashes with AttributeError: 'list' object has no attribute 'lower' on the first production Claude call (P02).
Multi-modal code that works on GPT-4o breaks on Claude because pre-1.0
image-block shapes differed across providers (P64). Multi-turn Claude
replay with extended thinking fails with
anthropic.BadRequestError: missing signature when prior thinking
blocks are stripped. Forced tool_choice prevents
stop_reason="end_turn" and loops forever (P63).
This is the deep-dive companion to langchain-model-inference. That
skill's references/content-blocks.md covers the str vs list[dict]
divergence and a safe text extractor. This skill goes further:
tool_use block iteration mechanics — IDs, args as dict vs JSON string, streaming deltas
At least one provider package: pip install langchain-anthropic langchain-openai
For extended thinking: langchain-anthropic >= 1.0 and Claude Sonnet 4+ / Opus 4+
For citations: anthropic >= 0.40 and Claude Sonnet 4+
Familiarity with langchain-model-inference (reads references/content-blocks.md first)
Instructions
Step 1 — Learn the block-type taxonomy
LangChain 1.0 defines six typed content blocks on AIMessage.content
(and on chunks during streaming):
Block type
Produced by
Notes
text
All providers
On Claude, always wrapped as [{"type":"text","text":"..."}]
tool_use
Claude, GPT-4o, Gemini
Always round-trip via msg.tool_calls, not hand-parsed
tool_result
You (via ToolMessage)
One per tool_use; tool_call_id must match byte-for-byte
image
Claude vision, GPT-4o, Gemini
Universal 1.0 shape; adapter handles wire format per provider
thinking
Claude extended thinking only
Must preserve signature for replay
document
Claude citations API (Sonnet 4+)
Input-side only; citations attach to output text blocks
See Block-Type Matrix for the full table
with streaming behavior and per-type gotchas.
Step 2 — Iterate mixed content safely
For most code, use the helpers:
text = msg.text() # concatenated text across all text blocks
tool_calls = msg.tool_calls # normalized list[ToolCall]
usage = msg.usage_metadata # input_tokens, output_tokens, cache_*
Hand-roll block iteration only when you need to (a) preserve order,
(b) extract thinking blocks for replay, or (c) read citations
metadata from text blocks. Order-preserving iteration:
Step 3 — Compose multi-modal messages with the universal image block
import base64
from pathlib import Path
from langchain_core.messages import HumanMessage
defimage_block(path: str) -> dict:
data = base64.standard_b64encode(Path(path).read_bytes()).decode("ascii")
mime = {"png": "image/png", "jpg": "image/jpeg",
"jpeg": "image/jpeg", "webp": "image/webp"}[
Path(path).suffix.lstrip(".").lower()]
return {
"type": "image",
"source_type": "base64", # or "url""data": data,
"mime_type": mime,
}
msg = HumanMessage(content=[
image_block("screenshot.png"), # put image FIRST
{"type": "text", "text": "What is broken here?"}, # instruction LAST
])
response = claude.invoke([msg])
Three invariants:
contentmust be list[dict] when including non-text blocks.
Put the image before the instruction — Claude attends most to trailing tokens.
Respect provider limits (Anthropic: 5 MB/image, up to 20 images; OpenAI: 20 MB/image; Gemini: 20 MB/request total).
LangChain's adapter translates the universal shape to each provider's
wire format. See Multi-Modal Composition
for the full adapter table, MIME-type compatibility, and the
document/citations pattern.
Step 4 — Iterate tool_use correctly across stream deltas
Canonical non-streaming:
for tc in msg.tool_calls:
output = tools[tc["name"]](**tc["args"])
history.append(ToolMessage(content=str(output), tool_call_id=tc["id"]))
tc["args"] is already a parsed dict — do not json.loads it.
tc["id"] is provider-shaped (toolu_* on Anthropic, call_* on
OpenAI, 24+ chars) and must be copied verbatim to the ToolMessage.
Streaming is different. tool_use.input arrives as partial JSON
fragments across on_chat_model_stream events. Buffer with
tool_call_chunks, parse once at on_chat_model_end:
from collections import defaultdict
import json
partial = defaultdict(str) # index -> accumulated JSON fragment
meta = {} # index -> {name, id}asyncfor event in model.astream_events({"messages": [...]}, version="v2"):
if event["event"] != "on_chat_model_stream":
continuefor tc_chunk ingetattr(event["data"]["chunk"], "tool_call_chunks", []) or []:
idx = tc_chunk["index"]
if tc_chunk.get("name"):
meta[idx] = {"name": tc_chunk["name"], "id": tc_chunk["id"]}
if tc_chunk.get("args"):
partial[idx] += tc_chunk["args"]
completed = [{**meta[i], "args": json.loads(partial[i])} for i in meta]
See Tool-Use Iteration for
multi-tool-per-turn handling, ToolMessage ordering, and the forced-
tool_choice infinite-loop trap (P63).
Step 5 — Preserve Claude thinking blocks for replay
Claude extended thinking (Sonnet 4+, Opus 4+) returns thinking blocks
carrying a cryptographic signature. The next turn must round-trip
those blocks intact or Anthropic rejects the request:
Pre-resize to < 5 MB (1024x1024 JPEG 85 is ~500 KB)
openai.BadRequestError: Invalid image data
Hand-rolled image_url with wrong prefix
Use the universal block; adapter emits the data:image/...;base64, prefix
Infinite agent loop
Forced tool_choice inside a loop (P63)
Use tool_choice="auto" for agents; forced-choice only for single-call extraction
json.JSONDecodeError inside stream loop
Parsing partial tool_use.input fragment
Buffer in a defaultdict(str); parse once at on_chat_model_end
Citations silently missing
Read via msg.text() which strips metadata
Iterate msg.content and read block["citations"] on text blocks
Examples
Single-shot multi-modal on Claude + GPT-4o with one message object
msg = HumanMessage(content=[
image_block("ui.png"),
{"type": "text", "text": "Identify the broken UI element."},
])
# Same message works on both providers via adapter translation
claude_resp = claude.invoke([msg])
gpt4o_resp = gpt4o.invoke([msg])
Multi-turn Claude replay with extended thinking
claude = ChatAnthropic(
model="claude-sonnet-4-6",
max_tokens=8192,
thinking={"type": "enabled", "budget_tokens": 4096},
)
ai_1 = claude.invoke([HumanMessage(content="What is the capital of France?")])
# ai_1.content == [{"type":"thinking",...,"signature":"..."}, {"type":"text",...}]# Turn 2 — pass ai_1 VERBATIM
ai_2 = claude.invoke([
HumanMessage(content="What is the capital of France?"),
ai_1, # thinking preserved
HumanMessage(content="And the population?"),
])
See Thinking Blocks for the full replay
invariants and persistence pattern.
Extracting Claude citations from document input
doc_block = {
"type": "document",
"source": {"type": "base64", "media_type": "application/pdf", "data": pdf_b64},
"title": "Q3 Earnings Report",
"citations": {"enabled": True},
}
resp = claude.invoke([HumanMessage(content=[
doc_block,
{"type": "text", "text": "What drove revenue this quarter?"},
])])
for block in resp.content:
if block.get("type") != "text":
continueprint(block["text"])
for c in block.get("citations", []):
print(f" -> {c['document_title']}: {c['cited_text']!r}")
msg.text() flattens this — you lose citations. See
Multi-Modal Composition for the
full document block reference including supported source types.
Streaming tool_use with live argument rendering
See Tool-Use Iteration for the
complete tool_call_chunks accumulator including multi-tool-per-turn
handling and the ToolMessage ordering invariant.