- name
- building-agents
- description
- Use when building or restructuring an LLM agent — provider adapter, tool calling, structured output, RAG, agent loop, eval gate, cost routing, tracing, MCP server — model-agnostic across OpenAI/Anthropic/Gemini/OSS so a model swap is a config change. NOT vector-store SQL alone (that is `postgresdb`) or service deployment (that is `deployment`).
- tags
- ["agents","llm","mcp","rag","evals","ai"]
- recommends
- ["secure-coding","deployment"]
- origin
- risco
# Building production LLM agents (model-agnostic)
A thin provider adapter, a disciplined agent loop, schema-validated tools, provider-neutral RAG, eval gates, OTel tracing, and optionally an MCP server — so swapping OpenAI ↔ Anthropic ↔ Gemini ↔ OSS is a config change, not a rewrite.
## The one rule
> Program against a **capability interface**, never a vendor SDK. Vendor specifics (model id, tool-schema shape, JSON mode, caching, token limits) live behind one adapter resolved from config. Model names and prices rot — if one appears in business logic it's a bug, and re-verify the dated tables before quoting a number.
**Hand off instead when:** a new non-trivial feature has no approved spec + plan under `02-DOCS/wiki/sdd/` → stop and run [`specify`](../specify/SKILL.md) first (method: [`sdd`](../sdd/SKILL.md)), which routes back here once the plan is approved; one-line/low-risk changes go straight through. Anthropic-SDK internals (caching, thinking, batch) in a file that *only* imports `anthropic` → `claude-api` if your environment has it, since this skill stays multi-provider. Workspace scaffolding → [`harness`](../harness/SKILL.md). Choosing *which* coding agent to use → agent-eval territory. Pure prompt-wording tuning with no architecture change → prompt engineering, not this. A one-shot throwaway prompt, or no retrieval/tools/loop/evals at all → you don't need an agent; call the SDK directly and say so.
## Decision rules (read before writing code)
1. **Adapter first** — define the `LLMProvider` Protocol before any provider call.
2. **Smallest loop that works** — single-agent before multi-agent; ReAct only when the path is uncertain; plan-execute when steps are knowable. Multi-agent means orchestrator-worker with a semaphore-bounded parallel fan-out, never a free-for-all.
3. **Tools are typed contracts** — schema + validation + idempotency key on every side-effecting tool; no catch-all tools.
4. **Retrieve, don't stuff** — RAG when ground truth lives in data; cite or refuse.
5. **Eval before ship** — a golden set + regression gate in CI, or it's not production.
6. **Cheapest model that passes the eval** — route/cascade up, never default to flagship.
## The provider adapter (the heart of the skill)
The one payload to internalize. Python 3.12+, Pydantic v2, **async** so it composes directly with the bounded loop (and orchestrator-worker fan-out) in [`references/agent-loops-and-harness.md`](references/agent-loops-and-harness.md). Structured output is the quirk that differs most per vendor: strict JSON Schema (OpenAI), tool-forcing (Anthropic), `response_json_schema` (Gemini). Streaming, the Gemini and OSS/litellm adapters, tool-result plumbing, and a `route()` registry live in [`references/provider-abstraction.md`](references/provider-abstraction.md) — this excerpt is the load-bearing core, not the whole interface.
```python
from __future__ import annotations
import os
from typing import Literal, Protocol, runtime_checkable
from pydantic import BaseModel, Field
class Message(BaseModel):
role: Literal["system", "user", "assistant", "tool"]
content: str
class ToolSpec(BaseModel):
name: str
description: str
parameters: dict # JSON Schema for the tool's arguments
class Usage(BaseModel):
input_tokens: int = 0
output_tokens: int = 0
cost_usd: float = 0.0
class CompletionRequest(BaseModel):
model: str # resolved from config, e.g. "claude-sonnet-4-6" — never literal in logic
messages: list[Message]
tools: list[ToolSpec] = Field(default_factory=list)
response_schema: dict | None = None # JSON Schema -> structured output
temperature: float = 0.0
max_tokens: int = 1024
class CompletionResponse(BaseModel):
text: str = ""
tool_calls: list[dict] = Field(default_factory=list) # [{id, name, arguments}]
usage: Usage = Field(default_factory=Usage)
raw: dict | None = None
@runtime_checkable
class LLMProvider(Protocol):
# Async so it drives the async agent loop directly. The full interface in
# references/provider-abstraction.md adds stream() and embed().
async def complete(self, req: CompletionRequest) -> CompletionResponse: ...
class OpenAIAdapter:
def __init__(self, model: str) -> None:
from openai import AsyncOpenAI
self.model, self.client = model, AsyncOpenAI()
async def complete(self, req: CompletionRequest) -> CompletionResponse:
# Chat Completions shape (universal, still current); references/provider-abstraction.md
# gives the preferred Responses-API adapter. system stays a `system` role message here.
kwargs: dict = {"model": self.model, "messages": [m.model_dump() for m in req.messages],
"temperature": req.temperature, "max_tokens": req.max_tokens}
if req.tools:
kwargs["tools"] = [{"type": "function", "function": {"name": t.name, "description": t.description, "parameters": t.parameters}} for t in req.tools]
if req.response_schema:
kwargs["response_format"] = {"type": "json_schema", "json_schema": {"name": "out", "schema": req.response_schema, "strict": True}}
r = await self.client.chat.completions.create(**kwargs)
msg = r.choices[0].message
calls = [{"id": c.id, "name": c.function.name, "arguments": c.function.arguments} for c in (msg.tool_calls or [])]
return CompletionResponse(text=msg.content or "", tool_calls=calls, raw=r.model_dump(),
usage=Usage(input_tokens=r.usage.prompt_tokens, output_tokens=r.usage.completion_tokens))
class AnthropicAdapter:
def __init__(self, model: str) -> None:
from anthropic import AsyncAnthropic
self.model, self.client = model, AsyncAnthropic()
async def complete(self, req: CompletionRequest) -> CompletionResponse:
# QUIRKS: system is a top-level param (not a message); tools use input_schema (not function).
system = "\n".join(m.content for m in req.messages if m.role == "system") or None
turns = [{"role": m.role, "content": m.content} for m in req.messages if m.role != "system"]
kwargs: dict = {"model": self.model, "system": system, "messages": turns, "max_tokens": req.max_tokens, "temperature": req.temperature}
if req.tools:
kwargs["tools"] = [{"name": t.name, "description": t.description, "input_schema": t.parameters} for t in req.tools]
if req.response_schema: # structured output via tool-forcing
kwargs["tools"] = [{"name": "out", "description": "Emit the result", "input_schema": req.response_schema}]
kwargs["tool_choice"] = {"type": "tool", "name": "out"}
r = await self.client.messages.create(**kwargs)
text = "".join(b.text for b in r.content if b.type == "text")
calls = [{"id": b.id, "name": b.name, "arguments": b.input} for b in r.content if b.type == "tool_use"]
return CompletionResponse(text=text, tool_calls=calls, raw=r.model_dump(),
usage=Usage(input_tokens=r.usage.input_tokens, output_tokens=r.usage.output_tokens))
def get_provider(spec: str | None = None) -> LLMProvider:
"""Parse 'provider:model' (default from env LLM) into a concrete adapter."""
provider, _, model = (spec or os.environ["LLM"]).partition(":")
if provider == "openai":
return OpenAIAdapter(model)
if provider == "anthropic":
return AnthropicAdapter(model)
raise ValueError(f"unknown provider: {provider!r}")
# Gemini + OSS/litellm adapters, streaming, tool-result plumbing, and route() registry
# -> references/provider-abstraction.md
```
## Good vs Bad
Call-sites use the adapter and never name a model: `provider = get_provider(settings.llm)` (e.g. `"anthropic:claude-sonnet-4-6"`), then `await provider.complete(req)`. The two failures that survive that discipline:
```python
# BAD — parse-and-pray; wrong shape fails silently at 3am.
raw = (await provider.complete(req)).text
try:
data = json.loads(raw)
except json.JSONDecodeError:
data = {} # the bug is now invisible
```
```python
# GOOD — strict structured output + schema validation that fails loudly on drift.
class Answer(BaseModel):
sentiment: Literal["pos", "neg", "neu"]
score: float
req.response_schema = Answer.model_json_schema()
ans = Answer.model_validate_json((await provider.complete(req)).text)
```
```python
# BAD — unbounded loop; no cap/timeout/idempotency. Burns budget, repeats side effects, wedges.
while True:
resp = await provider.complete(req)
if not resp.tool_calls:
break
for call in resp.tool_calls:
await run_tool(call)
```
```python
# GOOD — bounded loop: step cap + per-tool timeout + idempotency key (safe to retry).
for step in range(max_steps):
resp = await provider.complete(req)
if not resp.tool_calls:
break
for call in resp.tool_calls:
async with asyncio.timeout(tool_timeout_s):
await run_tool(call, idempotency_key=call["id"])
# full loop, budgets, recovery -> references/agent-loops-and-harness.md
```
## Tools & structured output (minimum viable)
```python
from typing import Callable, Literal
from pydantic import BaseModel, ConfigDict, Field, ValidationError
class CreateInvoiceArgs(BaseModel):
model_config = ConfigDict(extra="forbid") # reject unknown keys from the model
customer_id: str = Field(min_length=1)
amount_cents: int = Field(gt=0)
currency: Literal["EUR", "USD"] = "EUR"
class ToolResult(BaseModel):
status: Literal["success", "warning", "error"]
summary: str
data: dict | None = None
next_actions: list[str] = Field(default_factory=list)
def _create_invoice(args: CreateInvoiceArgs) -> ToolResult:
invoice_id = f"inv_{args.customer_id}_{args.amount_cents}" # real impl: DB insert + idempotency
return ToolResult(status="success", summary=f"Created {invoice_id}", data={"id": invoice_id})
TOOLS: dict[str, tuple[type[BaseModel], Callable]] = {
"create_invoice": (CreateInvoiceArgs, _create_invoice),
}
def dispatch(name: str, raw_args: dict) -> ToolResult:
spec = TOOLS.get(name)
if spec is None:
return ToolResult(status="error", summary=f"unknown tool {name!r}", next_actions=["pick a registered tool"])
args_model, handler = spec
try:
args = args_model.model_validate(raw_args) # validate BEFORE side effects
except ValidationError as e:
return ToolResult(status="error", summary="invalid args", data={"errors": e.errors()},
next_actions=["fix the arguments and retry"])
return handler(args)
```
Schema design, sandboxing, idempotency, DI-scoped DB sessions, plus the RAG internals below — chunking, hybrid RRF, rerank, the citation grader, memory — are in [`references/tools-and-rag.md`](references/tools-and-rag.md).
## RAG in 30 lines (provider-agnostic embeddings)
```sql
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE IF NOT EXISTS docs (
id bigserial PRIMARY KEY,
content text NOT NULL,
embedding vector(1536) NOT NULL,
meta jsonb NOT NULL DEFAULT '{}'
);
CREATE INDEX IF NOT EXISTS docs_embedding_hnsw
ON docs USING hnsw (embedding vector_cosine_ops);
```
```python
async def embed(texts: list[str]) -> list[list[float]]:
# Same provider interface as completions; impl in references/tools-and-rag.md.
return await provider.embed(texts) # returns one 1536-d vector per text
async def retrieve(query: str, k: int = 5, min_sim: float = 0.25) -> list[dict]:
[q] = await embed([query])
rows = await db.fetch( # cosine distance <=>; similarity = 1 - distance
"SELECT id, content, 1 - (embedding <=> $1) AS sim "
"FROM docs ORDER BY embedding <=> $1 LIMIT $2",
q, k,
)
return [dict(r) for r in rows if r["sim"] >= min_sim]
async def answer(query: str) -> str:
chunks = await retrieve(query)
if not chunks: # refuse rather than hallucinate
return "I don't have grounded information to answer that."
context = "\n".join(f"[{c['id']}] {c['content']}" for c in chunks)
req = CompletionRequest(
model=settings.model_id,
messages=[Message(role="system", content="Answer ONLY from context; cite chunk ids like [12]."),
Message(role="user", content=f"{context}\n\nQ: {query}")],
)
return (await provider.complete(req)).text
```
## Evals & cost gates (the production line)
```python
import json
import statistics
import sys
import time
async def run_eval(golden_path: str, graders: list, thresholds: dict[str, float]) -> None:
cases = [json.loads(line) for line in open(golden_path)] # {"input","expected","meta"}
results = []
for case in cases:
t0 = time.perf_counter()
out = await provider.complete(CompletionRequest(model=settings.model_id,
messages=[Message(role="user", content=case["input"])]))
scores = {g.name: g.grade(case, out) for g in graders} # exact / schema / LLM-judge
results.append({"scores": scores, "cost": out.usage.cost_usd,
"ms": (time.perf_counter() - t0) * 1000})
n = len(results)
metrics = {
GitHub에서 보기