| name | openai-agents-sdk |
| description | Use when building Python agents with the OpenAI Agents SDK, including runners, tools, handoffs, guardrails, tracing, memory, multi-agent topology, and deterministic orchestration. |
| metadata | {"portable":true,"compatible_with":["Codex","codex"]} |
OpenAI Agents SDK
Operating contract
Inputs
| Input | Required | Purpose |
|---|
| Domain evidence | yes | agent goal, SDK/runtime version, tool and handoff inventory, authority limits, data classification, and evaluation cases |
Outputs
- Produce: agent topology, tool schemas, guardrails, handoff rules, tracing configuration, and runnable validation plan.
Capability and permission boundaries
Default to read-only analysis. Read only scoped records; redact secrets and regulated data. Writes, execution, network calls, production configuration, customer communication, billing changes, and delegation require explicit authority and an identified owner. Never widen tenant, time-window, or system scope implicitly.
Degraded mode
When required telemetry, evidence, execution, network access, or write authority is unavailable, return a partial result with each unassessed item labelled, preserve the safest existing state, and state the evidence or approval needed to continue. Never convert missing evidence into a pass.
Decision rules
| Condition | Action |
|---|
| Scope, owner, or threshold is missing | Stop the affected decision and request it |
| Evidence is incomplete but read-only analysis is safe | Produce a qualified partial result and gap list |
| A mutation exceeds authority or tenant boundary | Block it and route for approval |
| Evidence meets the stated threshold | Issue the output with provenance and owner |
Anti-Patterns
- Treating absent evidence as success. Fix: mark the check unassessed and name the missing source.
- Expanding one tenant or workflow to all tenants. Fix: enforce supplied scope at every query and action.
- Performing a production write during analysis. Fix: emit a reviewed change plan until authority is explicit.
- Reporting a metric without population, window, or source. Fix: attach all three.
- Hiding a failed threshold inside an average. Fix: report failure slices and the remediation owner.
Acknowledgement: Shared by Peter Bamuhigire, techguypeter.com, +256 784 464178.
Use When
- Build production AI agents with the OpenAI Agents SDK (Python) — 6 core primitives (Agent, Runner, Tools, Handoff, Guardrails, Tracing), multi-agent patterns (Centralized, Hierarchical, Decentralized, Swarm), dynamic/deterministic orchestration...
Evidence Produced
| Category | Artifact | Format | Example |
|---|
| Correctness | OpenAI Agents SDK contract test plan | Markdown doc covering Agent, Runner, Tools, Handoff, and Guardrails primitive tests | docs/ai/openai-agents-tests.md |
| Security | Agent guardrail and key handling note | Markdown doc covering tool whitelisting, output filtering, and API key rotation | docs/ai/openai-agents-security.md |
References
- Use the links and companion skills already referenced in this file when deeper context is needed.
Minimal Python SDK for building AI agents. Six primitives: Agent, Runner, Tools, Handoff, Guardrails, Tracing.
pip install openai-agents
export OPENAI_API_KEY="sk-..."
1. Agent — The Core Primitive
A configurable wrapper around an LLM that can take actions.
from agents import Agent
customer_service_agent = Agent(
name="Customer Service Agent",
model="gpt-4o",
instructions="""
You are a helpful customer service agent.
Handle returns, orders, and billing queries.
Escalate complex issues to a human.
""",
tools=[get_order_status, process_refund, track_shipment],
handoffs=[billing_agent, technical_agent],
)
Agent parameters: name, instructions (system prompt), model, tools, handoffs, guardrails, output_type
2. Runner — The Agent Loop
Runs the agent's reasoning loop (think → act → observe → repeat).
from agents import Agent, Runner
agent = Agent(name="Assistant", instructions="You are a helpful assistant.", model="gpt-4o")
result = Runner.run_sync(agent, "What is the capital of France?")
print(result.final_output)
import asyncio
result = await Runner.run(agent, "Summarise this document.", max_turns=10)
print(result.final_output)
Key parameter: max_turns — safety valve against infinite loops. Always set it.
result = Runner.run_sync(agent, user_message, max_turns=15)
3. Tools — Extending Agent Capabilities
Custom Tools (Python Functions)
Any Python function becomes a tool with @function_tool. The SDK reads the name, docstring, and type hints automatically.
from agents import function_tool
@function_tool
def get_order_status(order_id: str) -> str:
"""Gets the current status of a customer order.
Args:
order_id: The unique order identifier (e.g. ORD-12345)
"""
order = db.query("SELECT status FROM orders WHERE id = ?", [order_id])
return f"Order {order_id} is: {order['status']}"
@function_tool
def calculate_refund(order_id: str, reason: str) -> dict:
"""Calculate and process a refund for an order.
Args:
order_id: The order to refund
reason: Customer-provided reason for return
"""
amount = get_order_total(order_id)
return {"refund_amount": amount, "processing_days": 3}
Hosted Tools (Built-in OpenAI)
from agents import Agent
from agents.tools import WebSearchTool, FileSearchTool, CodeInterpreterTool
research_agent = Agent(
name="Research Agent",
model="gpt-4o",
instructions="Research topics thoroughly and provide cited summaries.",
tools=[
WebSearchTool(),
FileSearchTool(
vector_store_ids=["vs_abc123"]
),
CodeInterpreterTool(),
],
)
Agent as Tool
analysis_tool = analysis_agent.as_tool(
tool_name="RunAnalysis",
tool_description="Run deep financial analysis on the provided data."
)
orchestrator = Agent(
name="Orchestrator",
tools=[analysis_tool, data_fetcher],
)
Handoff vs as_tool:
handoff — complete transfer of control, caller stops
as_tool — caller delegates subtask, caller resumes after
4. Handoffs — Multi-Agent Delegation
from agents import Agent, Runner
billing_agent = Agent(
name="Billing Agent",
instructions="Handle all billing, payment, and invoice queries.",
tools=[lookup_invoice, process_payment],
)
technical_agent = Agent(
name="Technical Agent",
instructions="Resolve technical issues, bugs, and connectivity problems.",
tools=[check_system_status, reset_connection],
)
triage_agent = Agent(
name="Triage Agent",
instructions="""
Triage the user's request. If billing-related, route to Billing Agent.
If technical, route to Technical Agent. Otherwise, answer directly.
""",
handoffs=[billing_agent, technical_agent],
)
result = Runner.run_sync(triage_agent, "My invoice shows the wrong amount")
print(result.final_output)
Handoff prompt best practices:
- Explicitly name which conditions trigger a handoff in each agent's instructions
- Each specialist agent must clearly state its purpose and domain
- Use
result.last_agent to track which agent handled the final response
5. Multi-Agent Patterns
Centralized (Most Common)
One triage/orchestrator routes to specialists. Best for: customer support, internal assistants, helpdesks.
triage_agent = Agent(
name="Triage",
handoffs=[billing_agent, technical_agent, sales_agent],
)
Hierarchical
Multi-tier routing: triage → managers → specialists. Best for: deep research, complex enterprise workflows.
science_manager = Agent(name="Science Manager",
handoffs=[physics_agent, chemistry_agent])
history_manager = Agent(name="History Manager",
handoffs=[politics_agent, warfare_agent])
triage = Agent(name="Research Triage",
handoffs=[science_manager, history_manager])
Dynamic vs Deterministic Orchestration
def orchestrate(message: str):
if "complaint" in message.lower():
return Runner.run_sync(complaints_agent, message)
return Runner.run_sync(inquiry_agent, message)
triage_agent = Agent(
name="Triage",
instructions="Route requests to the most appropriate specialized agent.",
handoffs=[complaints_agent, inquiry_agent],
)
Decentralized (Debate / Brainstorm)
No central agent; agents exchange turns. Best for: ideation, debate, negotiation.
agents = [Agent(name=f"{role}", instructions=f"You are a {role}...") for role in roles]
for i in range(rounds):
result = Runner.run_sync(agents[i % len(agents)], history, session=session)
Swarm (Parallel Exploration)
Many simple agents in parallel, results aggregated. Best for: creative generation, optimization.
import concurrent.futures
def run_agent(agent, prompt):
return Runner.run_sync(agent, prompt).final_output
with concurrent.futures.ThreadPoolExecutor() as executor:
results = list(executor.map(lambda a: run_agent(a, prompt), specialist_agents))
summary = Runner.run_sync(aggregator_agent, "\n".join(results))
6. Memory Management
from agents import SQLiteSession
session = SQLiteSession("my_app_db")
last_agent = triage_agent
while True:
user_input = input("You: ")
result = Runner.run_sync(last_agent, user_input, session=session)
print("Agent:", result.final_output)
last_agent = result.last_agent
Sliding window for long conversations:
if len(messages) > 20:
summary = Runner.run_sync(summarizer_agent, "\n".join(messages[-20:]))
messages = [{"role": "system", "content": f"Conversation so far: {summary.final_output}"}]
7. Guardrails — Safety Validation
from agents import Agent, Runner, GuardrailFunctionOutput, RunContextWrapper
from agents import input_guardrail, output_guardrail, InputGuardrailTripwireTriggered
from agents.types import TResponseInputItem
@input_guardrail
async def scope_guardrail(
ctx: RunContextWrapper[None],
agent: Agent,
input: str | list[TResponseInputItem]
) -> GuardrailFunctionOutput:
"""Only allow customer service queries."""
is_valid = any(kw in str(input).lower()
for kw in ['order', 'refund', 'account', 'billing', 'payment', 'delivery'])
return GuardrailFunctionOutput(
output_info="valid" if is_valid else "out of scope",
tripwire_triggered=not is_valid,
)
agent = Agent(
name="CS Agent",
instructions="Handle customer service queries.",
input_guardrails=[scope_guardrail],
)
try:
result = Runner.run_sync(agent, "What's the meaning of life?")
except InputGuardrailTripwireTriggered:
print("Sorry, I can only help with customer service queries.")
8. Third-Party Models (OpenAI-Compatible)
from agents.extensions.models.litellm_model import LitellmModel
agent = Agent(
name="Multi-Model Agent",
model=LitellmModel(model="deepseek/deepseek-chat", api_key="..."),
instructions="You are a helpful assistant.",
)
import agents
agents.set_default_openai_api("chat_completions")
agents.set_default_openai_client(AsyncOpenAI(
base_url="https://api.deepseek.com/v1",
api_key=os.environ["DEEPSEEK_API_KEY"],
))
Anti-Patterns
| Anti-Pattern | Fix |
|---|
No max_turns limit | Always set max_turns — prevents runaway agent loops |
| Vague agent instructions | Explicitly name routing conditions and domain boundaries |
| Too many tools per agent | Keep 5–8 tools max per agent — too many confuses the model |
| Write actions without approval | Add human approval gate before any irreversible action |
| No session management | Use SQLiteSession for multi-turn conversations |
| Missing error handling | Wrap Runner.run_sync() in try/except for guardrail errors |
Source: Habib — Building Agents with OpenAI Agents SDK (Packt, 2025)