Skip to main content

ai-agent-patterns

Guides expert-level ai agent patterns implementation: ai-ml and architecture decision frameworks, production-ready patterns, and concrete templates for ai agent patterns workflows. Use when the user asks about ai agent patterns, ai agent patterns configuration, or ai-ml best practices for ai projects. Do NOT use when the user needs a different ai ml engineering capability -- check sibling skills in the ai ml engineering subcategory.

Jump to install

Source facts

Repository
FerroxLabs/murage
Last source activity
September 1, 2026 at 13:26
Detected SKILL.md language
English
Stars
0
Forks
0

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

File Explorer
2 files

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
ai-agent-patterns
description
Guides expert-level ai agent patterns implementation: ai-ml and architecture decision frameworks, production-ready patterns, and concrete templates for ai agent patterns workflows. Use when the user asks about ai agent patterns, ai agent patterns configuration, or ai-ml best practices for ai projects. Do NOT use when the user needs a different ai ml engineering capability -- check sibling skills in the ai ml engineering subcategory.
license
Apache-2.0
metadata
{"author":"foundry-skills","version":"1.0.0","tags":"ai-ml architecture automation","category":"ai-machine-learning","subcategory":"ai-ml-engineering","depends":"","disclaimer":"none","difficulty":"advanced"}
# AI Agent Patterns ## When to Use **Use this skill when:** - A user needs to design or implement an autonomous or semi-autonomous AI agent system -- including single-agent, multi-agent, or hierarchical agent architectures - A user is choosing between ReAct, Plan-and-Execute, Reflexion, MRKL, or other agent control flow patterns and needs a principled decision framework - A user needs to architect the memory system for an agent -- including working memory, episodic memory, semantic memory (vector stores), or procedural memory (tool registries) - A user is building a tool-using LLM system and needs to design the tool interface, tool selection strategy, error handling, and safety rails - A user is debugging an agent that loops indefinitely, fails to complete tasks, hallucinates tool calls, or produces inconsistent results in production - A user needs to implement multi-agent coordination patterns -- orchestrator/worker, critic/generator, debate, or parallel execution with aggregation - A user is planning an agent system that must operate safely in production with guardrails, cost controls, and observability **Do NOT use this skill when:** - The user needs basic LLM prompt engineering without autonomous decision-making -- use the prompt-engineering skill instead - The user needs RAG (Retrieval-Augmented Generation) pipeline design without agent control flow -- use the rag-pipeline skill - The user wants fine-tuning or model training guidance -- use the model-fine-tuning skill - The user is building a simple LLM chain with no branching, tool use, or iteration -- that is a chain pattern, not an agent pattern - The user needs LLM evaluation and benchmarking framework design -- use the llm-evaluation skill - The user is asking about model selection, provider comparison, or API cost optimization in isolation -- use the llm-provider-selection skill --- ## Process ### 1. Classify the Agent Task and Capability Requirements Before touching any architecture, precisely define what the agent must do: - **Task type classification:** Is this a research/synthesis task (needs search + summarization), an execution task (needs code execution or API calls), a decision task (needs multi-step reasoning over structured data), or a conversational task (needs long-term memory and context management)? - **Autonomy level:** On a scale of Level 1 (human approves every step) to Level 5 (fully autonomous with no human review), determine the required autonomy for this use case. Most production systems should target Level 2-3. - **Tool inventory:** List every external capability the agent needs -- web search, code execution, database queries, API calls, file system access, calendar/email, etc. Each tool adds attack surface and failure modes. - **Context window budget:** Estimate the expected token budget per agent run. For long-horizon tasks, assume 10-50 tool call round trips, each consuming 500-2000 tokens in the context window. At GPT-4 pricing, a 50-step task with 128k context can cost $0.50-$2.00 per run. - **Latency tolerance:** Is this synchronous (user is waiting, target < 10 seconds per step) or asynchronous (background job, hours acceptable)? This determines whether you can afford slow reasoning models like o1 or need faster models like GPT-4o or Claude Haiku. ### 2. Select the Core Control Flow Pattern Choose the agent architecture based on task complexity and reliability requirements: - **ReAct (Reasoning + Acting):** The default for most agent use cases. The agent interleaves natural language reasoning (Thought:) with tool invocations (Action:) and observations (Observation:). Use for tasks with 3-15 steps where intermediate reasoning needs to be inspectable. Failure mode: reasoning loops and circular Thought chains. Add a hard step limit of 15-25 iterations. - **Plan-and-Execute:** The planner LLM produces a structured task plan (a list of steps with dependencies), then an executor LLM executes each step. Use when tasks are long-horizon (20+ steps), when individual steps are parallelizable, or when you need deterministic task structure for auditing. Drawback: rigid plans break on unexpected tool outputs -- build in a replanning trigger when step failure rate exceeds 2 consecutive failures. - **Reflexion:** After each attempt, the agent reflects on its output using a self-critique loop, generates a verbal reflection, stores it in an episodic memory buffer, and retries. Use when first-pass quality is consistently poor but retrying with feedback improves results. Add a maximum reflection depth of 3 to prevent unbounded cost escalation. - **LATS (Language Agent Tree Search):** Monte Carlo Tree Search applied to agent trajectories. The agent explores multiple parallel action branches, scores each branch with a value function (often another LLM call), and selects the best path. Use only for high-value, compute-tolerant tasks -- this is 5-20x more expensive than ReAct. Appropriate for code generation competitions, complex research synthesis, or high-stakes decision support. - **Orchestrator/Worker (Multi-Agent):** A central orchestrator LLM receives the task, decomposes it, routes subtasks to specialized worker agents, and aggregates results. Use when different subtasks require different tool sets or system prompts. Each worker should have a single, narrowly defined responsibility. - **Critic/Generator (Multi-Agent):** A generator agent produces a draft, a critic agent evaluates and annotates it, and the generator revises based on critique. Use for content generation, code review, or any task requiring quality evaluation. Limit to 3 critique rounds to control cost. ### 3. Design the Memory Architecture Memory is the most underengineered component in most agent systems: - **Working memory (context window):** The agent's active scratchpad. Manage it explicitly -- implement a context compression strategy when the context exceeds 70% of the model's context window limit. Summarize completed tool call chains rather than keeping raw outputs. Use structured formats (JSON or XML) for tool results to reduce token consumption by 20-40% compared to prose. - **Episodic memory (conversation/session history):** Store prior agent runs with their inputs, outputs, and intermediate steps in a structured database (PostgreSQL with JSONB or MongoDB). Use this for Reflexion patterns and for debugging. Implement a rolling window of the 5-10 most relevant prior episodes retrieved via embedding similarity. - **Semantic memory (vector store):** Long-term factual knowledge. Use a vector database (Pinecone, Qdrant, Weaviate, or pgvector) for domain knowledge retrieval. Chunk documents at 256-512 tokens with 10-15% overlap. Use embedding models appropriate to the domain -- text-embedding-3-large for general English, domain-specific models for code or scientific text. - **Procedural memory (tool registry):** The catalog of available tools with their descriptions, parameter schemas, and examples. Tool descriptions are part of the system prompt and directly influence tool selection accuracy. Write tool descriptions as imperative sentences that specify what the tool does, what inputs it requires, and when to use it versus alternatives. - **Cache layer:** Cache deterministic tool outputs (search results, database queries) with TTLs appropriate to data freshness requirements. A Redis cache with a 15-minute TTL on search results can reduce repeated search costs by 60-80% in research agents. ### 4. Design the Tool Interface and Safety Layer Tools are the most dangerous part of any agent system: - **Tool schema design:** Every tool must have a JSON Schema definition with required/optional fields, type constraints, and description fields. Use strict schema validation -- reject malformed tool calls before execution rather than letting the tool fail with a cryptic error. - **Tool execution sandbox:** Code execution tools must run in isolated environments -- Docker containers with no network access, read-only filesystem mounts, and CPU/memory limits (e.g., 2 CPU cores, 512MB RAM, 30-second timeout). Never execute LLM-generated code in the host environment. - **Idempotency requirements:** Mark tools as idempotent (read-only: search, query, calculate) or non-idempotent (write: send email, update database, deploy code). Non-idempotent tools require explicit confirmation gates at Level 1-3 autonomy. Log every non-idempotent tool call with full parameters and timestamps. - **Tool failure handling:** Each tool must return a structured response with a success flag, result payload, and error message. The agent's error handling loop should: (1) retry transient failures up to 3 times with exponential backoff, (2) attempt an alternative tool if available, (3) ask the user for guidance if all alternatives are exhausted, and (4) hard-stop if a safety constraint is violated. - **Rate limiting and cost controls:** Implement per-agent-run hard limits: maximum tool calls (e.g., 25 per run), maximum LLM tokens consumed (e.g., 100k tokens per run), maximum wall clock time (e.g., 5 minutes for synchronous, 60 minutes for async). These prevent runaway costs from looping agents. - **Dangerous action classification:** Maintain an explicit list of high-risk actions (deleting data, sending external communications, making financial transactions, modifying infrastructure) that require a human approval step regardless of autonomy level. Never make this list implicit. ### 5. Implement the Agent Loop with Observability The agent execution loop is the core runtime -- build observability in from line one: - **Structured trace logging:** Every agent run gets a unique run_id (UUID). Log every LLM call, every tool invocation, every tool result, and every reasoning step as a structured JSON event with: run_id, step_number, timestamp, event_type, model_name, input_tokens, output_tokens, latency_ms, cost_usd, and the full input/output payload (with PII redacted per your data handling policy). - **Step-level span tracing:** Use OpenTelemetry or a framework like LangSmith, Langfuse, or Helicone to create parent/child spans for the full agent run and each step. This enables waterfall visualization of multi-step runs and identification of slow tool calls. - **Intermediate state persistence:** Checkpoint agent state after each step to a durable store (Redis with persistence or a database). If the agent process crashes mid-run, resumption from the last checkpoint prevents duplicated work and non-idempotent tool re-execution. - **Streaming intermediate output:** For synchronous agents where a user is waiting, stream the reasoning steps in real time -- even if the final answer is not ready. Users tolerate 30-60 second waits far better when they see progress. Use server-sent events or WebSocket for streaming. - **Quality signal collection:** After each run, collect: task completion success (binary), user satisfaction rating (1-5 if user-facing), and tool call efficiency (percentage of tool calls that contributed to the final answer). These signals feed the evaluation loop. ### 6. Design the Guardrails and Safety System Production agents require explicit safety architecture, not bolt-on filters: - **Input guardrails:** Before the agent starts, validate the input against: prompt injection patterns (detect instructions to ignore system prompt or override tool permissions), scope validation (is this task within the agent's intended domain?), and PII/sensitive data detection. Use a fast, cheap classifier (a fine-tuned BERT-class model or rule-based patterns) for latency-sensitive guardrails -- do not use an LLM for input validation. - **Output guardrails:** Before returning results to the user or executing non-idempotent actions, validate outputs for: factual grounding (does the output reference sources from the context?), toxicity and policy violations, and schema conformance for structured outputs. Use structured output validation (Pydantic or Zod) for any downstream system that consumes agent output. - **Behavioral constraints in system prompt:** Encode behavioral constraints as explicit rules in the system prompt, not as vague instructions. Bad: "Be helpful and safe." Good: "You MUST NOT call the delete_record tool unless the user's request explicitly contains the word 'delete' and you have confirmed the specific record ID with the user. You MUST NOT send any external communications (email, Slack, SMS) without displaying the full message to the user and receiving explicit confirmation." - **Agent identity and scope limitation:** Each agent should have a narrowly scoped system prompt that defines its role, its available tools, and the boundaries of its authority. A research agent that also has access to email tools is a security risk -- separate the scopes. ### 7. Evaluate, Test, and Iterate Agent evaluation is fundamentally different from model evaluation: - **Task-level evaluation harness:** Create a test suite of 20-50 representative tasks with known correct solutions or evaluation rubrics. Run the full agent loop on each task. Measure: task completion rate, tool call accuracy (did it use the right tools?), step efficiency (steps used / minimum steps needed), and answer quality (scored by a judge LLM or human rater). - **Adversarial testing:** Test the agent against prompt injection attempts, tool call parameter boundary violations, and task inputs designed to trigger looping behavior. At least 20% of test cases should be adversarial or edge cases. - **Regression testing:** Run the full evaluation suite on every agent system prompt change, tool description change, or underlying model update. LLM behavior is sensitive to small prompt changes -- a rephrasing of a tool description can change tool selection accuracy by 10-30%. - **A/B evaluation for pattern changes:** When switching control flow patterns (e.g., from ReAct to Plan-and-Execute), run both patterns on the same task set and compare completion rate, cost per task, and latency. Do not switch patterns based on intuition alone. - **Cost and latency profiling:** After evaluation, compute cost per successful task completion (total LLM cost + tool API costs / successful completions). This is the key operational metric for production viability. A task that costs $0.50 to complete with 80% success rate is often worse than a task that costs $0.10 to complete with 70% success rate. --- ## Output Format ### Agent Architecture Decision Record ``` # Agent Architecture Decision Record (AADR) # Task: [One-sentence description of what the agent must accomplish] # Date: [ISO 8601] # Status: [Draft | Approved | Superseded] ## Task Classification - Task Type: [research | execution | decision | conversational] - Autonomy Level: [1-5, with justification] - Latency Class: [synchronous <10s | async <60min | batch <24h] - Estimated Steps: [expected tool call round trips per run] - Estimated Cost: [$X.XX per run at expected step count] ## Control Flow Pattern - Pattern Selected: [ReAct | Plan-and-Execute | Reflexion | LATS | Orchestrator-Worker | Critic-Generator] - Rationale: [2-3 sentences: why this pattern over alternatives] - Step Limit: [hard maximum iterations before forced termination] - Replanning Trigger: [conditions that cause plan revision] ## Memory Architecture - Working Memory: [context window budget, compression strategy, format] - Episodic Memory: [storage system, retention policy, retrieval strategy] - Semantic Memory: [vector store, embedding model, chunk size, retrieval k] - Cache: [technology, TTL, invalidation strategy] ## Tool Inventory | Tool Name | Type | Idempotent | Rate Limit | Timeout | |---------------------|---------------|------------|------------------|---------| | [tool_name] | [read/write] | [yes/no] | [X calls/min] | [Xs] | ## Safety and Guardrails - Input Guardrails: [specific checks: prompt injection, scope, PII] - Output Guardrails: [specific checks: grounding, toxicity, schema] - Dangerous Actions: [list of tool calls requiring human confirmation] - Hard Limits: [max_steps, max_tokens, max_cost, max_wall_time] ## Observability - Trace Platform: [LangSmith | Langfuse | Helicone | OpenTelemetry] - Log Store: [destination for structured trace logs] - Metrics: [completion_rate, cost_per_run, latency_p50/p99, step_efficiency] - Alerting: [conditions for alerts: failure rate >X%, cost spike] ``` ### Agent Implementation Template ```python # agent_core.py from __future__ import annotations import asyncio import uuid import time from dataclasses import dataclass, field from typing import Any, Callable, Optional from enum import Enum class StepType(Enum): THOUGHT = "thought" TOOL_CALL = "tool_call" TOOL_RESULT = "tool_result" FINAL_ANSWER = "final_answer" ERROR = "error" @dataclass class AgentStep: step_number: int step_type: StepType content: Any tool_name: Optional[str] = None tool_input: Optional[dict] = None input_tokens: int = 0 output_tokens: int = 0 latency_ms: float = 0.0 cost_usd: float = 0.0 @dataclass class AgentConfig: max_steps: int = 20 # Hard iteration limit max_tokens_per_run: int = 80_000 max_cost_per_run_usd: float = 1.00 max_wall_time_seconds: float = 300.0 reflection_depth: int = 3 # For Reflexion pattern require_confirmation_for: list[str] = field(default_factory=list) # Non-idempotent tools @dataclass class AgentRun: run_id: str = field(default_factory=lambda: str(uuid.uuid4())) task: str = "" steps: list[AgentStep] = field(default_factory=list) total_cost_usd: float = 0.0 total_tokens: int = 0 success: bool = False final_answer: Optional[str] = None termination_reason: str = "" # "success" | "step_limit" | "cost_limit" | "error" class AgentExecutor: """ ReAct-pattern agent executor with hard safety limits, structured tracing, and pluggable tool registry. Designed for production use. Key design decisions: - Hard limits on steps, cost, and wall time prevent runaway execution - Every step is logged as a structured event for observability - Tool calls are validated against JSON Schema before execution - Non-idempotent tools require explicit confirmation if configured """ def __init__( self, llm_client, # LLM client (OpenAI, Anthropic, etc.) tools: dict[str, Callable], tool_schemas: dict[str, dict], # JSON Schema for each tool system_prompt: str, config: AgentConfig, tracer=None, # OpenTelemetry tracer or LangSmith client ): self.llm = llm_client self.tools = tools self.tool_schemas = tool_schemas self.system_prompt = system_prompt self.config = config self.tracer = tracer async def run(self, task: str) -> AgentRun: agent_run = AgentRun(task=task) start_time = time.monotonic() messages = [ {"role": "system", "content": self.system_prompt}, {"role": "user", "content": task}, ] for step_num in range(self.config.max_steps): # Enforce wall time limit elapsed = time.monotonic() - start_time if elapsed > self.config.max_wall_time_seconds: agent_run.termination_reason = "wall_time_limit" break # Enforce cost limit if agent_run.total_cost_usd >= self.config.max_cost_per_run_usd: agent_run.termination_reason = "cost_limit" break # LLM call to get next action step_start = time.monotonic() response = await self.llm.chat( messages=messages, tools=list(self.tool_schemas.values()), ) step_latency = (time.monotonic() - step_start) * 1000 # Parse response: tool call or final answer if response.tool_calls: for tool_call in response.tool_calls: step = await self._execute_tool_call( tool_call=tool_call, step_number=step_num, latency_ms=step_latency, ) agent_run.steps.append(step) agent_run.total_cost_usd += step.cost_usd agent_run.total_tokens += step.input_tokens + step.output_tokens # Append tool result to message history messages.append({ "role": "tool", "tool_call_id": tool_call.id, "content": str(step.content), }) else: # No tool call -- this is the final answer agent_run.final_answer = response.content agent_run.success = True agent_run.termination_reason = "success" break self._emit_trace(agent_run) return agent_run async def _execute_tool_call( self, tool_call, step_number: int, latency_ms: float ) -> AgentStep: tool_name = tool_call.function.name tool_input = tool_call.function.arguments # dict after JSON parse # Validate tool exists if tool_name not in self.tools: return AgentStep( step_number=step_number, step_type=StepType.ERROR, content=f"Tool '{tool_name}' not found in registry.", tool_name=tool_name, latency_ms=latency_ms, ) # Execute tool with timeout try: tool_fn = self.tools[tool_name] result = await asyncio.wait_for(
View on GitHub
This SKILL.md is very large, so SkillsMP previews the first section here. View on GitHub