- name
- production-fast-api-template
- description
- Structure and execution rules for a unified, production-ready backend combining domain-oriented FastAPI, Retrieval-Augmented Generation (RAG) and Agentic AI in one codebase on AWS ECS Fargate. Use when scaffolding a new backend; adding a module, API domain, RAG component, agent, tool, workflow, LLM/provider integration or test; deciding where a file belongs; writing async/sync routes, services or clients; offloading blocking or CPU-heavy work (thread pools, process pools, AWS Lambda); designing background jobs, queues (SQS, Celery), workers and scaling; or writing Pydantic schemas, settings (pydantic-settings), serialization, structured LLM outputs, citation validation, agent plans or tool-call arguments; or writing FastAPI dependencies (request validation, auth, access scopes, chaining), REST paths, response models, OpenAPI docs, error responses, database naming and SQL queries, migrations, API tests (async client, dependency overrides) or lint/format setup (ruff); or adding content-safety guardrails (input/output checks, PII, prompt injection, toxicity, Bedrock Guardrails, AgentCore Policy/Gateway Cedar policies, Guardrails AI validators), fail-closed handling and guardrail evaluation; or evaluating RAG, retrieval stages, structured outputs, citations, agents, multi-agent workflows or conversations (DeepEval, Ragas, LLM-as-judge, G-Eval, golden datasets, regression tiers, thresholds, judge calibration); or securing the API (authentication, JWT/JWKS, OAuth 2.0/OIDC, RBAC/ABAC, BOLA/BFLA, tenant isolation, input/upload/SSRF validation, security headers, CORS, CSRF, rate limiting, audit logging, ECS/IAM/secrets hardening, CI security scanning, security tests) or LLM/agent security (OWASP LLM Top 10, OWASP Agentic Top 10, prompt injection, tool authorization, human approval, excessive agency).
# Production FastAPI Template: Rules
These rules define **where code lives** and **how it executes** in a unified FastAPI + RAG +
Agentic AI backend. Apply them whenever you generate or modify code in a project built from
this template.
**Companion references.** Read the relevant file before writing code:
| File | Read before writing… | Summarized in |
|---|---|---|
| [ASYNC_EXECUTION.md](ASYNC_EXECUTION.md) | Any route, service, client, agent tool, worker or background job | Section 5 |
| [PYDANTIC_STANDARDS.md](PYDANTIC_STANDARDS.md) | Any schema, settings class, serialization, LLM call with structured output, citation check, agent plan or tool argument model | Section 6 |
| [API_CONVENTIONS.md](API_CONVENTIONS.md) | Any dependency, route signature, auth/permission check, OpenAPI metadata, DB model or query, migration, API test or lint setup | Section 7 |
| [GUARDRAILS.md](GUARDRAILS.md) | Any user-facing LLM call, RAG answer, ingestion of untrusted documents, agent tool execution, gateway target or content-safety check | Section 8 |
| [AI_EVALUATION.md](AI_EVALUATION.md) | Any evaluation contract, metric, dataset, judge, runner, report or threshold; any new RAG stage, agent, tool or conversation feature that needs quality coverage | Section 9 |
| [API_SECURITY.md](API_SECURITY.md) | Any authentication, token, permission, tenant or ownership check; input/upload/outbound-URL handling; response headers, CORS, errors; rate limits; audit events; ECS/IAM/secrets; CI security; security tests | Section 10 |
| [AGENT_SECURITY.md](AGENT_SECURITY.md) | Any agent, tool, tool registry entry, approval flow, memory write, MCP connector or LLM-output sink | Section 10 |
**Status labels:**
- 🟢 **Convention.** A proposed rule. Follow it by default.
- 🟡 **Needs approval.** An open team decision. Surface the options to the user and never pick one silently.
These rules don't choose agent frameworks, LLM providers, models, vector stores or a queue
implementation. If a rule doesn't fit a real need, raise it with the user instead of silently
working around it.
---
## 1. Core structure rules 🟢
1. **One application.** FastAPI, RAG and Agentic AI are **not** separate apps. There is exactly one `src/` package, one FastAPI entry point (`src/main.py`) and one shared infrastructure layer.
2. **Never create a second entry point.** Don't add a second `src/`, a separate app folder (`rag_app/`, `agents_service/`, `worker_app/`) or an extra `FastAPI()` instance.
3. **Background workers and Lambda functions are extra entry points into the same codebase.** They are not separate applications, and they call the same service functions the API calls.
4. **One central API router.** All routers mount through `src/api/v1/router.py`. Domain, RAG and agent routers never mount themselves on the app.
5. **Domain-oriented modules.** Business code lives inside its domain package. `src/`'s root holds only application-wide concerns.
6. **No duplicated shared layers.** Model clients go in `src/llm/`, content-safety checks in `src/guardrails/` and external-service clients in `src/providers/`, each exactly once. RAG and agents **consume** them and never define their own copies (no `rag/llm_client.py`, no `agents/vector_db.py`).
7. **Provider-neutral naming.** Name files by responsibility, not vendor (`llm/embeddings.py`, not `openai_embeddings.py`). Vendor specifics stay behind `src/providers/` or `src/llm/`.
8. **`common/` is for genuinely reusable, domain-agnostic code only.** Never put business logic there.
9. **Keep agents, reasoning and execution apart.** Agent *definitions* go in `agents/`, *thinking* in `cognition/` and *doing, scheduling and offloading* in `execution/`.
10. **Tests mirror `src/`.** Quality evaluation (datasets, DeepEval/Ragas suites, deterministic metrics, runners, reports) lives in the top-level `evaluation/` directory, separate from `tests/`. `src/evaluation/` holds only framework-neutral contracts, the metric registry and adapters, with lazy framework imports (AI_EVALUATION.md §1).
11. **Every Python directory under `src/` has an `__init__.py`.** Non-Python directories that must survive in Git (prompt folders, data, empty test folders) get a `.gitkeep`.
12. **Explicit module imports across packages**, e.g. `from src.auth import constants as auth_constants`.
13. **Reuse existing files before adding new ones.** Extend `execution/background_worker.py` rather than creating `execution/worker2.py` or `rag/worker.py`.
---
## 2. Canonical structure
```text
<project-name>/
├── src/
│ ├── __init__.py
│ ├── main.py # the single FastAPI entry point (lifespan creates shared clients/limiters)
│ ├── config.py # application-wide configuration
│ ├── constants.py
│ ├── exceptions.py # base domain exception + handlers rendering ErrorResponse
│ ├── middleware.py
│ ├── database.py # shared metadata (naming convention), engine/session factory, get_db_session
│ ├── models.py # genuinely shared models only
│ ├── pagination.py
│ │
│ ├── api/
│ │ └── v1/
│ │ └── router.py # aggregates domain, RAG and agent routers
│ │
│ ├── auth/ # reference domain module
│ │ ├── router.py
│ │ ├── schemas.py # the domain's own Pydantic API/internal schemas
│ │ ├── models.py # persistence (ORM) models, not Pydantic
│ │ ├── dependencies.py # current_principal, require_<permission>, scope dependencies
│ │ ├── config.py # the domain's own pydantic-settings class (AUTH_ prefix: issuers, audiences, JWKS)
│ │ ├── constants.py # permission and scope names
│ │ ├── exceptions.py
│ │ ├── service.py
│ │ ├── tokens.py # JWT/JWKS validation, token types, claims → Principal
│ │ ├── permissions.py # RBAC/ABAC policy evaluation (default deny)
│ │ ├── security.py # password hashing, only if local accounts are approved 🟡
│ │ ├── audit.py # auth audit events (login, token rejected, session revoked)
│ │ └── utils.py
│ │
│ ├── example_domain/ # copy this layout for every new business domain
│ │ ├── router.py
│ │ ├── schemas.py
│ │ ├── models.py
│ │ ├── dependencies.py # valid_<entity>_id and other I/O-backed validation
│ │ ├── constants.py
│ │ ├── exceptions.py
│ │ ├── service.py
│ │ └── utils.py
│ │
│ ├── rag/
│ │ ├── routes.py # "routes", not "router": avoids clashing with rag/router/
│ │ ├── schemas.py # RAG API schemas + cross-stage contracts (RetrievedEvidence)
│ │ ├── service.py # container/lifecycle of long-lived RAG services
│ │ ├── dependencies.py
│ │ ├── config.py # RAG settings (RAG_ prefix)
│ │ ├── constants.py
│ │ ├── exceptions.py
│ │ ├── extraction.py # structure-preserving document extraction
│ │ ├── ingestion.py # coordinates processing and indexing (run by background jobs)
│ │ ├── semantic/
│ │ │ ├── metadata.py
│ │ │ ├── segmentation.py
│ │ │ └── chunks.py
│ │ ├── router/
│ │ │ ├── classifier.py
│ │ │ ├── fusion.py
│ │ │ └── evaluation.py
│ │ ├── preprocessing/
│ │ │ ├── query_expansion.py
│ │ │ └── hyde.py
│ │ ├── retrieval/
│ │ │ ├── semantic_search.py
│ │ │ ├── bm25.py
│ │ │ ├── rrf.py
│ │ │ ├── mmr.py
│ │ │ ├── evidence_filter.py
│ │ │ └── hybrid_retriever.py
│ │ ├── generation/
│ │ │ ├── context_builder.py
│ │ │ ├── citations.py
│ │ │ ├── streaming.py
│ │ │ ├── schemas.py # LLM-output contracts: GeneratedAnswer, Citation
│ │ │ ├── validation.py # application-level answer/citation validation
│ │ │ └── service.py
│ │ ├── chat/
│ │ │ ├── routes.py
│ │ │ ├── schemas.py
│ │ │ ├── service.py
│ │ │ └── storage.py
│ │ └── security/
│ │ ├── access_control.py # AccessScope model + vector-store/DB filter translation
│ │ ├── document_validation.py # ingestion file checks (type, signature, size, archives)
│ │ └── retrieval_policy.py # post-retrieval ACL re-check, trust labels, provenance
│ │
│ ├── agents/
│ │ ├── router.py
│ │ ├── schemas.py # AgentTask, ToolResult, AgentExecutionState
│ │ ├── structured_output.py # LLM-produced: ExecutionPlan, ToolCall union, FinalAgentResponse
│ │ ├── dependencies.py
│ │ ├── config.py # agent settings (AGENTS_ prefix)
│ │ ├── constants.py
│ │ ├── exceptions.py
│ │ ├── service.py
│ │ ├── base_agent.py
│ │ ├── autonomous_agent.py
│ │ ├── planner_agent.py
│ │ ├── agent_interface.py
│ │ ├── team_orchestrator.py
│ │ ├── step_handler.py
│ │ ├── task_manager.py
│ │ ├── tools/
│ │ │ ├── calculator.py
│ │ │ ├── file_manager.py
│ │ │ ├── search_tool.py
│ │ │ └── tool_registry.py
│ │ ├── workflows/
│ │ │ ├── code_review_chain.py
│ │ │ ├── research_chain.py
│ │ │ ├── multi_agent_workflow.yaml
│ │ │ └── workflow_executor.py
│ │ └── security/ # policies; enforced by execution/executor.py
│ │ ├── tool_permissions.py # default-deny tool authorization (user perms ∩ agent profile)
│ │ ├── execution_policy.py # step, tool-call, token, cost and time budgets
│ │ └── approval_policy.py # human approval bound to argument hash
│ │
│ ├── cognition/
│ │ ├── cognitive_loop.py
│ │ ├── decision_policy.py
│ │ ├── planner.py
│ │ ├── reasoner.py
│ │ ├── state_interpreter.py
│ │ └── memory/
│ │ ├── long_term_memory.py
│ │ ├── short_term_memory.py
│ │ └── memory_manager.py
│ │
│ ├── execution/
│ │ ├── action_resolver.py
│ │ ├── controller.py
│ │ ├── error_handler.py
│ │ ├── executor.py
│ │ ├── threadpool.py # thread-offload helpers and limiters (no custom executor yet)
│ │ ├── job_scheduler.py # enqueue/schedule background jobs
│ │ └── background_worker.py # queue consumer / job dispatch loop
│ │
│ ├── llm/ # shared by RAG and agents
│ │ ├── client.py
│ │ ├── embeddings.py
│ │ ├── reranker.py
│ │ ├── model_loader.py
│ │ ├── cache.py
│ │ ├── dependencies.py # accessors for lifespan-created model clients
│ │ ├── config.py # LLM settings (LLM_ prefix, SecretStr keys)
│ │ ├── schemas.py # structured-output modes, capabilities, failure kinds
│ │ ├── structured_output.py # generate → validate → bounded retry (provider-independent)
│ │ └── prompts/
│ │ ├── system/
│ │ ├── tasks/
│ │ └── templates/
│ │
│ ├── guardrails/ # shared content-safety layer, used by RAG, agents and workers
│ │ ├── schemas.py # GuardrailPhase, GuardrailOutcome, GuardrailFinding, GuardrailVerdict
│ │ ├── service.py # run a phase's checks, decide outcome, enforce / log-only, fail closed
│ │ ├── policies.py # which checks apply to which surface and phase
│ │ ├── config.py # GUARDRAILS_ settings (mode, fail mode, thresholds, policy version)
│ │ ├── dependencies.py
│ │ ├── exceptions.py # GuardrailBlocked, GuardrailUnavailable
│ │ └── validators/ # deterministic in-house checks (regex, lists, length)
│ │
│ ├── providers/ # external-service boundaries (MCP, vector DB, queues, Lambda, guardrails, …)
│ │ ├── guardrails/client.py # managed guardrail API / validator-library adapters
│ │ ├── mcp/client.py
│ │ ├── vector_store/client.py
│ │ └── external/client.py
│ │
│ ├── evaluation/ # runtime-safe evaluation contracts only (no deepeval/ragas at import)
│ │ ├── schemas.py # EvaluationSample, RetrievedContext, AgentTrajectory, EvaluationResult
│ │ ├── registry.py # metric key → framework, class, version, required fields, result kind
│ │ └── adapters/
│ │ ├── deepeval_adapter.py # contracts ↔ LLMTestCase / ConversationalTestCase (lazy import)
│ │ └── ragas_adapter.py # contracts ↔ Ragas collections inputs / samples (lazy import)
│ │
│ └── common/
│ ├── schemas/
│ │ ├── base.py # CustomModel, UtcDatetime, shared annotated types
│ │ ├── responses.py # ErrorResponse and shared response shapes
│ │ └── pagination.py # Page[T], PageParams (models only; logic in src/pagination.py)
│ ├── dependencies.py
│ ├── logging.py
│ ├── retry.py
│ ├── timers.py
│ ├── serialization.py
│ ├── validation.py
│ └── security/
│ ├── headers.py # security-headers ASGI middleware (registered in src/middleware.py)
│ ├── rate_limiting.py # distributed limiter interface + key builders (backend 🟡)
│ ├── request_validation.py # body size/depth, content type, filenames, outbound URL (SSRF) policy
│ └── audit_schemas.py # AuditEvent, AuditActor, AuditTarget
│
├── tests/
│ ├── conftest.py # async client (httpx + ASGITransport), dependency-override fixtures
│ ├── auth/
│ ├── rag/{ingestion,retrieval,generation}/
│ ├── agents/
│ ├── cognition/
│ ├── execution/
│ ├── guardrails/ # fake checkers; deny / suppress / fail-closed / log-only paths
│ ├── evaluation/ # test_schemas.py, test_adapters.py, test_deterministic.py (no LLM calls)
│ ├── security/
│ │ ├── authentication/ # JWT/JWKS validation, throttling
│ │ ├── authorization/ # BOLA, BFLA, BOPLA, cross-tenant, service-to-service
│ │ ├── api/ # headers, CORS, limits, uploads, SSRF, injection, CSRF, rate limits
│ │ ├── rag/ # unauthorized/cross-tenant retrieval, ACL revocation, citation integrity
│ │ └── agents/ # tool authorization, argument BOLA, approval bypass, budgets, injection
│ ├── integration/
│ ├── e2e/
│ └── fixtures/
├── evaluation/ # offline quality evaluation (never imported by src/)
│ ├── README.md
│ ├── config/{metrics.yaml,thresholds.yaml}
│ ├── datasets/{rag,agents,conversations,golden,guardrails}/
│ ├── deepeval/{rag,agents,conversations,custom}/
│ ├── ragas/{rag,agents,custom}/
│ ├── deterministic/{citations,retrieval,structured_output,tool_calls}.py
│ ├── tracing/{collectors,normalization}.py
│ ├── runners/{offline,regression,report}.py
│ ├── fixtures/
│ └── reports/ # generated, git-ignored 🟡
├── data/{documents,knowledge_bases,agent_state,fixtures}/
├── scripts/
├── docs/
│ ├── architecture/
│ ├── api/
│ ├── development/
│ ├── decisions/
│ └── standards/
│ ├── ASYNC_EXECUTION.md # copied from this skill
│ ├── PYDANTIC_STANDARDS.md # copied from this skill
│ ├── API_CONVENTIONS.md # copied from this skill
│ ├── GUARDRAILS.md # copied from this skill
│ ├── AI_EVALUATION.md # copied from this skill
│ ├── API_SECURITY.md # copied from this skill
│ └── AGENT_SECURITY.md # copied from this skill
├── .env.example
├── .gitignore
├── pyproject.toml
├── README.md
└── CLAUDE.md
```
`__init__.py` files are omitted above for brevity. Rule 11 still applies.
---
## 3. Where does new code go?
| You are adding… | Put it in | Not in |
|---|---|---|
| A new business capability (orders, billing…) | `src/<domain>/` using the `example_domain/` layout | `src/` root, `common/` |
| An HTTP endpoint | The owning module's `router.py` (`routes.py` inside `rag/`), mounted via `api/v1/router.py` | `main.py` |
| Request/response models | `<module>/schemas.py`, inheriting `CustomModel` | `src/models.py`, direct `BaseModel` |
| Shared base model, `UtcDatetime`, shared annotated types | `common/schemas/base.py` | Per-domain base classes |
| Error, message and page response shapes | `common/schemas/{responses,pagination}.py` | Each domain redefining them |
| Settings for a domain or component | `<module>/config.py` (own prefix, cached getter) | The global `src/config.py` |
| LLM-output schema for answer generation | `rag/generation/schemas.py` | `rag/schemas.py` (API) |
| LLM-output schema for one RAG stage (metadata, classification, expansion) | `rag/<stage>/schemas.py`, created when the stage is implemented | API schemas |
| Citation and answer checks against evidence | `rag/generation/validation.py` | Pydantic validators |
| Agent plan, tool-call union, final agent response | `agents/structured_output.py` | `agents/schemas.py` |
| A tool's argument model | That tool's module in `agents/tools/` | A central args file |
| Structured-output generation, validation and retry flow | `llm/structured_output.py` + `llm/schemas.py` | Re-implemented in RAG or agents |
| Evaluation contracts (`EvaluationSample`, `EvaluationResult`, …), metric registry | `src/evaluation/{schemas,registry}.py` | `evaluation/`, per-suite copies |
| Contract ↔ DeepEval / Ragas mapping | `src/evaluation/adapters/` (lazy framework imports) | Suites building test cases ad hoc |
| Metric suites, G-Eval/DAG/custom metrics, judge wrappers | `evaluation/{deepeval,ragas}/` | `src/`, `tests/` |
| Deterministic metrics (retrieval P/R/MRR/NDCG, citations, schema, tool calls) | `evaluation/deterministic/` | LLM judges |
| Evaluation datasets, metric/threshold config, runners, reports | `evaluation/{datasets,config,runners,reports}/` | `data/`, `tests/fixtures/` |
| Tests of the evaluation code | `tests/evaluation/` | `evaluation/` |
| DB models used by one domain | `<domain>/models.py` | `src/models.py` |
| Existence, ownership or uniqueness checks that need I/O | `<domain>/dependencies.py` (`valid_<entity>_id`) | Pydantic validators, repeated in each route |
| `current_principal`, `require_<permission>`, scope dependencies | `auth/dependencies.py` | Each domain parsing tokens |
| JWT/JWKS validation, claims → `Principal` | `auth/tokens.py` | `auth/service.py`, routes |
| RBAC/ABAC policy evaluation | `auth/permissions.py` | Routes, prompts |
| Password hashing (only if local accounts are approved 🟡) | `auth/security.py` | `common/` |
| Auth audit events; shared `AuditEvent` schema | `auth/audit.py`; `common/security/audit_schemas.py` | Free-form log lines |
| Security headers, rate limiting, request/upload/outbound-URL validation | `common/security/{headers,rate_limiting,request_validation}.py` | Per-domain copies |
View on GitHub