| name | agentic-ai-controls |
| description | Reviews and proposes controls for an agentic AI use case (an AI system that selects actions, calls tools, reads or writes systems of record, or operates with autonomy beyond single-turn generation) in a regulated financial-services firm. Output names the agent's scope and authority statement, the architecture and tool inventory with permissions and blast radius, the identity and authorisation posture, the human oversight points with effective oversight evidence, the prompt-injection-via-tool-output and tool-misuse threat assessment, the kill-switch and rollback procedures with drill evidence, the logging and audit posture, the incident response with regulator-notification triggers, the residual risk with accepted owners, and the recommended owner actions. Designed for second-line review of agents that book actions in real systems, not chat-only assistants.
Best for:
- A first-line owner has built or is building an agentic system that calls tools, reads from systems of record, or executes actions, and second-line needs control coverage before pre-prod or expansion.
- A previously deployed assistant is being upgraded with tool-use, retrieval, or multi-step planning, and the controls and gates need to be revisited.
- An enterprise-wide review of agentic deployments needs a consistent control inventory across sponsoring units (e.g., a CISO-led portfolio sweep, an AI risk committee inventory refresh, an exam-readiness sprint).
- An insurance carrier is deploying or has deployed an agentic system in claims handling, FNOL intake, agent / producer co-pilot, or policyholder-facing actuation; the insurance overlay carries the absorbed claims-AI control bar (NAIC #880, state UCSPA, NAIC AIS Program, HIPAA where PHI flows).
Not the right tool when:
- The system is a retrieval-augmented chat assistant with no tool calls and no actuation (use rag-evaluation-review and the model card; the agentic-controls overhead is wrong-shape for a no-actuation system).
- The work is the use-case intake or the risk tier itself (use ai-use-case-intake or ai-risk-tiering; this skill consumes their outputs).
- The system is a vendor-hosted agentic product still in procurement (use vendor-diligence in third-party-operational-resilience with the AI-vendor scenario mode; this skill assumes a deployment review on a system the firm is operating).
- The work is the prompt-injection-only deep-dive (use prompt-injection-risk; this skill folds prompt-injection into the broader agentic-control assessment but the focused memo lives there).
- The work is the firm-side governance card or the validation plan (use model-card-builder or validation-plan; this skill consumes them and produces the controls-focused artifact).
|
| argument-hint | [use-case ID, model-card record, validation-plan record, prompt-injection review, runbook references, IR drill records, or scope statement] |
Agentic AI controls review
An agentic-controls review is the second-line surfacing of one specific class of AI deployment: a system that selects actions, calls tools, reads or writes systems of record, or chains decisions across turns with autonomy beyond single-turn generation. The artifact names the scope and authority the agent operates within, the architecture and the tool inventory with blast-radius framing, the identity and authorisation posture, the human-oversight points with evidence the oversight is effective (not theatrical), the prompt-injection-via-tool-output and tool-misuse threat assessment, the kill-switch and rollback procedures with drill evidence, the logging and audit posture, the incident response with named regulator-notification triggers, the residual risk with accepted owners, and the recommended owner actions.
Agentic systems sit across multiple regulator perspectives at once and the controls work has to name them. For most current deployments the binding hooks are the cyber-supervision regimes (NYDFS Part 500 for covered entities, plus the October 2024 AI cybersecurity industry letter; bank computer-security incident notification; SEC cybersecurity disclosure for public registrants; state Insurance Data Security Model Law) where tool-use and external-system actuation materially widen the cyber surface. The AI-specific framings carry the controls vocabulary practitioners actually use: OWASP LLM Top 10 (with the 2025 update on LLM06 excessive agency and LLM08 vector misuse), the OWASP Agentic threat catalogue, MITRE ATLAS techniques, NIST AI 600-1 GenAI Profile, NCSC/CISA secure-AI-development guidance, ISO/IEC 42001. The human-oversight bar comes from sector consumer-protection regimes (fair-claims-settlement standards in insurance, fair-lending in credit) and, for EU-in-scope deployments, the EU AI Act human-oversight expectations. The bank model-risk regime (the 2026 joint interagency revised guidance) explicitly excludes GenAI and agentic AI from its scope; where the firm's model-risk programme treats the agent as a model by analogy the principles inform, but the binding hook is firm policy, not the bulletin. Be explicit in the artefact about which scopes are in play.
The audience reads from four angles at once. The AI Governance Lead owns the artefact and consolidates the analysis. The MRMO and the model owner inherit the model-risk seams; the model owner is the source of upstream evidence (model card, validation plan, prompt-injection review, vendor system card) and the MRMO challenges the residual-risk view. The CISO function is a co-reviewer for any agent with tool use, identity, or external-system actuation; for most agentic deployments that is the default. Sector and conduct co-reviewers (for insurance: privacy officer where PHI flows, claims operations leadership for claims AI; for capital markets: surveillance and trading-controls leadership; for banking: BSA officer for AML-touching flows) sit on the residual-risk acceptance.
The artifact is a draft until the human reviewer attests. The skill stops short of filing or approving it.
Ask first
Most of what the review needs is on the table by the time someone reaches for this skill. A few things to settle before drafting:
- What is the agent's architecture and what does it actually do. Planner-executor split, single-agent vs multi-agent topology, foundation-model dependence, RAG, tool inventory, memory mode, autonomy level, multimodal posture ... the answer drives every section. If
model-card-builder has run, the system-description and GenAI-overlay sections of the card name this; consume the card. If prompt-injection-risk has run, the threat-surface and mitigations sections feed straight in.
- What is the agent's authority. Read-only vs read-and-write, customer-facing vs internal-only, financial impact per action, reversibility per action, blast radius per tool. The authority statement is the spine of the rest of the review; without explicit out-of-scope actions the boundary becomes operational drift inside three months of deployment.
- Where is the use case in lifecycle, and what triggered this review. Pre-prod-gate review consumes upstream artifacts and projects coverage gaps. In-prod-periodic review consumes monitoring evidence and drill records. Post-incident review consumes the incident timeline, the IR runbook outputs, and the scope-creep audit. Expansion review (new tool, new sponsor, new geography, new customer segment) consumes the prior review and the delta.
- Who is co-reviewing. CISO function for any agent with tool use, identity, or external-system actuation (which is most agents). Sector and conduct co-reviewers per the loaded overlays. The reviewer-attestation block names them as functions, never individuals.
- What sector and cross-cutting overlays load. Banking, insurance, capital markets, or payments-fintech as sector; cyber as the default cross-cutting overlay for tool-using agents; privacy where regulated personal data flows through the agent's prompt or tool path; conduct where the agent's actions touch customers (insurance claims AI is the canonical example here, and the insurance overlay carries the absorbed claims-AI control bar).
When the scope record is supplied, the skill consumes it for institution, persona, source posture, sector and cross-cutting overlays, lifecycle stage, and architecture flags. Otherwise it asks the practitioner the few facts it needs, and source posture sets what the artefact can assert at high confidence and what carries [evidence needed].
How the controls review gets filled in
The artifact has the same spine across architectures. The order below is the dependency chain a senior practitioner walks; sections without dependencies fill in as evidence arrives.
Tier and authority drive depth, so consume the model card and the scope record before deciding how heavy any section sits. Authority statement and tool inventory must precede threat assessment and oversight design, because the attack surface and the oversight points are functions of what the agent is allowed to do and what tools sit underneath. Mitigation evidence must be in hand before residual-risk framing, because residual risk is the gap between the threat surface and the controls that hold it down.
Review metadata names the reviewer role, the review stage (pre-prod-gate, in-prod-periodic, post-incident, expansion-review, exam-readiness), the date, and the upstream artifact IDs (scope, model card, validation plan, prompt-injection review). Reviewer roles are functions, never named individuals.
Agent metadata and use case reference lands the canonical agent ID and version, the sponsor function, the tier from ai-risk-tiering, the lifecycle stage, and pointers to the upstream intake and use-case records. A pointer to the model card's system description is preferable to restating; the agentic-controls artifact is not the place to re-derive architecture.
Scope and authority statement is the most load-bearing section. It enumerates the in-scope actions (concrete verbs against concrete systems, not "helps the user"), the out-of-scope actions (at least three concrete entries; this is the firewall against scope drift), the geographic and customer scope, and the financial-impact and reversibility framing per action. Reviewers probe the out-of-scope list hardest; descriptive prose ("the agent helps representatives serve customers") with no enumeration of what the agent may not do fails the section.
Architecture summary records the planner-executor split, the foundation-model provider and version pinning posture, the RAG corpora with retrieval scoping rules, the multi-agent topology if any, the memory mode, the modalities, and the autonomy level on the standard enum (advisory, advisory-with-suggested-action, human-in-the-loop, human-on-the-loop, agentic-with-tool-use, fully-autonomous). Autonomy level is the input that drives the oversight section.
Tool inventory captures, per tool: tool name, system of record it acts on, action type (read / write / execute / communicate), permission level granted to the agent, blast radius if misused (named in concrete terms: "could cancel up to N policies in a single batch", "could send up to M outbound emails per minute", "could route an order up to value V"), rate limit, and tool owner (function). A tool inventory without blast radius is a list, not a control surface; the blast-radius column is required on every entry. Tools the agent does not have are explicit "not applicable" rather than omitted; the omission would otherwise read as oversight.
Identity and authorisation records the agent's service account (separate from any human identity), the scoping principle (least-privilege per tool, per customer, per case, per session), secret management (where the agent's credentials live, who can rotate them, who can read them), rate limits at the identity layer, and session boundaries. For agents acting on regulated systems, the cyber overlay carries the named control-family mappings (NIST SP 800-53 access-control families, NYDFS Part 500 access expectations) used in the firm's cyber programme.
Human oversight points are the operational seam between the autonomy level and the actuation. Per consequential decision class: when a human must approve before action (block-and-wait), when a human is notified post-action with reversal authority (notify-and-revocable), when a human is sampled post-action for quality (post-hoc), when the agent acts autonomously without human touch (autonomous, allowed only for low-stakes reversible actions). Each entry names the reviewer role, the SLA, the sampling expectation, and the evidence the oversight is effective. Effective oversight evidence is what fails most reviews: reviewers click through, approval gates become rubber stamps, sampling is not audited. The entry needs an evidence pointer (gate-decision log, sampling audit, drift signal) that supports the assertion that the oversight catches what it is meant to catch. The EU AI Act human-oversight standard is explicit on this for in-scope deployments; the bar applies by analogy elsewhere.
Guardrails record the input filtering, output filtering, refusal behaviour, escalation triggers, and structured-output enforcement that sit between the agent's reasoning and its tool calls. Each entry has a type (input filter, output filter, refusal, escalation, structured output, content provenance), an owner, an evidence pointer, and a flag for whether the guardrail was exercised in any tested scenario.
Threat assessment enumerates the prompt-injection-via-tool-output and tool-misuse attack classes the agent's surface presents, and records what was tested. Carriers and trust-basis framing follow prompt-injection-risk; if that skill has run, consume its threat-surface and tested-attacks tables and add agentic-specific classes here. The agentic-specific attack classes (per the OWASP Agentic catalogue and the OWASP LLM 2025 entries on excessive agency and vector misuse): tool hijack (instructions in the agent's input or in tool output cause the agent to invoke a different tool than intended); cross-scope tool query (the agent invokes a tool with arguments outside its allowed scope); excessive-agency exploitation (the agent chains low-stakes actions into a high-stakes outcome the per-action guardrails do not catch); permission-escalation via planner-executor seam (the planner formulates a plan that the executor runs without re-checking the planner's authority); multi-agent message injection where applicable; memory poisoning where persistent memory applies; sandbox escape on systems with code execution. Each entry names attack class, the source taxonomy (typically MITRE ATLAS technique ID, OWASP LLM/Agentic identifier, NIST AI 600-1 subcategory), method, dataset reference, success criterion, result, and tester. Vendor red-team results are recorded explicitly as vendor-tested; firm-internal red-team coverage gaps land as [evidence needed] and route to recommended actions.
Sandbox and execution-environment guarantees record what the agent's execution sandbox enforces and what it does not. For systems with code execution, browser automation, or shell access: process isolation, filesystem scoping, network-egress controls, secret isolation, time-bounded execution, resource quotas, snapshot-and-rollback posture. For tool-use-only systems without execution: the section names the boundary explicitly ("no code-execution surface; agent is restricted to enumerated tool calls") rather than omitting. Sandbox guarantees that exist as policy without enforcement evidence are routed to recommended actions, not recorded as coverage.
Kill switch and rollback are non-optional. Per agent-affecting incident class: the named procedure (runbook pointer), the owner role, the recovery time objective (how fast the agent goes offline once the decision is made), the last-drilled date, and the rollback path (how state is reverted, where compensating actions live, who signs off on resumption). The drill cadence ties to tier; if the last-drilled date is older than the tier's required cadence, the entry is flagged as a gap. A documented but never-drilled kill switch is theatre; the drill record is the load-bearing field.
Logging, traceability, and audit record what is logged on every agent invocation, every tool call, every plan step, every human-oversight event, and every kill-switch action; the retention period; the access controls on the logs; the immutability posture (append-only, hash-chained, off-system snapshot); and the audit-trail integrity expectations the firm's cyber programme already carries (NYDFS Part 500 audit-trail requirements for covered entities, the FFIEC IT Handbook for federally regulated banks). Logging that exists but is not actually being written to (a frequent finding in pre-prod reviews) is recorded as exercised_in_test = no and routed to recommended actions.
Incident response names each incident class (confirmed tool-misuse outside scope, confirmed cross-scope tool query, confirmed external-system unauthorised action, suspected agent compromise, foundation-model behaviour shift suggesting upstream issue, sandbox-escape attempt), the procedure pointer, the regulator-notification triggers that apply (the entity's existing cyber-incident regimes... NYDFS 72-hour notice, the bank 36-hour computer-security incident notification rule, SEC 8-K Item 1.05, GLBA Safeguards notification, sector-specific notification regimes loaded per overlay), and the off-switch criterion. The off-switch is the operational decision the runbook turns on; if there is no off-switch, that is itself a critical-severity recommended action.
Escalation triggers are the production-monitoring conditions that surface to the AI risk committee, the CISO function, sector and conduct co-reviewers, or the model owner without waiting for the periodic review. Each trigger names the signal, the threshold (in firm policy units), the frequency, the owner (function), and the escalation path (named committee, officer, or process). The standard agentic trigger set: tool-invocation-rate spike per tool, cross-scope tool query rate above zero, human-oversight rejection rate above baseline, refusal-rate drop, kill-switch activation, foundation-model version change, drift in any blast-radius-relevant metric.
Residual risk names each concrete residual risk with likelihood, impact, basis, accepted owner role, accepted date, and review cadence. "Low" without a basis is opinion. Cross-reference any partial or fail result in the threat-assessment table to a residual-risk row; an unaccepted partial result is itself a finding. For agents with tool use, residual risk on the planner-executor seam, on the cross-scope query class, and on the excessive-agency class is non-optional.
Recommended owner actions name each gap with owner role (function), deadline, severity, and any depends-on. Severity is the second-line judgement, not the owner's. Critical-severity items typically block pre-prod sign-off. The standard recommended-action items that fire for any agentic deployment: foundation-model swap re-validation flag in change management; kill-switch drill cadence ratification per tier; per-tool blast-radius re-validation on any tool addition; expansion-review trigger on any new sponsor, geography, customer segment, or tool added post-deployment.
Source trace and confidence records every material claim, its source, the evidence pointer, and a confidence label. Vendor red-team results carry vendor-self-attestation confidence (typically low to medium); firm-internal red-team results carry higher confidence. Drill records, gate-decision logs, and configuration evidence carry higher confidence than attestations. Do not collapse vendor and firm evidence into one line. Items without evidence carry [evidence needed] and route to recommended actions.
Depth flexes with tier and audience. A pre-prod-gate review for a tier-3 internal agent compresses to a few pages of substance; a tier-1 customer-facing agentic system review with cyber and conduct overlays runs long and dense. Empty named sections are not acceptable, but compression is.
Sector and cross-cutting overlays
When the scope names a sector (banking, insurance, capital markets, payments-fintech), load the matching references/sector-overlays/<sector>.md. Each overlay carries sector-specific authority-statement constraints, tool-inventory expectations, oversight-design bars, monitoring signals, regulator-notification triggers, and co-reviewer expectations. The overlay's named additions land in the artifact; treating the overlay as background reading is the failure mode.
The insurance overlay is written deeper than the others because this skill absorbs the deprecated claims-ai-controls skill. For any agentic system in claims handling, FNOL, agent / producer co-pilot, or policyholder-facing actuation, the insurance overlay carries: NAIC Unfair Trade Practices Act (Model #880) and state Unfair Claims Settlement Practices Acts as the fair-claims-settlement bar (which an agent cannot move); NAIC Model Bulletin on AI Systems (December 2023) AIS Program element-by-element coverage; state DOI AI bulletins (Colorado SB 21-169 implementing regulation, NY DFS Insurance Circular Letter on AI/ECDIS January 2024, California DOI bulletin); HIPAA Privacy and Security Rules where the agent reads or writes PHI; NAIC Insurance Data Security Model Law (Model #668) §4.F where a TPA or claims administrator runs the agent; UPL boundary check on consumer-facing claims chatbots; Marketing Rule analogues and state insurance advertising rules for any customer-facing communication; ORSA materiality check for enterprise-material agents. The insurance overlay also carries the claims-specific control families: prompt-acknowledgement, reasoned-denial drafting, complaint flow, supervisor-override path, audit trail, denial-letter human review.
The cyber cross-cutting overlay should be considered the default for any agent with tool use, identity, or external-system actuation. The overlay carries the cyber-side mitigations expected (authentication, audit logging, network controls, secrets handling, supply-chain controls, IR runbook integration, sandbox enforcement evidence), the cyber-tagged monitoring signals (authentication anomalies, tool-invocation patterns, cross-scope query anomaly, output-anomaly detection), and the regulator-notification triggers most likely to apply.
The privacy and conduct cross-cutting overlays load when the scope flags them: privacy where the agent reads or writes regulated personal data with potential for tool-misuse exfiltration; conduct where the agent's actions touch customers (the insurance claims AI flow is the canonical example, but also agent-facing communications, customer-facing communications, and decision actuation that influences a customer outcome). Climate is not applicable to agentic-controls reviews.
Load only the overlays the scope names. Gold-plating with overlays the engagement does not implicate adds noise without challenge value.
Quality bar
The controls review is only credible when these hold:
- Every material claim cites a source. Unsupported items carry
[evidence needed] and route to recommended actions, not silently into the artifact body.
- Evidence is separated from inference. Vendor red-team results are not the same line as firm-internal red-team results; vendor self-attestation confidence is recorded explicitly. Drill records, decision-forum logs, and configuration evidence carry higher confidence than attestations.
- No fabricated regulatory facts. Unknown section references carry
[verify section] in the source-anchors file (not in the artifact body).
- Authority statement enumerates out-of-scope actions concretely. Descriptive prose without enumeration fails the section.
- Tool inventory carries blast radius on every entry. A tool list without blast-radius framing is not a control surface.
- Human oversight entries carry an evidence pointer that supports the assertion the oversight is effective. Theatrical oversight is the failure mode the section exists to surface.
- Kill switch entries carry a
last_drilled date. A documented but never-drilled kill switch is theatre.
- Threat assessment names the attack class, the carrier (system prompt, user prompt, retrieved content, tool output, agent memory, multi-agent message), and the source taxonomy (MITRE ATLAS, OWASP LLM, NIST AI 600-1) on every entry. Generic "prompt injection" without naming the carrier or class fails the section.
- Sandbox and execution-environment guarantees are recorded with enforcement evidence, not policy assertion. Policy without enforcement routes to recommended actions.
- Logging entries are exercised: confirmed by reading a sample log line, not asserted from a policy document.
- Incident-response entries name the regulator-notification trigger per the loaded overlays and the off-switch criterion. The off-switch is non-optional on cyber-relevant and customer-impact-relevant classes.
- Foundation-model swap re-validation is a non-optional recommended action for any architecture depending on a third-party foundation model.
- No named institutions outside finalised public enforcement actions; examples are anonymised and public-source-derived.
- Reviewer roles are functions, never named individuals.
- The artifact is a draft until the human reviewer attests. The skill does not file the artifact, post to the AI risk committee, trigger an IR runbook, or activate a kill switch.
Adaptation
Tier drives depth. Lifecycle stage drives which sections lean heavy (pre-prod-gate emphasises tested attacks, sandbox guarantees, oversight design, and drill plans; in-prod-periodic emphasises monitoring evidence, drill records, scope-drift audit; post-incident emphasises the timeline, the IR runbook output, and the recommended actions; expansion-review emphasises the delta against the prior review). Audience drives tone (working group is plain, committee is structured, examiner response is formal, board distillation pulls residual risk and recommended actions to the front). Sector and cross-cutting overlays load from the scope. Source posture sets what the artifact can assert at high confidence and what carries [evidence needed]. Where firm-specific policy or taxonomy applies, it lives in references/firm-overlay.md (consumed when present) and never in the artifact directly.
Output
Default to drafting the controls review against templates/default-output.md. Render as Word for committee or CISO review, or another format the audience asks for. Produce the structured record at schemas/agentic-controls-review.schema.json when downstream consumers (genai-pre-prod-review, board-ai-risk-pack, ai-governance-reviewer, the firm cyber IR chain, the model inventory of record) need it. The reviewer-attestation block is filled by the human reviewer (AI Governance Lead with CISO function for cyber-flagged reviews; co-acceptors per sector and cross-cutting overlays where applicable; for insurance claims AI, the privacy officer co-signs where PHI flows and claims operations leadership co-signs the operational-control acceptance); the artifact is filed only after.
Downstream consumers: genai-pre-prod-review consumes the structured object for the gate decision; board-ai-risk-pack pulls the residual-risk summary, the recommended-actions list, the kill-switch posture, and material incident-response triggers; the ai-governance-reviewer agent pulls the structured object for second-line challenge; the firm cyber IR chain consumes the incident-response section as input to runbook design and IR drill scope; the model inventory of record pulls the agent metadata for the central registry. The schema is the input contract for those consumers; additive changes only, never silent renames. Breaking changes ship as a versioned migration with the consumers told in advance.
Pointers
references/source-anchors.md — citations and excerpts for the named anchors.
references/sector-overlays/{banking,insurance,capital-markets,payments-fintech}.md — sector overlays loaded from scope; the insurance overlay carries the absorbed claims-AI control bar.
references/cross-cutting/{cyber,privacy,conduct,ai-ethics}.md — cross-cutting overlays loaded from scope; cyber is the default for tool-using agents.
references/firm-overlay.md — firm policy, taxonomy, named owners (consumed when present).
templates/default-output.md — controls-review template.
schemas/agentic-controls-review.schema.json — structured-output contract.
examples/ — anonymised public-source-derived scenarios (customer-service agent in a regional bank; insurance claims-AI agent absorbing the claims-ai-controls scenario).
TROUBLESHOOTING.md — recurring defects.