Tools to Help Banks Comply With AI Agent Rules | Arthur

Best Practices for Building Agents Recap

Banks are deploying AI agents across customer support, fraud detection, KYC and AML, credit decisioning, and back-office operations. These agents do not just answer questions. They reason, call tools and APIs, access sensitive systems, and take multi-step actions on their own. That autonomy is exactly what makes them useful, and exactly what makes them a regulatory problem.

No single tool makes a bank compliant. In practice, banks assemble a stack of governance, discovery, observability, evaluation, and guardrail tools that together satisfy a fast-moving set of authorities: the EU AI Act, the NIST AI Risk Management Framework, DORA, consumer protection and fair lending law, and data protection rules such as GLBA and GDPR. Notably, as of an April 2026 interagency rewrite, SR 11-7 model risk guidance no longer covers generative and agentic AI, which reshapes how banks have to frame agent compliance. This post breaks down the categories of tools banks use, maps them to the regulations that actually apply, and explains how to choose them.

Why AI agents create new regulatory risk in banking

Traditional model risk management was built around a clear question: is this model accurate, fair, and explainable? Agentic AI shifts the question. The concern is no longer only whether a model is safe, but whether the bank can discover, observe, govern, and reconstruct every action an agent took.

Three properties of agents drive the new risk:

This is the "shadow agent" problem, and it is already here. According to McKinsey, 80% of organizations are reporting risky behavior from AI agents. For a regulated bank, an unmanaged agent with access to customer data or core systems is a compliance failure waiting to be found in an audit.

The compliance tool stack for AI agents

Because agents span discovery, runtime control, and oversight, banks need several categories of tooling that work together. Here is how the stack breaks down.

1. Agent discovery and inventory

You cannot govern what you cannot see. Discovery tools automatically scan cloud and compute environments to find and catalog agents as they appear, instead of relying on manual spreadsheets. Common discovery techniques include:

A complete inventory, with a named owner for every agent, is the foundation of any audit. An agent without an owner is an agent without accountability.

2. AI governance and policy enforcement

Governance tools turn a raw inventory into a controlled, auditable operation. They provide a unified policy framework across the enterprise, agnostic governance that works no matter the cloud, framework, or model, and customizable policies because one size does not fit all. A customer support agent for retail banking needs different controls than an internal AML investigation agent.

The most important capability is enforcement that adapts to each use case: which data an agent can access, which tools it can invoke, and what behaviors are acceptable.

3. Observability and audit trails

Regulators and auditors want to answer "what did the agent do, and why?" Observability tools trace every agent run end to end: prompts, completions, tool calls, retrievals, reasoning steps, token counts, and cost. Built on open standards like OpenTelemetry, this tracing produces the decision lineage and forensic replay that audits depend on. The teams that instrument early are the ones that can demonstrate control later.

4. Continuous evaluation and monitoring

Agents are non-deterministic. One that passes a test suite today can fail the same case tomorrow, and production inputs are far more varied than any handwritten test set. Continuous evaluations run automated checks against live traffic to catch issues like hallucination, incomplete answers, off-topic responses, and wrong tool use before customers or regulators do. Pairing evals with alerting means the compliance team is notified the moment behavior drifts, rather than after a complaint.

5. Guardrails for real-time policy enforcement

Discovery, observability, and evals are retrospective. Guardrails intercept agent behavior in real time. They fall into two types:

6. Model governance and explainability

Banks still need model inventory, validation, documentation, bias checks, and explainability for the agents they run. The newer agent tooling complements existing model governance platforms rather than replacing them, adding the action-level traceability that traditional model risk management was never designed to capture.

A word of caution on framing here. Many guides still cite SR 11-7 as the regulation that governs AI agents. As of the April 17, 2026 interagency rewrite, that is no longer accurate. The OCC, Federal Reserve, and FDIC formally placed generative and agentic AI outside the scope of SR 11-7, with a request for information (RFI) on AI-specific model risk expected to follow. SR 11-7 still applies to traditional quantitative models like credit scoring and VaR, but not to your loan-underwriting RAG pipeline or your KYC LLM classifier.

That carve-out does not create an unregulated zone. It removes one consolidated frame and leaves the agent governed by a different set of authorities.

What actually governs AI agents in banking now

The regulatory picture is more fragmented than most "AI compliance" content suggests. With SR 11-7 stepping back from GenAI, the controls banks build still have to satisfy a range of overlapping authorities. The tool categories above map onto them directly.

The throughline holds even as the specific pegs shift: banks need to demonstrate visibility, control, traceability, and accountability for every agent in production. The forthcoming RFI is also a reason to build defensible controls now, since the firms engaging with that process are the ones whose real-world architecture will shape whatever replaces SR 11-7's coverage of GenAI.

How Arthur helps banks govern AI agents

Arthur was built for exactly this challenge: giving teams the visibility and control to move agents from experimental pilots to governed production systems. It spans both the discovery and governance layer and the development lifecycle that produces a governable agent in the first place.

Agent Discovery and Governance (ADG). Arthur's ADG platform automatically discovers and catalogs agents across fragmented environments like Vertex AI, Bedrock, and others, then brings them under a single control plane. It provides a unified policy framework, agnostic governance across clouds and frameworks, and customizable policies so each agent gets controls that fit its use case. Compliance teams can see the tools, models, data sources, and subagents each agent uses, and assign a clear owner for accountability.

The Agent Development Lifecycle (ADLC). Beyond discovery, Arthur covers the practices that make an agent enterprise-ready:

Because Arthur runs natively inside your own cloud with a federated data plane and control plane, sensitive inference data, prompts, completions, retrieved documents, and PII stay inside your environment. Only lightweight, anonymized metrics flow to the control plane. For regulated industries, keeping production data in-VPC is often the difference between an agent that clears compliance review and one that does not.

What to look for when choosing tools

When evaluating tools to help your bank comply with AI regulations for agents, the questions that matter most are:

TLDR