Regulation explainerFinancial Services
LLM model risk management after SR 11-7: what supervisory texts now imply
Generative AI now sits outside the main US model risk guidance: SR 26-2, which replaced SR 11-7 in April 2026, excludes generative and agentic models and leaves their governance to each bank1. An LLM application still needs governing. This explainer shows how to bring one into a model risk framework using SR 26-2's principles, the PRA's SS1/23 and, for high-risk uses, the EU AI Act.
On this page
- What this explainer covers and what it cannot decide for you
- What changed when SR 26-2 replaced SR 11-7
- Texts that govern LLM applications at banks and insurers
- Registering an LLM application component by component
- Validating a model your bank did not train
- Sections a validation report for an LLM application should contain
- Where LLM validation programs go wrong
- A hypothetical briefing assistant moves from pilot to production
- Questions and answers
- Sources
What this explainer covers and what it cannot decide for you
What changed when SR 26-2 replaced SR 11-7
On April 17, 2026, the Federal Reserve, OCC and FDIC issued revised model risk guidance superseding SR 11-7 and the 2021 interagency statement on BSA/AML models1. The OCC's companion, Bulletin 2026-13, rescinded OCC Bulletin 2011-123.
Three changes matter for LLM work. The new text defines a model as a complex quantitative method applying statistical, economic or financial theories to turn input data into quantitative estimates2. A footnote states that generative and agentic AI models are novel and rapidly evolving and so fall outside the guidance, while the bank's own risk management and governance practices should decide controls for anything not covered2. And the guidance sets no enforceable standards, with most relevance to banking organizations above $30 billion in total assets2.
Supervisory action can still follow from unsafe or unsound practices caused by poor model risk management2, and the agencies plan a request for information covering generative and agentic AI3. Banks that already run LLM applications through model risk management have little reason to stop; the rest need a written internal standard, and the most defensible one borrows the guidance's vocabulary of materiality, conceptual soundness, outcomes analysis, ongoing monitoring and effective challenge.
One boundary case: a non-generative model fed with fields an LLM extracted, such as a credit model reading parsed bank statements, stays in scope, and its validation must examine the extraction as a data input.
Texts that govern LLM applications at banks and insurers
Four instruments do most of the work. Read the primary text before relying on any summary.
SR 26-2, Revised Guidance on Model Risk Management
United States (Federal Reserve, OCC, FDIC)Applies whenBanking organizations using models, most relevant above $30 billion in total assets; generative and agentic AI are outside its scope2.
- Sound practice for statistical models and non-generative AI across development, validation, monitoring and governance.
- For excluded tools, the bank's own governance practices determine appropriate controls2.
- For vendor models, understand conceptual soundness, development data and performance, then monitor outcomes2.
PRA Supervisory Statement SS1/23, Model risk management principles for banks
United Kingdom (Prudential Regulation Authority)Applies whenUK-incorporated banks, building societies and PRA-designated investment firms with internal model approval for regulatory capital; other firms may adopt it voluntarily4.
Artificial Intelligence Act, Regulation (EU) 2024/1689, with the Digital Omnibus amendments
European UnionApplies whenAn AI system evaluates the creditworthiness of natural persons or sets their credit score, or performs risk assessment and pricing for natural persons in life and health insurance (Annex III, point 5(b) and (c))5.
Interagency Guidance on Third-Party Relationships: Risk Management (SR 23-4)
United States (Federal Reserve, OCC, FDIC)Applies whenThe LLM comes from a model provider, cloud platform or software vendor7.
- Manage the relationship from planning and due diligence through contracting, monitoring and termination7.
- Negotiate notice of model version changes, limits on use of your data, and audit rights.
Registering an LLM application component by component
An entry naming only the vendor model misses most of what can change behavior.
| Component | What can change it | What the inventory records | Re-test trigger |
|---|---|---|---|
| Base model | Vendor releases, deprecations, provider-side updates | Provider, model identifier and version, hosting region, contract reference | Any version change, including one you only learn of afterwards |
| System prompt and templates | Developers and product owners, often frequently | Versioned text in source control with an owner and approval record | Every change that reaches production |
| Retrieval corpus | Uploads, deletions, re-indexing, chunking or embedding changes | Source systems, refresh cadence, access rules, embedding model version | New embedding model, new chunking or a large document addition |
| Tools and actions | New integrations and permission changes | Each tool, its permissions and whether it writes to a system of record | Any new tool or broader permission |
| Evaluation set | Validator additions and production incidents | Version, coverage by task and risk, last run and result | Not a trigger itself: it is the instrument re-run for every other trigger |
Rule of thumb: if a change could alter what a user sees for the same input, govern it as a model change.
Validating a model your bank did not train
Conceptual soundness work moves to the application around a model you cannot inspect. ColdAI's financial services process builds backtesting, sensitivity analysis and bias testing into model development rather than after it8.
Fix intended use and boundaries
Write down which decisions the application informs, who sees its output and what it must never be used for. Drafting meeting summaries carries a different risk from suggesting credit limits.
Assess the vendor evidence
Read the provider's documentation for training data scope, known limitations and published evaluations, and record what you could not obtain. SR 26-2 accepts that vendors may withhold code and data, and says the principles still apply2.
Review retrieval and prompt design
Check that retrieval returns the right passages, that entitlements survive retrieval and that the prompt keeps answers tied to sources. Many wrong answers start in retrieval rather than the base model, so test it on its own.
Run outcomes analysis on an evaluation set
Score answers against a versioned set of cases with known good answers, including adversarial prompts and questions the corpus cannot answer. See building an LLM evaluation set for construction and judge calibration.
Sample human review in production
Agree a review sample per period, stratified by task, segment and language, and compare quality across customer groups against conduct and fair lending standards. Record error categories, not one pass rate.
Set monitoring and change control
Monitor input mix, refusal rates, retrieval hit rates and reviewer error rates, and route every component change in the table above through re-testing before release.
Sections a validation report for an LLM application should contain
Where LLM validation programs go wrong
Validating the base model and ignoring the application
Early signalThe report discusses public benchmarks, not your prompts, corpus or tools.
MitigationScope validation to the deployed application and every component in the inventory.
A silent vendor update
Early signalAnswer style or refusal rates shift without any release on your side.
MitigationPin model versions where the provider allows it, monitor output drift and require advance notice in the contract.
Calling every use low risk because a person reviews the output
Early signalReviewers approve nearly every output without edits.
MitigationMeasure edit and override rates; treat rubber-stamping as an absent control and tier the application accordingly.
Reading the generative AI exclusion as an exemption
Early signalLLM applications have no inventory entry or owner.
MitigationAdopt an internal standard that applies materiality-based controls to excluded tools, which the guidance itself expects banks to determine2.
A hypothetical briefing assistant moves from pilot to production
Questions and answers
Is SR 11-7 still the US standard for validating generative AI models?
No. SR 26-2 superseded SR 11-7 on April 17, 2026, and places generative and agentic AI models outside its scope2. It says a bank's own risk management and governance practices should determine controls for tools it does not cover, so many banks keep LLM applications in the model inventory under an internal standard.
Do UK banks have to put LLM applications through SS1/23?
SS1/23 applies to UK banks, building societies and PRA-designated investment firms with internal model approval for regulatory capital4. Its model definition includes qualitative inputs and outputs, and it expects firms to manage risks from AI used in modeling, so an LLM application that informs business decisions can fall inside it. Each firm decides through its own definition and tiering.
Is a customer-facing banking chatbot a high-risk AI system under the EU AI Act?
Not merely for being a chatbot. It becomes high-risk if intended to evaluate the creditworthiness of natural persons, set their credit score, or assess risk and price life or health insurance, the categories in Annex III, point 55. Separately, Article 50 requires telling people they are interacting with an AI system unless that is obvious.
Who should validate an LLM application if validators have no LLM experience?
SR 26-2 describes effective challenge as work by people with the right expertise, enough independence to stay objective and the standing to force change2. Banks lacking LLM skills train validators, hire specialists or use external validators, keeping oversight of that work inside their own model risk management.
Sources
- SR 26-2: Revised Guidance on Model Risk Management — Board of Governors of the Federal Reserve System · checked 10 October 2026
- Supervisory Guidance on Model Risk Management (SR 26-2 attachment) — Federal Reserve, FDIC and OCC · checked 10 October 2026
- OCC Bulletin 2026-13: Model Risk Management: Revised Guidance — Office of the Comptroller of the Currency · checked 10 October 2026
- SS1/23 Model risk management principles for banks — Prudential Regulation Authority, Bank of England · checked 10 October 2026
- Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act) — EUR-Lex · checked 10 October 2026
- Regulation (EU) 2026/1744 (Digital Omnibus on AI) amending Regulation (EU) 2024/1689 — EUR-Lex · checked 10 October 2026
- SR 23-4: Interagency Guidance on Third-Party Relationships: Risk Management — Board of Governors of the Federal Reserve System · checked 10 October 2026
- Financial Services: delivery process and use cases — ColdAI