Regulation explainerFinancial Services

LLM model risk management after SR 11-7: what supervisory texts now imply

Generative AI now sits outside the main US model risk guidance: SR 26-2, which replaced SR 11-7 in April 2026, excludes generative and agentic models and leaves their governance to each bank1. An LLM application still needs governing. This explainer shows how to bring one into a model risk framework using SR 26-2's principles, the PRA's SS1/23 and, for high-risk uses, the EU AI Act.

Reviewed 8 min read

On this page
  1. What this explainer covers and what it cannot decide for you
  2. What changed when SR 26-2 replaced SR 11-7
  3. Texts that govern LLM applications at banks and insurers
  4. Registering an LLM application component by component
  5. Validating a model your bank did not train
  6. Sections a validation report for an LLM application should contain
  7. Where LLM validation programs go wrong
  8. A hypothetical briefing assistant moves from pilot to production
  9. Questions and answers
  10. Sources

What this explainer covers and what it cannot decide for you

What changed when SR 26-2 replaced SR 11-7

On April 17, 2026, the Federal Reserve, OCC and FDIC issued revised model risk guidance superseding SR 11-7 and the 2021 interagency statement on BSA/AML models1. The OCC's companion, Bulletin 2026-13, rescinded OCC Bulletin 2011-123.

Three changes matter for LLM work. The new text defines a model as a complex quantitative method applying statistical, economic or financial theories to turn input data into quantitative estimates2. A footnote states that generative and agentic AI models are novel and rapidly evolving and so fall outside the guidance, while the bank's own risk management and governance practices should decide controls for anything not covered2. And the guidance sets no enforceable standards, with most relevance to banking organizations above $30 billion in total assets2.

Supervisory action can still follow from unsafe or unsound practices caused by poor model risk management2, and the agencies plan a request for information covering generative and agentic AI3. Banks that already run LLM applications through model risk management have little reason to stop; the rest need a written internal standard, and the most defensible one borrows the guidance's vocabulary of materiality, conceptual soundness, outcomes analysis, ongoing monitoring and effective challenge.

One boundary case: a non-generative model fed with fields an LLM extracted, such as a credit model reading parsed bank statements, stays in scope, and its validation must examine the extraction as a data input.

Texts that govern LLM applications at banks and insurers

Four instruments do most of the work. Read the primary text before relying on any summary.

SR 26-2, Revised Guidance on Model Risk Management

United States (Federal Reserve, OCC, FDIC)

Applies whenBanking organizations using models, most relevant above $30 billion in total assets; generative and agentic AI are outside its scope2.

  • Sound practice for statistical models and non-generative AI across development, validation, monitoring and governance.
  • For excluded tools, the bank's own governance practices determine appropriate controls2.
  • For vendor models, understand conceptual soundness, development data and performance, then monitor outcomes2.

PRA Supervisory Statement SS1/23, Model risk management principles for banks

United Kingdom (Prudential Regulation Authority)

Applies whenUK-incorporated banks, building societies and PRA-designated investment firms with internal model approval for regulatory capital; other firms may adopt it voluntarily4.

  • Adopt the PRA model definition, which includes qualitative or expert-judgment inputs and qualitative outputs4.
  • Keep a firm-wide inventory and risk-based tiering; manage risks of AI used in modeling within the same framework4.
  • Validate your own use of vendor models against your own outcomes.

Artificial Intelligence Act, Regulation (EU) 2024/1689, with the Digital Omnibus amendments

European Union

Applies whenAn AI system evaluates the creditworthiness of natural persons or sets their credit score, or performs risk assessment and pricing for natural persons in life and health insurance (Annex III, point 5(b) and (c))5.

  • Providers: risk management, data governance, documentation, logging, human oversight, accuracy and robustness.
  • Deployers in these categories: a fundamental rights impact assessment before first use under Article 275.
  • Regulation (EU) 2026/1744 sets 2 December 2027 for these Annex III obligations6.

Interagency Guidance on Third-Party Relationships: Risk Management (SR 23-4)

United States (Federal Reserve, OCC, FDIC)

Applies whenThe LLM comes from a model provider, cloud platform or software vendor7.

  • Manage the relationship from planning and due diligence through contracting, monitoring and termination7.
  • Negotiate notice of model version changes, limits on use of your data, and audit rights.

Registering an LLM application component by component

An entry naming only the vendor model misses most of what can change behavior.

ComponentWhat can change itWhat the inventory recordsRe-test trigger
Base modelVendor releases, deprecations, provider-side updatesProvider, model identifier and version, hosting region, contract referenceAny version change, including one you only learn of afterwards
System prompt and templatesDevelopers and product owners, often frequentlyVersioned text in source control with an owner and approval recordEvery change that reaches production
Retrieval corpusUploads, deletions, re-indexing, chunking or embedding changesSource systems, refresh cadence, access rules, embedding model versionNew embedding model, new chunking or a large document addition
Tools and actionsNew integrations and permission changesEach tool, its permissions and whether it writes to a system of recordAny new tool or broader permission
Evaluation setValidator additions and production incidentsVersion, coverage by task and risk, last run and resultNot a trigger itself: it is the instrument re-run for every other trigger

Rule of thumb: if a change could alter what a user sees for the same input, govern it as a model change.

Validating a model your bank did not train

Conceptual soundness work moves to the application around a model you cannot inspect. ColdAI's financial services process builds backtesting, sensitivity analysis and bias testing into model development rather than after it8.

  1. Fix intended use and boundaries

    Write down which decisions the application informs, who sees its output and what it must never be used for. Drafting meeting summaries carries a different risk from suggesting credit limits.

    Output
    Intended-use statement and tier
    Owner
    Model owner
  2. Assess the vendor evidence

    Read the provider's documentation for training data scope, known limitations and published evaluations, and record what you could not obtain. SR 26-2 accepts that vendors may withhold code and data, and says the principles still apply2.

    Output
    Vendor evidence memo
    Owner
    Validator
  3. Review retrieval and prompt design

    Check that retrieval returns the right passages, that entitlements survive retrieval and that the prompt keeps answers tied to sources. Many wrong answers start in retrieval rather than the base model, so test it on its own.

    Output
    Design review findings
    Owner
    Validator with developers
  4. Run outcomes analysis on an evaluation set

    Score answers against a versioned set of cases with known good answers, including adversarial prompts and questions the corpus cannot answer. See building an LLM evaluation set for construction and judge calibration.

    Output
    Evaluation report against agreed thresholds
    Owner
    Validator
  5. Sample human review in production

    Agree a review sample per period, stratified by task, segment and language, and compare quality across customer groups against conduct and fair lending standards. Record error categories, not one pass rate.

    Output
    Review and fairness log
    Owner
    First-line business owner
  6. Set monitoring and change control

    Monitor input mix, refusal rates, retrieval hit rates and reviewer error rates, and route every component change in the table above through re-testing before release.

    Output
    Monitoring plan and change policy
    Owner
    Model owner and model risk management

Sections a validation report for an LLM application should contain

0 of 9 checked

Where LLM validation programs go wrong

Validating the base model and ignoring the application

Early signalThe report discusses public benchmarks, not your prompts, corpus or tools.

MitigationScope validation to the deployed application and every component in the inventory.

A silent vendor update

Early signalAnswer style or refusal rates shift without any release on your side.

MitigationPin model versions where the provider allows it, monitor output drift and require advance notice in the contract.

Calling every use low risk because a person reviews the output

Early signalReviewers approve nearly every output without edits.

MitigationMeasure edit and override rates; treat rubber-stamping as an absent control and tier the application accordingly.

Reading the generative AI exclusion as an exemption

Early signalLLM applications have no inventory entry or owner.

MitigationAdopt an internal standard that applies materiality-based controls to excluded tools, which the guidance itself expects banks to determine2.

A hypothetical briefing assistant moves from pilot to production

Questions and answers

Is SR 11-7 still the US standard for validating generative AI models?

No. SR 26-2 superseded SR 11-7 on April 17, 2026, and places generative and agentic AI models outside its scope2. It says a bank's own risk management and governance practices should determine controls for tools it does not cover, so many banks keep LLM applications in the model inventory under an internal standard.

Do UK banks have to put LLM applications through SS1/23?

SS1/23 applies to UK banks, building societies and PRA-designated investment firms with internal model approval for regulatory capital4. Its model definition includes qualitative inputs and outputs, and it expects firms to manage risks from AI used in modeling, so an LLM application that informs business decisions can fall inside it. Each firm decides through its own definition and tiering.

Is a customer-facing banking chatbot a high-risk AI system under the EU AI Act?

Not merely for being a chatbot. It becomes high-risk if intended to evaluate the creditworthiness of natural persons, set their credit score, or assess risk and price life or health insurance, the categories in Annex III, point 55. Separately, Article 50 requires telling people they are interacting with an AI system unless that is obvious.

Who should validate an LLM application if validators have no LLM experience?

SR 26-2 describes effective challenge as work by people with the right expertise, enough independence to stay objective and the standing to force change2. Banks lacking LLM skills train validators, hire specialists or use external validators, keeping oversight of that work inside their own model risk management.

Sources

  1. SR 26-2: Revised Guidance on Model Risk Management — Board of Governors of the Federal Reserve System · checked 10 October 2026
  2. Supervisory Guidance on Model Risk Management (SR 26-2 attachment) — Federal Reserve, FDIC and OCC · checked 10 October 2026
  3. OCC Bulletin 2026-13: Model Risk Management: Revised Guidance — Office of the Comptroller of the Currency · checked 10 October 2026
  4. SS1/23 Model risk management principles for banks — Prudential Regulation Authority, Bank of England · checked 10 October 2026
  5. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act) — EUR-Lex · checked 10 October 2026
  6. Regulation (EU) 2026/1744 (Digital Omnibus on AI) amending Regulation (EU) 2024/1689 — EUR-Lex · checked 10 October 2026
  7. SR 23-4: Interagency Guidance on Third-Party Relationships: Risk Management — Board of Governors of the Federal Reserve System · checked 10 October 2026
  8. Financial Services: delivery process and use cases — ColdAI

More in Financial Services

Back to Financial Services

Next step

Send one LLM application to tier and scope for validation

Share the use case, the model provider, how output reaches people and your current model risk policy. We will come back with a component inventory, a suggested tier and a validation scope to discuss.

Scope an LLM validation