
The evaluation questions most risk and compliance teams bring to AI vendor conversations are designed for technology, not for artificial intelligence. "What integrations do you support?" and "Can you show us a demo?" are reasonable starting points for software procurement. They are insufficient for AI in banking, where the regulatory exposure, model risk, and operational complexity are categorically different from what a standard enterprise software evaluation was built to surface.
The right framework starts with what regulators are watching, moves to what production deployment actually looks like at financial institutions similar to yours, and ends with the contractual commitments a vendor will either make or refuse to make. A vendor's willingness to answer specific questions with specifics, not slide decks and whitepapers, tells you more than the feature list does.
Third-party risk management processes at most financial institutions were not built for AI. They were designed to evaluate data vendors, SaaS platforms, and service providers: organizations with stable, auditable outputs. Machine learning models don't work that way. Their outputs shift with data distributions, their internal logic is often opaque, and the failure modes look nothing like what a standard IT risk assessment was built to catch.
The OCC, Federal Reserve, and FFIEC have each issued guidance making clear that financial institutions are responsible for AI models they deploy, regardless of whether those models are built in-house or procured from a vendor. SR 11-7, the Federal Reserve and OCC's guidance on model risk management, applies to vendor-supplied models as well as internally developed ones. Your validation obligations don't end at contract signature. They begin there.
Before you evaluate any AI vendor's capabilities, confirm that your risk assessment framework has been updated to account for this. If it hasn't, you're running vendor evaluation without the right measurement instrument.
"Is your AI explainable?" produces almost universally useless answers. Every vendor will say yes. The right question is: explainable to whom, in what format, and at what granularity?
For retail banking credit decisions and credit scoring applications, explainability has a specific legal meaning. The model must produce an adverse action notice that satisfies ECOA and FCRA requirements, and that explanation must reflect the model's actual decision logic, not a post-hoc approximation. Machine learning models using ensemble methods or deep learning architectures can satisfy this, but only with specific tooling built into the evaluation pipeline. Ask the vendor to walk through how adverse action notices are generated. If the answer involves a separate explanation layer that approximates the model's reasoning, document that as a compliance risk before you proceed.
For back-office and operational artificial intelligence in banking: document classification, regulatory compliance monitoring, anti-money laundering detection. The explainability standard here is different. Regulators want audit trails, not consumer-facing explanations. Ask what logs the system generates for every model decision, how long those logs are retained, who can access them, and whether they can be produced in a format your internal audit team can actually use.
Most AI vendors have case studies. Most of those case studies are marketing documents, not evidence of production readiness for regulated financial institutions.
Ask for three things a case study won't tell you. First, referenceable customers in comparable regulatory environments. A vendor's deployment at a regional credit union is not evidence of capacity to operate at a national bank under OCC supervision. Ask for references specifically from financial institutions under comparable oversight, and call those references. Second, ask about model drift. Every machine learning model experiences some degree of degradation as real-world data distributions shift away from training data. Ask how often the vendor monitors for drift, what the trigger thresholds are, and what remediation looks like when thresholds are crossed. A vendor that cannot describe this in operational terms has not been stress-tested in a demanding production environment. Third, ask what happens when the model is wrong. What is the escalation path when the AI's output contradicts a human reviewer's judgment? What human-in-the-loop architecture does the vendor actually run in production, not as an available option but as the standard deployment pattern at regulated financial institutions?
Fraud detection is the most polished use case in AI in banking vendor pitches. The demos are sophisticated. The benchmark statistics are carefully selected. Ask to see performance on your data, not theirs.
Standard vendor presentations show precision and recall from internal test sets. These figures are real and largely irrelevant. They tell you how the model performed on data it was designed to handle. The only meaningful test is a retrospective analysis on your institution's historical transaction data: records with labels you already have. If a vendor won't run this as part of the evaluation process, that is a material signal about how they manage production expectations.
For financial crime detection and anti-money laundering workflows, the evaluation question shifts from accuracy to auditability. Regulators and FinCEN are not primarily interested in whether your model flags the right alerts. They want to know whether you can explain, reconstruct, and defend every decision the model made during a regulatory exam. Ask the vendor to walk through what an examination of their AML output actually looks like, step by step. If they haven't supported a banking customer through one, you're pioneering that experience together.
Identity verification AI carries a specific risk that surfaces not in demos but in production data. Models trained on skewed demographic distributions produce higher error rates for underrepresented groups, a regulatory and reputational exposure that AI chatbots and other customer-facing tools compound. Ask whether the vendor has conducted disparate impact analysis on their identity verification models. Ask for the methodology, not a pass/fail result.
Generative AI in banking requires a different evaluation framework than predictive machine learning. The risk profile is different. The failure modes are different. And the regulatory guidance is still developing.
LLMs from providers including OpenAI, Google, and others are being embedded into financial services workflows at an accelerating rate: virtual assistants and AI chatbots for customer inquiries, internal knowledge tools for compliance teams, regulatory document summarization, and back-office process automation. The governance questions for generative AI in banking are not the same as for predictive models. A machine learning model produces a structured prediction. An LLM produces text — and text can be wrong, confidently and fluently, in ways that are hard to detect without human review.
For any agentic AI deployment, meaning AI agents that execute actions rather than produce outputs, ask the vendor to describe the human oversight architecture in specific operational terms. What decisions can an AI agent execute without human approval? What is the authority threshold? Who reviews escalated cases, and what is the SLA for that review? AI agents in banking operations can process documents, update records, and initiate workflows. Without explicit governance over what the agent can and cannot do autonomously, you are delegating operational authority to a system you cannot fully audit.
Natural language processing applications, including regulatory monitoring tools built on agentic AI or LLMs, carry a risk that is routinely underappreciated in early evaluations: hallucination. NLP systems built on generative models can produce confident summaries of regulatory text that misrepresent the source material. For compliance-critical workflows, ask how hallucination is detected, what the rates are for the vendor's specific use cases, and whether every AI-generated compliance output is subject to human verification before it influences a decision.
The credit decision use case illustrates the stakes clearly. A generative AI tool that assists with credit decisions and produces a well-structured but inaccurate summary of an applicant's financial history is not a software bug in the traditional sense. It's a model behavior that can trigger ECOA violations, CFPB examination findings, and class action exposure simultaneously. Predictive models and generative AI in banking need to be evaluated and governed as separate risk categories, even when they're sold as a single platform.
The data governance requirements for AI vendors are more demanding than for most software categories because your data is used to train and refine predictive models, not just to execute transactions, and the terms governing that use need to be explicit before you sign.
Ask where your data is stored, whether it is used to train models serving other customers, and what the contractual prohibition on data commingling looks like. AI vendors have materially different approaches to this. Some use customer data to continuously improve shared models across their client base. Others maintain dedicated model environments that isolate each institution's data entirely. Both architectures are defensible. The risk is not knowing which one you've agreed to.
Cybersecurity due diligence for AI vendors should go beyond standard SOC 2 and penetration testing attestations. Ask specifically whether your machine learning models can be probed through API access in ways that expose training data or proprietary decision logic. Model inversion and adversarial attack vectors are documented risks against deployed financial AI systems, not theoretical ones. A vendor operating in regulated environments should have specific, documented responses to these scenarios.
Unstructured data adds another layer. AI systems that process documents, communications, and unstructured records introduce data residency and classification questions that standard data governance frameworks handle inconsistently. Confirm that the vendor's data handling policies account for unstructured data at the same level of rigor as structured transaction records.
Customer onboarding workflows and digital banking experiences that incorporate artificial intelligence introduce additional complexity around consent, data minimization, and right-to-explanation obligations under federal and state consumer protection frameworks. If the vendor's product touches customer-facing processes, legal and compliance belong in the evaluation from the first meeting, not after you've selected a finalist.
Operational efficiency and digital transformation are what AI vendors pitch. Production reliability is what you need to evaluate. The gap between those two things is where vendor evaluation most commonly fails.
Ask for the vendor's uptime SLA and incident response protocols. Then ask for post-mortem reports from recent production incidents at banking customers: not a summary, the actual reports. Every AI vendor has incidents. The ones who can produce detailed post-mortems with root cause analysis, timeline reconstruction, and remediation documentation have built operational discipline into their systems. The ones who can't, or won't, are the ones where your institution becomes the post-mortem.
Predictive analytics and data analytics capabilities are commonly pitched as differentiators in banking AI vendor proposals. Ask whether those capabilities are genuinely integrated into the production system or whether they sit in a separate reporting layer requiring manual data export. A dashboard is not an integration.
For investment banking and institutional operations, ask specifically about the vendor's process for keeping AI outputs aligned with regulatory updates. SR 11-7 guidance evolves. The Consumer Financial Protection Bureau's interpretive positions shift. The AI systems you deploy today need to be maintained against the regulatory environment of the next five years, not just validated against requirements that existed at implementation. A vendor whose model update process requires a new engagement every time a regulatory requirement changes is a vendor whose total cost of ownership is being underrepresented in the initial proposal.
Fintech vendors entering the banking AI market often bring sophisticated technology and limited regulatory history. That's not disqualifying. It does mean your institution will be building the regulatory evidence base alongside them, and your risk management process needs to account for a vendor who has not yet been through a banking examination with an AI system in production. Budget for that accordingly.
Invisible works with financial institutions deploying AI into regulated operations, from model governance design to production oversight. See how we work with banking teams or get in touch.
Focus on explainability, production track record at regulated financial institutions, and what happens when the model is wrong. Ask vendors to walk through how audit trails are generated, whether model risk management documentation has been reviewed by a banking regulator, and what the escalation path looks like when AI output contradicts a human reviewer's judgment. Specifics in the answers separate production vendors from demo-ready ones.
Yes. SR 11-7 model risk management guidance from the Federal Reserve and OCC applies to vendor-supplied models as well as internally developed ones. Financial institutions are responsible for validating, monitoring, and documenting the performance of any model that influences a business decision, regardless of origin. Vendor contracts should specify who owns ongoing validation and how model changes are communicated to the institution.
Generative AI introduces hallucination risk — the system produces confident, fluent outputs that are factually incorrect. For banking applications, this is a compliance risk, not just an accuracy concern. Evaluate generative AI vendors on hallucination detection methodology, human verification requirements for compliance-sensitive outputs, and what governance architecture governs what AI agents can execute autonomously versus what requires explicit human approval.
At minimum: clarity on whether your data trains models serving other customers, a contractual prohibition on data commingling, data residency requirements specifying where your data is stored and processed, and audit rights allowing your institution to inspect data handling practices. For customer-facing AI, add requirements covering consent architecture, data minimization, and right-to-explanation obligations under applicable consumer protection law.
Request uptime SLAs, incident response protocols, and post-mortem reports from recent production incidents at banking customers. Every vendor has incidents; the question is how they document and remediate them. Also assess financial stability: AI vendors with limited banking deployments present third-party concentration risk that should surface in your vendor risk management review, especially for systems running fraud detection or compliance-critical workflows.
AI chatbots and virtual assistants in customer engagement contexts carry regulatory risk beyond standard software procurement. Conversational AI can generate fair lending concerns if responses vary by customer segment in ways that are not defensible under ECOA. Confirm that the vendor has conducted disparate impact analysis on their conversational outputs, and verify that human escalation is required, not just available, for defined categories of sensitive customer interactions.
Ask the vendor to run a retrospective analysis on your historical transaction data (records you already have labels for), not benchmark statistics from their internal test sets. Performance on your actual transaction patterns is the only meaningful predictor of production results. Also ask about false positive rates: over-flagging legitimate transactions creates operational costs and customer engagement damage that demo-environment precision-recall metrics will not capture.
