Thinking ·
AI vendor due diligence in financial services
Short answer
AI vendor due diligence in financial services has to answer three questions a standard software assessment does not: what happens when the underlying model changes, who owns the corrections your staff make, and what you can take with you if you exit. Regulated firms remain accountable for outsourced processes, so the vendor’s assurances are not a substitute for evidence you hold yourself.
Most regulated firms already have a competent third-party assessment process. Applied to an AI vendor it will cover security, resilience, data protection and financial stability perfectly well — and miss the three things that actually determine whether the arrangement survives contact with a regulator or with year two.
Why the existing questionnaire is not enough
A conventional supplier assessment assumes the thing you are buying is stable: it does what it did last month, and it will do it next month. AI vendors break that assumption in a specific way. The capability sits in a model the vendor usually does not own, which changes under them, and the quality of the output is probabilistic rather than deterministic.
None of the standard sections have a box for that. So it goes unasked, and it surfaces eighteen months later when someone asks why the output quality moved and nobody can produce a baseline.
The three questions to add
1. What happens when the model changes?
Underlying models are deprecated, retrained and reweighted on the provider’s schedule, not yours. Ask the vendor, in writing, which parts of the product survive a model swap, what their process is when a provider deprecates a version, and whether they will notify you before a change reaches your users.
The contractual version of this is a service-level commitment that survives model changes at no additional cost. Vendors with real engineering agree to it. Vendors selling prompts negotiate hard, which is itself the answer.
2. Who owns the corrections your staff make?
This is the clause that decides whether you are building an asset or renting one. Every time an analyst overrides an output, that is a labelled example — the most valuable data your firm will generate about this task.
- Does it improve the model for your tenant, or for the vendor’s entire customer base — which may include firms you compete with?
- Can you export the corrections in a usable, structured format, or only view them in their interface?
- If you leave, do the corrections leave with you?
Ambiguity here is not neutral. An unclear IP and feedback clause defaults, in practice, to the vendor.
3. What is the measured error rate on your task?
Not their benchmark. Yours. A vendor that will run an evaluation on your data during a proof of concept, and commit to acceptance criteria based on the result, is a different proposition from one offering a demo and a reference call.
If nobody measured it at the start, nobody can tell you it has degraded — and “we would have noticed” is not a control.
Accountability does not transfer
The point regulators keep making about outsourcing applies here without modification: a firm remains accountable for a process it has outsourced. You can delegate the capability. You cannot delegate the obligation.
Practically, that means the evidence has to be evidence you hold, not assurances the vendor holds on your behalf. If your only proof that the model performs acceptably is a vendor slide, you have a supplier relationship where you needed a control.
The direction of travel across the EU AI Act, DORA, and UK operational resilience and outsourcing expectations is consistent on this: more documented understanding of third-party dependencies, clearer exit arrangements, and evidence that the firm — not the supplier — understands what it has bought. A due diligence process that produces artefacts you can hand to a reviewer is doing double duty.
General framing, not legal advice. Your compliance and legal teams own the specific obligations that apply to your firm.
What the second line will ask you for
Worth assembling before you are asked, because assembling it is also how you find out whether the deal is any good:
- A plain description of what the model does, what it does not do, and where a human decides.
- A measured baseline of performance on your own data, dated, with the method described.
- The data flow: what leaves your estate, where it is processed, whether it is retained, and whether it trains anything.
- The exit plan: what you receive, in what format, within what period — and, honestly costed, what leaving would take.
- Concentration risk: which underlying model provider sits beneath this, and how many of your other vendors sit on the same one.
That last one is under-asked. Several AI vendors in a portfolio can look diversified while every one of them routes to the same two model providers.
The build-versus-buy question underneath
Due diligence in regulated firms tends to ask whether the vendor is safe. The more useful question is whether the vendor is worth it — whether what you are buying is a capability you could not reasonably assemble, or a prompt with an onboarding process attached.
Both answers are legitimate. Buying non-differentiating capability when your team is constrained is good judgement. Paying for a moat that does not exist, on a multi-year term, is the expensive mistake — and it is usually made by firms whose assessment process was thorough about security and silent about substance.
The Wrapper Test scores exactly that gap in about two minutes, and the rubric is published so it can be attached to an assessment and argued with rather than taken on trust.
Common questions
- What should regulated firms ask AI vendors?
- Beyond standard security and resilience questions: what survives a model change, whether your data trains a shared model, what the measured error rate is on your own task, and exactly what you receive on exit and in what format.
- Does outsourcing AI transfer regulatory accountability?
- No. A regulated firm stays accountable for outsourced and third-party processes. The vendor can hold the capability, but the firm holds the obligation, which is why evidence you can produce yourself matters more than vendor assurances.
- What is the biggest gap in most AI vendor assessments?
- Evaluation evidence. Security, resilience and data protection are usually well covered by existing frameworks; a measured error rate on the firm’s own task, and who owns the corrections that improve it, usually are not.