top of page

AI Vendor Due Diligence: The Questions to Ask Before You Sign

Writer: Matt Lazarus
Matt Lazarus
Sep 14
5 min read
Isometric illustration on a dark navy background of a glowing AI cube on a pedestal inspected through a large magnifying lens, with checklist panels and a shield orbiting it.
AI products deserve deeper inspection than the standard procurement checklist gives them.

Almost every software product you evaluate this year will claim to be an AI product. Some are genuine platforms with models and data pipelines at their core; others are thin wrappers around a third-party API, assembled in a quarter. From the sales deck, the two are indistinguishable.

 

Traditional procurement was built for a different question. Security certifications, uptime service levels and reference calls tell you whether a vendor can run software reliably. They tell you almost nothing about where your prompts travel, whether your data trains someone else's model, how often the product is simply wrong, or what happens when the model behind it changes overnight.

 

The gap matters because AI purchases fail differently. A slow implementation can be managed; a vendor that quietly retains your customer records for training, or an assistant that fabricates figures in front of a client, is a different category of problem. This guide sets out the due diligence questions that separate defensible AI purchases from expensive lessons - and they are questions any executive can ask, no data science background required.

 

Key Takeaways

 

  • AI procurement is different: data flows, model dependencies and output behaviour carry risks standard SaaS checklists never touch.

  • Follow the data first: where prompts and outputs travel, and whether they train models, decides your privacy and confidentiality exposure.

  • Contract for change: models, prices and vendors shift fast, so exit rights, change notices and data return terms matter most.

 

Why Is AI Vendor Due Diligence Different From Standard Procurement?

 

Because the risks live in different places. Standard due diligence tests whether a vendor can run software reliably and securely. AI due diligence must also test what the product does with your data, how its outputs behave when they are wrong, and how dependent the vendor is on models it does not control.

 

None of the standard checks become optional - certifications, insurance, financial health and references still matter. But each traditional check now has an AI-specific counterpart that most procurement templates simply do not ask about:

 

Standard check

AI-specific counterpart

Security certifications (ISO 27001, SOC 2)

Training-use rights over your prompts, files and outputs

Uptime and support service levels

Accuracy expectations and behaviour when the model is wrong

Product roadmap briefing

Notice obligations for model changes and deprecations

Customer reference calls

A structured evaluation on your own data before signing

 

The rest of this guide works through those counterparts one at a time, in the order they tend to matter.

 

Where Does Your Data Go When Staff Use the Product?

 

Follow the complete path: every prompt, uploaded file and generated output travels from your users to the vendor, often onward to subprocessors and foundation-model providers, and lands in storage in specific regions. Due diligence means getting that entire path in writing, not accepting a reassuring sentence on a trust page.

 

The questions to put to the vendor are concrete. Is our data used to train or improve any model, and is the answer contractual or just a settings toggle someone can change? How long are prompts and outputs retained, and can we set that window? Do the vendor's staff ever review customer prompts? Which subprocessors receive our data, and in which countries does it rest? For Australian organisations, offshore hosting is not automatically a problem, but it must be a known and accepted fact rather than a surprise - the mechanics are covered in our guide to data sovereignty and AI.

 

Answers you should treat as warning signs: training-use terms that differ between pricing tiers, retention windows the vendor cannot state, and subprocessor lists that are unavailable or last updated years ago.

 

Isometric illustration on a dark navy background of two conveyor lanes - sealed boxes passing one checkpoint beside an open glowing box passing several, with light streams traced back to a source database.
Standard procurement checks the box; AI due diligence follows the data all the way through.

Whose Model Is Actually Under the Hood?

 

Many AI products are orchestration built around a foundation model the vendor does not own. That is not a defect - most good products are built this way - but it means you inherit the upstream provider's terms, hosting regions and change cadence, so you need to know exactly who sits underneath.

 

Ask which model providers the product depends on, and what happens when one of them changes or retires a model. A silent model swap can change tone, accuracy and behaviour across your whole deployment overnight, so ask what notice you receive and whether you can pin or delay upgrades. Then ask the harder commercial question: what is proprietary here? A vendor whose only asset is a prompt in front of someone else's model can be replicated - or undercut - quickly. Durable value usually lives in workflow depth, evaluation discipline, data integrations and domain tuning.

 

How Should You Test Accuracy Before You Buy?

 

On your own data, with your own people, against tasks you defined - never on the vendor's demo environment alone. A structured pilot with written acceptance criteria is the single most revealing step in AI due diligence, and reputable vendors will agree to one.

 

Demo environments are curated; your document library, CRM and ticket history are not. Define a set of representative tasks before the pilot starts, including the awkward ones, and agree what good looks like: acceptable error rates, required citations or sources, and behaviour on questions the product cannot answer. Pay particular attention to how it fails. A product that says it does not know is workable; one that produces a fluent, plausible and wrong answer is a liability in front of customers or a boardroom.

 

Also decide who inside your organisation is accountable for checking outputs during the pilot, and pair the trial with clear internal rules on what staff may feed into it - the ground covered by an AI acceptable use policy.

 

What Contract Terms Matter Most in an AI Agreement?

 

The terms that manage change. A prohibition on training with your data, notice of material model changes, data return and deletion on exit, meaningful accountability for outputs, and price protection - these do more work than any warranty, because the product you buy will not be the product you are running in eighteen months.

 

Work through them in order. Training-use restrictions belong in the contract body, not in a linked policy the vendor can edit. Model-change notice should cover replacements and retirements of underlying models, not just user-interface updates. Exit terms should state the format and timeframe in which you get your data, configurations and prompt libraries back. On liability, be realistic: no vendor will underwrite every wrong answer, but they can commit to remediation processes, service credits and cooperation in incident investigations. Price protection matters because usage-based AI pricing can move sharply; caps or re-negotiation triggers keep the surprise manageable.

 

How Do You Judge Vendor Viability in a Market Moving This Fast?

 

Assume consolidation and plan for it. Ask about funding runway, customer concentration and what happens to your data and service if the vendor is acquired or winds down. The goal is not to avoid young vendors - it is to make sure an exit is survivable.

 

Traditional software escrow offers little comfort when the product is a hosted service wrapped around third-party models. The practical insurance is portability: keep your prompt libraries, workflow definitions, integration mappings and evaluation results in forms you can move, and confirm the contract lets you take them. Then size the commitment to the evidence - pilot with one team, expand when the results justify it, and avoid multi-year lock-in on a product category that is repricing itself every year.

 

Due Diligence Is Cheaper Than an Unwind

 

The questions in this guide take days to work through. Unwinding a poor AI purchase - migrating data out, retraining staff, explaining to customers or regulators where information went - takes quarters, and some of the cost is reputational and permanent. The discipline is not about slowing AI adoption down; it is about making the purchases you do approve defensible.

 

The strongest position to negotiate from is knowing your own estate before the pitch arrives. An AI Data Readiness Audit establishes what your data can actually support, which makes vendor claims much easier to test. And if the real question is which use cases deserve a vendor at all, a structured AI opportunity discovery and enablement engagement ranks the candidates on evidence before a dollar is committed.

 
 
bottom of page