AI Supplier Due-Diligence Checklist

A compact evidence checklist for buyers evaluating AI products and suppliers.

Use this checklist as a starting point for proportionate AI supplier evaluation. Not every question applies to every purchase, and higher-impact uses require deeper evidence.

Supplier and product identity

  • What legal entity will contract with us?
  • What exact product/service are we buying?
  • Who owns and operates it?
  • Which critical third parties are involved?
  • Is the proposed deployment materially different from the public product?

Intended use and limitations

  • What is the product designed to do?
  • What uses are unsupported or prohibited?
  • What known limitations or failure modes are documented?
  • What human oversight does the supplier expect?

Models and architecture

  • Which model(s) or model providers are used?
  • Can the supplier switch model/provider without notice?
  • Are retrieval systems, agents, tools or external APIs involved?
  • What material dependencies could affect service or risk?
  • How are model and system changes tested?

Data and privacy

  • What customer data is processed?
  • Where is it stored and processed?
  • Who can access it?
  • How long is it retained?
  • Is customer content used for training, fine-tuning or improvement?
  • Which subprocessors are involved?
  • How is deletion handled?
  • Can the supplier support the buyer's privacy impact assessment where needed?

Security

  • What security framework and secure-development controls are used?
  • Has the product undergone relevant security testing?
  • How are vulnerabilities managed?
  • How are AI-specific threats considered?
  • What logs and access controls are available?
  • How are supplier-chain dependencies secured?
  • What incident-response process applies?

The NCSC Guidelines for secure AI system development provide a useful benchmark for secure design, development, deployment and operation.

Testing and performance

  • What metrics are used to measure the product?
  • What datasets and scenarios were used in evaluation?
  • Do published benchmarks match our use case?
  • Can we run a buyer-side pilot or evaluation?
  • What failure modes have been identified?
  • How are regressions detected after model updates?

Governance and accountability

  • Who owns AI risk within the supplier?
  • How are material incidents escalated?
  • How are model/provider changes approved?
  • What policies are actually evidenced in operation?
  • Which external certifications or independent assessments apply to this exact service?

Human oversight and users

  • How are users told AI is involved?
  • Can users challenge, correct or escalate outputs?
  • What controls reduce automation bias?
  • What training or usage guidance is provided?
  • Are accessibility impacts considered?

Customer and deployment evidence

  • Can customer relationships and use cases be verified?
  • Are case studies current and relevant to the same service?
  • Is customer evidence independently sourced or purely marketing-selected?
  • Are incident and support claims supported by evidence?

Contract and change control

  • Are data-use restrictions contractual?
  • Are model/provider changes notified?
  • Are material subprocessor changes controlled?
  • What audit/evidence rights exist?
  • What incident-notification commitments apply?
  • What happens if performance degrades?
  • What termination and exit rights exist?

Post-award monitoring

  • What performance data will the supplier provide?
  • How are model changes communicated?
  • How are incidents and vulnerabilities reported?
  • What evidence expires or needs refreshing?
  • What events trigger re-evaluation?

Evidence rating

For each answer, classify the support as: 1. Claim only — supplier assertion without supporting evidence. 2. Documented — policy, technical documentation or contract evidence exists. 3. Demonstrated — evidence shows the control operating in practice. 4. Independently verified — a suitable independent party has checked the relevant evidence.

A strong procurement record distinguishes those levels rather than treating every “yes” answer as equivalent.

Return to the AI Procurement Knowledge Base.