Evaluate AI Suppliers

How to distinguish supplier claims from independently supportable evidence when evaluating AI services.

Direct answer

How should buyers evaluate an AI supplier? Separate every material supplier claim from the evidence supporting it, confirm that evidence applies to the exact product and deployment, assess architecture and dependencies, data practices, security, performance, human oversight and change controls, then record residual uncertainty and contractual conditions. Marketing claims alone are not verified evidence.

AI supplier evaluation should be evidence-led. A polished policy, security badge or benchmark slide is not the same as evidence that the proposed service is suitable for the buyer's intended use.

This page is the lifecycle home for evaluation. Use it to structure comparison; use the deeper guides below for verification, due diligence, security, dependencies, oversight and agentic systems. It is not a complete paid evaluation matrix.

1. Separate claims from evidence

For each material requirement, record:

  • the supplier claim;
  • the evidence supplied;
  • whether the evidence is current;
  • whether it applies to the exact product and deployment being procured;
  • whether it has been independently verified;
  • any residual uncertainty.

2. Verify the supplier and product

Confirm the contracting entity, product identity, ownership, support route and material subcontractors. Distinguish the supplier's own product from third-party models, APIs and infrastructure on which it depends.

See independent AI supplier verification and AI vendor due diligence.

3. Assess architecture and dependencies

Understand the system boundary. Ask which models, clouds, retrieval systems, data stores, safety systems and external APIs are used. Determine which dependencies can change without buyer approval and which failures could interrupt service or materially change risk.

See model, API and subprocessor dependencies.

4. Review data practices

Check data flows rather than relying only on a privacy policy. Evidence should address processing purpose, locations, retention, subprocessors, access, deletion, customer-content use, training or improvement use, and separation between tenants where relevant.

See ICO AI and data protection for procurement.

5. Evaluate security

Apply standard supplier security due diligence and AI-specific questions. The NCSC secure AI guidance recommends attention to threat modelling, supply-chain security, documentation, secure deployment, incident handling and continuous monitoring.

Useful evidence may include penetration-test summaries, secure-development controls, vulnerability management, access-control design, threat models, logging capabilities, incident procedures and relevant independent assurance.

See security evidence buyers should request.

6. Test performance in context

Supplier benchmarks may not represent the buyer's use case. For material deployments, test with representative tasks, realistic inputs and expected edge cases. Record the evaluation method, acceptance threshold and limitations.

Consider both average performance and harmful failure modes. The correct metric depends on the use case.

7. Review human oversight

Assess whether users can understand when AI is being used, identify uncertainty, challenge outputs and escalate issues. Check whether the proposed operating process creates automation bias or turns nominal human review into a rubber stamp.

See human oversight and consequential decisions and procuring agentic AI.

8. Examine governance and change control

Ask who owns AI risk inside the supplier, how material incidents are governed, how model changes are tested, how customers are notified, and whether a model/provider change could alter the agreed risk profile.

9. Check customer and deployment evidence

Customer evidence can help answer whether the supplier operates as claimed in real deployments. Give greater weight to evidence where the customer relationship, use case and source can be independently verified.

Do not treat testimonials, anonymous quotes and marketing case studies as equivalent to verified customer evidence.

10. Record residual risk

Evaluation should not force every uncertainty into a pass/fail box. Record unresolved issues, compensating controls, contractual conditions, pilot requirements and post-award monitoring actions.

Suggested evaluation record

For each requirement capture: 1. requirement; 2. supplier response; 3. evidence reference; 4. evidence strength; 5. evaluator finding; 6. residual risk; 7. contractual condition if required; 8. recheck date or trigger.

Public guidance explains the record structure. A complete organisation-specific evaluation matrix remains professional work.

How AI TrustMark fits

AI TrustMark's role is independent verification of supplier, product, customer and operational evidence. A TrustMark or assessment should complement, not replace, the buyer's own procurement decision.

Use the AI supplier due-diligence checklist for a compact review structure.

Next: Contract for AI services.