Monitor AI Suppliers After Award
A practical post-award monitoring framework for AI-enabled services.
Direct answer
How should an AI supplier be monitored after deployment? Track agreed outcomes, incidents, model and provider changes, security findings, data practices, performance drift, human-oversight effectiveness and evidence expiry. Define re-approval triggers so material changes to models, data, use cases, subprocessors or risk can force re-evaluation rather than silently changing the accepted risk profile.
AI procurement is a lifecycle activity. Models, provider dependencies, data practices and operating conditions can change after deployment, so assurance should not end at contract signature or go-live.
1. Monitor the agreed outcomes
Track the metrics that justified the procurement. These may include task quality, error rates, service availability, latency, complaint rates, override rates, user outcomes or other use-case-specific indicators.
2. Track model and supplier changes
Require the supplier to notify material changes to model families, providers, safety controls, hosting, subprocessors, data-use terms or significant functionality. Compare changes against the production baseline and trigger re-evaluation where necessary.
3. Review incidents and near misses
Maintain a route for users and operational teams to report harmful, unsafe, incorrect or suspicious behaviour. Review recurring patterns, not only major incidents.
4. Monitor security
The NCSC secure AI guidance recommends monitoring system behaviour and inputs, applying secure approaches to updates, and collecting lessons learned. Supplier monitoring should therefore include relevant vulnerability, incident and update information.
5. Recheck data practices
Confirm that processing locations, subprocessors, retention and training-use terms continue to match the contract. Changes to product terms or service architecture can alter the data-protection position even where the front-end product looks unchanged.
6. Watch for performance drift
Performance may change because the model changes, the buyer's data changes, user behaviour changes or the system is used outside its original scope. Re-test material use cases periodically and after significant changes.
7. Review human oversight
Check whether review and escalation controls are still functioning in practice. Monitor whether users routinely accept AI outputs without scrutiny, whether overrides occur, and whether operational pressures are weakening agreed safeguards.
8. Maintain evidence currency
Policies, certificates, test reports and customer evidence age. Record expiry dates and refresh triggers for material evidence rather than assuming procurement-time documents remain current indefinitely.
9. Reassess regulatory and policy obligations
AI regulation, procurement policy and sector guidance continue to evolve. Periodically review whether new obligations affect the service, the buyer or the supplier.
10. Define re-approval triggers
Examples include:
- material model/provider change;
- new high-impact use case;
- processing of a new sensitive data class;
- significant security incident;
- recurring harmful failures;
- major subprocessor change;
- regulatory change;
- ownership or control change;
- substantial product redesign.
Monitoring record
Maintain a simple evidence trail covering: 1. service and current model/configuration; 2. owner and review date; 3. current performance indicators; 4. incidents and complaints; 5. supplier changes; 6. evidence refreshes; 7. residual risks; 8. actions and deadlines; 9. decision to continue, condition, restrict or re-evaluate.
Return to the AI Procurement Knowledge Base.