AI SRE companies, products & suppliers

Incident and reliability products that diagnose production systems. Jobs: incident summarisation; runbook suggestions; root-cause hypotheses.

What is AI SRE?

Incident and reliability products that diagnose production systems. Jobs: incident summarisation; runbook suggestions; root-cause hypotheses.

What problems does it solve?

Incidents still start with a blank timeline and too many dashboards.

Typical business use cases

  • Incident summarisation
  • Runbook suggestions
  • Root-cause hypotheses

Important capabilities

  • Telemetry connectors
  • On-call context
  • Timeline

What buyers should evaluate

  • Access to production telemetry
  • Whether it can act
  • Retention of incident data

Risks and governance considerations

An SRE bot that restarts services or pages the wrong people.

Procurement checklist

  • Read vs act
  • Paging policy
  • Incident data retention

Relevant AI Trustmark assurance

AI Trustmark independent findings appear only when an assessment or certificate exists. Category membership does not imply verification.

Methodology · How verification works

Companies and providers

Claimed suppliers appear first so buyers can start with listings the company has taken ownership of. Payment does not buy this order.

  • Amazon Web Services is Amazon's cloud division selling compute, storage, Bedrock model hosting and related AI infrastructure. Public pages cover regional cloud services for builder

  • Discover how agentic AI in IT operations helps teams reduce response times & improve incident management. BigPanda publishes BigPanda AIOps as named AI products. BigPanda is used f

  • Ciroos is an AI SRE solution that delivers root cause clarity and SRE toil reduction, and delivers measurable ROI in complex enterprise environments. Ciroos publishes product infor

  • Harness is a unified AI software delivery platform to manage the SDLC using purpose-built AI agents. Harness publishes Harness AI SRE and Harness AIDA as named AI products. Harness

  • International Business Machines Corporation sells hybrid-cloud software, Granite models and enterprise AI tooling. Public pages distinguish IBM software and models from IBM Consult

  • The open-source AIOps (AI for IT operations) and alert management platform. Keep is used for AI SRE work. Keep publishes product information at keephq.dev. Keep is grouped with AI

  • Microsoft Corporation sells Windows, Azure, Microsoft 365 and Copilot AI products. Public pages cover cloud, productivity and developer APIs used by enterprises and consumers. Micr

  • PagerDuty's AI-first Operations Platform helps you automate, resolve, and prevent critical issues — from detection to resolution. PagerDuty publishes PagerDuty Advance and PagerDut

  • Resolve AI publishes Resolve AI as a named AI product. Resolve AI is used for AI SRE and DevOps assistants work. Resolve AI is recorded in San Francisco, United States. Resolve AI

Products

Claimed products appear first. Ranking packs and payment do not change this list.

Related categories

Relevant procurement and assurance guides

Frequently asked questions

What is AI SRE?

Incident and reliability products that diagnose production systems. Jobs: incident summarisation; runbook suggestions; root-cause hypotheses.

What should not be listed as AI SRE?

Products whose buyer job is DevOps pipeline assistants, observability platforms, or generic coding assistants. Those belong on their own category page so search queries are not split.

Has AI Trustmark independently assessed every AI SRE supplier?

No. A category listing is descriptive. Independent assessment is shown only on company or product pages that carry Trustmark evidence.

Should coding assistants be allowed to send private source code to third-party models?

That is a buyer policy decision. Procurement should require the default data-use terms, retention, whether snippets leave the organisation, and whether secrets are redacted. A listing under coding assistants does not mean code stays on-premises.

How should prompt injection and tool-output attacks be controlled?

Agents that read untrusted content or tool output can be instructed to exfiltrate data or take writes. Buyers should ask what is treated as untrusted, whether tool output can change the plan, and what tests were run. Scanner marketing is not the same as independent testing.

How can buyers tell whether customer data is used to train models?

Ask whether prompts, files, logs or outputs are used to train, fine-tune or evaluate models, including by subprocessors. Require the contractual default, any opt-out, and whether the setting can be changed silently. Treat marketing 'we do not train' claims as unverified until evidenced.

What model or provider changes should a buyer insist on being told about?

Material change usually includes a new model family, new region, new subprocessor, new write-capable tool, or a change that affects logging, privacy or human oversight. Those changes should trigger evidence refresh rather than a silent release.

How should buyers verify where AI customer data is processed?

Ask for the named processing locations, cloud regions and any subprocessors that see prompts, files or outputs. A directory listing is not evidence of residency. Independent assessment records the locations that were in scope on the assessment date.