Agent observability companies, products & suppliers

Traces and evals for tool-using agents, not generic LLM logs. Jobs: tool-call traces; trajectory eval; agent session replay.

What is Agent observability?

Traces and evals for tool-using agents, not generic LLM logs. Jobs: tool-call traces; trajectory eval; agent session replay.

What problems does it solve?

Agent actions are invisible, so failures look like 'the model was wrong'.

Typical business use cases

  • Tool-call traces
  • Trajectory eval
  • Agent session replay

Important capabilities

  • Span of tools
  • Session grouping
  • Action diffs

What buyers should evaluate

  • Whether actions are reconstructed
  • PII in tool payloads
  • Retention

Risks and governance considerations

Tool payloads that include customer records stored as traces.

Procurement checklist

  • Payload redaction
  • Retention
  • Who can replay

Relevant AI Trustmark assurance

AI Trustmark independent findings appear only when an assessment or certificate exists. Category membership does not imply verification.

Methodology · How verification works

Companies and providers

Claimed suppliers appear first so buyers can start with listings the company has taken ownership of. Payment does not buy this order.

  • AgentOps publishes AgentOps as a named AI product. AgentOps is used for Agent observability work. AgentOps publishes product information at agentops.ai. AgentOps is grouped with Ag

  • Continuously improve AI agents with agent observability, evaluation, tracing, and experimentation. Arize AI publishes Arize AX Evaluation, Arize Phoenix, and Arize Model Monitoring

  • Athina AI Inc. sells evaluation and monitoring for LLM applications, including traces and datasets. Public pages describe production quality checks, not image remixing. Athina is u

  • HoneyHive helps enterprises trace, evaluate, monitor, and improve production AI agents across teams, stacks, and sensitive data environments. HoneyHive is used for Agent observabil

  • LangChain enables every company to own their intelligence. Control, govern, and compound intelligence with an open agent engineering platform. Trusted by 7k+ organizations. LangCha

  • Langtrace publishes Langtrace as a named AI product. Langtrace is used for Agent observability work. Langtrace publishes product information at docs.langtrace.ai. Langtrace is grou

  • Unified observability for the whole team — developers, SREs, and AI agents. Single view for metrics, logs, traces, and APM. Last9 is used for Agent observability work. Last9 publis

  • Middleware offers full-stack observability with real-time monitoring and diagnostics, helping teams detect issues at scale while keeping data secure. Middleware is used for Agent o

  • OpenObserve publishes OpenObserve as a named AI product. OpenObserve is used for Agent observability work. OpenObserve publishes product information at openobserve.ai. OpenObserve

  • Empowering healthcare leaders with purpose-built AI Agent Suites for Radiology, Allergy, and Clinical Trials. RagaAI publishes RagaAI Catalyst as named AI products. RagaAI is used

  • Snowflake Cortex brings LLM and ML functions into the Snowflake Data Cloud for governed SQL/AI analytics. Public marketing also emphasises aI in-warehouse with Snowflake governance

Products

Claimed products appear first. Ranking packs and payment do not change this list.

Also used in this category

These products have a different primary category so they do not compete for the same ranking queries. They are listed here because buyers still encounter them in this job.

Related categories

Relevant procurement and assurance guides

Frequently asked questions

What is Agent observability?

Traces and evals for tool-using agents, not generic LLM logs. Jobs: tool-call traces; trajectory eval; agent session replay.

What should not be listed as Agent observability?

Products whose buyer job is AI observability without tools, LLMOps suites, or security audit logs. Those belong on their own category page so search queries are not split.

Has AI Trustmark independently assessed every Agent observability supplier?

No. A category listing is descriptive. Independent assessment is shown only on company or product pages that carry Trustmark evidence.

What logging should an AI agent provide?

Buyers should be able to see who the agent acted as, which tool was called, what data was sent, what changed, and when. Logs that omit write actions or store raw customer prompts without access control are incomplete evidence.

What model or provider changes should a buyer insist on being told about?

Material change usually includes a new model family, new region, new subprocessor, new write-capable tool, or a change that affects logging, privacy or human oversight. Those changes should trigger evidence refresh rather than a silent release.

How should buyers verify where AI customer data is processed?

Ask for the named processing locations, cloud regions and any subprocessors that see prompts, files or outputs. A directory listing is not evidence of residency. Independent assessment records the locations that were in scope on the assessment date.

Does a TrustMark on one product cover the rest of the company?

No. Independent assessment is scoped to the named organisation and, where relevant, the named product. Category pages list suppliers as a topic label. They do not imply that every listed company has been assessed.

What incident-handling evidence is useful for AI suppliers?

Buyers should see how AI-specific failures are detected, contained and notified — including unsafe outputs, data leakage and unauthorised agent actions. An incident policy that never mentions models, prompts or tools is incomplete for this class of product.