Evaluation platforms companies, products & suppliers

Workbenches to test model or app quality against datasets and graders. Jobs: offline evals; human grading; regression of quality across versions.

What are Evaluation platforms?

Workbenches to test model or app quality against datasets and graders. Jobs: offline evals; human grading; regression of quality across versions.

What problems does it solve?

Ship decisions rest on anecdotes and vendor leaderboards.

Typical business use cases

  • Offline evals
  • Human grading
  • Regression of quality across versions

Important capabilities

  • Datasets
  • Graders
  • Compare runs

What buyers should evaluate

  • Whether the vendor grades its own model
  • Dataset rights
  • Export

Risks and governance considerations

An eval suite that cannot be replayed independently.

Procurement checklist

  • Export of datasets
  • Grader identity
  • Replay

Relevant AI Trustmark assurance

AI Trustmark independent findings appear only when an assessment or certificate exists. Category membership does not imply verification.

Methodology · How verification works

Companies and providers

Claimed suppliers appear first so buyers can start with listings the company has taken ownership of. Payment does not buy this order.

  • Continuously improve AI agents with agent observability, evaluation, tracing, and experimentation. Arize AI publishes Arize AX Evaluation, Arize Phoenix, and Arize Model Monitoring

  • Ship quality agents at scale. Braintrust is the AI observability platform for tracing production, running evals, and catching regressions before they reach users. Braintrust Data,

  • Improve AI reliability with automated testing and monitoring with Citadel Lens and Citadel Radar. Citadel AI is used for Evaluation platforms work. Citadel AI publishes product inf

  • Confident AI is the AI quality platform for enterprise teams to standardize AI evals and observability across the org — one consistent bar for how every team measures and monitors

  • Freeplay is an evaluation and experimentation platform for teams building applications with large language models. It helps product and engineering teams test prompts and model beh

  • Future AGI publishes Future AGI as a named AI product. Future AGI is used for Evaluation platforms work. Future AGI publishes product information at futureagi.com. Future AGI is gr

  • Secure AI agents with Giskard’s continuous AI red teaming. Detect vulnerabilities, improve LLM security, and safeguard your AI systems. Giskard publishes Giskard Guards, Giskard Hu

  • LangChain enables every company to own their intelligence. Control, govern, and compound intelligence with an open agent engineering platform. Trusted by 7k+ organizations. LangCha

  • The fastest, most resilient, enterprise-grade LLM, MCP, and agent gateway. Routing, governance, guardrails, and observability via one control plane. Maxim AI publishes Bifrost AI G

  • Patronus AI develops simulation research and infrastructure to accelerate progress toward human-aligned AGI. Patronus AI publishes Lynx and Patronus AI Evaluator as named AI produc

  • From agentic skills to coding and AI safety — we build data solutions integrating human expertise and technology to accelerate AI development. Toloka publishes Toloka Platform as n

  • Vals AI Inc. sells evaluation suites that test language models on finance, legal and other domain tasks. Public pages describe benchmark runs and model comparisons, not voice cloni

Products

Claimed products appear first. Ranking packs and payment do not change this list.

Also used in this category

These products have a different primary category so they do not compete for the same ranking queries. They are listed here because buyers still encounter them in this job.

Related categories

Relevant procurement and assurance guides

Frequently asked questions

What is Evaluation platforms?

Workbenches to test model or app quality against datasets and graders. Jobs: offline evals; human grading; regression of quality across versions.

What should not be listed as Evaluation platforms?

Products whose buyer job is production monitoring, prompt editors, or red-teaming services. Those belong on their own category page so search queries are not split.

Has AI Trustmark independently assessed every Evaluation platforms supplier?

No. A category listing is descriptive. Independent assessment is shown only on company or product pages that carry Trustmark evidence.

What are Evaluation platforms?

Workbenches to test model or app quality against datasets and graders. Jobs: offline evals; human grading; regression of quality across versions.

What logging should an AI agent provide?

Buyers should be able to see who the agent acted as, which tool was called, what data was sent, what changed, and when. Logs that omit write actions or store raw customer prompts without access control are incomplete evidence.

What model or provider changes should a buyer insist on being told about?

Material change usually includes a new model family, new region, new subprocessor, new write-capable tool, or a change that affects logging, privacy or human oversight. Those changes should trigger evidence refresh rather than a silent release.

How should buyers verify where AI customer data is processed?

Ask for the named processing locations, cloud regions and any subprocessors that see prompts, files or outputs. A directory listing is not evidence of residency. Independent assessment records the locations that were in scope on the assessment date.

Does a TrustMark on one product cover the rest of the company?

No. Independent assessment is scoped to the named organisation and, where relevant, the named product. Category pages list suppliers as a topic label. They do not imply that every listed company has been assessed.