Experiment management companies, products & suppliers
Tracking training runs, metrics and artefacts for ML experiments. Jobs: run tracking; artefact store; compare experiments.
What is Experiment management?
Tracking training runs, metrics and artefacts for ML experiments. Jobs: run tracking; artefact store; compare experiments.
What problems does it solve?
Training results live in notebooks with no comparable history.
Typical business use cases
- Run tracking
- Artefact store
- Compare experiments
Important capabilities
- Metrics log
- Artefacts
- Collaboration
What buyers should evaluate
- Who can see experiments
- Whether datasets are uploaded
- Retention
Risks and governance considerations
Training datasets copied into the tracker.
Procurement checklist
- Dataset policy
- Access
- Retention
Relevant AI Trustmark assurance
AI Trustmark independent findings appear only when an assessment or certificate exists. Category membership does not imply verification.
Companies and providers
Claimed suppliers appear first so buyers can start with listings the company has taken ownership of. Payment does not buy this order.
Anaconda is the trusted foundation for AI-native development. Secure, orchestrate, and accelerate data and AI at scale, from first experiment to production. Anaconda publishes Meta
Powered by Ray, Anyscale helps AI builders run data-intensive workloads to build and deploy Foundation Models and AI at scale on any cloud. Anyscale publishes Anyscale Platform and
Curate and annotate vision, audio, and LLM datasets, track experiments, and manage models on a single platform. DagsHub is used for Experiment management work. DagsHub publishes pr
Mosaic AI on Databricks unifies model building, evaluation, serving, and governance on the lakehouse. Public marketing also emphasises lakehouse-native GenAI/ML. It is presented fo
Google is a US technology company that sells consumer and cloud AI products, including Gemini, Search, and related developer APIs on Google Cloud. Google publishes Imagen, Gemini f
Hugging Face operates a model hub, inference tooling and collaboration products for machine-learning teams. Public pages cover hosting, inference endpoints and open-model distribut
Organize machine learning experiments and monitor training progress from mobile. LabML is used for Experiment management work. LabML publishes product information at labml.ai. LabM
Ensure AI data governance, make AI training and agent runs reproducible, reduce access friction. lakeFS is used for Experiment management work. lakeFS publishes product information
Helping open technology projects build world class open source software, communities and companies. Linux Foundation publishes Monocle, vLLM, and Kubeflow as named AI products. Lin
neptune.ai provides experiment tracking and model metadata management for ML teams. neptune.ai is used for Experiment management work. neptune.ai is recorded in Poland. neptune.ai
AI control plane to schedule, run, debug, observe, and automate AI/ML workloads, agents, and sandboxes on your infra. Polyaxon is used for Experiment management work. Polyaxon publ
Learn why thousands of companies rely on W&B as their system of record for training AI models and developing AI applications with confidence. Weights & Biases publishes W&B Serverl
Products
Claimed products appear first. Ranking packs and payment do not change this list.
- DagsHub· DagsHub
- Flyte· Linux Foundation
- LabML· LabML
- lakeFS· lakeFS
- Metaflow· Anaconda
- neptune.ai· neptune.ai
- Polyaxon· Polyaxon
- Ray Tune· Anyscale
- TensorBoard· Google
- Trackio· Hugging Face
- W&B Experiments· Weights & Biases
Also used in this category
These products have a different primary category so they do not compete for the same ranking queries. They are listed here because buyers still encounter them in this job.
- MLflow· Databricks
Related categories
Relevant procurement and assurance guides
Frequently asked questions
What is Experiment management?
Tracking training runs, metrics and artefacts for ML experiments. Jobs: run tracking; artefact store; compare experiments.
What should not be listed as Experiment management?
Products whose buyer job is model monitoring in production, LLMOps, or evaluation platforms for apps. Those belong on their own category page so search queries are not split.
Has AI Trustmark independently assessed every Experiment management supplier?
No. A category listing is descriptive. Independent assessment is shown only on company or product pages that carry Trustmark evidence.
What logging should an AI agent provide?
Buyers should be able to see who the agent acted as, which tool was called, what data was sent, what changed, and when. Logs that omit write actions or store raw customer prompts without access control are incomplete evidence.
What model or provider changes should a buyer insist on being told about?
Material change usually includes a new model family, new region, new subprocessor, new write-capable tool, or a change that affects logging, privacy or human oversight. Those changes should trigger evidence refresh rather than a silent release.
How should buyers verify where AI customer data is processed?
Ask for the named processing locations, cloud regions and any subprocessors that see prompts, files or outputs. A directory listing is not evidence of residency. Independent assessment records the locations that were in scope on the assessment date.
Does a TrustMark on one product cover the rest of the company?
No. Independent assessment is scoped to the named organisation and, where relevant, the named product. Category pages list suppliers as a topic label. They do not imply that every listed company has been assessed.
What incident-handling evidence is useful for AI suppliers?
Buyers should see how AI-specific failures are detected, contained and notified — including unsafe outputs, data leakage and unauthorised agent actions. An incident policy that never mentions models, prompts or tools is incomplete for this class of product.