Model hosting / inference companies, products & suppliers
Model hosting and inference platforms run someone else's or your models as an API with scaling, isolation and observability. Buyers are purchasing infrastructure, not a chatbot. Tenancy, region and which models are actually served are the entity facts that matter.
What is Model hosting / inference?
Model hosting and inference platforms run someone else's or your models as an API with scaling, isolation and observability. Buyers are purchasing infrastructure, not a chatbot. Tenancy, region and which models are actually served are the entity facts that matter.
What problems does it solve?
Teams cannot serve models with isolation, regional control and cost limits.
Typical business use cases
- Production inference
- Private model APIs
- Burst GPU
Important capabilities
- Autoscaling
- VPC
- Routing
- Observability
What buyers should evaluate
- Tenancy
- Regions
- Supported runtimes
- Keys
Risks and governance considerations
Cross-tenant prompt leakage.
Procurement checklist
- Isolation evidence
- Regions
- KMS
- SLAs
Relevant AI Trustmark assurance
Hosting quality is operational evidence, not a model-quality score.
Companies and providers
Claimed suppliers appear first so buyers can start with listings the company has taken ownership of. Payment does not buy this order.
Amazon Web Services is Amazon's cloud division selling compute, storage, Bedrock model hosting and related AI infrastructure. Public pages cover regional cloud services for builder
Serve and scale open-source and custom AI models on the fastest, most reliable inference platform. Baseten publishes Baseten Model APIs as named AI products. Baseten is used for Mo
Cerebras powers the world's fastest AI inference on the biggest wafer chip. Cerebras CS-4 delivers up to 30x faster inference than GPUs. Cerebras publishes Cerebras Inference and C
Cloudflare publishes Cloudflare DLP, Cloudflare Agents SDK MCP Support, and Cloudflare Workers AI as named AI products. Cloudflare is used for Agent integration / MCP and AI agents
DeepInfra offers cost-effective, scalable, production-ready ML model hosting and inference infrastructure. DeepInfra is used for Model hosting / inference work. DeepInfra is record
fal publishes fal Serverless as named AI products. fal is used for Model hosting / inference work. fal publishes product information at fal.ai. fal is grouped with Model hosting /
Fireworks’ state of the art training and inference platform take you beyond the frontier, transforming open models into your specialized intelligence. Fireworks AI publishes Firewo
Groq is the premier neocloud for fast inference. One fully integrated platform for infrastructure, inference, and control. Millions of developers run trillions of tokens on Groq ev
Hugging Face operates a model hub, inference tooling and collaboration products for machine-learning teams. Public pages cover hosting, inference endpoints and open-model distribut
Hyperbolic provides GPU cloud and inference services aimed at affordable access to AI compute and open models. Hyperbolic Labs.xyz, Model hosting / inference. Hyperbolic is recorde
Microsoft Corporation sells Windows, Azure, Microsoft 365 and Copilot AI products. Public pages cover cloud, productivity and developer APIs used by enterprises and consumers. Micr
Novita AI markets GPU cloud and model APIs for generative AI workloads. Novita AI is used for Model hosting / inference work. Novita AI is recorded in Singapore. Novita AI publishe
NVIDIA sells GPUs, CUDA software and AI enterprise stacks used for training and inference. Public pages cover data-centre, cloud and on-prem AI compute. NVIDIA publishes NVIDIA cuO
Find the best models & prices for your prompts. OpenRouter is used for Model routing and Model hosting / inference work. OpenRouter is recorded in New York, United States. OpenRout
Run open-source machine learning models with a cloud API. Replicate publishes Replicate Training API and Replicate API as named AI products. Replicate is used for Model hosting / i
Full-stack AI platform for inference, fine-tuning, and GPU clusters — powered by cutting-edge research. Together AI publishes Together Fine-Tuning and Together Inference as named A
Products
Claimed products appear first. Ranking packs and payment do not change this list.
- Baseten Model APIs· Baseten
- Cerebras Inference· Cerebras
- Cloudflare Workers AI· Cloudflare
- DeepInfra· DeepInfra
- fal Serverless· fal
- Fireworks Serverless Inference· Fireworks AI
- GroqCloud· Groq
- Hugging Face Inference Endpoints· Hugging Face
- Hyperbolic· Hyperbolic
- Microsoft Foundry Models· Microsoft
- Novita AI· Novita AI
- NVIDIA NIM / NeMo· NVIDIA
- Replicate API· Replicate
- Together Inference· Together AI
Also used in this category
These products have a different primary category so they do not compete for the same ranking queries. They are listed here because buyers still encounter them in this job.
- Amazon Bedrock· Amazon Web Services
- NVIDIA AI Enterprise· NVIDIA
- OpenRouter· OpenRouter
Related categories
Relevant procurement and assurance guides
Frequently asked questions
Is this an LLM vendor?
Not necessarily. Many hosts serve third-party or customer weights.
What is Model hosting / inference?
Model hosting and inference platforms run someone else's or your models as an API with scaling, isolation and observability. Buyers are purchasing infrastructure, not a chatbot. Tenancy, region and which models are actually served are the entity facts that matter.
What model or provider changes should a buyer insist on being told about?
Material change usually includes a new model family, new region, new subprocessor, new write-capable tool, or a change that affects logging, privacy or human oversight. Those changes should trigger evidence refresh rather than a silent release.
How should buyers verify where AI customer data is processed?
Ask for the named processing locations, cloud regions and any subprocessors that see prompts, files or outputs. A directory listing is not evidence of residency. Independent assessment records the locations that were in scope on the assessment date.
Does a TrustMark on one product cover the rest of the company?
No. Independent assessment is scoped to the named organisation and, where relevant, the named product. Category pages list suppliers as a topic label. They do not imply that every listed company has been assessed.
What incident-handling evidence is useful for AI suppliers?
Buyers should see how AI-specific failures are detected, contained and notified — including unsafe outputs, data leakage and unauthorised agent actions. An incident policy that never mentions models, prompts or tools is incomplete for this class of product.
What security testing evidence should buyers request for an AI product?
Ask what was tested, against which version, whether prompt-injection, data-exfiltration and tenant isolation were in scope, and where failed prompts were stored. A generic ISO certificate or a vendor scanner screenshot is not by itself an AI TrustMark assessment.
How should buyers check AI subprocessors?
Ask for the named subprocessors that process prompts, files, embeddings or logs, their locations, and whether a change is notified. Related-organisation hosting should be disclosed. A privacy policy URL is not a current subprocessor list.