AI Infrastructure
51Companies categorized as AI Infrastructure.
Baseten offers an inference platform for deploying, managing, and scaling open-source and custom AI models in production. It is used by developers and engineers to run high‑performance model serving across cloud or on‑premises infrastructure.
Groq provides custom silicon processors and a cloud platform for AI inference, enabling developers to run models with low latency and reduced cost.
LangChain offers an engineering platform and open-source frameworks that developers use to build, test, and deploy AI agents.
OpenRouter provides a single API that lets developers access many large language models from different providers, handling routing, pricing, and data policies. It enables applications to send prompts to the chosen model without managing separate provider integrations.
Lambda provides cloud-based AI compute infrastructure with NVIDIA GPUs, offering on-demand instances and reserved clusters for training and inference workloads used by AI developers and research teams.
Crusoe operates AI infrastructure and cloud compute platforms that run AI workloads, offering managed inference and cloud services with an energy-first focus. Its platform is used by developers and enterprises to train and serve machine learning models.
E2B provides sandboxed cloud computers that AI agents can use to execute code, access the internet, and interact with real-world tools. It is used by enterprises and developers building agentic workflows.
Nebius operates an AI cloud platform that provides non-virtualized GPU hardware, storage, and MLOps tooling for training and inference workloads. It is used by AI developers and enterprises to build, scale, and deploy AI models.
Voltage Park provides on-demand and reserved GPU cloud infrastructure using NVIDIA H100, B200, and GB300 GPUs in Tier 3+ data centers for AI workloads. It serves AI research labs and startups that require compute resources.
Runpod provides on-demand GPU compute resources—including pods, serverless endpoints, and instant clusters—for developers and organizations to train, fine‑tune, and run AI models in the cloud. Users rent GPU capacity by the hour or second without owning hardware.
Weaviate is an open-source vector database that stores, indexes, and searches high-dimensional vectors at scale. It is used by developers and enterprises to build AI applications such as search, retrieval-augmented generation, and autonomous agents.
Hyperbolic provides a cloud platform that lets startups, researchers, and AI teams rent on-demand or reserved GPUs such as H100, H200, and B200 and access them via an OpenAI‑compatible inference API.
Massed Compute offers on-demand NVIDIA cloud GPU instances, including single-tenant bare-metal servers and GPU clusters. Customers can provision resources through a web console or API and are billed hourly for AI and other compute workloads.
Cirrascale offers a cloud platform for private AI training and inference that supports GPUs and other accelerators, providing managed services and high-bandwidth networking without data egress fees for enterprises and research institutions.
Radiant runs an integrated AI infrastructure platform that combines data center facilities, power, land, capital, compute hardware, and software to provide AI compute and cloud services. It serves technology companies and other organizations that require large‑scale AI compute resources.
Hugging Face operates a platform for developers and researchers to host, share, and collaborate on machine learning models, datasets, and applications. It also offers paid compute and enterprise services for model inference and team collaboration.
Novita AI offers an API that lets developers access over 200 AI models and launch on‑demand GPU instances and isolated agent sandboxes. It is aimed at developers and startups building AI applications.
Verda provides a cloud platform with GPU compute instances, storage, and serverless containers that developers and enterprises use to train and run AI models.
Mithril offers a platform that aggregates and orchestrates multi‑cloud GPU, CPU, and storage resources, letting users provision compute for machine‑learning workloads through a single interface with transparent pricing. It is used by AI startups, research labs, public companies, and academic institutions that need GPU compute.
TensorWave provides a cloud platform that offers bare‑metal access to AMD Instinct GPUs for training and inference of AI models. It is used by AI teams and enterprises that need high‑performance compute for large‑scale machine‑learning workloads.
Nestor provides dedicated GPU infrastructure for AI teams, offering managed or self‑run environments for inference, training, fine‑tuning, and reinforcement learning. It handles capacity, orchestration, networking, storage, and monitoring so teams can focus on models and data.
GMI Cloud is a cloud platform that lets AI teams run production inference workloads using serverless inference, dedicated GPU clusters, and bare metal GPU infrastructure. It provides APIs for large language and multimodal models and supports automatic scaling and cost management.
Domino Data Lab provides an enterprise AI platform that enables organizations to develop, deploy, and manage AI‑powered applications, offering tools for coding, governance, and collaboration for data science teams.
Vespa offers a distributed serving engine that integrates search, ranking, and machine-learned inference for AI applications. Developers and enterprises use it to build search, recommendation, and generative AI systems that handle billions of items and high query volumes.