AI inference isn’t one market. It’s an entire computing stack. Here’s how I think about the ecosystem — from atoms to tokens:
1. Memory / Foundry / Advanced Packaging
Samsung (HBM, DRAM, Foundry, Advanced Packaging), SK hynix (HBM, DRAM), Micron (HBM, DRAM), TSMC (Foundry, CoWoS), Intel (Foundry, Advanced Packaging), ASE (Advanced Packaging), Amkor (Advanced Packaging)
2. General-Purpose AI Accelerators
NVIDIA (GPU, Grace, Blackwell/Rubin), AMD (Instinct GPU), Google (TPU), AWS (Trainium, Inferentia), Microsoft (Maia), Intel (Gaudi), Meta (MTIA)
3. Inference-Specialized Silicon
Cerebras (Wafer-Scale Engine), d-Matrix (Corsair / Raptor), SambaNova (RDU), Tenstorrent (AI accelerators), Etched (Transformer ASIC), FuriosaAI (RNGD), Rebellions (AI accelerators), Taalas (hard-wired model compute), Positron (inference accelerators), MatX (LLM accelerators), Axelera AI (AI accelerators), SiMa.ai (MLSoC), Untether AI (at-memory compute), EnCharge AI (analog in-memory compute), Fractile (inference silicon)
4. Edge / On-Device Inference
Apple (Neural Engine), Qualcomm (Hexagon / Snapdragon AI), Samsung (Exynos NPU), MediaTek (APU), Google (Tensor), Hailo (edge AI accelerators), Kneron (edge NPUs), Ambarella (edge AI SoCs), NXP (edge AI processors), STMicroelectronics (MCU/edge AI), Infineon (edge AI MCUs), Axelera AI (edge accelerators)
5. Networking / Interconnect / Optical
NVIDIA (NVLink, NVSwitch, Spectrum-X, InfiniBand), Broadcom (Ethernet switching, custom silicon), Marvell (interconnect, custom silicon), Arista (AI Ethernet networking), Cisco (AI networking), Astera Labs (PCIe/CXL connectivity), AMD / Pensando (DPUs), Lightmatter (photonic interconnect), Ayar Labs (optical I/O), Celestial AI (photonic fabric)
6. Kernels / Compilers / Low-Level Software
NVIDIA (CUDA, cuDNN, CUTLASS, TensorRT), AMD (ROCm, HIP), Google (XLA), Meta (PyTorch, TorchInductor), OpenAI (Triton), Intel (oneAPI), Modular (Mojo, MAX), OctoML / Apache TVM ecosystem (TVM)
7. Inference Engines / Runtimes
NVIDIA (TensorRT-LLM), Inferact / vLLM ecosystem (vLLM), SGLang ecosystem (SGLang), Microsoft (ONNX Runtime, DeepSpeed), Hugging Face (TGI), ggml ecosystem (llama.cpp)
8. Inference Platforms / Model Serving
Together AI (Inference, Together Kernel), Fireworks AI (optimized inference), Baseten (Inference Stack, Truss), Modal (serverless GPU infrastructure), DeepInfra (model inference APIs), Anyscale (Ray-based AI infrastructure), BentoML (BentoCloud, model serving), Hugging Face (Inference Endpoints), Replicate (model APIs), fal (generative media inference), Lightning AI (AI infrastructure)
9. GPU Clouds / Neoclouds
CoreWeave (GPU cloud), Lambda (GPU cloud), Nebius (AI cloud), Crusoe (AI cloud + data centers), Nscale (AI infrastructure), Fluidstack (AI cloud), RunPod (GPU cloud), Voltage Park (GPU cloud), GMI Cloud (GPU cloud), Vultr (GPU cloud), DigitalOcean (GPU cloud)
10. Hyperscale Cloud / Managed AI Infrastructure
AWS (EC2, Bedrock, SageMaker, Trainium/Inferentia), Microsoft (Azure AI), Google (Google Cloud, Vertex AI, TPU) Oracle (OCI AI infrastructure), Alibaba (Alibaba Cloud), Tencent (Tencent Cloud)
11. Frontier Model Labs / Massive-Scale Inference
OpenAI (GPT), Anthropic (Claude), Google DeepMind (Gemini), Meta (Llama), xAI (Grok), Mistral AI (Mistral), Cohere (Command), DeepSeek (DeepSeek), Alibaba (Qwen), Z.ai (GLM), MiniMax (MiniMax), Moonshot AI (Kimi)
12. Inference Gateway / RoutingOpenRouter (multi-model routing)
Cloudflare (AI Gateway, Workers AI), Vercel (AI Gateway / AI SDK), Portkey (AI Gateway), LiteLLM (LLM gateway / proxy)