A Map of the AI Inference Stack

AI inference isn’t one market. It’s an entire computing stack. Here’s how I think about the ecosystem — from atoms to tokens:

1. Memory / Foundry / Advanced Packaging

Samsung (HBM, DRAM, Foundry, Advanced Packaging), SK hynix (HBM, DRAM), Micron (HBM, DRAM), TSMC (Foundry, CoWoS), Intel (Foundry, Advanced Packaging), ASE (Advanced Packaging), Amkor (Advanced Packaging)

2. General-Purpose AI Accelerators

NVIDIA (GPU, Grace, Blackwell/Rubin), AMD (Instinct GPU), Google (TPU), AWS (Trainium, Inferentia), Microsoft (Maia), Intel (Gaudi), Meta (MTIA)

3. Inference-Specialized Silicon

Cerebras (Wafer-Scale Engine), d-Matrix (Corsair / Raptor), SambaNova (RDU), Tenstorrent (AI accelerators), Etched (Transformer ASIC), FuriosaAI (RNGD), Rebellions (AI accelerators), Taalas (hard-wired model compute), Positron (inference accelerators), MatX (LLM accelerators), Axelera AI (AI accelerators), SiMa.ai (MLSoC), Untether AI (at-memory compute), EnCharge AI (analog in-memory compute), Fractile (inference silicon)

4. Edge / On-Device Inference

Apple (Neural Engine), Qualcomm (Hexagon / Snapdragon AI), Samsung (Exynos NPU), MediaTek (APU), Google (Tensor), Hailo (edge AI accelerators), Kneron (edge NPUs), Ambarella (edge AI SoCs), NXP (edge AI processors), STMicroelectronics (MCU/edge AI), Infineon (edge AI MCUs), Axelera AI (edge accelerators)

5. Networking / Interconnect / Optical

NVIDIA (NVLink, NVSwitch, Spectrum-X, InfiniBand), Broadcom (Ethernet switching, custom silicon), Marvell (interconnect, custom silicon), Arista (AI Ethernet networking), Cisco (AI networking), Astera Labs (PCIe/CXL connectivity), AMD / Pensando (DPUs), Lightmatter (photonic interconnect), Ayar Labs (optical I/O), Celestial AI (photonic fabric)

6. Kernels / Compilers / Low-Level Software

NVIDIA (CUDA, cuDNN, CUTLASS, TensorRT), AMD (ROCm, HIP), Google (XLA), Meta (PyTorch, TorchInductor), OpenAI (Triton), Intel (oneAPI), Modular (Mojo, MAX), OctoML / Apache TVM ecosystem (TVM)

7. Inference Engines / Runtimes

NVIDIA (TensorRT-LLM), Inferact / vLLM ecosystem (vLLM), SGLang ecosystem (SGLang), Microsoft (ONNX Runtime, DeepSpeed), Hugging Face (TGI), ggml ecosystem (llama.cpp)

8. Inference Platforms / Model Serving

Together AI (Inference, Together Kernel), Fireworks AI (optimized inference), Baseten (Inference Stack, Truss), Modal (serverless GPU infrastructure), DeepInfra (model inference APIs), Anyscale (Ray-based AI infrastructure), BentoML (BentoCloud, model serving), Hugging Face (Inference Endpoints), Replicate (model APIs), fal (generative media inference), Lightning AI (AI infrastructure)

9. GPU Clouds / Neoclouds

CoreWeave (GPU cloud), Lambda (GPU cloud), Nebius (AI cloud), Crusoe (AI cloud + data centers), Nscale (AI infrastructure), Fluidstack (AI cloud), RunPod (GPU cloud), Voltage Park (GPU cloud), GMI Cloud (GPU cloud), Vultr (GPU cloud), DigitalOcean (GPU cloud)

10. Hyperscale Cloud / Managed AI Infrastructure

AWS (EC2, Bedrock, SageMaker, Trainium/Inferentia), Microsoft (Azure AI), Google (Google Cloud, Vertex AI, TPU) Oracle (OCI AI infrastructure), Alibaba (Alibaba Cloud), Tencent (Tencent Cloud)

11. Frontier Model Labs / Massive-Scale Inference

OpenAI (GPT), Anthropic (Claude), Google DeepMind (Gemini), Meta (Llama), xAI (Grok), Mistral AI (Mistral), Cohere (Command), DeepSeek (DeepSeek), Alibaba (Qwen), Z.ai (GLM), MiniMax (MiniMax), Moonshot AI (Kimi)

12. Inference Gateway / RoutingOpenRouter (multi-model routing)

Cloudflare (AI Gateway, Workers AI), Vercel (AI Gateway / AI SDK), Portkey (AI Gateway), LiteLLM (LLM gateway / proxy)