Inference Stack
The reference for how AI infrastructure actually fits together
Architectures
Catalog
Where it runs
About
Back to the interactive map
Orchestration
Ray Serve
anyscale
Scalable model serving built on Ray with request routing and autoscaling
Official docs
Used by
vLLM
SGLang
llama.cpp
HF Diffusers
Faster Whisper
Leads to
CUDA
ROCm
XLA / JAX