Inference Stack
The reference for how AI infrastructure actually fits together
Architectures
Catalog
Where it runs
About
Back to the interactive map
Inference Frameworks
llama.cpp
C/C++ LLM inference engine using GGUF-quantized weights, running efficiently on CPU and GPU
Official docs
Used by
Text Generation
Code Completion
Speech (ASR / TTS)
Leads to
Ray Serve