Back to the interactive map
Inference Frameworks

llama.cpp

C/C++ LLM inference engine using GGUF-quantized weights, running efficiently on CPU and GPU