Back to the interactive map
Inference Frameworks

TensorRT-LLM

nvidia

NVIDIA's fused-kernel LLM inference library — in-flight batching, FP8/INT4 quantization, and speculative decoding, optimized for Hopper/Blackwell