Inference Frameworks
TensorRT-LLM
nvidiaNVIDIA's fused-kernel LLM inference library — in-flight batching, FP8/INT4 quantization, and speculative decoding, optimized for Hopper/Blackwell
NVIDIA's fused-kernel LLM inference library — in-flight batching, FP8/INT4 quantization, and speculative decoding, optimized for Hopper/Blackwell