Inference Frameworks
vLLM
PagedAttention-based LLM serving engine with continuous batching, speculative decoding, and an OpenAI-compatible API
PagedAttention-based LLM serving engine with continuous batching, speculative decoding, and an OpenAI-compatible API