DeepSpeed-Mii
OSSDeepSpeed-Mii is an open-source library designed to optimize large language model (LLM) inference by enabling low-latenc...
Comparable LLM Inference AI agents, ranked by popularity. Not sold on TensorRT-LLM? These are the closest options worth a look.
You're comparing against
TensorRT-LLM is an open-source library by NVIDIA designed to optimize and accelerate inference for large language models (LLMs) using TensorRT. It provides tools for efficient model deployment, including quantization, kernel optimization, and batch processing. The framework supports popular LLM architectures and is tailored for high-performance inference on NVIDIA GPUs.
DeepSpeed-Mii is an open-source library designed to optimize large language model (LLM) inference by enabling low-latenc...
deploy-llms-with-ansible is an open-source Ansible playbook that automates the deployment of large language models (LLMs...
exllama is a highly optimized implementation of the Llama language model for inference, specifically designed to work ef...
FastChat is an open-source platform for serving and interacting with multiple large language models (LLMs) efficiently. ...
FasterTransformer is an open-source library developed by NVIDIA for accelerating transformer-based large language model ...
Infinity is a Python-based tool designed for efficient text-embedding inference, enabling users to generate embeddings l...
Liger-Kernel provides optimized Triton kernels for efficient large language model (LLM) training, focusing on performanc...
LMDeploy is a high-performance inference and serving framework designed for large language models (LLMs) and vision-lang...
MInference is an open-source tool designed to optimize the inference process of long-context large language models (LLMs...
mistral.rs is a Rust-based implementation for efficient inference of Mistral language models, focusing on high performan...
prima.cpp is a distributed implementation of llama.cpp designed to enable efficient inference of large language models (...
SGLang is an open-source framework designed for efficient serving of large language models (LLMs) and vision language mo...