Home / Categories / LLM Inference

LLM Inference

Showing 17 agents

See our top picks: Best LLM Inference AI Agents (2026) →
e
LLM Inference

exllama is a highly optimized implementation of the Llama language model for inference, specifically designed to work ef...

View Details Visit
F
LLM Inference

FastChat is an open-source platform for serving and interacting with multiple large language models (LLMs) efficiently. ...

View Details Visit
I
LLM Inference

Infinity is a Python-based tool designed for efficient text-embedding inference, enabling users to generate embeddings l...

View Details Visit
L
LLM Inference

Liger-Kernel provides optimized Triton kernels for efficient large language model (LLM) training, focusing on performanc...

View Details Visit
L
LLM Inference

LMDeploy is a high-performance inference and serving framework designed for large language models (LLMs) and vision-lang...

View Details Visit
M
LLM Inference

MInference is an open-source tool designed to optimize the inference process of long-context large language models (LLMs...

View Details Visit
m
LLM Inference

mistral.rs is a Rust-based implementation for efficient inference of Mistral language models, focusing on high performan...

View Details Visit
p
LLM Inference

prima.cpp is a distributed implementation of llama.cpp designed to enable efficient inference of large language models (...

View Details Visit
S
LLM Inference

SGLang is an open-source framework designed for efficient serving of large language models (LLMs) and vision language mo...

View Details Visit
S
LLM Inference

SkyPilot is an open-source framework for running large language models (LLMs) and batch jobs across multiple cloud provi...

View Details Visit
T
LLM Inference

TensorRT-LLM is an open-source library by NVIDIA designed to optimize and accelerate inference for large language models...

View Details Visit
T
LLM Inference

TGI (Text Generation Inference) is a toolkit designed for deploying and serving Large Language Models (LLMs) efficiently...

View Details Visit
v

vLLM

OSS
LLM Inference

vLLM is a high-performance inference engine designed for large language models (LLMs), optimized for throughput and memo...

View Details Visit