Back to RepoInsider
Signal dossier

vllm-project/vllm

PythonAI Tools

A high-throughput and memory-efficient inference and serving engine for LLMs

Pulse

56

Velocity

+54%

Stars

78,907

Best forAmdBlackwellCuda

Extremely fast LLM inference and serving engine, optimized for throughput and memory.

A high-throughput and memory-efficient inference and serving engine for LLMs.

Serve LLMs at maximum speed with 2x higher throughput and significantly reduced memory usage.

Why Trending

Achieves 2x higher throughput than Hugging Face Transformers, significantly boosting LLM serving efficiency.

Target Audience

ML engineers and researchers deploying LLMs.

Similar Projects

Text Generation InferenceOpenLLM
RepoInsider · Spot breakout GitHub repos early