A high-throughput and memory-efficient inference and serving engine for LLMs
Pulse
56
Velocity
+54%
Stars
78,907
Extremely fast LLM inference and serving engine, optimized for throughput and memory.
A high-throughput and memory-efficient inference and serving engine for LLMs.
Serve LLMs at maximum speed with 2x higher throughput and significantly reduced memory usage.
Why Trending
Achieves 2x higher throughput than Hugging Face Transformers, significantly boosting LLM serving efficiency.
Target Audience
ML engineers and researchers deploying LLMs.
Similar Projects