vibehacker

vLLM

High-throughput, memory-efficient LLM inference and serving

by vLLM communityCoding & Dev ToolsAgents & Automation Listed Jun 1, 2023

Product badge

Share this product's name and rating in your README or on your website.

vLLM: rating on VibeHacker

Updates may be delayed by image caching.

vLLM cover
vLLM screenshot 2vLLM screenshot 3vLLM screenshot 4

About vLLM

vLLM is an inference and serving engine for large language models (LLMs). It is intended for developers and teams deploying models across different hardware platforms.

The engine provides a drop-in OpenAI-compatible API and supports a range of hardware, including NVIDIA CUDA GPUs, AMD ROCm GPUs, AWS Neuron accelerators, Google Cloud TPUs, Apple Silicon, and others. PagedAttention, advanced scheduling, and continuous batching are used to maximize throughput and GPU utilization.

vLLM can be installed with Python or Docker. The quick-start instructions require Python 3.10 or newer, with Python 3.12 or newer recommended; stable and nightly builds are available.

Used vLLM?

Log in to write a review.

Similar tools

View all