DEV Community

#vllm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Can vLLM Run GGUF? Yes — on GPU Only

Can vLLM Run GGUF? Yes — on GPU Only

1
Comments
4 min read
Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU

Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU

Comments
13 min read
KV Cache on 16 GB GPUs: Making Long Context Actually Fit

KV Cache on 16 GB GPUs: Making Long Context Actually Fit

Comments 1
22 min read
From adapter to deployment: merging LoRA weights and serving with vLLM or a Space

From adapter to deployment: merging LoRA weights and serving with vLLM or a Space

Comments
3 min read
Inside vLLM: Following One Request from the API to GPU Execution

Inside vLLM: Following One Request from the API to GPU Execution

1
Comments 2
24 min read
Gemma 4 on a Tesla T4, Part 2: The Minimum GCE VM and a Script to Drive It

Gemma 4 on a Tesla T4, Part 2: The Minimum GCE VM and a Script to Drive It

13
Comments
13 min read
vLLM v0.28.0: the breaking change small GPU users must read

vLLM v0.28.0: the breaking change small GPU users must read

Comments
5 min read
Serving Gemma 4 on an AMD MI300X: What $1.99 an Hour Buys

Real-world benchmarks and MCP tool management tips

Serving Gemma 4 on an AMD MI300X: What $1.99 an Hour Buys

14
Comments 7
9 min read
Gemma 4 on a Tesla T4: QAT Weights Decode 1.79x Faster Than bf16

Gemma 4 on a Tesla T4: QAT Weights Decode 1.79x Faster Than bf16

8
Comments 1
9 min read
FastMCP Is Now MCPServer: Migrating a Python MCP Server to the MCP SDK 2.x

Dependency upgrades break shared environments

FastMCP Is Now MCPServer: Migrating a Python MCP Server to the MCP SDK 2.x

14
Comments 10
12 min read
Deploying the 600GB Inkling-NVFP4 Model on Spot A3: A GKE and vLLM Deep Dive

Deploying the 600GB Inkling-NVFP4 Model on Spot A3: A GKE and vLLM Deep Dive

5
Comments
61 min read
Ollama to vLLM: When to Migrate Your Local LLM Server

Ollama to vLLM: When to Migrate Your Local LLM Server

Comments
15 min read
Deploying Inference Using NVIDIA Dynamo and vLLM

Deploying Inference Using NVIDIA Dynamo and vLLM

13
Comments
8 min read
The Cheapest CUDA GPU on AWS Has an Arm CPU — and You Probably Want the Intel One

The Cheapest CUDA GPU on AWS Has an Arm CPU — and You Probably Want the Intel One

3
Comments
11 min read
DGX Spark (GB10) bare-metal vLLM: the install that works, two landmines, measured timings

DGX Spark (GB10) bare-metal vLLM: the install that works, two landmines, measured timings

1
Comments
2 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.