DEV Community

#vllm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

10
Comments
9 min read
Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

Comments
9 min read
Ollama to vLLM: When to Migrate Your Local LLM Server

Ollama to vLLM: When to Migrate Your Local LLM Server

Comments
15 min read
The unofficial TPU migration guide: Cloud TPU API to Compute Engine

The unofficial TPU migration guide: Cloud TPU API to Compute Engine

6
Comments 2
17 min read
Self-Hosted Gemma 4 on TPU v6e: Deployment & SRE with Antigravity

Self-Hosted Gemma 4 on TPU v6e: Deployment & SRE with Antigravity

Comments
8 min read
Self-hosting a lite agent backend on one TPU: Gemma 4 E2B + vLLM on a v5e-1

Self-hosting a lite agent backend on one TPU: Gemma 4 E2B + vLLM on a v5e-1

15
Comments 1
21 min read
What 90% Line-Rate Utilization on a Single 100GbE Port Means: Analyzing Network Bottlenecks in Inference Storage

What 90% Line-Rate Utilization on a Single 100GbE Port Means: Analyzing Network Bottlenecks in Inference Storage

Comments
5 min read
Serving Gemma 4 E2B on a TPU v6e-1: what Trillium buys, and what it doesn't

Serving Gemma 4 E2B on a TPU v6e-1: what Trillium buys, and what it doesn't

2
Comments
20 min read
Does a Second GPU Increase Ollama's Context Window? (Quadro P2000 + RTX 3090 Tested)

Does a Second GPU Increase Ollama's Context Window? (Quadro P2000 + RTX 3090 Tested)

Comments
3 min read
vLLM vs llama.cpp vs Ollama: What Happens When Your Model Doesn't Fit in 24GB VRAM

vLLM vs llama.cpp vs Ollama: What Happens When Your Model Doesn't Fit in 24GB VRAM

Comments
6 min read
Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive

Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive

8
Comments 1
8 min read
tpu-management: a Claude Code skill for running Gemma 4 on Cloud TPUs

tpu-management: a Claude Code skill for running Gemma 4 on Cloud TPUs

7
Comments
3 min read
AI Inference Optimization in Cloud-Native Environments: GPU Orchestration, Edge Deployment, and Latency Reduction at Scale

AI Inference Optimization in Cloud-Native Environments: GPU Orchestration, Edge Deployment, and Latency Reduction at Scale

Comments
4 min read
Kimi K3 Open Weights Are Here: How to Self-Host the 2.8T-Parameter Model (Hardware, vLLM, and Data Sovereignty)

Kimi K3 Open Weights Are Here: How to Self-Host the 2.8T-Parameter Model (Hardware, vLLM, and Data Sovereignty)

4
Comments
8 min read
AI Inference at the Edge: Running Real-Time LLMs in Kubernetes Without a GPU Farm

AI Inference at the Edge: Running Real-Time LLMs in Kubernetes Without a GPU Farm

Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.