Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
vllm
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU
xbill
xbill
xbill
Follow
for
Google Developer Experts
Aug 13
Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU
#
aws
#
vllm
#
cuda
#
machinelearning
10
 reactions
Comments
Add Comment
9 min read
Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU
xbill
xbill
xbill
Follow
for
AWS Community Builders
Aug 13
Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU
#
aws
#
vllm
#
cuda
#
machinelearning
Comments
Add Comment
9 min read
Ollama to vLLM: When to Migrate Your Local LLM Server
Rost
Rost
Rost
Follow
Aug 2
Ollama to vLLM: When to Migrate Your Local LLM Server
#
ollama
#
vllm
#
llm
#
ai
Comments
Add Comment
15 min read
The unofficial TPU migration guide: Cloud TPU API to Compute Engine
xbill
xbill
xbill
Follow
for
Google Developer Experts
Aug 11
The unofficial TPU migration guide: Cloud TPU API to Compute Engine
#
tpu
#
gcp
#
vllm
#
devops
6
 reactions
Comments
2
 comments
17 min read
Self-Hosted Gemma 4 on TPU v6e: Deployment & SRE with Antigravity
xbill
xbill
xbill
Follow
for
Google Developer Experts
Jul 25
Self-Hosted Gemma 4 on TPU v6e: Deployment & SRE with Antigravity
#
tpu
#
llm
#
vllm
#
antigravity
Comments
Add Comment
8 min read
Self-hosting a lite agent backend on one TPU: Gemma 4 E2B + vLLM on a v5e-1
xbill
xbill
xbill
Follow
for
Google Developer Experts
Aug 9
Self-hosting a lite agent backend on one TPU: Gemma 4 E2B + vLLM on a v5e-1
#
tpu
#
vllm
#
llm
#
gcp
15
 reactions
Comments
1
 comment
21 min read
What 90% Line-Rate Utilization on a Single 100GbE Port Means: Analyzing Network Bottlenecks in Inference Storage
Mingxin Technology
Mingxin Technology
Mingxin Technology
Follow
Jul 22
What 90% Line-Rate Utilization on a Single 100GbE Port Means: Analyzing Network Bottlenecks in Inference Storage
#
kvcache
#
lmcache
#
vllm
#
ai
Comments
Add Comment
5 min read
Serving Gemma 4 E2B on a TPU v6e-1: what Trillium buys, and what it doesn't
xbill
xbill
xbill
Follow
for
Google Developer Experts
Aug 11
Serving Gemma 4 E2B on a TPU v6e-1: what Trillium buys, and what it doesn't
#
tpu
#
vllm
#
llm
#
gcp
2
 reactions
Comments
Add Comment
20 min read
Does a Second GPU Increase Ollama's Context Window? (Quadro P2000 + RTX 3090 Tested)
Arsen Apostolov
Arsen Apostolov
Arsen Apostolov
Follow
Jul 9
Does a Second GPU Increase Ollama's Context Window? (Quadro P2000 + RTX 3090 Tested)
#
llm
#
ollama
#
vllm
#
gpu
Comments
Add Comment
3 min read
vLLM vs llama.cpp vs Ollama: What Happens When Your Model Doesn't Fit in 24GB VRAM
Arsen Apostolov
Arsen Apostolov
Arsen Apostolov
Follow
Jul 5
vLLM vs llama.cpp vs Ollama: What Happens When Your Model Doesn't Fit in 24GB VRAM
#
llm
#
homelab
#
vllm
#
ai
Comments
Add Comment
6 min read
Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive
xbill
xbill
xbill
Follow
for
Google Developer Experts
Jul 21
Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive
#
tpu
#
llm
#
vllm
#
googlecloud
8
 reactions
Comments
1
 comment
8 min read
tpu-management: a Claude Code skill for running Gemma 4 on Cloud TPUs
xbill
xbill
xbill
Follow
for
Google Developer Experts
Jul 21
tpu-management: a Claude Code skill for running Gemma 4 on Cloud TPUs
#
googlecloud
#
tpu
#
vllm
#
claudecode
7
 reactions
Comments
Add Comment
3 min read
AI Inference Optimization in Cloud-Native Environments: GPU Orchestration, Edge Deployment, and Latency Reduction at Scale
The Cyber Sidekick
The Cyber Sidekick
The Cyber Sidekick
Follow
Jun 23
AI Inference Optimization in Cloud-Native Environments: GPU Orchestration, Edge Deployment, and Latency Reduction at Scale
#
kubernetesgpuscheduling
#
aiinferenceoptimization
#
llmdeployment
#
vllm
Comments
Add Comment
4 min read
Kimi K3 Open Weights Are Here: How to Self-Host the 2.8T-Parameter Model (Hardware, vLLM, and Data Sovereignty)
Lola Lin
Lola Lin
Lola Lin
Follow
Jul 27
Kimi K3 Open Weights Are Here: How to Self-Host the 2.8T-Parameter Model (Hardware, vLLM, and Data Sovereignty)
#
kimik3
#
openweights
#
selfhostai
#
vllm
4
 reactions
Comments
Add Comment
8 min read
AI Inference at the Edge: Running Real-Time LLMs in Kubernetes Without a GPU Farm
The Cyber Sidekick
The Cyber Sidekick
The Cyber Sidekick
Follow
Jun 18
AI Inference at the Edge: Running Real-Time LLMs in Kubernetes Without a GPU Farm
#
edgeai
#
kubernetes
#
llminference
#
vllm
Comments
Add Comment
3 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account