DEV Community

#inference

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
what a turn actually costs me

what a turn actually costs me

Comments
2 min read
LLM Inference Optimization: From Quantization to Speculative Decoding

LLM Inference Optimization: From Quantization to Speculative Decoding

1
Comments
2 min read
AMD's Move on Weight Storage: The Taalas Bet

AMD's Move on Weight Storage: The Taalas Bet

Comments
2 min read
AMD Bets Weight Storage Is the Real Bottleneck

AMD Bets Weight Storage Is the Real Bottleneck

Comments
2 min read
vLLM reinvented the operating system, and nobody told you

vLLM reinvented the operating system, and nobody told you

1
Comments
12 min read
You're Not Paying for Compute. You're Paying for Memory Bandwidth

You're Not Paying for Compute. You're Paying for Memory Bandwidth

Comments
4 min read
Inference Optimization for MiMo v2.5: Mastering Hybrid SWA Efficiency

Inference Optimization for MiMo v2.5: Mastering Hybrid SWA Efficiency

Comments
2 min read
local-llm: A Field Report on Running SOTA Models on Your Own Hardware

local-llm: A Field Report on Running SOTA Models on Your Own Hardware

1
Comments 1
3 min read
The KV cache, why LLM inference is memory-bound, not compute-bound

The KV cache, why LLM inference is memory-bound, not compute-bound

Comments
4 min read
Etched hits $5B and $1B in orders: why inference chips matter

Etched hits $5B and $1B in orders: why inference chips matter

Comments
4 min read
Two labs race to make AI write whole paragraphs at once instead of word by word

Two labs race to make AI write whole paragraphs at once instead of word by word

Comments
3 min read
96% of cuBLAS, no `unsafe`: what cuTile Rust proves

96% of cuBLAS, no `unsafe`: what cuTile Rust proves

Comments
8 min read
Extract Structured JSON from Messy Text with Telnyx AI Inference

Extract Structured JSON from Messy Text with Telnyx AI Inference

Comments
2 min read
Chạy LLM trên iGPU: Giới hạn VRAM của Intel Arc và Radeon 780M

Chạy LLM trên iGPU: Giới hạn VRAM của Intel Arc và Radeon 780M

Comments
3 min read
How to Build a Secure Homelab for LLM Inference

How to Build a Secure Homelab for LLM Inference

Comments
4 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.