Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
inference
Follow
Hide
Posts
Left menu
👋
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
what a turn actually costs me
Saltorious
Saltorious
Saltorious
Follow
Aug 13
what a turn actually costs me
#
engineering
#
inference
#
localmodels
Comments
Add Comment
2 min read
LLM Inference Optimization: From Quantization to Speculative Decoding
Aarush Karak
Aarush Karak
Aarush Karak
Follow
Aug 13
LLM Inference Optimization: From Quantization to Speculative Decoding
#
llm
#
inference
#
quantization
#
optimization
1
 reaction
Comments
Add Comment
2 min read
AMD's Move on Weight Storage: The Taalas Bet
Peremptory
Peremptory
Peremptory
Follow
Aug 12
AMD's Move on Weight Storage: The Taalas Bet
#
amd
#
compute
#
inference
#
aiinfrastructure
Comments
Add Comment
2 min read
AMD Bets Weight Storage Is the Real Bottleneck
Peremptory
Peremptory
Peremptory
Follow
Aug 11
AMD Bets Weight Storage Is the Real Bottleneck
#
amd
#
hardware
#
aiinfrastructure
#
inference
Comments
Add Comment
2 min read
vLLM reinvented the operating system, and nobody told you
Yathiskumar
Yathiskumar
Yathiskumar
Follow
Aug 11
vLLM reinvented the operating system, and nobody told you
#
systemdesign
#
llm
#
inference
#
operatingsystems
1
 reaction
Comments
Add Comment
12 min read
You're Not Paying for Compute. You're Paying for Memory Bandwidth
AI Explore
AI Explore
AI Explore
Follow
Jul 11
You're Not Paying for Compute. You're Paying for Memory Bandwidth
#
ai
#
llm
#
inference
#
mlops
Comments
Add Comment
4 min read
Inference Optimization for MiMo v2.5: Mastering Hybrid SWA Efficiency
Tamiz Uddin
Tamiz Uddin
Tamiz Uddin
Follow
Jul 11
Inference Optimization for MiMo v2.5: Mastering Hybrid SWA Efficiency
#
ai
#
inference
#
optimization
#
mimo
Comments
Add Comment
2 min read
local-llm: A Field Report on Running SOTA Models on Your Own Hardware
Reno Lu
Reno Lu
Reno Lu
Follow
Jul 20
local-llm: A Field Report on Running SOTA Models on Your Own Hardware
#
localllm
#
gpu
#
selfhosting
#
inference
1
 reaction
Comments
1
 comment
3 min read
The KV cache, why LLM inference is memory-bound, not compute-bound
I Want To Learn Programming
I Want To Learn Programming
I Want To Learn Programming
Follow
Jul 4
The KV cache, why LLM inference is memory-bound, not compute-bound
#
gpu
#
llm
#
inference
#
performance
Comments
Add Comment
4 min read
Etched hits $5B and $1B in orders: why inference chips matter
Induwara Ashinsana
Induwara Ashinsana
Induwara Ashinsana
Follow
Jul 1
Etched hits $5B and $1B in orders: why inference chips matter
#
aihardware
#
inference
#
cost
Comments
Add Comment
4 min read
Two labs race to make AI write whole paragraphs at once instead of word by word
Breach Protocol
Breach Protocol
Breach Protocol
Follow
Jul 1
Two labs race to make AI write whole paragraphs at once instead of word by word
#
diffusion
#
openweight
#
google
#
inference
Comments
Add Comment
3 min read
96% of cuBLAS, no `unsafe`: what cuTile Rust proves
Creeta
Creeta
Creeta
Follow
Jun 26
96% of cuBLAS, no `unsafe`: what cuTile Rust proves
#
cutile
#
rust
#
gpu
#
inference
Comments
Add Comment
8 min read
Extract Structured JSON from Messy Text with Telnyx AI Inference
Sonam
Sonam
Sonam
Follow
Jun 26
Extract Structured JSON from Messy Text with Telnyx AI Inference
#
ai
#
inference
#
telnyx
#
json
Comments
Add Comment
2 min read
Chạy LLM trên iGPU: Giới hạn VRAM của Intel Arc và Radeon 780M
Review Laptop
Review Laptop
Review Laptop
Follow
Jun 21
Chạy LLM trên iGPU: Giới hạn VRAM của Intel Arc và Radeon 780M
#
llama3
#
llm
#
ollama
#
inference
Comments
Add Comment
3 min read
How to Build a Secure Homelab for LLM Inference
Jay Grider
Jay Grider
Jay Grider
Follow
Jun 12
How to Build a Secure Homelab for LLM Inference
#
homelab
#
llmsecurity
#
inference
#
supplychain
Comments
Add Comment
4 min read
👋
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account