Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
llamacpp
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
My inference server decided my second GPU no longer exists. Here is how I got it back without upgrading a driver.
Rickesh T N
Rickesh T N
Rickesh T N
Follow
Aug 24
My inference server decided my second GPU no longer exists. Here is how I got it back without upgrading a driver.
#
gpu
#
llamacpp
#
ollama
#
homelab
Comments
Add Comment
3 min read
Fix Local LLM Quality: Context Stacking & Rope Freq Tweaks
Umair Bilal
Umair Bilal
Umair Bilal
Follow
Aug 23
Fix Local LLM Quality: Context Stacking & Rope Freq Tweaks
#
localllms
#
aiagents
#
ollama
#
llamacpp
Comments
Add Comment
8 min read
"V cache quantization requires flash_attn" — the llama.cpp error that quietly halves your context window
Jasur Yuldoshev
Jasur Yuldoshev
Jasur Yuldoshev
Follow
Aug 21
"V cache quantization requires flash_attn" — the llama.cpp error that quietly halves your context window
#
llamacpp
#
llm
#
performance
#
debugging
Comments
Add Comment
10 min read
Nine ways to talk to a local model
the kilted dev
the kilted dev
the kilted dev
Follow
Aug 11
Nine ways to talk to a local model
#
localllm
#
ollama
#
llamacpp
#
buildinpublic
Comments
Add Comment
9 min read
Running a 26B MoE on an 8 GB Jetson by streaming experts from SSD
Anish Shrestha
Anish Shrestha
Anish Shrestha
Follow
Aug 1
Running a 26B MoE on an 8 GB Jetson by streaming experts from SSD
#
moestream
#
llamacpp
#
mixtureofexperts
#
jetsonorin
Comments
Add Comment
6 min read
GnLOLot Releases MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF for Enhanced Local AI Development
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 18
GnLOLot Releases MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF for Enhanced Local AI Development
#
gguf
#
llamacpp
#
quantized
#
minicpm5
Comments
Add Comment
3 min read
Running a 1.5B-Parameter LLM Entirely On-Device for Mental Health — The NilaMind Architecture
sampathmannam
sampathmannam
sampathmannam
Follow
Jul 14
Running a 1.5B-Parameter LLM Entirely On-Device for Mental Health — The NilaMind Architecture
#
llamacpp
#
opensource
#
android
Comments
Add Comment
5 min read
llama-bench skipped FA on capable GPUs — b9437 corrects it
Creeta
Creeta
Creeta
Follow
Jun 18
llama-bench skipped FA on capable GPUs — b9437 corrects it
#
llamacpp
#
llm
#
gguf
#
flashattention
Comments
Add Comment
7 min read
Hot-swapping GGUF models kernel-panicked my M4 Mac: wired memory, llama.cpp, and why we restart the server instead
Jasur Yuldoshev
Jasur Yuldoshev
Jasur Yuldoshev
Follow
Jul 18
Hot-swapping GGUF models kernel-panicked my M4 Mac: wired memory, llama.cpp, and why we restart the server instead
#
llm
#
macos
#
debugging
#
llamacpp
Comments
Add Comment
5 min read
Hermes Agent Desktop Free With Local LLMs: The Claude Code Alternative Nobody's Billing You For [2026]
Kunal
Kunal
Kunal
Follow
Jun 5
Hermes Agent Desktop Free With Local LLMs: The Claude Code Alternative Nobody's Billing You For [2026]
#
hermesagent
#
localllm
#
claudecodealternative
#
llamacpp
Comments
Add Comment
8 min read
What secretly eats your local LLMs' speed as your context fills up - Part 2
Federico "SpeederX" Piana
Federico "SpeederX" Piana
Federico "SpeederX" Piana
Follow
Jul 4
What secretly eats your local LLMs' speed as your context fills up - Part 2
#
ai
#
machinelearning
#
locallm
#
llamacpp
1
 reaction
Comments
Add Comment
4 min read
Can a $2,000 Mini PC Replace Your AI Cloud Bill?
ZyVOP
ZyVOP
ZyVOP
Follow
Jun 22
Can a $2,000 Mini PC Replace Your AI Cloud Bill?
#
localai
#
strixhalo
#
hermesagent
#
llamacpp
Comments
Add Comment
9 min read
How to Tune llama.cpp --n-gpu-layers: A Practical VRAM Guide (2026)
Patrick Hughes
Patrick Hughes
Patrick Hughes
Follow
Jun 9
How to Tune llama.cpp --n-gpu-layers: A Practical VRAM Guide (2026)
#
localllm
#
llamacpp
#
gpu
#
vram
Comments
Add Comment
4 min read
How to Tune --n-gpu-layers for Your VRAM Budget
Patrick Hughes
Patrick Hughes
Patrick Hughes
Follow
Jun 8
How to Tune --n-gpu-layers for Your VRAM Budget
#
localllm
#
llamacpp
#
gpu
#
vram
Comments
Add Comment
4 min read
llama.cpp ngl: when -ngl 99 still runs on your CPU
Patrick Hughes
Patrick Hughes
Patrick Hughes
Follow
Jun 4
llama.cpp ngl: when -ngl 99 still runs on your CPU
#
llamacpp
#
localllm
#
gpuoffloading
#
ngpulayers
1
 reaction
Comments
Add Comment
5 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account