SteadyDetected Jun 23

Runing GLM-5.2 on local hardware

Steady
1.1 momentum
Hacker NewsReddit
18 stories across sources

What's happening

GLM-5.2 is now running locally across multiple setups and quantizations, with users reporting concrete speed and size numbers. Examples: unsloth/GLM-5.2-GGUF UD-IQ1_M on llama.cpp reached ~579.75 tokens/s prefill at 8k context and ~324.48 tokens/s prefill at 57k context on a RTX 5090 + RTX 3090 Ti setup, with decode around 10.6 t/s. Another user reports GLM-5.2 UD-IQ2_M at ~7.3 tok/s generation on 4×RTX 3090 with 192GB and RAM expert offload, and notes IQ2->IQ1 quant halving did nothing while increasing CPU threads 6->12 gave +22%. The 2-bit GGUF quant shrank the model from 1.51TB to 238GB (≈-84% size) while retaining ~82% accuracy, enabling runs on 256GB Mac or mixed RAM/VRAM rigs; community posts show CPU-only experiments (dual Xeon 6248R, 768GB) and multigpu configs (6×3090, 128GB) with generation rates in the single-digit to high single-digit tok/s range depending on engine and quant.

Why it's trending

Because unsloth released GGUF quant files (UD-IQ1_M/UD-IQ2_M and 2-bit) and users immediately benchmarked them across llama.cpp, ik_llama.cpp and offload setups, producing shareable speed/size wins.

Story volume

Stories per day
06-1306-1706-1806-1906-2006-22

Angles you could write

contrarian take

If you think GLM-5.2 is too huge to run locally, the 1.51TB->238GB 2-bit GGUF should force you to rethink your hardware bets right now.

+2 more angles for this topic with an account — all it takes is your email.

Original sources15

+12 more sources for this topic

Create an account to follow the full coverage in the live radar.

More rising in AI & Tech

All rising AI & Tech trends →