Runing GLM-5.2 on local hardware
What's happening
GLM-5.2 is now running locally across multiple setups and quantizations, with users reporting concrete speed and size numbers. Examples: unsloth/GLM-5.2-GGUF UD-IQ1_M on llama.cpp reached ~579.75 tokens/s prefill at 8k context and ~324.48 tokens/s prefill at 57k context on a RTX 5090 + RTX 3090 Ti setup, with decode around 10.6 t/s. Another user reports GLM-5.2 UD-IQ2_M at ~7.3 tok/s generation on 4×RTX 3090 with 192GB and RAM expert offload, and notes IQ2->IQ1 quant halving did nothing while increasing CPU threads 6->12 gave +22%. The 2-bit GGUF quant shrank the model from 1.51TB to 238GB (≈-84% size) while retaining ~82% accuracy, enabling runs on 256GB Mac or mixed RAM/VRAM rigs; community posts show CPU-only experiments (dual Xeon 6248R, 768GB) and multigpu configs (6×3090, 128GB) with generation rates in the single-digit to high single-digit tok/s range depending on engine and quant.
Why it's trending
Because unsloth released GGUF quant files (UD-IQ1_M/UD-IQ2_M and 2-bit) and users immediately benchmarked them across llama.cpp, ik_llama.cpp and offload setups, producing shareable speed/size wins.
Story volume
Stories per dayAngles you could write
If you think GLM-5.2 is too huge to run locally, the 1.51TB->238GB 2-bit GGUF should force you to rethink your hardware bets right now.
+2 more angles for this topic with an account — all it takes is your email.
Original sources15
- Idea for how to run GLM2 at a decent quant, need critique/feedback
I am currently running a 4x 5060 ti P2P rig (64 GB VRAM total)where each card is running at gen 3 with 4 pcie lanes per card. My use case is inference only. During my benchmarking the bottleneck was compute, not pcie bandwidth for low concurrency inference tasks, such as a single user use case. This gave me an idea, since my cards are already running at gen 3 pcie, I could pickup 512 GB of DDR3 16
r/LocalLLaMAJun 22 - GLM-5.2 UD-IQ1_M on llama.cpp — 5090 + 3090 Ti speed test (~ 579 t/s prefill @ 8k ctx, ~324 t/s prefill @ 57k ctx, ~10.6 t/s decode)
Just sharing some speed test numbers for GLM-5.2 running on llama.cpp. Setup: Model: unsloth/GLM-5.2-GGUF, UD-IQ1_M quant GPUs: RTX 5090 + RTX 3090 Ti 186 GB DDR5 used Debian 13 CUDA 13.3 128k context, q8_0 KV cache Prefill (prompt processing): n_tokens tokens/s 8,201 579.75 16,393 522.28 24,585 468.21 32,777 422.61 40,969 384.43 49,161 351.90 57,353 324.48 Decode (generation): Holds steady around
r/LocalLLaMAJun 22 - Runing GLM-5.2 on local hardwareHackerNewsJun 22
+12 more sources for this topic
Create an account to follow the full coverage in the live radar.
More rising in AI & Tech
- 'AI runs on semiconductors': Why chips have become the world's most valuable technologyClimbing2.4Climbing2.4 momentum
- AI leaders sign statement asking the government to do something about automated AIClimbing1.9Climbing1.9 momentum
- Google just had its first negative cash flow quarter due to massive AI spendingClimbing2.8Climbing2.8 momentum
- How AI guardrails are impeding the work of offensive cybersecurity researchersSteady0.8Steady0.8 momentum
- Korean chip stocks tumble with SK Hynix below US listing price amid China competition fearsClimbing1.9Climbing1.9 momentum
- AMD vs. Nvidia: What AMD’s Major $5 Billion AI-Chip Deal With Anthropic Means for InvestorsSteady0.5Steady0.5 momentum