Blog
Hands-on notes on LLMs and software engineering — what we learned by actually building things.
We cut DeepSeek-V4-Flash's experts from 256 to 128 per layer with REAP, shrinking the 167 GB checkpoint to 82 GB and serving it at 17 tok/s on one DGX Spark with stock vLLM. Calibration ran on the same 128 GB machine — and an English-only calibration set turned the model's Japanese into Chinese.