Block-wise diffusion language models run diffusion inside a block and attend causally across blocks—so, unlike full-attention dLLMs whose keys and values change at every denoising step, they reuse a stable prefix KV cache. That cache is re-read at every step, and at long context its traffic dominates inference latency. While sparse attention has begun to attack this cost, KV-cache quantization for dLLMs has gone essentially unstudied. We apply three existing AR KV-cache quantization methods—uniform, outlier-protected, and SVD-based mixed precision—to a block-diffusion model and its AR parent. At every matched budget, all three methods hurt the AR parent more; below approximately 2 bits, the AR model collapses while the block-diffusion model retains most of its accuracy. The reason is the attention pattern. The AR model assigns roughly a third of every long-context read to its attention sink, so that single token’s quantization error reaches the output at almost full strength; the block-diffusion model’s much weaker sink spreads reads across many tokens, whose independent errors largely cancel. Building on this observation, we propose QWaF, a query-aware mixed-precision KV-cache quantizer that preserves the key directions used by future queries. On SDAR-8B, QWaF retains 97% of the FP16 LongBench score at 2.22 effective bits, while 1-bit KIVI retains 25% at 2.15 bits. At 32K, QWaF reduces KV-cache storage by up to 10.8×, and a fused CUDA serving path reduces decoding latency by 65% on an A100. To our knowledge, QWaF is the first KV-cache quantization method for diffusion LLMs.