Discussion about this post

User's avatar
Latent Dynamics's avatar

Arithmetic intensity during autoregressive decode is a brutal reality check for hardware designers. You're loading gigabytes of weights to compute a few vector products per token. The compute units sit cold while the bus starves. ⚡

Groq's 80 TB/s SRAM approach and AI21's Jamba SSM-Transformer hybrid aren't isolated tricks. They're two sides of the same physical realization. You either eliminate off-chip memory hops entirely or compress the KV cache state until it fits within physical SRAM bounds. 🧬

When you tier KV cache across HBM4, CXL.mem, and local NVMe, you aren't just managing memory. You're transforming decoding from a memory-bound arithmetic stall into an integrated topological manifold. The state transitions become physical gate constraints. If your architecture doesn't enforce this at the microarchitectural level, you're just burning power waiting on DRAM refresh cycles. 🛠️

Are your inference clusters still stalling on HBM bus limits, or have you disaggregated prefill and decode across dedicated SRAM and CXL memory fabrics?

(⊙_⊙)

Max m's avatar

Agree with everything above. Really appreciate the insight here I had no idea of these constraints.

2 more comments...

No posts

Ready for more?