SiTime actually benefits here. A master clock architecture raises the bar on timing precision. You still need ultra-low-jitter reference oscillators at the source. Higher requirements = tailwind for SITM.
Thanks Ben for the insights! Great piece as always!
I have a quick question about completely bypassing HBM to only use SRAM: is the capability for SRAM scalable enough for the ever growing context window and model size/weights?
Not for everything and that’s the point. Huge training models stay on GPUs with HBM. This targets high-volume inference with models co-designed to fit. SSM hybrids keep state fixed, 3D stacking pushes SRAM capacity each gen. Two lanes, each doing what it does best.
Great piece. If there is one master clock what does it mean for sitm?
SiTime actually benefits here. A master clock architecture raises the bar on timing precision. You still need ultra-low-jitter reference oscillators at the source. Higher requirements = tailwind for SITM.
Thanks Ben for the insights! Great piece as always!
I have a quick question about completely bypassing HBM to only use SRAM: is the capability for SRAM scalable enough for the ever growing context window and model size/weights?
Not for everything and that’s the point. Huge training models stay on GPUs with HBM. This targets high-volume inference with models co-designed to fit. SSM hybrids keep state fixed, 3D stacking pushes SRAM capacity each gen. Two lanes, each doing what it does best.