Discussion about this post

User's avatar
Atlas.ML's avatar

So, theoretically, when moving to HBM5, with a hypothetical 1GB of HBM on one Nvidia AI GPU, the context window will be so large that it can solve a highly complex problem?

The constraints now lies in the HBM capacity and bandwidth for LLM companies to expand their processing capability? In exploring “better answer or solution” ?

roy's avatar
Mar 7Edited

You've argued the Groq co-design is additive TAM and that better models keep increasing memory bandwidth demand. But does that inflect? Even with multi-chip optical sync extending SRAM’s effective ceiling, there’s a crossover point where reasoning hcains outgrow what that architecture can economically serve - and the HBM’s density advantage reasserts. What share of total inference demand do you think stays below that crossover long term? Or do reasoning models become the default and the LPU-hybrid case shrinks to a latency-sensitive niche?

2 more comments...

No posts

Ready for more?