Discussion about this post

User's avatar
JP's avatar

Coming back to this because the Super just dropped and it validates everything you wrote about the enterprise economics angle. The Nano was the opening move; Super takes it further with 120B total parameters, 12B active, and the same hybrid Mamba-Transformer MoE architecture but at a scale that competes with frontier models on reasoning.

The LatentMoE routing in Super is the bit that wasn't obvious from the Nano release. It compresses tokens before routing across 512 experts, which is how they pack in 4x more expert capacity at the same compute cost. I went deep on the full architecture here https://reading.sh/inside-the-model-merging-three-ai-architectures-into-one-c5dcc7302528 because the engineering choices are genuinely different from anything else in the open-source space.

Synthetic already picked up Super for their flash tier. A 120B model at flash speeds. The enterprise cost calculus you outlined in this piece just got even more interesting.

The AI Architect's avatar

Outstanding analysis on the SaaS wrapper risk. The part about break-even thresholds matches what we saw deploying on-prem inference last year, costs dropped like 95% once we hit volume. I dunno if enough VCs are factoring this commoditization into their valuations tho, theres gonna be a reckoning for companies charging $100/seat for glorified API calls.

3 more comments...

No posts

Ready for more?