I think this is where the market’s current understanding of inference semiconductors becomes visible.
My summary of @cerebras is as follows.
The WSE 3 Turbo in the CS 4 does not replace HBM.
It is not superior across all inference workloads either.
However, the low latency