Cerebras CS-4 doubles inference performance with wafer-scale engine and SRAM for MoE models
Cerebras CS-4:性能翻倍 | WSE-3晶圆级引擎 | AI推理芯片 | 内存带宽 | tokens/s | 流水线并行 | 解耦推理 | MoE混合专家模型 | SRAM架构
The post only has a title with no body. Cerebras announced the CS-4 inference chip claiming 2x performance, powered by the WSE-3 wafer-scale engine and SRAM architecture. It targets high memory bandwidth, high tokens/s, pipeline parallelism, decoupled inference, and MoE models. Price, power, and availability are not disclosed.