https://arxiv.org/pdf/2412.19437
> The minimum deployment unit of the decoding stage consists of 40 nodes with 320 GPUs. The attention part employs TP4 with SP, combined with DP80, while the MoE part uses EP320.
EP320 means expert parallelism, each on 320 GPUs.
https://arxiv.org/pdf/2412.19437
> The minimum deployment unit of the decoding stage consists of 40 nodes with 320 GPUs. The attention part employs TP4 with SP, combined with DP80, while the MoE part uses EP320.
EP320 means expert parallelism, each on 320 GPUs.