Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Correct, but note that's exactly what inference providers do.

https://arxiv.org/pdf/2412.19437

> The minimum deployment unit of the decoding stage consists of 40 nodes with 320 GPUs. The attention part employs TP4 with SP, combined with DP80, while the MoE part uses EP320.

EP320 means expert parallelism, each on 320 GPUs.



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: