Hi, first of all thanks for this great resource.
I want to ask if you have any plan to add a chapter on Mixture of Experts and Expert Parallel.
I find there are quite a few non-obvious axes related to MoE such as EP+(TP/DP/CP), training and inference workload, expert choice vs token choices etc, such that having a holistic understanding of performance can get quite tricky.
If you don't have such plan, please considering adding it ;)
Thanks again.
Hi, first of all thanks for this great resource.
I want to ask if you have any plan to add a chapter on Mixture of Experts and Expert Parallel.
I find there are quite a few non-obvious axes related to MoE such as EP+(TP/DP/CP), training and inference workload, expert choice vs token choices etc, such that having a holistic understanding of performance can get quite tricky.
If you don't have such plan, please considering adding it ;)
Thanks again.