Seminar: Graduate Seminar

ECE Women Community

Topology and Routing Co-Design for Efficient Mixture-of-Experts Communication

Date: September,22,2026 Start Time: 10:00 - 11:00
Location: 506, New Zisapel Building
Add to:
Lecturer: Ori Cohen
Distributed training increasingly relies on specialized interconnects to meet the bandwidth, cost, and power demands of large machine-learning models. Low-degree direct-connect networks and optical circuit switches offer an attractive alternative to monolithic packet-switched fabrics, but traffic between non-neighboring accelerators must be relayed. Mixture-of-experts models make this limitation especially important because every token is dynamically dispatched to a sparse set of experts distributed across accelerators.
This thesis develops topology and routing techniques for efficient mixture-of-experts com-munication on direct-connect fabrics. Within the collaborative TLOS architecture, it intro-duces an expander-based topology for expert parallelism, a splittable expander construction that uses low-radix optical switches to support different communication-group sizes, and failure resilience for the expander dimension. It also develops end-to-end simulation support for re-configurable multi-dimensional topologies, mixture-of-experts communication, and one-forward-one-backward pipeline schedules.
The TLOS evaluation reveals that a low-diameter topology alone does not remove the cost of indirect routing for communication-intensive mixture-of-experts models. To address this gap, the thesis presents BATON, a routing system that retains the semantics hidden by aggregate AllToAll-V communication. BATON treats dispatch as per-token multicast, reverses the mul-ticast tree to partially reduce expert outputs during combine, and balances relay traffic using static per-link weights computed offline. It requires neither prediction of the deployed traffic matrix nor continuous topology reconfiguration.
Using token-level traces and end-to-end simulation, the thesis shows that BATON ap-
proaches ideal packet-switched performance on the evaluated 16-, 32-, and 64-node expanders and substantially outperforms conventional min-hop routing. It also reduces bottleneck cable load in a 1,024-accelerator Boardfly-style hierarchy, demonstrating that the approach applies beyond TLOS. These results show that exposing token semantics to the network can make specialized direct-connect fabrics effective for dynamic mixture-of-experts workloads.

M.Sc. student under the supervision of Prof. Mark  Silberstein.

 

All Seminars