VLDB 2026 Research / reviewers in the wild / expert
Yiming Lei 0002
dblp:77/7910-2
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-5124-7156ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 7 · 2 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SyncWise: Error-Aware Time Synchronization for Reconfigurable Data Center Networks
Yiming Lei 0002, Jialong Li 0006, Zhengqing Liu, Raj Joshi, Yiting Xia |
NSDI | 1 |
| 2026 | OpenOptics: Enabling Open Research and Implementation of Optical Data Center Networks
Yiming Lei 0002, Federico De Marchi 0002, Jialong Li 0006, Raj Joshi, Shu-Ting Wang, Balakrishnan Chandrasekaran 0002, Yiting Xia |
NSDI | 1 |
| 2026 | Unlocking Diversity of Fast-Switched Optical Data Center Networks With Unified RoutingabstractOptical data center networks (DCNs) are emerging as a promising solution for cloud infrastructure in the post-Moore’s Law era, particularly with the advent of “fast-switched” optical architectures capable of circuit reconfiguration at microsecond or even nanosecond scales. However, frequent reconfiguration of optical circuits introduces a unique challenge: in-flight packets risk loss during these transitions, hindering the deployment of many mature optical hardware designs due to the lack of suitable routing solutions. In this paper, we presentUnifiedRouting forOptical networks (URO), a general routing framework designed to support fast-switched optical DCNs across various hardware architectures. URO combines theoretical modeling of this novel routing problem with practical implementation on programmable switches, enabling precise, time-based packet transmission. Our prototype on Intel Tofino2 switches achieves a minimum circuit duration of$\mathrm {2~\mu \text {s} }$, ensuring end-to-end, loss-free application performance. Large-scale simulations using production DCN traffic validate URO’s generality across different hardware configurations, demonstrating its effectiveness and efficient system resource utilization. Jialong Li 0006, Federico De Marchi 0002, Yiming Lei 0002, Raj Joshi, Balakrishnan Chandrasekaran 0002, Yiting Xia |
IEEE Trans. Netw. | 3 |
| 2026 | Optimizing Mixture-of-Experts Inference Time via Model Deployment and Communication SchedulingabstractAs machine learning models scale in size and complexity, their computational requirements become a significant barrier. Mixture-of-Experts (MoE) models alleviate this issue by selectively activating relevant experts. Despite this, MoE models are hindered by high communication overhead from all-to-all operations, low GPU utilization, and complications from heterogeneous GPU environments. This paper presents Comet, which optimizes both model deployment and all-to-all communication scheduling to address these challenges in MoE inference. Comet achieves minimal communication times by strategically ordering token transmissions in all-to-all communications. It improves GPU utilization by colocating experts from different models on the same device, avoiding the limitations of all-to-all communication. We analyze Comet’s optimization strategies theoretically across four common GPU cluster settings: exclusive vs. colocated models on GPUs, and homogeneous vs. heterogeneous GPUs. Comet provides optimal solutions for three cases, and for the remaining NP-hard scenario, it offers a polynomial-time sub-optimal solution with only a 1.09× degradation from the optimal, as shown in the simulation results. Comet is the first approach to minimize MoE inference time via optimal model deployment and communication scheduling across various scenarios. Evaluations demonstrate that Comet significantly accelerates inference, achieving speedups of up to 2.63× in homogeneous clusters and 2.91× in heterogeneous environments. Moreover, Comet enhances GPU utilization by up to 2.38× compared to existing methods. Jialong Li 0006, Shreyansh Tripathi, Lakshay Rastogi, Yiming Lei 0002, Rui Pan 0003, Yiting Xia |
IEEE Trans. Netw. | 4 |
| 2024 | Uniform-Cost Multi-Path Routing for Reconfigurable Data Center NetworksabstractReconfigurable data center networks (RDCNs) are arising as a promising data center network (DCN) design in the post-Moore's law era. However, the constantly reconfigured network topology in RDCNs invalidates the assumption of using hop count as the cost metric for routing, e.g., the status quo Equal-Cost Multi-Path routing (ECMP) in traditional DCNs. Unfortunately, existing routing solutions in RDCNs stick to the old assumption and deliver suboptimal performance either high in latency or low in bandwidth efficiency. In this paper, we redefine the cost metric for RDCN routing with uniform cost to unify the effects of topology disruption and hop count on latency and bandwidth efficiency. We propose Uniform-Cost Multi-Path routing (UCMP), an ECMP equivalent for RDCNs, where minimizing uniform cost leads flows of various sizes to the right balance between latency and bandwidth efficiency. Our simulation shows that UCMP achieves 53% to 98% lower flow completion time (FCT) and 1.55× bandwidth efficiency compared to the state-of-the-art RDCN routing strategy, and our testbed implementation demonstrates sustainable switch resource usage of UCMP as RDCNs scale. Jialong Li 0006, Haotian Gong, Federico De Marchi 0002, Aoyu Gong, Yiming Lei 0002, Wei Bai 0001, Yiting Xia |
SIGCOMM | 5 |
| 2022 | Hop-On Hop-Off Routing: A Fast Tour across the Optical Data Center Network for Latency-Sensitive FlowsabstractOptical data center networks show promise to serve as the next-generation cloud infrastructure especially with their cost and power benefits. The need to set up dedicated optical circuits between endpoints before they can exchange data, however, delays latency-sensitive (“mice”) flows. We find the state-of-the-art solution to reducing flow latency produces sub-optimal paths. To address this issue, we leverage programmable switches to realize Hop-On Hop-Off (HOHO) routing, where mice flows are forwarded along the minimal-latency paths. We prove the optimality and robustness of our algorithm and sketch an implementation on programmable switches. In our packet-level simulations, HOHO routing reduces the flow-completion times for mice flows by up to 35% and the average path length by 15% compared to the state-of-the-art solution. Jialong Li 0006, Yiming Lei 0002, Federico De Marchi 0002, Raj Joshi, Balakrishnan Chandrasekaran 0002, Yiting Xia |
APNet | 2 |
| 2022 | Efficient flow scheduling in distributed deep learning training with echelon formationabstractThis paper discusses why flow scheduling does not apply to distributed deep learning training and presents EchelonFlow, the first network abstraction to bridge the gap. EchelonFlow deviates from the common belief that semantically related flows should finish at the same time. We reached the key observation, after extensive workflow analysis of diverse training paradigms, that distributed training jobs observe strict computation patterns, which may consume data at different times. We devise a generic method to model the drastically different computation patterns across training paradigms, and formulate EchelonFlow to regulate flow finish times accordingly. Case studies of mainstream training paradigms under EchelonFlow demonstrate the expressiveness of the abstraction, and our system sketch suggests the feasibility of an EchelonFlow scheduling system. Rui Pan 0003, Yiming Lei 0002, Jialong Li 0006, Binhang Yuan, Yiting Xia |
HotNets | 2 |