VLDB 2026 Research / reviewers in the wild / expert
Zhehao Lin
dblp:300/6728
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | xCCLTuner: Treating xCCL as Black-Box and Automatically Tuning
Chenxu Wang 0007, Zhehao Lin, Peirui Cao, Xiaohu Xu, Wan-Chun Dou, Guihai Chen, Chen Tian 0001 |
INFOCOM | 3 |
| 2026 | SwiftEP: Accelerating MoE Inference with Buffer Fusion and TMA Offloading
Xingyi Li 0004, Shangguang Wang, Zhehao Lin, Yinben Xia, Qihang Liu, Xiang Li 0067, Zekun He, Yachen Wang, Xianneng Zou |
NSDI | 7 |
| 2026 | OSCAR: O(1)-Step Convergence and Readily-deployable Congestion Control
Zhaochen Zhang, Feiyang Xue, Rui Ning, Keqiang He, Gianni Antichi, Zhimeng Yin 0001, Rui Li 0020, Zhengqi Cui, Zhehao Lin, Peirui Cao, Guihai Chen, Chen Tian 0001 |
NSDI | 11 |
| 2026 | Rail: ReArranging Inter-GPU Links for GPU-Centric ClustersabstractIn modern GPU-centric clusters, large-scale AI training relies on two distinct communication domains: a high-bandwidth intra-node domain using proprietary interconnects (e.g., NVLink), and a scale-out inter-node network domain (e.g., RDMA). We observe that the widely-used ring algorithm, often create a significant load imbalance across these domains. This leads to the counter-intuitive scenario where the expensive, high-bandwidth intra-node domain becomes a performance bottleneck, while the inter-node network remains underutilized. This inefficiency is further exacerbated by the disparity in bandwidth provisioning: inter-node network bandwidth is generally more cost-effective and accessible, whereas intra-node bandwidth is often proprietary and more costly to scale. To address this fundamental imbalance, we propose RAIL, aimed at resolving the intra-node bottleneck by strategically rearranging inter-GPU communication paths. This rebalancing ensures that traffic loads are appropriately matched with the distinct transmission capabilities of each domain, thereby maximizing overall communication performance. RAIL incorporates a Load Distributing Strategy (LDS) that can accurately partition physical nodes into logical nodes based on the a transmission capabilities of both domains, shifting excess traffic from the overloaded intra-node domain to the underutilized network domain. Additionally, the Intra-Rail Strategy (IRS) leverages topological characteristics to ensure optimal communication paths through the network domain between logical nodes. Our evaluation demonstrates that RAIL effectively mitigates congestion and achieves a 30.7% average increase in collective communication bus bandwidth compared to the widely-used NCCL solution. Haixin Nan, Jun Xu 0037, Peirui Cao, Zhaochen Zhang, Yizhi Wang 0004, Zhehao Lin, Yuhang Li 0002, Chengyuan Huang, Xiaohu Xu, Zhongming Ji, Shengju Zhang, Lingkun Meng, Rong Gu 0001, Guihai Chen, Chen Tian 0001 |
IEEE Trans. Netw. | 6 |
| 2024 | Research on Intelligent Recognition Algorithm of Container Numbers in Ports Based on Deep Learning
Zhehao Lin |
ICIC (7) | 1 |