VLDB 2026 Research / reviewers in the wild / expert
ChonLam Lao
dblp:291/3926
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-0214-664XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Your network doesn't end at the NIC: A case for unifying the inter-host and intra-host networks in (AI) datacentersabstractModern ML workloads increasingly rely on direct communication between host devices—such as GPUs, NVMe SSDs, and DRAM—spanning intra-host and inter-host networks. However, today's intra-host network lacks hardware-level primitives for routing across heterogeneous interconnects, hindering efficient use of alternative paths and leading to sub-optimal performance under failures or congestion. Furthermore, the inter-host network treats the NIC as the endpoint, with intra-host interconnects like PCIe running oblivious to inter-host network protocols. This prevents leveraging multiple paths for communication between host devices across different servers. To address these limitations, we propose expanding the datacenter network layer to encompass the intra-host network, making intra-host devices first-class network endpoints. Our scheme envisions hardware-level routing and forwarding across multiple intra-host interconnects and makes intra-host devices visible to the inter-host network. This unified approach provides a principled foundation for robust, efficient peer-to-peer communication between storage and compute hardware devices in AI datacenters. Raj Joshi, Saksham Agarwal, ChonLam Lao, Minlan Yu |
HotNets | 3 |
| 2025 | eTran: Extensible Kernel Transport with eBPF
Zhongjie Chen, Qingkai Meng 0001, ChonLam Lao, Fengyuan Ren, Minlan Yu, Yang Zhou 0008 |
NSDI | 3 |
| 2025 | Astral: A Datacenter Infrastructure for Large Language Model Training at ScaleabstractThe flourishing of Large Language Models (LLMs) calls for increasingly ultra-scale training. In this paper, we share our experience in designing, deploying, and operating our novel Astral datacenter infrastructure, along with operational lessons and evolutionary insights gained from its production use. Astral has three important innovations: (i) a same-rail interconnection network architecture on tier-2, which enables the scaling of LLM training. To physically deploy this high-density infrastructure, we introduce a distributed high-voltage direct current power system and a new air-liquid integrated cooling system. (ii) a full-stack monitoring system featuring cross-host and hierarchical logging correlation, which diagnoses failures at scale and precisely localizes root causes. (iii) an operator-granular forecasting component Seer that efficiently generates operator execution timelines with acceptable accuracy, aiding in fault diagnosis, model tuning, and network architecture upgrading. Astral infrastructure has been gradually deployed over 18 months, supporting LLM training and inference for multiple customers. Qingkai Meng 0001, Zhenhui Zhang, ChonLam Lao, Chengyuan Huang, Baojia Li 0002, Weizhen Dang, Zitong Lin, Yuanyuan Gong, Chunzhi He, Xiaoyuan Hu, Yinben Xia, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Kun Yang 0001, Gianni Antichi, Guihai Chen, Chen Tian 0001 |
SIGCOMM | 4 |
| 2024 | THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic Compression
Minghao Li 0003, Ran Ben-Basat, Shay Vargaftik, ChonLam Lao, Kevin Xu, Michael Mitzenmacher, Minlan Yu |
NSDI | 4 |
| 2023 | A Generic Service to Provide In-Network Aggregation for Key-Value StreamsabstractKey-value stream aggregation is a common operation in distributed systems, which requires intensive computation and network resources. We propose a generic in-network aggregation service for key-value streams, ASK, to accelerate the aggregation operations in diverse distributed applications. ASK is a switch-host co-designed system, where the programmable switch provides a best-effort aggregation service, and the host runs a daemon to interact with applications. ASK makes in-depth optimization tailored to traffic characteristics, hardware restrictions, and network unreliable natures: it vectorizes multiple key-value tuples’ aggregation of one packet in one switch pipeline pass, which improves the per-host’s goodput; it develops a lightweight reliability mechanism for key-value stream’s asynchronous aggregation, which guarantees computation correctness; it designs a hot-key agnostic prioritization for key-skewed workloads, which improves the switch memory utilization. We prototype ASK and use it to support Spark and BytePS. The evaluation shows that ASK could accelerate pure key-value aggregation tasks by up to 155 times and big data jobs by 3-5 times, and be backward compatible with existing INA-empowered distributed training solutions with the same speedup. Yongchao He, Wenfei Wu, Yanfang Le, Ming Liu 0027, ChonLam Lao |
ASPLOS (2) | 5 |
| 2023 | Preemptive Switch Memory Usage to Accelerate Training Jobs with Shared In-Network AggregationabstractRecent works introduce In-Network Aggregation (INA) for distributed training (DT), which moves the gradient summation into network programmable switches. INA can reduce the traffic volume and accelerate communication in DT jobs. However, switch memory is a scarce resource, unable to support massive DT jobs in data centers, and existing INA solutions have not utilized switch memory to the best extent. We propose DSA, an Efficient Data-Plane switch memory Scheduler for in-network Aggregation. DSA introduces preemption to the switch memory management for INA jobs. In the data plane, DSA allows gradient tensors with high priority to preempt the switch aggregators (basic computation unit in INA) from tensors with low priority, which avoids an aggregator wasting time in idle. In the control plane, DSA devises a priority policy which assigns high priority to gradient tensors that benefit overall job efficiency more, e.g., communication-intensive jobs. We prototype DSA and experiments show that DSA can improve the average JCT by up to 1.35x compared with baseline solutions. Yuxuan Qin, ChonLam Lao, Yanfang Le, Wenfei Wu |
ICNP | 3 |
| 2023 | In-Network Key-Value Cache with LinearizabilityabstractRecently, In-Network Cache (INC) systems have been proposed to promote the performance of remote storage systems. INC offloads cache onto programmable switches between the clients and the servers, responding to clients’ data queries within a sub-RTT time. However, most existing INC solutions do not take applications’ linearizability requirement into consideration, which could lead to query errors and storage state errors in the runtime. We propose a new INC system — NetKV-L, which preserves the high-performance I/O without compromising the linearizability. NetKV-L devises a sequentiality enforcement mechanism, a PSN correction mechanism, and a response memorization mechanism to guarantee linearizability under possible unreliable network conditions. Our prototype and experiments show that NetKV-L could achieve almost the same performance as the state-of-the-art systems while additionally guaranteeing linearizability. Yuxuan Qin, Weize Gao, ChonLam Lao, Wenfei Wu |
ICPADS | 3 |
| 2021 | ATP: In-network Aggregation for Multi-tenant Learning
ChonLam Lao, Yanfang Le, Kshiteej Mahajan, Wenfei Wu, Aditya Akella, Michael M. Swift |
NSDI | 1 |