Liangyu Zhao

dblp:33/1876 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
7 papers
High-performance computing · 34% Cloud and datacenter computing · 18% Interconnection networks and networks-on-chip · 16%
Artificial intelligence
3 papers
Optimization for machine learning · 82% Efficient and distributed learning · 18%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
collective communication
3.442026
ForestColl: Throughput-Optimal Collective Communications on Heterogeneous Network Fabrics · NSDI 2026
Efficient Direct-Connect Topologies for Collective Communications · NSDI 2025
Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow Problem · SIGCOMM 2024
Interconnection networks and networks-on-chip › interprocessor communication
all-to-all communication
1.012026
FAST: An Efficient Scheduler for All-to-All GPU Communication · NSDI 2026
GPUs and heterogeneous computing
GPU communication
1.012026
FAST: An Efficient Scheduler for All-to-All GPU Communication · NSDI 2026
Cloud and datacenter computing
inference serving
0.912025
NanoFlow: Towards Optimal Large Language Model Serving Throughput · OSDI 2025
Cloud and datacenter computing › inference serving
LLM serving
0.912025
NanoFlow: Towards Optimal Large Language Model Serving Throughput · OSDI 2025
Interconnection networks and networks-on-chip
network topology
0.912025
Efficient Direct-Connect Topologies for Collective Communications · NSDI 2025
Parallel and multicore computing › parallel scheduling
communication scheduling
0.812024
Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow Problem · SIGCOMM 2024
Electronic design automation › physical design › routing
multicommodity flow
0.812024
Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow Problem · SIGCOMM 2024
Machine learning › Optimization for machine learning › model-based optimization
bayesian optimization
0.512021
AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the Fly · ICLR 2021
Machine learning › Optimization for machine learning
hyperparameter optimization
0.512021
AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the Fly · ICLR 2021
Machine learning › Optimization for machine learning
learning rate schedule
0.512021
AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the Fly · ICLR 2021
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.412019
Nexus: a GPU cluster engine for accelerating DNN-based video analysis · SOSP 2019
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference
DNN inference scheduling
0.412019
Nexus: a GPU cluster engine for accelerating DNN-based video analysis · SOSP 2019
GPUs and heterogeneous computing › GPU resource management
GPU cluster resource management
0.412019
Nexus: a GPU cluster engine for accelerating DNN-based video analysis · SOSP 2019
GPUs and heterogeneous computing
GPU scheduling
0.412019
Nexus: a GPU cluster engine for accelerating DNN-based video analysis · SOSP 2019
Electronic design automation › high-level synthesis
scheduling
0.312026
FAST: An Efficient Scheduler for All-to-All GPU Communication · NSDI 2026
Machine learning › Efficient and distributed learning
distributed training
0.212024
Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow Problem · SIGCOMM 2024

Methods — techniques the papers use, named apart from their topics

traffic engineering · 1.5multicommodity flow · 0.8multi-commodity flow · 0.8co-scheduling · 0.8DNN fragment execution · 0.8learning rate schedule · 0.5bayesian optimization · 0.5
YearPublicationVenuePosition
2026 FAST: An Efficient Scheduler for All-to-All GPU Communication
Yiran Lei, Dongjoo Lee 0001, Liangyu Zhao, Daniar Kurniawan, Chanmyeong Kim, Heetaek Jeong, Changsu Kim 0004, Hyeonseong Choi, Liangcheng Yu, Arvind Krishnamurthy, Justine Sherry, Eriko Nurvitadhi
NSDI3
2026 ForestColl: Throughput-Optimal Collective Communications on Heterogeneous Network Fabrics
Liangyu Zhao, Saeed Maleki, Yuanhong Wang, Zezhou Wang, Hossein Pourreza, Arvind Krishnamurthy
NSDI1
2025 Efficient Direct-Connect Topologies for Collective Communications
Liangyu Zhao, Siddharth Pal, Tapan Chugh, Weiyang Wang, Jason Fantl, Prithwish Basu, Joud Khoury, Arvind Krishnamurthy
NSDI1
2025 NanoFlow: Towards Optimal Large Language Model Serving Throughput
Kan Zhu, Yilong Zhao 0002, Liangyu Zhao, Gefei Zuo, Yile Gu, Dedong Xie, Zihao Ye 0001, Keisuke Kamahori, Chien-Yu Lin, Ziren Wang, Stephanie Wang, Arvind Krishnamurthy, Baris Kasikci
OSDI4
2025 SHAA: Spatial Hybrid Attention Network With Adaptive Cross-Entropy Loss Function for UAV-View Geo-Localization
abstract
Cross-view geo-localization provides an offline visual positioning strategy for unmanned aerial vehicles (UAVs) in Global Navigation Satellite System (GNSS)-denied environments. However, it still faces the following challenges, leading to suboptimal localization performance: 1) Existing methods primarily focus on extracting global features or local features by partitioning feature maps, neglecting the exploration of spatial information, which is essential for extracting consistent feature representations and aligning images of identical targets across different views. 2) Cross-view geo-localization encounters the challenge of data imbalance between UAV and satellite images. To address these challenges, the Spatial Hybrid Attention Network with Adaptive Cross-Entropy Loss Function (SHAA) is proposed. To tackle the first issue, the Spatial Hybrid Attention (SHA) method employs a Spatial Shift-MLP (SSM) to focus on the spatial geometric correspondences in feature maps across different views, extracting both global features and fine-grained features. Additionally, the SHA method utilizes a Hybrid Attention (HA) mechanism to enhance feature extraction diversity and robustness by capturing interactions between spatial and channel dimensions, thereby extracting consistent cross-view features and aligning images. For the second challenge, the Adaptive Cross-Entropy (ACE) loss function incorporates adaptive weights to emphasize hard samples, alleviating data imbalance issues and improving training effectiveness. Extensive experiments on widely recognized benchmarks, including University-1652, SUES-200, and DenseUAV, demonstrate that SHAA achieves state-of-the-art performance, outperforming existing methods by over 3.92%.
Nanhua Chen, Dongshuo Zhang, Kai Jiang 0001, Yeqing Zhu, Tai-Shan Lou, Liangyu Zhao
IEEE Trans. Circuits Syst. Video Technol.7
2024 Efficient all-to-all Collective Communication Schedules for Direct-connect Topologies
abstract
The all-to-all collective communications primitive is widely used in machine learning (ML) and high performance computing (HPC) workloads, and optimizing its performance is of interest to both ML and HPC communities. All-to-all is a particularly challenging workload that can severely strain the underlying interconnect bandwidth at scale. This paper takes a holistic approach to optimize the performance of all-to-all collective communications on supercomputer-scale direct-connect interconnects. We address several algorithmic and practical challenges in developing efficient and bandwidth-optimal all-to-all schedules for any topology and lowering the schedules to various runtimes and interconnect technologies. We also propose a novel topology that delivers near-optimal all-to-all performance.
Prithwish Basu, Liangyu Zhao, Jason Fantl, Siddharth Pal, Arvind Krishnamurthy, Joud Khoury
HPDC2
2024 PLP-SLAM: Point-Line-Plane Simultaneous Localization and Mapping
abstract
For indoor environments, prior point-based visual SLAM cannot be processed in real time under low texture and illumination. To address this issue, this work proposes PLP-SLAM (Point-Line-Plane-SLAM) with RGB-D camera. Firstly, point and line features are detected in RGB images. For line features, establish length suppression and near line merge strategy to improve the line extraction quality. Secondly, plane features are extracted based on agglomerative hierarchical clustering method in point cloud obtained by RGB-D camera. Point clouds are divided into several nodes, unlike prior methods spend a lot of time to estimate the normal vector for each individual point, this work assumes that points within each node sharing the same plane normal vector, which can significantly improve the computational efficiency. Thirdly, sparse maps including points, lines and planes are established, meanwhile the scenes are reconstructed by creating the dense maps to show plan features directly. Finally, the performance of proposed method is compared against the state-of-the-art SLAM on public datasets to evaluate the pose estimation. All modules are run in real-time on a CPU, experiments clarify that PLP-SLAM can significantly enhance the robustness of 6DoF pose of the camera and simultaneously creating more detailed maps of the environment.
Yeqing Zhu, Liangyu Zhao, Qingjie Zhao, Zhenyu Wu 0001, Hongming Shen, Danwei Wang
ICARCV2
2024 Rethinking Machine Learning Collective Communication as a Multi-Commodity Flow Problem
abstract
Cloud operators utilize collective communication optimizers to enhance the efficiency of the single-tenant, centrally managed training clusters they manage. However, current optimizers struggle to scale for such use cases and often compromise solution quality for scalability. Our solution, TE-CCL, adopts a traffic-engineering-based approach to collective communication. Compared to a state-of-the-art optimizer, TACCL, TE-CCL produced schedules with 2× better performance on topologies TACCL supports (and its solver took a similar amount of time as TACCL's heuristic-based approach). TECCL additionally scales to larger topologies than TACCL. On our GPU testbed, TE-CCL outperformed TACCL by 2.14× and RCCL by 3.18× in terms of algorithm bandwidth.
Xuting Liu 0003, Behnaz Arzani, Siva Kesava Reddy K., Liangyu Zhao, Vincent Liu 0001, Srikanth Kandula, Luke Marshall
SIGCOMM4
2023 Adaptive fast desensitized ensemble Kalman filter for uncertain systems
Nanhua Chen, Liangyu Zhao, Tai-Shan Lou, Chuanjun Li
Signal Process.2
2021 AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the Fly
Tianyi Zhou 0001, Liangyu Zhao, Yibo Zhu 0001, Chuanxiong Guo, Marco Canini, Arvind Krishnamurthy
ICLR3
2019 Nexus: a GPU cluster engine for accelerating DNN-based video analysis
abstract
We address the problem of serving Deep Neural Networks (DNNs) efficiently from a cluster of GPUs. In order to realize the promise of very low-cost processing made by accelerators such as GPUs, it is essential to run them at sustained high utilization. Doing so requires cluster-scale resource management that performs detailed scheduling of GPUs, reasoning about groups of DNN invocations that need to be co-scheduled, and moving from the conventional whole-DNN execution model to executing fragments of DNNs. Nexus is a fully implemented system that includes these innovations. In large-scale case studies on 16 GPUs, when required to stay within latency constraints at least 99% of the time, Nexus can process requests at rates 1.8-12.7X higher than state of the art systems can. A long-running multi-application deployment stays within 84% of optimal utilization and, on a 100-GPU cluster, violates latency SLOs on 0.27% of requests.
Haichen Shen, Lequn Chen 0001, Liangyu Zhao, Bingyu Kong, Matthai Philipose, Arvind Krishnamurthy, Ravi Sundaram
SOSP4
2006 Visualization of Dynamic Brain Activities Based on the Single-Trial MEG and EEG Data Analysis
Jianting Cao, Liangyu Zhao, Andrzej Cichocki
ISNN (2)2