VLDB 2026 Research / reviewers in the wild / expert
Hongliang Li 0003
dblp:91/1905-3
· DBLP profile ↗
25ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0003-4377-0550ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 6 first-author · 8 since 2021Computer networks · 6 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rehabilitating over Recomputing: A Novel Failure Recovery Method for Large Model Training
Hongliang Li 0003, Jie Wu 0001, Zhewen Xu, Hairui Zhao 0002, Haixiao Xu |
INFOCOM | 2 |
| 2026 | Sift: Channel-Wise Historical Embedding for High Efficiency Distributed Graph Neural Network Training with Accuracy GuaranteeabstractDistributed Graph Neural Network (DGNN) is a powerful tool in large-scale graph representation learning. However, high data-transfer overhead among workers in a DGNN training job confines its scalability and thus the overall performance. Vertex-wise historical embedding methods have demonstrated high potential to alleviate the problems, but still suffer from severe accuracy loss and limited performance scalability, which has been attributed to the information loss of critical channels in historical vertices and redundant information in local channels. This article explores the optimization of channel level and construct a quantitative accuracy model for channel-wise historical embedding. We propose Sift, a novel DGNN training framework, supporting channel-wise partial historical embedding with accuracy guarantee. Sift has three components: a historical embedding evaluator with channel-wise quantitative accuracy model, a sawtooth-like matrix rearrangement for accelerating message passing, and a hybrid parallel framework for overlapping communication overhead. Comprehensive experimental results show that Sift achieves near-linear parallel convergence speedup, outperforming the state-of-the-art baselines by up to 72% in total training performance and up to 21% in convergence speed. Zhewen Xu, Hongliang Li 0003, Junze Han, Hengshan Yue, Hairui Zhao 0002, Dongyuan Tian, Zijian Li 0007, Xiaohui Wei 0002 |
ACM Trans. Archit. Code Optim. | 2 |
| 2025 | ArrayPipe: Introducing Job-Array Pipeline Parallelism for High Throughput Model Exploration
Hairui Zhao 0002, Hongliang Li 0003, Jie Wu 0001, Zhewen Xu, Xiang Li 0197, Haixiao Xu |
INFOCOM | 2 |
| 2025 | FlexPipe: Maximizing Training Efficiency for Transformer-based Models with Variable-Length Inputs
Hairui Zhao 0002, Hongliang Li 0003, Zizhong Chen |
USENIX ATC | 3 |
| 2025 | Alleviating straggler impacts for data parallel deep learning with hybrid parameter update
Hongliang Li 0003, Hairui Zhao 0002, Zhewen Xu |
Future Gener. Comput. Syst. | 1 |
| 2025 | Convergence-aware optimal checkpointing for exploratory deep learning training jobs
Hongliang Li 0003, Hairui Zhao 0002, Xiang Li 0197, Haixiao Xu |
Future Gener. Comput. Syst. | 1 |
| 2025 | Harnessing dynamic graph differential operators for efficient data-driven wind prediction
Xiaohui Wei 0002, Zhewen Xu, Hongliang Li 0003, Jieyun Hao, Hengshan Yue, Changzheng Liu |
GeoInformatica | 3 |
| 2024 | Visage: Visual-Aware Generation of Adversarial Examples in Black-Box for Text Classification
Hairui Zhao 0002, Hongliang Li 0003 |
NLPCC (4) | 3 |
| 2024 | HiRM: Hierarchical resource management for earth system models on many-core clusters
Zhewen Xu, Xiaohui Wei 0002, Jieyun Hao, Hongliang Li 0003, Zhaohui Ding |
CCF Trans. High Perform. Comput. | 5 |
| 2024 | DGFormer: a physics-guided station level weather forecasting model with dynamic spatial-temporal graph neural network
Zhewen Xu, Xiaohui Wei 0002, Jieyun Hao, Junze Han, Hongliang Li 0003, Changzheng Liu, Zijian Li 0007, Dongyuan Tian, Nong Zhang |
GeoInformatica | 5 |
| 2024 | Interference-aware opportunistic job placement for shared distributed deep learning clusters
Hongliang Li 0003, Hairui Zhao 0002, Xiang Li 0197, Haixiao Xu |
J. Parallel Distributed Comput. | 1 |
| 2023 | ExplSched: Maximizing Deep Learning Cluster Efficiency for Exploratory JobsabstractResource management for Deep Learning (DL) clusters is essential for system efficiency and model training quality. Existing schedulers provided by DL frameworks are mostly adaptations from traditional HPC clusters and usually work on jobs’ makespan, assuming that DL training jobs finish completely. Unfortunately, it is reported that a fair amount of training jobs are exploratory jobs and often finish unsuccessfully (over 30%) in production clusters. This is due to the distinct characteristic of Deep Neural Network (DNN) training that it is an exploratory process of frequent user interventions, such as adjusting model structures, tuning hyperparameters, and exploring feature validity. Existing DL cluster schedulers using offline algorithms are not suitable for exploratory jobs when unexpected early terminations can cause noticeable resource waste. Moreover, DL training jobs are iterative and usually yield diminishing returns as they progress. Equally allocating resource among training iterations is not efficient, especially when dealing with exploratory jobs where it can worsen the degradation of system efficiency. The fundamental goal of a DL training job is to gain model quality improvement, usually indicated by the loss reduction (job profit) of a DNN model. This paper introduces a novel scheduling problem for exploratory jobs that seeks to maximize the overall training profit of a DL cluster. We propose ExplSched, an online scheduling solution based on the primal-dual framework, resulting in a competitive ratio of 2α that belongs to O(ln n). It uses a resource price function that emphasizes the importance of job profit to resource consumption ratio to make quick resource allocation decisions. Experimental results show that ExplSched achieved an average system utility improvement of 87.28% compared with other related work. Hongliang Li 0003, Hairui Zhao 0002, Zhewen Xu, Xiang Li 0197, Haixiao Xu |
CLUSTER | 1 |
| 2021 | Coordinated process scheduling algorithms for coupled earth system modelsabstractAbstract It is becoming increasingly significant for humans to predict and understand future climate changes using coupled climate system models. Although the performance and scalability of individual physical components have improved over the past few years, coupled climate systems still suffer from low efficiency. This paper focuses on the process scheduling problem for the widely applied coupled earth system model (CESM). The proposed resource allocation strategies allow components to execute on a compromised suboptimal setup and still maintain approximately the best parallel speedup. With this flexible resource allocation strategy, we further propose a coordinated process scheduling algorithm (CPSA). More notably, we propose an upgraded version called CPSA‐B, which makes efficient resource sharing configurations, including resource allocation and process layout of components. We integrate CPSA and CPSA‐B as pre‐arrangement tools into the CESM program and deploy them on the Huawei Kunpeng platform. The speedup curves of the CESM components are prepared in advance, based on sampling tests. Experimental data show that CPSA‐B reduces up to 58% of the execution time compared with the CESM default strategy. The algorithm has low complexity and can efficiently find solutions for large input sizes. Xiaohui Wei 0002, Zhewen Xu, Hongliang Li 0003, Zhaohui Ding |
Concurr. Comput. Pract. Exp. | 3 |
| 2020 | CPSA: A Coordinated Process Scheduling Algorithm for Coupled Earth System ModelabstractCoupled climate system models are important tools for climatologists to predict and understand future climate. These models are usually resource-consuming due to the large number of processors required and long execution time. Although the performance and scalability of individual physical system model have been improved over the past years, coupled climate systems still suffer from low efficiency when sharing resource across models. This paper focuses on the process scheduling strategy of Coupled Earth System Model (CESM), a widely applied coupled system model. Instead of pursuing best speedup efficiency for individual component, the proposed resource allocation strategy allows components to execute on compromised sub-optimal setup and still maintains relatively high parallel speedup. With this flexible resource allocation strategy, we further propose a Coordinated Process Scheduling Algorithm (CPSA) to make efficient resource sharing configurations, including resource allocation and process layout of components. We integrate CPSA as a tool into CESM program, and deploy it on Huawei Kunpeng Platform. Speedup curves of CESM components are prepared in advance based on sampling tests. Experimental data show that our algorithm reduces up to 52.6% of execution time compared with CESM default strategy. We also present simulation data to show that our algorithm is efficient for the platforms with up to a million cores. Hongliang Li 0003, Zhewen Xu, Fangyu Tang, Xiaohui Wei 0002, Zhaohui Ding |
ICCCN | 1 |
| 2020 | Reducing Fault-tolerant Overhead for Distributed Stream Processing with Approximate BackupabstractThe stream processing model continuously processes online data in an on-pass fashion that can be more vulnerable to failures than other offline-data processing schemes. Checkpoint-based fault-tolerant methods have been widely used to enhance the reliability of stream processing systems. To ensure exact data recoveries upon failures, full-backup mechanisms are used to store a complete copy of data, which introduces substantial runtime overhead and increases output latency. In the meantime, a wide range of online processing applications prefer quick-and-dirty results with a slight degradation inaccuracy to delayed exact results. This paper introduces a novel approximate fault-tolerant problem (OAFP) with the objective of reducing the failure-free fault-tolerant overhead and ensuring user-defiled output accuracy requirement upon failure at the same time. We present an approximate fault-tolerant scheme based on sampling backup mechanism and study the trade-off between fault-tolerant overhead and output accuracy in stream processing systems. We proposed two algorithms to compute backup plans for both single-node failure and correlated failure scenarios. Extensive experiments with different types of stream topologies are conducted on our simulator to verify the correctness and effectiveness of our approach. We prove our solution guarantees the output accuracy requirement with minimum FT latency for directed acyclic graph (DAG) stream topologies with single-node failures. Yuan Zhuang 0003, Xiaohui Wei 0002, Hongliang Li 0003, Mingkai Hou, Yundi Wang |
ICCCN | 3 |
| 2019 | An optimal checkpointing model with online OCI adjustment for stream processing applicationsabstractSummary Checkpoint‐based fault‐tolerant (FT) methods have been widely used to enhance the reliability of stream processing systems, but a checkpointing process usually introduces considerable overhead. It is a critical issue to choose the optimal checkpoint interval (OCI) that maximizes the processing efficiency. Traditional OCI models consider the recovery time equals to the execution time from the last checkpoint to the failure moment. However, for stream processing jobs, the recovery time is related to reprocessing workloads, depending on the real‐time input data before a failure. A new model is needed to choose the OCI for stream processing applications. Moreover, the input data rate of a stream processing job fluctuates over time. To solve these problems, we present a novel DSPS OCI (DOCI) model in this paper. We prove that it maximizes the processing efficiency for a given time. We propose an approach to dynamically adjust the OCI for an application to accommodate the workload fluctuations. We conduct simulation experiments to verify the effectiveness of our DOCI model and the efficiency of the online OCI adjustment algorithm. Experimental results with a real‐world dataset show that DOCI achieves an improvement on system efficiency by up to 32%, compared with existing FT approaches. Yuan Zhuang 0003, Xiaohui Wei 0002, Hongliang Li 0003, Xubin He |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | Reducing the synchronizing communication overhead for distributed graph-parallel computingabstractA number of graph-parallel computing abstractions have been proposed to address the needs of solving complex and large-scale graph computing. However, unnecessary and excessive communication and state sharing between nodes in these frameworks not only reduce the network efficiency but may also caus e decrease in runtime performance. In this paper, we propose a mechanism called LightGraph, which reduces the synchronizing communication overhead for distributed graph-parallel computing abstractions. Besides identifying and eliminating the redundant synchronizing communications in existing systems, in order to minimize the required synchronizing communications LightGraph also proposes an edge direction-aware graph partitioning strategy. This new graph partitioning strategy optimally isolates the outgoing edges from the incoming edges of a vertex. We have conducted extensive experiments using real-world data, and our results verified the effectiveness of LightGraph. For example compared to PowerGraph LightGraph can not only reduce up to 31.5% synchronizing communication overhead for intra-graph synchronizations, but also cut up to 16.3% runtime for PageRank running on Livejournal dataset. Yue Zhao 0014, Kenji Yoshigoe, Hongliang Li 0003, Ke Xiong 0001 |
Intell. Data Anal. | 3 |
| 2019 | Pec: Proactive Elastic Collaborative Resource Scheduling in Data Stream ProcessingabstractIn the Distributed Parallel Stream Processing Systems (DPSPS), elastic resource allocation allows applications to dynamically response to workload fluctuations. However, resource provisioning can be particularly challenging, due to the unpredictability of the workload. In addition, unlike CPU resources, bandwidth resources are often ignored in resource allocation. Moreover, resource allocation and resource placement are considered separately. In this paper, we investigate the proactive elastic resource scheduling problem for computation-intensive and communication-intensive applications, which aims at meeting the latency requirement with the minimal energy cost, and propose a dynamic collaborative strategy from the systemic perspective. Specifically, we first model a collaborative workload prediction pattern to accurately predict the upcoming workload, and construct a latency estimation model to estimate the latency of the application. Then, we design an energy-efficient resource pre-allocation method, in which the CPU frequency adjustment and the stability of resource reconfigurations are both considered. Finally, we present a communication-aware resource placement approach. Simulation results show that, compared with the reactive strategies, our strategy achieves an obviously better latency performance, and effectively avoids unnecessary resource adjustments. Meanwhile, the energy consumption is about saved by 50 percent on average, and the communication cost is maintained at a very low level of 4 percent. Xiaohui Wei 0002, Xiang Li 0197, Xingwang Wang 0003, Shang Gao 0005, Hongliang Li 0003 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2018 | An Optimal Checkpointing Model with Online OCI Adjustment for Stream Processing ApplicationsabstractCheckpoint-based fault tolerant method has been widely used to enhance the reliability of Distributed Stream Processing Engines (DSPEs), but a checkpointing process usually introduces considerable overhead. It is a critical issue to choose the Optimal Checkpoint Interval (OCI) that maximizes the processing efficiency. Traditional OCI models consider the recovery time only related to the execution time from the last checkpoint to the moment of the failure. They are not suitable for stream processing jobs because the recovery time is related to the reprocessing workload, which depends on the realtime input data before a failure. A new model is needed to choose the OCI for stream processing applications. Moreover, the input data rate of an stream processing job fluctuates over time. The OCI of an application should also be adjusted dynamically according to the input workload. To solve these problems, we present a novel DSPS Optimal Checkpoint Interval (DOCI) model in this paper. We prove that it maximizes the processing efficiency for a given time period. We propose an approach to dynamically adjust the OCI for an application to accommodate the realtime workload fluctuations. We conduct simulation experiments to verify the effectiveness of DOCI model and the efficiency of the online OCI adjustment algorithm. Experimental results with a real-world dataset show DOCI achieves an improvement on system efficiency by up to 40%, comparing with existing fault-tolerant approaches. Yuan Zhuang 0003, Xiaohui Wei 0002, Hongliang Li 0003, Xubin He |
ICCCN | 3 |
| 2018 | A Task Allocation Method for Stream Processing with Recovery Latency Constraint
Hongliang Li 0003, Jie Wu 0001, Xiang Li 0197, Xiaohui Wei 0002 |
J. Comput. Sci. Technol. | 1 |
| 2017 | Task Allocation for Stream Processing with Recovery Latency GuaranteeabstractStream processing applications continuously process large amounts of online streaming data in real-time or near real-time. They have strict latency constraints, but they are also vulnerable to failures. Failure recoveries may slow down the entire processing pipeline and break latency constraints. Upstream backup is one of the most widely applied fault-tolerant schemes for stream processing systems. It introduces complex backup dependencies to tasks, and increases the difficulty of controlling recovery latencies. Moreover, when dependent tasks are located on the same processor, they fail at the same time in processor-level failures, bringing extra recovery latencies that increase the impacts of failures. This paper presents a correlated failure effect model to describe the recovery latency of a stream topology in processor-level failures for an allocation plan. We introduce a Recovery-latency-aware Task Allocation Problem (RTAP) that seeks task allocation plans for stream topologies that will achieve guaranteed recovery latencies. We present a heuristic algorithm with a computational complexity of O(nlog^2n) to solve the problem. Extensive experiments were conducted to verify the correctness and effectiveness of our approach. Hongliang Li 0003, Jie Wu 0001, Xiang Li 0197, Xiaohui Wei 0002 |
CLUSTER | 1 |
| 2017 | Integrated recovery and task allocation for stream processingabstractStream processing applications continueously process large-scale data streams online. The throughput of a stream processing application must match the input rate to avoid loss of data. Failures affect throughput because a task failure can suspend itself from producing new data and can even cause an application-level halt. The key motivation of this work is to mitigate the performance degradation caused by task-level failures. We introduce a novel Integrated Recovery Model (IRM) that allows resource sharing among both failure-free tasks and recovering tasks on a processor. The failure-free tasks slow down to accelerate a task recovery rather than suspending their actions and waiting for the recovery to finish; waiting causes a complete halt of the application. In this way, the recovery is seamless and does not suspend the entire system. The performance slowdown is related to both the failure-free processing cost and recovery cost on each processor. Moreover, the recovery cost of a task is related to the Fault-Tolerant Configuration (FTC) of the stream application. This paper introduces a novel task allocation problem that, given an FTC, can constrain processing performance during recoveries (i.e. throughput slowdown ratio) while minimizing the amount of resource occupied. We propose both a greedy algorithm and a heuristic algorithm with computational complexities of O(n log n) and O(n log2n), respectively, to solve the problem. Extensive experiments verify the correctness and effectiveness of our approach. Our approach enables continuous processing results and seamless failure recoveries with a constrained slowdown ratio. Hongliang Li 0003, Jie Wu 0001, Xiang Li 0197, Xiaohui Wei 0002, Yuan Zhuang 0003 |
IPCCC | 1 |
| 2017 | Minimum Backups for Stream Processing With Recovery Latency GuaranteesabstractThe stream processing model continuously processes online data in an on-pass fashion that can be more vulnerable to failures than other big-data processing schemes. Existing fault-tolerant (FT) approaches have been presented to enhance the reliability of stream processing systems. However, the fundamental tradeoff between recovery latency and FT overhead is still unclear, so these scheme cannot provide recovery latency guarantees. This paper introduces the FT Configuration (FTC) problem and presents a solution for guaranteed recovery latency with minimum backups. A failure effect model is presented to describe the relationship between recovery latency and FTC (the amount and locations of backups). With this model, we design an algorithm to compute FTCs for different types of stream topologies according to recovery latency requirements. Extensive experiments are conducted to verify the correctness and effectiveness of our approach. We prove that our algorithm guarantees recovery latencies for all directed acyclic graph (DAG) stream topologies. For line(s) and tree topologies, our algorithm solves the FTC problem with a time complexity of O(N). For a general DAG topology, a heuristic function is used to generate FTCs. This causes fewer than 10% more backups on average compared to the optimal solution with a time complexity of O(N2). Hongliang Li 0003, Jie Wu 0001, Xiang Li 0197, Xiaohui Wei 0002 |
IEEE Trans. Reliab. | 1 |
| 2014 | MapReduce delay scheduling with deadline constraintabstractSUMMARY MapReduce programming paradigm has been widely applied to solve large‐scale data‐intensive problems. Intensive studies of MapReduce scheduling have been carried out to improve MapReduce system performance. Delay scheduling is a common way to achieve high data locality and system performance. However, inappropriate delays can lead to low system throughput and potentially break the original job priority constraints. This paper proposes a deadline‐enabled delay (DLD) scheduling algorithm that optimizes job delay decisions according to real‐time resource availability and resource competition, while still meets job deadline constraints. Experimental results illustrate that the resource availability estimation method of DLD is accurate (92%). Compared with other approaches, DLD reduces job turnaround time by 22% in average while keeping a high locality rate (88%).Copyright © 2013 John Wiley & Sons, Ltd. Hongliang Li 0003, Xiaohui Wei 0002, Qingwu Fu |
Concurr. Comput. Pract. Exp. | 1 |
| 2014 | Topology-Aware Partial Virtual Cluster Mapping Algorithm on Shared Distributed InfrastructuresabstractNovel virtualized HPC centers provide isolated and configurable Virtual Clusters (VC) on shared distributed infrastructures as execution environments for parallel and distributed applications. These VCs are usually customized and deployed per job in runtime. Allocating physical resources for VC is known as Virtual Cluster Mapping (VCM) problem, which is a critical issue that affects both performance of the VC and resource utilization of the system. Most previous works treat all Virtual Machines (VMs) in a VC request equally. However, because sub-jobs in a parallel job usually perform different roles, the corresponding VMs in a VC that execute these sub-jobs respectively should have different levels of importance. Based on this argument, this paper introduces the concept of partial VC mapping in contrast to the full mapping methodology in the current literatures. To fulfill partial mapping, the important backbone communication structure of parallel job called Communication Skeleton (CS) is derived based on the network topology among virtual nodes. To generate the CS of a job, mechanisms for evaluating the importance of nodes are proposed. Eventually, a Topology-aware Partial Virtual Cluster Mapping algorithm (TOP-VCM) is proposed which is based on sub-graph isomorphism detection. TOP-VCM can fully satisfy the nodes/links requirements in CS to ensure the execution performance with only slight degradation of other trivial nodes/links to significantly reduce the mapping difficulty. Simulation results have shown that TOP-VCM has significantly improved the total revenue, the utilization of physical resources and the performance of mapping algorithm while satisfying the VC requirements. Xiaohui Wei 0002, Hongliang Li 0003, Kun Yang 0001, Lei Zou 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |