EDBT 2026 Demo / reviewers in the wild / expert
Xiang Li 0197
dblp:40/1491-197
· DBLP profile ↗
13ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-4930-0878ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 since 2021Computer networks · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ArrayPipe: Introducing Job-Array Pipeline Parallelism for High Throughput Model Exploration
Hairui Zhao 0002, Hongliang Li 0003, Jie Wu 0001, Zhewen Xu, Xiang Li 0197, Haixiao Xu |
INFOCOM | 7 |
| 2025 | GraphFT: A Lightweight Fault-tolerant Framework for Iterative Graph Processing
Xiaohui Wei 0002, Mengting Zhou, Nan Jiang 0013, Xiang Li 0197, Hengshan Yue |
WASA (3) | 4 |
| 2025 | Convergence-aware optimal checkpointing for exploratory deep learning training jobs
Hongliang Li 0003, Hairui Zhao 0002, Xiang Li 0197, Haixiao Xu |
Future Gener. Comput. Syst. | 5 |
| 2025 | ResCheckpointer: Building Program Error Resilience-Aware Checkpointing Mechanism for HPC Systems
Xiaohui Wei 0002, Shiyu Tong, Zhongao Sun, Xiang Li 0197, Hengshan Yue |
J. Comput. Sci. Technol. | 4 |
| 2024 | Interference-aware opportunistic job placement for shared distributed deep learning clusters
Hongliang Li 0003, Hairui Zhao 0002, Xiang Li 0197, Haixiao Xu |
J. Parallel Distributed Comput. | 4 |
| 2023 | ExplSched: Maximizing Deep Learning Cluster Efficiency for Exploratory JobsabstractResource management for Deep Learning (DL) clusters is essential for system efficiency and model training quality. Existing schedulers provided by DL frameworks are mostly adaptations from traditional HPC clusters and usually work on jobs’ makespan, assuming that DL training jobs finish completely. Unfortunately, it is reported that a fair amount of training jobs are exploratory jobs and often finish unsuccessfully (over 30%) in production clusters. This is due to the distinct characteristic of Deep Neural Network (DNN) training that it is an exploratory process of frequent user interventions, such as adjusting model structures, tuning hyperparameters, and exploring feature validity. Existing DL cluster schedulers using offline algorithms are not suitable for exploratory jobs when unexpected early terminations can cause noticeable resource waste. Moreover, DL training jobs are iterative and usually yield diminishing returns as they progress. Equally allocating resource among training iterations is not efficient, especially when dealing with exploratory jobs where it can worsen the degradation of system efficiency. The fundamental goal of a DL training job is to gain model quality improvement, usually indicated by the loss reduction (job profit) of a DNN model. This paper introduces a novel scheduling problem for exploratory jobs that seeks to maximize the overall training profit of a DL cluster. We propose ExplSched, an online scheduling solution based on the primal-dual framework, resulting in a competitive ratio of 2α that belongs to O(ln n). It uses a resource price function that emphasizes the importance of job profit to resource consumption ratio to make quick resource allocation decisions. Experimental results show that ExplSched achieved an average system utility improvement of 87.28% compared with other related work. Hongliang Li 0003, Hairui Zhao 0002, Zhewen Xu, Xiang Li 0197, Haixiao Xu |
CLUSTER | 4 |
| 2022 | MSSA-FL: High-Performance Multi-stage Semi-asynchronous Federated Learning with Non-IID Data
Xiaohui Wei 0002, Mingkai Hou, Chenghao Ren, Xiang Li 0197, Hengshan Yue |
KSEM (2) | 4 |
| 2022 | Cooperative task assignment in spatial crowdsourcing via multi-agent deep reinforcement learning
Xiang Li 0197, Shang Gao 0005, Xiaohui Wei 0002 |
J. Syst. Archit. | 2 |
| 2019 | Pec: Proactive Elastic Collaborative Resource Scheduling in Data Stream ProcessingabstractIn the Distributed Parallel Stream Processing Systems (DPSPS), elastic resource allocation allows applications to dynamically response to workload fluctuations. However, resource provisioning can be particularly challenging, due to the unpredictability of the workload. In addition, unlike CPU resources, bandwidth resources are often ignored in resource allocation. Moreover, resource allocation and resource placement are considered separately. In this paper, we investigate the proactive elastic resource scheduling problem for computation-intensive and communication-intensive applications, which aims at meeting the latency requirement with the minimal energy cost, and propose a dynamic collaborative strategy from the systemic perspective. Specifically, we first model a collaborative workload prediction pattern to accurately predict the upcoming workload, and construct a latency estimation model to estimate the latency of the application. Then, we design an energy-efficient resource pre-allocation method, in which the CPU frequency adjustment and the stability of resource reconfigurations are both considered. Finally, we present a communication-aware resource placement approach. Simulation results show that, compared with the reactive strategies, our strategy achieves an obviously better latency performance, and effectively avoids unnecessary resource adjustments. Meanwhile, the energy consumption is about saved by 50 percent on average, and the communication cost is maintained at a very low level of 4 percent. Xiaohui Wei 0002, Xiang Li 0197, Xingwang Wang 0003, Shang Gao 0005, Hongliang Li 0003 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | A Task Allocation Method for Stream Processing with Recovery Latency Constraint
Hongliang Li 0003, Jie Wu 0001, Xiang Li 0197, Xiaohui Wei 0002 |
J. Comput. Sci. Technol. | 4 |
| 2017 | Task Allocation for Stream Processing with Recovery Latency GuaranteeabstractStream processing applications continuously process large amounts of online streaming data in real-time or near real-time. They have strict latency constraints, but they are also vulnerable to failures. Failure recoveries may slow down the entire processing pipeline and break latency constraints. Upstream backup is one of the most widely applied fault-tolerant schemes for stream processing systems. It introduces complex backup dependencies to tasks, and increases the difficulty of controlling recovery latencies. Moreover, when dependent tasks are located on the same processor, they fail at the same time in processor-level failures, bringing extra recovery latencies that increase the impacts of failures. This paper presents a correlated failure effect model to describe the recovery latency of a stream topology in processor-level failures for an allocation plan. We introduce a Recovery-latency-aware Task Allocation Problem (RTAP) that seeks task allocation plans for stream topologies that will achieve guaranteed recovery latencies. We present a heuristic algorithm with a computational complexity of O(nlog^2n) to solve the problem. Extensive experiments were conducted to verify the correctness and effectiveness of our approach. Hongliang Li 0003, Jie Wu 0001, Xiang Li 0197, Xiaohui Wei 0002 |
CLUSTER | 4 |
| 2017 | Integrated recovery and task allocation for stream processingabstractStream processing applications continueously process large-scale data streams online. The throughput of a stream processing application must match the input rate to avoid loss of data. Failures affect throughput because a task failure can suspend itself from producing new data and can even cause an application-level halt. The key motivation of this work is to mitigate the performance degradation caused by task-level failures. We introduce a novel Integrated Recovery Model (IRM) that allows resource sharing among both failure-free tasks and recovering tasks on a processor. The failure-free tasks slow down to accelerate a task recovery rather than suspending their actions and waiting for the recovery to finish; waiting causes a complete halt of the application. In this way, the recovery is seamless and does not suspend the entire system. The performance slowdown is related to both the failure-free processing cost and recovery cost on each processor. Moreover, the recovery cost of a task is related to the Fault-Tolerant Configuration (FTC) of the stream application. This paper introduces a novel task allocation problem that, given an FTC, can constrain processing performance during recoveries (i.e. throughput slowdown ratio) while minimizing the amount of resource occupied. We propose both a greedy algorithm and a heuristic algorithm with computational complexities of O(n log n) and O(n log2n), respectively, to solve the problem. Extensive experiments verify the correctness and effectiveness of our approach. Our approach enables continuous processing results and seamless failure recoveries with a constrained slowdown ratio. Hongliang Li 0003, Jie Wu 0001, Xiang Li 0197, Xiaohui Wei 0002, Yuan Zhuang 0003 |
IPCCC | 4 |
| 2017 | Minimum Backups for Stream Processing With Recovery Latency GuaranteesabstractThe stream processing model continuously processes online data in an on-pass fashion that can be more vulnerable to failures than other big-data processing schemes. Existing fault-tolerant (FT) approaches have been presented to enhance the reliability of stream processing systems. However, the fundamental tradeoff between recovery latency and FT overhead is still unclear, so these scheme cannot provide recovery latency guarantees. This paper introduces the FT Configuration (FTC) problem and presents a solution for guaranteed recovery latency with minimum backups. A failure effect model is presented to describe the relationship between recovery latency and FTC (the amount and locations of backups). With this model, we design an algorithm to compute FTCs for different types of stream topologies according to recovery latency requirements. Extensive experiments are conducted to verify the correctness and effectiveness of our approach. We prove that our algorithm guarantees recovery latencies for all directed acyclic graph (DAG) stream topologies. For line(s) and tree topologies, our algorithm solves the FTC problem with a time complexity of O(N). For a general DAG topology, a heuristic function is used to generate FTCs. This causes fewer than 10% more backups on average compared to the optimal solution with a time complexity of O(N2). Hongliang Li 0003, Jie Wu 0001, Xiang Li 0197, Xiaohui Wei 0002 |
IEEE Trans. Reliab. | 4 |