EDBT 2026 Demo / reviewers in the wild / expert
Qinlong Wang
dblp:90/9593
· DBLP profile ↗
8ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Distributed systems · 50% High-performance computing · 33% Hardware reliability and fault tolerance · 9% | |
| Artificial intelligence
4 papers |
Efficient and distributed learning · 69% Language models and text generation · 21% Trustworthy machine learning · 10% | |
| Computer networks
2 papers |
Wireless networking · 45% Cellular and mobile networks · 34% Datacenter networks · 21% |
Topics — the 20 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing
collective communication |
1.9 | 2 | 2026 | Handling Network Faults in Distributed AI Training: Failover is Now an Option · EuroSys 2026 Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training · SOSP 2025 |
Distributed systems
fault tolerance |
1.7 | 2 | 2025 | Robust LLM Training Infrastructure at ByteDance · SOSP 2025 Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training · SOSP 2025 |
Machine learning › Efficient and distributed learning
distributed training |
1.0 | 2 | 2025 | DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloud · Proc. VLDB Endow. 2024 Robust LLM Training Infrastructure at ByteDance · SOSP 2025 |
Distributed systems › distributed machine learning
distributed training |
1.0 | 1 | 2026 | Handling Network Faults in Distributed AI Training: Failover is Now an Option · EuroSys 2026 |
Distributed systems › fault tolerance › high availability
failover |
1.0 | 1 | 2026 | Handling Network Faults in Distributed AI Training: Failover is Now an Option · EuroSys 2026 |
Hardware reliability and fault tolerance
network fault tolerance |
1.0 | 1 | 2026 | Handling Network Faults in Distributed AI Training: Failover is Now an Option · EuroSys 2026 |
Distributed systems › fault tolerance
failure recovery |
0.9 | 1 | 2025 | Robust LLM Training Infrastructure at ByteDance · SOSP 2025 |
High-performance computing › large-scale training
large language model training |
0.9 | 1 | 2025 | Robust LLM Training Infrastructure at ByteDance · SOSP 2025 |
Distributed systems
root cause analysis |
0.9 | 1 | 2025 | Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training · SOSP 2025 |
Machine learning › Efficient and distributed learning › distributed training
recommendation model training |
0.8 | 1 | 2024 | DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloud · Proc. VLDB Endow. 2024 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.8 | 1 | 2024 | DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloud · Proc. VLDB Endow. 2024 |
Natural language and speech › Language models and text generation › text generation
grammatical error correction |
0.3 | 1 | 2017 | A Nested Attention Neural Hybrid Model for Grammatical Error Correction · ACL (1) 2017 |
Machine learning › Trustworthy machine learning › robustness
fault tolerance |
0.3 | 1 | 2025 | Robust LLM Training Infrastructure at ByteDance · SOSP 2025 |
Natural language and speech › Language models and text generation
large language model training |
0.3 | 1 | 2025 | Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training · SOSP 2025 |
Wireless networking › cognitive radio › spectrum sharing
licensed-assisted access |
0.2 | 1 | 2016 | Design and Performance Analysis of a Fairness-Based License-Assisted Access and Resource Scheduling Scheme · IEEE J. Sel. Areas Commun. 2016 |
Cellular and mobile networks
radio resource management |
0.2 | 1 | 2016 | Design and Performance Analysis of a Fairness-Based License-Assisted Access and Resource Scheduling Scheme · IEEE J. Sel. Areas Commun. 2016 |
Cellular and mobile networks
resource scheduling |
0.2 | 1 | 2016 | Design and Performance Analysis of a Fairness-Based License-Assisted Access and Resource Scheduling Scheme · IEEE J. Sel. Areas Commun. 2016 |
Wireless networking › cognitive radio
spectrum sharing |
0.2 | 1 | 2016 | Design and Performance Analysis of a Fairness-Based License-Assisted Access and Resource Scheduling Scheme · IEEE J. Sel. Areas Commun. 2016 |
Wireless networking › cognitive radio › spectrum sharing › coexistence
wifi coexistence |
0.1 | 1 | 2016 | Design and Performance Analysis of a Fairness-Based License-Assisted Access and Resource Scheduling Scheme · IEEE J. Sel. Areas Commun. 2016 |
Wireless networking
WLAN |
0.1 | 1 | 2016 | Design and Performance Analysis of a Fairness-Based License-Assisted Access and Resource Scheduling Scheme · IEEE J. Sel. Areas Commun. 2016 |
Methods — techniques the papers use, named apart from their topics
intra-host GPU routing · 2.0dynamic channel load balancing · 2.0tracing · 1.7fault diagnosis · 1.7failure tolerance · 1.7data-driven localization · 1.7resource-performance models · 1.5heuristic resource allocation · 1.5sequence-to-sequence · 0.3nested attention · 0.3hybrid neural model · 0.3linear programming · 0.2contention window optimization · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Handling Network Faults in Distributed AI Training: Failover is Now an OptionabstractDistributed AI training often suffers from network faults. Network faults, especially at the last hop between a switch and a host, result in loss of connectivity, resulting in training job stalls and eventual failure. This is typically managed through a fail-stop mechanism, followed by a restart, incurring significant inefficiencies. We present ReCCL, the first network fault-tolerant collective communication library (CCL) that allows training progress to be preserved by seamlessly failing over to alternate paths when a network fault occurs. During failover, ReCCL keeps communication states synchronized while using dynamic channel load balancing and intra-host GPU routing to improve communication performance. Our evaluations demonstrate that ReCCL can perform failover seamlessly with minimal performance losses. Additionally, our simulations also demonstrate that failover can be effectively used to achieve significant savings in GPU hours for large-scale distributed AI training workloads. Xin Zhe Khooi, Zhuo Jiang, Pan Xie, Zhigang Cui, Meng Wang 0018, Yuze Jin, Pengfei Huo, Lulu Chen, Liaoyuan Feng, Qinlong Wang, Yongcan Wang, Jinshuai Sun, Yingkai Zhao, Haiquan Chen 0002, Yi Li 0098, Jianxi Ye, Mun Choon Chan |
EuroSys | 14 |
| 2025 | Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM TrainingabstractReliability is essential for ensuring efficiency in LLM training. However, many real-world reliability issues remain difficult to resolve, resulting in wasted resources and degraded model performance. Unfortunately, today's collective communication libraries operate as black boxes, hiding critical information needed for effective root cause analysis. Yangtao Deng, Qinlong Wang, Xiaoyun Zhi, Zhuo Jiang, Haohan Xu, Zuquan Song, Gaohong Liu, Shuguang Wang, Wencong Xiao, Jianxi Ye, Minlan Yu, Hong Xu 0001 |
SOSP | 3 |
| 2025 | Robust LLM Training Infrastructure at ByteDanceabstractThe training scale of large language models (LLMs) has reached tens of thousands of GPUs and is still continuously expanding, enabling faster learning of larger models. Accompanying the expansion of the resource scale is the prevalence of failures (CUDA error, NaN values, job hang, etc.), which poses significant challenges to training stability. Any large-scale LLM training infrastructure should strive for minimal training interruption, efficient fault diagnosis, and effective failure tolerance to enable highly efficient continuous training. This paper presents ByteRobust, a large-scale GPU infrastructure management system tailored for robust and stable training of LLMs. It exploits the uniqueness of LLM training process and gives top priorities to detecting and recovering failures in a routine manner. Leveraging parallelisms and characteristics of LLM training, ByteRobust enables high-capacity fault tolerance, prompt fault demarcation, and localization with an effective data-driven approach, comprehensively ensuring continuous and efficient training of LLM tasks. ByteRobust is deployed on a production GPU platform with over 200,000 GPUs and advances the state of the art in training robustness by achieving 97% ETTR for a three-month training job on 9,600 GPUs. Borui Wan, Gaohong Liu, Zuquan Song, Jun Wang 0039, Guangming Sheng, Shuguang Wang, Houmin Wei, Weiqiang Lou, Mofan Zhang, Kaihua Jiang, Cheng Ren, Xiaoyun Zhi, Menghan Yu, Zhe Nan, Zhuolin Zheng, Baoquan Zhong, Qinlong Wang, Jinxin Chi, Wang Zhang 0017, Zixian Du, Sida Zhao, Jingzhe Tang, Zherui Liu, Chuan Wu 0001, Yanghua Peng, Haibin Lin, Wencong Xiao, Xin Liu 0086 |
SOSP | 20 |
| 2024 | DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloudabstractDeep learning recommendation models (DLRM) rely on large embedding tables to manage categorical sparse features. Expanding such embedding tables can significantly enhance model performance, but at the cost of increased GPU/CPU/memory usage. Meanwhile, tech companies have built extensive cloud-based services to accelerate training DLRM models at scale. In this paper, we conduct a deep investigation of the DLRM training platforms at AntGroup and reveal two critical challenges: low resource utilization due to suboptimal configurations by users and the tendency to encounter abnormalities due to an unstable cloud environment. To overcome them, we introduce DLRover, an elastic training framework for DLRMs designed to increase resource utilization and handle the instability of a cloud environment. DLRover develops a resource-performance model by considering the unique characteristics of DLRMs and a three-stage heuristic strategy to automatically allocate and dynamically adjust resources for DLRM training jobs for higher resource utilization. Further, DLRover develops multiple mechanisms to ensure efficient and reliable execution of DLRM training jobs. Our extensive evaluation shows that DLRover reduces job completion times by 31%, increases the job completion rate by 6%, enhances CPU usage by 15%, and improves memory utilization by 20%, compared to state-of-the-art resource scheduling frameworks. DLRover has been widely deployed at AntGroup and processes thousands of DLRM training jobs on a daily basis. DLRover is open-sourced and has been adopted by 10+ companies. Qinlong Wang, Tingfeng Lan, Yinghao Tang, Bo Sang, Ziling Huang, Yiheng Du, Jian Sha, Hui Lu 0001, Yuanchun Zhou, Ke Zhang 0048, MingJie Tang |
Proc. VLDB Endow. | 1 |
| 2020 | Robust Point Set Registration Based on Semantic InformationabstractPoint cloud registration a challenging task in situations with poor initial value and scenarios with limited geometric structure. In these cases, the correct correspondence between two point clouds is unknown and difficult to establish. To cope with this problem, the semantic of partial points is introduced in this paper. Firstly, the semantic information is used to find more reasonable correspondence, i.e. semantic point pairs. Secondly, we formulate a novel objective function to integrate the matching error of semantic point pairs as guidance of registration. Thirdly, a hyperparameter is applied to balance the confidence of semantic point pairs. At last, a novel algorithm under the ICP framework is presented to optimize the rigid transformation iteratively. The evaluation of KITTI data set reveals the robustness and accuracy of our method in the complex scenes mentioned above. Qinlong Wang, Yang Yang 0066, Teng Wan, Shaoyi Du |
SMC | 1 |
| 2017 | A Nested Attention Neural Hybrid Model for Grammatical Error CorrectionabstractJianshu Ji, Qinlong Wang, Kristina Toutanova, Yongen Gong, Steven Truong, Jianfeng Gao. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017. Jianshu Ji, Qinlong Wang, Kristina Toutanova, Yongen Gong, Steven Truong, Jianfeng Gao 0001 |
ACL (1) | 2 |
| 2017 | QoS-guaranteed data rate allocation for mixed services on licensed and unlicensed bands in LTE and WiFi systemsabstractIn face of the exponentially growing demands for various mobile data services, the long term evolution-unlicensed (LTE-U) technology has been proposed to offload over-loaded traffics on the unlicensed spectrum band by leveraging carrier aggregation (CA) technology. However, the quality of service (QoS) requirements of different or mixed services have not been fully considered in existing traffic allocation strategies in terms of various spectrum qualities on both licensed and unlicensed spectrum bands. Therefore, a novel two-step traffic allocation approach is proposed for mixed services to maximize the overall utility of LTE and WiFi systems. And two data rate allocation algorithms for licensed band (DRA-L) and unlicensed band (DRA-UL) are designed based on theoretical analysis. Simulation results verify the enhanced network utility performance of the proposed algorithms compared to the conventional stand-alone DRA-L algorithm. Qixun Zhang, Lei Ji 0003, Qinlong Wang, Zhiyong Feng 0001 |
PIMRC | 4 |
| 2016 | Design and Performance Analysis of a Fairness-Based License-Assisted Access and Resource Scheduling SchemeabstractIn face of the explosive surge of mobile data services, spectrum aggregation or carrier aggregation technology has been proposed to improve system throughput and spectrum efficiency (SE) by aggregating licensed and unlicensed spectrum bands. However, the system performances would be severely deteriorated by the channel access collision if the channel access and resource scheduling approaches are not coordinated among different networks in the same spectrum band. Therefore, in order to improve the system throughput and the SE, a fairness-based license-assisted access and resource scheduling scheme are designed for the coexisting systems, incorporating long term evolution-advanced and WiFi systems in the unlicensed band. The optimal sizes of the contention window in the proposed fairness-based channel access approach are obtained in terms of various density ratios between these two systems. Furthermore, a novel resource scheduling approach employing linear programming is proposed to maximize the utility function with the goal of improving the service experience of users and the SE with various spectrum qualities. The theoretical proofs and simulation results verify the enhanced performances of the proposed approaches in terms of key metrics, such as throughput, SE, delay, and packet loss ratio. Qixun Zhang, Qinlong Wang, Zhiyong Feng 0001 |
IEEE J. Sel. Areas Commun. | 2 |