Qinlong Wang

dblp:90/9593 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Distributed systems · 50% High-performance computing · 33% Hardware reliability and fault tolerance · 9%
Artificial intelligence
4 papers
Efficient and distributed learning · 69% Language models and text generation · 21% Trustworthy machine learning · 10%
Computer networks
2 papers
Wireless networking · 45% Cellular and mobile networks · 34% Datacenter networks · 21%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing
collective communication
1.922026
Handling Network Faults in Distributed AI Training: Failover is Now an Option · EuroSys 2026
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training · SOSP 2025
Distributed systems
fault tolerance
1.722025
Robust LLM Training Infrastructure at ByteDance · SOSP 2025
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training · SOSP 2025
Machine learning › Efficient and distributed learning
distributed training
1.022025
DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloud · Proc. VLDB Endow. 2024
Robust LLM Training Infrastructure at ByteDance · SOSP 2025
Distributed systems › distributed machine learning
distributed training
1.012026
Handling Network Faults in Distributed AI Training: Failover is Now an Option · EuroSys 2026
Distributed systems › fault tolerance › high availability
failover
1.012026
Handling Network Faults in Distributed AI Training: Failover is Now an Option · EuroSys 2026
Hardware reliability and fault tolerance
network fault tolerance
1.012026
Handling Network Faults in Distributed AI Training: Failover is Now an Option · EuroSys 2026
Distributed systems › fault tolerance
failure recovery
0.912025
Robust LLM Training Infrastructure at ByteDance · SOSP 2025
High-performance computing › large-scale training
large language model training
0.912025
Robust LLM Training Infrastructure at ByteDance · SOSP 2025
Distributed systems
root cause analysis
0.912025
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training · SOSP 2025
Machine learning › Efficient and distributed learning › distributed training
recommendation model training
0.812024
DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloud · Proc. VLDB Endow. 2024
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.812024
DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloud · Proc. VLDB Endow. 2024
Natural language and speech › Language models and text generation › text generation
grammatical error correction
0.312017
A Nested Attention Neural Hybrid Model for Grammatical Error Correction · ACL (1) 2017
Machine learning › Trustworthy machine learning › robustness
fault tolerance
0.312025
Robust LLM Training Infrastructure at ByteDance · SOSP 2025
Natural language and speech › Language models and text generation
large language model training
0.312025
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training · SOSP 2025
Wireless networking › cognitive radio › spectrum sharing
licensed-assisted access
0.212016
Design and Performance Analysis of a Fairness-Based License-Assisted Access and Resource Scheduling Scheme · IEEE J. Sel. Areas Commun. 2016
Cellular and mobile networks
radio resource management
0.212016
Design and Performance Analysis of a Fairness-Based License-Assisted Access and Resource Scheduling Scheme · IEEE J. Sel. Areas Commun. 2016
Cellular and mobile networks
resource scheduling
0.212016
Design and Performance Analysis of a Fairness-Based License-Assisted Access and Resource Scheduling Scheme · IEEE J. Sel. Areas Commun. 2016
Wireless networking › cognitive radio
spectrum sharing
0.212016
Design and Performance Analysis of a Fairness-Based License-Assisted Access and Resource Scheduling Scheme · IEEE J. Sel. Areas Commun. 2016
Wireless networking › cognitive radio › spectrum sharing › coexistence
wifi coexistence
0.112016
Design and Performance Analysis of a Fairness-Based License-Assisted Access and Resource Scheduling Scheme · IEEE J. Sel. Areas Commun. 2016
Wireless networking
WLAN
0.112016
Design and Performance Analysis of a Fairness-Based License-Assisted Access and Resource Scheduling Scheme · IEEE J. Sel. Areas Commun. 2016

Methods — techniques the papers use, named apart from their topics

intra-host GPU routing · 2.0dynamic channel load balancing · 2.0tracing · 1.7fault diagnosis · 1.7failure tolerance · 1.7data-driven localization · 1.7resource-performance models · 1.5heuristic resource allocation · 1.5sequence-to-sequence · 0.3nested attention · 0.3hybrid neural model · 0.3linear programming · 0.2contention window optimization · 0.2
YearPublicationVenuePosition
2026 Handling Network Faults in Distributed AI Training: Failover is Now an Option
abstract
Distributed AI training often suffers from network faults. Network faults, especially at the last hop between a switch and a host, result in loss of connectivity, resulting in training job stalls and eventual failure. This is typically managed through a fail-stop mechanism, followed by a restart, incurring significant inefficiencies. We present ReCCL, the first network fault-tolerant collective communication library (CCL) that allows training progress to be preserved by seamlessly failing over to alternate paths when a network fault occurs. During failover, ReCCL keeps communication states synchronized while using dynamic channel load balancing and intra-host GPU routing to improve communication performance. Our evaluations demonstrate that ReCCL can perform failover seamlessly with minimal performance losses. Additionally, our simulations also demonstrate that failover can be effectively used to achieve significant savings in GPU hours for large-scale distributed AI training workloads.
Xin Zhe Khooi, Zhuo Jiang, Pan Xie, Zhigang Cui, Meng Wang 0018, Yuze Jin, Pengfei Huo, Lulu Chen, Liaoyuan Feng, Qinlong Wang, Yongcan Wang, Jinshuai Sun, Yingkai Zhao, Haiquan Chen 0002, Yi Li 0098, Jianxi Ye, Mun Choon Chan
EuroSys14
2025 Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training
abstract
Reliability is essential for ensuring efficiency in LLM training. However, many real-world reliability issues remain difficult to resolve, resulting in wasted resources and degraded model performance. Unfortunately, today's collective communication libraries operate as black boxes, hiding critical information needed for effective root cause analysis.
Yangtao Deng, Qinlong Wang, Xiaoyun Zhi, Zhuo Jiang, Haohan Xu, Zuquan Song, Gaohong Liu, Shuguang Wang, Wencong Xiao, Jianxi Ye, Minlan Yu, Hong Xu 0001
SOSP3
2025 Robust LLM Training Infrastructure at ByteDance
abstract
The training scale of large language models (LLMs) has reached tens of thousands of GPUs and is still continuously expanding, enabling faster learning of larger models. Accompanying the expansion of the resource scale is the prevalence of failures (CUDA error, NaN values, job hang, etc.), which poses significant challenges to training stability. Any large-scale LLM training infrastructure should strive for minimal training interruption, efficient fault diagnosis, and effective failure tolerance to enable highly efficient continuous training. This paper presents ByteRobust, a large-scale GPU infrastructure management system tailored for robust and stable training of LLMs. It exploits the uniqueness of LLM training process and gives top priorities to detecting and recovering failures in a routine manner. Leveraging parallelisms and characteristics of LLM training, ByteRobust enables high-capacity fault tolerance, prompt fault demarcation, and localization with an effective data-driven approach, comprehensively ensuring continuous and efficient training of LLM tasks. ByteRobust is deployed on a production GPU platform with over 200,000 GPUs and advances the state of the art in training robustness by achieving 97% ETTR for a three-month training job on 9,600 GPUs.
Borui Wan, Gaohong Liu, Zuquan Song, Jun Wang 0039, Guangming Sheng, Shuguang Wang, Houmin Wei, Weiqiang Lou, Mofan Zhang, Kaihua Jiang, Cheng Ren, Xiaoyun Zhi, Menghan Yu, Zhe Nan, Zhuolin Zheng, Baoquan Zhong, Qinlong Wang, Jinxin Chi, Wang Zhang 0017, Zixian Du, Sida Zhao, Jingzhe Tang, Zherui Liu, Chuan Wu 0001, Yanghua Peng, Haibin Lin, Wencong Xiao, Xin Liu 0086
SOSP20
2024 DLRover-RM: Resource Optimization for Deep Recommendation Models Training in the cloud
abstract
Deep learning recommendation models (DLRM) rely on large embedding tables to manage categorical sparse features. Expanding such embedding tables can significantly enhance model performance, but at the cost of increased GPU/CPU/memory usage. Meanwhile, tech companies have built extensive cloud-based services to accelerate training DLRM models at scale. In this paper, we conduct a deep investigation of the DLRM training platforms at AntGroup and reveal two critical challenges: low resource utilization due to suboptimal configurations by users and the tendency to encounter abnormalities due to an unstable cloud environment. To overcome them, we introduce DLRover, an elastic training framework for DLRMs designed to increase resource utilization and handle the instability of a cloud environment. DLRover develops a resource-performance model by considering the unique characteristics of DLRMs and a three-stage heuristic strategy to automatically allocate and dynamically adjust resources for DLRM training jobs for higher resource utilization. Further, DLRover develops multiple mechanisms to ensure efficient and reliable execution of DLRM training jobs. Our extensive evaluation shows that DLRover reduces job completion times by 31%, increases the job completion rate by 6%, enhances CPU usage by 15%, and improves memory utilization by 20%, compared to state-of-the-art resource scheduling frameworks. DLRover has been widely deployed at AntGroup and processes thousands of DLRM training jobs on a daily basis. DLRover is open-sourced and has been adopted by 10+ companies.
Qinlong Wang, Tingfeng Lan, Yinghao Tang, Bo Sang, Ziling Huang, Yiheng Du, Jian Sha, Hui Lu 0001, Yuanchun Zhou, Ke Zhang 0048, MingJie Tang
Proc. VLDB Endow.1
2020 Robust Point Set Registration Based on Semantic Information
abstract
Point cloud registration a challenging task in situations with poor initial value and scenarios with limited geometric structure. In these cases, the correct correspondence between two point clouds is unknown and difficult to establish. To cope with this problem, the semantic of partial points is introduced in this paper. Firstly, the semantic information is used to find more reasonable correspondence, i.e. semantic point pairs. Secondly, we formulate a novel objective function to integrate the matching error of semantic point pairs as guidance of registration. Thirdly, a hyperparameter is applied to balance the confidence of semantic point pairs. At last, a novel algorithm under the ICP framework is presented to optimize the rigid transformation iteratively. The evaluation of KITTI data set reveals the robustness and accuracy of our method in the complex scenes mentioned above.
Qinlong Wang, Yang Yang 0066, Teng Wan, Shaoyi Du
SMC1
2017 A Nested Attention Neural Hybrid Model for Grammatical Error Correction
abstract
Jianshu Ji, Qinlong Wang, Kristina Toutanova, Yongen Gong, Steven Truong, Jianfeng Gao. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017.
Jianshu Ji, Qinlong Wang, Kristina Toutanova, Yongen Gong, Steven Truong, Jianfeng Gao 0001
ACL (1)2
2017 QoS-guaranteed data rate allocation for mixed services on licensed and unlicensed bands in LTE and WiFi systems
abstract
In face of the exponentially growing demands for various mobile data services, the long term evolution-unlicensed (LTE-U) technology has been proposed to offload over-loaded traffics on the unlicensed spectrum band by leveraging carrier aggregation (CA) technology. However, the quality of service (QoS) requirements of different or mixed services have not been fully considered in existing traffic allocation strategies in terms of various spectrum qualities on both licensed and unlicensed spectrum bands. Therefore, a novel two-step traffic allocation approach is proposed for mixed services to maximize the overall utility of LTE and WiFi systems. And two data rate allocation algorithms for licensed band (DRA-L) and unlicensed band (DRA-UL) are designed based on theoretical analysis. Simulation results verify the enhanced network utility performance of the proposed algorithms compared to the conventional stand-alone DRA-L algorithm.
Qixun Zhang, Lei Ji 0003, Qinlong Wang, Zhiyong Feng 0001
PIMRC4
2016 Design and Performance Analysis of a Fairness-Based License-Assisted Access and Resource Scheduling Scheme
abstract
In face of the explosive surge of mobile data services, spectrum aggregation or carrier aggregation technology has been proposed to improve system throughput and spectrum efficiency (SE) by aggregating licensed and unlicensed spectrum bands. However, the system performances would be severely deteriorated by the channel access collision if the channel access and resource scheduling approaches are not coordinated among different networks in the same spectrum band. Therefore, in order to improve the system throughput and the SE, a fairness-based license-assisted access and resource scheduling scheme are designed for the coexisting systems, incorporating long term evolution-advanced and WiFi systems in the unlicensed band. The optimal sizes of the contention window in the proposed fairness-based channel access approach are obtained in terms of various density ratios between these two systems. Furthermore, a novel resource scheduling approach employing linear programming is proposed to maximize the utility function with the goal of improving the service experience of users and the SE with various spectrum qualities. The theoretical proofs and simulation results verify the enhanced performances of the proposed approaches in terms of key metrics, such as throughput, SE, delay, and packet loss ratio.
Qixun Zhang, Qinlong Wang, Zhiyong Feng 0001
IEEE J. Sel. Areas Commun.2