Borui Xu

dblp:301/1626 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021
YearPublicationVenuePosition
2025 ScalaGBM: Memory Efficient GBDT Training for High-Dimensional Data on GPU
abstract
Gradient Boosted Decision Trees (GBDTs) are classical machine learning algorithms widely employed in recommendation systems, database queries, etc. Due to the extensive memory access involved in histogram-based GBDT training methods, high-bandwidth GPUs have been widely adopted to accelerate the training. However, when handling millions of feature data, it requires significant memory to store the training data and histograms, posing challenges for training on limited GPU memories. In this paper, we develop a GPU-based GBDT framework named ScalaGBM, aiming to accelerate high-dimensional data training with less memory usage. We first employ a CSR-like data format and CSR-based histogram construction to reduce the memory occupation of the training data. Then, we reorganize the training workflow with a double buffer structure to reduce the overall memory consumption for the histogram. Finally, we develop multi-dimensional parallel histogram construction and global optimal split point reduction to speed up the training process. Experimental results demonstrate that ScalaGBM handles real-world datasets with over 100 million instances of 50 million features with a single commercial GPU while existing GBDT frameworks all run into out-of-memory errors. Meanwhile, ScalaGBM achieves a maximum speedup of 39× over state-of-the-art GBDT counterparts without sacrificing the training quality. The code is available at https://github.com/Xtra-Computing/thundergbm.
Borui Xu, Zeyi Wen, Yao Chen 0008, Weng-Fai Wong, Bingsheng He
KDD (1)1
2025 Evaluating Small Language Models for News Summarization: Implications and Factors Influencing Performance
abstract
Borui Xu, Yao Chen, Zeyi Wen, Weiguo Liu, Bingsheng He. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Borui Xu, Yao Chen 0008, Zeyi Wen, Bingsheng He
NAACL (Long Papers)1
2023 Leveraging Data Density and Sparsity for Efficient SVM Training on GPUs
abstract
Support Vector Machines (SVMs) are a widely adopted data mining algorithm for binary and multi-class classification due to their ability to handle high-dimensional and non-linearly separable problems. However, SVM training is computationally expensive because of the heavy kernel matrix computation on large training datasets. Although much effort has been made to accelerate the training of SVMs, we find that existing libraries still suffer from inappropriate matrix multiplication methods and inefficient memory access patterns. In this paper, we propose a series of optimization approaches to address these limitations, including (i) matrix partitioning based on column density to achieve efficient kernel matrix computation; (ii) optimizing high latency memory access patterns; and (iii) dynamically selecting more suitable matrix multiplication methods based on the training dataset characteristics. Our proposed methods demonstrate significant improvements in SVM training performance without sacrificing accuracy, achieving a maximum speedup of 52x over the state-of-the-art SVMs on GPUs. These results highlight the effectiveness of our optimization in improving SVM training efficiency.
Borui Xu, Zeyi Wen, Lifeng Yan, Zhan Zhao, Zekun Yin, Bingsheng He
ICDM1
2022 A Multi-scale Convolution and Gated Recurrent Unit Based Network for Limit Order Book Prediction
Borui Xu
KSEM (1)1