EDBT 2026 Demo / reviewers in the wild / expert
Xiaosong Peng
dblp:383/3261
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0005-4919-3274ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QTTARNN: A Highly-Efficient Attention Driven Quantized Tensor Train Recursive Neural Network for Cyber-Physical-Social IntelligenceabstractABSTRACT Cyber‐Physical‐Social System (CPSS) which refers to the complex interaction of cyber, physical and social systems, has the important purpose to provide personalized intelligent services. CPSS data, generated from every aspect of people living life, are mainly in the form of time series multimodal data with characteristic of high order and high dimension. How to efficiently process these CPSS data is one of the fundamental ways for the intelligent services. In this paper, a highly‐efficient attention driven Quantized Tensor Train Recursive Neural Network is proposed, in which the CPSS data is decomposed into the form of tensor train cores. In this way, the proposed method is composed of lightweight high‐order neural network units, which better preserves the multi‐attribute features of the original data and the correlation between different dimensions by using tensors with its calculations, and implicitly trims the dense vector‐matrix connections in the fully connected network by using the form of quantization tensor train decomposition, which greatly reduces the model parameters, shortens the training time and improves the efficiency. Also, an effective attentional feature enhancement module is constructed to assist the high‐order neural network, so that the overall model can achieve a balance between low parameter number and accuracy. The network structure proposed in this paper realizes an efficient and lossless high‐order tensor recurrent neural network model with a small number of parameters. Finally, experiments on the UCF50 action video dataset, CWRU bearing dataset, and image generation tasks are conducted. Comparative analyses with vanilla LSTM and other tensorized LSTMs in terms of training time, accuracy, error, and compression ratio validate the reliability of the proposed model. Tinghua Zhang, Junxin Li, Xiaosong Peng, Zhixuan Zhao, Biyuan Yao |
Concurr. Comput. Pract. Exp. | 3 |
| 2026 | Dynamic Incremental Tucker Decomposition for Sparse Tensors on Heterogeneous PlatformsabstractTucker decomposition approximates the original tensor using a set of factor matrices and a core tensor, thereby reducing the storage and computational cost. Existing incremental Tucker decomposition techniques have limitations in efficiency and are not applicable to diverse dynamic incremental tensors. This paper proposes a dynamic incremental Tucker decomposition method based on heterogeneous computing, which efficiently handles complex dynamic tensors while preserving computational efficiency. Specifically, the target optimization function of Tucker decomposition is reconstructed to decouple the intrinsic correlation between factor matrix rows and fixed tensor dimensions, and the optimal solutions for the factor matrix rows are derived. This enables the algorithm to adapt to arbitrary forms of incoming incremental data, including time-series tensors, without recalculating existing data, thus enhancing data utilization efficiency. Then, a lightweight precision optimization strategy is introduced, which iteratively updates factor matrix rows using the most element-rich dimensions of each tensor order based on recorded observable data volumes. This strategy improves precision while significantly reducing computational costs. Furthermore, a task-characteristic-based heterogeneous computing scheduling strategy is proposed. By extracting highly parallelizable intermediate operators and leveraging the high throughput of GPUs, the algorithm is accelerated, leading to a 10 to 40 times speedup in incremental tensor decomposition compared to state-of-the-art methods, while maintaining accuracy. Xiaosong Peng, Laurence T. Yang, Xiaokang Wang 0001, Shijie Lv |
IEEE Trans. Computers | 1 |
| 2026 | Based on Tensor Core Sparse Kernels Accelerating Deep Neural NetworksabstractLarge language models in deep learning have numerous parameters, requiring significant storage space and computational resources. Compression techniques are highly effective in addressing these challenges. With the development of hardware like Graphics Processing Unit (GPU), Tensor Core can accelerate low-precision matrix multiplication but achieve acceleration for sparse matrices is challenging. Due to its sparsity, the utilization of Tensor Cores is relatively low. To address this, we propose the based onTensorCoreCompressedSparseRow format (TC-CSR), which facilitates data loading on GPUs and matrix operations on Tensor Cores. Based on this format, we designed block Sparse Matrix-Matrix Multiplication (SpMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM) kernels, which are common operations in deep learning. Utilizing these designs, we achieved a$\mathbf {1.41\times }$speedup on Sputnik in scenarios of moderate sparsity and a$\mathbf {1.38\times }$speedup with large-scale highly sparse matrices. Benefit from our design, we achieved a$\mathbf {1.75\times }$speedup in end-to-end inference with sparse Transformers and save memory. Shijie Lv, Debin Liu, Laurence T. Yang, Xiaosong Peng, Ruonan Zhao, Zecan Yang, Jun Feng 0007 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | A faster heterogeneous parallel computing method for Tucker decomposition
Xiaosong Peng, Laurence T. Yang |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | A High-Efficiency Parallel Mechanism for Canonical Polyadic Decomposition on Heterogeneous Computing PlatformabstractCanonical Polyadic decomposition (CPD) obtains the low-rank approximation for high-order multidimensional tensors through the summation of a sequence of rank-one tensors, greatly reducing storage and computation overhead. It is increasingly being used in the lightweight design of artificial intelligence and big data processing. The existing CPD technology exhibits inherent limitations in simultaneously achieving high accuracy and high efficiency. In this paper, a heterogeneous computing method for CPD is proposed to optimize computing efficiency with guaranteed convergence accuracy. Specifically, a quasi-convex decomposition loss function is constructed and the extreme points of the Kruskal matrix rows have been solved. Further, the massively parallelized operators in the algorithm are extracted, a software-hardware integrated scheduling method is designed, and the deployment of CPD on heterogeneous computing platforms is achieved. Finally, the memory access strategy is optimized to improve memory access efficiency. We tested the algorithm on real-world and synthetic sparse tensor datasets, numerical experimental results show that compared with the state-of-the-art method, the proposed method has a higher convergence accuracy and computing efficiency. Compared to the standard CPD parallel library, the method achieves efficiency improvements of tens to hundreds of times while maintaining the same accuracy. Xiaosong Peng, Laurence T. Yang, Xiaokang Wang 0001, Debin Liu, Jie Li 0111 |
IEEE Trans. Computers | 1 |