Guanglin Xu

dblp:33/5589 · DBLP profile ↗
← Back
18ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 2 since 2021Systems, architecture and hardware · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 QiMeng-Tensify: Scaling Up Tensor Computation Optimization via Architecture-Aware LLM-Guided MCTS
Shouyang Dong, Jun Bi, Yuanbo Wen 0001, Xiyue Yu, Jianxing Xu, Guanglin Xu, Ling Li 0001, Xuehai Zhou, Tianshi Chen 0002, Qi Guo 0001
ISCA6
2026 FlashAttention-T: Towards Fully Tensorized Attention by Exploiting Tensor-Vector Parallelism
abstract
The attention mechanism is central to modern deep learning, particularly in large language models (LLMs), but suffers from quadratic computational complexity. To accelerate attention computation on GPUs, fused attention techniques (e.g., FlashAttention) consolidate the matrix multiplication (GEMM) and softmax computations into a single kernel. However, these operations remain computationally decoupled: the GEMM leverages high-performance tensor units (Tensor Cores), while the softmax executes on slower vector units (CUDA cores). This imbalance induces severe vector intervals—periods where tensor units sit idle awaiting vector unit completion—significantly underutilizing tensor units. Furthermore, ongoing hardware advancements delivering faster tensor units exacerbate this bottleneck.
Jianxing Xu, Yuanbo Wen 0001, Jun Bi, Ruibai Xu, Guanglin Xu, Rui Zhang 0040, Wei Li 0008, Ling Li 0001, Tianshi Chen 0002, Qi Guo 0001, Yunji Chen
PPoPP5
2026 AGON: Automated Design Framework for Customizing Processors From ISA Documents
abstract
Customized processors are essential for domain-specific applications such as the Internet of Things (IoT) and multi-media embedded systems, yet their design often requires extensive expert intervention. Traditional approaches, including hardware design using encapsulated abstractions (e.g., Chisel) and high-level synthesis (HLS) from languages like C or SystemC, reduce some manual efforts but remain either costly or suboptimal. Recent explorations into leveraging Large Language Models (LLMs) to generate RTL from natural language specifications have shown promise, but these methods still struggle with generating complex and high-performance processors mainly due to the complicated low-level details in the RTL code. In this work, we introduce AGON, a novel framework designed to facilitate the development of customized processor RTL from instruction set architecture (ISA) documents using LLMs. The framework comprises two layers: a functional description layer and a hardware implementation layer. At the functional layer, AGON employs a nano-operator (nOP)-based Intermediate Representation (IR) that abstracts basic instruction operations, thereby reducing the semantic gap between natural language and RTL code. This abstraction significantly shortens the descriptive code required for LLM generation, improving the generation accuracy in single-pass. At the hardware layer, AGON offers three abstraction levels (i.e. instruction, ISA, and processor) along with rule-based primitives to systematically lower the nOP-based IR into a fully optimized processor implementation. This decoupled design not only ensures correctness-by-construction but also enables automated, PPA-aware performance optimization. We evaluate AGON by designing high-performance out-of-order processors that correctly execute practical programs. Experimental results demonstrate that processors generated with AGON achieve an average speedup of 4.51× on specific tasks compared to expert-designed general-purpose CPUs while requiring minimal design effort.
Chongxiao Li, Pengwei Jin, Tianyun Ma, Husheng Han, Shuyao Cheng, Yifan Hao 0001, Yongwei Zhao 0001, Guanglin Xu, Zidong Du, Rui Zhang 0040, Xiaqing Li, Yuanbo Wen 0001, Xing Hu 0001, Qi Guo 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2025 Automated Superscalar Processor Design by Learning Data Dependencies
abstract
Automated processor design, which can significantly reduce human efforts and accelerate design cycles, has received considerable attention. While recent advancements have automatically designed single-cycle processors that execute one instruction per cycle, their performance cannot compete with modern superscalar processors that execute multiple instructions per cycle. Previous methods fail on superscalar processor design because they cannot address inter-instruction data dependencies, leading to inefficient sequential instruction execution. This paper proposes a novel approach to automatically designing superscalar processors using a hardware-friendly model called the Stateful Binary Speculation Diagram (State-BSD). We observe that processor parallelism can be enhanced through on-the-fly inter-instruction dependent data predictors, reusing the processor's internal states to learn the data dependency. To meet the challenge of both hardware-resource limitation and design functional correctness, State-BSD consists of two components: 1) a lightweight state-selector trained by simulated annealing method to detect the most reusable processor states and store them in a small buffer; and 2) a highly precise state-speculator trained by BSD expansion method to predict the inter-instruction dependent data using the selected states. It is the first work to achieve the automated superscalar processor design, i.e. QiMeng-CPU-v2, which improves the performance by about 380x than the state-of-the-art automated design and is comparable to human-designed superscalar processors such as ARM Cortex A53.
Shuyao Cheng, Rui Zhang 0040, Wenkai He, Pengwei Jin, Chongxiao Li, Zidong Du, Xing Hu 0001, Yifan Hao 0001, Guanglin Xu, Yuanbo Wen 0001, Ling Li 0001, Qi Guo 0001, Yunji Chen
IJCAI9
2024 Revisiting Automatic Pipelining: Gate-level Forwarding and Speculation
abstract
Pipelining is a widely applied micro-architectural performance optimization and requires non-trivial designs for better execution throughput. The key to pipeline throughput optimization is to resolve data hazards caused by read-after-write (RAW) dependencies, which are traditionally tackled by forwarding and speculation to avoid pipeline stalls. However, existing approaches are conducted based on high-level dataflow analysis, with potential loss of optimization opportunities for lack of analysis of the netlist structures.
Shuyao Cheng, Chongxiao Li, Zidong Du, Rui Zhang 0040, Xing Hu 0001, Xiaqing Li, Guanglin Xu, Yuanbo Wen 0001, Qi Guo 0001
DAC7
2024 A 1.19GHz 9.52Gsamples/sec Radix-8 FFT Hardware Accelerator in 28nm
abstract
•Dedicated FFT hardware is typically designed as a standalone block and then integrated into systems •Relatively fixed functionality once implemented •System-level integration with FFT-based application software becomes manual and challenging to optimize
Larry Tang 0003, Keshav Harisrikanth, Guanglin Xu, Franz Franchetti, Ken Mai
HCS4
2022 Diversity in news recommendations using contextual bandits
Alexander Semenov, Maciej Rysz, Gaurav Pandey 0003, Guanglin Xu
Expert Syst. Appl.4
2020 Robust rigid registration algorithm based on pointwise correspondence and correntropy
Shaoyi Du, Guanglin Xu, Sirui Zhang, Xuetao Zhang 0001, Yue Gao 0002, Badong Chen
Pattern Recognit. Lett.2
2019 Precise iterative closest point algorithm with corner point constraint for isotropic scaling registration
Shaoyi Du, Wenting Cui, Liyang Wu, Sirui Zhang, Xuetao Zhang 0001, Guanglin Xu, Meifeng Xu
Multim. Syst.6
2019 RGB-D point cloud registration via infrared and color camera
Teng Wan, Shaoyi Du, Yiting Xu, Guanglin Xu, Badong Chen, Yue Gao 0002
Multim. Tools Appl.4
2018 Precise Point Set Registration Using Point-to-Plane Distance and Correntropy for LiDAR Based Localization
abstract
In this paper, we propose a robust point set registration algorithm which combines correntropy and point-to-plane distance, which can register rigid point sets with noises and outliers. Firstly, as correntropy performs well in handling data with non-Gaussian noises, we introduce it to model rigid point set registration problem based on point-to-plane distance; Secondly, we propose an iterative algorithm to solve this problem, which repeats to compute correspondence and transformation parameters respectively in closed form solutions. Simulated experimental results demonstrate the high precision and robustness of the proposed algorithm. In addition, LiDAR based localization experiments on automated vehicle performs satisfactory for localization accuracy and time consumption.
Guanglin Xu, Shaoyi Du, Dixiao Cui, Sirui Zhang, Badong Chen, Xuetao Zhang 0001, Jianru Xue, Yue Gao 0002
Intelligent Vehicles Symposium1
2018 Precise Point Set Registration with Color Assisted and Correntropy for 3D Reconstruction
abstract
Iterative closest point (ICP) algorithm, as its accuracy and efficiency, is widely used in rigid registration. However, ICP algorithm is easily failed when point sets lack of structure variety, such as semicircles. To solve this problem, a precise point set registration method for RGB-D data is proposed. Firstly, the color information provides a new information for registration, and the correntropy is introduced to deal with the noises and outliers. With color assisted and correntropy, a more robust objective function is built. Secondly, a variant ICP algorithm is used to deal with optimization problem via multiple iterations. Finally, as shown in the experimental results and scene reconstruction, our method obtains more precise results than other ICP algorithms.
Teng Wan, Shaoyi Du, Yiting Xu, Guanglin Xu, Yang Yang 0066, Yue Gao 0002, Badong Chen
SMC4
2018 Robust Non-rigid Registration Based on Affine ICP Algorithm and Part-Based Method
Liyang Wu, Wenting Cui, Sirui Zhang, Guanglin Xu, Huaizhong Hu
Neural Process. Lett.5
2017 Study on Updating Algorithm of Attribute Coordinate Evaluation Model
Guanglin Xu, Jiali Feng
ICIC (3)2
2017 Robust non-rigid point set registration via building tree dynamically
Shaoyi Du, Bo Bi, Guanglin Xu, Jihua Zhu, Xuetao Zhang 0001
Multim. Tools Appl.3
2016 Precise 2D point set registration using iterative closest algorithm and correntropy
abstract
The iterative closest point (ICP) algorithm is fast and accurate for rigid point set registration, but it works badly when there are many outliers and noises in the point sets. This paper instead proposes a novel method based on the ICP algorithm to deal with this problem. Firstly, correntropy is introduced into the rigid registration problem and then a new energy function based on maximum correntropy criterion is proposed. After that, a new ICP algorithm based on correntropy is proposed, which performs well in dealing with rigid registration with noises and outliers. This new algorithm converges moronically from any given parameters, which is similar to the ICP algorithm. Experimental results demonstrate its accuracy and efficiency compared with the traditional ICP algorithm.
Guanglin Xu, Shaoyi Du, Jianru Xue
IJCNN1
2013 Restricted Bayesian classification networks
Shuangcheng Wang, Guanglin Xu, Ruijie Du
Sci. China Inf. Sci.2
2005 Qualitative Mapping, Inner Product Transformation of Qualitative Criterion, Artificial Neuron and Pattern Recognition
abstract
The qualitative mapping (QM) model for judging a property p(o) whose true value varies according to the qualitative criterion [/spl alpha/,/spl beta/], /spl tau//sub p/(x, [/spl alpha/,/spl beta/]) is presented in this paper. The inner product transformation of qualitative criterion w/spl I.bar/[/spl alpha/,/spl beta/] and the relation between w/spl I.bar/[/spl alpha/,/spl beta/] and artificial neuron is discussed. By the the intercept form of artificial neuron, we prove that an AN is just a boundary of a qualitative mapping, and if a neighborhood is closed by a group of artificial neurons induced by a group inner product transformations, then the qualitative mapping whose criterion is the closed neighborhood is equivalent to the artificial networks which consists of the group of them.
Jiali Feng, Qihuang Mao, Guanglin Xu, Jingjuan Feng
ISM3