EDBT 2026 Demo / reviewers in the wild / expert
Xuhang Wang
dblp:361/4828
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic reliability analysis for complex multi-state systems of more-electric aircraft: A lightweight DBN method based on E-TrMF and interval grey number
Jiayu Chen 0002, Xuhang Wang, Qinhua Lu, Min Xie 0001, Hongjuan Ge |
Adv. Eng. Informatics | 2 |
| 2025 | MHDiff: Memory- and Hardware-Efficient Diffusion Acceleration via Focal Pixel Aware QuantizationabstractDiffusion models have demonstrated superior performance in image generation tasks, thus becoming the mainstream model for generative visual tasks. Diffusion models need to execute multiple timesteps sequentially, resulting in a dramatic increase in workload. Existing accelerators leverage the data similarity between adjacent timesteps and perform mixed-precision differential quantization to accelerate diffusion models. However, merging differential values with raw inputs in each layer of each timestep to ensure computational correctness requires significant memory access for loading raw inputs, which creates a heavy memory burden. Moreover, mixed-precision computations may lead to low hardware utilization if not well designed. Unlike these works, we propose MHDiff, a tailored framework that identifies the focal pixels at the first layer and finetunes them to fit all layers, then represents focal pixels with high-precision while using low-precision for others, thereby accelerating diffusion models while minimizing memory burden. To improve hardware utilization, MHDiff employs a packing module that merges low-precision values into high-precision values to create full high-precision matrices and designs a processing element (PE) array to efficiently process the packed matrices. Extensive experiment results demonstrate that MHDiff can achieve satisfactory performance with negligible quality loss. Chunyu Qi, Xuhang Wang, Yuanzheng Yao, Naifeng Jing, Chen Zhang 0001, Jun Wang 0001, Zhihui Fu, Xiaoyao Liang, Zhuoran Song |
DAC | 2 |
| 2025 | RTSA: A Run-Through Sparse Attention Framework for Video TransformerabstractIn the realm of video understanding tasks, Video Transformer models (VidT) have recently exhibited impressive accuracy improvements in numerous edge devices. However, their deployment poses significant computational challenges for hardware. To address this, pruning has emerged as a promising approach to reduce computation and memory requirements by eliminating unimportant elements from the attention matrix. Unfortunately, existing pruning algorithms face a limitation in that they only optimize one of the two key modules on VidT's critical path: linear projection or self-attention. Regrettably, due to the variation in battery power in edge devices, the video resolution they generate will also change, which causes both linear projection and self-attention stages to potentially become bottlenecks, the existing approaches lack generality. Accordingly, we establish a Run-Through Sparse Attention (RTSA) framework that simultaneously sparsifies and accelerates two stages. On the algorithm side, unlike current methodologies conducting sparse linear projection by exploring redundancy within each frame, we extract extra redundancy naturally existing between frames. Moreover, for sparse self-attention, as existing pruning algorithms often provide either too coarse-grained or fine-grained sparsity patterns, these algorithms face limitations in simultaneously achieving high sparsity, low accuracy loss, and high speedup, resulting in either compromised accuracy or reduced efficiency. Thus, we prune the attention matrix at a medium granularity—sub-vector. The sub-vectors are generated by isolating each column of the attention matrix. On the hardware side, we observe that the use of distinct computational units for sparse linear projection and self-attention results in pipeline imbalances because of the bottleneck transformation between the two stages. To effectively eliminate pipeline stall, we design a RTSA architecture that supports sequential execution of both sparse linear projection and self-attention. To achieve this, we devised an atomic vector-scalar product computation underpinning all calculations in parse linear projection and self-attention, as well as evolving a spatial array architecture with augmented processing elements (PEs) tailored for the vector-scalar product. Experiments on VidT models show that RTSA can save 2.71$\boldsymbol{\times}$to 5.32$\boldsymbol{\times}$ideal computation with$ \lt 1\%$accuracy loss, achieving 105$\boldsymbol{\times}$, 56.8$\boldsymbol{\times}$, 3.59$\boldsymbol{\times}$, and 3.31$\boldsymbol{\times}$speedup compared to CPU, GPU, as well as the state-of-the-art ViT accelerators ViTCoD and HeatViT. Xuhang Wang, Zhuoran Song, Chunyu Qi, Fangxin Liu, Naifeng Jing, Li Jiang 0002, Xiaoyao Liang |
IEEE Trans. Computers | 1 |
| 2025 | Vision Transformer Acceleration via a Versatile Attention Optimization FrameworkabstractVision Transformers (ViTs) have achieved remarkable success across various tasks. However, their deployment is hindered by challenges, such as high memory requirements, long inference latency, and significant power consumption. To address these challenges, existing works optimize one of the two key stages on ViT’s critical path: linear projection or self-attention. Regrettably, we have noticed that both linear projection and self-attention can potentially become bottlenecks as the input image resolution varies, which makes the existing approaches lack generality. Accordingly, in this article, we propose a versatile attention optimization framework. On the algorithm side, we present a SpQuant algorithm that sparsifies weight matrices offline and input matrices online during linear projection as well as tunes the bit-width of the probabilities matrix according to their importance. On the hardware side, we design SQArch architectures to improve the performance of the SpQuant algorithm. The proposed SQArch architecture offers a low-cost preprocess module that predicts and prunes nonkey elements of the input matrix on the fly. Moreover, we design a compute module that supports sparse-sparse matrix multiplications (SpMSpM) and multiple precision computations on a single systolic array for generality. Furthermore, we can address the underutilization and workload imbalance problems by 1) decoupling the rows in the systolic array for enough flexibility and 2) proposing a workload balance scheme for SpMSpM that allows the array to accept data of similar sparsity, thereby reducing synchronization between computing units. Extensive experiment results demonstrate that SQArch can achieve satisfactory performance speedups and energy saving compared to state-of-the-art designs. Xuhang Wang, Qiyue Huang, Xing Li 0031, Haozhe Jiang, Qiang Xu 0001, Xiaoyao Liang, Zhuoran Song |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | InterArch: Video Transformer Acceleration via Inter-Feature Deduplication with Cube-based DataflowabstractIn the realm of video-oriented tasks, Video Transformer models (VidT), an evolution from vision Transformers (ViT), have demonstrated considerable success. However, their widespread application is constrained by substantial computational demands and high energy consumption. Addressing these limitations and thus improving VidT efficiency has become a hot topic. Current methodologies solve this challenge by dividing a video into several features and applying intra-feature sparsity. However, they neglect the crucial point of inter-feature redundancy and often entail prolonged latency in fine-tuning phases. In response, this paper introduces InterArch, a tailored framework designed to significantly enhance VidT efficiency. We first design a novel inter-feature sparsity algorithm consisting of hierarchical deduplication and recovery. The deduplication phase capitalizes on temporal similarities at both block and element levels, enabling the elimination of redundant computations across features in both coarse-grained and fine-grained manners. To prevent long-latency fine-tuning, we employ a lightweight recovery mechanism that constructs approximate features for the sparsified data. Furthermore, InterArch incorporates a regular dataflow strategy, which consolidates sparse features and effectively translates sparse computations into dense ones. Complementing this, we develop a spatial array architecture equipped with augmented processing elements (PEs), specifically optimized for our proposed dataflow. Extensive experiment results demonstrate that InterArch can achieve satisfactory performance speedups and energy saving. Xuhang Wang, Zhuoran Song, Xiaoyao Liang |
DAC | 1 |
| 2024 | Janus: A Flexible Processing-in-Memory Graph Accelerator Toward SparsityabstractGraph application is ever-growing in relational data analysis. However, the memory access patterns become the performance bottleneck in graph analytics and graph neural network (GNN) suffering from single-side and dual-side sparsity, separately. Existing resistive random access memory (RRAM)-based processing-in-memory accelerators reduce data movements but fail to handle both types of sparsity in graph data. To address these issues, our work introduces Janus, a flexible highly compact architecture that is capable of being configured to enable single-sparse mode and dual-sparse mode, to accelerate graph analytics and GNN workloads in compressed mapping, respectively. Upon performing graph analytics with single-side sparsity, Janus employs a tandem-isomorphic-crossbar design both to remove zero-stored footprint, and to eliminate redundant search and sequential indexing. To address the challenge of dual-side sparsity in GNN, Janus still takes a random index access mechanism to gather data rapidly and uses a semi-SPM2 compute paradigm to boost the RRAM-based analog multiplication-and-accumulation in the compressed format. Compared with the state-of-the-art works, Janus outperforms them in both performance and energy efficiency for graph analytics and GNN, respectively. Xing Li 0031, Zhuoran Song, Rachata Ausavarungnirun, Xiao Liu 0033, Xueyuan Liu 0001, Xuan Zhang 0001, Xuhang Wang, Jiayao Ling, Gang Li 0015, Naifeng Jing, Xiaoyao Liang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2023 | DEQ: Dynamic Element-wise Quantization for Efficient Attention ArchitectureabstractAttention-based models, such as transformers, have achieved remarkable success across various tasks. However, their deployment is hindered by challenges such as high memory requirements, long inference latency, and significant power consumption. Quantization has emerged as an effective approach to address these challenges by reducing the bit-width of the model. However, existing quantization algorithms suffer from too coarse-grained quantization granularity or statically determining the bit-width of tokens, lacking the flexibility needed to achieve maximum performance improvement. Accordingly, in this paper, we present a Dynamic Element-wise Quantization (DEQ) algorithm that dynamically tunes tokens’ bit-width according to the importance of elements in the attention possibilities matrix.On the hardware side, we design three versions of DEQ architectures to progressively improve the performance of the DEQ algorithm. The proposed DEQ architecture can address the under-utilization and workload imbalance problems by 1) supporting multiple precision computations on a single systolic array for generality, 2) decoupling the rows in the systolic array for enough flexibility, 3) identifying and parallelizing the independent computations within one systolic array for high parallelism. Extensive experiment results demonstrate that DEQ can achieve satisfactory performance speedups and energy saving compared to state-of-the-art designs. Xuhang Wang, Zhuoran Song, Qiyue Huang, Xiaoyao Liang |
ICCD | 1 |
| 2023 | RealArch: A Real-Time Scheduler for Mapping Multi-Tenant DNNs on Multi-Core AcceleratorsabstractNowadays, the significance of multi-tenant deep neural networks (DNNs) has grown exponentially, particularly for cloud providers who execute multiple DNN models on one server to fulfill the users’ requirements while reducing the computational overhead. To satisfy the heavy computation requirement of multi-tenant DNNs, a feasible approach is to establish a multi-core accelerator housing multiple sub-accelerators. Although many researchers have achieved a certain success by designing either offline schedulers for heterogeneous accelerators or real-time schedulers for homogeneous accelerators, they fail to schedule multi-tenant DNNs to both homogeneous and heterogeneous multi-core accelerators in real time given the large search space and restricted overhead constraint.In this paper, we propose RealArch, a novel real-time scheduler that efficiently schedules multi-tenant DNNs to both homogeneous and heterogeneous multi-core accelerators in real-time. The key idea of RealArch is to quickly find the minimal latency of mapping multi-tenant DNNs to sub-accelerators, considering the occupation of DRAM, sub-accelerators, and buffers. To support the key idea, we first establish lightweight estimation models for multiple sub-accelerators to evaluate the Data Movement (DM) and Execution (EX) time when mapping a layer to them. Then, we design a real-time scheduling algorithm to compute and select the mapping solution with minimal latency. Finally, we build a low-cost hardware scheduler to perform the estimation models and the real-time scheduling algorithm. Extensive experiment results verify that RealArch can exceed the baseline Round Robin scheduling algorithm and two state-of-the-art schedulers AI-MT and MAGMA with acceptable hardware overhead. Xuhang Wang, Zhuoran Song, Xiaoyao Liang |
ICCD | 1 |