VLDB 2026 Research / reviewers in the wild / expert
Zhican Wang
dblp:299/5510
· DBLP profile ↗
8ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient 3D Gaussian Splatting with Axis-Shared Rasterization and Order-independent Transmittanceabstract3D Gaussian Splatting (3DGS) has emerged as a powerful technique for novel view synthesis, combining high-quality reconstruction with efficient rendering. It has been widely adopted in domains such as AR/VR, robotics, and autonomous driving. However, achieving real-time performance on resource-constrained platforms remains challenging due to strict power and area budgets. Prior accelerators improve hardware performance but still overlook key inefficiencies, including insufficient rasterization efficiency, poor sorting scalability, and pipeline imbalance. This paper presents an architecture-algorithm co-design to address these challenges. First, we propose axis-shared rasterization, which precomputes and reuses common terms along the X- and Y-axes, reducing multiply-and-accumulate (MAC) operations by up to 38% while preserving high parallelism. Second, we develop a novel order-independent transmittance method that removes the need for explicit sorting by leveraging a lightweight multilayer perceptron (MLP) to directly approximate the transmittance of each Gaussian, enabling efficient alpha blending with negligible quality loss. Third, we design a unified reconfigurable PE array that supports both rasterization and MLP inference, sustaining high utilization without costly sorting hardware. Our experiments demonstrate that our design preserves rendering quality while achieving a 1.33 to 1.88x speedup over state-of-the-art 3DGS accelerators. Our code is open source at https://github.com/WangZhican/ISCA26_3DGS_Acc. Zhican Wang, Guanghui He 0002, Lingjun Gao, Dantong Liu, Shell Xu Hu, Chen Zhang 0001, Zhuoran Song, Nicholas D. Lane, Hongxiang Fan |
ISCA | 1 |
| 2025 | Exploring Code Language Models for Automated HLS-based Hardware Generation: Benchmark, Infrastructure and AnalysisabstractRecent advances in code generation have illuminated the potential of employing large language models (LLMs) for general-purpose programming languages such as Python and C++, opening new opportunities for automating software development and enhancing programmer productivity. The potential of LLMs in software programming has sparked significant interest in exploring automated hardware generation and automation. Although preliminary endeavors have been made to adopt LLMs in generating hardware description languages (HDLs) such as Verilog and SystemVerilog, several challenges persist in this direction. First, the volume of available HDL training data is substantially smaller compared to that for software programming languages. Second, the pre-trained LLMs, mainly tailored for software code, tend to produce HDL designs that are more error-prone. Third, the generation of HDL requires a significantly higher number of tokens compared to software programming, leading to inefficiencies in cost and energy consumption. To tackle these challenges, this paper explores leveraging LLMs to generate High-Level Synthesis (HLS)-based hardware design. Although code generation for domain-specific programming languages is not new in the literature, we aim to provide experimental results, insights, benchmarks, and evaluation infrastructure to investigate the suitability of HLS over low-level HDLs for LLM-assisted hardware design generation. To achieve this, we first finetune pretrained models for HLS-based hardware generation, using a collected dataset with text prompts and corresponding reference HLS designs. An LLM-assisted framework is then proposed to automate end-to-end hardware code generation, which also investigates the impact of chain-of-thought and feedback loops promoting techniques on HLS- design generation. Comprehensive experiments demonstrate the effectiveness of our methods. Jiahao Gai, Zhican Wang, Wanru Zhao, Nicholas D. Lane, Hongxiang Fan |
ASP-DAC | 3 |
| 2025 | VEDA: Efficient LLM Generation Through Voting-based KV Cache Eviction and Dataflow-flexible AcceleratorabstractLarge Language Models (LLMs) excel in natural language processing tasks but pose significant computational and memory challenges for edge deployment due to their intensive resource demands. This work addresses the efficiency of LLM inference by algorithm-hardwaredataflow tri-optimizations. We propose a novel voting-based KV cache eviction algorithm, balancing hardware efficiency and algorithm accuracy by adaptively identifying unimportant kv vectors. From a dataflow perspective, we introduce a flexible-product dataflow and a runtime reconfigurable PE array for matrix-vector multiplication. The proposed approach effectively handles the diverse dimensional requirements and solves the challenges of incrementally varying sequence lengths. Additionally, an element-serial scheduling scheme is proposed for nonlinear operations, such as softmax and layer normalization (layernorm). Results demonstrate a substantial reduction in latency, accompanied by a significant decrease in hardware complexity, from $O(N)$ to $O(1)$. The proposed solution is realized in a custom-designed accelerator, VEDA, which outperforms existing hardware platforms. This research represents a significant advancement in LLM inference on resource-constrained edge devices, facilitating real-time processing, enhancing data privacy, and enabling model customization. Zhican Wang, Hongxiang Fan, Haroon Waris, Gang Wang 0063, Jianfei Jiang 0001, Yanan Sun 0003, Guanghui He 0002 |
DAC | 1 |
| 2025 | DESA: Dataflow Efficient Systolic Array for Acceleration of TransformersabstractTransformers have become prevalent in various Artificial Intelligence (AI) applications, spanning natural language processing to computer vision. Owing to their suboptimal performance on general-purpose platforms, various domain-specific accelerators that explore and utilize the model sparsity have been developed. Instead, we conduct a quantitative analysis of Transformers. (Transformers can be categorized into three types: Encoder-Only, Decoder-Only, and Encoder-Decoder. This paper focuses on Encoder-Only Transformers.) to identify key inefficiencies and adopt dataflow optimization to address them. These inefficiencies arise from1)diverse matrix multiplication,2)multi-phase non-linear operations and their dependencies, and3)heavy memory requirements. We introduce a novel dataflow design to support decoupling with latency hiding, effectively reducing the dependencies and addressing the performance bottlenecks of nonlinear operations. To enable fully fused attention computation, we propose practical tiling and mapping strategies to sustain high throughput and notably decrease memory requirements from$O(N^{2}H)$to$O(N)$. A hybrid buffer-level reuse strategy is also introduced to enhance utilization and diminish off-chip access. Based on these optimizations, we propose a novel systolic array design, named DESA, with three innovations:1)A reconfigurable vector processing unit (VPU) and immediate processing units (IPUs) that can be seamlessly fused within the systolic array to support various normalization, post-processing, and transposition operations with efficient latency hiding.2)A hybrid stationary systolic array that improves the compute and memory efficiency for matrix multiplications with diverse operational intensity and characteristics.3)A novel tile fusion processing that efficiently addresses the low utilization issue in the conventional systolic array during the data setup and offloading. Across various benchmarks, extensive experiments demonstrate that DESA archives$5.0\boldsymbol{\times\thicksim}8.3\boldsymbol{\times}$energy saving over 3090 GPU and$25.6\boldsymbol{\times\thicksim}88.4\boldsymbol{\times}$than Intel 6226R CPU. Compared to the SOTA designs, DESA achieves$11.6\boldsymbol{\times\thicksim}15.0\boldsymbol{\times}$speedup and up to$2.3\times$energy saving over the SOTA accelerators. Zhican Wang, Hongxiang Fan, Guanghui He 0002 |
IEEE Trans. Computers | 1 |
| 2025 | COSA Plus: Enhanced Co-Operative Systolic Arrays for Attention Mechanism in TransformersabstractThe attention mechanism is becoming a vital building block across various modern neural networks, e.g., Transformers. However, it encounters low efficiency when deployed on the general-purpose GPU/CPU platform, which motivates the dedicated accelerator design. Existing accelerators are commonly devised by exploring the potential sparsity in attention mechanism using a hardware-software codesign scheme, which suffers from complicated training, fine-tuning processes, and possible accuracy degradation. More importantly, the sparse pattern only focuses on certain datasets with less generality, and the fine-grained sparse pattern could also bring hardware inefficiency. Instead, we try to solve these issues from another perspective: by systematically analysing the inherent dataflow characteristics of the attention mechanism, we propose the co-operative systolic arrays (COSAs) with an optimized dataflow to support the general purpose attention mechanism and pursue higher computational efficiency. COSA system exploits the high parallelism from the inherent model and leverages run-time configurable hybrid dataflows, i.e., weight and output stationary (OS) for a systolic array (SA) to support the varying matrix multiplication in the attention mechanism. Regarding the cascaded matrix multiplications, COSA proposes levels of fusion methodologies to reduce the off-chip access and enhance processing element (PE) utilization, such as directly using the result of OS as the weight of weight stationary SA by deep fusion. Additionally, the COSA system also provides the solution to hide the latency and radically save the buffer size related to the softmax. Experiment results show that, across various benchmarks, COSA can achieve$2.29-2.60\times $throughput improvement over the traditional SA of the same MAC number, with up to 94.7% PE utilization rate and$8.2\times $less off-chip memory access. Compared with the general-purpose platforms,$7.6-12.4\times $energy efficiency over NVIDIA GeForce 3090 GPU and$35.2-80.9\times $energy efficiency over Intel 6226R server CPU. Zhican Wang, Gang Wang 0063, Guanghui He 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | DEFA: Efficient Deformable Attention Acceleration via Pruning-Assisted Grid-Sampling and Multi-Scale Parallel ProcessingabstractMulti-scale deformable attention (MSDeformAttn) has emerged as a key mechanism in various vision tasks, demonstrating explicit superiority attributed to multi-scale grid-sampling. However, this newly introduced operator incurs irregular data access and enormous memory requirement, leading to severe PE under-utilization. Meanwhile, existing approaches for attention acceleration cannot be directly applied to MSDeformAttn due to lack of support for this distinct procedure. Therefore, we propose a dedicated algorithm-architecture co-design dubbed DEFA, the first-of-its-kind method for MSDeformAttn acceleration. At the algorithm level, DEFA adopts frequency-weighted pruning and probability-aware pruning for feature maps and sampling points respectively, alleviating the memory footprint by over 80%. At the architecture level, it explores the multi-scale parallelism to boost the throughput significantly and further reduces the memory access via fine-grained layer fusion and feature map reusing. Extensively evaluated on representative benchmarks, DEFA achieves 10.1-31.9X speedup and 20.3-37.7X energy efficiency boost compared to powerful GPU platforms. It also rivals the related accelerators by 2.2-3.7X energy efficiency improvement while providing pioneering support of MSDeformAttn. Dongxu Lyu, Zilong Wang 0030, Gang Wang 0063, Zhican Wang, Haomin Li 0002, Guanghui He 0002 |
DAC | 7 |
| 2023 | COSA:Co-Operative Systolic Arrays for Multi-head Attention Mechanism in Neural Network using Hybrid Data Reuse and Fusion MethodologiesabstractAttention mechanism acceleration is becoming increasingly vital to achieve superior performance in deep learning tasks. Existing accelerators are commonly devised dedicatedly by exploring the potential sparsity in neural network (NN) models, which suffer from complicated training, tuning processes, and accuracy degradation. By systematically analyzing the inherent dataflow characteristics of attention mechanism, we propose the Co-Operative Systolic Array (COSA) to pursue higher computational efficiency for its acceleration. In COSA, two systolic arrays that can be dynamically configured into weight or output stationary modes are cascaded to enable efficient attention operation. Thus, hybrid dataflows are simultaneously supported in COSA. Furthermore, various fusion methodologies and an advanced softmax unit are designed. Experimental results show that the COSA-based accelerator can achieve 2.95-28.82× speedup compared with the existing designs, with up to 97.4% PE utilization rate and less memory access. Zhican Wang, Gang Wang 0063, Honglan Jiang, Ningyi Xu, Guanghui He 0002 |
DAC | 1 |
| 2021 | An Energy Efficient Accelerator for Bidirectional Recurrent Neural Networks (BiRNNs) Using Hybrid-Iterative Compression With Error SensitivityabstractRecurrent Neural Networks (RNNs) have been widely used in many sequential applications, such as machine translation, speech recognition and sentiment analysis. Long Term Short Term Memory (LSTM) and Gated Recurrent Unit (GRU) are widely used variants of RNN due to their effectiveness in overcoming gradient vanishing and exploding problems; however, compared to conventional RNN, their massive storage and computation requirements hinder their application. In addition, the recurrent structure of RNNs makes them prone to accumulate errors, resulting in a severe loss of accuracy. In this work, we propose a hybrid-iterative compression (HIC) algorithm for LSTM/GRU. By exploiting the error sensitivity of RNN, the gating units are divided into error-sensitive and error-insensitive groups, that are compressed using different algorithms. By using this approach, a 37.1×/32.3× compression ratio is achieved with negligible accuracy loss for LSTM/GRU. Further, an energy efficient accelerator for bidirectional RNNs is proposed. In this accelerator, the data flow of the matrix operation unit based on the block structure matrix (MOU-S) is improved through rearranging weights; the utilization of BRAM is improved through a fine-grained parallelism configuration of matrix-vector multiplications (MVMs). Meanwhile, the timing matching strategy alleviates the load-imbalance problem between MOU-S and the matrix operation unit based on top- k pruning (MOU-P). When running at 200MHz on Xilinx ADM-PCIE-7V3 FPGA, the proposed design achieves an improvement in energy efficiency in a range of 5%-237% for LSTM networks, and an improvement of 58% for GRU networks compared with state-of-the-art designs. Guocai Nan, Zhengkuan Wang, Chenghua Wang, Bi Wu 0002, Zhican Wang, Weiqiang Liu 0001, Fabrizio Lombardi |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |