VLDB 2026 Research / reviewers in the wild / expert
Yang Wang 0089
dblp:181/2842-89
· DBLP profile ↗
14ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0002-8293-8881ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 2 first-author · 13 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts
Yushu Zhao, Yubin Qin, Yang Wang 0089, Huiming Han, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
ASP-DAC | 3 |
| 2026 | PADE: A Predictor-Free Sparse Attention Accelerator via Unified Execution and Stage FusionabstractAttention-based models have revolutionized AI, but the quadratic cost of self-attention incurs severe computational and memory overhead. Sparse attention methods alleviate this by skipping low-relevance token pairs. However, current approaches lack practicality due to the heavy expense of added sparsity predictor, which severely drops their hardware efficiency. This paper advances the state-of-the-art (SOTA) by proposing a bit-serial enable stage-fusion (BSF) mechanism, which eliminates the need for a separate predictor. However, it faces key challenges: 1) Inaccurate bit-sliced sparsity speculation leads to incorrect pruning; 2) Hardware under-utilization due to finegrained and imbalanced bit-level workloads. 3) Tiling difficulty caused by the row-wise dependency in sparsity pruning criteria. We propose PADE, a predictor-free algorithm-hardware codesign for dynamic sparse attention acceleration. PADE features three key innovations: 1) Bit-wise uncertainty interval-enabled guard filtering (BUI-GF) strategy to accurately identify trivial tokens during each bit round; 2) Bidirectional sparsity-based out-of-order execution (BS-OOE) to improve hardware utilization; 3) Interleaving-based sparsity-tiled attention (ISTA) to reduce both I/O and computational complexity. These techniques, combined with custom accelerator designs, enable practical sparsity acceleration without relying on an added sparsity predictor. Extensive experiments on 22 benchmarks show that PADE achieves$7.43 \times$speed up and$31.1 \times$higher energy efficiency than Nvidia H100 GPU. Compared to SOTA accelerators, PADE achieves$5.1 \times, 4.3 \times$and$3.4 \times$energy saving than Sanger, DOTA and SOFA. Huizheng Wang, Zichuan Wang, Zhiheng Yue, Yang Wang 0089, Chao Li 0009, Yang Hu 0001, Shouyi Yin |
HPCA | 5 |
| 2026 | An Energy-Efficient Transformer Fine-Tuning Processor for Personalized Edge ApplicationsabstractTransformer models have achieved remarkable success in various domains. Given concerns about user privacy, there is an urgent need for on-device fine-tuning of Transformer models at the edge. Transformer fine-tuning faces three key challenges: 1)$O(n^{3})$re-computations during BP/WG save only$O(n^{2})$storage, limiting batch size for fine-tuning speedup. 2)Weakly related tokens account for 87.9% of computations but contribute only 8.7% to accuracy. 3)89.1% of multiplications in matrix multiplications (MM) involve dual near-zero operands, leading to a$1.9\times $increase in logic toggling energy due to frequent exponent/mantissa variations near zero. This paper proposes a Transformer-based processor supporting energy-efficient fine-tuning with three key features to tackle the above challenges. 1)An exponent-stationary re-computing scheduler (ESRS) reduces 44.2% of the storage requirement for each batch. 2)An aggressive linear fitting unit (ALFU) saves 47.4% of the computations in each iteration. 3)A logarithmic domain processing element (LDPE) decreases 36.3% of energy for MM in fine-tuning. Fabricated with 22nm technology, the proposed processor has an area of 6.4 mm2. The proposed Transformer processor achieves a peak energy efficiency of 54.94 TFLOPS/W. It reduces fine-tuning energy by$4.27\times $and offers$3.57\times $speedup for GPT-2. Yang Wang 0089, Yubin Qin, Wende Xu, Zhiheng Yue, Huiming Han, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2025 | MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
Huizheng Wang, Zichuan Wang, Zhiheng Yue, Yousheng Long, Taiquan Wei, Jianxun Yang, Yang Wang 0089, Chao Li 0009, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
MICRO | 7 |
| 2025 | 3D-PATH: A Hierarchy LUT Processing-in-memory Accelerator with Thermal-aware Hybrid Bonding IntegrationabstractLUT-based processing-in-memory (PIM) architectures enable generalpurpose in-situ computing by retrieving precomputed results.However, they suffer from limited computing precision, redundancy, and high latency of off-table access.To address these challenges, we present 3D-PATH, a novel PIM architecture that employs 3D hybrid bonding to integrate a DRAM-LUT, enhancing system capacity and reducing access latency.To further optimize efficiency, 3D-PATH introduces a hierarchical fast-LUT design that reduces storage redundancy and accelerates computation.Additionally, 3D-PATH extends computing precision by efficiently supporting floating-point operations via representation transformation and parallel interleaving banks.While hybrid bonding offers significant benefits, it induces heat dissipation challenges.To address this, we implement thermal-aware hardware that ensures the DRAM Die temperature maintains below the threshold of 85°C.Evaluations on arithmetic and AI workloads demonstrate that 3D-PATH achieves up to 12.68× higher throughput than GPUs and 2.27-7.54×over prior LUT-PIMs, while delivering a 12.24× improvement in floating-point energy efficiency over GPU and 2.13× over a 3D baseline. Zhiheng Yue, Yang Wang 0089, Chao Li 0009, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
MICRO | 2 |
| 2024 | FQP: A Fibonacci Quantization Processor with Multiplication-Free Computing and Topological-Order RoutingabstractWith the continuous advancement of artificial intelligence, neural networks exhibit an escalating parameter size, demanding increased computational power and excessive memory access. Low bit-width quantization emerges as a viable solution to address this challenge. However, conventional low bit-width uniform quantization suffers from a mismatch with the weight and activation data distribution in neural networks, resulting in accuracy degradation. Yang Wang 0089, Yubin Qin, Jiachen Wang 0010, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
DAC | 2 |
| 2024 | MECLA: Memory-Compute-Efficient LLM Accelerator with Scaling Sub-matrix PartitionabstractLarge language models (LLMs) have been showing surprising performance in processing language tasks, bringing a new prevalence to deploy LLM from cloud to edge. However, being a scaling auto-regressive Transformer with a huge parameter amount and generating output one by one, LLM introduces overwhelming memory footprints and computation during its inference, especially from its linear layers. For example, generating 32 output tokens with LLaMA-7B LLM requires 14GB of weight data and performs over 400 billion operations (98% from linear layers), which is far beyond the capability of consumer-level GPU and traditional accelerators. To solve these issues, we propose a memory-compute-efficient LLM accelerator, MECLA, with a parameter-efficient scaling sub-matrix partition method (SSMP). It decomposes large weight matrices into several tiny-scale source sub-matrices (SS) and derived sub-matrices (DS). Each DS can be obtained by scaling the corresponding SS with a scalar. For memory issues, SSMP avoids accessing the full weight matrix but only requires small SS and DS scaling scalars. For computation issues, the proposed MECLA processor fully exploits the intermediate data reuse of matrix multiplication via on-chip matrix regrouping, inner-product multiplication re-association, and outer-product partial sum reuse. Experiments on 20 benchmarks show that MECLA reduces memory access and computation by 83.6% and 72.2%. It achieves an energy efficiency of 7088GOPS/W. Compared to V100 GPU and state-of-the-art Transformer accelerator SpAtten and FACT, MECLA saves 113.14×, 12.99×, and 1.62× higher energy efficiency. Yubin Qin, Yang Wang 0089, Zhiren Zhao, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
ISCA | 2 |
| 2024 | Exploiting Similarity Opportunities of Emerging Vision AI Models on Hybrid Bonding ArchitectureabstractWhile extensive research has focused on optimizing performance and efficiency in vision-based AI accelerators, an unexplored phenomenon, Clustering Similarity Effect, presents a significant opportunity for further improvement. This effect reveals that clusters of neighboring data points exhibit similar values, enabling the potential to skip redundant computations.To fully capitalize on the potential of the Clustering Similarity Effect (CSE), this work integrates hybrid bonding DRAM technology. We conduct a comprehensive analysis of the associated design considerations and integration overhead. Leveraging these insights, we propose a novel CSE-aware architecture specifically tailored for hybrid bonding memory. This architecture facilitates similarity detection and adapts to the inherent data characteristics associated with CSE.Compared with state-of-the-art 2D/2.5D AI accelerators, the hybrid bonding baseline demonstrates an average energy efficiency improvement of $2.89 \times \sim 14.28 \times$ and an area efficiency improvement of $2.67 \times \sim 7.68 \times$. Incorporating the similarity optimizations further enhances energy efficiency and area efficiency improvement to $5.69 \times \sim 28.13 \times$ and $3.82 \times \sim 10.98 \times$, respectively. Zhiheng Yue, Huizheng Wang, Jiahao Fang, Jinyi Deng, Guangyang Lu, Fengbin Tu, Yubin Qin, Yang Wang 0089, Chao Li 0009, Huiming Han, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
ISCA | 10 |
| 2024 | SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated TilingabstractBenefiting from the self-attention mechanism, Transformer models have attained impressive contextual comprehension capabilities for lengthy texts. The requirements of high-throughput inference arise as the large language models (LLMs) become increasingly prevalent, which calls for large-scale token parallel processing (LTPP). However, existing dynamic sparse accelerators struggle to effectively handle LTPP, as they solely focus on separate stage optimization, and with most efforts confined to computational enhancements. By re-examining the end-to-end flow of dynamic sparse acceleration, we pinpoint an ever-overlooked opportunity that the LTPP can exploit the intrinsic coordination among stages to avoid excessive memory access and redundant computation. Motivated by our observation, we present SOFA, a cross-stage compute-memory efficient algorithm-hardware co-design, which is tailored to tackle the challenges posed by LTPP of Transformer inference effectively. We first propose a novel leading zero computing paradigm, which predicts attention sparsity by using log-based add-only operations to avoid the significant overhead of prediction. Then, a distributed sorting and a sorted updating FlashAttention mechanism are proposed with cross-stage coordinated tiling principle, which enables fine-grained and lightweight coordination among stages, helping optimize memory access and latency. Further, we propose a SOFA accelerator to support these optimizations efficiently. Extensive experiments on 20 benchmarks show that SOFA achieves$9.5\times$speed up and$71.5\times$higher energy efficiency than Nvidia A100 GPU. Compared to eight SOTA accelerators, SOFA achieves an average$15.8\times$energy efficiency,$10.3\times$area efficiency and$9.3\times$speed up, respectively. Huizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue, Yubin Qin, Sihan Guan, Qinze Yang, Yang Wang 0089, Chao Li 0009, Yang Hu 0001, Shouyi Yin |
MICRO | 9 |
| 2023 | FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation PredictionabstractTransformer model is becoming prevalent in various AI applications with its outstanding performance. However, the high cost of computation and memory footprint make its inference inefficient. We discover that among the three main computation modules in a Transformer model (QKV generation, attention computation, FFN), it is the QKV generation and FFN that contribute to the most power cost. While the attention computation, focused by most previous works, only has decent power share when dealing with extremely long inputs. Therefore, in this paper, we propose FACT, an efficient algorithm-hardware co-design optimizing all three modules of Transformer. We first propose an eager prediction algorithm which predicts the attention matrix before QKV generation. It further detects the unnecessary computation in QKV generation and assigns mixed-precision FFN with the predicted attention, which helps improve the throughput. Further, we propose FACT accelerator to efficiently support eager prediction with three designs. It avoids the large overhead of prediction by using log-based add-only operations for prediction. It eliminates the latency of prediction through an out-of-order scheduler that makes the eager prediction and computation work in full pipeline. It additionally avoids memory access conflict in the mixed-precision FFN with a novel diagonal storage pattern. Experiments on 22 benchmarks show that our FACT improves the throughput of the whole Transformer by 3.59× on the geomean average. It achieves an enviable 47.64× and 278.1× energy saving when computing attention, compared to previous attention-optimization-only SOTA works ELSA and Sanger. Further, FACT achieves an energy efficiency of 4388 GOPS/W performing the whole Transformer layer on average, which is 94.98× higher than Nvidia V100 GPU. Yubin Qin, Yang Wang 0089, Dazheng Deng, Zhiren Zhao, Leibo Liu, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
ISCA | 2 |
| 2023 | Reconfigurability, Why It Matters in AI Tasks Processing: A Survey of Reconfigurable AI ChipsabstractNowadays, artificial intelligence (AI) technologies, especially deep neural networks (DNNs), play an vital role in solving many problems in both academia and industry. In order to simultaneously meet the demand of performance, energy efficiency and flexibility in DNN processing, various reconfigurable AI chips have been proposed in the past several years. They are based on FPGA or CGRA platforms and have domain-specific reconfigurability to customize the computing units and data paths for different DNN tasks without re-produce the chips. This paper surveys typical reconfigurable AI chips from three reconfiguration hierarchies: processing element level, processing element array level, and chip level. Each reconfiguration hierarchy covers a set of important optimization techniques for DNN computation which are frequently adopted in real life. This paper lists the reconfigurable AI chip works in chronological order, discusses the hardware development process for each optimization techniques, and analyzes the necessity of reconfigurability in AI tasks processing. The trends of each reconfiguration hierarchy and insights about the cooperation of techniques from different hierarchies are also proposed. Shaojun Wei, Xinhan Lin, Fengbin Tu, Yang Wang 0089, Leibo Liu, Shouyi Yin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | PL-NPU: An Energy-Efficient Edge-Device DNN Training Processor With Posit-Based Logarithm-Domain ComputingabstractEdge device deep neural network (DNN) training is practical to improve model adaptivity for unfamiliar datasets while avoiding privacy disclosure and huge communication cost. Nevertheless, apart from feed-forward (FF) as inference, DNN training still requires back-propagation (BP) and weight gradient (WG), introducing power-consuming floating-point computing requirements, hardware underutilization, and energy bottleneck from excessive memory access. This paper proposes a DNN training processor named PL-NPU to solve the above challenges with three innovations. First, a posit-based logarithm-domain processing element (PE) adapts to various training data requirements with a low bit-width format and reduces energy by transferring complicated arithmetics into simple logarithm domain operation. Second, a reconfigurable inter-intra-channel-reuse dataflow dynamically adjusts the PE mapping with a regrouping omega network to improve the operands reuse for higher hardware utilization. Third, a pointed-stake-shaped codec unit adaptively compresses small values to variable-length data format while compressing large values to fixed-length 8b posit format, reducing the memory access for breaking the training energy bottleneck. Simulated with 28nm CMOS technology, the proposed PL-NPU achieves a maximum frequency of 1040MHz with 343mW and 5.28mm$\mathbf {^{2}}$. The peak energy efficiency is 3.87TFLOPS/W for 0.6V at 60MHz. Compared with the state-of-the-art training processor, PL-NPU reaches$3.75\times $higher energy efficiency and offers$1.68\times $speedup when training ResNet18. Yang Wang 0089, Dazheng Deng, Leibo Liu, Shaojun Wei, Shouyi Yin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | SWPU: A 126.04 TFLOPS/W Edge-Device Sparse DNN Training Processor With Dynamic Sub-Structured Weight PruningabstractWhen deploying deep neural networks (DNNs), edge devices training is practical to improve model adaptivity for various user-specific scenarios while avoiding privacy disclosure. However, the training computation is intolerable for edge devices. It inspires sparse DNN training (SDT) into the limelight, which reduces training computation by dynamic weight pruning. Generally, SDT has two strategies based on the pruning granularity: the structured or the unstructured. Unfortunately, both of them suffer from limited training efficiency due to the gap between pruning granularity and hardware implementation. The former is hardware-friendly but has a low pruning ratio, indicating limited computation reduction. The latter has a high pruning ratio, but the unbalanced workload decreases utilization and irregular sparsity distribution causes considerable sparsity processing overhead. This paper proposes a software-hardware co- design to bridge the gap for improving the efficiency of SDT. On the algorithm side, a sub-structured pruning method, achieved with hybrid shape-wise and line-wise pruning, generates a high sparsity ratio and keeps the hardware-friendly property. On the hardware side, a sub-structured weight processing unit (SWPU) effectively handles the hybrid sparsity with three techniques. First, SWPU dynamically reorders the computation sequence with hamming-distance-based clustering, balancing the irregular workload. Second, SWPU performs runtime scheduling by exploiting the feature of sub-structured sparse convolution through a detect-before-load controller, which skips redundant memory access and sparsity processing. Third, SWPU performs sparse convolution by compressing operands with spatial disconnect log-based routing and recovers their location with bi-directional switching, avoiding the power-consumed routing logic. Synthesized with 28nm CMOS technology, SWPU can enable 0.56V-to-1.0V supply voltage with a maximum frequency of 675 MHz. It achieves a 50.1% higher pruning ratio than structured pruning and$1.53\times $higher energy efficiency than unstructured pruning. The peak energy efficiency of SWPU is 126.04TFLOPS/W, outperforming the state-of-the-art training processor by$1.67\times $. When training a ResNet-18 model, SWPU reduces$3.72\times $energy and offers$4.69\times $speedup than previous sparse training processors. Yang Wang 0089, Yubin Qin, Leibo Liu, Shaojun Wei, Shouyi Yin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | STC: Significance-aware Transform-based Codec Framework for External Memory Access ReductionabstractDeep convolutional neural networks (DCNNs), with extensive computation, require considerable external memory bandwidth and storage for intermediate feature maps. External memory accesses for feature maps become a significant energy bottleneck for DCNN accelerators. Many works have been done on quantizing feature maps into low precision to decrease the costs for computation and storage. There is an opportunity that the large amount of correlation among channels in feature maps can be exploited to further reduce external memory access. Towards this end, we propose a novel compression framework called Significance-aware Transform-based Codec (STC). In its compression process, significance-aware transform is introduced to obtain low-correlated feature maps in an orthogonal space, as the intrinsic representations of original feature maps. The transformed feature maps are quantized and encoded to compress external data transmission. For the next layer computation, the data will be reloaded with STC's reconstruction process. The STC framework can be supported with a small set of extensions to current DCNN accelerators. We implement STC extensions to the baseline TPU architecture for hardware evaluation. The strengthened TPU achieves average reduction of 2.57x in external memory access, 1.95x~2.78x improvement of system-level energy efficiency, with a negligible accuracy loss of only 0.5%. Fengbin Tu, Man Shi, Yang Wang 0089, Leibo Liu, Shaojun Wei, Shouyi Yin |
DAC | 4 |