VLDB 2026 Research / reviewers in the wild / expert
Shangshang Yao
dblp:309/4507
· DBLP profile ↗
12ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0001-7217-1712ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 9 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AutoCT: Hybrid Compressor Tree Optimization via Reinforcement Learning with Graph ModelingabstractCompressor tree optimization is a critical step in the design of high-performance arithmetic circuits, such as multipliers or Multiply-Accumulate (MAC) units, where the goal is to efficiently accumulate partial products. Traditional methods often rely on heuristic rules or manual design, which struggle to achieve optimal performance across diverse multiplier scales. In this paper, we propose AutoCT, a novel framework for hybrid compressor tree optimization that leverages reinforcement learning (RL) with graph neural networks (GNNs) to model and optimize compressor trees. By representing the compressor tree as a graph and employing a deep Q-network enhanced with graph attention networks, AutoCT dynamically selects compressor types and configurations to minimize both area and delay. Experimental results highlight that AutoCT reduces the area-delay product by up to $\mathbf{1 5. 0 3} \boldsymbol{\%}$ compared with commercial multiplier IPs for 32-bit designs. Our source code is publicly available at https://github.com/shangshnagyao/AutoCT. Shangshang Yao, Kunlong Li, Li Shen 0007 |
ASP-DAC | 1 |
| 2026 | Approx-L: An Error-Balanced Approximate Floating-Point Divider with Multi-Level Linear Compensation
Shangshang Yao, Huidong Ji, Zuoning Chen |
ACM Great Lakes Symposium on VLSI | 1 |
| 2026 | Enabling Efficient Vector Processing: A Heterogeneous Vector Architecture with in-SRAM ComputingabstractThe growing demand for high-performance, energy-efficient execution of data-parallel workloads has driven the resurgence of vector processors, yet their expanding instruction sets exacerbate the area-performance tradeoff of vector processing units (VPUs). Processing-in-memory (PIM) technique offers a promising path to mitigate this tradeoff by offloading vector operations near data. However, integrating the computing-capable SRAM (C-SRAM) with conventional VPUs introduces significant architectural challenges, including inefficient coordination between heterogeneous devices, the lack of a unified hardware/software interface, and underutilized parallelism within the C-SRAM arrays due to unoptimized data handling. To address these challenges, this article proposes a heterogeneous VPU (HVPU) that seamlessly integrates a standard VPU inside a vector processor with a C-SRAM for more efficient vector processing. HVPU introduces a standardized interface between the processor frontend and the C-SRAM, enabling instruction dispatch, dynamic hazard resolution, and concurrent execution. Furthermore, it employs multiple independent PIM blocks coupled with dual controlling pipelines inside the C-SRAM for further performance improvement. This design facilitates pipelined data loading and computing, effectively hiding memory latency and fully unlocking the parallel potential of C-SRAM. The system is supported by a user-friendly and generic programming model featuring a two-layer extended ISA system and a vector batch pipelining mechanism. Experimental results show that our HVPU-enhanced processor achieves significant speedups of 5.11× to 39.0× over the Xuantie-910 baseline on vector benchmarks, respectively, while reducing energy consumption by 75% on average. Meanwhile, it outperforms state-of-the-art PIM accelerators by 1.33× to 1.97× on the same benchmarks, with minimal area overhead. This demonstrates that our architectural co-design effectively alleviates the area-performance tradeoff in vector processors and offers a scalable heterogeneous architecture template for efficient vector processing. Dunbo Zhang, Shangshang Yao, Qingjie Lang, Junyi Zhu 0016, Li Shen 0007 |
ACM Trans. Archit. Code Optim. | 3 |
| 2026 | DyTopK: Accelerating Top-K SpMV on Embedded FPGAs via Dynamic Floating-Point QuantizationabstractTop-K sparse matrix–vector multiplication is the computational backbone of modern recommendation systems and graph neural networks (GNNs). While shifting these workloads to the edge offers the potential for ultralow latency and enhanced privacy, it presents significant architectural challenges. Recent HBM-based field-programmable gate array (FPGA) accelerators deployed in data centers have shown promise, but their designs are ill-suited for power-constrained embedded FPGAs, which suffer from the limited bandwidth of standard DDR interfaces and restricted on-chip logic resources. In this article, we propose DyTopK, a hardware–software codesigned accelerator tailored specifically for embedded edge devices. Unlike prior approaches that rely on coarse-grained blockwise quantization or complex filtering, DyTopK introduces a novel Dynamic FP4/FP8 Quantization scheme. This mechanism adapts the numerical format at the granularity of individual nonzero elements. To maximize the utility of restricted off-chip bandwidth, we further propose a bandwidth-maximal sparse matrix compression format named packet CSR (PCSR), perfectly aligning with the 64-bit data bus width of standard DDR4 interfaces. Complementing these data-centric optimizations, we design a high-efficiency hardware architecture featuring a scalable processing element array, allowing for flexible expansion based on available logic resources. Experimental results demonstrate that DyTopK achieves significant improvements, delivering$12.4\times $speedup in effective bandwidth utilization and$8.6\times $higher computational throughput compared to state-of-the-art CPU baselines. Furthermore, it outperforms prior FPGA implementations by$2.3\times $in energy efficiency while maintaining negligible accuracy loss (less than 0.5% recall drop) at K = 100. Shangshang Yao, Zuoning Chen |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | Weight-and-Load-Driven Approximation Methodology for Energy-Efficient Neural Network Accelerators
Zhiqiang Wen, Shangshang Yao |
ICA3PP (1) | 4 |
| 2025 | ILP-Driven FPGA Multiplier Synthesis: A Scalable Framework for Area-Latency Co-OptimizationabstractModern computing paradigms impose diverging requirements on arithmetic circuits, with cloud applications prioritizing throughput and edge devices demanding area efficiency under power constraints. While Field-Programmable Gate Arrays (FPGAs) leverage heterogeneous DSP-LUT fabrics for flexibility, their rigid DSP layouts and LUT-centric architectural constraints hinder scalable multiplier designs. Existing FPGA-based approaches face intrinsic scalability limitations from primitive cascading techniques and inflexible performance-resource tradeoffs. This paper proposes an Integer Linear Programming (ILP)-driven framework for Pareto-optimal multiplier synthesis, enabling arbitrary bit-widths via LUT-compressor modeling, application-aware configurations (performance-focused vs. area-minimized modes), and an automated toolchain translating ILP solutions to synthesizable Verilog. Evaluations on Xilinx UltraScale+ series FPGAs demonstrate 27.8% critical path delay reduction in high-performance mode and 21.4% LUT resource savings in areaefficient mode for 16-bit multipliers versus Xilinx LogiCORE™IP. The framework’s adaptive optimization bridges cloud-edge computational divergence, achieving a 15.5-34.2% area-delay product improvement across 8-16b designs. Shangshang Yao, Kunlong Li, Li Shen 0007 |
ICCAD | 1 |
| 2025 | Biased Compressor Based Approximate Multiplier Design Using Genetic AlgorithmabstractMultiplication is the cornerstone in a vast array of applications. The employment of approximate multipliers offers a significant advantage in improving system performance. In this work, we conduct an in-depth and comprehensive analysis of various compressors, analyze their error distributions and propose a set of compressors with biased errors. By configuring heterogeneous compressors via a genetic algorithm and optimizing the input order of compressors, we propose several multipliers. Experimental results show that these multipliers have lower power, delay and area consumption compared with the state-of-the-art approximate multipliers. Notably, the AM_4 multiplier, without reducing the accuracy in image recognition and with an image blending PSNR exceeding 50, has reduced the area overhead by 27.3%, power consumption by 49.2%, and delay by 15.2% compared to the precise multiplier. Zhiqiang Wen, Shangshang Yao, Weikang Xu |
ISCAS | 3 |
| 2025 | Approx-T: Design Methodology for Approximate Multiplication Units via Taylor-ExpansionabstractApproximate computing is emerging as a promising approach to devising energy-efficient IoT systems by exploiting the inherent error-tolerant nature of various applications. In this article, we present Approx-T, to tackle several major challenges according to the prior state-of-the-art (SOTA) approximate multiplication units (AMUs)—lack of comprehensive optimization formulation, asymmetric error distribution, nonadjustable runtime precision, and exponentially growing area complexity when adding up error compensation levels. We innovatively conduct an in-depth study on approximate multiplier via Taylor-expansion to address these issues. 1) Incorporate the Taylor’s theorem into the design concept of approximate arithmetic multipliers. 2) Leverage the inherent symmetrical error distribution of Taylor series to conduct unbiased approximations. 3) Present a runtime configurable error compensation architecture with low-complexity arithmetic operations. We implemented both approximate unsigned and signed integer and floating multiplication arithmetic units and compared with the SOTA works. The experimental results demonstrate that Approx-T surpasses other designs in all metrics, encompassing precision, area utilization, and power consumption. Furthermore, when deployed on an embedded field programmable gate array platform to assess a spectrum of edge computing tasks, Approx-T showcases remarkable performance. Particularly in CNN applications, it achieves up to a$10.7\times $enhancement in energy efficiency, while maintaining negligible impact on accuracy. Shangshang Yao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | ImprLM: An Improved Logarithmic Multiplier Design Approach via Iterative Linear-Compensation and Modified Dynamic SegmentabstractIn this paper, we present ImprLM, an improved logarithmic multiplier design approach via iterative linear-compensation and dynamic segment. Firstly, we have optimized the computational flow of the Mitchell-based logarithmic algorithm to devise a more concise and compact logarithmic multiplier. Subsequently, we introduce an innovative iterative linear error compensation technique, which significantly improves the precision of the logarithmic multiplier compared with the cutting-edge Look-Up Table (LUT) based error compensation strategies. For the purpose of achieving additional logic savings, we introduce modified dynamic segment methods for intermediate variables during multiplication. The synthesis results demonstrate that ImprLM can achieve up to a 22.5% reduction in the power-delay product, while simultaneously enhancing accuracy by 26.8% compared to the state-of-the-art Mitchell-based logarithmic multipliers. Furthermore, we incorporated ImprLM into image processing and the results reveal that ImprLM has a negligible effect on the output quality, while maintaining a low resource consumption. Shangshang Yao, Li Shen 0007 |
ICCD | 1 |
| 2022 | Hardware-Efficient FPGA-Based Approximate Multipliers for Error-Tolerant ComputingabstractWith the increasing demand for data processing, approximate computing is widely used in various fault-tolerant applications such as image processing, computer vision and machine learning. These applications also require a huge number of multiplication operations. In this paper, we are mainly oriented to the softcore approximate multiplier which is implemented on FPGA via encoding the INIT parameter values in the Look-Up-Table (LUT) primitives. Three approximate multipliers with associated carry chain are presented in the manner of reducing LUTs from proposed exact multiplier. An approximate multiplier without carry chain is also presented to further reduce the multiplier's critical path delay and power consumption. We also present an accuracy configurable adder to build high-order approximate multipliers for architectural space exploration. The resolution of the state-of-the-art Mean Relative Error Distance (MRED) and Power-Delay Product (PDP) pareto front is improved and the approximate multiplier we proposed achieves 24.4%, 52.9% and 56.4% reduction in latency, area, and power over the soft multiplier IP core, respectively. Finally, we apply the proposed approximate multiplier design to image processing and convolutional neural networks (CNNs). Compared to advanced approximate multipliers, it offers less energy consumption and area while remaining acceptable qualities. Our designs are open sourced at https://github.com/Yaoshangshang96/FPGA-based_approx_mult to assist further reproducing and development. Shangshang Yao |
FPT | 1 |
| 2022 | FHAM: FPGA-based High-Efficiency Approximate Multipliers via LUT EncodingabstractApproximate computing as a promising technique is widely used in a variety of error tolerance applications, such as computer vision, machine learning and image processing. Since multiplication is one of the basic arithmetic operations used extensively among these applications which creates a great potential to reduce the circuit’s area and power consumption by exploring approximate multipliers. In this paper, we present FHAM, a novel methodology for FPGA-based approximate softcore multiplier architecture via modifying INIT parameter values in LUT primitives. Our proposed approximate multipliers can achieve up to 24.4%, 52.9%, 56.4% improvement in delay, area and power over Xilinx soft multiplier IP core, respectively. We deployed the proposed approximate multiplier designs in the application of image blending, the results demonstrate that proposed multiplier design has a better accuracy-hardware tradeoff than other designs. Shangshang Yao |
ICCD | 1 |
| 2021 | An Efficient Hybrid Parallel Compression Approximate MultiplierabstractApproximate computing has been widely used in many fault-tolerant applications. Multiplication as a key kernel in such applications, it is significant to improve the efficiency of approximate multiplier to achieve high computational performance. This paper proposes a novel approximate multiplier design based on using different compressors for different regions of partial products. We designed two Preprocessing Units (PUs) to explore the best efficiency via increasing the number of sparse partial products. Multiple 8-bit multipliers are designed using Verilog and synthesized under the 45-nm CMOS technology. Compared with the conventional Wallace Tree multiplier, experimental results indicate that one of our proposed multipliers reduce Power-Delay Product (PDP) by 58.5% at most with 0.42% normalized mean error distance. Moreover, a case study of image processing applications is also investigated. Our proposed multipliers can achieve a high peak signal-to-noise ratio of 51.87dB. Compared to the state-of-the-art, the proposed multiplier has a better comprehensive performance in accuracy, area and power consumption. Shangshang Yao, Qiong Wang 0001, Li Shen 0007 |
ICCD | 1 |