VLDB 2026 Research / reviewers in the wild / expert
Weidong Yang 0007
dblp:67/4294-7
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0003-3610-2343ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Viper: An ILP-Based Vectorization Framework for Fully Homomorphic Encryption
Weidong Yang 0007, Xinmo Li, Xiangmin Guo, Jianfei Jiang 0001, Naifeng Jing, Qin Wang 0009, Zhigang Mao, Weiguang Sheng |
ASP-DAC | 1 |
| 2026 | HARMONY: A Hardware-Aware Mapping and Optimizing Framework for Computing-in-Memory AcceleratorsabstractThe increasing adoption of artificial intelligence has spurred the development of specialized deep neural network (DNN) accelerators. Among them, computing-in-memory (CIM) architectures are promising for their in-situ computation capability, which alleviates the computation and data movement bottlenecks of modern DNNs. However, the diversity of models and hardware designs makes it challenging to fully exploit CIM accelerators. Existing approaches often rely on manual mapping or provide limited automation, struggling to integrate general-purpose optimizations with CIM-specific features. In this work, we present HARMONY, a hardware-aware compilation framework for CIM accelerators. At its core is a hardware intermediate representation (IR) that unifies computational and memory abstractions. Based on this IR, HARMONY introduces an automatic mapping algorithm that identifies offloadable operators and constructs a hybrid software–hardware IR. This enables systematic integration of general-purpose and CIM-specific scheduling primitives within a unified search space, which is efficiently explored using reinforcement learning (RL). Extensive evaluations show that HARMONY supports a broader set of operators than existing CIM compilers and consistently delivers substantial performance and energy improvements across diverse DNN workloads. These results demonstrate that HARMONY provides both generality and efficiency, making it a practical compilation solution for CIM accelerators. Xinmo Li, Weidong Yang 0007, Naifeng Jing, Qin Wang 0009, Zhigang Mao, Weiguang Sheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | HEILP: An ILP-Based Scale Management Method for Homomorphic Encryption CompilerabstractRNS-CKKS, a fully homomorphic encryption (FHE) scheme, enabling secure computation on encrypted data, has widely be used in statistical analysis and data mining. However, developing RNS-CKKS programs requires substantial knowledge of cryptography, which is unfriendly to non-expert programmers. A critical obstacle is the scale management, which affects the complexity of programming and performance. Different FHE operations impose specific requirements on the scale and level, necessitating programmer intervention to ensure the recoverability of the results. Furthermore, operations at different levels have a significant impact on program performance. Existing methods rely on heuristic insights or iterative methods to manage the scales of ciphertexts. However, these methods lack a holistic understanding of the optimization space, leading to inefficient exploration and suboptimal performance. This work proposes HEILP, the first constrained-optimization-based approach for scale management in FHE. HEILP expresses node scale decision and scale management operation inserting as an integer linear programming model which can be solved with existing mathematical techniques in one shot. Our method creates a more comprehensive optimization space and enables a faster and more efficient exploration. Experimental results demonstrate that HEILP achieves an average performance improvement of 1.72 x over existing heuristic method, and outperforms a 1.19 x performance improvement with 48.65 x faster compilation time compared to the state-of-the-art iteration-based method. Weidong Yang 0007, Shuya Ji, Jianfei Jiang 0001, Naifeng Jing, Qin Wang 0009, Zhigang Mao, Weiguang Sheng |
DATE | 1 |
| 2025 | RVME: An Efficient Matrix Engine Design Based on Matrix Extension of RISC-VabstractThe rapid advancement of Deep Neural Networks (DNNs) continues to challenge traditional computing architectures, prompting the development of various hardware accelerators. However, insufficient software ecosystem support and limited programmability have significantly constrained the widespread deployment of such hardware accelerators. CPUs, owing to their versatility and widespread applicability, remain strong candidates for DNN acceleration. Several CPU vendors have proposed matrix extensions, but none has publicly disclosed the detailed microarchitecture, and researchers also lack simulation tools for evaluating architectures based on matrix extensions during early-stage design. To address these challenges, we propose RVME, an efficient matrix engine based on a matrix extension of RISC-V, designed as a CPU coprocessor, along with an open-source and configurable simulator built upon gem5. RVME introduces scale-out Outer Product Arrays (OPAs) that achieve bubble-free General Matrix Multiplication (GEMM) execution and superior power efficiency. We also introduce a cache-aware, loop-adaptive mapping framework for RVME that searches for mappings with optimal Energy-Delay Product (EDP). Experimental results show that RVME achieves up to$13.4 \times$speedup and over$21.7 \times$instruction count reduction compared to a RISC-V Vector Extension (RVV)-based design. Additionally, it delivers a peak energy-area efficiency of$1921.4 \text{GOPS} / \mathrm{W} / \text{mm}^{2}$, surpassing state-of-the-art DNN accelerators by more than$6 \times$. Wanqi Chen, Weidong Yang 0007, Renpei Wang, Jianfei Jiang 0001, Naifeng Jing, Qin Wang 0009 |
ICCD | 2 |
| 2025 | MACS: A Multidomain Collaborative Adaptive Clock Scheme for Large-Scale Reconfigurable Dataflow AcceleratorsabstractTo guarantee reliability and correctness, VLSI circuits are designed with conservative margins to maintain timing and power integrity against process, voltage, and temperature (PVT) variations across diverse workloads. However, worst-case PVT and workload conditions rarely occur in practice, resulting in significant timing slack and hence performance and energy loss, especially in reconfigurable dataflow accelerator RDA due to their large-scale and configurable features. Previous studies have attempted to exploit workload or PVT slack, yet achieving limited benefits for reconfigurable dataflow accelerator (RDAs) with large-scale processing element PE arrays. The key issues come from restricted scaling ranges for the clock, insufficient representations for the workload, and unbalanced workloads within processing elementss (PEs). To address these challenges, this article proposes the first multidomain collaborative adaptive clock scheme (MACS) to efficiently exploit both the workload and PVT timing slack for large-scale reconfigurable dataflow acceleratorss (RDAs). MACS partitions the RDA into several clock domains and allows constrained clock domain crossing, which enhances the hardware efficiency with minimal overhead and supports timing validation using conventional static timing analysis (STA) tools. In each domain, an operand-aware workload detection unit is developed, using both static configurations and dynamic operands to assess workload. The detected workload, combined with the monitored PVT conditions, determines the subsequent clock period. Additionally, to enable the exploration of timing slack over a broader range, the period range of the adaptive clock is extended. Experimental results show that MACS achieves a performance improvement of 76.3% or an energy saving of 36.6% with a hardware cost of 3.5%. Shuya Ji, Weidong Yang 0007, Jianfei Jiang 0001, Naifeng Jing, Honglan Jiang, Zhigang Mao, Qin Wang 0009 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | RecPIM: Efficient In-Memory Processing for Personalized Recommendation Inference Using Near-Bank ArchitectureabstractDeep learning (DL)-based personalized recommendation systems consume the major resources in modern AI data centers. The embedding layers with large memory capacity requirement and high bandwidth demand have been identified as the bottleneck of personalized recommendation inference. To mitigate the memory bandwidth bottleneck, near-memory processing (NMP) would be an effective solution which utilizes the through-silicon via (TSV) bandwidth within 3D-stacked DRAMs. However, existing NMP architectures suffer from the limited memory bandwidth caused by hard-to-scale TSVs. To overcome this obstacle, integrating the compute-logic near memory banks becomes a promising but challenging solution, since large memory capacity requirement limits the use of 3D-stacked DRAMs and irregular memory accesses lead to poor data locality, heavy TSV data traffic and low bank-level bandwidth utilization. To address this problem, we propose RecPIM, the first in-memory processing system for personalized recommendation inference using near-bank architecture based on 3D-stacked memory. From the hardware perspective, we introduce a heterogeneous memory system combined with 3D-stacked DRAM and DIMMs to accommodate large embedding tables and provide high bandwidth. By integrating processing logic units near memory banks on DRAM dies, our architecture can exploit the enormous bank-level bandwidth which is much higher than TSV bandwidth. Then, we integrate a small scratchpad memory to exploit the unique data reusability of DL-based personalized recommendation systems. Furthermore, we adopt a unidirectional data communication scheme to avoid additional cross-vault data transfer. From the software perspective, we present a customized programming model to facilitate memory management and task offloading. To reduce the data communication through TSVs and enhance the utilization of bank-level bandwidth, we develop an efficient data mapping scheme by partitioning the vector into smaller subvectors. Experimental results show that RecPIM achieves up to 2.58× speedup and 49.8% energy saving for data movement over the state-of-the-art NMP solution. Weidong Yang 0007, Shuya Ji, Jianfei Jiang 0001, Naifeng Jing, Qin Wang 0009, Zhigang Mao, Weiguang Sheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | An Efficient near-Bank Processing Architecture for Personalized Recommendation SystemabstractPersonalized recommendation systems consume the major resources in modern AI data centers. The memory-bound embedding layers with irregular memory access patterns have been identified as the bottleneck of recommendation systems. To overcome the memory challenges, near-memory processing (NMP) would be an effective solution which provides high bandwidth. Recent work proposes an NMP approach to accelerate the recommendation models by utilizing the through-silicon via (TSV) bandwidth in 3D-stacked DRAMs. However, the total bandwidth provided by TSVs is insufficient for a batch of embedding layers processed in parallel. In this paper, we propose a near-bank processing architecture to accelerate recommendation models. By integrating the compute-logic near memory banks on DRAM dies of the 3D-stacked DRAM, our architecture can exploit the enormous bank-level bandwidth which is much higher than TSV bandwidth. We also present a hardware/software interface for embedding layers offloading. Moreover, we propose an efficient mapping scheme to enhance the utilization of bank-level bandwidth. As a result, our architecture achieves up to 2.10X speedup and 31% energy saving for data movement over the state-of-the-art NMP solution for recommendation acceleration based on 3D-stacked memory. Weidong Yang 0007, Qin Wang 0009, Naifeng Jing, Jianfei Jiang 0001, Zhigang Mao, Weiguang Sheng |
ASP-DAC | 2 |
| 2023 | ACET: An Adaptive Clock Scheme Exploiting Comprehensive Timing Slack for Reconfigurable ProcessorsabstractTo ensure the correctness and reliability, digital circuits are designed with conservative timing margins to accommodate extreme variations in process, voltage, and temperature (PVT) and workload. However, worst-case scenarios rarely occur, leaving the reserved time margins unutilized, which leads to a waste of performance. This issue is particularly significant in reconfigurable processors, as they exhibit substantial workload timing slack in both spatial and temporal domains. Previous researches have mainly focused on either developing PVT slack or exploiting workload slack, but few have simultaneously considered both aspects. Additionally, directly applying existing timing enhancement techniques to reconfigurable processors is challenging due to their complex configurability and diminishing timing slack in array architectures.To address the above challenges, this paper introduces ACET, an Adaptive Clock scheme which Exploits Timing slack comprehensively through hardware-software co-optimization. On the hardware side, ACET incorporates an adaptive clock module that adjusts the clock period based on both workload and PVT conditions. The two conditions are obtained by employing a PVT delay monitor and encoding the workload-dependent delay into the configuration, respectively. Then timing information is transmitted to phase selection module for cycle-level adjustments, to leverage the temporal timing slack. On the software side, to further exploit the spatial timing slack, a scheduling algorithm is proposed, which heuristically rearranges the firing time of operations. Experiments demonstrate that ACET leads to an average performance increase of 70.1% or an equivalent energy saving of 35.6%, with the hardware overhead being only 0.56%. Shuya Ji, Weidong Yang 0007, Jianfei Jiang 0001, Naifeng Jing, Weiguang Sheng, Ang Li 0045, Qin Wang 0009 |
ICCD | 2 |