VLDB 2026 Research / reviewers in the wild / expert
Anni Lu
dblp:260/6893
· DBLP profile ↗
14ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0002-4415-0866ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Monolithic 3D FPGA Design and Synthesis with Back-End-of-Line Configuration MemoriesabstractThis work presents a novel monolithic 3D (M3D) FPGA architecture that leverages stackable back-end-of-line (BEOL) transistors to implement configuration memory and pass gates, significantly improving area, latency, and power efficiency. By integrating n-type (W -doped In2O3) and p-type (SnO) amorphous oxide semiconductor (AOS) transistors in the BEOL, Si SRAM configuration bits are substituted with a less leaky equivalent that can be programmed at logic-compatible voltages. BEOL-compatible AOS transistors are currently under extensive research and development in the device community, with investment by leading foundries, from which reported data is used to develop robust physics-based models in TCAD that enable circuit design. The use of AOS pass gates reduces the overhead of reconfigurable circuits by mapping FPGA switch block (SB) and connection block (CB) matrices above configurable logic blocks (CLBs), thereby increasing the proximity of logic elements and reducing latency. By interfacing with the latest Verilog-to-Routing (VTR) suite, an AOS-based M3D FPGA implemented in 7 nm technology is demonstrated with $3.4 \times$ lower area-time squared product ($\mathbf{A T}^{2}$), 27% lower critical path latency, and 26% lower reconfigurable routing block power on benchmarks including hyperdimensional computing and large language models (LLMs). Faaiq G. Waqar, Anni Lu, Zifan He, Jason Cong, Shimeng Yu |
DAC | 3 |
| 2025 | Digital Compute-in-Memory Ising Annealer with Ferroelectric Capacitor-Based nvSRAM for Combinatorial Optimization ProblemsabstractCombinatorial optimization problems (COPs) have a wide range of applications. The Ising model-based annealer is gaining attention for its efficiency and speed in finding approximate solutions. However, building an Ising machine that is area- and energy-efficient, scalable, and with low compute latency in CMOS is challenging. In this paper, we present a digital compute-in-memory (DCIM) Ising annealer that uses ferroelectric capacitor (FeCap)-based nvSRAM to solve COPs like the Traveling Salesman Problem (TSP). By using weak recall operations, our design eliminates the need to reload weights, significantly reducing energy consumption and speeding up processing compared to other approaches. Simulations using a 16nm PDK demonstrate that our nvSRAM-based DCIM array maintains accuracy while reducing latency by up to 55.0% and energy by 49.6% compared to prior work implemented with conventional SRAM DCIM array. Algorithm validation further shows that the random noise introduced by weak recall can be effectively utilized in the annealing process. Yuyao Kong, Jianwei Jia, Anni Lu, Faaiq G. Waqar, Yuan-Chun Luo, Hai Li 0001, Ian A. Young, Shimeng Yu |
ISCAS | 3 |
| 2024 | A Cross-layer Framework for Design Space and Variation Analysis of Non-Volatile Ferroelectric Capacitor-Based Compute-in-Memory AcceleratorsabstractUsing non-volatile “capacitive” crossbar arrays for compute-in-memory (CIM) offers higher energy and area efficiency compared to “resistive” crossbar arrays. However, the impact of device-to-device (D2D) variation and temporal noise on the system-level performance has not been explored yet. In this work, we provide an end-to-end methodology that incorporates experimentally measured D2D variation into the design space exploration from capacitive weight cell design, CIM array with peripheral circuits, to the inference accuracy of SwinV2-T vision transformer and ResNet-50 on the ImageNet dataset. Our framework further assesses the system’s power, performance, and area (PPA) by considering cell design, circuit structure, and model selection. We explore the design space using an early stopping algorithm to produce optimal designs while meeting strict inference accuracy requirements. Overall findings suggest that the capacitive CIM system is robust against D2D variation and noise, outperforming its resistive counterpart by $6.95 \times$ and $14.1 \times$ for the optimal design in the figure of merit (TOPS/W $\times {\mathrm {TOPS}}/\mathrm{mm}^{2}$) for ResNet-50 and SwinV2-T respectively. Yuan-Chun Luo, James Read, Anni Lu, Shimeng Yu |
ASPDAC | 3 |
| 2024 | Digital CIM with Noisy SRAM Bit: A Compact Clustered Annealer for Large-Scale Combinatorial OptimizationabstractCombinatorial optimization problems (COP) are NP-hard and intractable to solve using conventional computing. The Ising model-based annealer has gained increasing attention recently due to its efficiency and speed in finding approximate solutions. However, Ising solvers for travelling salesman problems (TSP) usually suffer from a scalability issue due to quadratically increasing number of spins. In this paper, we propose a digital computing-in-memory (CIM) based clustered annealer to solve tens of thousands of city-scale TSP with only a few mega-byte (MB) of static random access memory (SRAM), using hierarchical clustering to solve input sparsity and digital CIM flexibility to solve weight sparsity. The intrinsic process variations between SRAM devices are utilized to generate the noisy bit errors during pseudo-read under reduced supply voltage, realizing the annealing process. The design space of cluster size and programmability is explored to understand the trade-offs of solution quality and hardware cost, for TSP scale ranging from 3080 to 85900 cities. The proposed design speeds up the convergence by >109× with <25% solution quality overhead compared with the CPU baseline. The comparison with state-of-the-art scalable annealers shows a >1013× improvement on functionally normalized area and power. Anni Lu, Yuan-Chun Luo, Hai Li 0008, Ian A. Young, Shimeng Yu |
DAC | 1 |
| 2024 | A Heterogeneous Platform for 3D NAND-Based In-Memory Hyperdimensional Computing Engine for Genome Sequencing ApplicationsabstractHyperdimensional (HD) computing is a promising paradigm for large-scale genome sequencing. In prior work, we proposed a 3D NAND-based HD computing engine as an energy-efficient solution for sequencing several gigabytes or terabytes of genomic data. In this work, we introduce an improved HD computing engine for genome sequencing that leverages heterogeneous 3D integration techniques. We employ Cu-Cu hybrid bonding and CMOS under array (CuA) technologies to integrate the digital logic tier with the 3D NAND-based associative memory tier. We benchmark the performance of the proposed hardware design using a dataset of 10,575 microorganism genomes. The results indicate the robustness of the classification accuracy despite device non-idealities. Compared to conventional methods, the proposed design reduces the system-level energy consumption by$1000\times $, provided the data is not offloaded from the solid-state drive to the computing units. Po-Kai Hsu, Vaidehi Garg, Anni Lu, Shimeng Yu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | NeuroSim V1.4: Extending Technology Support for Digital Compute-in-Memory Toward 1nm NodeabstractOver the past decade, numerous compute-in-memory (CIM) platforms have been proposed in the literature. While emerging non-volatile memory based analog CIM (ACIM) has been widely studied, its silicon demonstrations are in the mature legacy node (22 nm or above). As an alternative, digital CIM (DCIM) based on static random access memory (SRAM) is recently drawing significant attention, as it enjoys the scaling benefits with the logic process to the leading-edge node (5 nm or below), and does not suffer from the accuracy loss due to process/voltage/temperature (PVT) variations. To assess the potential of DCIM in the future, we release NeuroSim V1.4, a CIM benchmark framework, which supports advanced technology nodes down to 1 nm node. We project the technology parameters (standard cell, transistor and interconnect) using TCAD device simulations, interconnect modeling, and the available industry/IRDS roadmaps. State-of-the-art technology trends such as fin-depopulation, buried power rail, stacked nanosheet, etc are captured in the updated parameters. Technology scaling down to 1 nm enables DCIM to achieve 1.4$\sim 1.8\times $and 44.1$\sim 63.1\times $higher system-level figure of merit than state-of-the-art 7 nm SRAM-based ACIM and 22 nm RRAM-based ACIM, respectively, for representative workloads such as ResNet18 and ResNet34 inference. Anni Lu, Wantong Li 0002, Shimeng Yu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | Endurance-Aware Compiler for 3-D Stackable FeRAM as Global Buffer in TPU-Like ArchitectureabstractEmerging nonvolatile memories as embedded memories offer low leakage power and high memory density, compared to the static random access memory (SRAM) and embedded dynamic random access memory (eDRAM) at the same technology node. However, the emerging memories generally suffer from limited cycling endurance. For read/write intensive applications, the limited endurance could become a bottleneck that limits the lifetime of the overall system. In this work, Intel’s reported prototype 3-D stackable ferroelectric random access memory (FeRAM) is considered as the global buffer memory of a tensor-processing-unit (TPU)-like architecture. An endurance-aware compiler is proposed to evaluate the maximum number of deep neural network (DNN) trainings considering the experimentally measured endurance limit. In addition, the proposed compiler applies two strategies to alleviate the endurance issue. The first strategy is wear leveling, and the second strategy is the dual-mode operation between volatile and nonvolatile modes. The maximum numbers of trainings increase by$6\times $to$300\times $and$4\times $to$58\times $thanks to the wear-leveling and dual-mode operations, respectively. Finally, a guideline of the system endurance (maximum number of trainings) is provided with given memory device endurance to bridge the gap between memory device engineers and system designers. Yuan-Chun Luo, Anni Lu, Yandong Luo, Sou-Chi Chang, Uygar Avci, Shimeng Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | Thermally Constrained Codesign of Heterogeneous 3-D Integration of Compute-in-Memory, Digital ML Accelerator, and RISC-V Cores for Mixed ML and Non-ML WorkloadsabstractHeterogeneous 3-D (H3D) integration not only reduces the chip form factor and fabrication cost but also allows the merging of diverse compute paradigms that suit different applications. This is especially attractive when modern algorithms, such as the augmented reality/virtual reality (AR/VR) workloads, consist of mixed machine learning (ML) and non-ML workloads. To date, codesign that considers the thermal, latency, and power constraints of H3D hardware is largely unexplored. In this work, a thermally aware framework for H3D hardware design is developed to evaluate the thermal, latency, and power trade-offs for a heterogeneous system with compute-in-memory (CIM), digital ML cores, and RISC-V cores. The framework solves for runtime tunable operating points described as the optimal speedup factor, the number of activated RISC-V cores, the cooling coefficient, and the activity rate based on user-defined criteria, achieving up to 135 TOPS and 215 TOPS/W under$74~^{\circ }$C for the AR/VR workloads. Yuan-Chun Luo, Anni Lu, Janak Sharda, Moritz Scherer 0001, Jorge Gomez 0001, Syed Shakib Sarwar, Ziyun Li 0001, Reid Frederick Pinkham, Barbara De Salvo, Shimeng Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2023 | CLUE: Cross-Layer Uncertainty Estimator for Reliable Neural Perception using Processing-in-Memory AcceleratorsabstractOne of the primary challenges of deploying deep neural networks (DNNs) is ensuring their reliable performance in unpredictable edge environments, which are often disrupted by a variety of uncertainties and variations. Estimating uncertainty is crucial in order to understand the reliability of task predictions and prevent system failures. However, quantifying uncertainty stemming from non-ideal properties of processing hardware has not yet been thoroughly studied. To address this, we present Cross-Layer Uncertainty Estimator (CLUE), which quantifies task uncertainty originating from both sensing/processing hardware variations and DNN algorithm uncertainty. Our experimental results demonstrate that CLUE provides uncertainty with up to 80.4% less calibration error and only 12% of energy overheads compared to using task DNN solely. Furthermore, CLUE is able to detect unreliable tasks that stem from processing hardware variations, which prior uncertainty estimators were unable to achieve. Finally, we demonstrate an adaptive control of processing hardware using CLUE, which allows a dynamic trade-off control between task accuracy and energy consumption. Minah Lee, Anni Lu, Mandovi Mukherjee, Shimeng Yu, Saibal Mukhopadhyay |
IJCNN | 2 |
| 2022 | Robust Processing-In-Memory With Multibit ReRAM Using Hessian-Driven Mixed-Precision ComputationabstractThis article presents an algorithmic approach to design reliable deep neural networks (DNNs) in the presence of stochastic variations in the network parameters induced by process variations in the bit cells in a processing-in-memory (PIM) architecture. We propose and derive a Hessian-based sensitivity metric that can be computed without computing or storing the full Hessian to identify and protect the “important” network parameters while allowing large variations in unprotected parameters. We also show that this metric can be used to aggressively quantize unprotected network parameters in the PIM for improved inference efficiency and compute density. Experiments on modern DNNs like ResNet, MobileNetv2, and DenseNet on CIFAR10 using measured RRAM device data shows the effectiveness of our approach. Saurabh Dash, Yandong Luo, Anni Lu, Shimeng Yu, Saibal Mukhopadhyay |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | A Runtime Reconfigurable Design of Compute-in-Memory based Hardware AcceleratorabstractCompute-in-memory (CIM) is an attractive solution to address the “memory wall” challenges for the extensive computation in machine learning hardware accelerators. Prior CIM-based architectures, though can adapt to different neural network models during the design time, they are implemented to different custom chips. Therefore, a specific chip instance is restricted to a specific network during runtime. However, the development cycle of the hardware is normally far behind the emergence of new algorithms. In this paper, a runtime reconfigurable design methodology of CIM-based accelerator is proposed to support a class of convolutional neural networks running on one pre-fabricated chip instance. First, several design aspects are investigated: 1) reconfigurable weight mapping method; 2) input side of data transmission, mainly about the weight reloading; 3) output side of data processing, mainly about the reconfigurable accumulation. Then, system-level performance benchmark is performed for the inference of different models like VGG-8 on CIFAR-10 dataset and AlexNet, GoogLeNet, ResNet-18 and DenseNet-121 on ImageNet dataset to measure the tradeoffs between runtime reconfigurability, chip area, memory utilization, throughput and energy efficiency. Anni Lu, Xiaochen Peng, Yandong Luo, Shanshi Huang, Shimeng Yu |
DATE | 1 |
| 2021 | DNN+NeuroSim V2.0: An End-to-End Benchmarking Framework for Compute-in-Memory Accelerators for On-Chip TrainingabstractDNN+NeuroSim is an integrated framework to benchmark compute-in-memory (CIM) accelerators for deep neural networks, with hierarchical design options from device-level, to circuit level and up to algorithm level. A python wrapper is developed to interface NeuroSim with a popular machine learning platform: Pytorch, to support flexible network structures. The framework provides automatic algorithm-to-hardware mapping, and evaluates chip-level area, energy efficiency and throughput for training or inference, as well as training/inference accuracy with hardware constraints. Our prior inference version of DNN+NeuroSim framework available athttps://github.com/neurosim/DNN_NeuroSim_V1.2was developed to estimate the impact of reliability in synaptic devices, and analog-to-digital converter (ADC) quantization loss on the accuracy and hardware performance of an inference engine. In this work, we further investigated the impact of the “analog” emerging nonvolatile memory (eNVM)’s nonideal device properties for on-chip training. By introducing the nonlinearity, asymmetry, device-to-device and cycle-to-cycle variation of weight update into the python wrapper, and peripheral circuits for error/weight gradient computation in NeuroSim core, we benchmarked CIM accelerators based on state-of-the-art SRAM and eNVM devices for VGG-8 on CIFAR-10 dataset, revealing the crucial specs of synaptic devices for on-chip training. The latest training version of the DNN+NeuroSim framework is available athttps://github.com/neurosim/DNN_NeuroSim_V2.1. Xiaochen Peng, Shanshi Huang, Hongwu Jiang, Anni Lu, Shimeng Yu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | A Runtime Reconfigurable Design of Compute-in-Memory-Based Hardware Accelerator for Deep Learning InferenceabstractCompute-in-memory (CIM) is an attractive solution to address the “memory wall” challenges for the extensive computation in deep learning hardware accelerators. For custom ASIC design, a specific chip instance is restricted to a specific network during runtime. However, the development cycle of the hardware is normally far behind the emergence of new algorithms. Although some of the reported CIM-based architectures can adapt to different deep neural network (DNN) models, few details about the dataflow or control were disclosed to enable such an assumption. Instruction set architecture (ISA) could support high flexibility, but its complexity would be an obstacle to efficiency. In this article, a runtime reconfigurable design methodology of CIM-based accelerators is proposed to support a class of convolutional neural networks running on one prefabricated chip instance with ASIC-like efficiency. First, several design aspects are investigated: (1) the reconfigurable weight mapping method; (2) the input side of data transmission, mainly about the weight reloading; and (3) the output side of data processing, mainly about the reconfigurable accumulation. Then, a system-level performance benchmark is performed for the inference of different DNN models, such as VGG-8 on a CIFAR-10 dataset and AlexNet GoogLeNet, ResNet-18, and DenseNet-121 on an ImageNet dataset to measure the trade-offs between runtime reconfigurability, chip area, memory utilization, throughput, and energy efficiency. Anni Lu, Xiaochen Peng, Yandong Luo, Shanshi Huang, Shimeng Yu |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2020 | Benchmark of the Compute-in-Memory-Based DNN Accelerator With Area ConstraintabstractCompute-in-memory (CIM) is a promising computing paradigm to accelerate the inference of deep neural network (DNN) algorithms due to its high processing parallelism and energy efficiency. Prior CIM-based DNN accelerators mostly consider full custom design, which assumes that all the weights are stored on-chip. For lightweight smart edge devices, this assumption may not hold. In this article, CIM-based DNN accelerators are designed and benchmarked under different chip area constraints. First, a scheduling strategy and dataflow for DNN inference is investigated when only part of the weights can be stored on-chip. Two weight reload schemes are evaluated: 1) reload partial weights and reuse input/output feature maps and 2) load a batch of input and reuse the partial weights on-chip across the batch. Then, system-level performance benchmark is performed for the inference of ResNet-18 on ImageNet data set. The design tradeoffs with different area constraints, dataflow, and device technologies [static random access memory (SRAM) versus ferroelectric field-effect transistor (FeFET)] are discussed. Anni Lu, Xiaochen Peng, Yandong Luo, Shimeng Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |