EDBT 2026 Demo / reviewers in the wild / expert
Wantong Li 0002
dblp:252/9640-2
· DBLP profile ↗
11ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-8288-393XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-author · 10 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DarkFlow: Hierarchical Digital SiPM Architecture with Low-Loss Dataflow Readout for Dark Matter DetectionabstractDirect dark matter detection experiments require large-scale photon sensing arrays with hundreds of thousands of synchronized readout channels. Silicon photomultipliers (SiPMs) have emerged as a leading candidate for these detectors due to their high integration density, low bias voltage, and superior radiopurity. However, existing digital SiPM readout architectures struggle to simultaneously preserve nanosecond-level temporal resolution for sparse scintillation events and sustain data integrity during high-intensity photon bursts. We propose DarkFlow, a hierarchical digital SiPM architecture that features local data aggregation, compact relative-time encoding, consumer-driven backpressure, and occupancy-aware eDRAM burst buffering within a unified dataflow framework. We show that DarkFlow maintains ultra-low packet loss at billion-photon event rates, where conventional architectures can exceed 80% data loss. Besides, the occupancy-aware refresh achieves a 2.14x improvement in effective refresh rate over conventional global refresh. Hardware evaluation in GlobalFoundries 22nm node confirms that the digital readout datapath accounts for less than 0.86% of the detector area and complies with the strict power budget in liquid argon environments. Aras Repond, Shawn Westerdale, Wantong Li 0002 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | RandEye: On-Sensor Stochastic Image Transformation for Backdoor-Resistant Edge InferenceabstractThe rising threat of backdoor attacks in artificial intelligence (AI) models poses significant risks across diverse applications, including medical diagnostics, autonomous navigation, and surveillance systems. Existing software-based defenses, though effective, are computationally demanding and impractical for low-latency, low-power edge AI inference. This work presents RandEye, the first hardware-based backdoor defense directly embedded into a CMOS image sensor (CIS) and adaptable to various CIS designs. RandEye applies stochastic transformations to captured images directly on the CIS to disrupt malicious input patterns crafted by adversaries, thus effectively deactivating backdoor triggers in compromised models. We introduce the design techniques of hardware-efficient affine transformations and pixel reverse transformation to achieve a lightweight on-sensor defense. RandEye integrated on a 256×256 pixel array is evaluated to achieve a low power profile of 0.35 mW and requires only 1.8% area overhead, while supporting frame rates exceeding 30 FPS for downstream image classification. For backdoored models, RandEye reduces the attack success rate to below 5% while maintaining model accuracy (ACC). When paired with model fine-tuning, RandEye further improves ACC by around 8.8% while still robustly suppressing ASR. Wantong Li 0002, Liuwan Zhu |
ACM Great Lakes Symposium on VLSI | 1 |
| 2025 | Folded Banks: 3D-Stacked HBM Design for Fine-Grained Random-Access BandwidthabstractDespite significant improvements in peak bandwidth, the HBM industry has neglected random-access (irregular) bandwidth, limiting performance in many real-world applications.Improving effective HBM bandwidth is challenging due to power-constrained activations and coarse-grained, long-distance data movement.Rather than addressing these issues directly, hardware vendors have opted for incremental changes, achieving a 6.4× increase in sequential access bandwidth over two generations while leaving the irregular bandwidth challenges unresolved.To remedy this, we introduce Folded Banks (FB-HBM), a novel 3D bank design that redistributes bank subarrays ("folds") across multiple dies and relocates command, control, and global sense amplifiers to an additional base layer.By implementing this logic in a new base layer, we eliminate the DRAM die overheads inherent in previous designs.This architecture enables vertical routing of intra-bank wires-column select lines (CSLs) and master data lines (MDLs)-through thin-pitch through-silicon vias (TSVs) and hybrid bonds, significantly reducing RC power losses.By employing self-timed sense amplifiers, we eliminate costly dummy subarrays * Work done while Wantong Li was an intern at AMD RAD. Vignesh Adhinarayanan, Bradford M. Beckmann, Wantong Li 0002, Mohammad Seyedzadeh, Sergey Blagodurov, Derrick Aguren, Hayden Hyungdong Lee |
ISCA | 3 |
| 2024 | NeuroSim V1.4: Extending Technology Support for Digital Compute-in-Memory Toward 1nm NodeabstractOver the past decade, numerous compute-in-memory (CIM) platforms have been proposed in the literature. While emerging non-volatile memory based analog CIM (ACIM) has been widely studied, its silicon demonstrations are in the mature legacy node (22 nm or above). As an alternative, digital CIM (DCIM) based on static random access memory (SRAM) is recently drawing significant attention, as it enjoys the scaling benefits with the logic process to the leading-edge node (5 nm or below), and does not suffer from the accuracy loss due to process/voltage/temperature (PVT) variations. To assess the potential of DCIM in the future, we release NeuroSim V1.4, a CIM benchmark framework, which supports advanced technology nodes down to 1 nm node. We project the technology parameters (standard cell, transistor and interconnect) using TCAD device simulations, interconnect modeling, and the available industry/IRDS roadmaps. State-of-the-art technology trends such as fin-depopulation, buried power rail, stacked nanosheet, etc are captured in the updated parameters. Technology scaling down to 1 nm enables DCIM to achieve 1.4$\sim 1.8\times $and 44.1$\sim 63.1\times $higher system-level figure of merit than state-of-the-art 7 nm SRAM-based ACIM and 22 nm RRAM-based ACIM, respectively, for representative workloads such as ResNet18 and ResNet34 inference. Anni Lu, Wantong Li 0002, Shimeng Yu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | RAWAtten: Reconfigurable Accelerator for Window Attention in Hierarchical Vision TransformersabstractAfter the success of the transformer networks on natural language processing (NLP), the application of transformers to computer vision has followed suit to deliver unprecedented performance gains on vision tasks including image recognition and object detection. The multi-head self-attention (MSA) is the key component in transformers, allowing the models to learn the amount of attention paid to each input position. In particular, hierarchical vision transformers (HVTs) utilize window-based MSA to capture the benefits of the attention mechanism at various scales for further accuracy enhancements. Despite its strong modeling capability, MSA involves complex operations that make transformers prohibitively costly for hardware deployment. Existing hardware accelerators have mainly focused on the MSA workloads in NLP applications, but HVTs involve different parameter dimensions, input sizes, and data reuse opportunities. Therefore, we design the RAWAtten architecture to target the window-based MSA workloads in HVT models. Each w-core in RAWAtten contains near-memory compute engines for linear layers, MAC arrays for intermediate matrix multiplications, and a lightweight reconfigurable softmax. The w-cores can be combined at runtime to perform hierarchical processing to accommodate varying model parameters. Compared to the baseline GPU, RAWAtten at 40nm provides 2.4x average speedup for running the window-MSA workloads in Swin transformer models while consuming only a fraction of GPU power. In addition, RAWAtten achieves 2x area efficiency compared to prior ASIC accelerator for window-MSA. Wantong Li 0002, Yandong Luo, Shimeng Yu |
DATE | 1 |
| 2023 | Enabling Long-Term Robustness in RRAM-based Compute-In-Memory Edge DevicesabstractThe states of resistive random-access memory (RRAM) have been shown to drift from read voltage stress. When using RRAM as weight elements in deep neural networks (DNNs), this drift degrades inference accuracy over time. To maintain accuracy, RRAM cells must be frequently reprogrammed back to their initial state. We propose a method for reprogramming RRAM memory arrays without accessing the nominal DNN parameters stored off-chip (e.g., in cloud). The method utilizes techniques in linear algebra to calculate the correct state of the RRAM as needed. Resistance drift of RRAM cells were measured from a 40 nm RRAM compute-in-memory (CIM) macro test chip, and the impact on inference accuracy was simulated. Then, the accuracy of the method was evaluated and the recovery in inference accuracy was recorded. Results suggest the method is capable of correctly calculating RRAM resistance states in the presence of resistance drift and other realistic chip noises. The method was also evaluated at reduced precision analog-to-digital conversion to examine its effectiveness in a practical CIM design. James Read, Wantong Li 0002, Shimeng Yu |
ISCAS | 2 |
| 2023 | ENNA: An Efficient Neural Network Accelerator Design Based on ADC-Free Compute-In-Memory SubarraysabstractCompute-in-memory (CIM) is an attractive solution for machine learning hardware acceleration since it merges computation directly into memory arrays, performing parallel multiply-and-accumulate (MAC) operations. The primary challenge in the reported CIM designs is the analog-to-digital converters (ADCs) that digitize analog MAC values for further processing, causing accuracy loss, excessive power dissipation, latency penalty, and area overhead. In this work, we propose ENNA, a novel CIM architecture based on an ADC-free sub-array design, implementing inter-array data processing in an analog manner. A lightweight input encoding scheme based on pulse-width modulation (PWM) is proposed to improve the throughput. We taped-out a prototype macro and validated the proposed ADC-free RRAM array design in TSMC 40nm process. Based on the measured silicon data, we explore the system-level performance with a partition between analog and digital processing at a level higher than the sub-array. The evaluation results show that the proposed accelerator can achieve 73.6~86.4 TOPS/W energy efficiency and 2.3~7 TOPS throughput (normalized to binary operation) tested on various DNN models. Furthermore, we project the proposed design using a heterogeneous 3D integration (H3D) scheme, showing a$3\times \sim 37\times $throughput improvement depending on different tasks and ~50% reduced area overhead compared to 2D design. Hongwu Jiang, Shanshi Huang, Wantong Li 0002, Shimeng Yu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | Temporal Frame Filtering for Autonomous Driving Using 3D-Stacked Global Shutter CIS With IWO Buffer Memory and Near-Pixel ComputeabstractWith the advancement of deep learning to solve autonomous driving problems, the computation and memory requirements have been growing rapidly. Near-pixel compute-based CMOS image sensors (CIS) have been investigated as a potential candidate to perform the initial computations of workloads close to the pixel and reduce data movement. In this work, we design a near-pixel compute CIS capable of implementing a temporal frame filtering network, which rejects redundant image frames targeting autonomous driving applications. To improve performance and avoid image distortion, 3D-stacked global shutter CIS is proposed. This architecture integrates photodiodes with memory and compute units using Cu-Cu hybrid bonding. We propose to use back-end-of-line (BEOL) compatible Tungsten-doped Indium Oxide Transistors (IWO FETs) based embedded DRAM as buffer memory to achieve refresh-free storage and high bandwidth connections between various components. Near-pixel compute circuit is optimized by including sparsity-aware adder tree and using NOR gates as data buffers. The two-tier system comprises photodiodes on tier-1 in 40 nm node, and near-pixel compute and buffer memory on tier-2 in 22 nm node. We perform simulations in Cadence, obtaining an energy efficiency of 65 TOPS/W and a compute density of 1.04 TOPS/mm2 for$8\times8\text{b}$MAC, with a total latency of 1.15 ms/frame. Janak Sharda, Wantong Li 0002, Qiucheng Wu, Shiyu Chang, Shimeng Yu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | H3DAtten: Heterogeneous 3-D Integrated Hybrid Analog and Digital Compute-in-Memory Accelerator for Vision Transformer Self-AttentionabstractAfter the success of the transformer networks on natural language processing (NLP), the application of transformers to computer vision (CV) has followed suit to deliver unprecedented performance gains on vision tasks, including image recognition and object detection. The multihead self-attention (MHSA) is the key component in transformers, allowing the models to learn the amount of attention paid to each input position. Despite its strong modeling capability, MHSA involves complex operations that make transformers prohibitively costly for hardware deployment. Existing acceleration efforts with conventional hardware platforms are challenged by the memory wall. To alleviate the memory wall problem, compute-in-memory (CIM) is a promising solution by storing all model parameters on-chip in compute-capable memory arrays. The footprint of 2-D CIM designs must, however, expand to accommodate the increasingly larger model sizes. In this work, we present a heterogeneous 3-D integrated (H3D) accelerator to target the MHSA workloads in vision transformers. H3D allows the proposed H3DAtten architecture to combine the merits of resistive random access memory (RRAM)-based analog CIM (ACIM) in 40 nm and static random access memory (SRAM)-based digital CIM (DCIM) in 16 nm. We perform comprehensive signaling and thermal analyses to examine the effects of 3-D stacking on the accelerator. Compared to iso-capacity 2-D baseline designs, the proposed 5-tier H3DAtten accelerator achieves$8.4\times $compute density without experiencing accuracy loss on the ImageNet-1k dataset. Wantong Li 0002, Madison Manley, James Read, Ankit Kaul, Muhannad S. Bakir, Shimeng Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | Secure XOR-CIM Engine: Compute-In-Memory SRAM Architecture With Embedded XOR EncryptionabstractCompute-in-memory (CIM), where information can be processed and stored at the same locations, is emerging as a promising paradigm to address the memory wall bottleneck in traditional Von Neumann architectures. Static random-access memory (SRAM) has been demonstrated as a mature candidate for CIM accelerator for deep neural networks (DNNs) due to its availability in advanced technology nodes. However, as SRAM is volatile and could not hold weight after power down, the necessity for downloading models from the cloud to inference engine causes potential threats such as model leaking. Also, saving raw weights of the DNN model stationary in the memory cells will increase the vulnerabilities. This work aims at developing a secure inference engine with a lightweight yet effective countermeasure to protect the DNN models in SRAM-based CIM architecture. We propose a secure XOR-CIM engine with a modified reverse secure sketch protocol to enable on-chip authentication and key processing for XOR-based stream cipher encrypted models. In the XOR-CIM core, we modify the six-transistor SRAM bit cell with dual wordlines to implement XOR decryption without sacrificing the parallel computation’s efficiency. The evaluations at 28 nm show that the XOR-CIM could enhance security, achieving comparable energy efficiency and no throughput loss, with negligible area overhead compared with the normal-CIM design without encryption. Shanshi Huang, Hongwu Jiang, Xiaochen Peng, Wantong Li 0002, Shimeng Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2020 | XOR-CIM: Compute-In-Memory SRAM Architecture with Embedded XOR EncryptionabstractCompute-in-memory (CIM) is a promising approach that exploits the analog computation inside the memory array to speed up the vector-matrix multiplication (VMM) for deep neural network (DNN) inference. SRAM has been demonstrated as a mature candidate for CIM architecture due to its availability in advanced technology node. However, as the weights of the DNN model are stationary in the memory cells, it causes potential threats and vulnerabilities for inference engine such as model leaking. This work aims at developing a lightweight yet effective countermeasure to protect the DNN model in CIM architecture. We modify the 6-transistor SRAM bit cell with dual wordlines to implement XOR cipher without sacrificing the parallel computation's efficiency. The evaluations at 28 nm show that XOR-CIM could provide enhanced security and achieve 1.4× energy efficiency improvement and no throughput loss, with only 2.5% area overhead compared to the normal-CIM design without encryption. Shanshi Huang, Hongwu Jiang, Xiaochen Peng, Wantong Li 0002, Shimeng Yu |
ICCAD | 4 |