VLDB 2026 Research / reviewers in the wild / expert
Lu Lu 0013
dblp:01/2086-13
· DBLP profile ↗
11ranked-venue papers
5as first author
7since 2021 · last 2023
0000-0001-6745-622XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 5 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | A 129.83 TOPS/W Area Efficient Digital SOT/STT MRAM-Based Computing-In-Memory for Advanced Edge AI ChipsabstractThis paper proposes a spin-orbit torque (SOT) magnetoresistive random access memory (MRAM)-based digital computing in memory (CIM) structure for advanced CIM edge AI chips. To avoid frequent data reloading and reduce the area overhead caused by the large transistor in the write path, 11 transistors and 4 shared heavy metal (HM) SOT MRAM bitcell is proposed. It contains 4b weight in a single cell in a compact area able to hold a large capacity with reduced latency. Compared to the previous SRAM+NOR digital CIM design, the SOT/STT MRAM CIM designs occupy only 40% and 30% bitcell area respectively. Additionally, SOT MRAM has a 1.16% leakage current of SRAM at the TT corner at room temperature. The proposed design is verified using statistic simulations in 28nm technology. It achieves 129.83 TOPS/W at 1b/1b/8b precision. Lu Lu 0013, Aarthy Mani, Anh-Tuan Do |
ISCAS | 1 |
| 2023 | 282-to-607 TOPS/W, 7T-SRAM Based CiM with Reconfigurable Column SAR ADC for Neural Network ProcessingabstractCompute in memory ($C$iM) is a promising solution for solving the bottleneck of frequent data interface between memory and processor in Von-Neumann architecture. In this work, a hybrid current/charge domain 7T-SRAM based CiM architecture is proposed to mitigate the PVT-induced RBL variation during computation and thus offer a better linearity without significant impact on the operating frequency and area efficiency. Additionally, a column-referenced 1b to 5b reconfigurable SAR ADC is proposed to support multi-bit output. The proposed design is verified by the Monte-Carlo simulations using 40nm CMOS technology. The 5b mode ADC transferred MAC curve's DNL (LSB) ranges from −0.025 to 0.02 and INL (LSB) ranges from −0.13 to 0.25. The largest RBL variation$(\sigma)$from MAC value −64 to MAC value +64 is 2.08 mV, resulting in a MNIST classification accuracy of 97.5%, which is only 0.1% degradation and Google Speech Command classification accuracy of 80.5%, which is only 0.5% degradation compared to the software baseline, respectively. The whole architecture offers energy efficiency of 282-to-607 TOPS/W for 1-5b output in the MAC operation, which is competitive when compared to other state-of-art$C$iM architectures. Qibang Zang, Wang Ling Goh, Lu Lu 0013, Chengshuo Yu, Junjie Mu, Tony Tae-Hyoung Kim, Bongjin Kim, Dongrui Li, Anh-Tuan Do |
ISCAS | 3 |
| 2023 | BP-SCIM: A Reconfigurable 8T SRAM Macro for Bit-Parallel Searching and Computing In-MemoryabstractThis work presents BP-SCIM: a reconfigurable 8T static random access memory (SRAM) macro for bit-parallel searching and computing in-memory (CIM). The decoupled read/write ports of the employed 8T SRAM bit-cell eliminate read disturbance during search and CIM operations. BP-SCIM can support both in-memory Boolean logic and arithmetic operations. Novel CIM-friendly algorithms and peripheral circuits are proposed to reduce the latency of complex arithmetic operations such as multiplication and division. In addition, BP-SCIM can be configured as either a binary content-addressable memory (CAM) or a ternary CAM for fast searching. A$256\times64$BP-SCIM test chip was implemented in 65-nm CMOS technology. The 8-bit addition and 8-bit multiplication operations can achieve the maximum energy efficiency of 3.11 TOPS/W and 0.17 TOPS/W, respectively at 0.7 V supply. For the binary CAM search operation, BP-SCIM can achieve the minimum energy consumption of 0.91 fJ/bit/search at 87 MHz and 0.8 V supply. Yuzong Chen 0001, Junjie Mu, Lu Lu 0013, Tony Tae-Hyoung Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | A Reconfigurable 8T SRAM Macro for Bit-Parallel Searching and Computing In-MemoryabstractThis work presents BP-SCIM: a reconfigurable 8T SRAM macro for bit-parallel searching and computing in-memory (CIM). BP-SCIM can perform in-memory Boolean logic, arithmetic, and content-addressable memory (CAM) operations. Novel peripheral circuits and algorithms are proposed to support complex arithmetic operations such as multiplication and division. A $256\times 64$ BP-SCIM test chip was implemented in 65-nm CMOS technology. The 8-bit addition and 8-bit multiplication operations can achieve the best energy efficiency of 3.11 TOPS/W and 0.17 TOPS/W, respectively at 0.7 V supply. For the binary CAM search operation, BP-SCIM can achieve the minimum energy consumption of 0.91 fJ/bit/search at 87 MHz and 0.8 V supply. Yuzong Chen 0001, Junjie Mu, Lu Lu 0013, Tony Tae-Hyoung Kim |
ISCAS | 4 |
| 2022 | A 6T SRAM Based Two-Dimensional Configurable Challenge-Response PUF for Portable DevicesabstractThis work proposes a 2-dimensional programable SRAM-based PUF. The selection of challenge groups, orders, and sequence lengths dominates the responses with challenge-response pairs (CRPs) by order of rows$^{\mathrm {(sequence~\textrm {}length- 1)}} \times $columns$^{\mathrm {(sequence~\textrm {}length - 1)}}$. The PUF bit cell has split word-lines with vertical and horizontal connections, the bit-lines are placed orthogonally to generate one-bit data with four cells, the entropy source is enriched to 24 transistors. The proposed PUF supports multiple data maps from a single chip. A test chip was fabricated in 65 nm CMOS technology. Under 0.8V and 20 °C (nominal point), the bit error rate reaches 3%. In a single chip, the hamming distance achieves 42.49% within the same group and different orders of challenges, and 47.32% within the different groups of challenges (when the sequence length is 5). The measured inter-hamming distance between chips is improved to 49.47%. Lu Lu 0013, Taegeun Yoo, Tony Tae-Hyoung Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2021 | A Multi-Functional 4T2R ReRAM Macro Enabling 2-Dimensional Access and Computing In-MemoryabstractThis paper presents a multi-functional resistive random access memory (ReRAM) macro using a novel 4T2R bit-cell. The proposed 4T2R ReRAM enables 2-dimensional (2D) memory access, which offers significant latency and energy reductions for many applications such as matrix operations. Besides non-volatile storage, the proposed 4T2R ReRAM can support two types of computing in-memory operations: ternary content-addressable memory (TCAM) and logic in-memory (LIM). Evaluations on various matrix operations show that the proposed 4T2R ReRAM with 2-D access capability can reduce memory access latency and energy by up to 88% and 82%, respectively compared with conventional 1T1R ReRAM. For TCAM, the proposed 4T2R bit-cell takes a smaller area than SRAM-based TCAM cell, while achieves a comparable search speed. For LIM, we propose an optimized LIM full adder (LIM- FA) that improves the delay and the power by 3.2* and 1.6*, respectively compared with prior LIM-FAs. Yuzong Chen 0001, Lu Lu 0013, Yuncheng Lu, Tony Tae-Hyoung Kim |
ISCAS | 2 |
| 2021 | A Configurable Randomness Enhanced RRAM PUF with Biased Current Sensing SchemeabstractThis paper explains a resistive nonvolatile memory (RRAM)-based physical unclonable function (PUF). The proposed PUF involves more variations and enhances the randomness by using four 1T1R cells to generate one-bit random data. Moreover, we propose a configurable replica column scheme and a biased current sensing amplifier to further improve the randomness. The proposed RRAM PUF is designed in 40nm CMOS technology. The simulated RRAM PUF with configurable replica columns achieves the randomness of 0.5001 with σ of 0.005. It also has a Hamming distance of 0.486 over various responses for a single chip and 0.498 between different chips. Lu Lu 0013, Yuzong Chen 0001, Tony Tae-Hyoung Kim |
ISCAS | 1 |
| 2020 | Reconfigurable 2T2R ReRAM with Split Word-Lines for TCAM Operation and In-Memory ComputingabstractThe increased latency and power consumption due to data movement between memory and ALU have become the major obstacle in modern big-data and machine learning applications. Beyond von-Neumann architectures, particularly in-memory computing, is under intensive research to overcome this memory access bottleneck. In this work, we propose a 2T2R ReRAM structure that supports ternary content addressable memory (TCAM), logic in-memory operations, and in-memory dot product for Deep Neural Networks (DNNs) besides the normal non-volatile memory (NVM) functionality. This is achieved by employing reconfigurable sense amplifiers and novel word-line drivers. The proposed architecture can serve as a high-density storage system as well as an accelerator for data-intensive applications. Simulation results verify that the proposed 2T2R structure functions correctly for TCAM search, logic in-memory operations and in-memory dot product. Yuzong Chen 0001, Lu Lu 0013, Bongjin Kim, Tony Tae-Hyoung Kim |
ISCAS | 2 |
| 2020 | Reconfigurable 2T2R ReRAM Architecture for Versatile Data Storage and Computing In-MemoryabstractNonvolatile memory (NVM)-based computing in-memory (CIM) is a promising solution to data-intensive applications. This work proposes a 2T2R resistive random access memory (ReRAM) architecture that supports three types of CIM operations: 1) ternary content addressable memory (TCAM); 2) logic in-memory (LiM) primitives and arithmetic blocks such as full adder (FA) and full subtractor; and 3) in-memory dot-product for neural networks. The proposed architecture allows the NVM operations in both 2T2R and conventional 1T1R configurations. The proposed LiM full adder (LiM-FA) improves the delay, the static power, and the dynamic power by$3.2\times $,$1.2\times $, and$1.6\times $, respectively, compared with state-of-the-art LiM-FAs. Furthermore, based on different optimization techniques and robustness analysis, a lower precharge voltage is set for each mode. This reduces the TCAM search energy and 1T1R ReRAM access energy by$1.6\times $and$1.14\times $, respectively, compared with the case without optimizations. Yuzong Chen 0001, Lu Lu 0013, Bongjin Kim, Tony Tae-Hyoung Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | A 0.506-pJ 16-kb 8T SRAM With Vertical Read Wordlines and Selective Dual Split Power LinesabstractThis article presents an 8T static random access memory (SRAM) macro with vertical read wordline (RWL) and selective dual split power (SDSP) lines techniques. The proposed vertical RWL reduces dynamic energy consumption during read operation by charging and discharging only selected read bitlines (RBLs). The data-aware SDSP technique combined with vertical write bitlines enhances both the write margin (WM) and the static noise margin (SNM). A 16-kb SRAM test chip fabricated in 65-nm CMOS technology demonstrates the minimum energy consumption of 0.506 pJ at 0.4 V and the minimum operating voltage of 0.26 V. Lu Lu 0013, Taegeun Yoo, Van Loi Le, Tony Tae-Hyoung Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | A Sequence-Dependent Configurable PUF Based on 6T SRAM for Enhanced Challenge Response SpaceabstractThis work proposes a 2D sequence-dependent PUF based on SRAM. It expands challenge to response pairs (CRPs) by the order of rows(sequence length − 1) × columns(sequence length − 1) for reliable authentication. This is achieved by configuring the sequence of SRAM cell selection. Each bit cell has a vertical word-line and a horizontal word-line to utilize the orthogonal word-lines to connect four cells simultaneously to generate one bit data. The proposed technique allows us to generate multiple data maps from one chip. This non-linear behavior also makes the chip more secure. A test chip was fabricated in 65 nm CMOS technology with the area of area is 12580 µm2. The bit error rate is 3% at the nominal point (0.8 V/20°C) and the inter-hamming distance between chips is 0.497. The hamming distance of 0.427 was measured when using the same sequence length with different orders. Lu Lu 0013, Tony Tae-Hyoung Kim |
ISCAS | 1 |