VLDB 2026 Research / reviewers in the wild / expert
Junjie Mu
dblp:283/1108
· DBLP profile ↗
10ranked-venue papers
2as first author
9since 2021 · last 2025
0000-0002-6496-6539ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 2 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A 1Mb RRAM Macro with Bipolar Forming for Improved Programming Yield and Cell-by-cell Write Verification SchemeabstractResistive RAM(RRAM) has emerged as a promising candidate for the next generation non-volatile memories (NVMs) due to its low write voltage, compact area, and CMOS compatibility. In this work, we propose a 1Mb RRAM macro with bipolar forming to reduce the forming voltage and improve the programming yield. Additionally, a cell-by-cell write verification scheme is introduced to protect RRAM cells from overstress and improve RRAM yield. The test chip, fabricated using 40nm CMOS technology, occupies a core area of 2.34 mm2. Byung-Kwon An, Junjie Mu, Putu Andhita Dananjaya, Weng Hong Lai, Wen Siang Lew, Tony Tae-Hyoung Kim |
ISCAS | 2 |
| 2025 | A Graph-Based Accelerator of Retinex Model With Bit-Serial Computing for Image EnhancementsabstractThis work proposes the Poisson equation formulation of the Retinex model for image enhancements using a low-power graph hardware accelerator performing finite difference updates on a lattice graph processing element (PE) array. By encapsulating the underlying algorithm in a graph hardware structure, a highly localized dataflow that takes advantage of the physical placement of the PEs is enabled to minimize data movement and maximize data reuse. The on-chip dataflow that achieves data sharing, and reuse among neighboring PEs during massively parallel updates is generated in each PE driven by two external control signals. Using a custom accumulator design intended for bit-serial computing, this work enables precision on demand and extensive on-chip data reuse with minimal area overhead, accommodating a non-overlap image mapping scheme in which a$20\times 20$image tile can be processed without external memory access at a time. With increasing user-configurable update count, image noise and shadow can be progressively removed with the inevitable loss of image details. Fabricated using a 65nm technology, the test chip occupies 0.2955mm2 core area and consumes 2.191mW operating at 1V, 25.6MHz, and a reconfigurable 10- or 14-bit precision. Zhengzhe Wei, Junjie Mu, Yuanjin Zheng, Tony Tae-Hyoung Kim, Bongjin Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | A Dual 7T SRAM-Based Zero-Skipping Compute- In-Memory Macro With 1-6b Binary Searching ADCs for Processing Quantized Neural NetworksabstractThis article presents a novel dual 7T static random-access memory (SRAM)-based compute-in-memory (CIM) macro for processing quantized neural networks. The proposed SRAM-based CIM macro decouples read/write operations and employs a zero-input/weight skipping scheme. A 65nm test chip with$528\times 128$integrated dual 7T bitcells demonstrated reconfigurable precision multiply and accumulate operations with$384\times $binary inputs (0/1) and$384\times 128$programmable multi-bit weights (3/7/15-levels). Each column comprises$384\times $bitcells for a dot product,$48\times $bitcells for offset calibration, and$96\times $bitcells for binary-searching analog-to-digital conversion. The analog-to-digital converter (ADC) converts a voltage difference between two read bitlines (i.e., an analog dot-product result) to a 1-6b digital output code using binary searching in 1-6 conversion cycles using replica bitcells. The test chip with 66Kb embedded dual SRAM bitcells was evaluated for processing neural networks, including the MNIST image classifications using a multi-layer perceptron (MLP) model with its layer configuration of 784-256-256-256-10. The measured classification accuracies are 97.62%, 97.65%, and 97.72% for the 3, 7, and 15 level weights, respectively. The accuracy degradations are only 0.58 to 0.74% off the baseline with software simulations. For the VGG6 model using the CIFAR-10 image dataset, the accuracies are 88.59%, 88.21%, and 89.07% for the 3, 7, and 15 level weights, with degradations of only 0.6 to 1.32% off the software baseline. The measured energy efficiencies are 258.5, 67.9, and 23.9 TOPS/W for the 3, 7, and 15 level weights, respectively, measured at 0.45/0.8V supplies. Chengshuo Yu, Haoge Jiang, Junjie Mu, Kevin Tshun Chuan Chai, Tony Tae-Hyoung Kim, Bongjin Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | A Graph-Based Accelerator of Retinex Model with Bit-Serial Computing for Image ProcessingabstractThis work implements the Poisson equation formulation of the Retinex model for image enhancements using a graph hardware accelerator performing finite difference updates on a 2D lattice graph PE array. A single clock gating control signal manages the data flow, data sharing, and reuse pattern among neighboring PEs during massively parallel updates. With increasing user-configurable update count, image noise and shadow can be progressively removed with the inevitable loss of image details. Accommodating a non-overlap image mapping scheme in which a$20\times 20$image tile can be processed without external memory access at a time, the proposed accelerator consists of 18$\times 18$regular PEs surrounded by$4\times 20$boundary PEs with reconfigurable data flow and 4 boundary cache registers. Fabricated using a 65nm technology, the test chip occupies 0.2955mm2core area, and consumes 2.191mW operating at 1V, 25.6MHz, and a reconfigurable 10- or 14-bit precision. Zhengzhe Wei, Junjie Mu, Zhongzhiguang Lu, Yuanjin Zheng, Tony Tae-Hyoung Kim, Bongjin Kim |
ISCAS | 2 |
| 2023 | 282-to-607 TOPS/W, 7T-SRAM Based CiM with Reconfigurable Column SAR ADC for Neural Network ProcessingabstractCompute in memory ($C$iM) is a promising solution for solving the bottleneck of frequent data interface between memory and processor in Von-Neumann architecture. In this work, a hybrid current/charge domain 7T-SRAM based CiM architecture is proposed to mitigate the PVT-induced RBL variation during computation and thus offer a better linearity without significant impact on the operating frequency and area efficiency. Additionally, a column-referenced 1b to 5b reconfigurable SAR ADC is proposed to support multi-bit output. The proposed design is verified by the Monte-Carlo simulations using 40nm CMOS technology. The 5b mode ADC transferred MAC curve's DNL (LSB) ranges from −0.025 to 0.02 and INL (LSB) ranges from −0.13 to 0.25. The largest RBL variation$(\sigma)$from MAC value −64 to MAC value +64 is 2.08 mV, resulting in a MNIST classification accuracy of 97.5%, which is only 0.1% degradation and Google Speech Command classification accuracy of 80.5%, which is only 0.5% degradation compared to the software baseline, respectively. The whole architecture offers energy efficiency of 282-to-607 TOPS/W for 1-5b output in the MAC operation, which is competitive when compared to other state-of-art$C$iM architectures. Qibang Zang, Wang Ling Goh, Lu Lu 0013, Chengshuo Yu, Junjie Mu, Tony Tae-Hyoung Kim, Bongjin Kim, Dongrui Li, Anh-Tuan Do |
ISCAS | 5 |
| 2023 | BP-SCIM: A Reconfigurable 8T SRAM Macro for Bit-Parallel Searching and Computing In-MemoryabstractThis work presents BP-SCIM: a reconfigurable 8T static random access memory (SRAM) macro for bit-parallel searching and computing in-memory (CIM). The decoupled read/write ports of the employed 8T SRAM bit-cell eliminate read disturbance during search and CIM operations. BP-SCIM can support both in-memory Boolean logic and arithmetic operations. Novel CIM-friendly algorithms and peripheral circuits are proposed to reduce the latency of complex arithmetic operations such as multiplication and division. In addition, BP-SCIM can be configured as either a binary content-addressable memory (CAM) or a ternary CAM for fast searching. A$256\times64$BP-SCIM test chip was implemented in 65-nm CMOS technology. The 8-bit addition and 8-bit multiplication operations can achieve the maximum energy efficiency of 3.11 TOPS/W and 0.17 TOPS/W, respectively at 0.7 V supply. For the binary CAM search operation, BP-SCIM can achieve the minimum energy consumption of 0.91 fJ/bit/search at 87 MHz and 0.8 V supply. Yuzong Chen 0001, Junjie Mu, Lu Lu 0013, Tony Tae-Hyoung Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | A 1-16b Reconfigurable 80Kb 7T SRAM-Based Digital Near-Memory Computing Macro for Processing Neural NetworksabstractThis work introduces a digital SRAM-based near-memory compute macro for DNN inference, improving on-chip weight memory capacity and area efficiency compared to state-of-the-art digital computing-in-memory (CIM) macros. A$20\times 256.1$-16b reconfigurable digital computing near-memory (NM) macro is proposed, supporting a reconfigurable 1-16b precision through the bit-serial computing scheme and the weight and input gating architecture for sparsity-aware operations. Each reconfigurable column MAC comprises$16\times $custom-designed 7T SRAM bitcells to store 1-16b weights, a conventional 6T SRAM for zero weight skip control, a bitwise multiplier, and a full adder with a register for partial-sum accumulations.$20\times $parallel partial-sum outputs are post-accumulated to generate a sub-partitioned output feature map, which will be concatenated to produce the final convolution result. Besides, pipelined array structure improves the throughput of the proposed macro. The proposed near-memory computing macro implements an 80Kb binary weight storage in a 0.473mm2 die area using 65nm. It presents the area/energy efficiency of 4329-270.6 GOPS/mm2 and 315.07-1.23TOPS/W at 1-16b precision. Junjie Mu, Chengshuo Yu, Tony Tae-Hyoung Kim, Bongjin Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | A Reconfigurable 8T SRAM Macro for Bit-Parallel Searching and Computing In-MemoryabstractThis work presents BP-SCIM: a reconfigurable 8T SRAM macro for bit-parallel searching and computing in-memory (CIM). BP-SCIM can perform in-memory Boolean logic, arithmetic, and content-addressable memory (CAM) operations. Novel peripheral circuits and algorithms are proposed to support complex arithmetic operations such as multiplication and division. A $256\times 64$ BP-SCIM test chip was implemented in 65-nm CMOS technology. The 8-bit addition and 8-bit multiplication operations can achieve the best energy efficiency of 3.11 TOPS/W and 0.17 TOPS/W, respectively at 0.7 V supply. For the binary CAM search operation, BP-SCIM can achieve the minimum energy consumption of 0.91 fJ/bit/search at 87 MHz and 0.8 V supply. Yuzong Chen 0001, Junjie Mu, Lu Lu 0013, Tony Tae-Hyoung Kim |
ISCAS | 2 |
| 2022 | SRAM-Based In-Memory Computing Macro Featuring Voltage-Mode Accumulator and Row-by-Row ADC for Processing Neural NetworksabstractThis paper presents a mixed-signal SRAM-based in-memory computing (IMC) macro for processing binarized neural networks. The IMC macro consists of$128\times 128$(16K) SRAM-based bitcells. Each bitcell consists of a standard 6T SRAM bitcell, an XNOR-based binary multiplier, and a pseudo-differential voltage-mode driver (i.e., an accumulator unit). Multiply-and-accumulate (MAC) operations between 64 pairs of inputs and weights (stored in the first 64 SRAM bitcells) are performed in 128 rows of the macro, all in parallel. A weight-stationary architecture, which minimizes off-chip memory accesses, effectively reduces energy-hungry data communications. A row-by-row analog-to-digital converter (ADC) based on 32 replica bitcells and a sense amplifier reduces the ADC area overhead and compensates for nonlinearity and variation. The ADC converts the MAC result from each row to an N-bit digital output taking 2N-1 cycles per conversion by sweeping the reference level of 32 replica bitcells. The remaining 32 replica bitcells in the row are utilized for offset calibration. In addition, this paper presents a pseudo-differential voltage-mode accumulator to address issues in the current-mode or single-ended voltage-mode accumulator. A test chip including a 16Kbit SRAM IMC bitcell array is fabricated using a 65nm CMOS technology. The measured energy- and area-efficiency is 741-87TOPS/W with 1-5bit ADC at 0.5V supply and 3.97TOPS/mm2, respectively. Junjie Mu, Bongjin Kim |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2020 | A 65nm Logic-Compatible Embedded and Flash Memory for In-Memory Computation of Artificial Neural NetworksabstractIn-memory computing using zero standby power nonvolatile memory is an attractive candidate for processing massively-parallel dot-products in artificial neural networks with high energy efficiency. A logic-compatible embedded flash (eFlash) is one of such nonvolatile memories. It has several advantages over the other candidates for in-memory computing: (1) programmable multi-level weight storage, (2) low power consumption with zero standby current, (3) low-cost fabrication using logic-process. Prior work proposed a 5T NOR-type eFlash cell for processing dot-products that are essential for neuromorphic computing. Despite the reliable dot-product computation using its carefully-designed program-and-verify sequence, the low memory density due to its 5T bitcell structure is one of the remaining challenges. In this work, we propose an AND eFlash cell structure for in-memory computation of artificial neural networks (ANNs) using multiple eFlash bitcells and shared access transistors. The proposed AND eFlash cell array is used for computing dot-products between binary inputs and reconfigurable bit-precision weights. The bitcell area has been reduced by 15% for the triple-cell structure compared to the prior 5T NOR eFlash. A multi-cycle program-and-verify operation is used to calibrate and improve the linearity. Junjie Mu, Bongjin Kim |
ISCAS | 1 |