VLDB 2026 Research / reviewers in the wild / expert
Jinyu Bai
dblp:224/5687
· DBLP profile ↗
15ranked-venue papers
3as first author
13since 2021 · last 2025
0000-0001-9369-0327ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 2 first-author · 11 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A 24.65 TOPS/W@INT8 Hybrid Analog-Digital Multi-core SRAM CIM Macro with Optimal Weight Dividing and Resource Allocation StrategiesabstractCompute-in-memory (CIM) technology integrates memory and computation to reduce memory bottlenecks in modern systems. However, current CIM architectures face challenges in balancing accuracy and energy efficiency. Analog-CIM (ACIM) is energy-efficient but less accurate, while Digital-CIM (DCIM) is accurate but consumes more energy. In this paper, we propose a novel multi-core hybrid analog-digital CIM macro that effectively addresses this trade-off. Our approach intelligently allocates computation tasks to ACIM and DCIM cores based on their accuracy requirements, achieving a balance of accuracy and efficiency. Additionally, we developed an optimization framework to determine the optimal weight divide ratio and computing resource allocation for the hybrid CIM. Experimental results demonstrate the efficacy of our approach. The proposed hybrid CIM achieves an outstanding energy efficiency of 24.65 TOPS/W at 8-bit precision, surpassing DCIM by a factor of 1.33 while maintaining a low error rate of only 0.4%, which is 30 times better than ACIM at the same precision. Wente Yi, Sifan Sun, Wenjia Wang 0011, Jinyu Bai, He Zhang 0011, Wang Kang 0001 |
ASP-DAC | 5 |
| 2025 | Model quantization for computing-in-memory: a survey
Sifan Sun, Jinyu Bai, Hanting Chen, Kaiwen Deng, Zhiwei Xie 0009, He Zhang 0011, Wang Kang 0001, Weisheng Zhao 0001 |
Sci. China Inf. Sci. | 2 |
| 2025 | Design and implementation of a real-time detection system for multi-token sandwich attacks in Ethereum based on Geth clientabstractThe Ethereum platform is booming with growing richness and variety in decentralized finance (DeFi) products. However, this progress comes with sophisticated threats, such as sandwich attacks, where attackers exploit the openness and certainty of blockchain technology to manipulate market prices and secure illegal financial rewards through a strategically planned series of transactions. The existing sandwich attack detection methods are ineffective at detecting multi-token transactions and fail to identify multi-token sandwich attacks. To tackle this challenge, this study improves the original detector’s algorithm to identify both traditional single-token and multi-token sandwich attacks. The enhanced system is not only responsive and accurate but also capable of detecting and alerting potential multi-token sandwich attacks. It has been successfully integrated with the go-Ethereum client (Geth). The system is performance-optimized with an average processing time of 0.81 seconds per block and an accuracy rate of 96.17%. The response time for detecting new blocks in real-time is usually no more than 4 seconds, with most between 2 and 3 seconds, which meets practical application requirements. By carefully analyzing the transaction data flow, this system is not only able to identify the traditional front-running attack and sandwich attack, but also extends to multi-currency complex attack strategies. The core innovation lies in the system’s ability to accurately detect and provide early warnings of multi-token sandwich attacks through real-time analysis of in-block transactions, all while maintaining the overall operational efficiency of the node. Jinyu Bai, Zhenxuan Jiang, Gang Du |
Discov. Comput. | 1 |
| 2024 | Series-Parallel Hybrid SOT-MRAM Computing-in-Memory Macro with Multi-Method Modulation for High Area and Energy EfficiencyabstractComputing-in-memory (CIM) shows its superiority in lots of applications like neural network inference. Recently, there are lots of exploration of the application of Magnetic Random-Access Memory (MRAM) in CIM. This paper aims to investigate the potential of Spin-Orbit-Torque-MRAM (SOT-MRAM) in CIM and proposes a high area and energy efficiency SOT-MRAM CIM macro based on a 6T-4J weight group. The bit-cell array adopts series-parallel hybrid architecture, which combines both serial and parallel configurations of Magnetic Tunnel Junction (MTJ) to solve the problem of high energy cost and low flexibility caused by MRAM-series and MRAM-parallel architecture, respectively. Additionally, the proposed SOT-MRAM CIM macro incorporates a multi-method modulation scheme, ranging from input unit to array, which meanwhile allows for configurable input precision (2/4/6/8-bit). The SOT-MRAM CIM macro is designed and verified in both 180nm and 28nm nodes, based on the verified electrical performance of the SOT-MRAM array in a 200-nm wafer pre-fabricated. The simulation results in 28nm show that this macro can achieve energy efficiency of 23.7~29.6 Tops/W at 8-bit input and output precision. Weiliang Huang, Jinyu Bai, Wang Kang 0001, Zhaohao Wang, Kaihua Cao, He Zhang 0011, Weisheng Zhao 0001 |
DAC | 2 |
| 2024 | MixMixQ: Quantization with Mixed Bit-Sparsity and Mixed Bit-Width for CIM AcceleratorsabstractQuantization is vital for deploying neural networks on Computing-In-Memory (CIM) based accelerators due to inherent limitations in memory devices and data interfaces’ representational capacities. However, traditional quantization algorithms often overlook CIM’s unique computing paradigm, leading to suboptimal performance. To address this, we introduce MixMixQ, a novel quantization algorithm specifically designed for CIM accelerators that strategically integrates mixed bit-sparsity and mixed bit-width, enhancing overall hardware efficiency while preserving high accuracy. Notably, our method can enhance hardware efficiency by up to 294% compared to traditional quantization methods, with only a minimal 0.13% decrease in accuracy compared to a full-precision network. Jinyu Bai, He Zhang 0011, Wang Kang 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2024 | CIMQ: A Hardware-Efficient Quantization Framework for Computing-In-Memory-Based Neural Network AcceleratorsabstractThe novel computing-in-memory (CIM) technology has demonstrated significant potential in enhancing the performance and efficiency of convolutional neural networks (CNNs). However, due to the low precision of memory devices and data interfaces, an additional quantization step is necessary. Conventional NN quantization methods fail to account for the hardware characteristics of CIM, resulting in inferior system performance and efficiency. This article proposes CIMQ, a hardware-efficient quantization framework designed to improve the efficiency of CIM-based NN accelerators. The holistic framework focuses on the fundamental computing elements in CIM hardware: inputs, weights, and outputs (or activations, weights, and partial sums in NNs) with four innovative techniques. First, bit-level sparsity induced activation quantization is introduced to decrease dynamic computation energy. Second, inspired by the unique computation paradigm of CIM, an innovative arraywise quantization granularity is proposed for weight quantization. Third, partial sums are quantized with a reparametrized clipping function to reduce the required resolution of analog-to-digital converters (ADCs). Finally, to improve the accuracy of quantized neural networks (QNNs), the post-training quantization (PTQ) is enhanced with a random quantization dropping strategy. The effectiveness of the proposed framework has been demonstrated through experimental results on various NNs and datasets (CIFAR10, CIFAR100, and ImageNet). In typical cases, the hardware efficiency can be improved up to 222% with a 58.97% improvement in accuracy compared to conventional quantization methods. Jinyu Bai, Sifan Sun, Weisheng Zhao 0001, Wang Kang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | CIM²PQ: An Arraywise and Hardware-Friendly Mixed Precision Quantization Method for Analog Computing-In-MemoryabstractComputing-in-memory (CIM) architecture is a promising convolutional neural network (CNN) accelerator known for its highly efficient matrix-vector multiplications (MVMs). However, due to the low-precision computation and limited size of CIM memory arrays, it is necessary to decompose the huge MVMs into smaller subsets. Conventional NN quantization methods overlook the characteristics of CIM hardware, resulting in diminished system performance and efficiency. This paper proposes a mixed precision quantization (MPQ) method based on evolutionary algorithm for CIM-based accelerators, while considering the hardware characteristics of CIM, called CIMPQ, which can automatically generate quantization strategies for NN model to improve the efficiency of CIM systems. Firstly, inspired by the CIM computing paradigm, an array-wise quantization granularity is introduced in the MPQ search space, which can jointly quantize the inputs, weights, and partial sums. Secondly, a production procedure containing fine-grained crossover and progressive adaptive mutation is proposed, which can efficiently explore the search space and speed up the search process. Thirdly, we propose a fast and efficient strategy evaluation method to obtain the performance of quantization strategy on the CIM platform, saving the evaluation time significantly without requiring fine-tuning. Finally, to protect CIM-friendly strategies with lower bit-widths but worse algorithm performance, we propose a strategy selection method based on multi-objective optimization, named qNSGA-III. The effectiveness of the proposed method has been demonstrated through experimental results of various NNs and datasets. For ResNet-18, the hardware efficiency and accuracy can be improved to 117% with 7.05%, 113% with 3.37%, and 119% with 5.78%, on CIFAR-10, CIFAR-100 and ImageNet, respectively, compared to the baseline MPQ method. Sifan Sun, Jinyu Bai, Zhaoyu Shi, Weisheng Zhao 0001, Wang Kang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | CiTST-AdderNets: Computing in Toggle Spin Torques MRAM for Energy-Efficient AdderNetsabstractRecently, Adder Neural Networks (AdderNets) have gained widespread attention as an alternative to traditional Convolutional Neural Networks (CNNs) for deep learning tasks. AdderNets use lightweight addition operations to replace multiplication and accumulation (MAC) operations, but can keep almost the same accuracy compared to other CNNs. Nevertheless, challenges still exist with regards to hardware resources, power consumption, and communication bandwidth, primarily due to the ‘Von-Neumann bottlenecks’. However, computing-in-memory (CIM) architecture based on magnetic random-access memory (MRAM) has great potential for edge DNN implementation. In this paper, we propose a novel CIM paradigm using a novel Toggle-Spin-Torques (TST) driven MRAM for energy-efficient AdderNets (called CiTST_AdderNets). In CiTST_AdderNets, MRAM is driven by the interplay of the field-free spin orbit torque (SOT) effect and the spin transfer torque (STT) effect, which offers a fascinating prospect for energy efficiency and speed. Furthermore, a novel CIM paradigm is proposed to implement the dominating subtraction and sum operations in AdderNets, reducing data transfer and the related energy. Meanwhile, a highly parallel array structure integrating computation and storage is designed to support CiTST_AdderNets. In addition, a mapping strategy is proposed to efficiently map the convolution layer on the array. Fully connected layers can also be efficiently computed. The CiTST-AdderNets macro is designed by using a 65-nm CMOS process. Results show that our CiTST-AdderNets consumes about 1.65 mJ, 9.29 mJ, and 42.46 mJ for running VGG8, ResNet-50, and ResNet-18 respectively at 8-bit fixed-point precision. Compared to state-of-the-art platforms, our macro achieves an energy efficiency improvement of 1.45 x to 66.78 x. Lichuan Luo, Erya Deng, Dijun Liu, Zhen Wang 0070, Weiliang Huang, He Zhang 0011, Xiao Liu 0051, Jinyu Bai, Junzhan Liu, Youguang Zhang, Wang Kang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2023 | Hierarchical Non-Structured Pruning for Computing-In-Memory Accelerators with Reduced ADC Resolution RequirementabstractThe crossbar architecture, which is comprised of novel nano-devices, enables high-speed and energy-efficient computing-in-memory (CIM) for neural networks. However, the overhead from analog-to-digital converters (ADCs) substantially degrades the energy efficiency of CIM accelerators. In this paper, we introduce a hierarchical non-structured pruning strategy where value-level and bit-level pruning are performed jointly on neural networks to reduce the resolution of ADCs by using the famous alternating direction method of multipliers (ADMM). To verify the effectiveness, we deployed the proposed method to a variety of state-of-the-art convolutional neural networks on two image classification benchmark datasets: CIFAR10, and ImageNet. The results show that our pruning method can reduce the required resolution of ADCs to 2 or 3 bits with only slight accuracy loss (~0.25 %), and thus can improve the hardware efficiency by 180%. Wenlu Xue, Jinyu Bai, Sifan Sun, Wang Kang 0001 |
DATE | 2 |
| 2022 | SpinCIM: spin orbit torque memory for ternary neural networks based on the computing-in-memory architecture
Lichuan Luo, Dijun Liu, He Zhang 0011, Youguang Zhang, Jinyu Bai, Wang Kang 0001 |
CCF Trans. High Perform. Comput. | 5 |
| 2022 | HD-CIM: Hybrid-Device Computing-In-Memory Structure Based on MRAM and SRAM to Reduce Weight Loading Energy of Neural NetworksabstractSRAM based computing-in-memory (SRAM-CIM) techniques have been widely studied for neural networks (NNs) to solve the “Von Neumann bottleneck”. However, as the scale of the NN model increasingly expands, the weight cannot be fully stored on-chip owing to the big device size (limited capacity) of SRAM. In this case, the NN weight data have to be frequently loaded from external memories, such as DRAM and Flash memory, which results in high energy consumption and low efficiency. In this paper, we propose a hybrid-device computing-in-memory (HD-CIM) architecture based on SRAM and MRAM (magnetic random-access memory). In our HD-CIM, the NN weight data are stored in on-chip MRAM and are loaded into SRAM-CIM core, significantly reducing energy and latency. Besides, in order to improve the data transfer efficiency between MRAM and SRAM, a high-speed pipelined MRAM readout structure is proposed to reduce the BL charging time. Our results show that the NN weight data loading energy in our design is only 0.242 pJ/bit, which is 289$\times $less in comparison with that from off-chip DRAM. Moreover, the energy breakdown and efficiency are analyzed based on different NN models, such as VGG19, ResNet18 and MobileNetV1. Our design can improve$\mathbf {58\times \,\,to\,\,124\times }$energy efficiency. He Zhang 0011, Junzhan Liu, Jinyu Bai, Lichuan Luo, Shaoqian Wei, Wang Kang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2021 | SpinLiM: Spin Orbit Torque Memory for Ternary Neural Networks Based on the Logic-in-Memory ArchitectureabstractLogic-in-memory architecture based on spintronic memories shows fascinating prospects in neural networks (NNs) for its high energy efficiency and good endurance. In this work, we leveraged two magnetic tunnel junctions (MTJs), which are driven by the interplay of field-free spin orbit torque (SOT) and spin transfer torque (STT) effects, to achieve a novel statefullogic-in-memory paradigm for ternary multiplication operations. Based on this paradigm, we further proposed a highly parallel array structure to serve for ternary neural networks (TNNs). Our results demonstrate the advantage of our design in power consumption compared with CPU, GPU and other state-of-the-art works. Lichuan Luo, He Zhang 0011, Jinyu Bai, Youguang Zhang, Wang Kang 0001, Weisheng Zhao 0001 |
DATE | 3 |
| 2021 | HSC: A Hybrid Spin/CMOS Logic Based In-Memory Engine with Area-Efficient Mapping StrategyabstractRecent advances in deep learning have shown that binary neural networks (BNNs) can provide a satisfying accuracy on various tasks with significant reduction in computation power and memory cost. Theoretically, the multiply-and-accumulate (MAC) operations of BNNs can be replaced by in-memory XNOR operations, thereby avoiding frequent data transfer between the buffer and the processor. However, devices supporting in-memory implementation of XNOR operations together with efficient weight-matrix mapping strategy is still an open research area. In this paper, a hybrid spin/CMOS cell (HSC) structure is proposed in which the XNOR operation can be simply realized in an in-memory computing manner by the non-volatile data from the spin component and the volatile data from the CMOS component. Given the time/spatial trade-off, a novel weight mapping method to break the large memory array and unroll the 3D kernel into 2D weight matrix is designed to cooperate with the proposed HSC structure in a time-division way. System-level simulation results show that the proposed BNN processor can achieve a 3.32* speedup and 11.9* improvement in throughput and energy efficiency, which could be attributed to the device and mapping method co-design. Erya Deng, Jinyu Bai, Wang Kang 0001, Biao Pan |
ISCAS | 3 |
| 2019 | Magnetic Skyrmion-Based Neural Recording System Design for Brain Machine InterfaceabstractNext-generation brain machine interface demand a high-channel-count neural recording system to wirelessly monitor activities of thousands of neurons. In order to achieve high-density neural recording, further development of single recording channel comprised of a neural amplifier front-end (AFE) and an analog-to-digit converter (ADC) is critical. Despite the great progress made in CMOS implementation of custom-designed neural recording system, hybrid limitations of increasing area and power consumption in line with Moore's law drove great demand for post-CMOS substitutes. Magnetic skyrmion with nano particle-like and non-volatile properties are of both fundamental and applied interests for future bio-inspired electronics. In this work, we propose a compact model including both AFE and ADC based on current-induced skyrmion motion. The proposed system achieved a power consumption of 0.63 pJ/channel with an area overhead of 0.14 μm2. The purpose of this work is to explore the feasibility of magnetic skyrmion for building large-scale, dense neuronal recording system which could pave a new way for future brain machine interface application. Biao Pan, Wang Kang 0001, Xing Chen 0012, Jinyu Bai, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 4 |
| 2019 | SR-WTA: Skyrmion Racing Winner-Takes-All Module for Spiking Neural ComputingabstractSpiking neural network (SNN) has emerged as one of the popular architectures in complex pattern recognition and classification tasks. However, hardware implementation of such algorithms using conventional CMOS based neuron consume resources and power that are orders of magnitude higher than that in human brain. This can be attributed to the mismatch of the computational architecture between biological brain and the current Boolean logic computing platform. Magnetic skyrmions have been intensively studied as a prospective information carrier in neuromorphic computing hardware design. In this work, a compact time-domain skyrmion-racing winner-takes-all (SR-WTA) leaky-integrate-fire (LIF) spiking neuron network is presented for the first time. The skyrmion motion dynamics in the LIF neuron and the behaviors of the neuron network was investigated comprehensively. Both SPICE and micromagnetic simulations are performed to evaluate the functionality and performance of the proposed SR-WTA based SNN. Biao Pan, Wang Kang 0001, Xing Chen 0012, Jinyu Bai, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 4 |