EDBT 2026 Demo / reviewers in the wild / expert
Yi Li 0049
dblp:59/871-49
· DBLP profile ↗
15ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0003-1930-8550ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MemSearch: An Efficient Memristive In-memory Search Engine with Configurable Similarity Measures
Yingjie Yu, Houji Zhou, Jiancong Li, Tong Hu, Jia Chen 0032, Yi Li 0049, Xiangshui Miao |
ASP-DAC | 6 |
| 2026 | OAH-CIM: Outlier-Aware Hybrid RRAM-SRAM CIM Accelerator with Variation-Robust Sparsity
Tong Hu, Han Bao 0013, Houji Zhou, Yuyang Fu, Jiancong Li, Jia Chen 0032, Yi Li 0049, Xiangshui Miao |
ASP-DAC | 8 |
| 2026 | A Scaling Annealing Method for Combinatorial Optimization in Asynchronous Memristive Hopfield Network
Han Bao 0013, Kehong Xu, Yibai Xue, Yuyang Fu, Jiancong Li, Jia Chen 0032, Yi Li 0049, Xiangshui Miao |
ISCAS | 8 |
| 2026 | Energy-Efficient Acceleration of Fourier-based Transformers on RRAM-CIM via Mixed-Precision and DFT Symmetry
Yuyang Fu, Jiancong Li, Yi Li 0049, Xiangshui Miao |
ISCAS | 4 |
| 2025 | ReSMiPS: A ReRAM-based Sparse Mixed-precision Solver with Fast Matrix Reordering AlgorithmabstractThe solution of sparse matrix equations is essential in scientific computing. However, traditional solvers on digital computing platforms are limited by memory bottlenecks in largescale sparse matrix storage and computation. Resistive Random Access Memory (ReRAM)-based computing-in-memory (CIM) offers a promising solution to this challenge but faces constraints in achieving high solution precision and energy efficiency simultaneously in sparse matrix computations. In this work, we propose ReSMiPS, a ReRAM-accelerated sparse mixed-precision solver. ReSMiPS incorporates a novel Fast Sparse Matrix Reordering (FSMR) algorithm and introduces an In-memory float64 (IF64) data format, enabling efficient floating-point sparse matrix computation directly within the analog ReRAM array. By combining our floating-point CIM macro design with a hybrid-domain solution framework, ReSMiPS achieves precision comparable to CPU and GPU-based BiCGSTAB solvers (with errors below $10^{-15}$) on real-world workloads, while demonstrating over two orders of magnitude improvement in both computational speed and energy efficiency. Yuyang Fu, Jiancong Li, Jia Chen 0032, Houji Zhou, Wenlong Peng, Yi Li 0049, Xiangshui Miao |
DAC | 7 |
| 2025 | Efficient Edge Vision Transformer Accelerator with Decoupled Chunk Attention and Hybrid Computing-In-MemoryabstractVision Transformers (ViTs) are new foundation models for vision applications. Edge-deploying ViTs to realize energy-saving, low-latency, and high-performance dense predictions have wide applications, such as autonomous driving and surveillance image analysis. However, the quadratic complexity of the self-attention mechanism renders ViTs slow and resource-intensive, particularly for pixel-level dense predictions that involve long contexts. Additionally, the pyramid-like architecture of modern ViT variants leads to an unbalanced workload, further reducing hardware utilization and decreasing the throughput of conventional edge devices. To this end, we propose an algorithm-hardware co-optimized edge ViT accelerator tailored for efficient dense predictions. At the algorithm level, we propose a decoupled chunk attention (DCA) mechanism implemented in a pipelined manner to reduce off-chip memory access, thereby enabling efficient dense predictions within limited on-chip memory. At the architecture level, we introduce a hybrid architecture that combines SRAM-based computing-in-memory (CIM) and nonvolatile RRAM storage to eliminate extensive off-chip memory access, with a fusion scheduling to balance workloads and minimize intermediate on-chip memory access. At the circuit level, a bit/element two-way-reconfigurable CIM macro is proposed to improve hardware utilization across pyramidal ViT blocks with varied matrix sizes. The experimental results on object detection, semantic segmentation, and depth estimation tasks demonstrate that our design can efficiently process patch lengths up to 16384 with a speedup of 18.5×-217.1×, a reduction in memory accesses of 1.7×-7.4×, and an improvement in energy efficiency of 1.8×, under less than 1% performance degradation. Yi Li 0049, Zijian Ye, Xiangqu Fu, Songqi Wang, Shucheng Du, Ning Lin, Dashan Shang, Jinshan Yue, Xiaojuan Qi 0001, Feng Zhang 0014 |
DAC | 1 |
| 2025 | Guarder: A Stable and Lightweight Reconfigurable RRAM-based PIM Accelerator for DNN IP ProtectionabstractDeploying deep neural networks (DNNs) on conventional digital edge devices faces significant challenges due to high energy consumption. A promising solution is the processing-inmemory (PIM) architecture with resistive random-access memory (RRAM), but RRAM-based systems suffer from imprecise weights due to programming stochasticity and cannot effectively utilize conventional weight encryption/decryption intellectual property (IP) protection schemes. To address these issues, we propose a novel software-hardware co-design Guarder. On the hardware side, we introduce 3T2R cells to achieve reliable multiply-accumulate (MAC) operations and use reconfigurable inverter operating voltages to encode keys for encrypting DNNs on RRAM. On the software side, we implement a contrastive training method that ensures high model accuracy on authorized chips while degrading performance on unauthorized ones. This approach protects DNN IP with minimal hardware overhead while significantly mitigating the effects of RRAM programming stochasticity. Extensive experiments on tasks such as image classification (using MLP, ResNet, and ViT), segmentation (using SegFormer), and image generation (using DiT) validate the effectiveness of our method. The proposed contrastive training ensures negligible performance degradation on authorized chips, while performance on unauthorized chips drops to random guessing or generation. Compared to traditional RRAM accelerators, the 3T2R-based accelerator achieves a $1.41 \times$ reduction in area overhead and a $2.28 \times$ reduction in energy consumption. Ning Lin, Yi Li 0049, Jiankun Li, Jichang Yang, Yangu He, Yukui Luo, Dashan Shang, Xiaoming Chen 0003, Xiaojuan Qi 0001 |
DAC | 2 |
| 2025 | 3D self-rectifying memristive ternary content addressable memory for massive and exact in-memory search
Yingjie Yu, Shengguang Ren, Yi Li 0049, Xiangshui Miao |
Sci. China Inf. Sci. | 4 |
| 2025 | ArPCIM: An Arbitrary-Precision Analog Computing-in-Memory Accelerator With Unified INT/FP ArithmeticabstractAnalog Computing-in-memory (ACIM) breaks the von Neumann bottleneck and significantly improves energy efficiency by enabling parallel matrix-vector-multiplication (MVM) operations. However, most existing ACIM accelerators are highly customized for specific precision formats, lacking the generality to support efficiently arbitrary precision in both integer (INT) and floating-point (FP) formats. In this work, we present an arbitrary precision analog computing-in-memory (ArPCIM) accelerator with unified INT/FP arithmetic to address this limitation. We introduce a CIM-friendly INT/FP arithmetic to convert FP numbers into INT numbers for efficient execution on CIM, minimizing precision loss through local pre-alignment and dynamic bit-weight slicing methods. In addition, we implement multi-level reconfigurable precision circuits, featuring both intra- and inter-processing element (PE) reconfigurability, which supports precision ranging from 1-bit to 47-bit. Experimental results show that our ArPCIM accelerator achieves up to$4.09\times $and$5.37\times $improvement in energy efficiency and area efficiency, respectively, compared with state-of-the-art arbitrary precision digital CIM. Our ArPCIM accelerator offers the flexibility to meet diverse computational needs while maintaining high energy efficiency and accuracy, paving the way for versatile CIM acceleration across various fields. Jiancong Li, Han Jia, Houji Zhou, Han Bao 0013, Yuyang Fu, Yi Li 0049, Xiangshui Miao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2024 | LSMR: Synergy Randomness in Liquid State Machine and RRAM-based Analog-digital AcceleratorabstractBio-inspired event sensors are gaining popularity at the edge, such as in robots and wearable electronics. This trend necessitates learning vast amounts of sensory data on the edge, often in few-shot or even zero-shot scenarios, posing challenges in both software and hardware. This paper presents a novel software-hardware co-design to address these issues. Software-wise, we develop an SNN-ANN model, where the SNN encoder is a liquid state machine (LSM) that naturally processes events and significantly reduces learning complexity at the edge due to fixed random weights. The lightweight trainable ANN projection heads are optimized through contrastive learning, enabling zero-shot learning of multimodal events. Hardware-wise, we propose a hybrid analog (RRAM)-digital (CMOS) accelerator - LSMR. The analog in-memory computing core physically implements the LSM by leveraging RRAM stochasticity to generate fixed random weights. The digital core utilizes innovative reconfigurable systolic arrays to accelerate the contrastive learning of ANN projection heads. Extensive experimental outcomes from six neuromorphic datasets, encompassing visual, tactile, and auditory modalities, demonstrate that LSMR considerably improves energy efficiency by a range of 1.65× to 23.70×, in comparison to state-of-the-art edge devices. Simultaneously, it reduces training complexity by a range of 152.83× to 20,587.77× across various edge learning tasks. Ning Lin, Songqi Wang, Xinyuan Zhang 0008, Shaocong Wang 0001, Yangu He, Woyu Zhang, Bo Wang 0153, Jiankun Li, Mingzi Li, Binbin Cui, Yi Li 0049, Jia Chen 0032, Chunwei Xia, Xiaoming Chen 0003, Dashan Shang |
ICCAD | 11 |
| 2024 | Energy Efficient Memristive Transiently Chaotic Neural Network for Combinatorial OptimizationabstractThe utilization of memristive analog-digital mixed in-memory computing has significantly tackled the issues of massive computing resources and time delays in solving combinatorial optimization problems. However, further improvements in computing energy efficiency are still desirable for resource-constrained conditions and practical applications. Therefore, in this work, a memristive analog transiently chaotic neural network (TCNN) system is proposed to solve the traveling salesman problems (TSPs), which is consisted of 1) a memristor array to perform matrix-vector multiplication for the network iteration; 2) neuronal modules and nonlinear activation function modules to emulate the basic functions of the network; 3) chaotic simulated annealing (CSA) modules to improve the solution performance at extremely low hardware overhead. Based on these, after mapping the TSPs onto the memristor array, the proposed TCNN can self-iterate to convergence states to solve the problems with high performance. Compared to prior analog-digital mixed ones, the analog system can eliminate the extra control and data conversions during the network solving process, and achieve an$8\times $reduction in the total energy consumption in consecutive solving tasks. Han Bao 0013, Kehong Xu, Houji Zhou, Jiancong Li, Yi Li 0049, Xiangshui Miao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2023 | A memristive neural network based matrix equation solver with high versatility and high energy efficiency
Jiancong Li, Houji Zhou, Yi Li 0049, Xiangshui Miao |
Sci. China Inf. Sci. | 3 |
| 2022 | Low-time-complexity document clustering using memristive dot product engine
Houji Zhou, Yi Li 0049, Xiangshui Miao |
Sci. China Inf. Sci. | 2 |
| 2022 | Complementary Memtransistor-Based Multilayer Neural Networks for Online Supervised Learning Through (Anti-)Spike-Timing-Dependent PlasticityabstractWe propose a complete hardware-based architecture of multilayer neural networks (MNNs), including electronic synapses, neurons, and periphery circuitry to implement supervised learning (SL) algorithm of extended remote supervised method (ReSuMe). In this system, complementary (a pair of n- and p-type) memtransistors (C-MTs) are used as an electrical synapse. By applying the learning rule of spike-timing-dependent plasticity (STDP) to the memtransistor connecting presynaptic neuron to the output one whereas the contrary anti-STDP rule to the other memtransistor connecting presynaptic neuron to the teacher one, extended ReSuMe with multiple layers is realized without the usage of those complicated supervising modules in previous approaches. In this way, both the C-MT-based chip area and power consumption of the learning circuit for weight updating operation are drastically decreased comparing with the conventional single memtransistor (S-MT)-based designs. Two typical benchmarks, the linearly nonseparable benchmark XOR problem and Mixed National Institute of Standards and Technology database (MNIST) recognition have been successfully tackled using the proposed MNN system while impact of the nonideal factors of realistic devices has been evaluated. Nuo Xu 0002, Bin Gao 0006, Fuwei Zhuge, Zijian Tang, Xinchen Deng, Yi Li 0049, Yuhui He, Xiangshui Miao |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2018 | Memcomputing: fusion of memory and computing
Yi Li 0049, Yaxiong Zhou, Zhuorui Wang, Xiangshui Miao |
Sci. China Inf. Sci. | 1 |