EDBT 2026 Demo / reviewers in the wild / expert
Can Li 0024
dblp:94/10021-24
· DBLP profile ↗
17ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0003-3795-2008ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 1 first-author · 12 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Robustness of Content-Addressable Memories for In-Memory Search
Liu Liu 0023, Tomas Sousa Pereira, Mohammad Mehdi Sharifi, Can Li 0024, Kai Ni 0004, Michael T. Niemier, Xiaobo Sharon Hu |
VTS | 6 |
| 2026 | EvaCAM: A Circuit-Level Evaluation Tool for General Content Addressable MemoriesabstractContent addressable memories (CAMs) are special-purpose in-memory computing units that support parallel searches directly in memory. There is growing interest in CAMs for data-intensive applications such as machine learning, data mining, and bioinformatics, which has led to a rapidly growing CAM design space. CAM cells can be implemented exclusively by CMOS or with various non-volatile memory (NVM) devices. In addition to traditional binary and ternary CAMs (BCAMs and TCAMs), analog CAM (ACAM) and multi-bit CAM (MCAM) designs have recently been introduced, which could further improve density, and also support unique in-memory distance functions. Furthermore, aside from the widely-used exact match function, CAM-based approximate match functions, such as threshold match and best match, have been proposed to further extend the utility of CAMs to new application spaces. As the CAM design space is large, evaluating different CAM design options for a given application is both crucial and challenging. This paper presents EvaCAM, a circuit-level modeling and evaluation tool for CAMs. EvaCAM supports TCAM, ACAM, and MCAM designs implemented in either CMOS or NVMs, for both exact and approximate match functions. It also allows for the exploration of different CAM designs under various optimization targets. EvaCAM has been validated against measured data from fabricated chips and detailed SPICE simulations. A comprehensive design space exploration for CAMs is provided to illustrate the impact of various design decisions and to demonstrate the use cases of EvaCAM. Liu Liu 0023, Mohammad Mehdi Sharifi, Kunshi Wang, Ruibin Mao, Kai Ni 0004, Can Li 0024, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2026 | WireLightning: Harnessing Capacitances for In-Transit Massively Parallel Matrix MultiplicationabstractAnalog computing-in-memory accelerators promise ultra-low-power, on-device AI by reducing data transfer and energy usage. Yet inherent device variations and high energy consumption for analog-digital conversion continue to hinder their wide-scale adoption in mainstream systems. To address these issues, this paper introduces WireLightning, a novel capacitive-computing accelerator featuring a mixed-signal architecture that rethinks analog AI acceleration. Unlike conventional analog crossbars that encode weights in programmable devices, WireLightning exploits intrinsic charge dynamics in passive capacitors, encoding matrix multiplication through spike amplitude and timing. This design addresses critical limitations such as weight drift, stochasticity, and power-intensive ADC bottlenecks. Key innovations include: amplitude-temporal dual encoding that enables constant-time analog dot-products; time-based decoding scheme that significantly reduces reliance on power-intensive ADCs; row-wise parallel architecture for concurrent dot-product calculations across multiple rows to enhance throughput; and value repetition exploitation in low-bit quantized vectors to reduce multiplications to constant time complexity. A PCB pro totype achieved a range-normalized RMSE of 1.80%—69.2% of the error in RRAM crossbar implementations—and a normalized error of 9.18%, corresponding to 77.1% of that in leading PCM crossbars. Implemented in a 40-nm CMOS technology, WireLightning macro delivers 465.47 TOPS/W at 4-bit precision, outperforming state-of-the-art analog accelerators while maintaining 0.23% range-normalized RMSE. By integrating algorithm-circuit co-design with physical computing, this work establishes capacitive computing as a promising path toward combining digital precision and analog efficiency in next-generation edge AI. Song Wang 0023, Zhu Wang 0014, Can Li 0024, Hayden Kwok-Hay So |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | Invited Paper: Circuit and Architecture Design with Emerging Computing ParadigmsabstractAs emerging computing paradigms push beyond the limitations of traditional CMOS-based computing using Von Neumann architectures, there is a growing need to rethink and extend Electronic Design Automation (EDA) methodologies to support their unique characteristics. These paradigms—including Approximate Computing, In-Memory Computing, Reconfigurable Field-Effect Transistors (RFETs), and Photonic Computing—represent diverse and promising directions beyond conventional digital design. Collectively, they offer transformative potential for achieving significant improvements in energy efficiency, computational speed, and architectural scalability. For example, application-specific approximate computing enables the design of custom arithmetic circuits that exploit application-level error resilience, allowing for optimized accuracy–power–performance–area (PPA) trade-offs in error-tolerant applications. Similarly, processing-in-non-volatile memories, such as those based on Ferroelectric Field-effect Transistors (FeFETs), enhances energy efficiency by enabling analog computation—particularly for operations like matrix multiplication—directly within the memory arrays. The intrinsic polymorphism of RFETs supports compact, multifunctional logic gates and introduces new opportunities for circuit-level obfuscation and security-aware design. Likewise, photonic analog wavefront computing offers substantial gains in latency and energy efficiency by encoding and processing information in the analog optical domain, leveraging phenomena such as diffraction and interference to perform computation at the speed of light. However, they also introduce a host of new challenges in circuit and architecture design, such as vast and irregular design spaces, analog and non-Boolean behavior, and new device-level constraints that existing EDA tools are not capable of handling. To this end, the current article focuses on the development of efficient and robust EDA frameworks that can enable the practical realization of circuits and architectures in these emerging domains. Salim Ullah, Siva Satyendra Sahoo, Can Li 0024, Chao Li 0065, Liu Liu 0023, Tomas Sousa Pereira, Xunzhao Yin, Armin Darjani, Nima Kavand, Chakravarthy Bodla, Rupa Yashaswi Panduga, Aniruddh Holemadlu, Johannes Maly, Jonathan Förste, Samarth Vadia, Xiaobo Sharon Hu, Akash Kumar 0001 |
ICCAD | 3 |
| 2025 | High-Performance In-Memory Bayesian Inference With Multi-Bit Ferroelectric FETabstractConventional neural network-based machine learning algorithms often encounter difficulties in data-limited scenarios or where interpretability is critical. Conversely, Bayesian inference-based models excel with reliable uncertainty estimates and explainable predictions. Recently, many in-memory computing (IMC) architectures achieve exceptional computing capacity and efficiency for neural network tasks leveraging emerging nonvolatile memory (NVM) technologies. However, their application in Bayesian inference remains limited because the operations in Bayesian inference differ substantially from those in neural networks. In this article, we introduce a compact in-memory Bayesian inference engine with high efficiency and performance utilizing a multi-bit ferroelectric field-effect transistor (FeFET). This design encodes a Bayesian model within a compact FeFETbased crossbar by mapping quantized probabilities to discrete FeFET states. Consequently, the crossbar’s outputs naturally represent the output posteriors of the Bayesian model. Our design facilitates efficient Bayesian inference, accommodating various input types and probability precisions, without additional calculation circuitry. As the first FeFET-based in-memory Bayesian inference engine, our design demonstrates a notable storage density of 26.32 Mb/mm2and a computing efficiency of 581.40 TOPS/W in a representative Bayesian classification task, indicating a 10.7×/43.4× compactness/efficiency improvement compared to the state-of-the-art alternative. Utilizing the proposed Bayesian inference engine, we develop a feature selection system that efficiently addresses a representative NP-hard optimization problem, showcasing our design’s capability and potential to enhance various Bayesian inference-based applications. Test results suggest that our design identifies the essential features, enhancing the model’s performance while reducing its complexity, surpassing the latest implementation in operation speed and algorithm efficiency by 2.9×/2.0×, respectively. Chao Li 0065, Xuchu Huang, Ruibin Mao, Thomas Kämpfe, Kai Ni 0004, Can Li 0024, Xunzhao Yin, Cheng Zhuo |
IEEE Trans. Computers | 9 |
| 2024 | Physics-Informed Learning for Versatile RRAM Reset and Retention SimulationabstractResistive random-access memory (RRAM) constitutes an emerging and promising platform for compute-inmemory (CIM) edge AI. However, the switching mechanism and controllability of RRAM are still under debate owing to the influence of multiphysics. Although physics-informed neural networks (PINNs) are successful in achieving mesh-free multiphysics solutions in many applications, the resultant accuracy is not satisfactory in RRAM analyses. This work investigates the characteristics of RRAM devices - retention and reset transition which are described in terms of the dissolution of a conductive filament (CF) in 3-D axis-symmetric geometry. Specifically, we provide a novel neural network characterization of ion migration, Joule heating, and carrier transport, governed by the solutions of partial differential equations (PDEs). Motivated by physics-informed learning, the separation of variables (SOV) method and the neural tangent kernel (NTK) theory, we propose a customized 3-channel fully-connected network and a modified random Fourier feature (mRFF) embedding strategy to capture multiscale properties and appropriate frequency features of the self-consistent multiphysics solutions. The proposed model eliminates the need for grid meshing and temporal iterations widely used in RRAM analysis. Experiments then confirm its superior accuracy over competing physics-informed methods. Tianshu Hou, Wenyong Zhou, Can Li 0024, Haibao Chen, Ngai Wong 0001 |
ASPDAC | 4 |
| 2024 | FeBiM: Efficient and Compact Bayesian Inference Engine Empowered with Ferroelectric In-Memory ComputingabstractIn scenarios with limited training data or where explainability is crucial, conventional neural network-based machine learning models often face challenges. In contrast, Bayesian inference-based algorithms excel in providing interpretable predictions and reliable uncertainty estimation in these scenarios. While many state-of-the-art in-memory computing (IMC) architectures leverage emerging non-volatile memory (NVM) technologies to offer unparalleled computing capacity and energy efficiency for neural network workloads, their application in Bayesian inference is limited. This is because the core operations in Bayesian inference, i.e., cumulative multiplications of prior and likelihood probabilities, differ significantly from the multiplication-accumulation (MAC) operations common in neural networks, rendering them generally unsuitable for direct implementation in most existing IMC designs. In this paper, we propose FeBiM, an efficient and compact Bayesian inference engine powered by multi-bit ferroelectric field-effect transistor (FeFET)-based IMC. FeBiM effectively encodes the trained probabilities of a Bayesian inference model within a compact FeFET-based crossbar. It maps quantized logarithmic probabilities to discrete FeFET states. As a result, the accumulated outputs of the crossbar naturally represent the posterior probabilities, i.e., the Bayesian inference model's output given a set of observations. This approach enables efficient in-memory Bayesian inference without the need for additional calculation circuitry. As the first FeFET-based in-memory Bayesian inference engine, FeBiM achieves an impressive storage density of 26.32 Mb/mm2 and a computing efficiency of 581.40 TOPS/W in a representative Bayesian classification task. These results demonstrate 10.7×/43.4× improvement in compactness/efficiency compared to the state-of-the-art hardware implementation of Bayesian inference. Chao Li 0065, Ruibin Mao, Can Li 0024, Thomas Kämpfe, Kai Ni 0004, Xunzhao Yin |
DAC | 5 |
| 2024 | FeReX: A Reconfigurable Design of Multi-Bit Ferroelectric Compute-in-Memory for Nearest Neighbor SearchabstractRapid advancements in artificial intelligence have given rise to transformative models, profoundly impacting our lives. These models demand massive volumes of data to operate effectively, exacerbating the data-transfer bottleneck inherent in the conventional von-Neumann architecture. Compute-in-memory (CIM), a novel computing paradigm, tackles these issues by seam-lessly embedding in-memory search functions, thereby obviating the need for data transfers. However, existing non-volatile memory (NVM)-based accelerators are application specific. During the similarity based associative search operation, they only support a single, specific distance metric, such as Hamming, Manhattan, or Euclidean distance in measuring the query against the stored data, calling for reconfigurable in-memory solutions adaptable to various applications. To overcome such a limitation, in this paper, we present FeReX, a reconfigurable associative memory (AM) that accommodates various distance metrics including Hamming, Manhattan, and Euclidean distances. Leveraging multi-bit ferroelectric field-effect transistors (FeFETs) as the proxy and a hardware-software co-design approach, we introduce a constrained satisfaction problem (CSP)-based method to automate AM search input voltage and stored voltage configurations for different distance based search functions. Device-circuit co-simulations first validate the effectiveness of the proposed FeReX methodology for reconfigurable search distance functions. Then, we benchmark FeReX in the context of k-nearest neighbor (KNN) and hyperdimensional computing (HDC), which highlights the robustness of FeReX and demonstrates up to 250× speedup and 104energy savings compared with GPU. Che-Kai Liu, Chao Li 0065, Ruibin Mao, Jianyi Yang 0003, Thomas Kämpfe, Mohsen Imani, Can Li 0024, Cheng Zhuo, Xunzhao Yin |
DATE | 8 |
| 2024 | ShiftCAM: A Time-Domain Content Addressable Memory Utilizing Shifted Hamming Distance for Robust Genome AnalysisabstractFast and efficient genome analysis can have a significant impact in areas such as scientific discovery and personalized medicine. Given the extensive data produced by sequencing machines, in-memory computing is considered a strong candidate to tackle the frequent data movement issue. Previous research has introduced many designs based on Content Addressable Memories (CAM), mainly optimized for tolerating edit distance; however, these systems struggle when there are a few insertion or deletion errors. This limitation presents a significant challenge for genome analysis, as current Third-Generation Sequencing still has high error rates. In this work, we introduce ShiftCAM, a time-domain Content Addressable Memory, designed to accommodate the high error rates in practical scenarios. Utilizing time-domain comparison, ShiftCAM effectively calculates the Shifted Hamming Distance to better approximate the computationally expensive edit distance. Additionally, the Modification to Accidental Match strategy specially designed for hardware implementation is introduced to eliminate accidental matches of single base pairs, further reducing false positives and improving edit distance approximation. Monte Carlo simulations based on physical ReRAM device statistical measurements and commercial PDK are also conducted to validate the robustness of the ShiftCAM design. Our experiments demonstrate that ShiftCAM can achieve an average of 2.1× (from 40.1% to 83.8%) higher F1 score in contamination analysis, 21.3% estimation error in relative abundance analysis, 51.2% reduction in cell area, 29.5× speed up, and 9.4× higher energy efficiency, compared to state-of-the-art in-memory DNA classification accelerators. Peiyi He, Ruibin Mao, Keyi Shan, Yunwei Tong, Muyuan Peng, Ruibang Luo, Can Li 0024 |
ICCAD | 8 |
| 2024 | Realizing In-Memory Baseband Processing for Ultrafast and Energy-Efficient 6GabstractTo support emerging applications ranging from holographic communications to extended reality, next-generation mobile wireless communication systems require ultrafast and energy-efficient baseband processors. Traditional complementary metal-oxide-semiconductor (CMOS)-based baseband processors face two challenges in transistor scaling and the von Neumann bottleneck. To address these challenges, in-memory computing-based baseband processors using resistive random-access memory (RRAM) present an attractive solution. In this article, we propose and demonstrate RRAM-implemented in-memory baseband processing for the widely adopted multiple-input–multiple-output orthogonal frequency division multiplexing (MIMO-OFDM) air interface. Its key feature is to execute the key operations, including discrete Fourier transform (DFT) and MIMO detection, using linear minimum mean square error (L-MMSE) and zero forcing (ZF), in one-step. In addition, RRAM-based channel estimation module is proposed and discussed. By prototyping and simulations, we demonstrate the feasibility of RRAM-based full-fledged communication system in hardware, and reveal it can outperform state-of-the-art baseband processors with a gain of$91.2\times $in latency and$671\times $in energy efficiency by large-scale simulations. Our results pave a potential pathway for RRAM-based in-memory computing to be implemented in the era of the sixth generation (6G) mobile communications. Qunsong Zeng, Mingrui Jiang, Yi Gong 0001, Yida Li 0004, Can Li 0024, Jim Ignowski, Kaibin Huang |
IEEE Internet Things J. | 8 |
| 2023 | ReRAM-based graph attention network with node-centric edge searching and hamming similarityabstractThe graph attention network (GAT) has demonstrated its advantages via local attention mechanism but suffered from low energy and latency efficiency when implemented on conventional von-Neumann hardware. This work proposes and experimentally demonstrates an algorithm-hardware co-designed GAT that runs efficiently and reliably in ReRAM-based hardware. The neighborhood information is retrieved from trained node embeddings stored on crossbars in a single time step, and attention is implemented by efficient hashing and hamming similarity for higher robustness. Our scaled simulation based on the experimentally-validated model shows only 0.9% accuracy loss with over 35,500x energy improvement on the Cora dataset compared with GPU, and 1.1% accuracy improvement with 2× energy improvement compared with state-of-the-art ReRAM-based GNN accelerator. Ruibin Mao, Xia Sheng, Catherine Graves, Can Li 0024 |
DAC | 5 |
| 2023 | Cross Layer Design for the Predictive Assessment of Technology-Enabled ArchitecturesabstractThere is great interest in “end-to-end” analysis that captures how innovation at the materials, device, and/or archi-tectural levels will impact figures of merit at the application-level. However, there are numerous combinations of devices and architectures to study, and we must establish systematic ways to accurately explore and cull a vast design space. We aim to capture how innovations at the materials/device-level may ultimately impact figures of merit associated with both existing and emerging technologies that may be employed for either logic and/or memory. We will highlight how collaborations with researchers at these levels of the design hierarchy - as well as efforts to help construct well-calibrated device models - can in-turn support architectural design space explorations that will help to identify the most promising ways to use new technologies to support application-level workloads of interest. For given compute workloads, we can then quantitatively assess the potential benefits of technology-driven architectures to identify the most promising paths forward. Because of the large number of potentially interesting device-architecture combinations, it is of the utmost importance to develop well-calibrated analytical modeling tools to more rapidly assess the potential value of a given (likely heterogeneous) solution. We highlight recent efforts and needs in this space. Michael T. Niemier, Xiaobo Sharon Hu, Liu Liu 0023, Mohammad Mehdi Sharifi, Ian O'Connor, David Atienza 0001, Giovanni Ansaloni, Can Li 0024, Daniel C. Ralph |
DATE | 8 |
| 2021 | Mixed Precision Quantization for ReRAM-based DNN Inference AcceleratorsabstractReRAM-based accelerators have shown great potential for accelerating DNN inference because ReRAM crossbars can perform analog matrix-vector multiplication operations with low latency and energy consumption. However, these crossbars require the use of ADCs which constitute a significant fraction of the cost of MVM operations. The overhead of ADCs can be mitigated via partial sum quantization. However, prior quantization flows for DNN inference accelerators do not consider partial sum quantization which is not highly relevant to traditional digital architectures. To address this issue, we propose a mixed precision quantization scheme for ReRAM-based DNN inference accelerators where weight quantization, input quantization, and partial sum quantization are jointly applied for each DNN layer. We also propose an automated quantization flow powered by deep reinforcement learning to search for the best quantization configuration in the large design space. Our evaluation shows that the proposed mixed precision quantization scheme and quantization flow reduce inference latency and energy consumption by up to 3.89x and 4.84x, respectively, while only losing 1.18% in DNN inference accuracy. Sitao Huang, Aayush Ankit, Plínio Silveira, Rodrigo Antunes, Sai Rahul Chalamalasetti, Izzat El Hajj, Dong Eun Kim, Glaucimar Aguiar, Pedro Bruel, Sergey Serebryakov, Can Li 0024, Paolo Faraboschi, John Paul Strachan, Deming Chen, Kaushik Roy 0001, Wen-Mei W. Hwu, Dejan S. Milojicic |
ASP-DAC | 12 |
| 2018 | Large Memristor Crossbars for Analog ComputingabstractMemristor with tunable non-volatile resistance offers in-memory computing capability that avoids the von-Neumann bottleneck. However, large-scale experimental demonstration to this end is yet to be implemented due to the immaturity of the device and integration technologies. Here in this paper we report our recent process in analog computing using analog-voltage-amplitude-vector input and analog-memristor-conductance matrix, with applications in signal and image processing. The vector matrix multiplication is processed in the memristor crossbars in one step, with 5-8 bit precision depending on the array size. The demonstration is made possible by high memristor yield (99.8%), stable multilevel memresistance states, linear current-voltage (IV) relation in the operation range, and low wire resistance between the cells. Can Li 0024, Yunning Li, Hao Jiang 0017, J. Joshua Yang, Qiangfei Xia, Miao Hu 0002, Eric Montgomery, Noraica Dávila, Catherine Graves, John Paul Strachan, R. Stanley Williams, Ning Ge 0001, Mark Barnell, Qing Wu 0002 |
ISCAS | 1 |
| 2018 | Unconventional computing with diffusive memristorsabstractDiffusive memristors with Ag active metal species are volatile threshold switches featuring spontaneous rupture of conduction channels at small electrical bias. The unique temporal dynamics of the conductance evolution originates from the underlying electrochemical and diffusive dynamics of the active metals in dielectrics, which can be explored for a variety of novel applications in unconventional computing. The superior I-V nonlinearity enables large crossbar arrays for high density non-volatile memories. The relaxation dynamics and the delay dynamics of the conductance evolution lead to faithful synaptic emulators and single-device threshold logic neurons, respectively. Unsupervised learning has been demonstrated with a fully memristive neural network consisting of these artificial synapses and neurons for the first time. In addition, the intrinsic stochasticity of the delay mechanism has been used to realize a true random number generators for security solutions. Rivu Midya, Saumil Joshi, Hao Jiang 0017, Can Li 0024, Mingyi Rao, Yunning Li, Mark Barnell, Qing Wu 0002, Qiangfei Xia, J. Joshua Yang |
ISCAS | 5 |
| 2018 | Artificial neural networks based on memristive devices
Vignesh Ravichandran, Can Li 0024, Ali BanaGozar, J. Joshua Yang, Qiangfei Xia |
Sci. China Inf. Sci. | 2 |
| 2014 | Device engineering and CMOS integration of nanoscale memristorsabstractOur group focuses on developing better nanoscale memristor with improved performance, understanding the underlying device physics, and exploring new applications for this novel device. This paper introduces our recent work on memristor device engineering and CMOS integration. We have fabricated the smallest memristors (8 nm × 8 nm) in a crossbar array, with each of the device consumes orders of magnitude lower energy per switch event than their larger counterparts. We have demonstrated that a very thin layer of chemically produced silicon oxide can be used to make memristors that only need ~0.5 V to switch. We have also proved that with multiple oxides as switching layer, both high ON/OFF ratio and high endurance can be achieved in the same device. Finally, we successfully integrated planar memristors with CMOS substrates, implementing hybrid memristor-CMOS integrated circuit with lower switching voltages and more uniform performance. Shuang Pi, Hao Jiang 0017, Can Li 0024, Qiangfei Xia |
ISCAS | 4 |