EDBT 2026 Demo / reviewers in the wild / expert
Wei Lu 0003
dblp:160/1674 · also Wei D. Lu
· DBLP profile ↗
22ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0003-4731-1976ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Noise-Tolerant CIM-DNNs ExplainedabstractCompute-in-memory (CIM) systems implemented with resistive random access memory (RRAM) crossbars are a promising approach for accelerating deep neural network (DNN) computations. However, it is noteworthy that RRAM-based CIM systems are susceptible to computational errors. Unlike digital computation, the nature of analog computing introduces the risk of error accumulation throughout the computation process. Various techniques have been proposed to help deal with the errors in CIM systems, among which, training methods to create noise-tolerant CIM-based DNNs (CIM-DNNs) models that are insensitive to weight variations are the most promising due to their simplicity and low implementation cost. Although promising empirical results of variation-aware training (VAT) showcasing DNN models with high tolerance to device nonidealities have been demonstrated, there remains a significant gap in the understanding of noise tolerance properties in VAT-trained CIM-DNNs and how to improve VAT based on these understandings. The exploration of these theoretical aspects represents an area requiring further investigation and research. This work endeavors to explore the fundamental properties of noise tolerance in DNNs for CIM systems. We encapsulate our contributions into three key points. First, we identify factors that influence DNNs' performance when subjected to noise through a series of training experiments. Second, we offer both theoretical insights and practical demonstrations illustrating how VAT operates to yield solutions with heightened resistance to noise during the training process. Finally, leveraging these insights, we provide guidelines for implementing VAT to obtain optimal noise tolerance in CIM-DNNs. Our objective is to establish a theoretical foundation for VAT, and building on these insights, we aim to offer general and straightforward guidelines for DNN training, experimenting with factors such as hyperparameter choices for optimizers and weight clamping. Ultimately, our aim is to contribute to general and practical solutions for the development of reliable CIM systems. Our studies focus on analyzing how noise injection and different optimizers affect the convergence dynamics during training to reach a more noise-tolerant solution through VAT. Future studies could incorporate advanced regularizers reflecting the flatness of the solution into the cost function, which may be necessary for models beyond DNNs studied here. Combined together, these techniques can potentially lead to practical solutions for the development of reliable CIM systems. Fan-Hsuan Meng, Eric Yeu-Jer Lee, Wei Lu 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | FHENDI: A Near-DRAM Accelerator for Compiler-Generated Fully Homomorphic Encryption ApplicationsabstractFully homomorphic encryption (FHE) is a powerful cryptographic technique that enables computation on encrypted data without needing to decrypt it. It has broad applications in scenarios where sensitive data needs to be processed in the cloud or in other untrusted environments. FHE applications are both compute- and memory-intensive, owing to expensive operations on large data. While prior works address the challenges of efficient compute using dedicated hardware, expensive memory transfers still remain a major limiting factor. In this work, we propose a hierarchical near-DRAM processing (NDP) solution for FHE applications, called FHENDI, that harnesses the massive DRAM bank bandwidth. We observe various data access patterns in FHE that reveal distinct levels of parallelism: element-wise, limb-wise, coefficient-wise, and ciphertext-wise. FHENDI exploits these levels of parallelism to map FHE operations and data onto different hierarchies of our design, while addressing three major challenges with NDP for FHE: (i) the lack of bank-to-bank communication support, (ii) limited die-to-die bandwidth, and (iii) large memory access latencies. We resolve the first problem through a novel, conflict-free mapping algorithm built atop localized permutation networks that enables efficient element-wise and butterfly operations in FHE. The second problem is addressed by pipelining the execution of parallel bootstrap operations observed in compiled FHE workloads. Finally, we hide the memory access latency behind computation latency by exploiting a dual-banking scheme and subarray-level parallelism (SLP) of the DRAM banks. We evaluate FHENDI using representative workloads in the domains of privacy-preserving machine learning inference on CNNs and Transformers, database range query, and sorting, that are obtained using a compiler framework called HElayers. We compare FHENDI with a server-class CPU and GPU running the state-of-the-art HEaaN library, and an FHE accelerator ASIC, and report mean speedups of $2145.8 \times, 118.29 \times$, and $2.45 \times$, respectively. Yongmo Park, Aporva Amarnath, Subhankar Pal, Karthik Swaminathan, Alper Buyuktosunoglu, Hayim Shaul, Ehud Aharoni, Nir Drucker, Wei Lu 0003, Omri Soceanu, Pradip Bose |
HPCA | 9 |
| 2024 | TT-CIM: Tensor Train Decomposition for Neural Network in RRAM-Based Compute-in-Memory SystemsabstractCompute-in-Memory (CIM) implemented with Resistive-Random-Access-Memory (RRAM) crossbars is a promising approach for accelerating Convolutional Neural Network (CNN) computations. The growing size in the number of parameters in state-of-the-art CNN models, however, creates challenge for on-chip weight storage for CIM implementations, and CNN compression becomes a crucial topic of exploration. Tensor Train (TT) decomposition can be used to decompose a tensor into smaller ones with fewer parameters, at the cost of increased number of computations. In this work we propose a technique to minimize intermediate operations across the full convolution operation and improve hardware utilization to implement TT-CNNs in CIM systems. We first use an iterative decompose-and-fine-tune method to prepare TT-CNNs. We then propose an inter-convolutional-step reuse scheme to reduce the required operation count and post-mapping RRAM count for TT-CNN implementation in tiled-CIM architecture. We demonstrate that through proper mapping, pipelining, and reuse, effective compression ratio of 12 and 20 with 0.8% and 1.4% accuracy drop, respectively for WRN; and effective compression ratio of 6 and 11 with 0.9% and 1.2% accuracy drop for VGG8. We also show that around 30% higher hardware utilization than the original CNN format can be achieved using the proposed TT-CIM approaches. Fan-Hsuan Meng, Zhengya Zhang, Wei Lu 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | Training Spiking Neural Networks Using Lessons From Deep LearningabstractThe brain is the perfect place to look for inspiration to develop more efficient neural networks. The inner workings of our synapses and neurons provide a glimpse at what the future of deep learning might look like. This article serves as a tutorial and perspective showing how to apply the lessons learned from several decades of research in deep learning, gradient descent, backpropagation, and neuroscience to biologically plausible spiking neural networks (SNNs). We also explore the delicate interplay between encoding data as spikes and the learning process; the challenges and solutions of applying gradient-based learning to SNNs; the subtle link between temporal backpropagation and spike timing-dependent plasticity; and how deep learning might move toward biologically plausible online learning. Some ideas are well accepted and commonly used among the neuromorphic engineering community, while others are presented or justified for the first time here. A series of companion interactive tutorials complementary to this article using our Python package,snnTorch, are also made available: https://snntorch.readthedocs.io/en/latest/tutorials/index.html. Jason Kamran Eshraghian, Max Ward 0001, Emre Neftci, Xinxin Wang 0002, Gregor Lenz, Girish Dwivedi, Mohammed Bennamoun, Doo Seok Jeong, Wei Lu 0003 |
Proc. IEEE | 9 |
| 2022 | Design Space Exploration of Dense and Sparse Mapping Schemes for RRAM ArchitecturesabstractThe impact of device and circuit-level effects in mixed-signal Resistive Random Access Memory (RRAM) accelerators typically manifest as performance degradation of Deep Learning (DL) algorithms, but the degree of impact varies based on algorithmic features. These include network architecture, capacity, weight distribution, and the type of inter-layer connections. Techniques are continuously emerging to efficiently train sparse neural networks, which may have activation sparsity, quantization, and memristive noise. In this paper, we present an extended Design Space Exploration (DSE) methodology to quantify the benefits and limitations of dense and sparse mapping schemes for a variety of network architectures. While sparsity of connectivity promotes less power consumption and is often optimized for extracting localized features, its performance on tiled RRAM arrays may be more susceptible to noise due to under-parameterization, when compared to dense mapping schemes. Moreover, we present a case study quantifying and formalizing the trade-offs of typical non-idealities introduced into l-Transistor-l-Resistor (ITIR) tiled memristive architectures and the size of modular crossbar tiles using the CIFAR-10 dataset. Corey Lammie, Jason Kamran Eshraghian, Chenqi Li, Amirali Amirsoleimani, Roman Genov, Wei Lu 0003, Mostafa Rahimi Azghadi |
ISCAS | 6 |
| 2022 | Spatiotemporal Spike Pattern Detection with Second-order Memristive SynapsesabstractSpatiotemporal patterns of spike trains convey critical information in a biological neural network. Second-order memristive devices, whose internal state variables offer short- and long-term temporal dynamics, have been employed to natively decode the temporal correlation of spiking patterns through bio-realistic implementation of synaptic learning rules. In this work, we demonstrate that a single artificial postsynaptic neuron equipped with an array of second-order memristive synapses can localize a precise spatiotemporal firing pattern, which repeats irregularly within an equally dense background of Poisson spiking events, in an unsupervised fashion. Sangmin Yoo, Fan-Hsuan Meng, Wei Lu 0003 |
ISCAS | 4 |
| 2021 | Device Non-Ideality Effects and Architecture-Aware Training in RRAM In-Memory Computing ModulesabstractWe studied factors that could degrade model performance in analog RRAM in-memory-computing (IMC) systems, including limited array size, ADC resolution, on/off ratio, and device conductance variations. Different levels of architecture-aware training methods were developed to mitigate these factors and allow the system to achieve accuracy comparable to floating-point baseline with realistic device parameters. Yongmo Park, Wei Lu 0003 |
ISCAS | 3 |
| 2021 | Neural connectivity inference with spike-timing dependent plasticity network
John Moon, Xiaojian Zhu, Wei Lu 0003 |
Sci. China Inf. Sci. | 4 |
| 2021 | How to Build a Memristive Integrate-and-Fire Model for Spiking Neuronal Signal GenerationabstractWe present and experimentally validate two minimal compact memristive models for spiking neuronal signal generation using commercially available low-cost components. The first neuron model is called the Memristive Integrate-and-Fire (MIF) model, for neuronal signaling with two voltage levels: the spike-peak, and the rest-potential. The second model MIF2 is also presented, which promotes local adaptation by accounting for a third refractory voltage level during hyperpolarization. We show both compact models are minimal in terms of the number of circuit elements and integration area. Using the MIF and MIF2 models, we postulate the design of a memristive solid-state brain with an estimation of its surface area and power consumption. Analytical projections show that a memristive solid-state brain could be realized within (i) the surface area of the median human brain, 2,400cm2, (ii) the same volume of the median human brain, and (iii) a total power budget of approximately 20 W using a 3.5 nm technology. Distinct from the past decade of memristive neuron literature, our benchmarks are attained using generic commercially available memristors that are reproducible using off-the-shelf components. We expect this work can promote more experimental demonstrations of memristive circuits that do not rely on prohibitively expensive fabrication processes. Sung-Mo Kang 0001, Jason Kamran Eshraghian, Peng Zhou 0017, Bai-Sun Kong, Xiaojian Zhu, Ahmet Samil Demirkol, Alon Ascoli, Ronald Tetzlaff, Wei Lu 0003, Leon O. Chua |
IEEE Trans. Circuits Syst. I Regul. Pap. | 11 |
| 2020 | Stabilization of Mode-Dependent Impulsive Hybrid Systems Driven by DFA With Mixed-Mode EffectsabstractThis paper is concerned with mode-dependent impulsive hybrid systems driven by deterministic finite automaton (DFA) with mixed-mode effects. In the hybrid systems, a complex phenomenon called mixed mode, caused in time-varying delay switching systems, is considered explicitly. Furthermore, mode-dependent impulses, which can exist not only at the instants coinciding with mode switching but also at the instants when there is no system switching, are also taken into consideration. First, we establish a rigorous mathematical equation expression of this class of hybrid systems. Then, several criteria of stabilization of this class of hybrid systems are presented based on semi-tensor product (STP) techniques, multiple Lyapunov-Krasovskii functionals, as well as the average dwell time approach. Finally, an example is simulated to illustrate the effectiveness of the obtained results. Anni Li, Wei Lu 0003, Jitao Sun |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Feature extraction and analysis using memristor networksabstractApproaches toward feature extraction and image analysis using memristor networks will be discussed. Through hardware implementation of a sparse coding algorithm in a fabricated 32×32 memristor array, lateral inhibition among neurons is obtained and allows the network to settle to a more optimal, sparse solution from many possible solutions. This capability enables the network to identify the hidden features that constitute the input real-world image. By utilizing the internal dynamics of the device, temporal features can also be learned and processed. Fuxi Cai, Wei Lu 0003 |
ISCAS | 2 |
| 2018 | Neuromorphic computing with memristive devices
Mohammed Affan Zidan, Wei Lu 0003 |
Sci. China Inf. Sci. | 3 |
| 2016 | Feature Extraction Using Memristor NetworksabstractCrossbar arrays of memristive elements are investigated for the implementation of dictionary learning and sparse coding of natural images. A winner-take-all training algorithm, in conjunction with Oja's rule, is used to learn an overcomplete dictionary of feature primitives that resemble Gabor filters. The dictionary is then used in the locally competitive algorithm to form a sparse representation of input images. The impacts of device nonlinearity and parameter variations are evaluated and a compensating procedure is proposed to ensure the robustness of the sparsification. It is shown that, with proper compensation, the memristor crossbar architecture can effectively perform sparse coding with distortion comparable with ideal software implementations at high sparsity, even in the presence of large device-to-device variations in the excess of 100%. Patrick Sheridan, Wei Lu 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | 3D ReRAM with Field Assisted Super-Linear Threshold (FASTTM) Selector technology for super-dense, low power, low latency data storage systemsabstract3D Resistive Ram (ReRAM) technology exhibits the best attributes to suit present and emerging non-volatile memory storage applications. However, the major challenge to make ReRAM work in a 3D crossbar array is the integration of a selector device with a ReRAM device. The selector device will need to solve the so called “sneak path” barrier and enable large density memory arrays with low power consumption. Here, we report a Field Assisted Superlinear Threshold (FASTTM) Selector technology that overcomes the sneak path barrier with a selectivity ratio of 10E10. The switching and recover speed, on/off ratio, switching slope, program, erase, and read endurance, and variability of the FASTTMselector will be discussed. Prototype 1S1R devices with the FASTTM selector integrated with a low current ReRAM cell have been demonstrated and characterized. Figure 1 shows the representative I-V characteristics of a ReRAM cell with integrated FASTTMselector. Sung Hyun Jo, Tanmay Kumar, Mehdi Asnaashari, Wei Lu 0003, Hagop Nazarian |
ASP-DAC | 4 |
| 2015 | A Low-Power Variation-Aware Adaptive Write Scheme for Access-Transistor-Free Memristive MemoryabstractRecent advances in access-transistor-free memristive crossbars have demonstrated the potential of memristor arrays as high-density and ultra-low-power memory. However, with considerable variations in the write-time characteristics of individual memristors, conventional fixed-pulse write schemes cannot guarantee reliable completion of the write operations and waste significant amount of energy. We propose an adaptive write scheme that adaptively adjusts the write pulses to address such variations in memristive arrays, resulting in 7×--11× average energy saving in our case studies. Our scheme embeds an online monitor to detect the completion of a write operation and takes into account the parasitic effect of line-shared devices in access-transistor-free crossbars. This feature also helps shorten the test time of memory march algorithms by eliminating the need of a verifying read right after a write, which is commonly employed in the test sequences of march algorithms. Amirali Ghofrani, Miguel Angel Lastras-Montaño, Siddharth Gaba, Melika Payvand, Wei Lu 0003, Luke Theogarajan, Kwang-Ting Cheng |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2014 | Memristive devices for stochastic computingabstractWe show resistive switching effects in memristive devices exhibit significant stochasticity. When the switching is dominated by a single filament, the switching time is fully random and shows a broad distribution. However, the switching distribution can be predicted and responds well to controlled changes in the programming conditions. The native stochastic characteristic can be used to generate random bit streams with predictable biases that can lead to efficient and error-tolerant computing. Siddharth Gaba, Phil Knag, Zhengya Zhang, Wei Lu 0003 |
ISCAS | 4 |
| 2014 | Analog signal processing on a FPAA/memristor hybrid circuitabstractWe show a Field-Programmable Analog Array complemented with post-processed memristors (FPAA/memristor hybrid circuit), and present it as a platform for analog signal processing. The FPAA is fabricated on CMOS and uses floating-gate transistors (FGT) to realize programmable wiring fabrics and analog computing resources. The memristors, post-processed on top of the FPAA, are analog, which means that their conductance can be programmed in a continuous fashion. We highlight the key differences between FGTs and memristors, and discuss how analog signal processing tasks could be divided between them. Experimental results of the FPAA-memristor hybrid circuit are presented. Mika Laiho, Eero Lehtonen, Jennifer Hasler, Jiantao Zhou 0003, Wei Lu 0003, Jussi H. Poikonen |
ISCAS | 6 |
| 2014 | Pattern recognition with memristor networksabstractIn this paper we develop the concept of implementing pattern recognition algorithms in analog memristor networks. First, a device model is presented with experimental results demonstrating the feasibility of using WOx-based memristors to represent the tunable weights in a neural network. Next, simulation results demonstrate that an array of these memristors can be used to implement an unsupervised learning algorithm for pattern recognition. Handwritten digits are classified as an example problem while the concept is developed for more general use. Patrick Sheridan, Wei Lu 0003 |
ISCAS | 3 |
| 2012 | Memristive analog arithmetic within cellular arraysabstractThis paper describes how memristors, together with CMOS transistors (CMOS-memristor hybrids), could be used for analog arithmetic within cellular (locally connected, regular) computing arrays, i.e., cellular nanoscale networks (CNN). Elementary analog programming of memristors and copying memristance values are described. Also, we show how memristors could be used for addition, subtraction, multiplication and division. Furthermore, we demonstrate how a local memristive computing unit could be mapped into a CNN cell. Relevant memristor models are shown and key simulations are presented. Mika Laiho, Eero Lehtonen, Wei Lu 0003 |
ISCAS | 3 |
| 2011 | Two-terminal resistive switches (memristors) for memory and logic applicationsabstractWe review the recent progress on the development of two-terminal resistive devices (memristors). Devices based on solid-state electrolytes (e.g. a-Si) have been shown to possess a number of promising performance metrics such as yield, on/off ratio, switching speed, endurance and retention suitable for memory or reconfigurable circuit applications. In addition, devices with incremental resistance changes have been demonstrated and can be used to emulate synaptic functions in hardware based neuromorphic circuits. Device and SPICE modeling based on a properly chosen internal state variable have been carried out and will be useful for large-scale circuit simulations. Wei Lu 0003, Kuk-Hwan Kim, Ting Chang, Siddharth Gaba |
ASP-DAC | 1 |
| 2011 | Time-dependency of the threshold voltage in memristive devicesabstractWe describe a generic exponential model with four parameters for thin-film memristive devices. This model is used to analyze the time dependency of the threshold voltage which defines the transition between non-programming and programming phases of the device. A relationship between timescale of operation and threshold voltage is derived. Furthermore, self-terminating programming is considered using the results of this analysis. Finally, the effect of parameter variations on the threshold voltage is analyzed. Eero Lehtonen, Jussi H. Poikonen, Mika Laiho, Wei Lu 0003 |
ISCAS | 4 |
| 2010 | Si Memristive devices applied to memory and neuromorphic circuitsabstractWe report studies on nanoscale Si-based memristive devices for memory and neuromorphic applications. The devices are based on ion motion inside an insulating a-Si matrix. Digital devices show excellent performance metrics including scalability, speed, ON/OFF ratio, endurance and retention. High density non-volatile memory arrays based on a crossbar structure have been fabricated and tested. Devices inside a 1kb array can be individually addressed with excellent reproducibility and reliability. By adjusting the device and material structures, nanoscale analog memristor devices have also been demonstrated. The analog memristor devices exhibit incremental conductance changes that are controlled by the charge flown through the device. The performances of the digital and analog devices are thought to be determined by the formation of a dominant conducting filament and the continuous motion of a uniform conduction front, respectively. Sung Hyun Jo, Kuk-Hwan Kim, Ting Chang, Siddharth Gaba, Wei Lu 0003 |
ISCAS | 5 |