VLDB 2026 Research / reviewers in the wild / expert
Jianan Wen
dblp:271/6709
· DBLP profile ↗
9ranked-venue papers
5as first author
8since 2021 · last 2026
0009-0003-7733-8907ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | End-to-End Design Flow for Resistive Neural AcceleratorsabstractNeural hardware accelerators have demonstrated notable energy efficiency in tackling tasks, which can be adapted to artificial neural network (ANN) structures. Research is currently directed toward leveraging resistive random-access memories (RRAMs) among various memristive devices. In conjunction with complementary metal-oxide semiconductor (CMOS) technologies within integrated circuits (ICs), RRAM devices are used to build such neural accelerators. In this study, we present a neural accelerator hardware design and verification flow, which uses a lookup table (LUT)-based Verilog-A model of IHP’s one-transistor-one-RRAM (1T1R) cell. In particular, we address the challenges of interfacing between abstract ANN simulations and circuit analysis by including a tailored Python wrapper into the design process for resistive neural hardware accelerators. To demonstrate our concept, the efficacy of the proposed design flow, we evaluate an ANN for the MNIST handwritten digit recognition task, as well as for the CIFAR-10 image recognition task, with the last layer verified through circuit simulation. Additionally, we implement different versions of a 1T1R model, based on quasi-static measurement data, providing insights on the effect of conductance level spacing and device-to-device variability. The circuit simulations tackle both schematic and physical layout assessment. The resulting recognition accuracies exhibit significant differences between the purely application-level PyTorch simulation and our proposed design flow, highlighting the relevance of circuit-level validation for the design of neural hardware accelerators. Max Uhlmann, Tommaso Rizzi, Jianan Wen, Emilio Pérez-Bosch Quesada, Bakr Al Beattie, Karlheinz Ochs, Philip Ostrovskyy, Corrado Carta, Christian Wenger, Gerhard Kahmen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2026 | ReFFT: An Energy-Efficient RRAM-Based FFT AcceleratorabstractThe fast Fourier transform (FFT) is a highly efficient algorithm for computing the discrete Fourier transform (DFT). It is widely employed in various applications, including digital communication, image processing, and signal analysis. Recently, in-memory computing architectures based on emerging technologies, such as resistive RAM (RRAM), have demonstrated promising performance with low hardware cost for data-intensive applications. However, directly mapping FFT onto RRAM crossbars is challenging because the algorithm relies on many small, sequential butterfly operations, while cross-bars are optimized for large-scale, highly parallel vector–matrix multiplications (VMMs). In this paper, we introduce ReFFT, a system architecture that reformulates FFT computations for efficient execution on RRAM crossbars. ReFFT combines the reduced computational complexity of FFT with the parallel VMM capability of RRAM. We incorporate measured device data into our framework to analyze the effect of variability and develop an adaptive mapping scheme that improves twiddle-factor programming accuracy, leading to a 9.9 dB peak signal-to-noise ratio (PSNR) improvement for a 256-point FFT. Compared with prior RRAM-based DFT designs, ReFFT achieves up to 4.6× and 19.5× higher energy efficiency for 256- and 2048-point FFTs, respectively. The system is further validated in digital communication and satellite image compression tasks. Jianan Wen, Andrea Baroni, Max Uhlmann, Christian Wenger, Milos Krstic |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2026 | RRAM-Based Spectral-Domain Convolution Accelerator for Reliable and Energy-Efficient CNN InferenceabstractThe growing computational demands of convolutional neural networks (CNNs) have motivated the use of spectral-domain inference as an alternative to costly spatial-domain convolutions. In this work, we propose a resistive RAM (RRAM)-based spectral-domain convolutional layer that exploits in-memory computing (IMC) for low energy consumption and high parallelism. Both the 2-D Fourier transform and the elementwise multiplications are directly executed on RRAM crossbar arrays, while Hermitian symmetry is leveraged to further enhance the energy efficiency of the transform and subsequent spectral processing. To ensure robustness, the measured RRAM device data are incorporated into system-level simulations to evaluate inference accuracy under the impact of device variability. Furthermore, we introduce a layer-wise mapping framework that adaptively selects between spatial- and spectral-domain execution based on the tradeoff between energy efficiency and accuracy. Simulation results show that the proposed design achieves up to a$2.18\times $improvement in energy efficiency across various convolutional layer configurations compared with the spatial-domain design. For VGG-8 on CIFAR-100, the proposed architecture with the layer-wise mapping scheme reduces the energy-delay product (EDP) by 45% while incurring negligible accuracy loss. This work presents the first complete RRAM-based spectral-domain convolutional layer that accounts for device variability, providing a promising solution for edge CNN inference. Jianan Wen, Andrea Baroni, Christian Wenger, Milos Krstic, Letícia Maria Veiras Bolzani |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2025 | ReDiM: An Efficient Strategy for Read Disturb Mitigation in RRAM-Based AcceleratorsabstractResistive RAM (RRAM) has emerged as a promising non-volatile memory technology for implementing energy-efficient hardware accelerators within the in-memory computing (IMC) paradigm. However, due to the immature fabrication process and inherent material instabilities, frequent read operations during computations can induce read disturb effects, leading to unintended resistance drift and potential data corruption. Existing mitigation approaches primarily focus on detecting read disturb effects and triggering memory refresh operations. In this work, we propose an architecture-level solution that mitigates read disturb in RRAM-based accelerators. Our strategy employs crossbar duplication and decomposes the single high input pulse into two lower-amplitude pulses, effectively minimizing the risk of read disturb. To validate our approach, we develop a simulation framework that incorporates measurement data from characterized RRAM devices under read disturb stress conditions. Experimental results on VGG-8 with CIFAR-10 demonstrate that the proposed method significantly mitigates inference accuracy degradation caused by read disturb in RRAM-based accelerators, while incurring modest area and energy overheads of 12.32% and 2.15%, respectively. This work provides a practical and scalable solution for enhancing the robustness of RRAM-based accelerators in edge and high-performance computing applications. Jianan Wen, Andrea Baroni, Alberto Mistroni, Cristian Zambelli, Christian Wenger, Milos Krstic, Letícia Maria Veiras Bolzani |
IOLTS | 1 |
| 2025 | A Compact One-Transistor-Multiple-RRAM Characterization PlatformabstractEmerging non-volatile memories (eNVMs) such as resistive random-access memory (RRAM) offer an alternative solution compared to standard CMOS technologies for implementation of in-memory computing (IMC) units used in artificial neural network (ANN) applications. Existing measurement equipment for device characterisation and programming of such eNVMs are usually bulky and expensive. In this work, we present a compact size characterization platform for RRAM devices, including a custom programming unit IC that occupies less than 1 mm2of silicon area. Our platform is capable of testing one-transistor-one-RRAM (1T1R) as well as one-transistor-multiple-RRAM (1TNR) cells. Thus, to the best knowledge of the authors, this is the first demonstration of an integrated programming interface for 1TNR cells. The 1T2R IMC cells were fabricated in the IHP’s 130 nm BiCMOS technology and, in combination with other parts of the platform, are able to provide more synaptic weight resolution for ANN model applications while simultaneously decreasing the energy consumption by 50 %. The platform can generate programming voltage pulses with a 3.3 mV accuracy. Using the incremental step pulse with verify algorithm (ISPVA) we achieve 5 non-overlapping resistive states per 1T1R device. Based on those 1T1R base states we measure 15 resulting state combinations in the 1T2R cells. Max Uhlmann, Milosz Krysik, Jianan Wen, Max Frohberg, Andrea Baroni, Keerthi Dorai Swamy Reddy, Philip Ostrovskyy, Krzysztof Piotrowski, Corrado Carta, Christian Wenger, Gerhard Kahmen |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | RISC-V CPU Design Using RRAM-CMOS Standard CellsabstractThe breakdown of Dennard scaling has been the driver for many innovations such as multicore CPUs and has fueled the research into novel devices such as resistive random access memory (RRAM). These devices might be a means to extend the scalability of integrated circuits since they allow for fast and nonvolatile operation. Unfortunately, large analog circuits need to be designed and integrated in order to benefit from these cells, hindering the implementation of large systems. This work elaborates on a novel solution, namely, creating digital standard cells utilizing RRAM devices. Albeit this approach can be used both for small gates and large macroblocks, we illustrate it for a 2T2R-cell. Since RRAM devices can be vertically stacked with transistors, this enables us to construct anandstandard cell, which merely consumes the area of two transistors. This leads to a 25% area reduction compared to an equivalent CMOSnandgate. We illustrate achievable area savings with a half-adder circuit and integrate this novel cell into a digital standard cell library. A synthesized RISC-V core using RRAM-based cells results in a 10.7% smaller area than the equivalent design using standard CMOS gates. Markus Fritscher, Max Uhlmann, Philip Ostrovskyy, Daniel Reiser, Junchao Chen 0001, Jianan Wen, Carsten Schulze, Gerhard Kahmen, Dietmar Fey, Marc Reichenbach, Milos Krstic, Christian Wenger |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | Towards Reliable and Energy-Efficient RRAM Based Discrete Fourier Transform AcceleratorabstractThe Discrete Fourier Transform (DFT) holds a prominent place in the field of signal processing. The development of DFT accelerators in edge devices requires high energy efficiency due to the limited battery capacity. In this context, emerging devices such as resistive RAM (RRAM) provide a promising solution. They enable the design of high-density crossbar arrays and facilitate massively parallel and in situ computations within memory. However, the reliability and performance of the RRAM-based systems are compromised by the device non-idealities, especially when executing DFT computations that demand high precision. In this paper, we propose a novel adaptive variability-aware crossbar mapping scheme to address the computational errors caused by the device variability. To quantitatively assess the impact of variability in a communication scenario, we implemented an end-to-end simulation framework integrating the modulation and demodulation schemes. When combining the presented mapping scheme with an optimized architecture to compute DFT and inverse DFT(IDFT), compared to the state-of-the-art architecture, our simulation results demonstrate energy and area savings of up to 57 % and 18 %, respectively. Meanwhile, the DFT matrix mapping error is reduced by 83% compared to conventional mapping. In a case study involving 16-quadrature amplitude modulation (QAM), with the optimized architecture prioritizing energy efficiency, we observed a bit error rate (BER) reduction from 1.6e-2 to 7.3e-5. As for the conventional architecture, the BER is optimized from 2.9e-3 to zero. Jianan Wen, Andrea Baroni, Max Uhlmann, Markus Fritscher, Karthik KrishneGowda, Markus Ulbricht 0002, Christian Wenger, Milos Krstic |
DATE | 1 |
| 2021 | Behavioral Model of Dot-Product Engine Implemented with 1T1R Memristor Crossbar Including AssessmentabstractMemristor is an emerging electrical device that enables non-volatile storage and in-memory computing. The memristive crossbar with high memory density and low energy consumption has drawn much attention for the implementation of dot-product engines, which can be deployed in power-hungry applications with intensive multiply-accumulate operations. However, simulating the crossbar containing a group of memristors based on the device-level modeling is time consuming. In this paper, we propose a model to simulate the memristive crossbar with high flexibility and automation at the behavioral level to perform the vector-matrix multiplication. This system-level model captures the non-linearity of memristors aiming for fast and accurate simulation. With the significantly reduced simulation time, this model enables simulating the systems containing memristive crossbar with large scale like neural networks in a more practical way. Moreover, this model can be exploited to analyze the effects of variations, which provides a condition and contributes to revealing potential computational errors. A multilayer perceptron detecting breast cancer is simulated based on this model to assess the classification accuracy with the presence of variabilities. Jianan Wen, Markus Ulbricht 0002, Xin Fan 0003, Milos Krstic |
DDECS | 1 |
| 2020 | Low-Cost DNN Hardware Accelerator for Wearable, High-Quality Cardiac Arrythmia DetectionabstractThis work implements a digital signal processing (DSP) accelerator for ECG signal classification. Targeting the integration into wearable devices for 24/7 monitoring, low energy consumption per classification is a key requirement, while maintaining a high classification accuracy at the same time. Co-optimization on algorithm and hardware level led to an architecture consisting mostly of convolution operations in the processing pipeline. The realized discrete wavelet transform and convolutional neural network (CNN) is utilized for continuous time-sequence classification in a sliding-window approach moving away from sample/batch-based processing typical for CNNs. In contrast to previous hardware realizations in this domain, the proposed design was validated using the benchmark dataset from the demanding CinC challenge 2017. The architecture achieves a competitive 0.781 Fl-score with only 5597 trainable parameters reducing the computational complexity of state-of-the-art ECGDNN software solutions by three orders of magnitude. Synthesis in a 22-nm FDSOI-CMOS technology features 0.783 $\mu$J per solution meeting requirements for edge device operation at high-end classification performance. Johnson Loh, Jianan Wen, Tobias Gemmeke |
ASAP | 2 |