Adrian Kneip

dblp:290/7753 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-3152-1947ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2025 An Event-Based Digital Compute-In-Memory Accelerator with Flexible Operand Resolution and Layer-Wise Weight/Output Stationarity
abstract
Compute-in-memory (CIM) accelerators for spiking neural networks (SNNs) are promising solutions to enable µs-level inference latency and ultra-low energy in edge vision applications. Yet, their current lack of flexibility at both the circuit and system levels prevents their deployment in a wide range of real-life scenarios. In this work, we propose FlexSpIM, a novel digital CIM macro that supports arbitrary operand resolution and shape within a unified CIM storage for weights and membrane potentials. These circuit-level techniques enable a hybrid weight- and output-stationary dataflow at the system level to maximize operand reuse, thereby minimizing costly on- and off-chip data movements during the SNN execution. Measurement results of a fabricated FlexSpIM prototype in 40-nm CMOS demonstrate a 2× increase in 1-bit-normalized energy efficiency compared to prior fixed-precision digital CIM-based SNNs, while providing resolution reconfiguration with bitwise granularity. Our approach can save up to 90% energy in large-scale systems, while reaching a state-of-the-art classification accuracy of 95.8% on the IBM DVS gesture dataset.
Nicolas Chauvaux, Adrian Kneip, Christoph Posch, Kofi A. A. Makinwa, Charlotte Frenkel
ISCAS2
2023 A 7T-NDR Dual-Supply 28-nm FD-SOI Ultra-Low Power SRAM With 0.23-nW/kB Sleep Retention and 0.8 pJ/32b Access at 64 MHz With Forward Back Bias
abstract
This work presents a 16kB ultra-low power (ULP) SRAM macro in 28nm FD-SOI with high energy efficiency in active mode and ultra-low leakage (ULL) in sleep mode, embedded in the SleepRider micro-controller unit (MCU) intended for IoT edge applications. The proposed SRAM integrates custom 7T ULL bitcells based on negative differential resistance (NDR) structures and a pMOS-only write port, achieving$2.1\times $lower area than previous NDR-based bitcells. A dual-supply strategy combined with negative-wordline write-assist concurrently provides worst-case data retention and correct write operations, up to the 64-MHz MCU target frequency. The SRAM macro periphery combines several low-power techniques to extract the full potential of the novel 7T bitcells, reaching an unprecedented speed-energy-leakage optimum with only 2.5% area overhead. Adaptive forward body biasing (FBB) further improves active mode performance while ensuring robustness against PVT variations. Measurement results showcase a minimum energy point of 0.78pJ per 32b access (assuming 50% read/write) at 0.5V and 64MHz. Moreover, leakage power drops from 296nW/kB at 0.5V in idle conditions to 0.23nW/kB in sleep at the 0.46V data retention voltage (DRV), yielding more than$1000\times $leakage reduction. As such, the proposed SRAM achieves an excellent trade-off between area, leakage and energy in the 10-to-100MHz frequency range.
Adrian Kneip, David Bol
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 Impact of Analog Non-Idealities on the Design Space of 6T-SRAM Current-Domain Dot-Product Operators for In-Memory Computing
abstract
In-memory computing provides unprecedented power and area efficiency for the execution of convolutional neural networks by using memory bitcells to perform dot-product (DP) operations in the analog domain. Yet, these operators suffer from analog non-idealities (ANIs) that degrade the inference accuracy. This paper proposes design guidelines inferred from a holistic simulation-based analysis of the impact of ANIs on the accuracy-efficiency trade-off that affects current-domain DP operators based on conventional 6T-SRAM bitcell arrays. We define a custom SNR metric aware of the DP operand distribution to quantify decision errors associated with various ANIs, over ranges of input/output resolution and hardware design parameters. We find out that non-linearity and local mismatch are the dominant ANIs limiting the design space, while IR drops turn out to be critical only when targeting high parallelism. We then quantify the accuracy-efficiency trade-off related to these dominant ANIs across the design space and propose optimal design choices. We notably identify that using larger operators can either improve or worsen the SNR depending on the target output resolution. Furthermore, we show that hardware calibration techniques which mitigate mismatch help to recover a fraction of the lost SNR, with greater effectiveness when scaling down the supply voltage.
Adrian Kneip, David Bol
IEEE Trans. Circuits Syst. I Regul. Pap.1