Stefan Dünkel

dblp:253/8944 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Ferroelectric Digital In-Memory Computing for Scalable, Reliable, and Efficient Similarity Computation
abstract
Classification-based learning in deep neural networks, particularly few-shot learning, demands efficient similarity metrics such as Hamming distance. Conventional architectures suffer from high energy overheads due to frequent data movement between memory and processing units, hindering scalability. In-memory computing addresses this by integrating computation within memory, yet analog-based systems rely on power-hungry analog-to-digital converters (ADCs) and face scalability challenges due to device variability, especially in emerging memories. This work presents a fully digital Ferroelectric FET (FeFET)-based Logic-in-Memory (LiM) XOR cell, designed using GlobalFoundries’ 28 nm technology, eliminating ADCs and ensuring robust, energy-efficient, and scalable operation. Our 2T FeFET XOR cell, applied to 4096-bit Hamming distance calculations, achieves$23\times $lower energy,$3\times $faster latency, and$14\times $area reduction over state-of-the-art designs. Delivering 2337 Gsamples/(s$\cdot $W$\cdot $mm2) — a$300\times $improvement — this architecture offers a compelling solution for energy-efficient, reliable, and scalable AI hardware, driving sustainable computing.
Anirban Kar, Albi Mema, Thorgund Nemec, Stefan Dünkel, Halid Mulaosmanovic, Sven Beyer, Yogesh Singh Chauhan, Hussam Amrouch
IEEE Trans. Circuits Syst. I Regul. Pap.4
2026 A Reconfigurable Time-Domain In-Memory Computing Macro Using FeFET-Based CAM With Multilevel Delay Calibration in 28-nm CMOS
abstract
Time-domain nonvolatile in-memory computing (TD-nvIMC) offers a promising pathway to reduce data movement and improve energy efficiency by encoding computation in delay rather than voltage or current. This work presents a fully integrated and reconfigurable TD-nvIMC macro, fabricated in 28 nm CMOS, that combines a ferroelectric FET (FeFET)-based content-addressable memory array, a cascaded delay element chain, and a time-to-digital converter. The architecture supports binary multiply-and-accumulate (MAC) operations using XOR- and AND-based matching, as well as in-memory Boolean logic and arithmetic functions. Sub-nanosecond MAC resolution is achieved through experimentally demonstrated 550 ps delay steps, representing a$2000\times $improvement over prior FeFET TD-nvIMC work, enabled by multilevel-state calibration with$\leq 100$ps resolution. Write-disturb resilience is ensured via isolated triple-well bulks. The proposed macro achieves a measured throughput of 222.2 MOPS/cell and energy efficiency of 1887TOPS/W at 0.85 V, establishing a viable path toward scalable, energy-efficient TD-nvIMC accelerators.
Jeries Mattar, Mor M. Dahan, Stefan Dünkel, Halid Mulaosmanovic, Gunda Beernink, Sven Beyer, Eilam Yalon, Nicolás Wainstein
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 Towards Uncertainty-aware Robotic Perception via Mixed-signal BNN Engine Leveraging Probabilistic Quantum Tunneling
abstract
Integrating deep learning with environmental perception enhances robotic adaptability to complex tasks. However, its “black-box” nature, such as the lack of uncertainty quantification, poses challenges for safety-critical applications, particularly in unstructured and noisy environments. Bayesian neural networks (BNNs) offer uncertainty quantification but are limited by high hardware overhead, restricting real-time implementation on resource-constrained robots. This paper presents a mixedsignal hardware accelerator for BNNs, utilizing probabilistic quantum tunneling in fully depleted silicon-on-insulator (FDSOI) transistors to enable efficient, real-time uncertainty quantification. Device measurements indicate high-quality Gaussian random variable generation, validated through quantile-quantile plot analysis, with a high correlation coefficient ($r=0.997$) at $200 \mathrm{fJ} /$ sample. Leveraging such compact randomness, the parallel architecture achieved $10^{3}-10^{4} \times$ latency reduction at less than $2 \times$ area cost. Finally, in uncertainty-aware visual localization application of autonomous underwater vehicles, the BNN model effectively distinguishes data noise from model uncertainty, yielding significant information gain and enhancing the resampling efficiency by $4.5 \times$ at same accuracy.
Likai Pei, Xingtian Wang, Xueji Zhao, Wanxin Huang, Boyang Cheng, Halid Mulaosmanovic, Stefan Dünkel, Dominik Kleimaier, Sven Beyer, Kai Ni 0004, Mengxue Hou, Michael T. Niemier, Ningyuan Cao
DAC8