EDBT 2026 Demo / reviewers in the wild / expert
Stefan Dünkel
dblp:253/8944
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ferroelectric Digital In-Memory Computing for Scalable, Reliable, and Efficient Similarity ComputationabstractClassification-based learning in deep neural networks, particularly few-shot learning, demands efficient similarity metrics such as Hamming distance. Conventional architectures suffer from high energy overheads due to frequent data movement between memory and processing units, hindering scalability. In-memory computing addresses this by integrating computation within memory, yet analog-based systems rely on power-hungry analog-to-digital converters (ADCs) and face scalability challenges due to device variability, especially in emerging memories. This work presents a fully digital Ferroelectric FET (FeFET)-based Logic-in-Memory (LiM) XOR cell, designed using GlobalFoundries’ 28 nm technology, eliminating ADCs and ensuring robust, energy-efficient, and scalable operation. Our 2T FeFET XOR cell, applied to 4096-bit Hamming distance calculations, achieves$23\times $lower energy,$3\times $faster latency, and$14\times $area reduction over state-of-the-art designs. Delivering 2337 Gsamples/(s$\cdot $W$\cdot $mm2) — a$300\times $improvement — this architecture offers a compelling solution for energy-efficient, reliable, and scalable AI hardware, driving sustainable computing. Anirban Kar, Albi Mema, Thorgund Nemec, Stefan Dünkel, Halid Mulaosmanovic, Sven Beyer, Yogesh Singh Chauhan, Hussam Amrouch |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2026 | A Reconfigurable Time-Domain In-Memory Computing Macro Using FeFET-Based CAM With Multilevel Delay Calibration in 28-nm CMOSabstractTime-domain nonvolatile in-memory computing (TD-nvIMC) offers a promising pathway to reduce data movement and improve energy efficiency by encoding computation in delay rather than voltage or current. This work presents a fully integrated and reconfigurable TD-nvIMC macro, fabricated in 28 nm CMOS, that combines a ferroelectric FET (FeFET)-based content-addressable memory array, a cascaded delay element chain, and a time-to-digital converter. The architecture supports binary multiply-and-accumulate (MAC) operations using XOR- and AND-based matching, as well as in-memory Boolean logic and arithmetic functions. Sub-nanosecond MAC resolution is achieved through experimentally demonstrated 550 ps delay steps, representing a$2000\times $improvement over prior FeFET TD-nvIMC work, enabled by multilevel-state calibration with$\leq 100$ps resolution. Write-disturb resilience is ensured via isolated triple-well bulks. The proposed macro achieves a measured throughput of 222.2 MOPS/cell and energy efficiency of 1887TOPS/W at 0.85 V, establishing a viable path toward scalable, energy-efficient TD-nvIMC accelerators. Jeries Mattar, Mor M. Dahan, Stefan Dünkel, Halid Mulaosmanovic, Gunda Beernink, Sven Beyer, Eilam Yalon, Nicolás Wainstein |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | Towards Uncertainty-aware Robotic Perception via Mixed-signal BNN Engine Leveraging Probabilistic Quantum TunnelingabstractIntegrating deep learning with environmental perception enhances robotic adaptability to complex tasks. However, its “black-box” nature, such as the lack of uncertainty quantification, poses challenges for safety-critical applications, particularly in unstructured and noisy environments. Bayesian neural networks (BNNs) offer uncertainty quantification but are limited by high hardware overhead, restricting real-time implementation on resource-constrained robots. This paper presents a mixedsignal hardware accelerator for BNNs, utilizing probabilistic quantum tunneling in fully depleted silicon-on-insulator (FDSOI) transistors to enable efficient, real-time uncertainty quantification. Device measurements indicate high-quality Gaussian random variable generation, validated through quantile-quantile plot analysis, with a high correlation coefficient ($r=0.997$) at $200 \mathrm{fJ} /$ sample. Leveraging such compact randomness, the parallel architecture achieved $10^{3}-10^{4} \times$ latency reduction at less than $2 \times$ area cost. Finally, in uncertainty-aware visual localization application of autonomous underwater vehicles, the BNN model effectively distinguishes data noise from model uncertainty, yielding significant information gain and enhancing the resampling efficiency by $4.5 \times$ at same accuracy. Likai Pei, Xingtian Wang, Xueji Zhao, Wanxin Huang, Boyang Cheng, Halid Mulaosmanovic, Stefan Dünkel, Dominik Kleimaier, Sven Beyer, Kai Ni 0004, Mengxue Hou, Michael T. Niemier, Ningyuan Cao |
DAC | 8 |