EDBT 2026 Demo / reviewers in the wild / expert
Florian Freye
dblp:323/4964
· DBLP profile ↗
7ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0003-3025-8910ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A 1.27 fJ/B/transition Digital Compute-in-Memory Architecture for Non-Deterministic Finite Automata Evaluation
Christian Lanius, Florian Freye, Tobias Gemmeke |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | A 22nm 96.83-TOPS/W Time-Domain Compute-in-Memory Engine Utilizing Mixed-Fidelity for Edge-AI Applications
Jie Lou, Florian Freye, Christian Lanius, Tobias Gemmeke |
ACM Great Lakes Symposium on VLSI | 2 |
| 2025 | An All-Digital Time-Domain Compute-in-Memory Engine for Convolutional Neural Networks in 22nmabstractThis paper presents a standard cell (SC) based time-domain compute-in-memory (TDCIM) macro for convolutional neural networks (CNNs), supporting 4b×4b multiplication. A basic cell is proposed to enable bitwise multiplication with 1-bit weights and 2-bit activations, along with double-edge computing. An 8-stage successive approximation register time-to-digital converter (SAR-TDC) is employed to convert time-domain signals into the digital domain. We leverage the inherent features of the network to enhance throughput and have fabricated the TDCIM macro in a 22nm technology. We present the measured delay and variation of the basic cell, the integral nonlinearity (INL) of the TDC, and the total computation error. The proposed macro achieves an energy efficiency of 98.44 TOPS/W at 0.55V for 4b-input and 4b-weight MAC computations. Jie Lou, Florian Freye, Christian Lanius, Tobias Gemmeke |
ISCAS | 2 |
| 2024 | An Energy Efficient All-Digital Time-Domain Compute-in-Memory Macro Optimized for Binary Neural NetworksabstractThe deployment of neural networks on edge devices has created a growing need for energy-efficient computing. In this paper, we propose an all-digital standard cell-based time-domain compute-in-memory (TDCIM) macro for binary neural networks (BNNs) that is compatible with commercial digital design flow. The TDCIM macro utilizes multiple computing chains that share one threshold chain, and supports double-edge operation, parallel computing and data reuse. Time-domain wave-pipelining technique is introduced to enhance throughput while preserving accuracy. Regular placement (RP) and custom routing (CR) are employed during place and route (P&R) to reduce systematic variations. We show computing delay, POOL computation accuracy, and network test accuracy at different voltages, indicating that the proposed TDCIM macro can maintain high accuracy under PVT variations. We implemented two versions of the TDCIM macro in 22nm FDSOI technology using foundry-provided delay cells DLY40 and DLY60, respectively. At a voltage of 0.5V, the TDCIM macro achieved an energy efficiency of 1.2 (1.05) POPS/W for DLY40 (DLY60), while maintaining a baseline accuracy of 98.9% on the MNIST dataset for both designs. Jie Lou, Florian Freye, Christian Lanius, Tobias Gemmeke |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | Scalable Time-Domain Compute-in-Memory BNN Engine with 2.06 POPS/W Energy Efficiency for Edge-AI DevicesabstractTime-domain (TD) computing has attracted attention for its high computing efficiency and suitability for applications on energy-constrained edge devices. In this paper, we present a time-domain compute-in-memory (TDCIM) macro for binary neural networks (BNNs) realized by standard as well as custom delay cells. Multiply-and-accumulate (MAC) operations, batch normalization (BN) and binarization (Bin) are all processed in the time-domain, avoiding costly digital domain post-processing. In addition, it supports flexible mapping for different kernel sizes, achieving 100% utilization. Starting from a standard cell-based implementation, we propose two custom cells that provide interesting trade-offs between energy efficiency, area and accuracy. The two proposed custom designs can achieve 1.5 and 2.06 POPS/W energy efficiencies at 0.5V and 0.6V with less cell area while maintaining model test accuracy. Jie Lou, Florian Freye, Christian Lanius, Tobias Gemmeke |
ACM Great Lakes Symposium on VLSI | 2 |
| 2023 | Hardware Trojans in fdSOIabstractWith shortening turn-around times and increasing complexity for digital circuits, design reuse, third party IP and today even physical chiplets has increased. Malicious actors have more options to introduce hardware backdoors to packaged systems, which will leak data if triggered. In this work, we show two novel approaches to introduce such backdoors, that are possible due to the specifics of fully depleted silicon on insulator (fdSOI) technology. The first method relies on modifying the doping profile of an antenna cell to introduce a covert short between the back gate and logic signals. The second method constructs specific illegal states which are latched when the clock is running with the trigger frequency. Basic test structures have been designed such that they are DRC and STA clean. LVS does not reveal the hidden structure, while measurements in silicon confirm their operation. Christian Lanius, Florian Freye, Tobias Gemmeke |
ISLPED | 2 |
| 2023 | An Energy-Efficient and Area-Efficient Depthwise Separable Convolution Accelerator with Minimal On-Chip Memory AccessabstractDepthwise separable convolution (DSC) has emerged as a crucial building block for developing lightweight convolutional neural networks (CNNs). In this paper, we present a hardware accelerator for DSC that enables 100% utilization of the processing element (PE) array for depthwise convolution (DWC) and achieves up to 98% utilization for pointwise convolution (PWC), while also reducing latency. By partitioning the input feature map (ifmap) SRAM of the DWC into three banks, we minimize memory access and maximize data reuse. The input activations and weights only need to be loaded once from SRAM to PE for both DWC and PWC. Additionally, to support efficient operations across different layers, we present a layerwise matching method. The proposed DSC accelerator is implemented in 22nm FDSOI technology and validated using MobileNetV1 on the CIFAR10 dataset. The post-layout results demonstrate that the proposed accelerator can operate at 1GHz and achieve an energy efficiency of 5.07 (3.96) TOPS/W and an area efficiency of 519.2 (461.52) GOPS/mm2for DWC (PWC) at 0.8V. After scaling the supply voltage down to 0.5V, the energy efficiency for the proposed accelerator increases to 13.64 TOPS/W for DWC and 10.64 TOPS/W for PWC, respectively. Jie Lou, Christian Lanius, Florian Freye, Johnson Loh, Tobias Gemmeke |
VLSI-SoC | 4 |