EDBT 2026 Demo / reviewers in the wild / expert
Christian Lanius
dblp:240/3419
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0001-7107-3782ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ConfASR: A Conformer Block Accelerator for Speech Recognition Optimized for Edge DevicesabstractAttention-based neural networks, like transformers, have significantly improved automatic speech recognition (ASR). Adding convolution operations to transformers results in conformers, which enable better learning of local dependencies and reduce word error rate. We introduce ConfASR, the first conformer block accelerator designed for efficient ASR inference on edge devices. Our system is optimized both in terms of algorithms and hardware to support all transformer operations and additional features required for conformers, including depthwise-separable convolution and learned positional encoding. We propose a hardware-friendly normalization, shared scaling factors for non-linear functions, and an efficient dataflow with a shared MAC array that keeps all activations on chip. Implemented in a 22 nm FDSOI technology, ConfASR operates at 250 MHz with a power consumption of 359 mW, and a die area of $1.19 \mathrm{~mm}^{2}$. It performs over 900 times faster than necessary for real-time streaming requirements. This makes the architecture suitable not only for ASR but also for other transformer-based applications. ConfASR reduces latency by over $4 \times$ and power consumption by $16 \times$ during real-time use compared to previous solutions, while supporting more functionality. Malte Wabnitz, Max Nilovic, Finn Scholz, Dominik Friedrich, Christian Lanius, Jie Lou, Tobias Gemmeke |
ASP-DAC | 5 |
| 2025 | A 1.27 fJ/B/transition Digital Compute-in-Memory Architecture for Non-Deterministic Finite Automata Evaluation
Christian Lanius, Florian Freye, Tobias Gemmeke |
ACM Great Lakes Symposium on VLSI | 1 |
| 2025 | A 22nm 96.83-TOPS/W Time-Domain Compute-in-Memory Engine Utilizing Mixed-Fidelity for Edge-AI Applications
Jie Lou, Florian Freye, Christian Lanius, Tobias Gemmeke |
ACM Great Lakes Symposium on VLSI | 3 |
| 2025 | An All-Digital Time-Domain Compute-in-Memory Engine for Convolutional Neural Networks in 22nmabstractThis paper presents a standard cell (SC) based time-domain compute-in-memory (TDCIM) macro for convolutional neural networks (CNNs), supporting 4b×4b multiplication. A basic cell is proposed to enable bitwise multiplication with 1-bit weights and 2-bit activations, along with double-edge computing. An 8-stage successive approximation register time-to-digital converter (SAR-TDC) is employed to convert time-domain signals into the digital domain. We leverage the inherent features of the network to enhance throughput and have fabricated the TDCIM macro in a 22nm technology. We present the measured delay and variation of the basic cell, the integral nonlinearity (INL) of the TDC, and the total computation error. The proposed macro achieves an energy efficiency of 98.44 TOPS/W at 0.55V for 4b-input and 4b-weight MAC computations. Jie Lou, Florian Freye, Christian Lanius, Tobias Gemmeke |
ISCAS | 3 |
| 2024 | An Energy Efficient All-Digital Time-Domain Compute-in-Memory Macro Optimized for Binary Neural NetworksabstractThe deployment of neural networks on edge devices has created a growing need for energy-efficient computing. In this paper, we propose an all-digital standard cell-based time-domain compute-in-memory (TDCIM) macro for binary neural networks (BNNs) that is compatible with commercial digital design flow. The TDCIM macro utilizes multiple computing chains that share one threshold chain, and supports double-edge operation, parallel computing and data reuse. Time-domain wave-pipelining technique is introduced to enhance throughput while preserving accuracy. Regular placement (RP) and custom routing (CR) are employed during place and route (P&R) to reduce systematic variations. We show computing delay, POOL computation accuracy, and network test accuracy at different voltages, indicating that the proposed TDCIM macro can maintain high accuracy under PVT variations. We implemented two versions of the TDCIM macro in 22nm FDSOI technology using foundry-provided delay cells DLY40 and DLY60, respectively. At a voltage of 0.5V, the TDCIM macro achieved an energy efficiency of 1.2 (1.05) POPS/W for DLY40 (DLY60), while maintaining a baseline accuracy of 98.9% on the MNIST dataset for both designs. Jie Lou, Florian Freye, Christian Lanius, Tobias Gemmeke |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | Fully Digital, Standard-Cell-Based Multifunction Compute-in-Memory Arrays for Genome SequencingabstractThe rapid advancement in genome sequencing technology has led to a significant increase in the number of genomic reads in recent years. Due to the immense size of reference genomes, which can be up to 3 billion bases, finding optimal solutions for through approximate string matching proves to be computationally challenging. Current alignment algorithms address this by performing a preprocessing step to efficiently calculate likely matching regions and only aligning at the base level within these regions. This article demonstrates the acceleration of sorting and searching in memories, both crucial components of genome alignment algorithms. We designed a compute-in-memory (CIM) array using standard cells, which is capable of sorting datastreams blockwise, merging sorted blocks, as well as operating as a content addressable memory (CAM) while also being able to perform multiword logic operations. We address the problem of datasets not fitting into on-chip memory by reusing the CIM array for a merge sorting step, enabling arbitrarily sized sorting. Our 2.6-$\mu \text{m}^{2}$/bit design, fabricated using 22-nm fully depleted silicon-on-insulator (FDSOI) technology, yields a throughput of up to 4.28 GB/s at$f_{\text {max}}$and 4.97 nJ/sort at the minimum energy point (MEP) when executing sort operations. Christian Lanius, Tobias Gemmeke |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | Scalable Time-Domain Compute-in-Memory BNN Engine with 2.06 POPS/W Energy Efficiency for Edge-AI DevicesabstractTime-domain (TD) computing has attracted attention for its high computing efficiency and suitability for applications on energy-constrained edge devices. In this paper, we present a time-domain compute-in-memory (TDCIM) macro for binary neural networks (BNNs) realized by standard as well as custom delay cells. Multiply-and-accumulate (MAC) operations, batch normalization (BN) and binarization (Bin) are all processed in the time-domain, avoiding costly digital domain post-processing. In addition, it supports flexible mapping for different kernel sizes, achieving 100% utilization. Starting from a standard cell-based implementation, we propose two custom cells that provide interesting trade-offs between energy efficiency, area and accuracy. The two proposed custom designs can achieve 1.5 and 2.06 POPS/W energy efficiencies at 0.5V and 0.6V with less cell area while maintaining model test accuracy. Jie Lou, Florian Freye, Christian Lanius, Tobias Gemmeke |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | Hardware Trojans in fdSOIabstractWith shortening turn-around times and increasing complexity for digital circuits, design reuse, third party IP and today even physical chiplets has increased. Malicious actors have more options to introduce hardware backdoors to packaged systems, which will leak data if triggered. In this work, we show two novel approaches to introduce such backdoors, that are possible due to the specifics of fully depleted silicon on insulator (fdSOI) technology. The first method relies on modifying the doping profile of an antenna cell to introduce a covert short between the back gate and logic signals. The second method constructs specific illegal states which are latched when the clock is running with the trigger frequency. Basic test structures have been designed such that they are DRC and STA clean. LVS does not reveal the hidden structure, while measurements in silicon confirm their operation. Christian Lanius, Florian Freye, Tobias Gemmeke |
ISLPED | 1 |
| 2023 | Automatic Generation of Structured Macros Using Standard Cells ‒ Application to CIMabstractRegularity can be exploited to efficiently describe, place and route logic blocks with a repetitive structure. We present a design flow to automatically generate regular, standard-cell based designs, which can be seamlessly integrated into a traditional digital flow in commercial EDA software. The generated arrays can be seamlessly integrated into a “sea of gates”. with no guard-rings or keep-out areas. The flow takes a description of a regular design as an input and generates netlist, placement, constraints, routing, initial parasitics estimates and timing information. We show that, in example designs, the run-time of EDA tooling is up to 2.5x faster, reduces the critical path by 47%, reduces the metal utilization by 45% and achieves a utilization of 93%. Christian Lanius, Jie Lou, Johnson Loh, Tobias Gemmeke |
ISLPED | 1 |
| 2023 | An Energy-Efficient and Area-Efficient Depthwise Separable Convolution Accelerator with Minimal On-Chip Memory AccessabstractDepthwise separable convolution (DSC) has emerged as a crucial building block for developing lightweight convolutional neural networks (CNNs). In this paper, we present a hardware accelerator for DSC that enables 100% utilization of the processing element (PE) array for depthwise convolution (DWC) and achieves up to 98% utilization for pointwise convolution (PWC), while also reducing latency. By partitioning the input feature map (ifmap) SRAM of the DWC into three banks, we minimize memory access and maximize data reuse. The input activations and weights only need to be loaded once from SRAM to PE for both DWC and PWC. Additionally, to support efficient operations across different layers, we present a layerwise matching method. The proposed DSC accelerator is implemented in 22nm FDSOI technology and validated using MobileNetV1 on the CIFAR10 dataset. The post-layout results demonstrate that the proposed accelerator can operate at 1GHz and achieve an energy efficiency of 5.07 (3.96) TOPS/W and an area efficiency of 519.2 (461.52) GOPS/mm2for DWC (PWC) at 0.8V. After scaling the supply voltage down to 0.5V, the energy efficiency for the proposed accelerator increases to 13.64 TOPS/W for DWC and 10.64 TOPS/W for PWC, respectively. Jie Lou, Christian Lanius, Florian Freye, Johnson Loh, Tobias Gemmeke |
VLSI-SoC | 3 |
| 2022 | NEUROTEC I: Neuro-inspired Artificial Intelligence Technologies for the Electronics of the FutureabstractThe field of neuromorphic computing is approaching an era of rapid adoption driven by the urgent need of a substitute for the von Neumann computing architecture. NEUROTEC I: “Neuro-inspired Artificial Intelligence Technologies for the Elec-tronics of the Future” project is an initiative sponsored by the German Federal Ministry of Education and Research (BMBF for its initials in German), that aims to effectively advance the foundations for the utilization and exploitation of neuromorphic computing. NEUROTEC I stands at its successful “final stage” driven by the collaboration from more than 8 institutes from the Jiilich Research Center and the RWTH Aachen University, as well as collaboration from several high-tech industry partners. The NEUROTEC I project considers the field interplay among materials, circuits, design and simulation tools. This paper provides an overview of the project's overall structure and discusses the scientific achievements of its individual activities. Melvin Galicia, Stephan Menzel, Farhad Merchant, Maximilian Müller, Qing-Tai Zhao, Felix Cüppers, Abdur R. Jalil, Qi Shu, Peter Schüffelgen, Gregor Mussler, Carsten Funck, Christian Lanius, Stefan Wiefels, Moritz von Witzleben, Christopher Bengel, Nils Kopperberg, Tobias Ziegler 0005, R. Walied Ahmad, Alexander Krüger, Letícia Maria Veiras Bolzani, Regina Dittmann, Susanne Hoffmann-Eifert, Vikas Rana, Detlev Grützmacher, Matthias Wuttig, Dirk J. Wouters, Andrei Vescan, Tobias Gemmeke, Joachim Knoch, Max Christian Lemme, Rainer Leupers, Rainer Waser |
DATE | 13 |