Johnson Loh

dblp:271/6804 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
4since 2021 · last 2024
0009-0001-2659-5255ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Stream Processing Architectures for Continuous ECG Monitoring Using Subsampling- Based Classifiers
abstract
Monitoring of biomedical data, such as electrocardiogram (ECG) signals, requires accelerators, which can process data streams in a continuous manner. Especially, wearable monitoring systems require both ultralow power consumption and sufficiently complex deep neural network (DNN) classifiers to identify asymptomatic and critical health conditions, such as atrial fibrillation (AF). Such continuous data streams pose unique constraints on the processing pipeline for classification systems, which can be addressed in the design methodology of application-specific integrated circuits (ASICs). In this work, we identify specific constraints to define common operating conditions, which guide the design of ECG accelerators in an algorithm–hardware codesign methodology. In specific, we show that the input frame size and the number of classifications per time frame play a significant role for the computational complexity (CC) of the classifier, as well as the ECG accelerator executing the classifier in a continuous manner. As an example, the constraints are applied in a top-down algorithm–hardware codesign flow. Here, an ECG accelerator is designed starting from an AF classifier, while proposed constraints are considered in an early design stage to estimate costs for the hardware design. In the end, it is essential for future ECG accelerators to adhere to common constraints in the design process to handle increasingly complex DNN classifiers for continuous data streams with ultralow power targets.
Johnson Loh, Tobias Gemmeke
IEEE Trans. Very Large Scale Integr. Syst.1
2023 Lossless Sparse Temporal Coding for SNN-based Classification of Time-Continuous Signals
abstract
Ultra-low power classification systems using spiking neural networks (SNN) promise efficient processing for mobile devices. Temporal coding represents activations in an artificial neural network (ANN) as binary signaling events in time, thereby minimizing circuit activity. Discrepancies in numeric results are inherent to common conversion schemes, as the atomic computing unit, i.e. the neuron, performs algorithmically different operations and, thus, potentially degrading SNN's quality of service (QoS). In this work, a lossless conversion method is derived in a top-down design approach for continuous time signals using electrocardiogram (ECG) classification as an example. As a result, the converted SNN achieves identical results compared to its fixed-point ANN reference. The computations, implied by proposed method, result in a novel hybrid neuron model located in between the integrate-and-fire (IF) and conventional ANN neuron, which numerical result is equivalent to the latter. Additionally, a dedicated SNN accelerator is implemented in 22 nm FDSOI CMOS suitable for continuous real-time classification. The direct comparison with an equivalent ANN counterpart shows that power reductions of$2.32\times$and area reductions of$7.22\times$are achievable without loss in QoS.
Johnson Loh, Tobias Gemmeke
DATE1
2023 Automatic Generation of Structured Macros Using Standard Cells ‒ Application to CIM
abstract
Regularity can be exploited to efficiently describe, place and route logic blocks with a repetitive structure. We present a design flow to automatically generate regular, standard-cell based designs, which can be seamlessly integrated into a traditional digital flow in commercial EDA software. The generated arrays can be seamlessly integrated into a “sea of gates”. with no guard-rings or keep-out areas. The flow takes a description of a regular design as an input and generates netlist, placement, constraints, routing, initial parasitics estimates and timing information. We show that, in example designs, the run-time of EDA tooling is up to 2.5x faster, reduces the critical path by 47%, reduces the metal utilization by 45% and achieves a utilization of 93%.
Christian Lanius, Jie Lou, Johnson Loh, Tobias Gemmeke
ISLPED3
2023 An Energy-Efficient and Area-Efficient Depthwise Separable Convolution Accelerator with Minimal On-Chip Memory Access
abstract
Depthwise separable convolution (DSC) has emerged as a crucial building block for developing lightweight convolutional neural networks (CNNs). In this paper, we present a hardware accelerator for DSC that enables 100% utilization of the processing element (PE) array for depthwise convolution (DWC) and achieves up to 98% utilization for pointwise convolution (PWC), while also reducing latency. By partitioning the input feature map (ifmap) SRAM of the DWC into three banks, we minimize memory access and maximize data reuse. The input activations and weights only need to be loaded once from SRAM to PE for both DWC and PWC. Additionally, to support efficient operations across different layers, we present a layerwise matching method. The proposed DSC accelerator is implemented in 22nm FDSOI technology and validated using MobileNetV1 on the CIFAR10 dataset. The post-layout results demonstrate that the proposed accelerator can operate at 1GHz and achieve an energy efficiency of 5.07 (3.96) TOPS/W and an area efficiency of 519.2 (461.52) GOPS/mm2for DWC (PWC) at 0.8V. After scaling the supply voltage down to 0.5V, the energy efficiency for the proposed accelerator increases to 13.64 TOPS/W for DWC and 10.64 TOPS/W for PWC, respectively.
Jie Lou, Christian Lanius, Florian Freye, Johnson Loh, Tobias Gemmeke
VLSI-SoC5
2020 Low-Cost DNN Hardware Accelerator for Wearable, High-Quality Cardiac Arrythmia Detection
abstract
This work implements a digital signal processing (DSP) accelerator for ECG signal classification. Targeting the integration into wearable devices for 24/7 monitoring, low energy consumption per classification is a key requirement, while maintaining a high classification accuracy at the same time. Co-optimization on algorithm and hardware level led to an architecture consisting mostly of convolution operations in the processing pipeline. The realized discrete wavelet transform and convolutional neural network (CNN) is utilized for continuous time-sequence classification in a sliding-window approach moving away from sample/batch-based processing typical for CNNs. In contrast to previous hardware realizations in this domain, the proposed design was validated using the benchmark dataset from the demanding CinC challenge 2017. The architecture achieves a competitive 0.781 Fl-score with only 5597 trainable parameters reducing the computational complexity of state-of-the-art ECGDNN software solutions by three orders of magnitude. Synthesis in a 22-nm FDSOI-CMOS technology features 0.783 $\mu$J per solution meeting requirements for edge device operation at high-end classification performance.
Johnson Loh, Jianan Wen, Tobias Gemmeke
ASAP1