Jens Karrenbauer

dblp:247/3335 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-5060-0843ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Noise Reduction in Hearing-Aid Processors: Traditional Methods vs. Neural Networks
abstract
Many deep neural networks (DNNs) have been applied lately in the field of speech enhancement. One particular subfield, where DNNs have shifted the boundaries of what is considered possible, is noise reduction, where the degrading effects of sounds interfering with speech are minimized. This is especially relevant for hearing impaired listeners, as their ability to understand speech in noisy circumstances is reduced. In contrast to traditional methods, which are known to improve speech quality, DNNs promise to also improve speech intelligibility. Due to the high computational complexity, DNNs have not yet been deployed on a hearing aid processor, constrained by frequencies up to 50 MHz and memory up to 2 MB. In this work we deploy a convolutional neural network (CNN) trained for noise reduction to a hearing-aid system-on-chip (SoC) developed at our institute. Real time capability is achieved by thorough optimization of the C -Code, leading to a speed up by a factor of 88 for the inference relevant layers when compared to a naïve C-Code implementation. The CNN approach is compared to an implementation of a traditional noise reduction method regarding their speech enhancement performance on white and complex noise and their computational cost. While both methods improve the speech quality measured with Perceptual Evaluation of Speech Quality (PESQ), only the CNN achieves a Short-Time Objective Intelligibility (STOI) improvement of 0.077 for complex noise. On the other hand, the CNN has a higher processor utilization of 60.1% compared to 23.5% for the traditional approach. Nonetheless, both methods are real time capable and consume only 3.3 mW for the CNN and 1.78 mW for the traditional approach, respectively.
Simon C. Klein, Lando Rossol, Finn Venema, Sven Schönewald, Jens Karrenbauer, Holger Blume
ASAP5
2024 Enhancing a Hearing Aid Processor with ISA Extensions Supporting Flexible Fixed-Point Formats
abstract
As the number of individuals experiencing hearing problems rises, research in this area is increasing. In particular, algorithms for enhancing sound quality are becoming more advanced. However, the energy consumption of hearing aids is restricted to a few milliwatts. New hearing aid processors and hardware must be developed to tackle this issue. Therefore, this paper presents and evaluates custom hardware units for hearing aids suitable for flexible fixed-point formats. The units are added to the instruction set architecture (ISA) of a Tensilica Fusion G6. This is one of two high-level programmable application-specific instruction-set processors (ASIP) integrated into the Smart Hearing Aid Processor (SmartHeaP), a hearing aid system-on-chip (SoC) fabricated in 22nm fully depleted silicon on insulator (FD-SOI) technology. As many audio algorithms operate in the frequency domain, complex-domain units like a complex multiply-accumulate (CMAC) unit are introduced first. Additionally, coordinate rotation digital computer (CORDIC) operations have been added to speed up nonlinear functions such as logarithms, another function frequently used in hearing aid applications. The implemented extensions are integrated into a MATLAB fixed-point framework to simplify access to the ISA extensions for algorithm developers. It automatically generates fixed-point C code with direct access to the added instructions, reducing development time and complexity. Integrating the seven proposed instructions with corresponding register files increases the core area by 20% or 0.065 mm2in a 22nm front-end synthesis. On the other hand, these extensions reduce the cycle count for two evaluated hearing aid algorithms, a beamformer and a loudness compensator, by 93% and 38%. A reduction of up to 84% is achieved for standalone mathematical functions, like logarithmic calculations. These performance improvements directly translate into power savings by allowing a reduced clock frequency.
Jens Karrenbauer, Sven Schönewald, Simon C. Klein, Holger Blume
ASAP1
2024 Design Space Exploration of Semantic Segmentation CNN SalsaNext for Constrained Architectures
abstract
The growing use of LiDAR systems and constrained computing resources in the automotive sector require efficient LiDAR processing. SalsaNext, a convolutional neural network for semantic segmentation, is a promising candidate for deployment in that area. To extend the research regarding its quantization and investigate its adaptability to constrained resources, a design space exploration is performed. The design space, defined by model size, topology, and compute precision, is evaluated on a Jetson AGX Orin regarding classification accuracy, latency, and energy efficiency. The results display a trade-off between classification accuracy and runtime. The smallest model evaluated in INT8 on the GPU provides the smallest latency of 14.48 ms with a mloU score of 43.2%. A mloU score of 47.7% at a latency of 26.92 ms can be achieved with the medium-sized model and modified topology evaluated in INT8 on the DLA. The medium-sized model with modified topology provides good classification accuracy evaluated in FP32 on the GPU with a mloU score of 55.2% in 67.85 ms.
Oliver Renke, Christoph Riggers, Jens Karrenbauer, Holger Blume
ASAP3