Hans Jakob Damsgaard

dblp:286/5373 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0001-8409-0282ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-author · 9 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Novel Direction-of-Arrival-Based Localization in Massive DECT-2020 5G NR Networks
abstract
This research investigates an affordable, energy-efficient direction-of-arrival (DOA)-based localization solution for digital enhanced cordless telecommunications (DECTs) 2020 new radio (NR), a new standard lacking a native positioning feature. This standard enables massive Internet of Things (IoT) networks, a vast 5G network interconnecting an unparalleled number of low-cost and battery-operated smart sensors. However, integrating DOA localization into such networks is challenging due to cost constraints and power limitations. We propose a potentially cost-effective solution using a single radio-frequency (RF) chain for uniform L-shaped antenna arrays. Each antenna takes turns sampling the orthogonal frequency division multiplexing (OFDM) signal via an RF switch, enabled by time-dividing the OFDM signal into sample and switch slots. Further, we introduce a novel DOA method optimized for single Line-of-Sight (LOS) OFDM signals and array sequential sampling. This method leverages the dual shift-invariant properties of L-shaped antenna arrays and the array frequency response to estimate the azimuth and elevation angles. Experiments in an indoor environment reveal that at a signal-to-noise ratio (SNR) of 15 dB, over 50% of data achieve subdegree angular accuracy, increasing to 75% at 20 dB. Thus, over 50% of position estimations fall below the submeter error level at 15 dB SNR, rising to nearly 75% at 25 dB SNR. Our findings also indicate that halving the slot rate by proportionately reducing active subcarriers does not compromise accuracy. Experiments on the nRF52480 system-on-chip show the new DOA method is both fast and energy-efficient, taking only 0.76–2.26 ms and consuming 5.08–15.1 nWh.
Tiago Troccoli, Hans Jakob Damsgaard, Juho Pirskanen, Elena Simona Lohan, Aleksandr Ometov, Jorge Morte Palacios, Jari Nurmi, Ville Kaseva
IEEE Internet Things J.2
2025 Parallel Accurate Minifloat MACCs for Neural Network Inference on Versal FPGAs
abstract
Machine learning (ML) is ubiquitous in contemporary applications. Its need for efficient acceleration has driven vast research efforts into the quantization of neural networks with low-precision numerical formats. Models quantized with minifloat formats of eight or fewer bits have proven capable of outperforming models quantized into same-size integers. However, unlike integers, minifloats require accurate accumulation to prevent the introduction of rounding errors. We explore the design space of parallel accurate minifloat multiply-accumulators (MACCs) targeting the AMD VersalTM FPGA fabric. We experiment with three variations of the multiply-and-shift and adder tree components of a minifloat MACC. For comparison, we apply similar alterations to a parallel integer MACC. Our results show that custom compressor trees with external sign-inversion gates reduce the mean area of the minifloat MACCs by 17.7% and increase their clock frequency by 16.2%. In comparison, custom compressor trees with absorbed partial product generation gates reduce the mean area of integer MACCs by 28.1% and increase their clock frequency by 3.60%. Comparing the best-performing designs, we observe that minifloat MACCs consume 20% to 180% more resources than integer ones with same-size operands without accounting for a conversion back into a floating-point format, and 60% to 300% more resources when including it. Our data enable engineers to make informed decisions in their designs of deeply integrated embedded ML solutions when trading off training and fine-tuning effort versus resource cost.
Hans Jakob Damsgaard, Konstantin Hoßfeld, Jari Nurmi, Thomas B. Preußer
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 Hardware Generators with Chisel
abstract
Most digital hardware is described in hardware description languages, such as VHDL and (System)Verilog. These languages provide limited programming models for hardware construction despite receiving regular updates and extensions. Chisel defines itself as a hardware construction language, which means it shall permit more than the mere description of digital circuits. However, programmatic hardware generation is not new. Scripting languages like Perl generate VHDL or Verilog code from sources like Excel spreadsheets. Chisel, embedded in the general-purpose language Scala, lends itself to writing hardware generators in that language. We consider this Chisel-Scala ecosystem an ideal starting point for programming hardware generators and illustrate this point with examples using various programming models. We are confident that proven technologies from the software development world can be leveraged in the hardware design domain to improve hardware designers' productivity to build the next billion transistor chips.
Martin Schoeberl, Hans Jakob Damsgaard, Luca Pezzarossa, Oliver Keszöcze, Erling Rennemo Jellum
DSD2
2024 Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
abstract
Post-training quantization (PTQ) is a powerful technique for model compression, reducing the numerical precision in neural networks without additional training overhead. Recent works have investigated adopting 8 -bit floating-point formats (FP8) in the context of PTQ for model inference. However, floating-point formats smaller than 8 bits and their relative comparison in terms of accuracy-hardware cost with integers remains unexplored on FPGAs. In this work, we present minifloats, which are reduced-precision floating-point formats capable of further reducing the memory footprint, latency, and energy cost of a model while approaching full-precision model accuracy. We implement a custom FPGA-based multiply-accumulate operator library and explore the vast design space, comparing minifloat and integer representations across 3 to 8 bits for both weights and activations. We also examine the applicability of various integer-based quantization techniques to minifloats. Our experiments show that minifloats offer a promising alternative for emerging workloads such as vision transformers.
Shivam Aggarwal, Hans Jakob Damsgaard, Alessandro Pappalardo, Giuseppe Franco, Thomas B. Preußer, Michaela Blott, Tulika Mitra
FPL2
2024 Adaptive approximate computing in edge AI and IoT applications: A review
abstract
Recent advancements in hardware and software systems have been driven by the deployment of emerging smart health and mobility applications. These developments have modernized the traditional approaches by replacing conventional computing systems with cyber-physical and intelligent systems combining the Internet of Things (IoT) with Edge Artificial Intelligence. Despite the many advantages and opportunities of these systems within various application domains, the scarcity of energy, extensive computing needs, and limited communication must be considered when orchestrating their deployment. Inducing savings in these directions is central to the Approximate Computing (AxC) paradigm, in which the accuracy of some operations is traded off with energy, latency, and/or communication reductions. Unfortunately, the dynamics of the environments in which AxC-equipped IoT systems operate have been paid little attention. We bridge this gap by surveying adaptive AxC techniques applied to three emerging application domains, namely autonomous driving, smart sensing and wearables, and positioning, paying special attention to hardware acceleration. We discuss the challenges of such applications, how adaptive AxC can aid their deployment, and which savings it can bring based on traits of the data and devices involved. Insights arising thereof may serve as inspiration to researchers, engineers, and students active within the considered domains.
Hans Jakob Damsgaard, Antoine Grenier, Dewant Katare, Zain Taufique, Salar Shakibhamedan, Tiago Troccoli, Georgios Chatzitsompanis, Anil Kanduri, Aleksandr Ometov, Aaron Yi Ding, Nima Taherinejad, Georgios Karakonstantis, Roger F. Woods, Jari Nurmi
J. Syst. Archit.1
2024 High-efficiency Compressor Trees for Latest AMD FPGAs
abstract
High-fan-in dot product computations are ubiquitous in highly relevant application domains, such as signal processing and machine learning. Particularly, the diverse set of data formats used in machine learning poses a challenge for flexible efficient design solutions. Ideally, a dot product summation is composed from a carry-free compressor tree followed by a terminal carry-propagate addition. On FPGA, these compressor trees are constructed from generalized parallel counters whose architecture is closely tied to the underlying reconfigurable fabric. This work reviews known counter designs and proposes new ones in the context of the new AMD Versal™ fabric. On this basis, we develop a compressor generator featuring variable-sized counters, novel counter composition heuristics, explicit clustering strategies, and case-specific optimizations like logic gate absorption. In comparison to the Vivado™ default implementation, the combination of such a compressor with a novel, highly efficient quaternary adder reduces the LUT footprint across different bit matrix input shapes by 45% for a plain summation and by 46% for a terminal accumulation at a slight cost in critical path delay still allowing an operation well above 500 MHz. We demonstrate the aptness of our solution at examples of low-precision integer dot product accumulation units.
Konstantin Hoßfeld, Hans Jakob Damsgaard, Jari Nurmi, Michaela Blott, Thomas B. Preußer
ACM Trans. Reconfigurable Technol. Syst.2
2023 Generating CGRA Processing Element Hardware with CGRAgen
abstract
The popularity of the Internet of Things and next-generation wireless networks calls for a greater distribution of small but high-performance and energy-efficient compute devices at the networks' Edge. These devices must integrate hardware acceleration to meet the latency requirements of relevant use cases. Existing work has highlighted Coarse-Grained Reconfigurable Arrays (CGRAs) as suitable compute architectures for this purpose. However, like other modern hardware design, research and design space exploration into CGRAs is hindered by long development times needed for Register Transfer Level implementation. In this paper, we propose mitigating these by extending the open-source CGRAgen tool with a Chisel-based hardware backend capable of transforming abstract Processing Element (PE) descriptions into synthesizable Verilog code. We present how CGRAgen's internal module representation is transformed to Chisel modules and demonstrate this on a selection of PE architectures from the literature. Finally, we outline future work on extending this flow to generate entire CGRAs.
Hans Jakob Damsgaard, Aleksandr Ometov, Jari Nurmi
DSD1
2023 Towards Coarse-Grained Reconfigurable Approximate Computing with CGRAgen
abstract
Modern Edge Computing devices execute applications that must meet strict latency requirements as per traditional standardization activities. Achieving the needed performance implies a need for efficiency in all aspects, thus, flexible solutions are needed. In this Ph.D. project, we address this issue for error-tolerant applications by using Coarse-Grained Reconfigurable Arrays (CGRAs) enriched with Approximate Computing (AxC) features. To do so, we aim to develop a CGRA architecture modeling, mapping, and hardware generation flow complete with AxC hardware primitives and significance analysis.
Hans Jakob Damsgaard, Aleksandr Ometov, Jari Nurmi
FPL1
2023 Approximate computing in B5G and 6G wireless systems: A survey and future outlook
abstract
As modern 5G systems are being deployed, researchers question whether they are sufficient for the oncoming decades of technological evolution.Growing numbers of interconnected intelligent devices put these networks under tremendous pressure, demanding their development.Paving the way for beyond 5G and 6G systems, commonly denoted by B5G herein, therefore means seeking enablers to increase efficiency from different perspectives.One novel look on this is the application of inexact computations where nine 9s reliability is not needed, for example, in non-critical mobile broadband traffic.The paradigm of Approximate Computing (AxC) focuses on such areas where constrained quality degradation results in savings that benefit the users and operators.This paper surveys the state-of-the-art publications on the intersection of AxC and B5G systems, identifying and emphasizing trends and tendencies in existing work and directions for future research.The work highlights resource allocation algorithms as particularly mesmerizing in the former, while research related to Intelligent Reflective Surfaces appears the most prominent in the latter.In both, problems are often NP-hard and, thus, only solvable using heuristics or approximations, Successive Convex Approximation and Reinforcement Learning are most frequently applied.
Hans Jakob Damsgaard, Aleksandr Ometov, Md. Munjure Mowla, Adam Flizikowski, Jari Nurmi
Comput. Networks1
2022 Enabling Coverage-Based Verification in Chisel
abstract
Ever-increasing performance demands are pushing hardware designers towards designing domain-specific accelerators. This has created a demand for improving the overall efficiency of the hardware design and verification cycles. The design efficiency was improved with the introduction of Chisel. However, verification efficiency has yet to be tackled. One method that can increase verification efficiency is the use of various types of coverage measures. In this paper, we present our open-source, coverage-related verification tools targeting digital designs described in Chisel. Specifically, we have created a new method allowing for statement coverage at an intermediate representation of Chisel, and several methods for gathering functional coverage directly on a Chisel description.
Andrew Dobis, Hans Jakob Damsgaard, Enrico Tolotto, Kasper Juul Hesse Rasmussen, Tjark Petersen, Martin Schoeberl
ETS2
2022 Comparing timed-division multiplexing and best-effort networks-on-chip
abstract
Best-effort (BE) networks-on-chips (NOCs) are usually preferred over time-division multiplexed (TDM) NOCs in multi-core platforms because they are work-conserving and have lower (zero-load) latency. On the other hand, BE NOCs are significantly more expensive to implement than TDM NOCs because of their virtual channel buffers, allocators/arbiters, and (credit-based) flow control; functionality that a TDM NOC avoids altogether. The objective of this paper is to compare the performance of BE and TDM NOCs, taking hardware cost into consideration. The networks are compared using graphs showing average latency as a function of offered load. For the BE NOCs, we use the BookSim simulator, and for the TDM NOCs, we derive a queuing theory model and an associated TDM NOC simulator. Through experiments with both router architectures, packet length, link width, and different traffic patterns, we show that for the same hardware cost, a TDM NOC can provide higher bandwidth and comparable latency. We also show that the packet length is the most important factor affecting the TDM period, which again is the primary factor affecting latency. The best TDM NOC design for BE traffic uses single flit packets, wide links/flits, and a router with two pipeline stages: link and router traversal.
Jens Sparsø, Hans Jakob Damsgaard, Dimitrios Katsamanis, Martin Schoeberl
J. Syst. Archit.2