Dania Susanne Mosuli

dblp:397/7236 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0005-8821-3290ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Electronic design automation · 67% High-performance computing · 33%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation › design automation tools › FPGA CAD
FPGA design tools
1.012026
A Hierarchical Methodology for Hardware Design Comparison in HPC Workloads · FPGA 2026
Electronic design automation
high-level synthesis
1.012026
A Hierarchical Methodology for Hardware Design Comparison in HPC Workloads · FPGA 2026
YearPublicationVenuePosition
2026 A Hierarchical Methodology for Hardware Design Comparison in HPC Workloads
abstract
As Moore's law slows down, developers face difficult choices between low-level HDLs (Verilog, VHDL) offering fine-grained control and higher-level tools (HLS and Chisel) promising improved productivity. While high-level tools accelerate development, performance gaps persist compared to expert HDL implementations. Prior studies emphasize end-to-end performance, offering limited insight into why tools excel or where performance diverges in the design hierarchy. We introduce a hierarchical framework for comparing hardware generation tools by decomposing HPC kernels (FFT, GEMM, QR factorization) into reusable primitives (MAC arrays, butterflies, permutations, reduction trees). Across Verilog, Chisel, and Vivado HLS, we built an automated tool flow and synthesized ~1, 500 variants on AMD Alveo U250, measuring resource utilization and frequency. We derived theoretical bounds for validation. Verilog achieves the highest frequency and lowest resource usage; Chisel performs comparably (5--15% gap), while HLS shows a 20--40% gap. All tools operate within bounds for well-structured designs. Crucially, performance divergence arises during primitive assembly, indicating that high-level tools require better composition optimization. This reproducible framework provides actionable insights and is extensible to other tools, domains, and FPGA architectures.
Doru-Thom Popovici, Mario Vega, Angelos Ioannou, Fabien Chaix, Dania Susanne Mosuli, Blair Reasoner, Tan Nguyen 0001, Xiaokun Yang, John Shalf
FPGA5
2024 Hardware Generation on Trigonometric Functions
abstract
This paper presents hardware generation for accelerating various floating-point (FP) trigonometric functions, including sine, cosine, and arctangent. The Chisel Hardware Construction Language (HCL) is used to develop parameterized and flexible designs for these functions. The hardware generator supports multiple design architectures with configurable parameters, such as precision (e.g., 16-bit, 32-bit, 64-bit, and 128-bit), iteration count, and pipeline depth, allowing for customization of hardware resource utilization, latency, speed, and accuracy.
Paul Wong, Dania Susanne Mosuli, Xuechen Zhang 0001, Xiaokun Yang
IEEE Big Data2