EDBT 2026 Demo / reviewers in the wild / expert
Sonia Gonzalez-Navarro
dblp:49/107 · also Sonia González-Navarro
· DBLP profile ↗
17ranked-venue papers
4as first author
3since 2021 · last 2026
0000-0002-5636-7042ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 2 first-author · 3 since 2021Theory of computation · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CASTM: An API for Accelerating Zero-Knowledge Proof Kernels on CGRA Architectures
Cristian Campos, Angeles G. Navarro, Sonia Gonzalez-Navarro |
Euro-Par (1) | 3 |
| 2021 | Genome Sequence Alignment - Design Space Exploration for Optimal Performance and Energy ArchitecturesabstractNext generation workloads, such as genome sequencing, have an astounding impact in the healthcare sector. Sequence alignment, the first step in genome sequencing, has experienced recent breakthroughs, which resulted in next generation sequencing (NGS). As NGS applications are memory bounded with random memory access patterns, we propose the use of high bandwidth memories like 3D stacked HBM2, instead of traditional DRAMs like DDR4, along with energy efficient compute cores to improve both performance and energy efficiency. Three state-of-the-art NGS applications, Bowtie2, BWA-MEM, and HISAT2 are used as case studies to explore and optimize NGS computing architectures. Then, using the gem5-X architectural simulator, we obtain an overall 68 percent performance improvement and 71 percent energy savings using HBM2 instead of DDR4. Furthermore, we propose an architecture based on ARMv8 cores and demonstrate that 16 ARMv8 64-bit OoO cores with HBM2 outperforms 32-cores of Intel Xeon Phi Knights Landing (KNL) processor with 3D stacked memory. Moreover, we show that by using frequency scaling we can achieve up to 59 percent and 61 percent energy savings for ARM in-order and OoO cores, respectively. Lastly, we show that many ARMv8 in-order cores at 1.5GHz match the performance of fewer OoO cores at 2GHz, while attaining 4.5x energy savings. Yasir Mahmood Qureshi, Jose Manuel Herruzo, Marina Zapater, Katzalin Olcoz, Sonia Gonzalez-Navarro, Oscar G. Plata, David Atienza 0001 |
IEEE Trans. Computers | 5 |
| 2021 | Enabling fast and energy-efficient FM-index exact matching using processing-near-memory
Jose Manuel Herruzo, Ivan Fernandez, Sonia Gonzalez-Navarro, Oscar G. Plata |
J. Supercomput. | 3 |
| 2020 | Floating-Point Fused Multiply-Add under HUB FormatabstractThe Half-Unit-Biased (HUB) format has interesting advantages for implementing floating-point arithmetic which has been proved for the four basic arithmetic operations as well as square root. Nevertheless, although Floating-point Fused Multiply-add (FMA) operation (AxB + C) is one of the most important and complex arithmetic instructions in modern processors, FMA operation for HUB numbers has not been confronted yet. In this paper, we present a design to deal with this operation under HUB format. The key points to turn the conventional FMA architecture into a HUB unit are explained. Comparing the ASIC implementation of a HUB FMA unit with the conventional one, the former reduces the required area and power up to 38% and 35%, respectively, for single-precision. For BFloat16, the HUB FMA increases the speed a 15%, and even then, reduces the area and power by 26% and 12%, respectively. Javier Hormigo, Julio Villalba, Sonia Gonzalez-Navarro |
ARITH | 3 |
| 2020 | New Results on Non-Normalized Floating-Point FormatsabstractCompulsory normalization of the represented numbers is a key requirement of the floating-point standard. This requirement contributes to fundamental characteristics of the standard, such as taking the most of the precision available, reproducibility and facilitation of comparison and other operations. However, it also imposes a high restriction in effectiveness of basic arithmetic operation implementation. In many embedded applications may be worth to sacrifice the benefits of normalization for gaining in implementation metrics. This paper analyzes and measures the effect of removing the normalization requirement in terms of precision and implementation savings for embedded applications. We propose several adder and multiplier architectures to deal with non-normalized floating-point numbers, and quantify the accuracy loss and the improvements in hardware implementation. Our experiments show that it is possible to reduce the area and power consumption up to 78 percent in ASIC and 50 percent in FPGA implementations with a reasonable accuracy loss. Sonia Gonzalez-Navarro, Javier Hormigo |
IEEE Trans. Computers | 1 |
| 2020 | Accelerating Sequence Alignments Based on FM-Index Using the Intel KNL ProcessorabstractFM-index is a compact data structure suitable for fast matches of short reads to large reference genomes. The matching algorithm using this index exhibits irregular memory access patterns that cause frequent cache misses, resulting in a memory bound problem. This paper analyzes different FM-index versions presented in the literature, focusing on those computing aspects related to the data access. As a result of the analysis, we propose a new organization of FM-index that minimizes the demand for memory bandwidth, allowing a great improvement of performance on processors with high-bandwidth memory, such as the second-generation Intel Xeon Phi (Knights Landing, or KNL), integrating ultra high-bandwidth stacked memory technology. As the roofline model shows, our implementation reaches 95 percent of the peak random access bandwidth limit when executed on the KNL and almost all of the available bandwidth when executed on other Intel Xeon architectures with conventional DDR memory. In addition, the obtained throughput in KNL is much higher than the results reported for GPUs in the literature. Jose Manuel Herruzo, Sonia Gonzalez-Navarro, Pablo Ibáñez 0001, Víctor Viñals, Jesús Alastruey-Benedé, Oscar G. Plata |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2019 | Boosting Backward Search Throughput for FM-Index Using a Compressed EncodingabstractThe rapid development of DNA sequencing technologies has demanded for compressed data structures supporting fast pattern matching queries. FM-index is a widely-used compressed data structure that also supports fast pattern matching queries. It is common for the exact matching algorithm to be memory bound, resulting in poor performance. We propose a new data-layout of FM-index that compacts all data needed to perform the searching process. This results in an improvement of the search computing time for genomic data. Jose Manuel Herruzo, Sonia Gonzalez-Navarro, Pablo Ibáñez 0001, Víctor Viñals, Jesús Alastruey-Benedé, Oscar G. Plata |
DCC | 2 |
| 2018 | Unbiased Rounding for HUB Floating-Point AdditionabstractHalf-Unit-Biased (HUB) is an emerging format based on shifting the represented numbers by half Unit in the Last Place. This format simplifies two's complement and round-to-nearest operations by preventing any carry propagation. This saves power consumption, time and area. Taking into account that the IEEE floating-point standard uses an unbiased rounding as the default mode, this feature is also desirable for HUB approaches. In this paper, we study the unbiased rounding for HUB floating-point addition in both as standalone operation and within FMA. We show two different alternatives to eliminate the bias when rounding the sum results, either partially or totally. We also present an error analysis and the implementation results of the proposed architectures to help the designers to decide what their best option are. Julio Villalba, Javier Hormigo, Sonia Gonzalez-Navarro |
IEEE Trans. Computers | 3 |
| 2017 | Normalizing or Not Normalizing? An Open Question for Floating-Point Arithmetic in Embedded SystemsabstractEmerging embedded applications lack of a specific standard when they require floating-point arithmetic. In this situation they use the IEEE-754 standard or ad hoc variations of it. However, this standard was not designed for this purpose. This paper aims to open a debate to define a new extension of the standard to cover embedded applications. In this work, we only focus on the impact of not performing normalization. We show how eliminating the condition of normalized numbers, implementation costs can be dramatically reduced, at the expense of a moderate loss of accuracy. Several architectures to implement addition and multiplication for non-normalized numbers are proposed and analyzed. We show that a combined architecture (adder-multiplier) can halve the area and power consumption of its counterpart IEEE-754 architecture. This saving comes at the cost of reducing an average of about 10 dBs the Signal-to-Noise Ratio for the tested algorithms. We think these results should encourage researchers to perform further investigation in this issue. Sonia Gonzalez-Navarro, Javier Hormigo |
ARITH | 1 |
| 2016 | Decimal Multiformat Online AdditionabstractThis paper presents and analyzes two different strategies for designing multiformat online decimal adders ($\text{olDFA}_{\text{Mformat}}$). The first strategy uses a code conversion stage plus an online Decimal Full Adder (olDFA); the second one involves designing specific adders by modifying the architecture of the olDFA. These strategies are applied in the design of specific architectures to deal with financial analysis calculations. We use synthesis results to verify the theoretical aspects of the designs and to analyze the robustness and lacks of the strategies. The guidelines presented in the paper are valuable to designers of online multiformat-based solutions. Carlos Garcia-Vega, Sonia Gonzalez-Navarro, Pedro Balboa-La Chica, Julio Villalba |
IEEE Trans. Computers | 2 |
| 2013 | Binary Integer Decimal-Based Floating-Point MultiplicationabstractThis paper presents a multiplier that operates on binary integer decimal (BID) encoded decimal floating-point (DFP) numbers. It uses a single binary multiplier with carry-save feedback for both significand multiplication and rounding, and it is compliant with the IEEE 754-2008 Standard. Optimizations decrease the BID multiplier's area and critical path delay. Sonia Gonzalez-Navarro, Charles Tsen, Michael J. Schulte |
IEEE Trans. Computers | 1 |
| 2012 | On-line Decimal Adder with RBCD RepresentationabstractIn this paper we present the design of an on-line adder dealing with two RBCD numbers. This basic element is intended to be be used in any on-line system in which the addition is involved. We obtain the on-line adder by serialization of a recent parallel RBCD adder with minimum latency.To reduce the cycle time a pipelined version is proposed. To deal with data stream the throughput has been reduced to its theoretical minimum possible value by a negligible cost hardware modification. Finally, actual implementation results for 16 digits (i.e. decimal64 format) are presented. Carlos Garcia-Vega, Sonia Gonzalez-Navarro, Julio Villalba, Emilio L. Zapata |
ASAP | 2 |
| 2012 | A study of decimal left shifters for binary numbers
Sonia Gonzalez-Navarro, Javier Hormigo, Michael J. Schulte |
Inf. Comput. | 1 |
| 2011 | Hardware Designs for Binary Integer Decimal-Based RoundingabstractDecimal floating-point (DFP) arithmetic is becoming increasingly important and specifications for it are included in the revised IEEE 754 standard for floating-point arithmetic (IEEE 754-2008). The binary encoding of DFP numbers specified in IEEE 754-2008 is commonly referred to as Binary-Integer Decimal (BID). BID uses a binary integer to encode the significant, which allows it to leverage existing high-speed binary circuits. However, performing decimal rounding on these binary significant is challenging. In this paper, we propose and evaluate several approaches to perform decimal rounding in hardware for DFP numbers that use the BID encoding. We summarize several rounding techniques, present the theory and design of each proposed rounding unit, and use synthesis results to evaluate the critical path delay, latency, and area of rounding units for 64-bit BID numbers. Our results indicate that the bulk of each rounder design is occupied by a binary fixed-point multiplier that can be shared with other integer and floating-point operations. This is the first paper to present and compare a variety of techniques for BID-based rounding hardware. These techniques are valuable to designers of BID-based DFP solutions. Samuel Tsen, Sonia Gonzalez-Navarro, Michael J. Schulte, Katherine Compton |
IEEE Trans. Computers | 2 |
| 2009 | A Combined Decimal and Binary Floating-Point MultiplierabstractIn this paper, we describe the first hardware design of a combined binary and decimal floating-point multiplier, based on specifications in the IEEE 754-2008 floating-point standard. The multiplier design operates on either (1) 64-bit binary encoded decimal floating-point (DFP) numbers or (2) 64-bit binary floating-point (BFP) numbers. It returns properly rounded results for the rounding modes specified in IEEE 754-2008. The design shares the following hardware resources between the two floating-point datatypes: a 54-bit by 54-bit binary multiplier, portions of the operand encoding/decoding, a 54-bit right shifter, exponent calculation logic, and rounding logic. Our synthesis results show that hardware sharing is feasible and has a reasonable impact on area, latency, and delay. The combined BFP and DFP multiplier occupies only 58% of the total area that would be required by separate BFP and DFP units. Furthermore, the critical path delay of a combined multiplier has a negligible increase over a standalone DFP multiplier, without increasing the number of cycles to perform either BFP or DFP multiplication. Charles Tsen, Sonia Gonzalez-Navarro, Michael J. Schulte, Brian J. Hickmann, Katherine Compton |
ASAP | 2 |
| 2007 | Hardware Design of a Binary Integer Decimal-based IEEE P754 Rounding UnitabstractBecause of the growing importance of decimal floating-point (DFP) arithmetic, specifications for it were recently added to the draft revision of the IEEE 754 Standard (IEEE P754). In this paper, we present a hardware design for a rounding unit for 64-bit DFP numbers (decimal 64) that use the IEEE P754 binary encoding of DFP numbers, which is widely known as the Binary Integer Decimal (BID) encoding. We summarize the technique used for rounding, present the theory and design of the BID rounding unit, and evaluate its critical path delay, latency, and area for combinational and pipelined designs. Over 86% of the rounding unit's area is due to a 55-bit by 54-bit binary multiplier, which can be shared with a double-precision binary floating-point multiplier. To our knowledge, this is the first hardware design for rounding IEEE P754 BID-encoded DFP numbers. Charles Tsen, Michael J. Schulte, Sonia Gonzalez-Navarro |
ASAP | 3 |
| 2007 | Hardware design of a Binary Integer Decimal-based floating-point adderabstractBecause of the growing importance of decimal floating-point (DFP) arithmetic, specifications for it are included in the IEEE Draft Standard for Floating-point Arithmetic (IEEE P754). In this paper, we present a novel algorithm and hardware design for a DFP adder. The adder performs addition and subtraction on 64-bit operands that use the IEEE P754 binary encoding of DFP numbers, widely known as the binary integer decimal (BID) encoding. The BID adder uses a novel hardware component for decimal digit counting and an enhanced version of a previously published BID rounding unit. By adding more sophisticated control, operations are performed with variable latency to optimize for common cases. We show that a BID-based DFP adder design can be achieved with a modest area increase compared to a single 2-stage pipelined 64-bit fixed-point multiplier. Over 70% of the BID adderpsilas area is due the 64-bit fixed-point multiplier, which can be shared with a binary floating-point multiplier and hardware for other DFP operations. To our knowledge, this is the first hardware design for adding and subtracting IEEE P754 BID-encoded DFP numbers. Charles Tsen, Sonia Gonzalez-Navarro, Michael J. Schulte |
ICCD | 2 |