EDBT 2026 Demo / reviewers in the wild / expert
Francesc Moll
dblp:27/5846 · also Francesc Moll Echeto
· DBLP profile ↗
20ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0002-1290-3253ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 9 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GAVINA: flexible aggressive undervolting for bit-serial mixed-precision DNN accelerationabstractVoltage overscaling, or undervolting, is an enticing approximate technique in the context of energy-efficient Deep Neural Network (DNN) acceleration, given the quadratic relationship between power and voltage. Nevertheless, its very high error rate has thwarted its general adoption. Moreover, recent undervolting accelerators rely on 8-bit arithmetic and cannot compete with state-of-the-art low-precision (<8b) architectures. To overcome these issues, we propose a new technique called Guarded Aggressive underVolting (GAV), which combines the ideas of undervolting and bit-serial computation to create a flexible approximation method based on aggressively lowering the supply voltage on a select number of least significant bit combinations. Based on this idea, we implement GAVINA (GAV mIxed-Precision Accelerator), a novel architecture that supports arbitrary mixed precision and flexible undervolting, with an energy efficiency of up to 89 TOP/sW in its most aggressive configuration. By developing an error model of GAVINA, we show that GAV can achieve an energy efficiency boost of 20% via undervolting, with negligible accuracy degradation on ResNet-18. Jordi Fornt, Pau Fontova, Adrian Gras, Omar Lahyani, Martí Caro, Jaume Abella 0001, Francesc Moll, Josep Altet |
ISLPED | 7 |
| 2025 | Mix-GEMM: Extending RISC-V CPUs for Energy-Efficient Mixed-Precision DNN Inference Using Binary SegmentationabstractEfficiently computing Deep Neural Networks (DNNs) has become a primary challenge in today's computers, especially on devices targeting mobile or edge applications. Recent progress on Post-Training Quantization (PTQ) and Quantization-Aware Training (QAT) has shown that the key to high energy efficiency lies in executing deep learning models with low- (8- to 5-bit) or ultra-low-precision (4- to 2-bit). Unfortunately, current Central Processing Unit (CPU) architectures and Instruction Set Architectures (ISAs) present severe limitations on the range of data sizes supported to compute DNN kernels. In this work, we presentMix-GEMM, a hardware-software co-designed architecture that enables RISC-V processors to efficiently compute arbitrary mixed-precision DNN kernels, supporting all data size combinations from 8- to 2-bit. By applyingbinary segmentation, our architecture can scale its throughput by decreasing the data size of the operands, resulting in a flexible approach capable of leveraging state-of-the-art QAT and PTQ to achieve high energy efficiency at a very low cost. Evaluating ourMix-GEMMarchitecture in a dual-issue in-order RISC-V processor shows that we are able to boost its performance and energy efficiency by up to$44\times$and$11\times$with respect to the baseline processor, with an area overhead of only 2%. This allows our extended processor to execute state-of-the-art DNNs with significantly higher performance and energy efficiency than the standard FP32 precision, while retaining almost the same model accuracy. Jordi Fornt, Enrico Reggiani, Pau Fontova, Narcís Rodas, Alessandro Pappalardo, Osman S. Unsal, Adrián Cristal, Josep Altet, Francesc Moll, Jaume Abella 0001 |
IEEE Trans. Computers | 9 |
| 2024 | QUETZAL: Vector Acceleration Framework for Modern Genome Sequence Analysis AlgorithmsabstractGenome sequence analysis is fundamental to medical breakthroughs such as developing vaccines, enabling genome editing, and facilitating personalized medicine. The exponentially expanding sequencing datasets and complexity of sequencing algorithms necessitate performance enhancements. While the performance of software solutions is constrained by their underlying hardware platforms, the utility of fixed-function accelerators is restricted to only certain sequencing algorithms.This paper presents QUETZAL, the first general-purpose vector acceleration framework designed for high efficiency and broad applicability across a diverse set of genomics algorithms. While a commercial CPU’s vector datapath is a promising candidate to exploit the data-level parallelism in genomics algorithms, our analysis finds that its performance is often limited due to long-latency scatter/gather memory instructions. QUETZAL introduces a hardware-software co-design comprising an accelerator microarchitecture closely integrated with the CPU’s vector datapath, alongside novel vector instructions to fully capitalize on the proposed hardware. QUETZAL integrates a set of scratchpad-style buffers meticulously designed to minimize latency associated with scatter/gather instructions during the retrieval of input genome sequences data. QUETZAL supports both short and long reads, and different types of sequencing data formats. A combination of hardware and software techniques enables QUETZAL to reduce the latency of memory instructions, perform complex computation using a single instruction, and transform data representations at runtime, resulting in overall efficiency gain. QUETZAL significantly accelerates a vectorized CPU baseline on modern genome sequence analysis algorithms by 5.7×, while incurring a small area overhead of 1.4% post place-and-route at the 7nm technology node compared to an HPC ARM CPU. Julian Pavon, Iván Vargas Valdivieso, Carlos Rojas 0001, César Hernández, Mehmet Aslan, Roger Figueras, Yichao Yuan, Joël Lindegger, Mohammed Alser, Francesc Moll, Santiago Marco-Sola, Oguz Ergin, Nishil Talati, Onur Mutlu, Osman S. Unsal, Mateo Valero, Adrián Cristal |
ISCA | 10 |
| 2024 | Energy and relevance-aware adaptive monitoring method for wireless sensor nodes with hard energy constraintsabstractTraditional dynamic energy management methods optimize the energy usage in wireless sensor nodes adjusting their behavior to the operating conditions. However, this comes at the cost of losing the predictability in the operation of the sensor nodes. This loss of predictability is particularly problematic for the battery life, as it determines when the nodes need to be serviced. In this paper, we propose an energy and relevance-aware monitoring method, which leverages the principles of self-awareness to address this challenge. On one hand, the relevance-aware behavior optimizes how the monitoring efforts are allocated to maximize the monitoring accuracy; while on the other hand, the power-aware behavior adjusts the overall energy consumption of the node to achieve the target battery life. The proposed method is able to balance both behaviors so as to achieve the target battery life, at the same time is able to exploit variations in the collected data to maximize the monitoring accuracy. Furthermore, the proposed method coordinates two different adaptive schemes, a dynamic sampling period scheme, and a dual prediction scheme, to adjust the behavior of the sensor node. The evaluation results show that the proposed method consistently meets its battery lifetime goal, even when the operating conditions are artificially changed, and is able to improve the mean square error of the collected signal by up to 20% with respect to the same method with the relevance-aware behavior disabled, and of up to 16% with respect the same algorithm with just the adaptive sampling period or the dual prediction scheme enabled. Consequently showing the ability of the proposed method of making appropriate decisions to balance the competing interest of its two behaviors and coordinate the two adaptive schemes to improve their performance. David Arnaiz, Francesc Moll, Eduard Alarcón, Xavier Vilajosana |
Integr. | 2 |
| 2023 | Relating Context and Self Awareness in the Internet of Things
David Arnaiz, Marc Vila 0001, Eduard Alarcón, Francesc Moll, Maria-Ribera Sancho, Ernest Teniente |
CoopIS | 4 |
| 2023 | VAQUERO: A Scratchpad-based Vector Accelerator for Query ProcessingabstractDatabase Management Systems (DBMS) have be-come an essential tool for industry and research and are often a significant component of data centers. There have been many efforts to accelerate DBMS application performance. One of the most explored techniques is the use of vector processing. Unfortunately, conventional vector architectures have not been able to exploit the full potential of DBMS acceleration.In this paper, we present VAQUERO, our Scratchpad-based Vector Accelerator for QUEry pROcessing. VAQUERO improves the efficiency of vector architectures for DBMS operations such as data aggregation and hash joins featuring lookup tables. Lookup tables are significant contributors to the performance bottlenecks in DBMS processing suffering from insufficient ISA support in the form of scatter-gather instructions. VAQUERO introduces a novel Advanced Scratchpad Memory specifically designed with two mapping modes — direct- and associative-mode. These map-ping modes enable VAQUERO to accelerate real-world databases with workload sizes that significantly exceed the scratchpad memory capacity. Additionally, the associative-mode allows to use VAQUERO with DBMS operators that use hashed keys, e.g. hash-join and hash-aggregate. VAQUERO has been designed considering general DBMS algorithm requirements instead of being based on a particular database organization. For this reason, VAQUERO is capable to accelerate DBMS operators for both row- and column-oriented databases.In this paper, we evaluate the efficiency of VAQUERO using two highly optimized popular open-source DBMS, namely the row-based PostgreSQL and column-based MonetDB. We imple-mented VAQUERO at the RTL level and prototype it, by performing Place&Route, at the 7nm technology node. VAQUERO incurs a modest 0.15% area overhead compared with an Intel Ice Lake processor. Our evaluation shows that VAQUERO significantly outperforms PostgreSQL and MonetDB by 2.09× and 3.32× respectively, when processing operators and queries from the TPC-H benchmark. Julian Pavon, Iván Vargas Valdivieso, Joan Marimon, Roger Figueras, Francesc Moll, Osman S. Unsal, Mateo Valero, Adrián Cristal |
HPCA | 5 |
| 2023 | An automotive case study on the limits of approximation for object detection
Martí Caro, Hamid Tabani, Jaume Abella 0001, Francesc Moll, Enric Morancho, Ramon Canal, Josep Altet, Antonio Calomarde, Francisco J. Cazorla, Antonio Rubio 0001, Pau Fontova, Jordi Fornt |
J. Syst. Archit. | 4 |
| 2023 | Vitruvius+: An Area-Efficient RISC-V Decoupled Vector Coprocessor for High Performance Computing ApplicationsabstractThe maturity level of RISC-V and the availability of domain-specific instruction set extensions, like vector processing, make RISC-V a good candidate for supporting the integration of specialized hardware in processor cores for the High Performance Computing (HPC) application domain. In this article, 1 we present Vitruvius+, the vector processing acceleration engine that represents the core of vector instruction execution in the HPC challenge that comes within the EuroHPC initiative. It implements the RISC-V vector extension (RVV) 0.7.1 and can be easily connected to a scalar core using the Open Vector Interface standard. Vitruvius+ natively supports long vectors: 256 double precision floating-point elements in a single vector register. It is composed of a set of identical vector pipelines (lanes), each containing a slice of the Vector Register File and functional units (one integer, one floating point). The vector instruction execution scheme is hybrid in-order/out-of-order and is supported by register renaming and arithmetic/memory instruction decoupling. On a stand-alone synthesis, Vitruvius+ reaches a maximum frequency of 1.4 GHz in typical conditions (TT/0.80V/25°C) using GlobalFoundries 22FDX FD-SOI. The silicon implementation has a total area of 1.3 mm 2 and maximum estimated power of ∼920 mW for one instance of Vitruvius+ equipped with eight vector lanes. Francesco Minervini, Oscar Palomar, Osman S. Unsal, Enrico Reggiani, Josue V. Quiroga, Joan Marimon, Carlos Rojas 0001, Roger Figueras, Abraham Ruiz, Alberto González 0004, Jonnatan Mendoza, Iván Vargas 0001, César Hernández, Joan Cabre, Lina Khoirunisya, Mustapha Bouhali, Julian Pavon, Francesc Moll, Mauro Olivieri, Mario Kovac, Mate Kovac, Leon Dragic, Mateo Valero, Adrián Cristal |
ACM Trans. Archit. Code Optim. | 18 |
| 2023 | An Energy-Efficient GeMM-Based Convolution Accelerator With On-the-Fly im2colabstractSystolic array architectures have recently emerged as successful accelerators for deep convolutional neural network (CNN) inference. Such architectures can be used to efficiently execute general matrix–matrix multiplications (GeMMs), but computing convolutions with this primitive involves transforming the 3-D input tensor into an equivalent matrix, which can lead to an inflation of the input data, increasing the OFF-chip memory traffic which is critical for energy efficiency. In this work, we propose a GeMM-based systolic array accelerator that uses a novel data feeder architecture to perform ON-chip, on-the-fly convolution lowering (also known as im2col), supporting arbitrary tensor and kernel sizes as well as strided and dilated (or atrous) convolutions. By using our data feeder, we reduce memory transactions and required bandwidth on state-of-the-art CNNs by a factor of two, while only adding an area and power overhead of 4% and 7%, respectively. Application specific integrated circuit (ASIC) implementation of our accelerator in 22-nm technology fits in less than 1.1 mm 2 and reaches an energy efficiency of 1.10 TFLOP/sW with 16-bit floating-point arithmetic. Jordi Fornt, Pau Fontova, Martí Caro, Jaume Abella 0001, Francesc Moll, Josep Altet, Christoph Studer |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2021 | VIA: A Smart Scratchpad for Vector Units with Application to Sparse Matrix ComputationsabstractSparse matrix operations are critical kernels in multiple application domains such as High Performance Computing, artificial intelligence and big data. Vector processing is widely used to improve performance on mathematical kernels with dense matrices. Unfortunately, existing vector architectures do not cope well with sparse matrix computations, achieving much lower performance in comparison with their dense counterparts.To overcome this limitation, we present the Vector Indexed Architecture (VIA), a novel hardware vector architecture that accelerates applications with irregular memory access patterns such as sparse matrix computations. There are two main bottlenecks when computing with sparse matrices: irregular memory accesses and index matching. VIA addresses these two bottlenecks with a smart scratchpad that is tightly coupled to the Vector Functional Units within the core.Thanks to this structure, VIA improves locality for sparse-dense computations and improves the index matching search process for sparse computations. As a result, VIA achieves significant performance speedup over highly optimized state-of-the-art C++ algebra libraries. On average, VIA outperforms sparse matrix vector multiplication, sparse matrix addition and sparse matrix matrix multiplication kernels by 4.22 ×, 6.14 × and 6.00 ×, respectively, when evaluated over a thousand sparse matrices that arise in real applications. In addition, we prove the generality of VIA by showing that it can accelerate histogram and stencil applications by 4.5 × and 3.5 ×, respectively. Julian Pavon, Iván Vargas Valdivieso, Adrián Barredo, Joan Marimon, Miquel Moretó, Francesc Moll, Osman S. Unsal, Mateo Valero, Adrián Cristal |
HPCA | 6 |
| 2020 | Mechanical Energy Harvesting Taxonomy for Industrial Environments: Application to the Railway IndustryabstractTraditional industry is experiencing a worldwide evolution with Industry 4.0. Wireless sensor networks (WSNs) have a main role in this evolution as an essential part of data acquisition. The way in which WSNs are powered is one of the main challenges to face if industry wants to achieve the digital transformation. Energy harvesting technologies are one of the possible solutions to this challenge. The main purpose of this paper is to present a novel method to taxonomize knowledge in the field of mechanical energy harvesting to enhance the use of energy harvesting technologies in industrial applications. The methodology is based on the analysis of key parameters and performance metrics for existing technologies. The taxonomy is applied to rail axles in order to select the energy harvesting technology that is more appropriate for this specific location, demonstrating the potential of mechanical energy harvesting technologies (MEHTs) for the railway industry, as a use case of industrial environment. In addition, the taxonomy allows to identify the upcoming challenges for research purposes while analyzing the compatibility among mechanical energy harvesting technologies in order to create hybrid harvesters. Pablo López Díez, Iosu Gabilondo, Eduard Alarcón, Francesc Moll |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2018 | A Comprehensive Method to Taxonomize Mechanical Energy Harvesting TechnologiesabstractTraditional industry is experiencing a worldwide development with Industry 4.0. Wireless sensor networks (WSNs) have a main role in this revolution as an essential part of data acquisition. The way in which WSNs are powered is one of the main challenges to face if Industry wants to achieve the digital transformation. Energy harvesting taechnologies are one of the possible solutions to this challenge. The main purpose of this paper is to present a novel method to taxonomize knowledge in the field of mechanical energy harvesting in order to enhance the use of energy harvesting technologies for industrial applications. Additionally, the taxonomy allows to identify upcoming challenges for research purposes. Pablo López Díez, Iosu Gabilondo, Eduard Alarcón, Francesc Moll |
ISCAS | 4 |
| 2017 | An on-line test strategy and analysis for a 1T1R crossbar memoryabstractMemristors are emerging devices known by their nonvolability, compatibility with CMOS processes and high density in circuits density in circuits mostly owing to the crossbar nanoarchitecture. One of their most notable applications is in the memory system field. Despite their promising characteristics and the advancements in this emerging technology, variability and reliability are still principal issues for memristors. For these reasons, exploring techniques that check the integrity of circuits is of primary importance. Therefore, this paper proposes a method to perform an on-line test capable to detect a single failure inside the memory crossbar array. Manuel Escudero-Lopez, Francesc Moll, Antonio Rubio 0001, Ioannis Vourkas |
IOLTS | 2 |
| 2017 | Insights Into Tunnel FET-Based Charge Pumps and Rectifiers for Energy Harvesting ApplicationsabstractIn this paper, the electrical characteristics of tunnel field-effect transistor (TFET) devices are explored for energy harvesting front-end circuits with ultralow power consumption. Compared with conventional thermionic technologies, the improved electrical characteristics of TFET devices are expected to increase the power conversion efficiency of front-end charge pumps and rectifiers powered at sub-μW power levels. However, under reverse bias conditions the TFET device presents particular electrical characteristics due to its different carrier injection mechanism. In this paper, it is shown that reverse losses in TFET-based circuits can be attenuated by changing the gate-to-source voltage of reverse-biased TFETs. Therefore, in order to take full advantage of the TFETs in front-end energy harvesting circuits, different circuit approaches are required. In this paper, we propose and discuss different topologies for TFET-based charge pumps and rectifiers for energy harvesting applications. David Cavalheiro, Francesc Moll, Stanimir Stoyanov Valtchev |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | ASIC Implementation of An All-digital Self-adaptive PVTA Variation-aware Clock Generation SystemabstractAn all-digital self-adaptive clock generation system capable of autonomously adapt the clock frequency to compensate the effects of static spatially heterogeneous (SSHet) PVTA variations is presented. The design uses time-to-digital converters (TDCs) as delay sensors and a variable length ring oscillator (VLRO) as clock generator. The VLRO naturally adapts its frequency to the PVTA variations suffered by its logic gates while the TDCs are used to track these variations across the chip and modify the VLRO length in order to allocate them. The proposed system has been implemented in a silicon chip using a 65nm process. The fabricated chip has been used to test the system adaptive capabilities under SSHet voltage variations. Measurement results show that it effectively adapts the VLRO length, and hence the clock frequency, to the supply voltage variations. Jordi Perez-Puigdemont, Francesc Moll |
ACM Great Lakes Symposium on VLSI | 2 |
| 2016 | Feasibility of Embedded DRAM Cells on FinFET TechnologyabstractIn this paper, we analyze the suitability of implementing embedded DRAM (eDRAM) cells on FinFET technology compared to classical planar MOSFETs. The results show a significant improvement in overall cell performance for multi-gate devices. While pFinFET-based memories showed better cell behavior and variability robustness, mixed n/pFinFET cells had the highest working frequency and a negligible impact on degradation. Finally, we show that a multiple fin-height strategy can be used to reduce the layout area of the eDRAM cells (>10%). Esteve Amat, Antonio Calomarde, Francesc Moll, Ramon Canal, Antonio Rubio 0001 |
IEEE Trans. Computers | 3 |
| 2015 | Pespectives of TFET devices in ultra-low power charge pumps for thermo-electric energy sourcesabstractThe superior electrical characteristics of the heterojunction III-V Tunnel FET (TFET) devices can outperform current technologies in the process of energy harvesting conversion at ultra-low power supply voltage operation (sub-0.25 V). In this work, it is shown by simulations that a cross-coupled switched-capacitor topology with GaSb-InAs TFET devices present better conversion performance compared to the use of Si FinFET technology at low temperature variations (ΔT <; 3 °C) when considering a thermo-electric energy harvesting source (with α = 80 mV/K). At higher ΔT, the conversion process is degraded with the increase of the transistor losses. Considering a ΔT of 1 °C (2 °C), one cross-coupled stage with TFET devices can achieve 74 % (69 %) of power conversion efficiency when considering an output load of 0.4 μA (6 μA). At the same conditions, the FinFET charge pump is shown inefficient. David Cavalheiro, Francesc Moll, Stanimir Stoyanov Valtchev |
ISCAS | 2 |
| 2014 | A Boolean Rule-Based Approach for Manufacturability-Aware Cell RoutingabstractAn approach for cell routing using gridded design rules is proposed. It is technology-independent and parameterizable for different fabrics and design rules, including support for multiple-patterning lithography. The core contribution is a detailed-routing algorithm based on a Boolean formulation of the problem. The algorithm uses a novel encoding scheme, graph theory to support floating terminals, efficient heuristics to reduce the computational cost, and minimization of the number of unconnected pins in case the cell is unroutable. The versatility of the algorithm is demonstrated by routing single- and double-height cells. The efficiency is ascertained by synthesizing a library with 127 cells in about one hour and a half of CPU time. The layouts derived by the implemented tool have also been compared with the ones from a commercial library; thus, showing the competitiveness of the approach for gridded geometries. Jordi Cortadella, Jordi Petit, Francesc Moll |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2013 | Design and implementation of an adaptive proactive reconfiguration technique for SRAM cachesabstractScaling of device dimensions toward nano-scale regime has made it essential to innovate novel design techniques for improving the circuit robustness. This work proposes an implementation of adaptive proactive reconfiguration methodology that can first monitor process variability and BTI aging among 6T SRAM memory cells and then apply a recovery mechanism to extend the SRAM lifetime. Our proposed technique can extend the memory lifetime between 2X to 4.5X times with a silicon area overhead of around 10% for the monitoring units, in a 1kB 6T SRAM memory chip. Peyman Pouyan, Esteve Amat, Francesc Moll, Antonio Rubio 0001 |
DATE | 3 |
| 2010 | VCTA: A Via-Configurable Transistor Array regular fabricabstractLayout regularity is introduced progressively by integrated circuit manufacturers to reduce the increasing systematic process variations in the deep sub-micron era. In this paper we focus on a scenario where layout regularity must be pushed to the limit to deal with severe systematic process variations in future technology nodes. With this objective, we propose and evaluate a new regular layout style called Via-Configurable Transistor Array (VCTA) that maximizes regularity at device and interconnect levels. In order to assess VCTA maximum layout regularity tradeoffs, we implement 32-bit adders in the 90 nm technology node for VCTA and compare them with implementations that make use of standard cells. For this purpose we study the impact of photolithography proximity and coma effects on channel length variations, and the impact of shallow trench isolation mechanical stress on threshold voltage variations. We demonstrate that both variations, that are important sources of energy and delay circuit variability, are minimized through VCTA regularity. Marc Pons 0001, Francesc Moll, Antonio Rubio 0001, Jaume Abella 0001, Xavier Vera, Antonio González 0001 |
VLSI-SoC | 2 |