EDBT 2026 Demo / reviewers in the wild / expert
Guillermo Payá-Vayá
dblp:64/1540 · also Guillermo Payá Vayá
· DBLP profile ↗
25ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-3503-8386ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Software engineering, systems software and programming languages · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OpenFI4ASIC: An Open-Source Fault Injection Framework for ASIC Designs via FPGA-Based Rapid Prototyping
Jasper Homann, Eike Trumann, Umut Durak, Guillermo Payá-Vayá |
SAFECOMP | 4 |
| 2025 | DCMA: Accelerating Parallel DMA Transfers with a Multi-Port Direct Cached Memory Access in a Massive-Parallel Vector ProcessorabstractState-of-the-art applications, such as convolutional neural networks, demand specialized hardware accelerators that address performance and efficiency constraints. An efficient memory hierarchy is mandatory for such hardware systems. While the memory architectures of general-purpose processors (e.g., CPU or GPUs) are based on cache systems, dedicated accelerators have mostly adopted the DMA (Direct Memory Access) concept due to the application field of image processing. DMA features like 2D data transfers or data padding can optimize the memory accesses of image processing. However, DMA lacks the capability to exploit temporal and spatial data reuse, a feature common in cache systems, particularly when multiple DMAs operate in parallel. This article proposes a novel Direct Cached Memory Access (DCMA) architecture, combining both DMA and cache methodologies and their respective advantages. Optimized for image-based AI algorithms, the DCMA architecture facilitates enhanced memory access by integrating multiple, parallel DMA ports with caching capabilities. This design allows for efficient data reuse and parallel memory access. Optimal parameters for the DCMA are determined through a comprehensive design space exploration. The DCMA is evaluated on a state-of-the-art Xilinx UltraScale+ FPGA board coupled with a massive-parallel vertical vector co-processor, called V 2 PRO. The results show the mitigation of the vector processor’s memory bottleneck. By using the proposed DCMA, speedups of up to ×17 for the ResNet-50 CNN can be achieved. Gia Bao Thieu, Sven Gesper, Guillermo Payá-Vayá |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | An Open-Source NanoController v2 Featuring Microcoded Instruction Set Redefinition in 22-nm FDSOI-CMOS for Autonomous Ultralow-Power SoCsabstractThe realization of autonomous, wearable, and implantable system-on-chip (SoC) for health monitoring applications poses several challenges, such as achieving ultralow size, cost, and power consumption, yet offering sufficient flexibility to reprogram and adapt the autonomously operating SoC during the course of treatment. Commonly, programmability is not considered for ultralow-power (ULP) biomedical SoCs, since a dedicated finite state machine (FSM) fixes the operation sequence, and instruction memory presents significant contributions to silicon area and power consumption. Based on a previously published tiny, programmable microarchitecture, this work proposes the strongly enhancedNanoController v2, for potential use in ultralow-power biomedical SoCs, and integrates it as a prototype chip in a 22-nm FDSOI-CMOS technology. By implementing a novel microcoded control unit and an automated design space exploration framework, which are made available open-source, the instruction set can be freely redefined to exploit application-specific properties. The benefits are increased code compaction, performance gain, and, consequently, decreased power consumption. In an extensive measurement campaign, a glucose sensor control application achieves 13.1% higher performance and 15.6% less code size in the best case, resulting in an extremely low power consumption of 660 nW (9% less than the reference),only by a different instruction setwithout hardware changes. Compared with other state-of-the-art small programmable microcontrollers, between 38% and 82% smaller code size and between 33% and 77% smaller silicon area and averaged power consumption could be shown. Based on the prototype results, a fully integrated glucose sensor chip will be evaluated in currently ongoing work. Moritz Weißbrich, Adilet Dossanov, Yerzhan Kudabay, Alexander Meyer, Vadim Issakov, Guillermo Payá-Vayá |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | Multi-Level Prototyping of a Vertical Vector AI Processing SystemabstractModern embedded systems must be designed carefully to cope with the complexity and real-time requirements of modern AI (Artificial Intelligence) driven automotive applications, such as Advanced Driver-Assistance Systems (ADAS). Despite increasing complexity, the time to market is decreasing. In this work, a SystemC-based Virtual Prototype of a neural network processing platform is exploited to bypass the limitations of standalone instruction set simulators (ISS) and FPGA prototyping. The processing platform under test is based on a novel massive parallel vector processor architecture coupled with a RISC- V control core that runs widely used convolutional neural networks (CNNs) for object detection. The paper discusses the variations and appropriateness of the three prototyping methods outlined, demonstrating how the Virtual Prototype can address the aforementioned constraints, resulting in a 2.07x increase in accuracy, 16x greater configurations, and more profound insights into the system compared to standalone and FPGA prototyping. Frederik Kautz, Sven Gesper, Gia Bao Thieu, Hans-Martin Blüthgen, Holger Blume, Guillermo Payá-Vayá |
ASAP | 6 |
| 2023 | Exploiting Subword Permutations to Maximize CNN Compute Performance and EfficiencyabstractNeural networks (NNs) are quantized to decrease their computational demands and reduce their memory foot-print. However, specialized hardware is required that supports computations with low bit widths to take advantage of such optimizations. In this work, we propose permutations on subword level that build on top of multi-bit-width multiply-accumulate operations to effectively support low bit width computations of quantized NNs. By applying this technique, we extend the data reuse and further improve compute performance for convolution operations compared to simple vectorization using SIMD (single-instruction-multiple-data). We perform a design space exploration using a cycle accurate simulation with MobileNet and VGG16 on a vector-based processor. The results show a speedup of up to$3.7\times$and a reduction of up to$1.9\times$for required data transfers. Additionally, the control overhead for orchestrating the computation is decreased by up to$3.9\times$. Michael Beyer, Sven Gesper, Andre Guntoro, Guillermo Payá-Vayá, Holger Blume |
ASAP | 4 |
| 2023 | ZuSE Ki-Avf: Application-Specific AI Processor for Intelligent Sensor Signal Processing in Autonomous DrivingabstractModern and future AI-based automotive applications, such as autonomous driving, require the efficient real-time processing of huge amounts of data from different sensors, like camera, radar, and LiDAR. In the ZuSE-KI-AVF project, multiple university, and industry partners collaborate to develop a novel massive parallel processor architecture, based on a cus-tomized RISC-V host processor, and an efficient high-performance vertical vector coprocessor. In addition, a software development framework is also provided to efficiently program AI-based sensor processing applications. The proposed processor system was verified and evaluated on a state-of-the-art UltraScale+ FPGA board, reaching a processing performance of up to 126.9 FPS, while executing the YOLO-LITE CNN on 224x224 input images. Further optimizations of the FPGA design and the realization of the processor system on a 22nm FDSOI CMOS technology are planned. Gia Bao Thieu, Sven Gesper, Guillermo Payá-Vayá, Christoph Riggers, Oliver Renke, Till Fiedler, Jakob Marten, Tobias Stuckenberg, Holger Blume, Christian Weis, Lukas Steiner, Chirag Sudarshan, Norbert Wehn, Lennart M. Reimann, Rainer Leupers, Michael Beyer, Daniel Köhler, Alisa Jauch, Jan Micha Borrmann, Setareh Jaberansari, Tim Berthold, Meinolf Blawat, Markus Kock, Gregor Schewior, Jens Benndorf, Frederik Kautz, Hans-Martin Blüthgen, Christian Sauer 0001 |
DATE | 3 |
| 2019 | KAVUAKA: A Low Power Application Specific Hearing Aid ProcessorabstractThe integration of application specific instruction set processors (ASIPs) in hearing aids requires various architectural customizations and software-side optimizations in order to meet the stringent power consumption constraints and processing performance demands. This paper presents the KAVUAKA application specific hearing aid processor and its ASIC integration as a system on chip (SoC). The final system contains four KAVUAKA processor cores and ten co-processors. Each of these processors and co-processors were individually customized and differ in their data path width. The processors are organized in two clusters, which share memories, an audio interface, co-processors and a serial interface. With this system, different hearing aid systems are evaluated in terms of performance, power and area by activating different processor and co-processor combinations. A 40 nm low power technology was used to build this research hearing aid system. The die size is 3.6 mm2with less than 1 mm2per core. The measured average power consumption is less than 1 mW per core. Lukas Gerlach 0001, Guillermo Payá-Vayá, Holger Blume |
VLSI-SoC | 2 |
| 2019 | FLINT+: A runtime-configurable emulation-based stochastic timing analysis framework
Moritz Weißbrich, Lukas Gerlach 0001, Holger Blume, Ardalan Najafi, Alberto García Ortiz, Guillermo Payá-Vayá |
Integr. | 6 |
| 2019 | Dynamic self-reconfiguration of a MIPS-based soft-core processor architecture
Stephan Nolting, Guillermo Payá-Vayá, Florian Giesemann, Holger Blume, Sebastian Niemann, Christian Müller-Schloer |
J. Parallel Distributed Comput. | 2 |
| 2019 | Online stereo camera calibration for automotive vision based on HW-accelerated A-KAZE-feature extraction
Nico Mentzer, Jannik Mahr, Guillermo Payá-Vayá, Holger Blume |
J. Syst. Archit. | 3 |
| 2019 | Comparing vertical and horizontal SIMD vector processor architectures for accelerated image feature extraction
Moritz Weißbrich, Alberto García Ortiz, Guillermo Payá-Vayá |
J. Syst. Archit. | 3 |
| 2019 | DNN-based performance measures for predicting error rates in automatic speech recognition and optimizing hearing aid parameters
Angel Mario Castro Martinez, Lukas Gerlach 0001, Guillermo Payá-Vayá, Hynek Hermansky, Jasper Ooster, Bernd T. Meyer |
Speech Commun. | 3 |
| 2018 | Cross-layer fault-space pruning for hardware-assisted fault injectionabstractWith shrinking structure sizes, soft-error mitigation has become a major challenge in the design and certification of safety-critical embedded systems. Their robustness is quantified by extensive fault-injection campaigns, which on hardware level can nevertheless cover only a tiny part of the fault space. Christian Dietrich 0001, Achim Schmider, Oskar Pusz, Guillermo Payá-Vayá, Daniel Lohmann |
DAC | 4 |
| 2017 | Application-specific soft-core vector processor for advanced driver assistance systemsabstractImplementing convolutional neural networks for scene labelling is a current hot topic in the field of advanced driver assistance systems. The massive computational demands under hard real-time and energy constraints can only be tackled using specialized architectures. Also, cost-effectiveness is an important factor when targeting lower quantities. In this PhD thesis, a vector processor architecture optimized for FPGA devices is proposed. Amongst other hardware mechanisms, a novel complex operand addressing mode and an intelligent DMA are used to increase perfromance. Also, a C-compiler support for creating applications is introduced. Stephan Nolting, Florian Giesemann, Julian Hartig, Achim Schmider, Guillermo Payá-Vayá |
FPL | 5 |
| 2017 | Two-LUT-based synthesizable temperature sensor for Virtex-6 FPGA devicesabstractThis paper proposes a new synthesizable oscillator-based temperature sensor with minimal footprint for use in contemporary Xilinx FPGA devices. In contrast to previously published ring-oscillator architectures, based on inverters mapped onto single LUTs, the proposed oscillator uses an asynchronous Gray-coded 4-bit counter requiring only two 6-input LUTs. Due to its reduced hardware requirements, the feedback path can be implemented using local routing signals only. Therefore, the impact of the routing on the oscillator frequency is slightly reduced making the oscillator less prone to placement-caused routing deviations. The proposed temperature sensor is calibrated using a two-point calibration approach, resulting in a mean accuracy of ±0.71° C with a mean resolution of 0.0081° C. As a further case study, a methodology to characterize on-chip semiconductor variations between identical FPGAs and within the same FPGA chip is presented using a sensor array, including up to 196 of the proposed two-LUT-based oscillators. Stephan Nolting, Guillermo Payá-Vayá |
FPL | 3 |
| 2017 | Real-time implementation of a GMM-based binaural localization algorithm on a VLIW-SIMD processorabstractLocalization algorithms have become of considerable interest for robot audition, acoustic navigation, teleconferencing, speaker localization, and many other applications over the last decade. In this paper, we present a real-time implementation of a Gaussian mixture model (GMM) based probabilistic sound source localization algorithm for a low-power VLIW-SIMD processor for hearing devices. The algorithm has been proven to allow for robust localization of multiple sound sources simultaneously in reverberant and noisy environments. Real-time computation for audio frames of 512 samples at 16 kHz was achieved by introducing algorithmic optimizations and hardware customizations. To the best of our knowledge, this is the first real-time capable implementation of a computationally complex GMM-based sound source localization algorithm on a low-power processor. The resulting estimated core area without consideration of memory in 40nm low-power TSMC technology is 188,511 pm2. Christopher Seifert, Joachim Thiemann, Lukas Gerlach 0001, Tobias Volkmar, Guillermo Payá-Vayá, Holger Blume, Steven van de Par |
ICME | 5 |
| 2017 | Small footprint synthesizable temperature sensor for FPGA devices
Guillermo Payá-Vayá, Christopher Bartels, Holger Blume |
J. Syst. Archit. | 1 |
| 2016 | Performance monitoring for automatic speech recognition in noisy multi-channel environmentsabstractIn many applications of machine listening it is useful to know how well an automatic speech recognition system will do before the actual recognition is performed. In this study we investigate different performance measures with the aim of predicting word error rates (WERs) in spatial acoustic scenes in which the type of noise, the signal-to-noise ratio, parameters for spatial filtering, and the amount of reverberation are varied. All measures under consideration are based on phoneme posteriorgrams obtained from a deep neural net. While frame-wise entropy exhibits only medium predictive power for factors other than additive noise, we found the medium temporal distance between posterior vectors (M-Measure) as well as matched phoneme filters (MaP) to exhibit excellent correlations with WER across all conditions. Since our results were obtained with simulated behind-the-ear hearing aid signals, we discuss possible applications for speech-aware hearing devices. Bernd T. Meyer, Sri Harish Reddy Mallidi, Angel Mario Castro Martinez, Guillermo Payá-Vayá, Hendrik Kayser, Hynek Hermansky |
SLT | 4 |
| 2015 | FLINT: layout-oriented FPGA-based methodology for fault tolerant ASIC design
Rochus Nowosielski, Lukas Gerlach 0001, Stephan Bieband, Guillermo Payá-Vayá, Holger Blume |
DATE | 4 |
| 2014 | ASEV - Automatic situation assessment for event-driven video analysisabstractMany complex maneuvers involving aircraft, vehicles and persons are carried out at airport aprons. Manual video surveillance used for safety and security purposes is inefficient and privacy protection must be guaranteed. In this paper, we propose a system named ASEV that automatically assesses situations for airport surveillance. It combines four main components: a low-level image processing unit based on a new hardware implementation to extract features in real time, a high-level image processing unit for scene analysis, a real-time inference engine for scene understanding, and a data protection stage for log encryption. In addition, four often neglected aspects are successfully addressed: two-way communication between system and operator, power consumption, monitored people privacy and operator activity control. Extensive evaluation at a real airport shows that the proposed system improves the operator performance with sound and visual alerts based on the automatic assessment of various events. Michele Fenzi, Jörn Ostermann, Nico Mentzer, Guillermo Payá-Vayá, Holger Blume, Tu Ngoc Nguyen, Thomas Risse 0001 |
AVSS | 4 |
| 2010 | A forwarding-sensitive instruction scheduling approach to reduce register file constraints in VLIW architecturesabstractThis paper presents a forwarding-based approach to increase the code compaction and consequently the processing performance of VLIW media-processors that implement monolithic or partitioned register file (RF) organizations with reduced number of read/write ports. This approach exploits the forwarding mechanism implemented in common pipelined VLIW architectures to reduce the number of RF accesses, which is one of the main limiting factors of the code compaction process. This RF access reduction enables a higher instruction scheduling efficiency and eventually decreases the power consumption, without requiring extra hardware. A forwarding-sensitive code generation algorithm based on an enhanced list scheduling algorithm is described in detail. In addition, three case studies are presented, where the proposed scheduling algorithm leads to performance improvements of up to 8.4% when running common image and video codec tasks on a generic VLIW architecture. This is attractively close to the maximum performance improvement (11.4%) that can be achieved when investing in hardware by using a RF with twice the number of ports. Guillermo Payá-Vayá, Javier Martín-Langerwerf, Holger Blume, Peter Pirsch |
ASAP | 1 |
| 2007 | RAPANUI: A case study in Rapid Prototyping for Multiprocessor System-on-ChipabstractThis paper describes a case study in a new rapid prototyping-based design framework for exploring and validating complex multiprocessor architectures for multimedia applications. The goal of the presented methodology is to speed up and improve the verification flow of a multiprocessor system that will finally be implemented as an ASIC. The case study consists of a 64-bit compatible AMBA AHB system bus which connects up to 14 32-Bit RISC processors to a host interface. A typical parallel computing application has been implemented for the parameterized multiprocessor system. The employed FPGA emulation environment increases by up to 200 the simulation frequency of the global system on a workstation (2.2 GHz AMD Dual Opteron with 8 GB RAM). Moreover a standalone emulation can be performed at the maximum achievable frequency (65 MHz). Guillermo Payá-Vayá, Javier Martín-Langerwerf, Peter Pirsch |
DSD | 1 |
| 2004 | FPGA Custom DSP for ECG Signal Analysis and Compression
Marcos Martínez Peiró, Francisco José Ballester-Merelo, Guillermo Payá-Vayá, Ricardo José Colom-Palero, Rafael Gadea Gironés, José Belenguer |
FPL | 3 |
| 2004 | Architectures for ICT on FPGAabstractWe evaluate some architectures for the implementation on FPGA of the one and two dimensions integer cosine transform (ICT). The area and speed synthesis results are shown. The ICT is the transformation used in the newest video compression standard H.264/AVC. Arturo Méndez Patiño, Marcos Martínez Peiró, Francisco José Ballester-Merelo, Guillermo Payá-Vayá |
FPT | 4 |
| 2003 | Fully Parameterized Discrete Wavelet Packet Transform Architecture Oriented to FPGA
Guillermo Payá-Vayá, Marcos Martínez Peiró, Francisco José Ballester-Merelo, Francisco José Mora Mas |
FPL | 1 |