Guillermo Payá-Vayá

dblp:64/1540 · also Guillermo Payá Vayá · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-3503-8386ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Software engineering, systems software and programming languages · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 OpenFI4ASIC: An Open-Source Fault Injection Framework for ASIC Designs via FPGA-Based Rapid Prototyping
Jasper Homann, Eike Trumann, Umut Durak, Guillermo Payá-Vayá
SAFECOMP4
2025 DCMA: Accelerating Parallel DMA Transfers with a Multi-Port Direct Cached Memory Access in a Massive-Parallel Vector Processor
abstract
State-of-the-art applications, such as convolutional neural networks, demand specialized hardware accelerators that address performance and efficiency constraints. An efficient memory hierarchy is mandatory for such hardware systems. While the memory architectures of general-purpose processors (e.g., CPU or GPUs) are based on cache systems, dedicated accelerators have mostly adopted the DMA (Direct Memory Access) concept due to the application field of image processing. DMA features like 2D data transfers or data padding can optimize the memory accesses of image processing. However, DMA lacks the capability to exploit temporal and spatial data reuse, a feature common in cache systems, particularly when multiple DMAs operate in parallel. This article proposes a novel Direct Cached Memory Access (DCMA) architecture, combining both DMA and cache methodologies and their respective advantages. Optimized for image-based AI algorithms, the DCMA architecture facilitates enhanced memory access by integrating multiple, parallel DMA ports with caching capabilities. This design allows for efficient data reuse and parallel memory access. Optimal parameters for the DCMA are determined through a comprehensive design space exploration. The DCMA is evaluated on a state-of-the-art Xilinx UltraScale+ FPGA board coupled with a massive-parallel vertical vector co-processor, called V 2 PRO. The results show the mitigation of the vector processor’s memory bottleneck. By using the proposed DCMA, speedups of up to ×17 for the ResNet-50 CNN can be achieved.
Gia Bao Thieu, Sven Gesper, Guillermo Payá-Vayá
ACM Trans. Archit. Code Optim.3
2025 An Open-Source NanoController v2 Featuring Microcoded Instruction Set Redefinition in 22-nm FDSOI-CMOS for Autonomous Ultralow-Power SoCs
abstract
The realization of autonomous, wearable, and implantable system-on-chip (SoC) for health monitoring applications poses several challenges, such as achieving ultralow size, cost, and power consumption, yet offering sufficient flexibility to reprogram and adapt the autonomously operating SoC during the course of treatment. Commonly, programmability is not considered for ultralow-power (ULP) biomedical SoCs, since a dedicated finite state machine (FSM) fixes the operation sequence, and instruction memory presents significant contributions to silicon area and power consumption. Based on a previously published tiny, programmable microarchitecture, this work proposes the strongly enhancedNanoController v2, for potential use in ultralow-power biomedical SoCs, and integrates it as a prototype chip in a 22-nm FDSOI-CMOS technology. By implementing a novel microcoded control unit and an automated design space exploration framework, which are made available open-source, the instruction set can be freely redefined to exploit application-specific properties. The benefits are increased code compaction, performance gain, and, consequently, decreased power consumption. In an extensive measurement campaign, a glucose sensor control application achieves 13.1% higher performance and 15.6% less code size in the best case, resulting in an extremely low power consumption of 660 nW (9% less than the reference),only by a different instruction setwithout hardware changes. Compared with other state-of-the-art small programmable microcontrollers, between 38% and 82% smaller code size and between 33% and 77% smaller silicon area and averaged power consumption could be shown. Based on the prototype results, a fully integrated glucose sensor chip will be evaluated in currently ongoing work.
Moritz Weißbrich, Adilet Dossanov, Yerzhan Kudabay, Alexander Meyer, Vadim Issakov, Guillermo Payá-Vayá
IEEE Trans. Very Large Scale Integr. Syst.6
2024 Multi-Level Prototyping of a Vertical Vector AI Processing System
abstract
Modern embedded systems must be designed carefully to cope with the complexity and real-time requirements of modern AI (Artificial Intelligence) driven automotive applications, such as Advanced Driver-Assistance Systems (ADAS). Despite increasing complexity, the time to market is decreasing. In this work, a SystemC-based Virtual Prototype of a neural network processing platform is exploited to bypass the limitations of standalone instruction set simulators (ISS) and FPGA prototyping. The processing platform under test is based on a novel massive parallel vector processor architecture coupled with a RISC- V control core that runs widely used convolutional neural networks (CNNs) for object detection. The paper discusses the variations and appropriateness of the three prototyping methods outlined, demonstrating how the Virtual Prototype can address the aforementioned constraints, resulting in a 2.07x increase in accuracy, 16x greater configurations, and more profound insights into the system compared to standalone and FPGA prototyping.
Frederik Kautz, Sven Gesper, Gia Bao Thieu, Hans-Martin Blüthgen, Holger Blume, Guillermo Payá-Vayá
ASAP6
2023 Exploiting Subword Permutations to Maximize CNN Compute Performance and Efficiency
abstract
Neural networks (NNs) are quantized to decrease their computational demands and reduce their memory foot-print. However, specialized hardware is required that supports computations with low bit widths to take advantage of such optimizations. In this work, we propose permutations on subword level that build on top of multi-bit-width multiply-accumulate operations to effectively support low bit width computations of quantized NNs. By applying this technique, we extend the data reuse and further improve compute performance for convolution operations compared to simple vectorization using SIMD (single-instruction-multiple-data). We perform a design space exploration using a cycle accurate simulation with MobileNet and VGG16 on a vector-based processor. The results show a speedup of up to$3.7\times$and a reduction of up to$1.9\times$for required data transfers. Additionally, the control overhead for orchestrating the computation is decreased by up to$3.9\times$.
Michael Beyer, Sven Gesper, Andre Guntoro, Guillermo Payá-Vayá, Holger Blume
ASAP4
2023 ZuSE Ki-Avf: Application-Specific AI Processor for Intelligent Sensor Signal Processing in Autonomous Driving
abstract
Modern and future AI-based automotive applications, such as autonomous driving, require the efficient real-time processing of huge amounts of data from different sensors, like camera, radar, and LiDAR. In the ZuSE-KI-AVF project, multiple university, and industry partners collaborate to develop a novel massive parallel processor architecture, based on a cus-tomized RISC-V host processor, and an efficient high-performance vertical vector coprocessor. In addition, a software development framework is also provided to efficiently program AI-based sensor processing applications. The proposed processor system was verified and evaluated on a state-of-the-art UltraScale+ FPGA board, reaching a processing performance of up to 126.9 FPS, while executing the YOLO-LITE CNN on 224x224 input images. Further optimizations of the FPGA design and the realization of the processor system on a 22nm FDSOI CMOS technology are planned.
Gia Bao Thieu, Sven Gesper, Guillermo Payá-Vayá, Christoph Riggers, Oliver Renke, Till Fiedler, Jakob Marten, Tobias Stuckenberg, Holger Blume, Christian Weis, Lukas Steiner, Chirag Sudarshan, Norbert Wehn, Lennart M. Reimann, Rainer Leupers, Michael Beyer, Daniel Köhler, Alisa Jauch, Jan Micha Borrmann, Setareh Jaberansari, Tim Berthold, Meinolf Blawat, Markus Kock, Gregor Schewior, Jens Benndorf, Frederik Kautz, Hans-Martin Blüthgen, Christian Sauer 0001
DATE3
2019 KAVUAKA: A Low Power Application Specific Hearing Aid Processor
abstract
The integration of application specific instruction set processors (ASIPs) in hearing aids requires various architectural customizations and software-side optimizations in order to meet the stringent power consumption constraints and processing performance demands. This paper presents the KAVUAKA application specific hearing aid processor and its ASIC integration as a system on chip (SoC). The final system contains four KAVUAKA processor cores and ten co-processors. Each of these processors and co-processors were individually customized and differ in their data path width. The processors are organized in two clusters, which share memories, an audio interface, co-processors and a serial interface. With this system, different hearing aid systems are evaluated in terms of performance, power and area by activating different processor and co-processor combinations. A 40 nm low power technology was used to build this research hearing aid system. The die size is 3.6 mm2with less than 1 mm2per core. The measured average power consumption is less than 1 mW per core.
Lukas Gerlach 0001, Guillermo Payá-Vayá, Holger Blume
VLSI-SoC2
2019 FLINT+: A runtime-configurable emulation-based stochastic timing analysis framework
Moritz Weißbrich, Lukas Gerlach 0001, Holger Blume, Ardalan Najafi, Alberto García Ortiz, Guillermo Payá-Vayá
Integr.6
2019 Dynamic self-reconfiguration of a MIPS-based soft-core processor architecture
Stephan Nolting, Guillermo Payá-Vayá, Florian Giesemann, Holger Blume, Sebastian Niemann, Christian Müller-Schloer
J. Parallel Distributed Comput.2
2019 Online stereo camera calibration for automotive vision based on HW-accelerated A-KAZE-feature extraction
Nico Mentzer, Jannik Mahr, Guillermo Payá-Vayá, Holger Blume
J. Syst. Archit.3
2019 Comparing vertical and horizontal SIMD vector processor architectures for accelerated image feature extraction
Moritz Weißbrich, Alberto García Ortiz, Guillermo Payá-Vayá
J. Syst. Archit.3
2019 DNN-based performance measures for predicting error rates in automatic speech recognition and optimizing hearing aid parameters
Angel Mario Castro Martinez, Lukas Gerlach 0001, Guillermo Payá-Vayá, Hynek Hermansky, Jasper Ooster, Bernd T. Meyer
Speech Commun.3
2018 Cross-layer fault-space pruning for hardware-assisted fault injection
abstract
With shrinking structure sizes, soft-error mitigation has become a major challenge in the design and certification of safety-critical embedded systems. Their robustness is quantified by extensive fault-injection campaigns, which on hardware level can nevertheless cover only a tiny part of the fault space.
Christian Dietrich 0001, Achim Schmider, Oskar Pusz, Guillermo Payá-Vayá, Daniel Lohmann
DAC4
2017 Application-specific soft-core vector processor for advanced driver assistance systems
abstract
Implementing convolutional neural networks for scene labelling is a current hot topic in the field of advanced driver assistance systems. The massive computational demands under hard real-time and energy constraints can only be tackled using specialized architectures. Also, cost-effectiveness is an important factor when targeting lower quantities. In this PhD thesis, a vector processor architecture optimized for FPGA devices is proposed. Amongst other hardware mechanisms, a novel complex operand addressing mode and an intelligent DMA are used to increase perfromance. Also, a C-compiler support for creating applications is introduced.
Stephan Nolting, Florian Giesemann, Julian Hartig, Achim Schmider, Guillermo Payá-Vayá
FPL5
2017 Two-LUT-based synthesizable temperature sensor for Virtex-6 FPGA devices
abstract
This paper proposes a new synthesizable oscillator-based temperature sensor with minimal footprint for use in contemporary Xilinx FPGA devices. In contrast to previously published ring-oscillator architectures, based on inverters mapped onto single LUTs, the proposed oscillator uses an asynchronous Gray-coded 4-bit counter requiring only two 6-input LUTs. Due to its reduced hardware requirements, the feedback path can be implemented using local routing signals only. Therefore, the impact of the routing on the oscillator frequency is slightly reduced making the oscillator less prone to placement-caused routing deviations. The proposed temperature sensor is calibrated using a two-point calibration approach, resulting in a mean accuracy of ±0.71° C with a mean resolution of 0.0081° C. As a further case study, a methodology to characterize on-chip semiconductor variations between identical FPGAs and within the same FPGA chip is presented using a sensor array, including up to 196 of the proposed two-LUT-based oscillators.
Stephan Nolting, Guillermo Payá-Vayá
FPL3
2017 Real-time implementation of a GMM-based binaural localization algorithm on a VLIW-SIMD processor
abstract
Localization algorithms have become of considerable interest for robot audition, acoustic navigation, teleconferencing, speaker localization, and many other applications over the last decade. In this paper, we present a real-time implementation of a Gaussian mixture model (GMM) based probabilistic sound source localization algorithm for a low-power VLIW-SIMD processor for hearing devices. The algorithm has been proven to allow for robust localization of multiple sound sources simultaneously in reverberant and noisy environments. Real-time computation for audio frames of 512 samples at 16 kHz was achieved by introducing algorithmic optimizations and hardware customizations. To the best of our knowledge, this is the first real-time capable implementation of a computationally complex GMM-based sound source localization algorithm on a low-power processor. The resulting estimated core area without consideration of memory in 40nm low-power TSMC technology is 188,511 pm2.
Christopher Seifert, Joachim Thiemann, Lukas Gerlach 0001, Tobias Volkmar, Guillermo Payá-Vayá, Holger Blume, Steven van de Par
ICME5
2017 Small footprint synthesizable temperature sensor for FPGA devices
Guillermo Payá-Vayá, Christopher Bartels, Holger Blume
J. Syst. Archit.1
2016 Performance monitoring for automatic speech recognition in noisy multi-channel environments
abstract
In many applications of machine listening it is useful to know how well an automatic speech recognition system will do before the actual recognition is performed. In this study we investigate different performance measures with the aim of predicting word error rates (WERs) in spatial acoustic scenes in which the type of noise, the signal-to-noise ratio, parameters for spatial filtering, and the amount of reverberation are varied. All measures under consideration are based on phoneme posteriorgrams obtained from a deep neural net. While frame-wise entropy exhibits only medium predictive power for factors other than additive noise, we found the medium temporal distance between posterior vectors (M-Measure) as well as matched phoneme filters (MaP) to exhibit excellent correlations with WER across all conditions. Since our results were obtained with simulated behind-the-ear hearing aid signals, we discuss possible applications for speech-aware hearing devices.
Bernd T. Meyer, Sri Harish Reddy Mallidi, Angel Mario Castro Martinez, Guillermo Payá-Vayá, Hendrik Kayser, Hynek Hermansky
SLT4
2015 FLINT: layout-oriented FPGA-based methodology for fault tolerant ASIC design
Rochus Nowosielski, Lukas Gerlach 0001, Stephan Bieband, Guillermo Payá-Vayá, Holger Blume
DATE4
2014 ASEV - Automatic situation assessment for event-driven video analysis
abstract
Many complex maneuvers involving aircraft, vehicles and persons are carried out at airport aprons. Manual video surveillance used for safety and security purposes is inefficient and privacy protection must be guaranteed. In this paper, we propose a system named ASEV that automatically assesses situations for airport surveillance. It combines four main components: a low-level image processing unit based on a new hardware implementation to extract features in real time, a high-level image processing unit for scene analysis, a real-time inference engine for scene understanding, and a data protection stage for log encryption. In addition, four often neglected aspects are successfully addressed: two-way communication between system and operator, power consumption, monitored people privacy and operator activity control. Extensive evaluation at a real airport shows that the proposed system improves the operator performance with sound and visual alerts based on the automatic assessment of various events.
Michele Fenzi, Jörn Ostermann, Nico Mentzer, Guillermo Payá-Vayá, Holger Blume, Tu Ngoc Nguyen, Thomas Risse 0001
AVSS4
2010 A forwarding-sensitive instruction scheduling approach to reduce register file constraints in VLIW architectures
abstract
This paper presents a forwarding-based approach to increase the code compaction and consequently the processing performance of VLIW media-processors that implement monolithic or partitioned register file (RF) organizations with reduced number of read/write ports. This approach exploits the forwarding mechanism implemented in common pipelined VLIW architectures to reduce the number of RF accesses, which is one of the main limiting factors of the code compaction process. This RF access reduction enables a higher instruction scheduling efficiency and eventually decreases the power consumption, without requiring extra hardware. A forwarding-sensitive code generation algorithm based on an enhanced list scheduling algorithm is described in detail. In addition, three case studies are presented, where the proposed scheduling algorithm leads to performance improvements of up to 8.4% when running common image and video codec tasks on a generic VLIW architecture. This is attractively close to the maximum performance improvement (11.4%) that can be achieved when investing in hardware by using a RF with twice the number of ports.
Guillermo Payá-Vayá, Javier Martín-Langerwerf, Holger Blume, Peter Pirsch
ASAP1
2007 RAPANUI: A case study in Rapid Prototyping for Multiprocessor System-on-Chip
abstract
This paper describes a case study in a new rapid prototyping-based design framework for exploring and validating complex multiprocessor architectures for multimedia applications. The goal of the presented methodology is to speed up and improve the verification flow of a multiprocessor system that will finally be implemented as an ASIC. The case study consists of a 64-bit compatible AMBA AHB system bus which connects up to 14 32-Bit RISC processors to a host interface. A typical parallel computing application has been implemented for the parameterized multiprocessor system. The employed FPGA emulation environment increases by up to 200 the simulation frequency of the global system on a workstation (2.2 GHz AMD Dual Opteron with 8 GB RAM). Moreover a standalone emulation can be performed at the maximum achievable frequency (65 MHz).
Guillermo Payá-Vayá, Javier Martín-Langerwerf, Peter Pirsch
DSD1
2004 FPGA Custom DSP for ECG Signal Analysis and Compression
Marcos Martínez Peiró, Francisco José Ballester-Merelo, Guillermo Payá-Vayá, Ricardo José Colom-Palero, Rafael Gadea Gironés, José Belenguer
FPL3
2004 Architectures for ICT on FPGA
abstract
We evaluate some architectures for the implementation on FPGA of the one and two dimensions integer cosine transform (ICT). The area and speed synthesis results are shown. The ICT is the transformation used in the newest video compression standard H.264/AVC.
Arturo Méndez Patiño, Marcos Martínez Peiró, Francisco José Ballester-Merelo, Guillermo Payá-Vayá
FPT4
2003 Fully Parameterized Discrete Wavelet Packet Transform Architecture Oriented to FPGA
Guillermo Payá-Vayá, Marcos Martínez Peiró, Francisco José Ballester-Merelo, Francisco José Mora Mas
FPL1