EDBT 2026 Demo / reviewers in the wild / expert
Katzalin Olcoz
dblp:08/4627
· DBLP profile ↗
26ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-1821-124XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 2Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Solving the task scheduling and GPU reconfiguration problem on MIG devices via deep reinforcement learningabstract• Multi-Instance GPU (MIG) technology enables adaptive co-execution of tasks, greatly improving the efficiency of computational resources in a flexible manner. • Prior methods simplify the MIG scheduling and dynamic reconfiguration challenge by reducing problem complexity, but can yield markedly suboptimal solutions in certain scenarios. • Modeling the problem with Reinforcement Learning (RL), and training with Deep Learning techniques, allows to approach it successfully without great simplifications, despite its high dimensionality. The design of the RL agent also needs to be carefully refined, with analysis such as that detailed in the manuscript, which can serve as a guide for similar resource management work. • Our refined RL agent reduces makespan by 2–7% over state-of-the-art on a wide set of benchmarks and synthetic workloads, with improvements of up to 30% in specific cases. Additional benefits include enhanced flexibility and adaptability of the scheduling framework. Recent advances in dynamic GPU partitioning, such as NVIDIA’s Multi-Instance GPU (MIG) technology, have enhanced resource utilization by enabling task co-execution without contention. However, existing MIG schedulers remain limited to static or task-agnostic methods that sacrifice optimality for tractability. This paper presents a Deep Reinforcement Learning framework that seeks to minimize the completion time of a task queue by holistically addressing the dimensions of the problem: task molding, GPU reconfiguration and execution order. To manage the vast solution space, we apply optimizations such as discrete and canonical representation of states, unification of equivalent configurations, action masking, or promoting the exploration of reconfigurations; this offers insights for similar resource management scenarios. The proposed models are extensively evaluated with widely used benchmarks of the Rodinia and Altis suites, and synthetic workloads generated to emulate a wide range of plausible real situations. The final model improves to the state-of-the-art, especially in workloads that clearly contradict the assumptions of previous proposals, achieving a difference of less than 20% to the optimum. Additionally, two different approaches to the problem are faced (offline vs. online), discussing their theoretical advantages and disadvantages, and evaluating them experimentally for the final model. Jorge Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz |
Future Gener. Comput. Syst. | 4 |
| 2026 | A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologies
Jorge Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz |
J. Supercomput. | 4 |
| 2025 | Leveraging Multi-Instance GPUs through moldable task schedulingabstractNVIDIA MIG (Multi-Instance GPU) allows partitioning a physical GPU into multiple logical instances with fully-isolated resources, which can be dynamically reconfigured. This work highlights the untapped potential of MIG through moldable task scheduling with dynamic reconfigurations . Specifically, we propose a makespan minimization problem for multi-task execution under MIG constraints. Our profiling shows that assuming monotonicity in task work with respect to resources is not viable, as is usual in multicore scheduling. Relying on a state-of-the-art proposal that does not require such an assumption, we present FAR , a 3-phase algorithm to solve the problem. Phase 1 of FAR builds on a classical task moldability method, phase 2 combines Longest Processing Time First and List Scheduling with a novel repartitioning tree heuristic tailored to MIG constraints, and phase 3 employs local search via task moves and swaps. FAR schedules tasks in batches offline, concatenating their schedules on the fly in an improved way that favors resource reuse. Excluding reconfiguration costs, the List Scheduling proof shows an approximation factor of 7/4 on the NVIDIA A30 model. We adapt the technique to the particular constraints of an NVIDIA A100/H100 to obtain an approximation factor of 2. Including the reconfiguration cost, our real-world experiments reveal a makespan with respect to the optimum no worse than 1.22× for a well-known suite of benchmarks, and 1.10× for synthetic inputs inspired by real kernels. We obtain good experimental results for each batch of tasks, but also in the concatenation of batches, with large improvements over the state-of-the-art and proposals without GPU reconfiguration. Moreover, we show that the proposed heuristics allow a correct adaptation to tasks of very different characteristics. Beyond the specific algorithm, the paper demonstrates the research potential of the MIG technology and suggests useful metrics, workload characterizations and evaluation techniques for future work in this field. Jorge Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz |
J. Parallel Distributed Comput. | 4 |
| 2025 | Balanced segmentation of CNNs for multi-TPU inference
John S. Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz |
J. Supercomput. | 4 |
| 2023 | Distributed training and inference of deep learning solar energy forecasting modelsabstractDifferent accurate predictive models have been developed to forecast the amount of solar energy produced in a given area. These models are usually run in a centralized manner, considering irradiance inputs taken from a set of sensors that are deployed in that area. CAIDE is a framework that supports the deployment and analysis of solar plants following Model Based System Engineering (MBSE) and Internet of Things (IoT) methodologies. However, the current solution performs the training and inference phases of the solar energy forecasting models in a central way, not taking advantage of the distributed environment modeled by means of CAIDE. This work presents an extension of CAIDE that allows us to distribute the training and inference phases, obtaining performance improvements, and achieving a greater adaptation to the inherently distributed topology of the deployment of the sensors. Javier Campoy, Ignacio-Iker Prado-Rujas, José Luis Risco-Martín, Katzalin Olcoz, María S. Pérez 0001 |
PDP | 4 |
| 2023 | Improving inference time in multi-TPU systems with profiled model segmentationabstractIn this paper, we systematically evaluate the inference performance of the Edge TPU by Google for neural networks with different characteristics. Specifically, we determine that, given the limited amount of on-chip memory on the Edge TPU, accesses to external (host) memory rapidly become an important performance bottleneck. We demonstrate how multiple devices can be jointly used to alleviate the bottleneck introduced by accessing the host memory. We propose a solution combining model segmentation and pipelining on up to four TPUs, with remarkable performance improvements that range from 6x for neural networks with convolutional layers to 46x for fully connected layers, compared with single-TPU setups. Jorge Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz |
PDP | 4 |
| 2021 | Gem5-X: A Many-core Heterogeneous Simulation Platform for Architectural Exploration and OptimizationabstractThe increasing adoption of smart systems in our daily life has led to the development of new applications with varying performance and energy constraints, and suitable computing architectures need to be developed for these new applications. In this article, we present gem5-X, a system-level simulation framework, based on gem-5, for architectural exploration of heterogeneous many-core systems. To demonstrate the capabilities of gem5-X, real-time video analytics is used as a case-study. It is composed of two kernels, namely, video encoding and image classification using convolutional neural networks (CNNs). First, we explore through gem5-X the benefits of latest 3D high bandwidth memory (HBM2) in different architectural configurations. Then, using a two-step exploration methodology, we develop a new optimized clustered-heterogeneous architecture with HBM2 in gem5-X for video analytics application. In this proposed clustered-heterogeneous architecture, ARMv8 in-order cluster with in-cache computing engine executes the video encoding kernel, giving 20% performance and 54% energy benefits compared to baseline ARM in-order and Out-of-Order systems, respectively. Furthermore, thanks to gem5-X, we conclude that ARM Out-of-Order clusters with HBM2 are the best choice to run visual recognition using CNNs, as they outperform DDR4-based system by up to 30% both in terms of performance and energy savings. Yasir Mahmood Qureshi, William Andrew Simon, Marina Zapater, Katzalin Olcoz, David Atienza 0001 |
ACM Trans. Archit. Code Optim. | 4 |
| 2021 | Genome Sequence Alignment - Design Space Exploration for Optimal Performance and Energy ArchitecturesabstractNext generation workloads, such as genome sequencing, have an astounding impact in the healthcare sector. Sequence alignment, the first step in genome sequencing, has experienced recent breakthroughs, which resulted in next generation sequencing (NGS). As NGS applications are memory bounded with random memory access patterns, we propose the use of high bandwidth memories like 3D stacked HBM2, instead of traditional DRAMs like DDR4, along with energy efficient compute cores to improve both performance and energy efficiency. Three state-of-the-art NGS applications, Bowtie2, BWA-MEM, and HISAT2 are used as case studies to explore and optimize NGS computing architectures. Then, using the gem5-X architectural simulator, we obtain an overall 68 percent performance improvement and 71 percent energy savings using HBM2 instead of DDR4. Furthermore, we propose an architecture based on ARMv8 cores and demonstrate that 16 ARMv8 64-bit OoO cores with HBM2 outperforms 32-cores of Intel Xeon Phi Knights Landing (KNL) processor with 3D stacked memory. Moreover, we show that by using frequency scaling we can achieve up to 59 percent and 61 percent energy savings for ARM in-order and OoO cores, respectively. Lastly, we show that many ARMv8 in-order cores at 1.5GHz match the performance of fewer OoO cores at 2GHz, while attaining 4.5x energy savings. Yasir Mahmood Qureshi, Jose Manuel Herruzo, Marina Zapater, Katzalin Olcoz, Sonia Gonzalez-Navarro, Oscar G. Plata, David Atienza 0001 |
IEEE Trans. Computers | 4 |
| 2020 | Leveraging knowledge-as-a-service (KaaS) for QoS-aware resource management in multi-user video transcoding
Luis Costero, Francisco D. Igual, Katzalin Olcoz, Francisco Tirado |
J. Supercomput. | 3 |
| 2020 | Resource Management for Power-Constrained HEVC Transcoding Using Reinforcement LearningabstractThe advent of online video streaming applications and services along with the users' demand for high-quality contents require High Efficiency Video Coding (HEVC), which provides higher video quality and more compression at the cost of increased complexity. On one hand, HEVC exposes a set of dynamically tunable parameters to provide trade-offs among Quality-of-Service (QoS), performance, and power consumption of multi-core servers on the video providers' data center. On the other hand, resource management of modern multi-core servers is in charge of adapting system-level parameters, such as operating frequency and multithreading, to deal with concurrent applications and their requirements. Therefore, efficient multi-user HEVC streaming necessitates joint adaptation of application-and system-level parameters. Nonetheless, dealing with such a large and dynamic design space is challenging and difficult to address through conventional resource management strategies. Thus, in this work, we develop a multi-agent Reinforcement Learning framework to jointly adjust application-and system-level parameters at runtime to satisfy the QoS of multi-user HEVC streaming in power-constrained servers. In particular, the design space, composed of all design parameters, is split into smaller independent sub-spaces. Each design sub-space is assigned to a particular agent so that it can explore it faster, yet accurately. The benefits of our approach are revealed in terms of adaptability and quality (with up to to 4× improvements in terms of QoS when compared to a static resource management scheme), and learning time (6× fasterthan an equivalent mono-agent implementation). Finally, we show that the power-capping techniques formulated outperform the hardware-based power capping with respect to quality. Luis Costero, Arman Iranfar, Marina Zapater, Francisco D. Igual, Katzalin Olcoz, David Atienza 0001 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2019 | MAMUT: Multi-Agent Reinforcement Learning for Efficient Real-Time Multi-User Video TranscodingabstractReal-time video transcoding has recently raised as a valid alternative to address the ever-increasing demand for video contents in servers' infrastructures in current multi-user environments. High Efficiency Video Coding (HEVC) makes efficient online transcoding feasible as it enhances user experience by providing the adequate video configuration, reduces pressure on the network, and minimizes inefficient and costly video storage. However, the computational complexity of HEVC, together with its myriad of configuration parameters, raises challenges for power management, throughput control, and Quality of Service (QoS) satisfaction. This is particularly challenging in multi-user environments where multiple users with different resolution demands and bandwidth constraints need to be served simultaneously. In this work, we present MAMUT, a multi-agent machine learning approach to tackle these challenges. Our proposal breaks the design space composed of run-time adaptation of the transcoder and system parameters into smaller sub-spaces that can be explored in a reasonable time by individual agents. While working cooperatively, each agent is in charge of learning and applying the optimal values for internal HEVC and system-wide parameters. In particular, MAMUT dynamically tunes Quantization Parameter, selects number of threads per video, and sets the operating frequency with throughput and video quality objectives under compression and power consumption constraints. We implement MAMUT on an enterprise multicore server and compare equivalent scenarios to state-of-the-art alternative approaches. The obtained results reveal that MAMUT consistently attains up to 8× improvement in terms of FPS violations (and thus Quality of Service), 24% power reduction, as well as faster and more accurate adaptation both to the video contents and available resources. Luis Costero, Arman Iranfar, Marina Zapater, Francisco D. Igual, Katzalin Olcoz, David Atienza 0001 |
DATE | 5 |
| 2019 | A Machine Learning-Based Framework for Throughput Estimation of Time-Varying Applications in Multi-Core ServersabstractAccurate workload prediction and throughput estimation are keys in efficient proactive power and performance management of multi-core platforms. Although hardware performance counters available on modern platforms contain important information about the application behavior, employing them efficiently is not straightforward when dealing with time-varying applications even if they have iterative structures. In this work, we propose a machine learning-based framework for workload prediction and throughput estimation using hardware events. Our framework enables throughput estimation over various available system configurations, namely, number of parallel threads and operating frequency. In particular, we first employ workload clustering and classification techniques along with Markov chains to predict the next workload for each available system configuration. Then, the predicted workload is used to estimate the next expected throughput through a machine learning-based regression model. The comparison with state of the art demonstrates that our framework is able to improve Quality of Service (QoS) by 3.4x, while consuming 15% less power thanks to the more accurate throughput estimation. Arman Iranfar, Wellington Silva de Souza, Marina Zapater, Katzalin Olcoz, Samuel Xavier de Souza, David Atienza 0001 |
VLSI-SoC | 4 |
| 2019 | A QoS and Container-Based Approach for Energy Saving and Performance Profiling in Multi-Core ServersabstractIn this work we present ContainEnergy, a new performance evaluation and profiling tool that uses software containers to perform application runtime assessment, providing energy and performance profiling data. It is focused on energy efficiency for next generation workloads and IT infrastructure. Wellington Silva de Souza, Arman Iranfar, Anderson B. N. da Silva, Marina Zapater, Samuel Xavier de Souza, Katzalin Olcoz, David Atienza 0001 |
VLSI-SoC | 6 |
| 2018 | Level-Spread: A New Job Allocation Policy for Dragonfly NetworksabstractThe dragonfly network topology has attracted attention in recent years owing to its high radix and constant diameter. However, the influence of job allocation on communication time in dragonfly networks is not fully understood. Recent studies have shown that random allocation is better at balancing the network traffic, while compact allocation is better at harnessing the locality in dragonfly groups. Based on these observations, this paper introduces a novel allocation policy called Level-Spread for dragonfly networks. This policy spreads jobs within the smallest network level that a given job can fit in at the time of its allocation. In this way, it simultaneously harnesses node adjacency and balances link congestion. To evaluate the performance of Level-Spread, we run packet-level network simulations using a diverse set of application communication patterns, job sizes, and communication intensities. We also explore the impact of network properties such as the number of groups, number of routers per group, machine utilization level, and global link bandwidth. Level-Spread reduces the communication overhead by 16% on average (and up to 71%) compared to the state-of-the-art allocation policies. Yijia Zhang 0002, Ozan Tuncer, Fulya Kaplan, Katzalin Olcoz, Vitus J. Leung, Ayse K. Coskun |
IPDPS | 4 |
| 2017 | User-profile-based analytics for detecting cloud security breachesabstractWhile the growth of cloud-based technologies has benefited the society tremendously, it has also increased the surface area for cyber attacks. Given that cloud services are prevalent today, it is critical to devise systems that detect intrusions. One form of security breach in the cloud is when cyber-criminals compromise Virtual Machines (VMs) of unwitting users and, then, utilize user resources to run time-consuming, malicious, or illegal applications for their own benefit. This work proposes a method to detect unusual resource usage trends and alert the user and the administrator in real time. We experiment with three categories of methods: simple statistical techniques, unsupervised classification, and regression. So far, our approach successfully detects anomalous resource usage when experimenting with typical trends synthesized from published real-world web server logs and cluster traces. We observe the best results with unsupervised classification, which gives an average F1-score of 0.83 for web server logs and 0.95 for the cluster traces. Trishita Tiwari, Ata Turk, Alina Oprea, Katzalin Olcoz, Ayse K. Coskun |
IEEE BigData | 4 |
| 2017 | Revisiting conventional task schedulers to exploit asymmetry in multi-core architectures for dense linear algebra operations
Luis Costero, Francisco D. Igual, Katzalin Olcoz, Sandra Catalán, Rafael Rodríguez-Sánchez 0001, Enrique S. Quintana-Ortí |
Parallel Comput. | 3 |
| 2014 | A novel energy-driven computing paradigm for e-health scenarios
Marina Zapater, Patricia Arroba, José Luis Ayala, José Manuel Moya, Katzalin Olcoz |
Future Gener. Comput. Syst. | 5 |
| 2014 | VLSI for the new era
José Luis Ayala, Katzalin Olcoz |
Integr. | 2 |
| 2012 | Memory power optimization of Java-based embedded systems exploiting garbage collection information
José Manuel Velasco, David Atienza 0001, Katzalin Olcoz |
J. Syst. Archit. | 3 |
| 2009 | Exploration of memory hierarchy configurations for efficient garbage collection on high-performance embedded systemsabstractModern embedded devices (e.g., PDAs, mobile phones) are now incorporating Java as a very popular implementation language in their designs. These new embedded systems include multiple applications that are dynamically launched by the user, which can produce very energy-hungry systems if the interactions between the applications and the garbage collectors (GCs) are not properly understood. In this paper we present a complete exploration, from an energy viewpoint, of the different possibilities of memory hierarchies for high-performance embedded systems when used by state-of-the-art GCs. Moreover, we explore the potential peformance improvement and energy reductions of using a scratchpad memory directed by the virtual machine to store critical code and data structures of the GCs; thus, enabling up to 40% performance improvements and 41% leakage reduction with respect to classical cache-based memory architectures. Our experimental results show that the key for an efficient low-power implementation of Java Virtual Machines (JVM) for high-performance embedded systems is the synergy between the GC choice, the memory architecture tuning, and the inclusion of power management schemes controlled by the JVM, exploiting knowledge of the used GC. José Manuel Velasco, David Atienza 0001, Katzalin Olcoz |
ACM Great Lakes Symposium on VLSI | 3 |
| 2004 | Adaptive Tuning of Reserved Space in an Appel Collector
José Manuel Velasco, Katzalin Olcoz, Francisco Tirado |
ECOOP | 2 |
| 1999 | Unified data path allocation and BIST intrusion
Katzalin Olcoz, Francisco Tirado, Hortensia Mecha |
Integr. | 1 |
| 1996 | A method for area estimation of data-path in high level synthesisabstractThis paper describes a new method to estimate the area of data paths generated during a High Level Synthesis (HLS) process, when the information concerning the circuit is not yet complete. Our method is more accurate and considers more factors than those used by other HLS systems of which we are aware. Our main concern is the interconnection area, often neglected by HLS systems, which has a strong influence on the final circuit area being optimized, as well as a high dependency on the technology used and on the circuit area itself. Predicting the area of a design layout with accuracy is important because it allows one to foresee whether the design will satisfy the area constraints, and will lend the allocator towards the best design among several possibilities with guarantees. Our estimations of the final standard-cell layout area are similar, or even more accurate, than those obtained following methods used by low-level design systems, which have much more information available. Due to the performance penalty their relatively high complexity will produce, these methods are unusable in an HLS system exploring a wide design space. Our estimation, on the contrary, has a low complexity and can be repeated time and again as the HLS design space is searched. Hortensia Mecha, Milagros Fernández, Francisco Tirado, Julio Septién, D. Motes, Katzalin Olcoz |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 1994 | Clock cycle estimation based on dead time and control unit area minimization
Hortensia Mecha, Milagros Fernández, Román Hermida, Daniel Mozos, Katzalin Olcoz |
Microprocess. Microprogramming | 5 |
| 1993 | Global hardware synthesis guided by realistic probability computation
Román Hermida, Daniel Mozos, Katzalin Olcoz |
Microprocess. Microprogramming | 4 |
| 1993 | Data path structures and heuristics for testable allocation in high level synthesis
Katzalin Olcoz, Francisco Tirado, Daniel Mozos, Julio Septién |
Microprocess. Microprogramming | 1 |