EDBT 2026 Demo / reviewers in the wild / expert
Luis Costero
dblp:168/8810 · also Luis Costero Valero
· DBLP profile ↗
18ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-6922-2520ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 4 first-author · 12 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CloudFormer: An Attention-Based Performance Prediction for Public Clouds with Unknown WorkloadabstractCloud platforms are increasingly relied upon to host diverse, resource-intensive workloads due to their scalability, flexibility, and cost-efficiency. In multi-tenant cloud environments, virtual machines are consolidated on shared physical servers to improve resource utilization. While virtualization guarantees resource partitioning for CPU, memory, and storage, it cannot ensure performance isolation. Competition for shared resources such as last-level cache, memory bandwidth, and network interfaces often leads to severe performance degradation. Existing management techniques, including VM scheduling and resource provisioning, require accurate performance prediction to mitigate interference. However, this remains challenging in public clouds due to the black-box nature of VMs and the highly dynamic nature of workloads. To address these limitations, we propose CloudFormer, a dual-branch Transformer-based model designed to predict VM performance degradation in black-box environments. CloudFormer jointly models temporal dynamics and system-level interactions, leveraging 206 system metrics at one-second resolution across both static and dynamic scenarios. This design enables the model to capture transient interference effects and adapt to varying workload conditions without scenario-specific tuning. Complementing the methodology, we provide a fine-grained dataset that significantly expands the temporal resolution and metric diversity compared to existing benchmarks. Experimental results demonstrate that CloudFormer consistently outperforms state-of-the-art baselines across multiple evaluation metrics, achieving robust generalization across diverse and previously unseen workloads. Notably, CloudFormer attains a mean absolute error (MAE) of just 7.8%, representing a substantial improvement in predictive accuracy and outperforming existing methods at least by 28%. Amirhossein Shahbazinia, Darong Huang 0003, Luis Costero, David Atienza 0001 |
CCGrid | 3 |
| 2026 | ETLA-3D: Equivalent Thin Layer Aggregation based Thermal FEM for Hybrid Bonding F2F 3D ICsabstractIn 3D face-to-face (F2F) hybrid bonding ICs, sub-micrometer thin layers lead to an extreme aspect ratio between the lateral dimensions and the vertical thickness. This poses major challenges for finite element method (FEM) thermal simulation. To address this, we introduce ETLA-3D, a thermal FEM methodology based on equivalent thin-layer aggregation, designed specifically for hybrid bonding F2F 3D ICs. The method consolidates the physical properties of thin layers into their neighboring layers by introducing new integral terms into the FEM weak form, greatly reducing the complexity of meshing, the simulation degrees of freedom (DoFs) and the computational cost, while preserving accuracy. Experimental results show that ETLA-3D achieves up to 695.8 × faster runtime compared to the commercial FEM tool (COMSOL Multiphysics), with a maximum absolute error of less than 1.1°C. By combining high accuracy with exceptional efficiency, ETLA-3D establishes a reliable and efficient FEM framework to model the thermal behavior of F2F 3D ICs. Zhen Zhuang, Darong Huang 0003, Luis Costero, Rongmei Chen, David Atienza 0001, Tsung-Yi Ho |
DATE | 5 |
| 2026 | 3D-ICE 4.0: Accurate and efficient thermal modeling for 2.5D/3D heterogeneous chiplet systemsabstractThe increasing power densities and intricate heat dissipation paths in advanced 2.5D/3D chiplet systems necessitate thermal modeling frameworks that deliver detailed thermal maps with high computational efficiency. Traditional compact thermal models (CTMs) often struggle to scale with the complexity and heterogeneity of modern architectures. This work introduces 3D-ICE 4.0, designed for heterogeneous chip-based systems. Key innovations include: (i) preservation of material heterogeneity and anisotropy directly from industrial layouts, integrated with OpenMP and SuperLU MT-based parallel solvers for scalable performance, (ii) adaptive vertical layer partitioning to accurately model vertical heat conduction, and (iii) temperature-aware non-uniform grid generation. The results with different benchmarks demonstrate that 3D-ICE 4.0 achieves speedups ranging from 3.61x-6.46x over state-of-the-art tools, while reducing grid complexity by more than 23.3% without compromising accuracy. Compared to the commercial software COMSOL, 3D-ICE 4.0 effectively captures both lateral and vertical heat flows, validating its precision and robustness. These advances demonstrate that 3D-ICE 4.0 is an efficient solution for thermal modeling in emerging heterogeneous 2.5D/3D integrated systems. Darong Huang 0003, Luis Costero, David Atienza 0001 |
DATE | 3 |
| 2026 | Solving the task scheduling and GPU reconfiguration problem on MIG devices via deep reinforcement learningabstract• Multi-Instance GPU (MIG) technology enables adaptive co-execution of tasks, greatly improving the efficiency of computational resources in a flexible manner. • Prior methods simplify the MIG scheduling and dynamic reconfiguration challenge by reducing problem complexity, but can yield markedly suboptimal solutions in certain scenarios. • Modeling the problem with Reinforcement Learning (RL), and training with Deep Learning techniques, allows to approach it successfully without great simplifications, despite its high dimensionality. The design of the RL agent also needs to be carefully refined, with analysis such as that detailed in the manuscript, which can serve as a guide for similar resource management work. • Our refined RL agent reduces makespan by 2–7% over state-of-the-art on a wide set of benchmarks and synthetic workloads, with improvements of up to 30% in specific cases. Additional benefits include enhanced flexibility and adaptability of the scheduling framework. Recent advances in dynamic GPU partitioning, such as NVIDIA’s Multi-Instance GPU (MIG) technology, have enhanced resource utilization by enabling task co-execution without contention. However, existing MIG schedulers remain limited to static or task-agnostic methods that sacrifice optimality for tractability. This paper presents a Deep Reinforcement Learning framework that seeks to minimize the completion time of a task queue by holistically addressing the dimensions of the problem: task molding, GPU reconfiguration and execution order. To manage the vast solution space, we apply optimizations such as discrete and canonical representation of states, unification of equivalent configurations, action masking, or promoting the exploration of reconfigurations; this offers insights for similar resource management scenarios. The proposed models are extensively evaluated with widely used benchmarks of the Rodinia and Altis suites, and synthetic workloads generated to emulate a wide range of plausible real situations. The final model improves to the state-of-the-art, especially in workloads that clearly contradict the assumptions of previous proposals, achieving a difference of less than 20% to the optimum. Additionally, two different approaches to the problem are faced (offline vs. online), discussing their theoretical advantages and disadvantages, and evaluating them experimentally for the final model. Jorge Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz |
Future Gener. Comput. Syst. | 2 |
| 2026 | A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologies
Jorge Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz |
J. Supercomput. | 2 |
| 2025 | QoS-aware workload scheduling on heterogeneous and dynamic edge-to-cloud deploymentsabstractWith the advent of the edge-to-cloud continuum paradigm, the heterogeneity of compute entities imposes new challenges to achieve proper levels of Quality of Service (QoS) in the mapping of workloads and requests to nodes, mainly in terms of response times, but also in terms of application-specific metrics. In addition, for deployments in which computing elements are dynamic by nature (in terms of availability, but also compute capabilities or latency, to name only a few parameters), placing workloads in the most suitable compute element becomes a major challenge. Current generic orchestrators, such as Kubernetes, have demonstrated to be a valid option in homogeneous and static deployments, in which QoS-oblivious scheduling policies typically pursue exclusively load balancing, and ignore other parameters such as latency reduction or limitation of application-level metrics. In this paper, we demonstrate that QoS-oblivious policies in a generic orchestrator such as Kubernetes are not enough for heterogeneous and dynamic edge-to-cloud deployments. We propose a new framework that integrates seamlessly into Kubernetes by means of a service. This framework is equipped with a set of QoS-aware scheduling policies to tackle the heterogeneity and dynamic character of many edge-to-cloud deployments. Our experimental results for a specific use case (a deployment of inference servers across heterogeneous nodes) reveal significant gains in terms of response times under different dynamic scenarios that include computing devices with different capabilities (multi-core CPUs and different types of GPUs). Julián Cámara-Miró, Luis Costero, Francisco D. Igual |
PDP | 2 |
| 2025 | Leveraging Multi-Instance GPUs through moldable task schedulingabstractNVIDIA MIG (Multi-Instance GPU) allows partitioning a physical GPU into multiple logical instances with fully-isolated resources, which can be dynamically reconfigured. This work highlights the untapped potential of MIG through moldable task scheduling with dynamic reconfigurations . Specifically, we propose a makespan minimization problem for multi-task execution under MIG constraints. Our profiling shows that assuming monotonicity in task work with respect to resources is not viable, as is usual in multicore scheduling. Relying on a state-of-the-art proposal that does not require such an assumption, we present FAR , a 3-phase algorithm to solve the problem. Phase 1 of FAR builds on a classical task moldability method, phase 2 combines Longest Processing Time First and List Scheduling with a novel repartitioning tree heuristic tailored to MIG constraints, and phase 3 employs local search via task moves and swaps. FAR schedules tasks in batches offline, concatenating their schedules on the fly in an improved way that favors resource reuse. Excluding reconfiguration costs, the List Scheduling proof shows an approximation factor of 7/4 on the NVIDIA A30 model. We adapt the technique to the particular constraints of an NVIDIA A100/H100 to obtain an approximation factor of 2. Including the reconfiguration cost, our real-world experiments reveal a makespan with respect to the optimum no worse than 1.22× for a well-known suite of benchmarks, and 1.10× for synthetic inputs inspired by real kernels. We obtain good experimental results for each batch of tasks, but also in the concatenation of batches, with large improvements over the state-of-the-art and proposals without GPU reconfiguration. Moreover, we show that the proposed heuristics allow a correct adaptation to tasks of very different characteristics. Beyond the specific algorithm, the paper demonstrates the research potential of the MIG technology and suggests useful metrics, workload characterizations and evaluation techniques for future work in this field. Jorge Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz |
J. Parallel Distributed Comput. | 2 |
| 2025 | Balanced segmentation of CNNs for multi-TPU inference
John S. Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz |
J. Supercomput. | 2 |
| 2024 | Is the powersave governor really saving power?abstractA frequency scaling governor is critical for the performance management of cloud servers, as it enhances energy efficiency and helps to control operational temperatures, thereby ensuring system reliability. However, our in-depth analysis of the application’s performance and Dynamic Voltage and Frequency Scaling (DVFS) actions, alongside assessments of server power consumption and operating temperature, indicates that existing Linux scaling governors often fall into non-optimal DVFS strategies, especially for cloud applications with varying workloads and requests. This shortfall comes from the misleading CPU load metrics, which fail to accurately capture the applications’ true performance requirements and demands. In this context, we introduce a novel scaling governor named GreenDVFS. First, it identifies the optimal frequencies for the application in a range of workload scenarios. Optimal frequencies are used to maintain application performance, reduce server power consumption, and maintain a balanced operating temperature in different workload scenarios. Furthermore, we design a long short-term memory (LSTM)-based time series methodology to detect the real-time workloads of cloud applications accurately and timely. Building on these foundations, the proposed method takes optimal DVFS actions, tailored for cloud applications under different workload conditions, to optimize performance, energy efficiency, and temperature. The experimental results highlight the effectiveness of the proposed GreenDVFS, with up to 18% savings in energy consumption and a 30% decrease in operational temperature by comparing against the default Linux governor, all while not compromising the application’s performance. Such improvements help to optimize cloud computing operations for enhanced efficiency and sustainability. Darong Huang 0003, Luis Costero, David Atienza 0001 |
CCGrid | 2 |
| 2024 | Intermediate Address Space: virtual memory optimization of heterogeneous architectures for cache-resident workloadsabstractThe increasing demand for computing power and the emergence of heterogeneous computing architectures have driven the exploration of innovative techniques to address current limitations in both the compute and memory subsystems. One such solution is the use of Accelerated Processing Units (APUs), processors that incorporate both a central processing unit (CPU) and an integrated graphics processing unit (iGPU). However, the performance of both APU and CPU systems can be significantly hampered by address translation overhead, leading to a decline in overall performance, especially for cache-resident workloads. To address this issue, we propose the introduction of a new intermediate address space (IAS) in both APU and CPU systems. IAS serves as a bridge between virtual address (VA) spaces and physical address (PA) spaces, optimizing the address translation process. In the case of APU systems, our research indicates that the iGPU suffers from significant translation look-aside buffer (TLB) misses in certain workload situations. Using an IAS, we can divide the initial address translation into front- and back-end phases, effectively shifting the bottleneck in address translation from the cache side to the memory controller side, a technique that proves to be effective for cache-resident workloads. Our simulations demonstrate that implementing IAS in the CPU system can boost performance by up to 40% compared to conventional CPU systems. Furthermore, we evaluate the effectiveness of APU systems, comparing the performance of IAS-based systems with traditional systems, showing up to a 185% improvement in APU system performance with our proposed IAS implementation. Furthermore, our analysis indicates that over 90% of TLB misses can be filtered by the cache, and employing a larger cache within the system could potentially result in even greater improvements. The proposed IAS offers a promising and practical solution to enhance the performance of both APU and CPU systems, contributing to state-of-the-art research in the field of computer architecture. Qunyou Liu, Darong Huang 0003, Luis Costero, Marina Zapater, David Atienza 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | An Evaluation Framework for Dynamic Thermal Management Strategies in 3D MultiProcessor System-on-Chip Co-DesignabstractDynamic thermal management (DTM) has been widely adopted to improve the energy efficiency, reliability, and performance of modern Multi-Processor SoCs (MPSoCs). However, the evolving industry trends and heterogeneous architecture designs have introduced significant challenges in state-of-the-art DTM methods. Specifically, the emergence of heterogeneous design has led to increased localized and non-uniform hotspots, necessitating accurate and responsive DTM strategies. Additionally, the increased number of cores to be managed requires the DTM to optimize and coordinate the whole system. However, existing methodologies fail in both precise thermal modeling in localized hotspots and fast architecture simulation. To tackle these existing challenges, we first introduce the latest version of 3D-ICE 3.1, with a novel non-uniform thermal modeling technique to support customized discretization levels of thermal grids. 3D-ICE 3.1 improves the accuracy of thermal analysis and reduces simulation overhead. Then, in conjunction with an efficient and fast offline application profiling strategy utilizing the architecture simulator gem5-X, we propose a novel DTM evaluation framework. This framework enables us to explore novel DTM methods to optimize the energy efficiency, reliability, and performance of contemporary 3D MPSoCs. The experimental results demonstrate that 3D-ICE 3.1 achieves high accuracy, with only 0.3K mean temperature error. Subsequently, we evaluate various DTM methods and propose a Multi-Agent Reinforcement Learning (MARL) control to address the demanding thermal challenges of 3D MPSoCs. Our experimental results show that the proposed DTM method based on MARL can reduce power consumption by 13% while maintaining a similar performance level to the comparison methods. Darong Huang 0003, Luis Costero, David Atienza 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | CloudProphet: A Machine Learning-Based Performance Prediction for Public CloudsabstractComputing servers have played a key role in developing and processing emerging compute-intensive applications in recent years. Consolidating multiple virtual machines (VMs) inside one server to run various applications introduces severe competence for limited resources among VMs. Many techniques such as VM scheduling and resource provisioning are proposed to maximize the cost-efficiency of the computing servers while alleviating the performance inference between VMs. However, these management techniques require accurate performance prediction of the application running inside the VM, which is challenging to get in the public cloud due to the black-box nature of the VMs. From this perspective, this paper proposes a novel machine learning-based performance prediction approach for applications running in the cloud. To achieve high-accuracy predictions for black-box VMs, the proposed method first identifies the running application inside the virtual machine. It then selects highly correlated runtime metrics as the input of the machine learning approach to accurately predict the performance level of the cloud application. Experimental results with state-of-the-art cloud benchmarks demonstrate that our proposed method outperforms existing prediction methods by more than 2× in terms of the worst prediction error. In addition, we successfully tackle the challenge of performance prediction for applications with variable workloads by introducing the performance degradation index, which other comparison methods fail to consider. The workflow versatility of the proposed approach has been verified with different modern servers and VM configurations. Darong Huang 0003, Luis Costero, Ali Pahlevan, Marina Zapater, David Atienza 0001 |
IEEE Trans. Sustain. Comput. | 2 |
| 2023 | Improving inference time in multi-TPU systems with profiled model segmentationabstractIn this paper, we systematically evaluate the inference performance of the Edge TPU by Google for neural networks with different characteristics. Specifically, we determine that, given the limited amount of on-chip memory on the Edge TPU, accesses to external (host) memory rapidly become an important performance bottleneck. We demonstrate how multiple devices can be jointly used to alleviate the bottleneck introduced by accessing the host memory. We propose a solution combining model segmentation and pipelining on up to four TPUs, with remarkable performance improvements that range from 6x for neural networks with convolutional layers to 46x for fully connected layers, compared with single-TPU setups. Jorge Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz |
PDP | 2 |
| 2022 | Reinforcement Learning-Based Joint Reliability and Performance Optimization for Hybrid-Cache Computing ServersabstractComputing servers play a key role in the development and process of emerging compute-intensive applications in recent years. However, they need to operate efficiently from an energy perspective viewpoint, while maximizing the performance and lifetime of the hottest server components (i.e., cores and cache). Previous methods focused on either improving energy efficiency by adopting new hybrid-cache architectures including the resistive random-access memory (RRAM) and static random-access memory (SRAM) at the hardware level, or exploring tradeoffs between lifetime limitation and performance of multicore processors under stable workloads conditions. Therefore, no work has so far proposed a co-optimization method with hybrid-cache-based server architectures for real-life dynamic scenarios taking into account scalability, performance, lifetime reliability, and energy efficiency at the same time. In this article, we first formulate a reliability model for the hybrid-cache architecture to enable precise lifetime reliability management and energy efficiency optimization. We also include the performance and energy overheads of cache switching, and optimize the benefits of hybrid-cache usage for better energy efficiency and performance. Then, we propose a runtime$q$-learning-based reliability management and performance optimization approach for multicore microprocessors with the hybrid-cache architecture, jointly incorporated with a dynamic preemptive priority queue management method to improve the overall tasks’ performance by targeting to respect their end time limits. Experimental results show that our proposed method achieves up to 44% average performance (i.e., tasks execution time) improvement, while maintaining the whole system design lifetime longer than five years, when compared to the latest state-of-the-art energy efficiency optimization and reliability management methods for computing servers. Darong Huang 0003, Ali Pahlevan, Luis Costero, Marina Zapater, David Atienza 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Leveraging knowledge-as-a-service (KaaS) for QoS-aware resource management in multi-user video transcoding
Luis Costero, Francisco D. Igual, Katzalin Olcoz, Francisco Tirado |
J. Supercomput. | 1 |
| 2020 | Resource Management for Power-Constrained HEVC Transcoding Using Reinforcement LearningabstractThe advent of online video streaming applications and services along with the users' demand for high-quality contents require High Efficiency Video Coding (HEVC), which provides higher video quality and more compression at the cost of increased complexity. On one hand, HEVC exposes a set of dynamically tunable parameters to provide trade-offs among Quality-of-Service (QoS), performance, and power consumption of multi-core servers on the video providers' data center. On the other hand, resource management of modern multi-core servers is in charge of adapting system-level parameters, such as operating frequency and multithreading, to deal with concurrent applications and their requirements. Therefore, efficient multi-user HEVC streaming necessitates joint adaptation of application-and system-level parameters. Nonetheless, dealing with such a large and dynamic design space is challenging and difficult to address through conventional resource management strategies. Thus, in this work, we develop a multi-agent Reinforcement Learning framework to jointly adjust application-and system-level parameters at runtime to satisfy the QoS of multi-user HEVC streaming in power-constrained servers. In particular, the design space, composed of all design parameters, is split into smaller independent sub-spaces. Each design sub-space is assigned to a particular agent so that it can explore it faster, yet accurately. The benefits of our approach are revealed in terms of adaptability and quality (with up to to 4× improvements in terms of QoS when compared to a static resource management scheme), and learning time (6× fasterthan an equivalent mono-agent implementation). Finally, we show that the power-capping techniques formulated outperform the hardware-based power capping with respect to quality. Luis Costero, Arman Iranfar, Marina Zapater, Francisco D. Igual, Katzalin Olcoz, David Atienza 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2019 | MAMUT: Multi-Agent Reinforcement Learning for Efficient Real-Time Multi-User Video TranscodingabstractReal-time video transcoding has recently raised as a valid alternative to address the ever-increasing demand for video contents in servers' infrastructures in current multi-user environments. High Efficiency Video Coding (HEVC) makes efficient online transcoding feasible as it enhances user experience by providing the adequate video configuration, reduces pressure on the network, and minimizes inefficient and costly video storage. However, the computational complexity of HEVC, together with its myriad of configuration parameters, raises challenges for power management, throughput control, and Quality of Service (QoS) satisfaction. This is particularly challenging in multi-user environments where multiple users with different resolution demands and bandwidth constraints need to be served simultaneously. In this work, we present MAMUT, a multi-agent machine learning approach to tackle these challenges. Our proposal breaks the design space composed of run-time adaptation of the transcoder and system parameters into smaller sub-spaces that can be explored in a reasonable time by individual agents. While working cooperatively, each agent is in charge of learning and applying the optimal values for internal HEVC and system-wide parameters. In particular, MAMUT dynamically tunes Quantization Parameter, selects number of threads per video, and sets the operating frequency with throughput and video quality objectives under compression and power consumption constraints. We implement MAMUT on an enterprise multicore server and compare equivalent scenarios to state-of-the-art alternative approaches. The obtained results reveal that MAMUT consistently attains up to 8× improvement in terms of FPS violations (and thus Quality of Service), 24% power reduction, as well as faster and more accurate adaptation both to the video contents and available resources. Luis Costero, Arman Iranfar, Marina Zapater, Francisco D. Igual, Katzalin Olcoz, David Atienza 0001 |
DATE | 1 |
| 2017 | Revisiting conventional task schedulers to exploit asymmetry in multi-core architectures for dense linear algebra operations
Luis Costero, Francisco D. Igual, Katzalin Olcoz, Sandra Catalán, Rafael Rodríguez-Sánchez 0001, Enrique S. Quintana-Ortí |
Parallel Comput. | 1 |