Lucia Pons

dblp:224/8716 · DBLP profile ↗
← Back
10ranked-venue papers
9as first author
8since 2021 · last 2026
0000-0002-4582-7744ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 8 first-author · 7 since 2021
YearPublicationVenuePosition
2026 WAPA: A Microarchitecture- and Workload-Agnostic Universal SMT Scheduler
Marta Navarro 0001, Vicent Pallardó-Julià, Lucia Pons, Salvador Petit, María Engracia Gómez, Julio Sahuquillo
IEEE Trans. Parallel Distributed Syst.3
2025 Advanced resource management: A hands-on master course in HPC and cloud computing
abstract
Resource management has become a major concern in dealing with performance and fairness in recent computing servers, including a wide variety of shared resources. To achieve high-performing and efficient systems, both hardware and software engineers must be thoroughly trained in effective resource management techniques. This paper introduces the GRE master course (Spanish acronym for Resource Management and Performance Evaluation in Cloud and High-Performance Workloads), which is being offered since Fall 2023. The course is taught by instructors with broad research expertise in resource management and performance evaluation. Subjects covered in this course include workload characterization, state-of-the-art resource management approaches, and performance evaluation tools and methodologies used in production systems. Management techniques are studied both in the context of HPC and cloud computing, where resource efficiency is becoming a primary concern. To enhance the learning experience, the course integrates theoretical concepts with a wide set of hands-on tasks carried out on recent real platforms. A real cloud virtualized environment is mimicked using typical software deployed in production systems such as Proxmox Virtual Environment. Students learn to use tools such as Linux Perf and Intel Vtune Profiler, which are commonly employed by researchers and practitioners to carry out typical tasks like performance bottleneck analysis from a microarchitectural perspective. Overall, the GRE course provides students with a solid foundation and skills in resource management by addressing current hot topics both in the industry and academia. Student satisfaction and learning outcomes prove the success of the GRE course and encourage us to continue in this direction.
Lucia Pons, Salvador Petit, Julio Sahuquillo
J. Parallel Distributed Comput.1
2024 A modular approach to build a hardware testbed for cloud resource management research
Lucia Pons, Salvador Petit, Julio Pons, María Engracia Gómez, Julio Sahuquillo
J. Supercomput.1
2023 Dynamic Allocation of Processor Cores to Graph Applications on Commodity Servers
abstract
Graph processing is increasingly adopted to solve problems that span many application domains, including scientific computing, social networks, and big-data analytics. These applications present particular features (huge working sets and irregular scalability) that make the default Linux scheduler, which adopts a time-sharing policy to provide a fair scheduler, perform poorly when co-locating multiple graph applications in the same processor. This work focuses on maximizing processor utilization, which is a major concern of current data centers. To this end, we propose AFAIR, a flexible scheduling policy that allocates multiple graph applications on the same processor and assigns a fraction of the cores exclusively to each application instead of sharing them. Moreover, AFAIR dynamically adds/removes cores to the running applications, adapting the number of threads used for parallel execution to balance memory load. This allows AFAIR to achieve almost perfect fairness, on average 95%.
Lucia Pons, Julio Sahuquillo, Timothy M. Jones 0001
PACT1
2023 Stratus: A Hardware/Software Infrastructure for Controlled Cloud Research
abstract
Cloud systems deploy a wide variety of shared resources and host a large number of tenant applications. To perform cloud research, a small experimental platform is commonly used, which hides the huge system complexity and provides flexibility. Despite being simpler, this platform should include the main cloud system components (hardware and software) to provide representative results. A wide set of platforms have spread in recent years; however, most of them only include a major cloud component or lack the deployment of virtual machines (VMs) to provide isolation. This paper presents Stratus, an experimental platform that is currently being used to carry out cloud research. To the best of our knowledge, Stratus is the only platform that jointly provides three main features: uses VMs to isolate tenant applications, deploys the three types of cloud nodes (server, client, and storage), and manages all main shared system resources (CPUs, LLC space, memory, network, and disk bandwidth). Moreover, Stratus implements a software manager to ease the research and aid the design of QoS-aware policies. The manager integrates three main functionalities: management and control of the execution of VMs and running applications, monitoring of hardware performance counters and system resource utilization, and partitioning of the main shared system resources by using technologies available in commercial processors.
Lucia Pons, Salvador Petit, Julio Pons, María Engracia Gómez, Chaoyi Huang, Julio Sahuquillo
PDP1
2023 Cloud White: Detecting and Estimating QoS Degradation of Latency-Critical Workloads in the Public Cloud
abstract
The increasing popularity of cloud computing has forced cloud providers to build economies of scale to meet the growing demand. Nowadays, data-centers include thousands of physical machines, each hosting many virtual machines (VMs), which share the main system resources, causing interference that can significantly impact on performance. Frequently, these data-centers run latency-critical workloads, whose performance is determined by tail latency, which is very sensitive to the interference of co-running workloads. To prevent QoS violations, cloud providers adopt overprovisioning strategies but they reduce the server utilization and increase the costs. A mechanism that accurately estimates performance degradation dynamically in a production system would allow cloud providers to improve the servers’ utilization. In this work we propose Cloud White, an approach that is able to detect the inter-VM interference in scenarios with multiple co-located latency-critical VMs and estimate the performance degradation using multi-variable regression models. Unlike previous proposals, Cloud White is built taking into account the limitations of a public cloud production system. Experimental results show that Cloud White is able to estimate performance degradation with a small overall prediction error of 5%.
Lucia Pons, Josué Feliu, Julio Sahuquillo, María Engracia Gómez, Salvador Petit, Julio Pons, Chaoyi Huang
Future Gener. Comput. Syst.1
2022 Cache-Poll: Containing Pollution in Non-Inclusive Caches Through Cache Partitioning
abstract
Current server processors have redistributed the cache hierarchy space over previous generations. The private L2 cache has been made larger and the shared last level caches (LLC) smaller but designed as non-inclusive to reduce the number of replicated blocks. As a result, the new organization shrinks the per-core cache area.
Lucia Pons, Julio Sahuquillo, Salvador Petit, Julio Pons
ICPP1
2022 Effect of Hyper-Threading in Latency-Critical Multithreaded Cloud Applications and Utilization Analysis of the Major System Resources
abstract
Multithreaded latency-critical applications represent an important subset of workloads running on public cloud systems. Most of these systems deploy powerful computing servers including Intel Hyper-Threading processors. Understanding how performance is affected by the consumption of the main system resources is a major concern for cloud providers in order to devise virtualization strategies that improve the system efficiency. With this aim, this paper first characterizes the impact of QPS on tail latency, analyzing different scenarios varying the number of threads and the thread-to-core allocation (single-task and multi-task execution) policy. The characterization study reveals that the performance of some applications does not scale with the number of threads, and the performance of some others is insensitive to the Hyper-Threading technology, so they can be allocated in less physical cores and improve system utilization. Identifying these applications, however, at run-time is challenging. Despite identifying these applications at run-time is challenging, this paper shows that they can be successfully detected at run-time by analyzing the utilization trend of the major system resources. In addition to CPU, we have also studied how assigning the share of each application of other major shared system resources impacts on performance. We outline considerations cloud providers should take into account to improve performance and resource utilization.
Lucia Pons, Josué Feliu, José Puche, Chaoyi Huang, Salvador Petit, Julio Pons, María Engracia Gómez, Julio Sahuquillo
Future Gener. Comput. Syst.1
2020 Phase-Aware Cache Partitioning to Target Both Turnaround Time and System Performance
abstract
The Last Level Cache (LLC) plays a key role in the system performance of current multi-cores by reducing the number of long latency main memory accesses. The inter-application interference at this shared resource, however, can lead the system to undesired situations regarding performance and fairness. Recent approaches have successfully addressed fairness and turnaround time (TT) in commercial processors. Nevertheless, these approaches must face sustaining system performance, which is challenging. This work makes two main contributions. LLC behaviors regarding cache performance, data reuse and cache occupancy, that adversely impact on the final performance are identified. Second, based on these behaviors, we propose the Critical-Phase Aware Partitioning Approach (CPA), which reduces TT while sustaining (and even improving) IPC by making an effective use of the LLC space. Experimental results show that CPA outperforms CA, Dunn and KPart state-of-the-art approaches, and improves TT (over 40 percent in some workloads) over Linux default behavior while sustaining or even improving IPC by more than 3 percent in several mixes.
Lucia Pons, Julio Sahuquillo, Vicent Selfa, Salvador Petit, Julio Pons
IEEE Trans. Parallel Distributed Syst.1
2018 Improving System Turnaround Time with Intel CAT by Identifying LLC Critical Applications
Lucia Pons, Vicent Selfa, Julio Sahuquillo, Salvador Petit, Julio Pons
Euro-Par1