EDBT 2026 Demo / reviewers in the wild / expert
Vinicius Petrucci
dblp:80/2060
· DBLP profile ↗
22ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-5671-7575ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 4 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bandwidth Speaks, We Listen: Dynamic Memory Interleaving for TieringabstractCXL memory devices increase the memory capacity and bandwidth available to a server, at the cost of higher access latency. Prior research focused on how to make use of the expanded memory capacity provided by CXL while minimizing the impact of its higher access latency. However, most current memory management techniques to achieve this, such as memory tiering, do not make effective use of the expanded bandwidth provided by CXL memory. These techniques place frequently accessed data in local memory, so bandwidth intensive applications will saturate the local bandwidth and leave the remote bandwidth unutilized. Bijan Tabatabai, Eishan Mirakhur, Ravi Shankar Jonnalagadda, Vinicius Petrucci, Rohit Sehgal, Jus Singh, Michael M. Swift |
ISMM | 4 |
| 2022 | The Effect of Animations Using Real-world Analogies on Diverse Computer Systems StudentsabstractIt is a challenge to engage students when teaching them abstract and complex computer systems concepts, such as buffer overflow, memory management, concurrent execution, and process synchronization. Past research has shown that interactive animation and real-life analogies make STEM concepts more approachable and help students achieve better learning outcomes. Based on these findings, we introduce interactive analogies into learning the concept of buffer overflow. More specifically, we created a dry-cleaning shop animation tool (https://scratch.mit.edu/projects/571317697/) targeting K-12 and undergraduate students. To assess the effectiveness of our tool, we are in the process of conducting a user study, in which students use our animation tool to learn about buffer overflow and take pre- and post-assessment on the concept. Our goal is to make CS learning more accessible to diverse students, regardless of their background and age. Rachel Puckett, Wonsun Ahn, Sherif M. Khattab, Luis Oliveira 0002, Vinicius Petrucci |
SIGCSE (2) | 6 |
| 2022 | Laundry Overflow: Engaging Diverse Students in CyberSecurity using Interactive AnalogiesabstractPast research has shown interactive animations, and those that use real-life analogies in particular, can play an important role in providing the intuition required to understand Computer Science concepts. Nonetheless, the use of analogies continues to be under-explored in CS education compared to other STEM fields. To break the impasse, we aim to create and evaluate a set of interactive animations based on analogies to understand their efficacy. As our first addition, we have created an animation explaining the Buffer Overflow computer systems concept. Explaining the concept abstractly has had a track record of ineffectiveness in our department, since the concept of computer memory as a series of contiguous storage locations is so foreign to students. Instead, the animation uses the analogy of a dry-cleaning shop with a series of hangers to provide a concrete mental picture of computer memory. Students explore various dry-cleaning scenarios, in which customers drop off and pick up their laundry, to understand at their own pace when buffer overflows cause harm and when they are silently ignored. This animation: https://scratch.mit.edu/projects/571317697/ (and others) will be provided as an open educational resource to instructors to encourage the use of interactive analogies in their teaching, and to undergraduate and K-12 students. Rachel Puckett, Wonsun Ahn, Sherif M. Khattab, Luis Oliveira 0002, Vinicius Petrucci |
SIGCSE (2) | 6 |
| 2021 | Profile-guided Frequency Scaling for Latency-Critical Search WorkloadsabstractDynamic frequency scaling is a technique to reduce power consumption in computer systems. However, this technique poses challenges when adopted in latency-critical applications. Prior work on dynamic frequency scaling is application agnostic and coarse-granulated in the sense that it considers the entire application process utilization for decision making, without the distinction between individual threads or functions.This work proposes a finer-grained dynamic frequency scaling approach for multi-core processors that leverages information about the computational intensity of certain functions in a latency-critical web search application. First, our approach profiles the running application to identify hot functions for typical workloads. Next, a run-time scheme is devised to adapt the individual core frequency whenever a compute-intensive thread enters or exits a hot function. We implemented and evaluated our proposal in a real multi-core system. We observed energy consumption savings up to 28% when compared to the recent Linux's Ondemand frequency scaling governor, while attaining acceptable levels of tail latency constraints. Daniel Araújo de Medeiros, Denilson das Mercês Amorim, Vinicius Petrucci |
CCGRID | 3 |
| 2021 | Improving HPC System Throughput and Response Time using Memory DisaggregationabstractHPC clusters are cost-effective, well understood, and scalable, but the rigid boundaries between compute nodes may lead to poor utilization of compute and memory resources. HPC jobs may vary, by orders of magnitude, in memory consumption per core. Thus, even when the system is provisioned to accommodate normal and large capacity nodes, a mismatch between the system and the memory demands of the scheduled jobs can lead to inefficient usage of both memory and compute resources. Disaggregated memory has recently been proposed as a way to mitigate this problem by flexibly allocating memory capacity across cluster nodes. This paper presents a simulation approach for at-scale evaluation of job schedulers with disaggregated memories and it introduces a new disaggregated-aware job allocation policy for the Slurm resource manager. Our results show that using disaggregated memories, depending on the imbalance between the system and the submitted jobs, a similar throughput and job response time can be achieved on a system with up to 33% less total memory provisioning. Felippe Vieira Zacarias, Paul M. Carpenter, Vinicius Petrucci |
ICPADS | 3 |
| 2021 | Heterogeneous Quasi-Partitioned SchedulingabstractWe consider the problem of scheduling a set of preemptible independent periodic implicit-deadline hard real-time tasks on heterogeneous processors. We divide this problem into two sub-problems: (a) assigning portions of each processor (offline) to each task without jeopardizing schedulability; and (b) generating a schedule satisfying the assigned portions using an online semi-partitioned scheduler, called Heterogeneous Quasi-Partitioned Scheduling (hQPS). The scheduler handles task servers at run-time for ensuring that the processor shares assigned to tasks are timely available to them. Assessments indicate that the proposed solution (i) has good scalability (up to 64 tasks, 64 processors), (ii) is effective in generating schedules with few preemptions and few migrations, and (iii) is effective in managing resources; for task sets where an extra processor speed is required, our solution needs at most 10% extra compared to an optimal scheduler. Ernesto Massa, George Lima 0001, Björn Andersson, Vinicius Petrucci |
RTSS | 4 |
| 2021 | Intelligent colocation of HPC workloads
Felippe Vieira Zacarias, Vinicius Petrucci, Rajiv Nishtala, Paul M. Carpenter, Daniel Mossé |
J. Parallel Distributed Comput. | 2 |
| 2021 | Mapping Computations in Heterogeneous Multicore Systems with Statistical Regression on Program InputsabstractA hardware configuration is a set of processors and their frequency levels in a multicore heterogeneous system. This article presents a compiler-based technique to match functions with hardware configurations. Such a technique consists of using multivariate linear regression to associate function arguments with particular hardware configurations. By showing that this classification space tends to be convex in practice, this article demonstrates that linear regression is not only an efficient tool to map computations to heterogeneous hardware, but also an effective one. To demonstrate the viability of multivariate linear regression as a way to perform adaptive compilation for heterogeneous architectures, we have implemented our ideas onto the Soot Java bytecode analyzer. Code that we produce can predict the best configuration for a large class of Java and Scala benchmarks running on an Odroid XU4 big.LITTLE board; hence, outperforming prior techniques such as ARM’s GTS and CHOAMP, a recently released static program scheduler. Junio Cezar R. da Silva, Lorena Leão, Vinicius Petrucci, Abdoulaye Gamatié, Fernando Magno Quintão Pereira |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2020 | Twig: Multi-Agent Task Management for Colocated Latency-Critical Cloud ServicesabstractMany of the important services running on data centres are latency-critical, time-varying, and demand strict user satisfaction. Stringent tail-latency targets for colocated services and increasing system complexity make it challenging to reduce the power consumption of data centres. Data centres typically sacrifice server efficiency to maintain tail-latency targets resulting in an increased total cost of ownership. This paper introduces Twig, a scalable quality-of-service (QoS) aware task manager for latency-critical services co-located on a server system. Twig successfully leverages deep reinforcement learning to characterise tail latency using hardware performance counters and to drive energy-efficient task management decisions in data centres. We evaluate Twig on a typical data centre server managing four widely used latency-critical services. Our results show that Twig outperforms prior works in reducing energy usage by up to 38% while achieving up to 99% QoS guarantee for latency-critical services. Rajiv Nishtala, Vinicius Petrucci, Paul M. Carpenter, Magnus Själander |
HPCA | 2 |
| 2019 | Compiler-assisted adaptive program scheduling in big.LITTLE systems: posterabstractEnergy-aware architectures provide applications with a mix of low and high frequency cores. Selecting the best core configurations for running programs is very challenging. Here, we leverage compilation, runtime monitoring and machine learning to map program phases to their best matching configurations. As a proof-of-concept, we devise the Astro system to show that our approach can outperform a state-of-the-art Linux scheduler for heterogeneous architectures. Marcelo Novaes, Vinicius Petrucci, Abdoulaye Gamatié, Fernando Magno Quintão Pereira |
PPoPP | 2 |
| 2019 | Intelligent Colocation of Workloads for Enhanced Server EfficiencyabstractMany server applications achieve only a fraction of their theoretical peak performance due to bottlenecks in the shared caches, instruction execution units, I/O or memory bandwidth, even though the remaining resources may be underutilized. It is very hard for developers and runtime systems to ensure that all these critical resources are fully exploited by a single application. An attractive technique for increasing server system utilization is to colocate multiple applications on the same server. When applications share critical resources, however, these applications may adversely affect each other, due to contention on the shared resources. In this paper, we show that server efficiency can be improved by modeling the expected performance degradation of colocated applications from measured hardware performance counters, and exploiting such a model to determine an optimized mix of colocated applications. This paper presents a novel resource management approach and makes the following contributions: (1) a new machine learning model to predict the performance degradation of colocated applications from hardware counters and (2) an intelligent scheduling scheme deployed on an existing resource manager to enable application co-scheduling with minimum performance degradation. Our results show that our approach achieves performance improvements of 15 % (avg) and 26 % (max) compared to the standard policy commonly used by existing job managers. Felippe Vieira Zacarias, Vinicius Petrucci, Rajiv Nishtala, Paul M. Carpenter, Daniel Mossé |
SBAC-PAD | 2 |
| 2017 | Hipster: Hybrid Task Manager for Latency-Critical Cloud WorkloadsabstractIn 2013, U.S. data centers accounted for 2.2% of the country's total electricity consumption, a figure that is projected to increase rapidly over the next decade. Many important workloads are interactive, and they demand strict levels of quality-of-service (QoS) to meet user expectations, making it challenging to reduce power consumption due to increasing performance demands. This paper introduces Hipster, a technique that combines heuristics and reinforcement learning to manage latency-critical workloads. Hipster's goal is to improve resource efficiency in data centers while respecting the QoS of the latency-critical workloads. Hipster achieves its goal by exploring heterogeneous multi-cores and dynamic voltage and frequency scaling (DVFS). To improve data center utilization and make best usage of the available resources, Hipster can dynamically assign remaining cores to batch workloads without violating the QoS constraints for the latency-critical workloads. We perform experiments using a 64-bit ARM big.LITTLE platform, and show that, compared to prior work, Hipster improves the QoS guarantee for Web-Search from 80% to 96%, and for Memcached from 92% to 99%, while reducing the energy consumption by up to 18%. Rajiv Nishtala, Paul M. Carpenter, Vinicius Petrucci, Xavier Martorell |
HPCA | 3 |
| 2017 | The Hipster Approach for Improving Cloud System EfficiencyabstractIn 2013, U.S. data centers accounted for 2.2% of the country’s total electricity consumption, a figure that is projected to increase rapidly over the next decade. Many important data center workloads in cloud computing are interactive, and they demand strict levels of quality-of-service (QoS) to meet user expectations, making it challenging to optimize power consumption along with increasing performance demands. This article introduces Hipster, a technique that combines heuristics and reinforcement learning to improve resource efficiency in cloud systems. Hipster explores heterogeneous multi-cores and dynamic voltage and frequency scaling for reducing energy consumption while managing the QoS of the latency-critical workloads. To improve data center utilization and make best usage of the available resources, Hipster can dynamically assign remaining cores to batch workloads without violating the QoS constraints for the latency-critical workloads. We perform experiments using a 64-bit ARM big.LITTLE platform and show that, compared to prior work, Hipster improves the QoS guarantee for Web-Search from 80% to 96%, and for Memcached from 92% to 99%, while reducing the energy consumption by up to 18%. Hipster is also effective in learning and adapting automatically to specific requirements of new incoming workloads just enough to meet the QoS and optimize resource consumption. Rajiv Nishtala, Paul M. Carpenter, Vinicius Petrucci, Xavier Martorell |
ACM Trans. Comput. Syst. | 3 |
| 2016 | REPP-H: Runtime Estimation of Power and Performance on Heterogeneous Data CentersabstractOne of the main challenges in data center systems is operating under certain Quality of Service (QoS) while minimizing power consumption. Increasingly, data centers are adopting heterogeneous server architectures with different power-performance trade-offs. This requires careful understanding of the application behavior across multiple architectures at runtime so as to enable meeting specified power and performance requirements. In this work, we present and evaluate REPP-H (Runtime Estimation of Performance and Power on Heterogeneous data centers). REPP-H leverages hardware performance counters available on all major server architectures to ensure a highly responsive power capping mechanism and delivering a minimum performance in a single step. We experimentally show that REPP-H can successfully estimate power and performance of several single-threaded and multiprogrammed workloads. The average errors on ARM, AMD and Intel architectures are, respectively, 7.1%, 9.0%, 7.1% when predicting performance, and 6.0%, 6.5%, 8.1% when predicting power on those heterogeneous servers. Rajiv Nishtala, Xavier Martorell, Vinicius Petrucci, Daniel Mossé |
SBAC-PAD | 3 |
| 2016 | Designing Future Warehouse-Scale Computers for Sirius, an End-to-End Voice and Vision Personal AssistantabstractAs user demand scales for intelligent personal assistants (IPAs) such as Apple’s Siri, Google’s Google Now, and Microsoft’s Cortana, we are approaching the computational limits of current datacenter (DC) architectures. It is an open question how future server architectures should evolve to enable this emerging class of applications, and the lack of an open-source IPA workload is an obstacle in addressing this question. In this article, we present the design of Sirius, an open end-to-end IPA Web-service application that accepts queries in the form of voice and images, and responds with natural language. We then use this workload to investigate the implications of four points in the design space of future accelerator-based server architectures spanning traditional CPUs, GPUs, manycore throughput co-processors, and FPGAs. To investigate future server designs for Sirius, we decompose Sirius into a suite of eight benchmarks (Sirius Suite) comprising the computationally intensive bottlenecks of Sirius. We port Sirius Suite to a spectrum of accelerator platforms and use the performance and power trade-offs across these platforms to perform a total cost of ownership (TCO) analysis of various server design points. In our study, we find that accelerators are critical for the future scalability of IPA services. Our results show that GPU- and FPGA-accelerated servers improve the query latency on average by 8.5× and 15×, respectively. For a given throughput, GPU- and FPGA-accelerated servers can reduce the TCO of DCs by 2.3× and 1.3×, respectively. Johann Hauswald, Michael Laurenzano, Hailong Yang 0002, Yiping Kang, Austin Rovinski, Arjun Khurana, Ronald G. Dreslinski, Trevor N. Mudge, Vinicius Petrucci, Lingjia Tang, Jason Mars |
ACM Trans. Comput. Syst. | 11 |
| 2015 | Sirius: An Open End-to-End Voice and Vision Personal Assistant and Its Implications for Future Warehouse Scale ComputersabstractAs user demand scales for intelligent personal assistants (IPAs) such as Apple's Siri, Google's Google Now, and Microsoft's Cortana, we are approaching the computational limits of current datacenter architectures. It is an open question how future server architectures should evolve to enable this emerging class of applications, and the lack of an open-source IPA workload is an obstacle in addressing this question. In this paper, we present the design of Sirius, an open end-to-end IPA web-service application that accepts queries in the form of voice and images, and responds with natural language. We then use this workload to investigate the implications of four points in the design space of future accelerator-based server architectures spanning traditional CPUs, GPUs, manycore throughput co-processors, and FPGAs. Johann Hauswald, Michael Laurenzano, Austin Rovinski, Arjun Khurana, Ronald G. Dreslinski, Trevor N. Mudge, Vinicius Petrucci, Lingjia Tang, Jason Mars |
ASPLOS | 9 |
| 2015 | Octopus-Man: QoS-driven task management for heterogeneous multicores in warehouse-scale computersabstractHeterogeneous multicore architectures have the potential to improve energy efficiency by integrating power-efficient wimpy cores with high-performing brawny cores. However, it is an open question as how to deliver energy reduction while ensuring the quality of service (QoS) of latency-sensitive web-services running on such heterogeneous multicores in warehouse-scale computers (WSCs). In this work, we first investigate the implications of heterogeneous multicores in WSCs and show that directly adopting heterogeneous multicores without re-designing the software stack to provide QoS management leads to significant QoS violations. We then present Octopus-Man, a novel QoS-aware task management solution that dynamically maps latency-sensitive tasks to the least power-hungry processing resources that are sufficient to meet the QoS requirements. Using carefully-designed feedback-control mechanisms, Octopus-Man addresses critical challenges that emerge due to uncertainties in workload fluctuations and adaptation dynamics in a real system. Our evaluation using web-search and memcached running on a real-system Intel heterogeneous prototype demonstrates that Octopus-Man improves energy efficiency by up to 41% (CPU power) and up to 15% (system power) over an all-brawny WSC design while adhering to specified QoS targets. Vinicius Petrucci, Michael Laurenzano, John Doherty, Daniel Mossé, Jason Mars, Lingjia Tang |
HPCA | 1 |
| 2015 | Energy-Efficient Thread Assignment Optimization for Heterogeneous Multicore SystemsabstractThe current trend to move from homogeneous to heterogeneous multicore systems provides compelling opportunities for achieving performance and energy efficiency goals. Running multiple threads in multicore systems poses challenges on meeting limited shared resources, such as memory bandwidth. We propose an optimization approach that includes an Integer Linear Programming (ILP) optimization model and a scheme to dynamically determine thread-to-core assignment. We present simulation analysis that shows energy savings and performance gains for a variety of workloads compared to state-of-the-art schemes. We implemented and evaluated a prototype of our thread assignment approach at user level, leveraging Linux scheduling and performance-monitoring capabilities. Vinicius Petrucci, Orlando Loques, Daniel Mossé, Rami G. Melhem, Neven Abou Gazala, Sameh Gobriel |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2013 | Energy-aware thread co-location in heterogeneous multicore processorsabstractGiven the wide variety of performance demands for various workloads, the trend in embedded systems is shifting from homogeneous to heterogeneous processors, which have been shown to yield performance and energy saving benefits. A typical heterogeneous processor has cores with different performance and power characteristics, that is, high performance and power hungry (“big”) cores, and low power and performance (“small”) cores. In order to satisfy the memory bandwidth and computation demands of various threads, it is important (albeit challenging) to map threads to cores. Such assignment should take into account that threads could potentially be harmful to each other in the usage of shared resources (e.g., cache, memory). We propose a scheme for dynamic energy-efficient assignment of threads to big/small cores, DIO-E (Distributed Intensity Online-Energy), which is an enhancement of the previously proposed DIO. In contrast to DIO, we take into account both CPU and memory demands of threads to characterize the performance of threads when co-running on the same core at run-time. Our results show that DIO-E improves the energy-delay-squared product (ED2) by 9% (average) over DIO, running on a performance-asymmetric multicore system. Both DIO and DIO-E show about 50% improvement in ED2over a state-of-the-art solution. Rajiv Nishtala, Daniel Mossé, Vinicius Petrucci |
EMSOFT | 3 |
| 2013 | An iterated local search heuristic for multi-capacity bin packing and machine reassignment problems
Renaud Masson, Thibaut Vidal, Julien Michallet, Puca Huachi Vaz Penna, Vinicius Petrucci, Anand Subramanian 0001, Hugues Dubedout |
Expert Syst. Appl. | 5 |
| 2012 | Thread Assignment Optimization with Real-Time Performance and Memory Bandwidth Guarantees for Energy-Efficient Heterogeneous Multi-core SystemsabstractThe current trend to move from homogeneous to heterogeneous multi-core systems promises further performance and energy-efficiency benefits. A typical future heterogeneous multi-core system includes two distinct types of cores, such as high performance sophisticated ("large'') cores and simple low-power ("small'') cores. In those heterogeneous platforms, execution phases of application threads that are CPU-intensive can take best advantage of large cores, whereas I/O or memory intensive execution phases are best suited and assigned to small cores. However, it is crucial that the assignment of threads to cores satisfy both the computational and memory bandwidth constraints of the threads. We propose an optimization approach to determine and apply the most energy efficient assignment of threads with soft real-time performance and memory bandwidth constraints in a multi-core system. Our approach includes an ILP (Integer Linear Programming) optimization model and a scheme to dynamically change thread-to-core assignment, since thread execution phases may change over time. In comparison to state-of-art dynamic thread assignment schemes, we show energy savings and performance gains for a variety of workloads, while respecting thread performance and memory bandwidth requirements. Vinicius Petrucci, Orlando Loques, Daniel Mossé, Rami G. Melhem, Neven Abou Gazala, Sameh Gobriel |
IEEE Real-Time and Embedded Technology and Applications Symposium | 1 |
| 2011 | Optimized Management of Power and Performance for Virtualized Heterogeneous Server ClustersabstractThis paper proposes and evaluates an approach for power and performance management in virtualized server clusters. The major goal of our approach is to reduce power consumption in the cluster while meeting performance requirements. The contributions of this paper are: (1) a simple but effective way of modeling power consumption and capacity of servers even under heterogeneous and changing workloads, and (2) an optimization strategy based on a mixed integer programming model for achieving improvements on power-efficiency while providing performance guarantees in the virtualized cluster. In the optimization model, we address application workload balancing and the often ignored switching costs due to frequent and undesirable turning servers on/off and VM relocations. We show the effectiveness of the approach applied to a server cluster test bed. Our experiments show that our approach conserves about 50% of the energy required by a system designed for peak workload scenario, with little impact on the applications' performance goals. Also, by using prediction in our optimization strategy, further QoS improvement was achieved. Vinicius Petrucci, Enrique V. Carrera, Orlando Loques, Julius C. B. Leite, Daniel Mossé |
CCGRID | 1 |