Christian Pinto

dblp:79/8360 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0001-7060-2742ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Page Migration for Hardware Memory Disaggregation Across a Network
abstract
Hardware memory disaggregation (HMD) is an emerging technology that enables access to remote memory, thereby creating expansive memory pools and reducing memory underutilization in datacenters.However, a significant challenge arises when accessing remote memory over a network: increased contention that can lead to severe application performance degradation.To reduce the performance penalty of using remote memory, the operating system uses page migration to promote frequently accessed pages closer to the processor.However, previously proposed page migration mechanisms do not achieve the best performance in HMD systems because of obliviousness to variable page transfer costs that occur due to network contention.To address these limitations, we present INDIGO: a network-aware page migration framework that uses novel page telemetry and a learning-based approach for network adaptation.We implemented INDIGO in the Linux kernel and evaluated it with common cloud and HPC applications on a real disaggregated memory system prototype.Our evaluation shows that INDIGO offers up to 50-70% improvement in application performance compared to other state-of-the-art page migration policies and reduces network traffic up to 2×.
Archit Patke, Christian Pinto, Saurabh Jha, Haoran Qiu, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
ICS2
2024 Queue Management for SLO-Oriented Large Language Model Serving
abstract
Large language model (LLM) serving is becoming an increasingly critical workload for cloud providers. Existing LLM serving systems focus on interactive requests, such as chatbots and coding assistants, with tight latency SLO requirements. However, when such systems execute batch requests that have relaxed SLOs along with interactive requests, it leads to poor multiplexing and inefficient resource utilization. To address these challenges, we propose QLM, a queue management system for LLM serving. QLM maintains batch and interactive requests across different models and SLOs in a request queue. Optimal ordering of the request queue is critical to maintain SLOs while ensuring high resource utilization. To generate this optimal ordering, QLM uses a Request Waiting Time (RWT) Estimator that estimates the waiting times for requests in the request queue. These estimates are used by a global scheduler to orchestrate LLM Serving Operations (LSOs) such as request pulling, request eviction, load balancing, and model swapping. Evaluation on heterogeneous GPU devices and models with real-world LLM serving dataset shows that QLM improves SLO attainment by 40-90% and throughput by 20-400% while maintaining or improving device utilization compared to other state-of-the-art LLM serving systems. QLM's evaluation is based on the production requirements of a cloud provider. QLM is publicly available at https://www.github.com/QLM-project/QLM.
Archit Patke, Dhemath Reddy, Saurabh Jha, Haoran Qiu, Christian Pinto, Chandrasekhar Narayanaswami 0001, Zbigniew T. Kalbarczyk, Ravishankar K. Iyer
SoCC5
2024 Disaggregated RDDs: Extending and Analyzing Apache Spark for Memory Disaggregated Infrastructures
abstract
Apache Spark has become essential in large-scale data processing as the demand for scalable data analytics grows. With memory costs constituting a significant portion of server expenses, the under-utilization and fragmentation of resources pose a substantial challenge for data center operators reliant on economies of scale. Memory disaggregation emerges as a solution to this challenge, by leveraging remote memory pools to reduce resource fragmentation and under-utilization. Yet, these advantages are not without cost. Disaggregated memory systems introduce increased latency and reduced bandwidth, significantly impacting job execution latency. This necessitates careful optimization and management strategies to effectively balance the trade-offs between accessibility and performance. This paper introduces cache-remote, a custom Apache Spark configuration balancing memory disaggregation benefits with execution efficiency. Cache-remote uses remote memory for RDD caching (Disaggregated RDDs) and local memory for latency-sensitive computations. Our work includes a comprehensive evaluation of different memory allocation policies and Spark configurations on a hardware setup that supports memory disaggregation. We expand upon prior work by exploring a range of solutions that cater to varying tolerances for job completion latency, introducing new points to the latency-memory usage Pareto. Notably, our cache-remote approach enhances the efficiency of current disaggregated memory allocation strategies. It achieves a substantial reduction in local memory utilization—up to ${2 4. 8 \%}$—while incurring minimal execution time overhead of merely $7 \%$, compared to local-only policies.
Achilleas Tzenetopoulos, Michele Gazzetti, Dimosthenis Masouros, Christian Pinto, Sotirios Xydis, Dimitrios Soudris
IC2E4
2023 Adrias: Interference-Aware Memory Orchestration for Disaggregated Cloud Infrastructures
abstract
Workload co-location has become the de-facto approach for hosting applications in Cloud environments, leading, however, to interference and fragmentation in shared resources of the system. To this end, hardware disaggregation is introduced as a novel paradigm, that allows fine-grained tailoring of cloud resources to the characteristics of the deployed applications. Towards the realization of hardware disaggregated clouds, novel orchestration frameworks must provide additional knobs to manage the increased scheduling complexity.We present Adrias, a memory orchestration framework for disaggregated cloud systems. Adrias exploits information from low-level performance events and applies deep learning techniques to effectively predict the system state and performance of arriving workloads on memory disaggregated systems, thus, driving cognitive scheduling between local and remote memory allocation modes. We evaluate Adrias on a state-of-art disaggregated testbed and show that it achieves 0.99 and 0.942 R2score for system state and application’s performance prediction on average respectively. Moreover, Adrias manages to effectively utilize disaggregated memory, by offloading almost 1/3 of deployed applications with less than 15% performance overhead compared to a conventional local memory scheduling, while clearly outperforms naive scheduling approaches (random and round-robin), by providing up to ×2 better performance.
Dimosthenis Masouros, Christian Pinto, Michele Gazzetti, Sotirios Xydis, Dimitrios Soudris
HPCA2
2022 EVOLVE: Towards Converging Big-Data, High-Performance and Cloud-Computing Worlds
abstract
EVOLVE is a pan European Innovation Action that aims to fully-integrate High-Performance-Computing (HPC) hardware with state-of-the-art software technologies under a unique testbed, that enables the convergence of HPC, Cloud and Big-Data worlds and increases our ability to extract value from massive and demanding datasets. EVOLVE's advanced compute platform combines HPC-enabled capabilities, with transparent deployment in high abstraction level, and a versatile Big-Data processing stack for end-to-end workflows. Hence, domain experts have the potential to improve substantially the efficiency of existing services or introduce new models in the respective domains, e.g., automotive services, bus transportation, maritime surveillance and others. In this paper, we describe EVOLVE's testbed, and evaluate the performance of the integrated pilots from different domains.
Achilleas Tzenetopoulos, Dimosthenis Masouros, Konstantina Koliogeorgi, Sotirios Xydis, Dimitrios Soudris, Antony Chazapis, Christos Kozanitis, Angelos Bilas, Christian Pinto, Huy-Nam Nguyen, Stelios Louloudakis, Georgios Gardikis, George Vamvakas, Michelle Aubrun, Christi Symeonidou, Vassilis Spitadakis, Konstantinos F. Xylogiannopoulos, Bernhard Peischl, Tahir Emre Kalayci, Alexander Stocker, Jean-Thomas Acquaviva
DATE9
2021 A Holistic Approach to Data Access for Cloud-Native Analytics and Machine Learning
abstract
Cloud providers offer a variety of storage solutions for hosting data, both in price and in performance. For Analytics and machine learning applications, object storage services are the go-to solution for hosting the datasets that exceed tens of gigabytes in size. However, such a choice results in performance degradation for these applications and requires extra engineering effort in the form of code changes to access the data on remote storage. In this paper, we present a generic end-to-end solution that offers seamless data access for remote object storage services, transparent data caching within the compute infrastructure, and data-aware topologies that boost the performance of applications deployed in Kubernetes. We introduce a custom-implemented cache mechanism that supports all the requirements of the former and we demonstrate that our holistic solution leads up to 48% improvement for Spark implementation of the TPC-DS benchmark and up to 191% improvement for the training of deep learning models from the MLPerf benchmark suite.
Panos K. Koutsovasilis, Srikumar Venugopal, Yiannis Gkoufas, Christian Pinto
CLOUD4
2021 A Holistic System Software Integration of Disaggregated Memory for Next-Generation Cloud Infrastructures
abstract
Modern cloud computing workloads are becoming day by day more demanding, in terms of computational resources, as they feature multiple complex components, utilize heterogeneous hardware, and require tremendous amounts of memory. Such attributes of the emerging workloads disrupt the traditional design of cloud infrastructures, which are bound to decisions at the design time of the infrastructure, and mandate the dynamic composability of the next-generation cloud infrastructures.Scaling beyond the physical boundaries of the server trays, and minimizing over-provisioning of the nodes by disaggregating computational resources, such as CPUs, memory, etc., is a timely research problem. However, the prior line of research focuses on either NVM technologies or other disk-related approaches that try to relax the resource pressure induced by over-provisioning.In this work, we design, implement, and evaluate all necessary changes in the system software stack to support the dynamic allocation of disaggregated memory, depending on the needs of the workloads at hand. As a result, we can transparently increase the available memory of a system and achieve on average 57% better performance, compared to Swap, the de-facto mechanism used today for over-provisioning. Furthermore, we present a memory balancing policy that autonomously migrates memory pages across local and disaggregated memory. Our experimental results show that, even with a limited amount of local memory, memory balancing improves the performance by 41% on average, compared with using only disaggregated memory.
Panos K. Koutsovasilis, Michele Gazzetti, Christian Pinto
CCGRID3
2021 EVOLVE: HPC and cloud enhanced testbed for extracting value from large-scale diverse data
abstract
EVOLVE is a pan-European Innovation Action building a converged infrastructure to bring together the HPC, Cloud, and Big Data worlds. EVOLVE's platform and software stack supports large-scale, data-intensive applications, driven primarily by industry requirements set by pilot and proof-of-concept use cases from diverse fields. Given the unprecedented data growth we are experiencing, EVOLVE's infrastructure is key in enabling the cost-effective processing of massive amounts of data and the adaptation of multiple high-end technologies, in an environment that fosters interoperability and enforces increased security.
Antony Chazapis, Jean-Thomas Acquaviva, Angelos Bilas, Georgios Gardikis, Christos Kozanitis, Stelios Louloudakis, Huy-Nam Nguyen, Christian Pinto, Arno Scharl, Dimitrios Soudris
CF8
2020 ThymesisFlow: A Software-Defined, HW/SW co-Designed Interconnect Stack for Rack-Scale Memory Disaggregation
abstract
With cloud providers constantly seeking the best infrastructure trade-off between performance delivered to customers and overall energy/utilization efficiency of their data-centres, hardware disaggregation comes in as a new paradigm for dynamically adapting the data-centre infrastructure to the characteristics of the running workloads. Such an adaptation enables an unprecedented level of efficiency both from the standpoint of energy and the utilization of system resources. In this paper, we present - ThymesisFlow - the first, to our knowledge, full-stack prototype of the holy-grail of disaggregation of compute resources: pooling of remote system memory. Thymesis-Flow implements a HW/SW co-designed memory disaggregation interconnect on top of the POWER9 architecture, by directly interfacing the memory bus via the OpenCAPI port. We use ThymesisFlow to evaluate how disaggregated memory impacts a set of cloud workloads, and we show that for many of them the performance degradation is negligible. For those cases that are severely impacted, we offer insights on the underlying causes and viable cross-stack mitigation paths.
Christian Pinto, Dimitris Syrivelis, Michele Gazzetti, Panos K. Koutsovasilis, Andrea Reale, Kostas Katrinis, H. Peter Hofstee
MICRO1
2019 Exploiting CPU Voltage Margins to Increase the Profit of Cloud Infrastructure Providers
abstract
Energy efficiency is a major concern for cloud computing, with CPUs accounting a significant fraction of datacenter nodes power consumption. CPU manufacturers introduce voltage margins to guarantee correct operation. However, these are unnecessarily wide for real-world execution scenarios, and translate to increased power consumption. In this paper, we investigate how such margins can be exploited by infrastructure operators, by selectively undervolting nodes, at the controlled risk of inducing failures and activating service-level agreement (SLA) violation penalties. We model the problem in a formal way, capturing the most important aspects that drive VM management and system configuration decisions. Then, we introduce XM-VFS policy that reduces infrastructure operator costs by reducing voltage margins, and compare it with the state-of-the-art which employs dynamic voltage-frequency scaling (DVFS) and workload consolidation. We perform simulations to quantify the cost reduction, considering the energy consumption and potential SLA violations. Our results show significant gains, up to 17.35% and 16.32% for the energy and cost reduction respectively. In our simulations, we use realistic assumptions for voltage margins, energy consumption and performance degradation of applications due to frequency scaling, based on the characterization of commercial Intel-and ARM-based machines. Our model and scheduling policy are generic and scalable.
Christos Kalogirou, Panos K. Koutsovasilis, Christos D. Antonopoulos, Nikolaos Bellas, Spyros Lalis, Srikumar Venugopal, Christian Pinto
CCGRID7
2017 Paving the Way Towards a Highly Energy-Efficient and Highly Integrated Compute Node for the Exascale Revolution: The ExaNoDe Approach
abstract
Power consumption and high compute density are the key factors to be considered when building a compute node for the upcoming Exascale revolution. Current architectural design and manufacturing technologies are not able to provide the requested level of density and power efficiency to realise an operational Exascale machine. A disruptive change in the hardware design and integration process is needed in order to cope with the requirements of this forthcoming computing target. This paper presents the ExaNoDe H2020 research project aiming to design a highly energy efficient and highly integrated heterogeneous compute node targeting Exascale level computing, mixing low-power processors, heterogeneous co-processors and using advanced hardware integration technologies with the novel UNIMEM Global Address Space memory system.
Alvise Rigo, Christian Pinto, Kevin Pouget, Daniel Raho, Denis Dutoit, Pierre-Yves Martinez, Chris Doran, Luca Benini, Iakovos Mavroidis, Manolis Marazakis, Valeria Bartsch, Guy Lonsdale, Antoniu Pop, John Goodacre, Annaik Colliot, Paul M. Carpenter, Petar Radojkovic, Dirk Pleiter, Dominique Drouin, Benoît Dupont de Dinechin
DSD2
2016 Rack-scale disaggregated cloud data centers: The dReDBox project vision
Kostas Katrinis, Dimitris Syrivelis, Dionisios N. Pnevmatikatos, Georgios Zervas, Dimitris Theodoropoulos 0001, Iordanis Koutsopoulos, K. Hasharoni, Daniel Raho, Christian Pinto, Felix Espina, Sergio López-Buedo, Qianqiao Chen, Mario Nemirovsky, Damian Roca, H. Klos, T. Berends
DATE9
2015 Reducing energy consumption in microcontroller-based platforms with low design margin co-processors
Andres Gomez 0001, Christian Pinto, Andrea Bartolini, Davide Rossi 0001, Luca Benini, Hamed Fatemi, José Pineda de Gyvez
DATE2
2015 GPU Acceleration for Simulating Massively Parallel Many-Core Platforms
abstract
Abstract—Emerging massively parallel architectures such as a general-purpose processor plus many-core programmable accelerators are creating an increasing demand for novel methods to perform their architectural simulation. Most state-of-the-art simulation technologies are exceedingly slow and the need to model full system many-core architectures adds further to the complexity issues. This paper presents a novel methodology to accelerate the simulation of many-core coprocessors using GPU platforms. We demonstrate the challenges, feasibility and benefits of our idea to use heterogeneous system (CPU and GPU) to simulate future architecture of many-core heterogeneous platforms. The target architecture selected to evaluate our methodology consists of an ARM general purpose CPU coupled with many-core coprocessor with thousands of simple in-order cores connected in a tile network. This work presents optimization techniques used to parallelize the simulation specifically for acceleration on GPUs. We partition the full system simulation between CPU and GPU, where the target general purpose CPU is simulated on the host CPU, whereas the many-core coprocessor is simulated on the NVIDIA Tesla 2070 GPU platform. Our experiments show performance of up to 50 MIPS when simulating the entire heterogeneous chip, and high scalability with increasing cores on coprocessor.
Shivani Raghav, Martino Ruggiero, Andrea Marongiu, Christian Pinto, David Atienza 0001, Luca Benini
IEEE Trans. Parallel Distributed Syst.4
2014 Exploring DMA-assisted prefetching strategies for software caches on multicore clusters
abstract
Modern many-core programmable accelerators are often composed by several computing units grouped in clusters, with a shared per-cluster scratchpad data memory. The main programming challenge imposed by these architectures is to hide the external memory to on-chip scratchpad memory transfer latency, trying to overlap as much as possible memory transfers with actual computation. This problem is usually tackled using complex DMA-based programming patterns (e.g. double buffering), which require a heavy refactoring of applications. Software caches are an alternative to hand-optimized DMA programming. However, even if a software cache can reduce the programming effort, it is still relying on synchronous memory transfers. In fact in case of a cache miss, the new line is copied in cache and the requesting processor has to wait for the completion of the transfer. While waiting, processors are not able to perform any other computation. Cache lines prefetching can be used to reduce the number of synchronous memory transfers, and increase the active time of each processor, by loading cache lines before they are actually needed. In this work we explore various DMA-based prefetching techniques applied to a software cache implementation, presenting both automatic and programmer assisted prefetch mechanisms applied to computer vision kernels.
Christian Pinto, Luca Benini
ASAP1
2013 A highly efficient, thread-safe software cache implementation for tightly-coupled multicore clusters
abstract
A widely adopted design paradigm for many-core accelerators features processing elements grouped in clusters. Due to area, power and design simplicity, processors in the same clusters are often not equipped with data-caches but rather share a tightly coupled data memory (TCDM). Even if the use of a TCDM is more energy and area efficient than a cache it requires a higher programming effort because memory needs to be explicitly managed with DMA-based L3 to TCDM copies. In this context Software Caches can be used to automatically transfer data between the local TCDM and the external memory, simplifying the task of the programmer. In this paper we present an implementation of Software Cache for the STMicroelectronics STHORM many-core accelerator, featuring a L1 TCDM shared by 16 processors in a cluster. Our main contribution is the design of a fast and thread-safe cache allowing parallel access from different processing elements inside the same cluster. We evaluate our implementation with micro-benchmarks as well as a real world application from the computer vision domain. Results show that a software cache provides major performance improvements with respect to L3 allocation of large data structures even when it is aggressively shared among many parallel threads.
Christian Pinto, Luca Benini
ASAP1
2013 SIMinG-1k: A thousand-core simulator running on general-purpose graphical processing units
abstract
SUMMARY This paper introduces SIMinG‐1k—a manycore simulator infrastructure. SIMinG‐1k is a graphics processing unit accelerated, parallel simulator for design‐space exploration of large‐scale manycore systems. It features an optimal trade‐off between modeling accuracy and simulation speed. Its main objectives are high performance, flexibility, and ability to simulate thousands of cores. SIMinG‐1k can model different architectures (currently, we support ARM (Available from: http://infocenter.arm.com/help/index.jsp?topic=/com.arm.doc.ddi0100i/index.html ) and Intel x86) using two‐step approac where architecture specific front end is decoupled from a fast and parallel manycore virtual machine running on graphical processing unit platform. We evaluate the simulator for target architecture with up to 4096 cores. Our results demonstrate very high scalability and almost linear speedup with simulation of increasing number of cores.Copyright © 2012 John Wiley & Sons, Ltd.
Shivani Raghav, Andrea Marongiu, Christian Pinto, Martino Ruggiero, David Atienza 0001, Luca Benini
Concurr. Comput. Pract. Exp.3
2011 GPGPU-Accelerated Parallel and Fast Simulation of Thousand-Core Platforms
abstract
The multicore revolution and the ever-increasing complexity of computing systems is dramatically changing system design, analysis and programming of computing platforms. Future architectures will feature hundreds to thousands of simple processors and on-chip memories connected through a network-on-chip. Architectural simulators will remain primary tools for design space exploration, software development and performance evaluation of these massively parallel architectures. However, architectural simulation performance is a serious concern, as virtual platforms and simulation technology are not able to tackle the complexity of thousands of core future scenarios. The main contribution of this paper is the development of a new simulation approach and technology for many core processors which exploit the enormous parallel processing capability of low-cost and widely available General Purpose Graphic Processing Units (GPGPU). The simulation of many-core architectures exhibits indeed a high level of parallelism and is inherently parallelizable, but GPGPU acceleration of architectural simulation requires an in-depth revision of the data structures and functional partitioning traditionally used in parallel simulation. We demonstrate our GPGPU simulator on a target architecture composed by several cores (i.e. ARM ISA based), with instruction and data caches, connected through a Network-on-Chip (NoC). Our experiments confirm the feasibility of our approach.
Christian Pinto, Shivani Raghav, Andrea Marongiu, Martino Ruggiero, David Atienza 0001, Luca Benini
CCGRID1