EDBT 2026 Demo / reviewers in the wild / expert
Ana Lucia Varbanescu
dblp:62/3310
· DBLP profile ↗
60ranked-venue papers
6as first author
20since 2021 · last 2026
0000-0002-4932-1900ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 39 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 15 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How Much Energy Is Wasted in LLM operations? Evidence from Kernel-Level DVFSabstractThe rapid growth of AI has fueled the expansion of accelerator- or GPU-based data centers. However, the rising operational energy consumption has emerged as a critical bottleneck and a major sustainability concern. Dynamic Voltage and Frequency Scaling (DVFS) is a well-known technique used to reduce energy consumption, and thus improve energy-efficiency, since it requires little effort and works with existing hardware. Reducing the energy consumption of training and inference of Large Language Models (LLMs) through DVFS or power capping is feasible: related work has shown energy savings can be significant, but at the cost of significant slowdowns. Jeffrey Spaan, Kuan-Hsun Chen, Ana Lucia Varbanescu |
CF | 3 |
| 2026 | POSTER: Enabling Fine-Grain DVFS for Multi-Kernel GPU WorkloadsabstractApplication-level DVFS for GPUs has succeeded in improving energy efficiency, using standard tools like nvidia-smi or rocm-smi. However, GPU workloads with many kernels, like AI training or inference, pose additional challenges: fine-grain DVFS could deliver additional energy savings, but requires specific triggers to change frequency settings at the right time. We propose an automated approach for such fine-grained DVFS for GPU workloads. We discuss possible trigger mechanisms and policies, we further analyze the requirements for the triggers and the overhead they may incur, and estimate feasible savings. Finally, we assess the approach for an NVIDIA Blackwell GPU, highlighting challenges towards a full prototype. Jeffrey Spaan, Kuan-Hsun Chen, Ana Lucia Varbanescu |
CF | 3 |
| 2026 | Assessing the Performance Impact of Data Layouts: a Benchmarking Approach
Jolly Chen, Ana Lucia Varbanescu, Axel Naumann |
Euro-Par (1) | 2 |
| 2025 | Empowering Sustainability: Energy Labeling of Digital Services Using SimulationabstractThe energy consumption of digital services has become a concern for stakeholders committed to sustainability. Raising awareness of this consumption is essential to improve the energy efficiency of digital services. However, expressing the energy usage of digital services in an easily understandable and actionable way remains a challenge. We address this challenge by proposing a first operational energy labeling method for digital services in the computing continuum. Our approach enables stakeholders, including cloud and network providers, application developers, researchers, and end-users of digital services, to better understand and improve the energy efficiency of their applications. Focusing on video surveillance digital services, and using the enhanced iFogSim framework, we propose an energy labeling scheme, and demonstrate its merits with extensive scenario analysis and simulation. We further discuss how our approach can help reduce energy consumption and/or improve performance, all without modifying the application's functional parameters or system architecture. Saeedeh Baneshi, Anuj Pathania, Benny Akesson, Andy D. Pimentel, Ana Lucia Varbanescu |
CCGrid | 5 |
| 2025 | ROOT's RNTuple and the Case for Custom Scientific Data FormatsabstractEach month, CERN collects and stores multiple petabytes of measurement data from the Large Hadron Collider (LHC) for further analysis. These data are stored in a custom columnar format provided by CERN’s open-source software framework ROOT. However, the increased availability of general-purpose columnar data formats raises the question whether developing and maintaining a custom data format for high-energy physics (HEP) is still worth it. To address this question, we compare ROOT’s RNTuple data format against Apache ORC and Apache Parquet. We find that RNTuple stores data 6-21% more efficiently, and processes physics data 3-6x faster. Florine Willemijn de Geus, Stijn Jongbloed, Vincenzo Eduardo Padulano, Ana Lucia Varbanescu |
eScience | 4 |
| 2025 | EVENTSETPROCESSOR: An Engine for Efficiently Combining High-Energy Physics DataabstractCERN’s Large Hadron Collider (LHC), the world’s largest high-energy physics (HEP) instrument, collects tens of petabytes of data per year. The LHC’s next phase is expected to produce up to ten times more data, which calls for novel, more efficient ways of storing and processing these data.HEP collider data are prepared and provided to physicists as read-only data sets, stored in a custom columnar data format. While traditionally all data needed for a particular analysis were captured in a single data set, the increasing scale of the LHC and the advent of modern analysis techniques now requires analysis workflows to use data from different data sets. However, the processing model established across the HEP community does not yet provide a straightforward way to achieve this and currently relies heavily on data duplication to produce the desired data sets. This leads to significant overhead in analysis workflows, both in runtime and storage.To reduce this overhead, we propose more efficient ways to combine HEP data sets. Specifically, we design union and join operations, as defined in relational algebra, to combine HEP data sets at runtime, eliminating therefore the need for data duplication. In this paper, we specify these operations for HEP data and introduce EVENTSETPROCESSOR – an engine that implements these operations for HEP data processing. Through a first prototype, we show that this engine integrates well in existing HEP workflows, and that it can perform up to twice as fast as the current approach. Florine Willemijn de Geus, Vincenzo Eduardo Padulano, Jakob Blomer, Hannes Mühleisen, Ana Lucia Varbanescu |
eScience | 5 |
| 2025 | PEIR: Modeling Performance in Neural Information Retrieval
Pooya Khandel, Andrew Yates, Ana Lucia Varbanescu, Maarten de Rijke, Andy D. Pimentel |
ECIR (2) | 3 |
| 2025 | Wedge-Parallel Triangle Counting for GPUs
Jeffrey Spaan, Kuan-Hsun Chen, David A. Bader, Ana Lucia Varbanescu |
Euro-Par (3) | 4 |
| 2025 | Component-Based Analytical Modeling of GPU Runtime Performance: a Case-Study in Scientific ComputingabstractAnalytical performance models are excellent tools for fast performance prediction and can be used effectively for designing and tuning parallel algorithms. However, such models are non-trivial to build, especially when both the application and the system are very complex. Jolly Chen, Ana Lucia Varbanescu, Monica Dessole |
ICPE | 2 |
| 2024 | Using Evolutionary Algorithms to Find Cache-Friendly Generalized Morton Layouts for ArraysabstractThe layout of multi-dimensional data can have a significant impact on the efficacy of hardware caches and, by extension, the performance of applications. Common multi-dimensional layouts include the canonical row-major and column-major layouts as well as the Morton curve layout. In this paper, we describe how the Morton layout can be generalized to a very large family of multi-dimensional data layouts with widely varying performance characteristics. We posit that this design space can be efficiently explored using a combinatorial evolutionary methodology based on genetic algorithms. To this end, we propose a chromosomal representation for such layouts as well as a methodology for estimating the fitness of array layouts using cache simulation. We show that our fitness function correlates to kernel running time in real hardware, and that our evolutionary strategy allows us to find candidates with favorable simulated cache properties in four out of the eight real-world applications under consideration in a small number of generations. Finally, we demonstrate that the array layouts found using our evolutionary method perform well not only in simulated environments but that they can effect significant performance gains---up to a factor ten in extreme cases---in real hardware. Stephen Nicholas Swatman, Ana Lucia Varbanescu, Andy D. Pimentel, Andreas Salzburger, Attila Krasznahorkay |
ICPE | 2 |
| 2023 | The Graph-Massivizer Approach Toward a European Sustainable Data Center Digital TwinabstractModeling and understanding an expensive next-generation data center operating at a sustainable exascale performance remains a challenge yet to solve. The paper presents the approach taken by the Graph-Massivizer project, funded by the European Union, towards a sustainable data center, targeting a massive graph representation and analysis of its digital twin. We introduce five interoperable open-source tools that support this undertaking, creating an automated, sustainable loop of graph creation, analytics, optimization, sustainable resource management, and operation, emphasizing state-of-the-art progress. We plan to employ the tools for designing a massive data center graph, representing a digital twin describing spatial, semantic, and temporal relationships between the monitoring metrics, hardware nodes, cooling equipment, and jobs. The project aims to strengthen Bologna Technopole as a leading European supercomputing and big data hub offering sustainable green computing for improved societally relevant science throughput. Martin Molan, Junaid Ahmed Khan, Andrea Bartolini, Roberta Turra, Giorgio Pedrazzi, Michael Cochez, Alexandru Iosup, Dumitru Roman, Joze M. Rozanec, Ana Lucia Varbanescu, Radu Prodan |
COMPSAC | 10 |
| 2023 | Analyzing Digital Services Across the Compute Continuum Using iFogSimabstractDigital services enable users to interact with a broad range of applications and, as such, have become an essential part of our daily lives. Although convenient, their ubiquity comes at a significant cost in energy, raising sustainability concerns. We access these services by triggering a computing continuum, spanning from the device to the edge, fog, and cloud. Scheduling decisions made at each layer impact the overall quality of service (QoS) and energy consumption of digital services. Saeedeh Baneshi, Ana Lucia Varbanescu, Anuj Pathania, Benny Akesson, Andy D. Pimentel |
RTCSA | 2 |
| 2023 | Systematically Exploring High-Performance Representations of Vector Fields Through Compile-Time CompositionabstractWe present a novel benchmark suite for implementations of vector fields in high-performance computing environments to aid developers in quantifying and ranking their performance. We decompose the design space of such benchmarks into access patterns and storage backends, the latter of which can be further decomposed into components with different functional and non-functional properties. Through compile-time meta-programming, we generate a large number of benchmarks with minimal effort and ensure the extensibility of our suite. Our empirical analysis, based on real-world applications in high-energy physics, demonstrates the feasibility of our approach on CPU and GPU platforms, and highlights that our suite is able to evaluate performance-critical design choices. Finally, we propose that our work towards composing vector fields from elementary components is not only useful for the purposes of benchmarking, but that it naturally gives rise to a novel library for implementing such fields in domain applications. Stephen Nicholas Swatman, Ana Lucia Varbanescu, Andy D. Pimentel, Andreas Salzburger, Attila Krasznahorkay |
ICPE | 2 |
| 2022 | Efficient trimming for strongly connected components calculationabstractStrongly Connected Components (SCCs) are useful for many applications, such as community detection and personalized recommendation. Determining the SCCs of a graph, however, can be very expensive, and parallelization is not an easy way out: the paral-lelization itself is challenging, and its performance impact varies non-trivially with the input graph structure. This variability is due to trivial components, i.e., SCCs consisting of a single vertex, which lead to significant workload imbalance. Trimming is an effective method to remove trivial components, but is inefficient when used on graphs with few trivial components. Dante Niewenhuis, Ana Lucia Varbanescu |
CF | 2 |
| 2022 | Design-Space Exploration for Decision-Support SoftwareabstractPresence monitoring or intrusion detection in a location/area are examples of decision-support applications. Decision-support applications are applications where monitoring is used to collect (heterogeneous) data and create situational awareness, which further requires decisions and/or actions. As such, decision-support software consists of different interconnected components with very diverse roles, whose communication and synchronization are essential for the application functionality and performance. Despite this complexity, software design for decision-support is often driven by short-term functional requirements and only supported by designers’ previous experience. In the current non-systematic approach, mistakes can be easily made, and can be very difficult to repair. Ate Penders, Ana Lucia Varbanescu, Gregor Pavlin, Henk J. Sips |
ASE | 2 |
| 2022 | Modelling Performance Loss due to Thread Imbalance in Stochastic Variable-Length SIMT WorkloadsabstractWhen designing algorithms for single-instruction multiple-thread (SIMT) devices such as general purpose graphics processing units (GPGPUs), thread imbalance is an important performance consideration. Thread imbalance can emerge in iterative applications where workloads are of variable length, because threads processing larger amounts of work will cause threads with less work to idle. This form of thread imbalance influences the design space of algorithms-particularly in terms of processing granularity-but we lack models to quantify its impact on application performance. In this paper, we present a statistical model for quantifying the performance loss due to thread imbalance for iterative SIMT applications with stochastic, variable-length workloads. Our model is designed to operate with minimal knowledge of the implementation details of the algorithm, relying solely on an understanding of the probability distribution of the lengths of the workloads. We validate our model against a synthetic benchmark based on a Monte Carlo simulation of matrix exponentiation, and show that our model achieves nearly perfect accuracy. Compared to empirical data extracted from real hardware, our model maintains a high degree of accuracy, predicting mean performance loss within a margin of 2%. Stephen Nicholas Swatman, Ana Lucia Varbanescu, Attila Krasznahorkay, Andy D. Pimentel |
MASCOTS | 2 |
| 2022 | The Cost of Reinforcement Learning for Game Engines: The AZ-Hive Case-studyabstractAlthough utilising computers to play board games has been a topic of research for many decades, the recent rapid developments in the field of reinforcement learning - like AlphaZero and variants - brought unprecedented progress in games such as chess and Go. However, the efficiency of this process remains unknown. In this work, we analyse the cost and efficiency of the AlphaZero approach when building a new game engine. Thus, we present our experience building AZ-Hive, an AlphaZero-based playing engine for the game of Hive. Using only the rules of the game and a quality of play assessment, AZ-Hive learns to play the game from scratch. Getting AZ-Hive up and running requires encoding the game in AlphaZero, i.e., capturing the board, the game state, the rules and the assessment of play-quality. And different encodings lead to significantly different AZ-Hive engines, with very different performance results. Thus, we propose a design space for configuring AZ-Hive, and demonstrate the costs and benefits of different configurations in this space. We find that different configurations lead to a less or more competitive playing-engine, but the training and evaluation for different such engines is prohibitively expensive. Moreover, no systematic, efficient exploration or pruning of the space is possible. In turn, an exhaustive exploration can easily take tens of training-years. Danilo de Goede, Duncan Kampert, Ana Lucia Varbanescu |
ICPE | 3 |
| 2022 | Isolating GPU Architectural Features Using Parallelism-Aware MicrobenchmarksabstractGPUs develop at a rapid pace, with new architectures emerging every 12 to 18 months. Every new GPU architecture introduces new features, expecting to improve on previous generations. However, the impact of these changes on the performance of GPGPU applications may not be directly apparent; it is often unclear to developers how exactly these features will affect the performance of their code. In this paper we propose a suite of microbenchmarks to uncover the performance of novel GPU hardware features in isolation. We target features in both the memory system and the arithmetic cores. We further ensure, by design, that our microbenchmarks capture the massively parallel nature of the GPUs, while providing fine-grained timing information at the level of individual compute units. Using this benchmarking suite, we study the differences between three of the most recent NVIDIA architectures: Pascal, Turing, and Ampere. We find that the architecture differences can have a meaningful impact on both synthetic and more realistic applications. This impact is visible both in terms of outright performance, but also affects the choice of execution parameters for realistic applications. We conclude that microbenchmarking, adapted to massive GPU parallelism, can expose differences between GPU generations, and discuss how it can be adapted for future architectures. Rico van Stigt, Stephen Nicholas Swatman, Ana Lucia Varbanescu |
ICPE | 3 |
| 2022 | ParClick: A Scalable Algorithm for EM-based Click ModelsabstractResearch on click models usually focuses on developing effective approaches to reduce biases in user clicks. However, one of the major drawbacks of existing click models is the lack of scalability. In this work, we tackle the scalability of Expectation-Maximization (EM)-based click models by introducing ParClick, a new parallel algorithm designed by following the Partitioning-Communication-Aggregation-Mapping (PCAM) method. To this end, we first provide a generic formulation of EM-based click models. Then, we design an efficient parallel version of this generic click model following the PCAM approach: we partition user click logs and model parameters into separate tasks, analyze communication among them, and aggregate these tasks to reduce communication overhead. Finally, we provide a scalable, parallel implementation of the proposed design, which maps well on a multi-core machine. Our experiments on the Yandex relevance prediction dataset show that ParClick scales well when increasing the amount of training data and computational resources. In particular, ParClick is 24.7 times faster to train with 40 million search sessions and 40 threads compared to the standard sequential version of the Click Chain Model (CCM) without any degradation in effectiveness. Pooya Khandel, Ilya Markov, Andrew Yates, Ana Lucia Varbanescu |
WWW | 4 |
| 2021 | Analytical Performance Estimation for Large-Scale Reconfigurable Dataflow PlatformsabstractNext-generation high-performance computing platforms will handle extreme data- and compute-intensive problems that are intractable with today’s technology. A promising path in achieving the next leap in high-performance computing is to embrace heterogeneity and specialised computing in the form of reconfigurable accelerators such as FPGAs, which have been shown to speed up compute-intensive tasks with reduced power consumption. However, assessing the feasibility of large-scale heterogeneous systems requires fast and accurate performance prediction. This article proposes Performance Estimation for Reconfigurable Kernels and Systems (PERKS), a novel performance estimation framework for reconfigurable dataflow platforms. PERKS makes use of an analytical model with machine and application parameters for predicting the performance of multi-accelerator systems and detecting their bottlenecks. Model calibration is automatic, making the model flexible and usable for different machine configurations and applications, including hypothetical ones. Our experimental results show that PERKS can predict the performance of current workloads on reconfigurable dataflow platforms with an accuracy above 91%. The results also illustrate how the modelling scales to large workloads, and how performance impact of architectural features can be estimated in seconds. Ryota Yasudo, José Gabriel F. Coutinho, Ana Lucia Varbanescu, Wayne Luk, Hideharu Amano, Tobias Becker, Ce Guo 0002 |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2020 | μ-Genie: A Framework for Memory-Aware Spatial Processor Architecture Co-Design ExplorationabstractSpatial processor architectures are essential to meet the increasing demand in performance and energy efficiency of both embedded and high performance computing systems. Due to the growing performance gap between memories and processors, the memory system of ten determines the overall performance and power consumption in silicon. The interdependency between memory system and spatial processor architectures suggests that they should be co-designed. For the same reason, state-of-the-art design methodologies for processor architectures are ineffective for spatial processor architectures because they do not include the memory system. In this paper, we present μ -Genie: an automated framework for co-design-space exploration of spatial processor architecture and the memory system, starting from an application description in a high-level programming language. In addition, we propose a spatial processor architecture template that can be configured at design-time for optimal hardware implementation. To demonstrate the effectiveness of our approach, we show a case study of co-designing a spatial processor using different memory technologies. Giulio Stramondo, Manil Dev Gomony, Bartek Kozicki, Cees T. A. M. de Laat, Ana Lucia Varbanescu |
DSD | 5 |
| 2020 | A Sampling-Based Tool for Scaling Graph Datasets
Ahmed Musaafir, Alexandru Uta, Henk Dreuning, Ana Lucia Varbanescu |
ICPE | 4 |
| 2020 | Designing and building application-centric parallel memoriesabstractSummary Memory bandwidth is a critical performance factor for many applications and architectures. Intuitively, a parallel memory could be a good solution for any bandwidth‐limited application, yet building application‐centric custom parallel memories remains a challenge. In this work, we present a comprehensive approach to tackle this challenge and demonstrate how to systematically design and implement application‐centric parallel memories. Specifically, our approach (1) analyzes the application memory access traces to extract parallel accesses, (2) configures our parallel memory for maximum performance, and (3) builds the actual application‐centric memory system. We further provide a simple performance prediction model for the constructed memory system. We evaluate our approach with two sets of experiments. First, we demonstrate how our parallel memories provide performance benefits for a broad range of memory access patterns. Second, we prove the feasibility of our approach and validate our performance model by implementing and benchmarking the designed parallel memories using FPGA hardware and a sparse version of the STREAM benchmark. Giulio Stramondo, Catalin Bogdan Ciobanu, Cees T. A. M. de Laat, Ana Lucia Varbanescu |
Concurr. Comput. Pract. Exp. | 4 |
| 2018 | Exploring HPC and Big Data Convergence: A Graph Processing Study on Intel Knights LandingabstractThe question "Can big data and HPC infrastructure converge?" has important implications for many operators and clients of modern computing. However, answering it is challenging. The hardware is currently different, and fast evolving: big data uses machines with modest numbers of fat cores per socket, large caches, and much memory, whereas HPC uses machines with larger numbers of (thinner) cores, non-trivial NUMA architectures, and fast interconnects. In this work, we investigate the convergence of big data and HPC infrastructure for one of the most challenging application domains, the highly irregular graph processing. We contrast through a systematic, experimental study of over 300,000 core-hours the performance of a modern multicore, Intel Knights Landing (KNL) and of traditional big data hardware, in processing representative graph workloads using state-of-the-art graph analytics platforms. The experimental results indicate KNL is convergence-ready, performance-wise, but only after extensive and expert-level tuning of software and hardware parameters. Alexandru Uta, Ana Lucia Varbanescu, Ahmed Musaafir, Chris Lemaire, Alexandru Iosup |
CLUSTER | 2 |
| 2018 | Performance Prediction for Large-Scale Heterogeneous PlatformsabstractThis paper presents an approach for analysing, modelling and predicting application performance of large-scale heterogeneous platforms. Our approach combines analytical and statistical modelling techniques, and aims to: (1) identify and characterise code regions that are the most promising candidates to benefit from acceleration; (2) provide statistical models that predict application behaviour for unobserved inputs; and (3) predict performance gain with different system architectures. Ryota Yasudo, Ana Lucia Varbanescu, José Gabriel F. Coutinho, Wayne Luk, Hideharu Amano |
FCCM | 2 |
| 2018 | Performance Estimation for Exascale Reconfigurable Dataflow PlatformsabstractThe next generation high-performance computing platforms will need to support exascale computing. A promising path in achieving exascale is to embrace heterogeneity and specialised computing in the form of reconfigurable accelerators. However, assessing the feasibility of heterogeneous exascale systems requires fast and accurate performance prediction. This paper proposes PERKS, a novel performance estimation frame-work for reconfigurable dataflow platforms (RDPs). PERKS uses machine and application parameters to build an analytical model for predicting the performance of multi-accelerator systems. Moreover, model calibration is automatic, making the model flexible and usable for different machine configurations and applications. Our experimental results demonstrate that PERKS can predict the performance of current workloads and RDPs with an accuracy above 95%. We also demonstrate how the modelling scales to exascale workloads and exascale platforms. Ryota Yasudo, José Gabriel F. Coutinho, Ana Lucia Varbanescu, Wayne Luk, Hideharu Amano, Tobias Becker |
FPT | 3 |
| 2018 | HLS Support for Polymorphic Parallel MemoriesabstractThe importance of High-Level Languages in abstracting machine language to enhance productivity has been proved in many sectors, and has recently encouraged the spread of reconfigurable hardware for general purpose computing. At the same time, Field Programmable Gate Arrays (FPGAs) become popular for data-intensive applications, because they promise customized hardware accelerators and achieve high-performance with low power consumption. However, taking advantage of parallel accesses to the local memories of FPGAs remains difficult, as it currently requires application re-engineering. A solution to this challenge is PolyMem, an easy-to-use parallel memory. In this work, we investigate the implementation, integration, and performance of PolyMem for HLS applications. To this end, we present a novel open-source implementation of PolyMem, optimized for the Xilinx Design Suite. We further demonstrate the use of PolyMem for three different case studies, implemented using both the Vivado workflow with a Virtex-7 VC707, and the SDx workflow with a Kintex Ultrascale 3 ADM-PCIE. Finally, we provide a thorough empirical analysis of these three cases studies in terms of latency, hardware resources, and productivity. Our results demonstrate that PolyMem delivers the expected performance, while enhancing productivity at the cost of a small increase in resources. Luca Stornaiuolo, Marco Rabozzi, Donatella Sciuto, Marco D. Santambrogio, Giulio Stramondo, Catalin Bogdan Ciobanu, Ana Lucia Varbanescu |
VLSI-SoC | 7 |
| 2017 | A NoC-based custom FPGA configuration memory architecture for ultra-fast micro-reconfigurationabstractRun-time reconfiguration in FPGAs is an important feature that offers design flexibility under low-cost silicon area and power budgets, at the cost of reconfiguration overhead. The reconfiguration time overhead produced by the conventional configuration ports (such as ICAP) is too high for the reconfiguration technology to be embraced as a standard. Furthermore, the current FPGA configuration memory architecture restricts the access of configuration data to the frame level; this significantly delays the reconfiguration process. The work presented in this paper explores the design space of the configuration memory architecture that fits the design of large FPGA's and is suitable to accomplish needs for ultra-fast reconfiguration. Therefore, the proposed method could be a stepping stone for next generation FPGA configuration memory architecture. Our simulation results show a reconfiguration speed gain of a factor of at least 1000 for substantially big parameterized applications that come with the cost of extra auxiliary hardware used on top of the column-based FPGA architecture. Amit Kulkarni 0002, Poona Bahrebar, Dirk Stroobandt, Giulio Stramondo, Catalin Bogdan Ciobanu, Ana Lucia Varbanescu |
FPT | 6 |
| 2017 | A Performance-centric Approach for Complex Decision SupportabstractMany situations in the security domain require decision-making based on complex data, i.e., many variables which need to be taken into account before adequate decisions can be made. For example, in a surveillance scenario, the size and complexity of the area of interest, the mix of objects, and the unexpected behavior of suspects are just a few examples of complex variables to be analyzed in the process. Existing decision support systems provide some analysis, but are typically limited in the complexity they can handle. Therefore, users end up with simplified models which often suffer in the accuracy of their decisions and, ultimately, may lead to incorrect decisions. In this work, we present a framework that can scale to cope with the complexity and time requirements of real-world scenarios, while remaining flexible to handle the ad-hoc adaptation to the situation. We discuss the challenges and solutions for such a scalable and flexible system, and validate it using a target tracking scenario in urban environments of different sizes. Ate Penders, Ana Lucia Varbanescu, Gregor Pavlin, Henk J. Sips |
ICPE | 2 |
| 2016 | Design and Experimental Evaluation of Distributed Heterogeneous Graph-Processing SystemsabstractGraph processing is increasingly used in a variety of domains, from engineering to logistics and from scientific computing to online gaming. To process graphs efficiently, GPU-enabled graph-processing systems such as TOTEM and Medusa exploit the GPU or the combined CPU+GPU capabilities of a single machine. Unlike scalable distributed CPU-based systems such as Pregel and GraphX, existing GPU-enabled systems are restricted to the resources of a single machine, including the limited amount of GPU memory, and thus cannot analyze the increasingly large-scale graphs we see in practice. To address this problem, we design and implement three families of distributed heterogeneous graph-processing systems that can use both the CPUs and GPUs of multiple machines. We further focus on graph partitioning, for which we compare existing graph-partitioning policies and a new policy specifically targeted at heterogeneity. We implement all our distributed heterogeneous systems based on the programming model of the single-machine TOTEM, to which we add (1) a new communication layer for CPUs and GPUs across multiple machines to support distributed graphs, and (2) a workload partitioning method that uses offline profiling to distribute the work on the CPUs and the GPUs. We conduct a comprehensive real-world performance evaluation for all three families. To ensure representative results, we select 3 typical algorithms and 5 datasets with different characteristics. Our results include algorithm run time, performance breakdown, scalability, graph partitioning time, and comparison with other graph-processing systems. They demonstrate the feasibility of distributed heterogeneous graph processing and show evidence of the high performance that can be achieved by combining CPUs and GPUs in a distributed environment. Ana Lucia Varbanescu, Dick H. J. Epema, Alexandru Iosup |
CCGrid | 2 |
| 2016 | Heterogeneous computing with accelerators: an overview with examplesabstractAccelerator-based platforms are heterogeneous in nature, yet most applications avoid heterogeneity, and focus on acceleration alone. Platform-level heterogeneity can bring significant performance improvement, as it essentially means using additional resources for the same computation. But is the performance gained using these additional resources worth the effort to program and deploy heterogeneous applications? In this work, we present a taxonomy of the existing programming models and tools available for heterogeneous computing with accelerators, and give examples of systems fitting different classes. We further provide guidelines for efficiently navigating this landscape in the search for a suitable tool for designing and deploying a new application. Ana Lucia Varbanescu, Jie Shen 0003 |
FDL | 1 |
| 2016 | The landscape of GPGPU performance modeling tools
Souley Madougou, Ana Lucia Varbanescu, Cees T. A. M. de Laat, Rob van Nieuwpoort |
Parallel Comput. | 2 |
| 2016 | Workload Partitioning for Accelerating Applications on Heterogeneous PlatformsabstractHeterogeneous platforms composed of multi-core CPUs and different types of accelerators, like GPUs and Xeon Phi, are becoming popular for data parallel applications. The heterogeneity of the hardware mix and the diversity of the applications pose significant challenges to exploiting such platforms. In this situation, an effective workload partitioning between processing units is critically important for improving application performance. This partitioning is a function of the hardware capabilities as well as the application and the dataset to be used. In this work, we present a systematic approach to solve the partitioning problem. Specifically, we use modeling, profiling, and prediction techniques to quickly and correctly predict the optimal workload partitioning and the right hardware configuration to use. Our approach effectively characterizes the platform heterogeneity, efficiently determines the accurate partitioning, and easily adapts to new platforms, different application types, and different datasets. Experimental evaluation on 13 applications shows that our approach delivers excellent performance improvement of 1.2$\times$–14.6$\times$over a single-processor execution, and accurate partitioning with in most cases below 10 percent performance gap versus an oracle-based partitioning. Jie Shen 0003, Ana Lucia Varbanescu, Yutong Lu, Henk J. Sips |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | An Empirical Performance Evaluation of GPU-Enabled Graph-Processing SystemsabstractGraph processing is increasingly used in knowledge economies and in science, in advanced marketing, social networking, bioinformatics, etc. A number of graph-processing systems, including the GPU-enabled Medusa and Totem, have been developed recently. Understanding their performance is key to system selection, tuning, and improvement. Previous performance evaluation studies have been conducted for CPU-based graph-processing systems, such as Graph and GraphX. Unlike them, the performance of GPU-enabled systems is still not thoroughly evaluated and compared. To address this gap, we propose an empirical method for evaluating GPU-enabled graph-processing systems, which includes new performance metrics and a selection of new datasets and algorithms. By selecting 9 diverse graphs and 3 typical graph-processing algorithms, we conduct a comparative performance study of 3 GPU-enabled systems, Medusa, Totem, and MapGraph. We present the first comprehensive evaluation of GPU-enabled systems with results giving insight into raw processing power, performance breakdown into core components, scalability, and the impact on performance of system-specific optimization techniques and of the GPU generation. We present and discuss many findings that would benefit users and developers interested in GPU acceleration for graph processing. Ana Lucia Varbanescu, Alexandru Iosup, Dick H. J. Epema |
CCGRID | 2 |
| 2015 | Improving Application Performance by Efficiently Utilizing Heterogeneous Many-core PlatformsabstractHeterogeneous platforms integrating different types of processing units (such as multi-core CPUs and GPUs) are in high demand in high performance computing. Existing studies have shown that using heterogeneous platforms can improve application performance and hardware utilization. However, systematic methods to design, implement, and map applications to efficiently use heterogeneous computing resources are only very few. The goal of my PhD research is therefore to study such heterogeneous systems and propose systematic methods to allow many (classes of) applications to efficiently use them. After 3.5 years of PhD study, my contributions are (1) a thorough evaluation of a suitable programming model for heterogeneous computing, (2) a workload partitioning framework to accelerate parallel applications on heterogeneous platforms, (3) a modelling-based prediction method to determine the optimal workload partitioning, (4) a systematic approach to decide the best mapping between the application and the platform by choosing the best performing hardware configuration (Only-CPU, Only-GPU, or CPU+GPU with the workload partitioning). In the near future, I plan to apply my approach to large-scale applications and platforms to expand its usability and applicability. Jie Shen 0003, Ana Lucia Varbanescu, Henk J. Sips |
CCGRID | 2 |
| 2015 | Matchmaking Applications and Partitioning Strategies for Efficient Execution on Heterogeneous PlatformsabstractHeterogeneous platforms are mixes of different processing units. The key factor to their efficient usage is workload partitioning. Both static and dynamic partitioning strategies have been defined in previous work, but their applicability and performance differ significantly depending on the application to execute. In this paper, we propose an application-driven method to select the best partitioning strategy for a given workload. To this end, we define an application classification based on the application kernel structure -- i.e., The number of kernels in the application and their execution flow. We also enable five different partitioning strategies, which mix the best features of both static and dynamic approaches. We further define the performance-driven ranking of all suitable strategies for each application class. Finally, we match the best partitioning to a given application by simply determining its class and selecting the best ranked strategy for that class. We test the matchmaking on six representative applications, and demonstrate that the defined performance ranking is correct. Moreover, by choosing the best performing partitioning strategy, we can significantly improve application performance, leading to average speedup of 3.0x/5.3x over the Only-GPU/Only-CPU execution, respectively. Jie Shen 0003, Ana Lucia Varbanescu, Xavier Martorell, Henk J. Sips |
ICPP | 2 |
| 2015 | Can Portability Improve Performance?: An Empirical Study of Parallel Graph AnalyticsabstractDue to increasingly large datasets, graph analytics - traversals, all-pairs shortest path computations, centrality measures, etc. - are becoming the focus of high-performance computing (HPC). Because HPC is currently dominated by many-core architectures (both CPUs and GPUs), new graph processing solutions have to be defined to efficiently use such computing resources. Prior work focuses on platform-specific performance studies and on platform-specific algorithm development, successfully proving that algorithms highly tuned to GPUs or multi-core CPUs can provide high performance graph analytics. However, the portability of such algorithms remains an important concern for many users, especially the many companies without the resources to invest in HPC or concerned about lock-in in single-use parallel techniques. Ana Lucia Varbanescu, Merijn Verstraaten, Cees T. A. M. de Laat, Ate Penders, Alexandru Iosup, Henk J. Sips |
ICPE | 1 |
| 2015 | Evaluating vector data type usage in OpenCL kernelsabstractSummary Open Computing Language (OpenCL) is an open, functionally portable programming model for a large range of highly parallel processors. To provide users with access to the underlying platforms, OpenCL has explicit support for features such as local memory and vector data types (VDTs). However, these are often low‐level, hardware‐specific features, which can be detrimental to performance on different platforms. In this paper, we focus on VDTs and investigate their usage in a systematic way. First, we propose two different approaches (inter‐vdtandintra‐vdt) to use VDTs in OpenCL kernels, and show how to translate scalar OpenCL kernels to vectorized ones. After obtaining vectorized code, we evaluate the performance effects of using VDTs with two types of benchmarks: micro‐benchmarks and macro‐benchmarks. With micro‐benchmarks, we study the execution model of VDTs and the role of the compiler‐aided vectorizer on five devices. With macro‐benchmarks, we explore the changes of memory access patterns before and after using VDTs, and the resulting performance impact. Not only our evaluation provides insights into how OpenCL's VDTs are mapped on different processors, but it also indicates that using such data types introduces changes in both computation and memory accesses. Based on the lessons learned, we discuss how to deal with performance portability in the presence of VDTs. Copyright © 2014 John Wiley & Sons, Ltd. Jianbin Fang, Ana Lucia Varbanescu, Xiangke Liao, Henk J. Sips |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Grover: Looking for Performance Improvement by Disabling Local Memory Usage in OpenCL KernelsabstractDue to the diversity of processor architectures and application memory access patterns, the performance impact of using local memory in OpenCL kernels has become unpredictable. For example, enabling the use of local memory for an OpenCL kernel can be beneficial for the execution on a GPU, but can lead to performance losses when running on a CPU. To address this unpredictability, we propose an empirical approach: by disabling the use of local memory in OpenCL kernels, we enable users to compare the kernel versions with and without local memory, and further choose the best performing version for a given platform. To this end, we have designed Grover, a method to automatically remove local memory usage from OpenCL kernels. In particular, we create a correspondence between the global and local memory spaces, which is used to replace local memory accesses by global memory accesses. We have implemented this scheme in the LLVM framework as a compiling pass, which automatically transforms an OpenCL kernel with local memory to a version without it. We have validated Grover with 11 applications, and found that it can successfully disable local memory usage for all of them. We have compared the kernels with and without local memory on three different processors, and found performance improvements for more than a third of the test cases after Grover disabled local memory usage. We conclude that such a compiler pass can be beneficial for performance, and, because it is fully automated, it can be used as an auto-tuning step for OpenCL kernels. Jianbin Fang, Henk J. Sips, Pekka Jääskeläinen, Ana Lucia Varbanescu |
ICPP | 4 |
| 2014 | Improving performance by matching imbalanced workloads with heterogeneous platformsabstractAlthough GPUs are considered ideal to accelerate massively data-parallel applications, there are still exceptions to this rule. For example, imbalanced applications cannot be efficiently processed by GPUs: despite the massive data parallelism, a varied computational workload per data point remains GPU-unfriendly. To efficiently process imbalanced applications, we exploit the use of heterogeneous platforms (GPUs and CPUs) by partitioning the workload to fit the usage patterns of the processors. In this work, we present our flexible and adaptive method that predicts the optimal partitioning. Our method aims to match a quantitative model of the application with the hardware capabilities of the platform, and calculates the optimal match according to a user-given criterion. We evaluate our method in terms of overall performance gain, prediction accuracy, flexibility and adaptivity. Our results, gathered from both synthetic and real-world workloads, show performance gains of up to 60%, accurate predictions for more than 90% of all the 1395 imbalanced workloads we have tested, and confirm that the method adapts correctly to application, dataset, and platform changes (both hardware and software). We conclude that model-based prediction of workload partitioning for heterogeneous platforms is feasible and useful for performance improvement. Jie Shen 0003, Ana Lucia Varbanescu, Yutong Lu, Henk J. Sips |
ICS | 2 |
| 2014 | How Well Do Graph-Processing Platforms Perform? An Empirical Performance Evaluation and AnalysisabstractGraph-processing platforms are increasingly used in a variety of domains. Although both industry and academia are developing and tuning graph-processing algorithms and platforms, the performance of graph-processing platforms has never been explored or compared in-depth. Thus, users face the daunting challenge of selecting an appropriate platform for their specific application. To alleviate this challenge, we propose an empirical method for benchmarking graph-processing platforms. We define a comprehensive process, and a selection of representative metrics, datasets, and algorithmic classes. We implement a benchmarking suite of five classes of algorithms and seven diverse graphs. Our suite reports on basic (user-lever) performance, resource utilization, scalability, and various overhead. We use our benchmarking suite to analyze and compare six platforms. We gain valuable insights for each platform and present the first comprehensive comparison of graph-processing platforms. Marcin Biczak, Ana Lucia Varbanescu, Alexandru Iosup, Claudio Martella, Theodore L. Willke |
IPDPS | 3 |
| 2014 | Test-driving Intel Xeon PhiabstractBased on Intel's Many Integrated Core (MIC) architecture, Intel Xeon Phi is one of the few truly many-core CPUs - featuring around 60 fairly powerful cores, two levels of caches, and graphic memory, all interconnected by a very fast ring. Given its promised ease-of-use and high performance, we took Xeon Phi out for a test drive. In this paper, we present this experience at two different levels: (1) the microbenchmark level, where we stress "each nut and bolt" of Phi in the lab, and (2) the application level, where we study Phi's performance response in a real-life environment. At the microbenchmarking level, we show the high performance of five components of the architecture, focusing on their maximum achieved performance and the prerequisites to achieve it. Next, we choose a medical imaging application (Leukocyte Tracking) as a case study. We observed that it is rather easy to get functional code and start benchmarking, but the first performance numbers can be far from satisfying. Our experience indicates that a simple data structure and massive parallelism are critical for Xeon Phi to perform well. When compiler-driven parallelization and/or vectorization fails, programming Xeon Phi for performance can become very challenging. Jianbin Fang, Henk J. Sips, Lilun Zhang, Chuanfu Xu, Yonggang Che, Ana Lucia Varbanescu |
ICPE | 6 |
| 2014 | Benchmarking graph-processing platforms: a visionabstractProcessing graphs, especially at large scale, is an increasingly useful activity in a variety of business, engineering, and scientific domains. Already, there are tens of graph-processing platforms, such as Hadoop, Giraph, GraphLab, etc., each with a different design and functionality. For graph-processing to continue to evolve, users have to find it easy to select a graph-processing platform, and developers and system integrators have to find it easy to quantify the performance and other non-functional aspects of interest. However, the state of performance analysis of graph-processing platforms is still immature: there are few studies and, for the few that exist, there are few similarities, and relatively little understanding of the impact of dataset and algorithm diversity on performance. Our vision is to develop, with the help of the performance-savvy community, a comprehensive benchmarking suite for graph-processing platforms. In this work, we take a step in this direction, by proposing a set of seven challenges, summarizing our previous work on performance evaluation of distributed graph-processing platforms, and introducing our on-going work within the SPEC Research Group's Cloud Working Group. Ana Lucia Varbanescu, Alexandru Iosup, Claudio Martella, Theodore L. Willke |
ICPE | 2 |
| 2014 | Cross-Loop Optimization of Arithmetic Intensity for Finite Element Local AssemblyabstractWe study and systematically evaluate a class of composable code transformations that improve arithmetic intensity in local assembly operations, which represent a significant fraction of the execution time in finite element methods. Their performance optimization is indeed a challenging issue. Even though affine loop nests are generally present, the short trip counts and the complexity of mathematical expressions, which vary among different problems, make it hard to determine an optimal sequence of successful transformations. Our investigation has resulted in the implementation of a compiler (called COFFEE) for local assembly kernels, fully integrated with a framework for developing finite element methods. The compiler manipulates abstract syntax trees generated from a domain-specific language by introducing domain-aware optimizations for instruction-level parallelism and register locality. Eventually, it produces C code including vector SIMD intrinsics. Experiments using a range of real-world finite element problems of increasing complexity show that significant performance improvement is achieved. The generality of the approach and the applicability of the proposed code transformations to other domains is also discussed. Fabio Luporini, Ana Lucia Varbanescu, Florian Rathgeber, Gheorghe-Teodor Bercea, J. Ramanujam, David A. Ham, Paul H. J. Kelly |
ACM Trans. Archit. Code Optim. | 2 |
| 2013 | Sesame: A User-Transparent Optimizing Framework for Many-Core ProcessorsabstractWith the integration of more computational cores and deeper memory hierarchies on modern processors, the performance gap between naively parallel zed code and optimized code becomes much larger than ever before. Very often, bridging the gap involves architecture-specific optimizations. These optimizations are difficult to implement by application programmers, who typically focus on the basic functionality of their code. Therefore, in this thesis, I focus on answering the following research question: "How can we address architecture-specific optimizations in a programmer-friendly way?'' As an answer, I propose an optimizing framework for parallel applications running on many-core processors (\textit{Sesame}). Taking a simple parallel zed code provided by the application programmers as input, Sesame chooses and applies the most suitable architecture-specific optimizations, aiming to improve the overall application performance in a user-transparent way. In this short paper, I present the motivation for designing and implementing Sesame, its structure and its modules. Furthermore, I describe the current status of Sesame, discussing our promising results in source-to-source vectorization, automated usage of local memory, and auto-tuning for implementation-specific parameters. Finally, I discuss my work-in-progress and sketch my ideas for finalizing Sesame's development and testing. Jianbin Fang, Ana Lucia Varbanescu, Henk J. Sips |
CCGRID | 2 |
| 2013 | Topic 9: Parallel and Distributed Programming - (Introduction)
Michael Philippsen, Domenico Talia, Ana Lucia Varbanescu |
Euro-Par | 4 |
| 2013 | ELMO: A User-Friendly API to Enable Local Memory in OpenCL KernelsabstractRecent parallel architectures are equipped with local memory, which simplifies hardware design at the cost of increased program complexity due to explicit management. To simplify this extra-burden that programmers have, we introduce an easy-to-use API, ELMO, that improves productivity while preserving high performance of local memory operations. Specifically, ELMO is a generic API that covers different local memory use-cases. We also present prototype implementations for these APIs and perform multiple GPU-inspired optimizations to maximize their performance. Experimental results on the NVIDIA Quadro5000 GPU show that performance is significantly improved by using ELMO on native implementations: the achieved speedup ranges from 1.3x to 3.7x. Furthermore, using ELMO we still achieve performance comparable (if not better) with that of hand-tuned applications, while the code is shorter, clearer, and safer. Jianbin Fang, Ana Lucia Varbanescu, Jie Shen 0003, Henk J. Sips |
PDP | 2 |
| 2013 | Performance Traps in OpenCL for CPUsabstractWith its design concept of cross-platform portability, OpenCL can be used not only on GPUs (for which it is quite popular), but also on CPUs. Whether porting GPU programs to CPUs, or simply writing new code for CPUs, using OpenCL brings up the performance issue, usually raised in one of two forms: "OpenCL is not performance portable!" or "Why using OpenCL for CPUs after all?!". We argue that both issues can be addressed by a thorough study of the factors that impact the performance of OpenCL on CPUs. This analysis is the focus of this paper. Specifically, starting from the two main architectural mismatches between many-core CPUs and the OpenCL platform-parallelism granularity and the memory model-we identify eight such performance "traps" that lead to performance degradation in OpenCL for CPUs. Using multiple code examples, from both synthetic and real-life benchmarks, we quantify the impact of these traps, showing how avoiding them can give up to 10 times better performance. Furthermore, we point out that the solutions we provide for avoiding these traps are simple and generic code transformations, which can be easily adopted by either programmers or automated tools. Therefore, we conclude that a certain degree of OpenCL inter-platform performance portability, while indeed not a given, can be achieved by simple and generic code transformations. Jie Shen 0003, Jianbin Fang, Henk J. Sips, Ana Lucia Varbanescu |
PDP | 4 |
| 2013 | An application-centric evaluation of OpenCL on multi-core CPUs
Jie Shen 0003, Jianbin Fang, Henk J. Sips, Ana Lucia Varbanescu |
Parallel Comput. | 4 |
| 2012 | Accelerating Cost Aggregation for Real-Time Stereo MatchingabstractReal-time stereo matching, which is important in many applications like self-driving cars and 3-D scene reconstruction, requires large computation capability and high memory bandwidth. The most time-consuming part of stereo-matching algorithms is the aggregation of information (i.e. costs) over local image regions. In this paper, we present a generic representation and suitable implementations for three commonly used cost aggregators on many-core processors. We perform typical optimizations on the kernels, which leads to significant performance improvement (up to two orders of magnitude). Finally, we present a performance model for the three aggregators to predict the aggregation speed for a given pair of input images on a given architecture. Experimental results validate our model with an acceptable error margin (an average of 10.4%). We conclude that GPU-like many-cores are excellent platforms for accelerating stereo matching. Jianbin Fang, Ana Lucia Varbanescu, Jie Shen 0003, Henk J. Sips, Gorkem Saygili, Laurens van der Maaten |
ICPADS | 2 |
| 2012 | Radio Astronomy Beam Forming on Many-Core ArchitecturesabstractTraditional radio telescopes use large steel dishes to observe radio sources. The largest radio telescope in the world, LOFAR, uses tens of thousands of fixed, omni-directional antennas instead, a novel design that promises ground-breaking research in astronomy. Where traditional tele-scopes use custom-built hardware, LOFAR uses software to do signal processing in real time. This leads to an instrument that is inherently more flexible. However, the enormous data rates and processing requirements (tens to hundreds of teraflops) make this extremely challenging. The next-generation telescope, the SKA, will require exaflops. Unlike traditional instruments, LOFAR and SKA can observe in hundreds of directions simultaneously, using beam forming. This is useful, for example, to search the sky for pulsars (i.e. rapidly rotating highly magnetized neutron stars). Beam forming is an important technique in signal processing: it is also used in WIFI and 4G cellular networks, radar systems, and health-care microwave imaging instruments. We propose the use of many-core architectures, such as 48-core CPU systems and Graphics Processing Units (GPUs), to accelerate beam forming. We use two different frameworks for GPUs, CUDA and Open CL, and present results for hardware from different vendors (i.e. AMD and NVIDIA). Additionally, we implement the LOFAR beam former on multi-core CPUs, using Open MP with SSE vector instructions. We use auto-tuning to support different architectures and implementation frameworks, achieving both platform and performance portability. Finally, we compare our results with the production implementation, written in assembly and running on an IBM Blue Gene/P supercomputer. We compare both computational and power efficiency, since power usage is one of the fundamental challenges modern radio telescopes face. Compared to the production implementation, our auto-tuned beam former is 45-50 times faster on GPUs, and 2-8 times more power efficient. Our experimental results lead to the conclusion that GPUs are an attractive solution to accelerate beam forming. Alessio Sclocco, Ana Lucia Varbanescu, Jan David Mol, Rob van Nieuwpoort |
IPDPS | 2 |
| 2012 | Parallel application characterization with quantitative metricsabstractSUMMARY When computer architects reinvented parallelism through multi‐core processors, application parallelization became a problem. Now that multi‐cores have penetrated from handhelds to supercomputers, parallelization becomes a large‐scale challenge. A lot of research is going into compiler improvements, language extensions, frameworks and application/platform case studies. Whereas fairly successful, these solutions are based on experimental tools, trial‐and‐error, and expert knowledge, and do not bring multi‐core programming into reach for the whole software industry. We believe that the challenge of “mass parallelization” must be tackled more systematically. Development begins at application specification and algorithm design, followed by application characterization with trade‐offs in parallelization strategies and data layouts. Only with a proper software design, implementation and optimization can start. In this article, we focus on quantitative application characterization for such a systematic approach. We introduce a set of metrics to characterize applications and show how they can be used. We present our interpretation of the results and suggest ways to use them to guide design decisions. We conclude that metrics can be used to understand applications and design decisions early on. Therefore, this characterization brings us closer to effective parallel applications development for multi‐core processors. Copyright © 2011 John Wiley & Sons, Ltd. Alexander S. van Amesfoort, Ana Lucia Varbanescu, Henk J. Sips |
Concurr. Comput. Pract. Exp. | 2 |
| 2011 | A Comprehensive Performance Comparison of CUDA and OpenCLabstractThis paper presents a comprehensive performance comparison between CUDA and OpenCL. We have selected 16 benchmarks ranging from synthetic applications to real-world ones. We make an extensive analysis of the performance gaps taking into account programming models, ptimization strategies, architectural details, and underlying compilers. Our results show that, for most applications, CUDA performs at most 30% better than OpenCL. We also show that this difference is due to unfair comparisons: in fact, OpenCL can achieve similar performance to CUDA under a fair comparison. Therefore, we define a fair comparison of the two types of applications, providing guidelines for more potential analyses. We also investigate OpenCL's portability by running the benchmarks on other prevailing platforms with minor modifications. Overall, we conclude that OpenCL's portability does not fundamentally affect its performance, and OpenCL can be a good alternative to CUDA. Jianbin Fang, Ana Lucia Varbanescu, Henk J. Sips |
ICPP | 2 |
| 2011 | OCL-BodyScan: A Case Study for Application-centric Programming of Many-Core ProcessorsabstractApplication development for many-core processors is predominately hardware-centric: programmers design, implement, and optimize applications for a pre-chosen target platform. While this approach may deliver very good performance, it lacks portability, being inefficient for applications that aim to use multiple architectures or large-scale parallel platforms with heterogeneous many-core nodes. In this work, we focus on application portability. Therefore, we propose an application-centric approach for developing parallel workloads for many-cores, and we make use of OpenCL to preserve portability until the very last optimization stages. We validate our application-centric approach using 3D body scan, a data intensive application with soft real-time constraints. Thus, we design and implement OCL-body scan (the portable OpenCL-based version of 3D Body scan), and we evaluate its performance on three families of platforms - general purpose multi-cores, graphical processing units, and the Cell/B.E.. Our experiments show that our application-centric strategy enables portability and leads to good performance results. Additionally, typical platform-specific optimizations can be applied in the final implementation stages, leading to performance results similar to those obtained using the native tool-chains. Milos Raskovic, Ana Lucia Varbanescu, Wouter Vlothuizen, Maarten Ditzel, Henk J. Sips |
ICPP | 2 |
| 2009 | Evaluating application mapping scenarios on the Cell/B.EabstractAbstract Applications running on multicore platforms are difficult to program, and even more difficult to optimize, mainly due to (1) the several layers where the optimizations occur and (2) the multitude of available resources to be exploited in parallel. Although low‐level optimizations only target code running on individual cores, high‐level optimizations (e.g. data‐ and task‐parallelism) target the overall application performance. In this paper, we focus on the latter, by evaluating possible mapping scenarios of a real application on a heterogeneous multicore processor. Specifically, we focus on analyzing the impact of combining data‐ and task‐parallelism for a multimedia analysis application running on the Cell Broadband Engine (Cell/B.E.). We find that both low‐level and high‐level optimizations are important for the overall application speed‐up. However, we show that a speed‐up factor of over 20 for the application running on Cell/B.E. can only be obtained if core utilization is increased by combining data‐ and task‐parallelism. Thus, we consider this case study essential for building expertise in both application optimization and performance analysis for multicore platforms. Copyright © 2008 John Wiley & Sons, Ltd. Ana Lucia Varbanescu, Henk J. Sips, Kenneth A. Ross, Apostol Natsev, John R. Smith, Lurng-Kuo Liu |
Concurr. Comput. Pract. Exp. | 1 |
| 2008 | Radioastronomy Image Synthesis on the Cell/B.E
Ana Lucia Varbanescu, Alexander S. van Amesfoort, Tim Cornwell, Andrew Mattingly, Bruce G. Elmegreen, Rob van Nieuwpoort, Ger van Diepen, Henk J. Sips |
Euro-Par | 1 |
| 2007 | Digital Media Indexing on the Cell ProcessorabstractWe present a case study of developing a digital media indexing application, code-named MARVEL, on the STI cell broadband engine (CBE) processor. There are two aspects of the target application that require significant computing power: image analysis for feature extraction, and support vector machine (SVM) based pattern classification for concept detection. We discuss the mapping of a large application like MARVEL onto a multicore processor, and show how feature extraction and concept detection can be implemented on the CBE. We discuss how the synergistic processing units of a CBE can be used to gain dramatic performance improvements. The empirical results of our experiments, conducted on a Cell blade running at 3.2 GHz, show that the CBE provides a significant performance speed-up in our digital media indexing application. Lurng-Kuo Liu, Apostol Natsev, Kenneth A. Ross, John R. Smith, Ana Lucia Varbanescu |
ICME | 6 |
| 2007 | An Effective Strategy for Porting C++ Applications on CellabstractIn this paper we present a solution for efficient porting of sequential C++ applications on the Cell B.E. processor. We present our step-by-step approach, focusing on its generality, we provide a set of code templates and optimization guidelines to support the porting, and we include a set of equations to estimate the performance gain of the new application. As a case-study, we show the use of our solution on a multimedia content analysis application, named MARVEL. The results of our experiments with MARVEL prove the significant performance increase in favor of the application running on Cell when compared with the reference implementation. Ana Lucia Varbanescu, Henk J. Sips, Kenneth A. Ross, Lurng-Kuo Liu, Apostol Natsev, John R. Smith |
ICPP | 1 |
| 2007 | Multicore Surprises: Lessons Learned from Optimizing Sweep3D on the Cell Broadband EngineabstractThe Cell Broadband Engine (BE) processor provides the potential to achieve an impressive level of performance for scientific applications. This level of performance can be reached by exploiting several dimensions of parallelism, such as thread-level parallelism using several synergistic processing elements, data streaming parallelism, vector parallelism in the form of 128-bit SIMD operations, and pipeline parallelism by issuing multiple instructions in the same clock cycle. In our exploration to achieve the optimum level of performance for Sweep3D, we have enjoyed many pleasant surprises, such as a very high floating point performance, reaching 64% of the theoretical peak in double precision, and an over all performance speedup ranging from 4.5 times when compared with "heavy iron" processors, up to over 20 times with conventional processors. Fabrizio Petrini, Gordon C. Fossum, Juan Fernández Peinador, Ana Lucia Varbanescu, Michael Kistler, Michael Perrone |
IPDPS | 4 |
| 2006 | PAM-SoC: A Toolchain for Predicting MPSoC Performance
Ana Lucia Varbanescu, Henk J. Sips, Arjan J. C. van Gemund |
Euro-Par | 1 |