EDBT 2026 Demo / reviewers in the wild / expert
Gabriele Capannini
dblp:05/530
· DBLP profile ↗
15ranked-venue papers
7as first author
4since 2021 · last 2024
0000-0002-2558-5354ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Predicting Cache Behaviour of Concurrent ApplicationsabstractModern digital solutions are built around a variety of applications. The continuous integration of these applications brings advancements in technology. Therefore, it is essential to understand how these applications will behave when they run together. However, this can be challenging to interpret due to the increasing complexity of the execution details. One such fundamental detail is the utilization of shared cache as it goes hand in hand with the computation capacity of computer systems. Since cache utilization behavior is not simple enough to translate with few assumptions we have investigated if this complex behavior can be predicted with the help of machine learning. We trained the deep neural network with enough examples that represent the cache behavior when applications were running alone and when they were running concurrently on the same core. The Long Short-Term Memory (LSTM) network learns the entire execution period of each application in the training set. As a result, without running two applications together in reality, provided with the L1 cache misses of two applications (running alone), it can predict how the cache will look like if two applications wish to run together. The model returns a time series that reflects the cache behavior in concurrency. Shamoona Imtiaz, Moris Behnam, Gabriele Capannini, Jan Carlson, Marcus Jägemar |
ETFA | 3 |
| 2023 | Automatic Clustering of Performance EventsabstractModern hardware and software are becoming increasingly complex due to advancements in digital and smart solutions. This is why industrial systems seek efficient use of resources to confront the challenges caused by the complex resource utilization demand. The demand and utilization of different resources show the particular execution behavior of the applications. One way to get this information is by monitoring performance events and understanding the relationship among them. However, manual analysis of this huge data is tedious and requires experts’ knowledge. This paper focuses on automatically identifying the relationship between different performance events. Therefore, we analyze the data coming from the performance events and identify the points where their behavior changes. Two events are considered related if their values are changing at "approximately" the same time. We have used the Sigmoid function to compute a real-value similarity between two sets (representing two events). The resultant value of similarity is induced as a similarity or distance metric in a traditional clustering algorithm. The proposed solution is applied to 6 different software applications that are widely used in industrial systems to show how different setups including the selection of cost functions can affect the results. Shamoona Imtiaz, Gabriele Capannini, Jan Carlson, Moris Behnam, Marcus Jägemar |
ETFA | 2 |
| 2022 | Reliable Visibility Algorithms for Emergency Stop Systems in Smart IndustriesabstractAutomated machinery and robots working with humans are the norm in modern smart industries. A previous work in this area proposed a tool for improving the safety of such work places: an emergency system which halts those machines that are visible from an emergency stop button when it is pressed [1]. The solution presented in this paper improves the reliability of the aforementioned one at the expense of a higher computational complexity. Furthermore, two algorithmic optimizations are presented to mitigate the extra computational cost as it is shown by the results collected from the set of experiments conducted. Gabriele Capannini, Jan Carlson, Roger Mellander |
COMPSAC | 1 |
| 2021 | Automatic Platform-Independent Monitoring and Ranking of Hardware Resource UtilizationabstractIn this paper, we discuss a method for automatic monitoring of hardware and software events using performance monitoring counters. Computer applications are complex and utilize a broad spectra of the available hardware resources, where multiple performance counters can be of significant interest to understand. The number of performance counters that can be captured simultaneously is, however, small due to hardware limitations in most modern computers. We suggest a platform independent solution to automatically retrieve hardware events from an underlying architecture. Moreover, to mitigate the hardware limitations we propose a mechanism that pinpoints the most relevant performance counters for an application's performance. In our proposal, we utilize the Pearson's correlation coefficient to rank the most relevant performance counters and filter out those that are most relevant and ignore the rest. Shamoona Imtiaz, Jakob Danielsson, Moris Behnam, Gabriele Capannini, Jan Carlson, Marcus Jägemar |
ETFA | 4 |
| 2018 | Adaptive Collision Culling for Massive Simulations by a Parallel and Context-Aware Sweep and Prune AlgorithmabstractWe present an improved parallel Sweep and Prune algorithm that solves the dynamic box intersection problem in three dimensions. It scales up to very large datasets, which makes it suitable for broad phase collision detection in complex moving body simulations. Our algorithm gracefully handles high-density scenarios, including challenging clustering behavior, by using a double-axis sweeping approach and a cache-friendly succinct data structure. The algorithm is realized by three parallel stages for sorting, candidate generation, and object pairing. By the use of temporal coherence, our sorting stage runs with close to optimal load balancing. Furthermore, our approach is characterized by a work-division strategy that relies on adaptive partitioning, which leads to almost ideal scalability. In addition, for scenarios that involves intense clustering along several axes simultaneously, we propose an enhancement that increases the context-awareness of the algorithm. By exploiting information gathered along three orthogonal axes, an efficient choice of what range query to perform can be made per object during run-time. Experimental results show high performance for up to millions of objects on modern multi-core CPUs. Gabriele Capannini, Thomas Larsson |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2016 | Parallel computation of optimal enclosing balls by iterative orthant scan
Thomas Larsson, Gabriele Capannini, Linus Källberg |
Comput. Graph. | 2 |
| 2016 | Quality versus efficiency in document scoring with learning-to-rank models
Gabriele Capannini, Claudio Lucchese, Franco Maria Nardini, Salvatore Orlando 0001, Raffaele Perego 0001, Nicola Tonellotto |
Inf. Process. Manag. | 1 |
| 2013 | A multi-criteria job scheduling framework for large computing farms
Ranieri Baraglia, Gabriele Capannini, Patrizio Dazzi, Giancarlo Pagano |
J. Comput. Syst. Sci. | 2 |
| 2012 | Sorting on GPUs for large scale datasets: A thorough comparison
Gabriele Capannini, Fabrizio Silvestri, Ranieri Baraglia |
Inf. Process. Manag. | 1 |
| 2011 | A Parallel Code for Time Independent Quantum Reactive Scattering on CPU-GPU Platforms
Ranieri Baraglia, Malko Bravi, Gabriele Capannini, Antonio Laganà, Edoardo Zambonini |
ICCSA (3) | 3 |
| 2011 | Efficient Diversification of Web Search ResultsabstractIn this paper we analyze the efficiency of various search results diversification methods. While efficacy of diversification approaches has been deeply investigated in the past, response time and scalability issues have been rarely addressed. A unified framework for studying performance and feasibility of result diversification solutions is thus proposed. First we define a new methodology for detecting when, and how, query results need to be diversified. To this purpose, we rely on the concept of "query refinement" to estimate the probability of a query to be ambiguous . Then, relying on this novel ambiguity detection method, we deploy and compare on a standard test set, three different diversification methods: IASelect, xQuAD, and OptSelect. While the first two are recent state-of-the-art proposals, the latter is an original algorithm introduced in this paper. We evaluate both the efficiency and the effectiveness of our approach against its competitors by using the standard TREC Web diversification track testbed. Results shown that OptSelect is able to run two orders of magnitude faster than the two other state-of-the-art approaches and to obtain comparable figures in diversification effectiveness. Gabriele Capannini, Franco Maria Nardini, Raffaele Perego 0001, Fabrizio Silvestri |
Proc. VLDB Endow. | 1 |
| 2011 | A multi-level scheduler for batch jobs on grids
Marco Pasquali, Ranieri Baraglia, Gabriele Capannini, Laura Ricci, Domenico Laforenza |
J. Supercomput. | 3 |
| 2010 | K-Model: A New Computational Model for Stream ProcessorsabstractWe introduce K-model, a computational model to evaluate the algorithms designed for graphic processors, and other architectures adhering to the stream programming model. We address the lack of a formal complexity model that properly accounts for memory contention, address coalescing in memory accesses, or the serial control of instruction flows. We study the impact of K-model rules on algorithm design. We devise a coalesced and low contention data access technique for Batcher's networks, and we evaluate the effectiveness of this technique within our K-model. To evaluate the benefits in using K-model in evaluating solutions for streaming architectures, we compare the complexity of a sorting network built using our technique, and quick sort. Although in theory quick sort is more efficient than bitonic sort, empirically, our bitonic sorting network has been shown to be faster than the state-of-the-art implementation of quick sort on graphics processing units (GPUs). We use our K-model to prove that this observation should generally hold. As a side result, our technique to perform a Batcher's network on GPUs improves the performance of one the fastest comparison-based solution for integers sorting. Gabriele Capannini, Fabrizio Silvestri, Ranieri Baraglia |
HPCC | 1 |
| 2008 | A two-level scheduler to dynamically schedule a stream of batch jobs in large-scale gridsabstractThis paper describes the study conducted to design and evaluate a two-level on-line scheduler to dynamically schedule a stream of sequential and multi-threaded batch jobs on large scale grids, made up of interconnected clusters of heterogeneous machines. The scheduler aims to schedule arriving jobs respecting their computational and deadline requirements, and optimizing the utilization of hardware resources as well as software resources. Marco Pasquali, Ranieri Baraglia, Gabriele Capannini, Laura Ricci, Domenico Laforenza |
HPDC | 3 |
| 2007 | A job scheduling framework for large computing farmsabstractIn this paper, we propose a new method, called Convergent Scheduling, for scheduling a continuous stream of batch jobs on the machines of large-scale computing farms. This method exploits a set of heuristics that guide the scheduler in making decisions. Each heuristics manages a specific problem constraint, and contributes to carry out a value that measures the degree of matching between a job and a machine. Scheduling choices are taken to meet the QoS requested by the submitted jobs, and optimizing the usage of hardware and software resources. We compared it with some of the most common job scheduling algorithms, i.e. Backfilling, and Earliest Deadline First. Convergent Scheduling is able to compute good assignments, while being a simple and modular algorithm. Gabriele Capannini, Ranieri Baraglia, Diego Puppin, Laura Ricci, Marco Pasquali |
SC | 1 |