Stefano Bortoli

dblp:88/4859 · DBLP profile ↗
← Back
18ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0003-1565-3007ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Operator Rebinding for Stream Processing on NUMA Machines
abstract
ABSTRACT Introduction Modern stream processing engines are increasingly deployed on high‐core‐count servers with Non‐Uniform Memory Access (NUMA) architectures, where the cost of inter‐socket memory access poses a significant challenge to achieving low latency and high throughput. Existing approaches to operator placement either rely on static assignments that degrade under workload variations or employ dynamic migrations that incur excessive overhead due to blocking synchronization or global barriers. Methods This paper introduces a lock‐free, NUMA‐aware operator rebinding mechanism that dynamically reallocates operator tasks across threads with minimal disruption. The mechanism uses an autonomic controller to detect imbalance in per‐thread queues and enacts rebinding via control messages and atomic updates, ensuring correctness without stalling execution. A two‐level policy is proposed, combining NUMA‐level partitioning with intra‐node thread‐level refinements, triggered by latency thresholds. Results Extensive experiments using a 300‐query urban traffic analytics workload demonstrate that the proposed method achieves non‐negligible throughput improvement and reduces latency compared to state‐of‐the‐art static and METIS‐based approaches. Furthermore, it reduces latency variance by an order of magnitude, illustrating the importance of fine‐grained NUMA‐aware scheduling in memory‐bound stream processing.
Xiaorui Du, Andrea Piccione, Adriano Pimpini, Stefano Bortoli, Alessandro Pellegrini 0001, Alois C. Knoll
Softw. Pract. Exp.4
2025 Autonomic Partition-Aware Malleable Microscopic Traffic Simulation
abstract
In this work, we present Autonomic CityMoS, a malleable parallel, and distributed traffic simulator engine that can automatically adapt the number of computing nodes in response to dynamic computational demands. We combine a snapshot system that enables data distribution, two predictive cost models that estimate the system speedup based on key metrics, including partitioning characteristics, and a policy that leverages these models to maintain a steady simulation pace. Autonomic CityMoS is able to effectively keep the simulation pace under varying traffic pattern flows, all without prior knowledge of traffic conditions. Although adaptation times are significant, we still observe an improvement in resource utilization compared with a run with static allocation of compute resources. The work presented in this paper should serve as an example of malleable simulation execution, where the objective is not to maximize performance but rather ensure a sustainable execution of large distributed simulation targeting to optimize the trade-off between target speed-up and overall compute resource utilization. Target applications include, but are not limited to, very large visual and interactive simulations.
Anibal Siguenza-Torres, Santiago Narvaez Rivas, Alexander Wieder, Andrea Piccione, Stefano Bortoli, Wentong Cai 0001, Hans-Joachim Bungartz, Alois C. Knoll
SIGSIM-PADS5
2024 HUILLY: A Non-Blocking Ingestion Buffer for Timestepped Simulation Analytics
abstract
We present HUILLY, a non-blocking data ingestion buffer designed for parallel applications built relying on time-stepped, fork-join computational paradigm. It provides complete data separation, reducing the intricacies of multi-threaded data structures and provides high operational efficiency. The effectiveness of HUILLY as a non-blocking ingestion buffer is demonstrated through an extensive experimental evaluation, providing significant improvements in throughput and latency for time-stepped, fork-join applications.
Xiaorui Du, Andrea Piccione, Adriano Pimpini, Stefano Bortoli, Alois C. Knoll, Alessandro Pellegrini 0001
CCGrid4
2024 Online Analytics with Local Operator Rebinding for Simulation Data Stream Processing
abstract
Leveraging multiple threads to process high volumes of simulation data is a prevalent strategy in modern streaming data processing systems. Statically binding operators to specific threads is the most common design employed due to its simplicity in implementation and initial system configuration. However, this approach often fails to effectively account for the inherently dynamic nature of simulation data, potentially leading to inefficient resource utilisation and processing bottlenecks. To address these limitations, we present a novel mechanism for stream-processing operator rebinding that enables lock-free, dynamic workload rebalancing between worker threads. The rebinding is driven by an autonomic policy that captures workload imbalance in the stream-processing pipeline when multiple queries are computed and reacts to it by moving computation around the different threads. We evaluate our proposal using data generated from large-scale traffic simulations on which multiple queries are executed. The volume and organisation of the data we feed to the stream-processing pipeline significantly change over time, providing excellent grounds to evaluate our rebinding policy. The evaluation confirms that the performance of stream processing pipelines can be greatly improved using local operator rebinding.
Xiaorui Du, Andrea Piccione, Adriano Pimpini, Stefano Bortoli, Alessandro Pellegrini 0001, Alois C. Knoll
DS-RT4
2024 Serialization-Oriented Data Layout for Distributed and Real-Time Agent-Based Simulation
abstract
Data transfer efficiency is a frequent bottleneck of distributed (co-)simulations and X-in-the-loop systems. One of the key reasons, particularly in Agent-Based Simulation (ABS), is related to the low serialization performance caused by non-optimal data layout in memory. To address these challenges, this paper explores the potentials of a Serialization-oriented Data Layout (SoDaLa) approach for ABS building on Data Oriented Design (DOD) principles. In this work we also introduce ABS_M, a model-based abstraction of memory access in ABS. Using this model, we evaluate the impact of SoDaLa. This is done also to promote the adoption of SoDaLa and to ease the assessment of data layout strategies prior to full implementation in an existing codebases. The results indicate that our proposed approach enhances serialization efficiency and highlights the trade-offs between serialization efficiency and simulation performance across different model specifics and hardware conditions.
Zhuoxiao Meng, Mingyue Gao, Stefano Bortoli, Christoph Sommer 0001, Alois C. Knoll
DS-RT3
2024 Benchmarking Stream Join Algorithms on GPUs: A Framework and its Application to the State-of-the-art
Dwi P. A. Nugroho, Philipp M. Grulich, Steffen Zeuch, Clemens Lutz, Stefano Bortoli, Volker Markl
EDBT5
2024 Fuzzy modelling and inference for physics-aware road vehicle driver behaviour model calibration
Cristian Axenie, Wolfgang Scherr, Alexander Wieder, Anibal Siguenza-Torres, Zhuoxiao Meng, Xiaorui Du, Paolo Sottovia, Daniele Foroni, Margherita Grossi, Stefano Bortoli, Goetz Brasche
Expert Syst. Appl.10
2023 Towards Discrete-Event, Aggregating, and Relational Control Interfaces for Traffic Simulation
abstract
The use of IoT and AI/ML to extract insights for Data-Driven Decision-Making (DDDM) in Intelligent Traffic Systems (ITS) is becoming increasingly popular. While simulation is a cost-effective and safe way to evaluate such approaches, existing simulators are often impractical due to inefficient control interfaces. In this work, we propose a Discrete-Event, Aggregating, and Relational Control Interfaces (DAR-CI) framework for achieving efficient traffic management simulations through a coupled approach. It enables a non-blocking interaction mode based on a discrete-event synchronization architecture. The overhead caused by data exchange is substantially reduced by supporting the direct retrieval of temporal metrics, data batch processing and customized in-situ aggregation. Combined with flexible, extendable, easy-to-understand, and implementation-friendly semantic specifications, we propose DAR-CI to serve as a universal tool for the traffic simulation community, taking the use and control of traffic simulation to a new level. A proof-of-concept study on the simulation of an adaptive traffic light control system demonstrates a 9.53X speedup compared to TraCI, a widely used protocol for controlling traffic simulators.
Zhuoxiao Meng, Anibal Siguenza-Torres, Mingyue Gao, Margherita Grossi, Alexander Wieder, Xiaorui Du, Stefano Bortoli, Christoph Sommer 0001, Alois C. Knoll
SIGSIM-PADS7
2021 Querying Top-k Dominant Traffic Flows on Large Urban Road Networks
Stella Maropaki, Paolo Sottovia, Stefano Bortoli
EDBT3
2021 OBELISC: Oscillator-Based Modelling and Control Using Efficient Neural Learning for Intelligent Road Traffic Signal Calculation
Cristian Axenie, Rongye Shi, Daniele Foroni, Alexander Wieder, Mohamad Al Hajj Hassan, Paolo Sottovia, Margherita Grossi, Stefano Bortoli, Goetz Brasche
ECML/PKDD (4)8
2020 Real-time Traffic Jam Detection and Congestion Reduction Using Streaming Graph Analytics
abstract
Traffic congestion is a problem in day to day life, especially in big cities. Various traffic control infrastructure systems have been deployed to monitor and improve the flow of traffic across cities. Real-time congestion detection can serve for many useful purposes that include sending warnings to drivers approaching the congested area and daily route planning. Most of the existing congestion detection solutions combine historical data with continuous sensor readings and rely on data collected from multiple sensors deployed on the road, measuring the speed of vehicles. While in our work we present a framework that works in a pure streaming setting where historic data is not available before processing. The traffic data streams, possibly unbounded, arrive in real-time. Moreover, the data used in our case is collected only from sensors placed on the intersections of the road. Therefore, we investigate in creating a real-time congestion detection and reduction solution, that works on traffic streams without any prior knowledge. The goal of our work is 1) to detect traffic jams in real-time, and 2) to reduce the congestion in the traffic jam areas.In this work, we present a real-time traffic jam detection and congestion reduction framework: 1) We propose a directed weighted graph representation of the traffic infrastructure network for capturing dependencies between sensor data to measure traffic congestion; 2) We present online traffic jam detection and congestion reduction techniques built on a modern stream processing system, i.e., Apache Flink; 3) We develop dynamic traffic light policies for controlling traffic in congested areas to reduce the travel time of vehicles. Our experimental results indicate that we are able to detect traffic jams in real-time and deploy new traffic light policies which result in 27% less travel time at the best and 8% less travel time on average compared to the travel time with default traffic light policies. Our scalability results show that our system is able to handle high-intensity streaming data with high throughput and low latency.
Zainab Abbas, Paolo Sottovia, Mohamad Al Hajj Hassan, Daniele Foroni, Stefano Bortoli
IEEE BigData5
2019 NARPCA: Neural Accumulate-Retract PCA for Low-Latency High-Throughput Processing on Datastreams
Cristian Axenie, Radu Tudoran, Stefano Bortoli, Mohamad Al Hajj Hassan, Goetz Brasche
ICANN (1)3
2019 Dimensionality Reduction for Low-Latency High-Throughput Fraud Detection on Datastreams
abstract
Given the exponential data growth and the recent focus on understanding high-dimensional "in-motion" data, fundamental machine learning tools, such as Principal Component Analysis (PCA), require computation-efficient streaming algorithms that operate near-real-time. Despite the different streaming PCA flavors, there is no algorithm that provably recovers the principal components in the same precision regime as the batch PCA algorithm does, while maintaining low-latency and high-throughput processing. This work, introduces a novel temporal accumulate / retract learning framework for streaming PCA. We consider the accumulate / retract framework implementation of several competitive PCA algorithms with proven theoretical advantages. We benchmark the improved PCA algorithms on real-world streams (i.e. bank transactions fraud detection) and prove their low-latency (millisecond level) and high-throughput (thousands events/second) processing guarantees.
Cristian Axenie, Radu Tudoran, Stefano Bortoli, Mohamad Al Hajj Hassan, Carlos Salort Sánchez, Goetz Brasche
ICMLA3
2019 An Online Incremental Clustering Framework for Real-Time Stream Analytics
abstract
With the evolution of data acquisition methods, our ability to collect real time data has increased. This requires the development of real-time analytics, using the most recent data to generate valuable insights. One example is customer profiling, where we want to identify groups of similar clients who were active recently, and improve the quality of the suggestions. Traditional clustering algorithms perform well on finite datasets, but their execution is often not compatible with real-time requirements, especially for rapid changing trends. In this context, we propose a novel approach for the definition of incremental clustering algorithms to work within real-time constraints, in an online fashion, while preserving accuracy. We show the general applicability of the framework by employing this method to three different clustering algorithms. We compare the experimental results between traditional and online approaches evaluating accuracy and computational cost. The results show that algorithms executed in our framework are comparable to their offline implementation in terms of accuracy and with a high gain in execution time, up to three orders of magnitude on average.
Carlos Salort Sánchez, Radu Tudoran, Mohamad Al Hajj Hassan, Stefano Bortoli, Goetz Brasche, Jan Baumbach, Cristian Axenie
ICMLA4
2018 KerA: Scalable Data Ingestion for Stream Processing
abstract
Big Data applications are increasingly moving from batch-oriented execution models to stream-based models that enable them to extract valuable insights close to real-time. To support this model, an essential part of the streaming processing pipeline is data ingestion, i.e., the collection of data from various sources (sensors, NoSQL stores, filesystems, etc.) and their delivery for processing. Data ingestion needs to support high throughput, low latency and must scale to a large number of both data producers and consumers. Since the overall performance of the whole stream processing pipeline is limited by that of the ingestion phase, it is critical to satisfy these performance goals. However, state-of-art data ingestion systems such as Apache Kafka build on static stream partitioning and offset-based record access, trading performance for design simplicity. In this paper we propose KerA, a data ingestion framework that alleviate the limitations of state-of-art thanks to a dynamic partitioning scheme and to lightweight indexing, thereby improving throughput, latency and scalability. Experimental evaluations show that KerA outperforms Kafka up to 4x for ingestion throughput and up to 5x for the overall stream processing throughput. Furthermore, they show that KerA is capable of delivering data fast enough to saturate the big data engine acting as the consumer.
Ovidiu-Cristian Marcu, Alexandru Costan, Gabriel Antoniu, María S. Pérez 0001, Bogdan Nicolae, Radu Tudoran, Stefano Bortoli
ICDCS7
2018 STARLORD: Sliding Window Temporal Accumulate-Retract Learning for Online Reasoning on Datastreams
abstract
Nowadays, data sources, such as IoT devices, financial markets, and online services, continuously generate large amounts of data. Such data is usually generated at high frequencies and is typically described by non-stationary distributions. Querying these data sources brings new challenges for machine learning algorithms, which now need to be considered from the perspective of an evolving stream and not a static dataset. Under such scenarios, where data flows continuously, the challenge is how to transform the vast amount of data into information and knowledge, and how to adapt to data changes (i.e. drifts) and accumulate experience over time to support online decision-making. In this paper, we introduce STARLORD, a novel incremental computation method and system acting on data streams and capable of achieving low-latency (millisecond level) and high-throughput (thousands events/second/core) when learning from data streams. Moreover, the approach is able to adapt to data drifts and accumulate experience over time, and to use such knowledge to improve future learning and prediction performance, with resource usage guarantees. This is proven by our preliminary experiments where we built-in the framework in an open source stream engine (i.e. Apache Flink).
Cristian Axenie, Radu Tudoran, Stefano Bortoli, Mohamad Al Hajj Hassan, Daniele Foroni, Goetz Brasche
ICMLA3
2017 Towards a unified storage and ingestion architecture for stream processing
abstract
Big Data applications are rapidly moving from a batch-oriented execution model to a streaming execution model in order to extract value from the data in real-time. However, processing live data alone is often not enough: in many cases, such applications need to combine the live data with previously archived data to increase the quality of the extracted insights. Current streaming-oriented runtimes and middlewares are not flexible enough to deal with this trend, as they address ingestion (collection and pre-processing of data streams) and persistent storage (archival of intermediate results) using separate services. This separation often leads to I/O redundancy (e.g., write data twice to disk or transfer data twice over the network) and interference (e.g., I/O bottlenecks when collecting data streams and writing archival data simultaneously). In this position paper, we argue for a unified ingestion and storage architecture for streaming data that addresses the aforementioned challenge. We identify a set of constraints and benefits for such a unified model, while highlighting the important architectural aspects required to implement it in real life. Based on these aspects, we briefly sketch our plan for future work that develops the position defended in this paper.
Ovidiu-Cristian Marcu, Alexandru Costan, Gabriel Antoniu, María S. Pérez 0001, Radu Tudoran, Stefano Bortoli, Bogdan Nicolae
IEEE BigData6
2010 Maskkot - An Entity-centric Annotation Platform
Armando Stellato, Heiko Stoermer, Stefano Bortoli, Noemi Scarpato, Andrea Turbati, Paolo Bouquet, Maria Teresa Pazienza
LREC3