VLDB 2026 Research / reviewers in the wild / expert
Benjamin Rouxel
dblp:168/9566
· DBLP profile ↗
14ranked-venue papers
5as first author
8since 2021 · last 2024
0000-0002-7178-4768ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1Security and privacy · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | High-Performance Feature Extraction for GPU -Accelerated ORB-SLAMxabstractIn the autonomous vehicles field, localization is a crucial aspect. While the ORB-SLAM algorithm is a recognized solution for these tasks, it poses challenges due to its computational intensity. Although accelerated implementation exists, a bottleneck persists in the Point Filtering phase which relies on the Distribute Octree algorithm that is not suitable for GPU processing. In this paper, we introduce a novel GPU-suitable algorithm designed to enhance the Point Filtering step, surpassing Distribute Octree. We conducted a comprehensive comparison with state-of-the-art CPU and GPU implementations, considering both computational time and trajectory accuracy. Our experimental results, demonstrate significant speed-ups up to 3x compared to previous contributions. Filippo Muzzini, Nicola Capodieci, Roberto Cavicchioli, Benjamin Rouxel |
DATE | 4 |
| 2023 | The TeamPlay Project: Analysing and Optimising Time, Energy, and Security for Cyber-Physical SystemsabstractNon-functional properties, such as energy, time, and security (ETS) are becoming increasingly important in Cyber-Physical Systems (CPS) programming. This article describes TeamPlay, a research project funded under the EU Horizon 2020 programme between January 2018 and June 2021. TeamPlay aimed to provide the system designer with a toolchain for developing embedded applications where ETS properties are first-class citizens, allowing the developer to reflect directly on energy, time and security properties at the source code level. In this paper we give an overview of the TeamPlay methodology, introduce the challenges and solutions of our approach and summarise the results achieved. Overall, applying our TeamPlay methodology led to an improvement of up to 18% performance and 52% energy usage over traditional approaches. Benjamin Rouxel, Christopher Brown 0002, Emad Samuel Malki Ebeid, Kerstin Eder, Heiko Falk, Clemens Grelck, Jesper Holst, Shashank Jadhav, Yoann Marquer, Marcos Martinez de Alejandro, Kris Nikov, Ali Sahafi, Ulrik Pagh Schultz Lundquist, Adam Seewald, Vangelis Vassalos, Simon Wegener, Olivier Zendra |
DATE | 1 |
| 2023 | The IMOCO4.E reference framework for intelligent motion control systemsabstractIntelligent motion control is integral to modern cyber-physical systems. However, smart integration of intelligent motion control with commercial and industrial systems requires domain expertise, industrial ‘know-how’ of the production processes, and resilient adaptation for the various engineering phases. The challenge is amplified with the adoption of advanced digital twin approaches, big data and artificial intelligence in the various industrial domains. This paper proposes the IMOCO4.E reference framework for the smart integration of intelligent motion control with commercial platforms (e.g. from SMEs) and industrial systems. The IMOCO4.E reference framework brings together the architecture, data management, artificial intelligence and digital twin viewpoints from the industrial users of the large-scale ‘Intelligent Motion Control under Industry4.E’ (IMOCO4.E) consortium. The framework envisions a generic platform for designing, developing, and implementing novice and complex motion-controlled industrial systems. Refinements and instantiations of the framework for the IMOCO4.E industrial cases validate the framework’s applicability for various industrial domains throughout the engineering phases and under different constraints imposed on the industrial cases. Sajid Mohamed, Gijs van der Veen, Hans Kuppens, Matias Vierimaa, Tassos Kanellos, Henry Stoutjesdijk, Riccardo Masiero, Kalle Määttä, Jan Wytze van der Weit, Gabriel Ribeiro, Ansgar Bergmann, Davide Colombo, Javier Arenas, Alphonsus Keary, Martin Goubej, Benjamin Rouxel, Pekka Kilpeläinen, Roberts Kadikis, Mikel Armendia, Petr Blaha, Joep Stokkermans, Martin Cech, Arend-Jan Beltman |
ETFA | 16 |
| 2023 | Memory-Aware Latency Prediction Model for Concurrent Kernels in Partitionable GPUs: Simulations and Experiments
Alessio Masola, Nicola Capodieci, Roberto Cavicchioli, Ignacio Sanudo Olmedo, Benjamin Rouxel |
JSSPP | 5 |
| 2023 | Machine Learning Techniques for Understanding and Predicting Memory Interference in CPU-GPU Embedded SystemsabstractNowadays, heterogeneous embedded platforms are extensively used in various low-latency applications, including the automotive industry, real-time IoT systems, and automated factories. These platforms utilize specific components, such as CPUs, GPUs, and neural network accelerators for efficient task processing and to solve specific problems with a lower power consumption compared to more traditional systems. However, since these accelerators share resources such as the global memory, it is crucial to understand how workloads behave under high computational loads to determine how parallel computational engines on modern platforms can interfere and adversely affect the system's predictability and performance. One area that remains unclear is the interference effect on shared memory resources between the CPU and GPU: more specifically, the latency degradation experienced by GPU kernels when memory-intensive CPU applications run concurrently. In this work, we first analyze the metrics that characterize the behavior of different kernels under various board conditions caused by CPU memory-intensive workloads on a Nvidia Jetson Xavier. Then, we exploit various machine learning methodologies aiming to estimate the latency degradation of kernels based on their metrics. As a result of this, we are able to identify the metrics that could potentially have the most significant impact when predicting the kernels completion latency degradation. Alessio Masola, Nicola Capodieci, Benjamin Rouxel, Giorgia Franchini, Roberto Cavicchioli |
RTCSA | 3 |
| 2023 | Brief Announcement: Optimized GPU-accelerated Feature Extraction for ORB-SLAM SystemsabstractReducing the execution time of ORB-SLAM algorithm is a crucial aspect of autonomous vehicles since it is computationally intensive for embedded boards. We propose a parallel GPU-based implementation, able to run on embedded boards, of the Tracking part of the ORB-SLAM2/3 algorithm. Our implementation is not simply a GPU port of the tracking phase. Instead, we propose a novel method to accelerate image Pyramid construction on GPUs. Comparison against state-of-the-art CPU and GPU implementations, considering both computational time and trajectory errors shows improvement on execution time in well-known datasets, such as KITTI and EuRoC. Filippo Muzzini, Nicola Capodieci, Roberto Cavicchioli, Benjamin Rouxel |
SPAA | 4 |
| 2023 | Time-sensitive autonomous architectures
Donato Ferraro, Luca Palazzi, Federico Gavioli, Michele Guzzinati, Andrea Bernardi, Benjamin Rouxel, Paolo Burgio, Marco Solieri |
Real Time Syst. | 6 |
| 2021 | YASMIN: a real-time middleware for COTS heterogeneous platformsabstractCommercial-off-the-shelf (COTS) heterogeneous platforms provide immense computational power, but are difficult to program and to correctly use when real-time requirements come into play: A sound configuration of the operating system scheduler is needed, and a suitable mapping of tasks to computing units must be determined. Flawed designs lead to sub-optimal system configurations and, thus, to wasted resources or even to deadline misses and system failures. Benjamin Rouxel, Sebastian Altmeyer, Clemens Grelck |
Middleware | 1 |
| 2020 | Q-learning for Statically Scheduling DAGsabstractData parallel frameworks (e.g. Hive, Spark or Tez) can be used to execute complex data analyses consisting of many dependent tasks represented by a Directed Acylical Graph (DAG). Minimising the job completion time (i.e. makespan) is still an open problem for large graphs.We propose a novel deep Q-learning (DQN) approach to statically scheduling DAGs and minimising the makespan. Our approach learns to schedule DAGs from scratch instead of learning how to imitate some heuristic. We show that our current approach learns fast and steadily. Furthermore, our approach can schedule DAGs almost 15 times faster than a Forward List Scheduling (FLS) heuristic. Julius Roeder, Benjamin Rouxel, Clemens Grelck |
IEEE BigData | 2 |
| 2020 | Towards Energy-, Time- and Security-Aware Multi-core Coordination
Julius Roeder, Benjamin Rouxel, Sebastian Altmeyer, Clemens Grelck |
COORDINATION | 2 |
| 2020 | PReGO: a generative methodology for satisfying real-time requirements on COTS-based systems: definition and experience reportabstractSatisfying real-time requirements in cyber-physical systems is challenging as timing behaviour depends on the application software, the embedded hardware, as well as the execution environment. This challenge is exacerbated as real-world, industrial systems often use unpredictable hardware and software libraries or operating systems with timing hazards and proprietary device drivers. All these issues limit or entirely prevent the application of established real-time analysis techniques. Benjamin Rouxel, Ulrik Pagh Schultz Lundquist, Benny Akesson, Jesper Holst, Ole Jørgensen 0001, Clemens Grelck |
GPCE | 1 |
| 2019 | Hiding Communication Delays in Contention-Free Execution for SPM-Based Multi-Core ArchitecturesabstractMulti-core systems using ScratchPad Memories (SPMs) are attractive architectures for executing time-critical embedded applications, because they provide both predictability and performance. In this paper, we propose a scheduling technique that jointly selects SPM contents off-line, in such a way that the cost of SPM loading/unloading is hidden. Communications are fragmented to augment hiding possibilities. Experimental results show the effectiveness of the proposed technique on streaming applications and synthetic task-graphs. The overlapping of communications with computations allows the length of generated schedules to be reduced by 4% on average on streaming applications, with a maximum of 16%, and by 8% on average for synthetic task graphs. We further show on a case study that generated schedules can be implemented with low overhead on a predictable multi-core architecture (Kalray MPPA). Benjamin Rouxel, Stefanos Skalistis, Steven Derrien, Isabelle Puaut |
ECRTS | 1 |
| 2017 | Tightening Contention Delays While Scheduling Parallel Applications on Multi-core ArchitecturesabstractMulti-core systems are increasingly interesting candidates for executing parallel real-time applications, in avionic, space or automotive industries, as they provide both computing capabilities and power efficiency. However, ensuring that timing constraints are met on such platforms is challenging, because some hardware resources are shared between cores. Assuming worst-case contentions when analyzing the schedulability of applications may result in systems mistakenly declared unschedulable, although the worst-case level of contentions can never occur in practice. In this paper, we present two contention-aware scheduling strategies that produce a time-triggered schedule of the application’s tasks. Based on knowledge of the application’s structure, our scheduling strategies precisely estimate the effective contentions, in order to minimize the overall makespan of the schedule. An Integer Linear Programming (ILP) solution of the scheduling problem is presented, as well as a heuristic solution that generates schedules very close to ones of the ILP (5% longer on average), with a much lower time complexity. Our heuristic improves by 19% the overall makespan of the resulting schedules compared to a worst-case contention baseline. Benjamin Rouxel, Steven Derrien, Isabelle Puaut |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2015 | CoDisasm: Medium Scale Concatic Disassembly of Self-Modifying Binaries with Overlapping InstructionsabstractFighting malware involves analyzing large numbers of suspicious binary files. In this context, disassembly is a crucial task in malware analysis and reverse engineering. It involves the recovery of assembly instructions from binary machine code. Correct disassembly of binaries is necessary to produce a higher level representation of the code and thus allow the analysis to develop high-level understanding of its behavior and purpose. Nonetheless, it can be problematic in the case of malicious code, as malware writers often employ techniques to thwart correct disassembly by standard tools. In this paper, we focus on the disassembly of x86 self-modifying binaries with overlapping instructions. Current state-of-the-art disassemblers fail to interpret these two common forms of obfuscation, causing an incorrect disassembly of large parts of the input. We introduce a novel disassembly method, called concatic disassembly, that combines CONCrete path execution with stATIC disassembly. We have developed a standalone disassembler called CoDisasm that implements this approach. Guillaume Bonfante, José M. Fernandez 0001, Jean-Yves Marion, Benjamin Rouxel, Fabrice Sabatier, Aurélien Thierry |
CCS | 4 |