VLDB 2026 Research / reviewers in the wild / expert
Stéphane Rubini
dblp:56/4743
· DBLP profile ↗
23ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0002-3206-0310ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cache Less to Save More: A Cost-Based Distributed Caching Strategy for ICNabstractThe rapid growth of global data traffic has exposed limitations in traditional content delivery architectures. Information-Centric Networking (ICN) addresses these challenges by leveraging in-network caching to enhance scalability, reduce latency, and improve overall performance. However, existing caching strategies either optimize single-node cache management without considering network-wide costs, or address distribution without hardware-aware cost modeling. We propose a unified, cost-aware distributed caching strategy that integrates multi-tier caching at each node with network-wide replication, guided by a comprehensive cost model including resource depreciation, bandwidth, energy, and Service Level Agreement compliance. Our approach minimizes redundant replication on the network while maximizing cache hit rates and reducing latency. Experiments show on average 19.15 %, and up to 45.19 %, cost reduction, 8.11 %, and up to 32.15 %, cache hit ratio increase, and 9.01 %, and up to 27.21 %, latency improvement over other methods, offering a cost-effective solution for next-generation ICN systems. Lydia Ait-Oucheggou, Stéphane Rubini, Abdella Battou, Jalil Boukhobza |
CLUSTER | 2 |
| 2025 | q-AMC: Integrating Quality Management in Mixed Criticality SchedulingabstractModern real-time embedded systems increasingly integrate software with varying criticality levels, which increases the interest in mixed criticality scheduling (MCS). MCS provides runtime adaptation mechanisms when low criticality tasks exceed their allocated execution budgets in order to guarantee the timing constraints of high criticality tasks. Most of the current research on MCS adaptation mechanisms focuses on guaranteeing timing constraints by interrupting and discarding low criticality tasks when their budgets are exceeded. They consider only the temporal dimension, without taking into account the quality of the results obtained. Quality is defined as the accuracy level of the results computed by a task within a given execution time. In this article, we propose an approach to integrate quality in a new task model to establish a relationship between quality and scheduling design. We propose the q-AMC scheduling algorithm to validate our task model. This algorithm integrates quality degradation into the scheduling adaptation mechanism. Simulation-based experiments show that our approach increases the quality up to 44.6% compared to the original AMC approach. Alan Le Boudec, Hai Nam Tran, Stéphane Rubini, Alexandre Skrzyniarz, Frank Singhoff |
ETFA | 3 |
| 2025 | Poster: Reusable Software Components to Prototype and Evaluate Mixed-Criticality Scheduling Policies
Alan Le Boudec, Hai Nam Tran, Stéphane Rubini, Alexandre Skrzyniarz, Frank Singhoff |
RTCSA | 3 |
| 2025 | QM-ARC: QoS-aware Multi-tier Adaptive Cache Replacement Strategy
Lydia Ait-Oucheggou, Stéphane Rubini, Abdella Battou, Jalil Boukhobza |
Future Gener. Comput. Syst. | 2 |
| 2023 | Investigating Multi-Tier and QoS-Aware Caching Based on ARCabstractMemory caching is a common practice to reduce application latencies by buffering relevant data in high speed memory. When the volume of data to cache is too large or a DRAM - based solution too expensive, several technologies such as NVM or high speed SSDs could complement DRAM to form a multi-tier cache. Additionally, most existing policies focus on categorizing the data based on factors like recency and frequency, setting aside the fact that applications/customers have varying Quality-of-Service requirements. This concept is well established in Cloud environment with Service Level Agreement (SLA). In this paper, by extending the Adaptive Replacement Cache (ARC), that uses recency and frequency lists, we propose a QoS-aware Multi-tier Adaptive Replacement Cache (QM-ARC) policy with the ability to take into account data applications/customers priorities through the concept of penalty borrowed from the Cloud. QM-ARC is generic, as it can be applied whatever the number of tiers and can accommodate different penalty functions. Using synthetic and real traces, our solution improved QoS as compared to state-of-the-art work. Lydia Ait-Oucheggou, Stéphane Rubini, Abdella Battou, Jalil Boukhobza |
MASCOTS | 2 |
| 2023 | Work-In-Progress: Could Tensorflow Applications Benefit from a Mixed-Criticality Approach?abstractIn this article, we investigate the interest in applying a mixed-criticality approach to schedule convolutional neural network (CNN) applications on multicore architectures. We deal with software composed of real-time interactive applications and CNNs that have different criticality levels. A classical means to schedule software with various criticality levels is to apply partitioning methods to enforce spatial and temporal isolation, which may be inefficient if application execution times have a high level of variability. In that case, applying a mixed-criticality approach may improve resource usage. We conducted a measurement campaign to assess the variability of CNN execution time and investigate whether this kind of application could benefit from a mixed-criticality approach. The results show that the execution times of the chosen CNN application vary with an average execution time of 109 ms and a worst case of 252 ms. Furthermore, they indicate a potential save of computing resources up to 73 % when applying a mixed-criticality approach instead of partitioning methods. Alan Le Boudec, Frank Singhoff, Hai Nam Tran, Stéphane Rubini, Sébastien Levieux, Alexandre Skrzyniarz |
RTSS | 4 |
| 2023 | Accelerating Random Forest on Memory-Constrained Devices Through Data Storage OptimizationabstractRandom forests is a widely used classification algorithm. It consists of a set of decision trees each of which is a classifier built on the basis of a random subset of the training data-set. In an environment where the memory work-space is low in comparison to the data-set size, when training a decision tree, a large proportion of the execution time is related to I/O operations. These are caused by data blocks transfers between the storage device and the memory work-space (in both directions). Our analysis of random forests training algorithms showed that there are two major issues :(1)Block Under-utilization: data blocks are poorly used when loaded into memory and have to be reloaded multiple times, meaning that the algorithm exhibits a poor spatial locality;(2)Data Over-read: the data-set is supposed to be fully loaded in memory whereas a large proportion of data are not effectively useful when building a decision tree. Our proposed solution is structured to address these two issues. First, we propose to reorganize the data-set in such a way to enhance spatial locality and second, to remove the assumption that the data-set is entirely loaded into memory and access data only when effectively needed. Our experiments show that this method made it possible to reduce random forest building time by 51 to 95% in comparison to a state-of-the-art method. Camélia Slimani, Chun-Feng Wu, Stéphane Rubini, Yuan-Hao Chang 0001, Jalil Boukhobza |
IEEE Trans. Computers | 3 |
| 2022 | Specification of schedulability assumptions to leverage multiprocessor AnalysisabstractIn order to ease the early verification of uniprocessor real-time systems, the tool Cheddar provides a service that guarantees the applicability of a schedulability analysis method for a given architecture model. This verification service uses a catalog of design patterns. In this article, we propose to extend these patterns to multiprocessor architectures. Designing such extension is a challenge because the knowledge of both the software and the hardware architectures are essential to decide on the schedulability of a task set in that context. Indeed, parallel execution of tasks involves hardware resource sharing, that has in turn an effect on the task execution times. Currently, no general method is able to assess the schedulability of a high-performance multicore system with a limited level of pessimism, except if assumptions or usage restrictions are set to simplify the system analysis. So, the research community is developing multiple schedulability tests based on various assumptions which constrain the task models and their execution platforms. In this article, we propose a framework based on Prolog that allows engineers to verify the conditions to apply a test are met. Prolog facts model the software and hardware architecture, and the inference engine checks whether these facts conform to a design pattern associated to a given verification method. The design pattern compliance framework is integrated with the Cheddar tool. Three examples of multiprocessor analyses illustrate the proposal. A scalability analysis shows the tool is able to verify the compliance of architectures composed of 600 tasks and 60 cores, in less than 140s on a desktop computer. Stéphane Rubini, Valérie-Anne Nicolas, Frank Singhoff, Alain Plantec, Hai Nam Tran, Pierre Dissaux |
J. Syst. Archit. | 1 |
| 2021 | ECTM: A network-on-chip communication model to combine task and message schedulability analysisabstractNetwork-on-Chips (NoC) are widely used in industrial applications since they provide communication parallelism and reduce energy consumption. The use of NoC has been recently extended to real-time systems, whose execution has to meet temporal constraints. Communication delays introduced by the network make the scheduling analysis challenging. In this article, we propose a new NoC communication model called ECTM. The main goal of this model is to assess the schedulability of dependent periodic tasks exchanging messages on a NoC. ECTM is a model allowing schedulability analysis of messages and tasks of the NoC. To achieve schedulability, ECTM produces an analysis model by transforming NoC messages to tasks in order to take into account communication delays during the scheduling analysis. Schedulability of the system is assessed using simulation over the feasibility interval with a list scheduling, ECTM supports Store-And-Forward and Wormhole NoC. In this article, we have demonstrated the correctness of the transformations of ECTM. ECTM has been implemented in a real-time scheduling analysis tool called Cheddar and we performed experiments to assess its efficiency. ECTM is more efficient than existing solutions with an improvement of 30% for Store-And-Forward NoCs and up to 100% for Wormhole NoCs, while the proposed model requires a larger computation time about 17% for Store-And-Forward NoCs. Mourad Dridi, Frank Singhoff, Stéphane Rubini, Jean-Philippe Diguet |
J. Syst. Archit. | 3 |
| 2021 | Feasibility interval and sustainable scheduling simulation with CRPD on uniprocessor platform
Hai Nam Tran, Stéphane Rubini, Jalil Boukhobza, Frank Singhoff |
J. Syst. Archit. | 2 |
| 2019 | K -MLIO: Enabling K -Means for Large Data-Sets and Memory Constrained Embedded SystemsabstractMachine Learning (ML) algorithms are increasingly used in embedded systems to perform different tasks such as clustering and pattern recognition. These algorithms are both compute and memory intensive whilst embedded devices offer lower hardware capabilities as compared to traditional ML platforms. K-means clustering is one of the widely used ML algorithms. In the case of large data-sets, our analysis showed that on average, more than 70% of the execution time is spent on I/Os. In this paper, we present a version of K-means that drastically reduces the number of I/Os by spanning the data-set only once as compared to the traditional version that reads it several times according to the number of iterations performed. Our evaluation showed that the proposed strategy reduces the overall execution time on large data-sets by 60% on average while lowering the number I/Os operations by 90% with a comparable precision to the traditional K-means implementation. Camélia Slimani, Stéphane Rubini, Jalil Boukhobza |
MASCOTS | 2 |
| 2019 | Design and Multi-Abstraction-Level Evaluation of a NoC Router for Mixed-Criticality Real-Time SystemsabstractA Mixed Criticality System (MCS) combines real-time software tasks with different criticality levels. In a MCS, the criticality level specifies the level of assurance against system failure. For high-critical flows of messages, it is imperative to meet deadlines; otherwise, the whole system might fail, leading to catastrophic results, like loss of life or serious damage to the environment. In contrast, low-critical flows may tolerate some delays. Furthermore, in MCS, flow performances such as the Worst Case Communication Time (WCCT) may vary depending on the criticality level of the applications. Then execution platforms must provide different operating modes for applications with different levels of criticality. To conclude, in Network-On-Chip (NoC), sharing resources between communication flows can lead to unpredictable latencies and subsequently turns the implementation of MCS in many-core architectures challenging. In this article, we propose and evaluate a new NoC router to support MCS based on an accurate WCCT analysis for high-critical flows. The proposed router, called Double Arbiter and Switching router (DAS), jointly uses Wormhole and Store And Forward communication techniques for low- and high-critical flows, respectively. It ensures that high-critical flows meet their deadlines while maximizing the bandwidth remaining for the low-critical flows. We also propose a new method for high-critical communication time analysis, applied to Store And Forward switching mode with virtual channels. For low-critical flows communication time analysis, we adapt an existing wormhole communication time analysis with share policy to our context. The second contribution of this article is a multi-abstraction-level evaluation of DAS. We evaluate the communication time of flows, the system mode change, the cost, and four properties of DAS. Simulations with a cycle-accurate SystemC NoC simulator show that, with a 15% network use rate, the communication delay of high-critical flows is reduced by 80% while communication delay of low-critical flow is increased by 18% compared to solutions based on routers with multiple virtual channels. For 10% of network interferences, using system mode change, DAS reduces the high-critical communication delays about 66%. We synthesize our router with a 28nm SOI technology and show that the size overhead is limited of 2.5% compared to the solution based on virtual channel router. Finally, we applied model checking verification techniques to automatically prove several DAS properties required by critical systems designers. Mourad Dridi, Stéphane Rubini, Mounir Lallali, Martha Johanna Sepúlveda, Frank Singhoff, Jean-Philippe Diguet |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2018 | Emerging NVM: A Survey on Architectural Integration and Research ChallengesabstractThere has been a surge of interest in Non-Volatile Memory (NVM) in recent years. With many advantages, such as density and power consumption, NVM is carving out a place in the memory hierarchy and may eventually change our view of computer architecture. Many NVMs have emerged, such as Magnetoresistive random access memory (MRAM), Phase Change random access memory (PCM), Resistive random access memory (ReRAM), and Ferroelectric random access memory (FeRAM), each with its own peculiar properties and specific challenges. The scientific community has carried out a substantial amount of work on integrating those technologies in the memory hierarchy. As many companies are announcing the imminent mass production of NVMs, we think that it is time to have a step back and discuss the body of literature related to NVM integration. This article surveys state-of-the-art work on integrating NVM into the memory hierarchy. Specially, we introduce the four types of NVM, namely, MRAM, PCM, ReRAM, and FeRAM, and investigate different ways of integrating them into the memory hierarchy from the horizontal or vertical perspectives. Here, horizontal integration means that the new memory is placed at the same level as an existing one, while vertical integration means that the new memory is interleaved between two existing levels. In addition, we describe challenges and opportunities with each NVM technique. Jalil Boukhobza, Stéphane Rubini, Renhai Chen, Zili Shao |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2017 | DAS: An Efficient NoC Router for Mixed-Criticality Real-Time SystemsabstractMixed-Criticality Systems (MCS) are real-time systems characterized by two or more distinct levels of criticality. In MCS, it is imperative that high-critical flows meet their deadlines while low critical flows can tolerate some delays. Sharing resources between flows in Network-On-Chip (NoC) can lead to different unpredictable latencies and subsequently complicate the implementation of MCS in many-core architectures. This paper proposes a new virtual channel router designed for MCS deployed over NoCs. The first objective of this router is to reduce the worst-case communication latency of high-critical flows. The second aim is to improve the network use rate and reduce the communication latency for low-critical flows. The proposed router, called DAS (Double Arbiter and Switching router), jointly uses Wormhole and Store And Forward techniques for low and high-critical flows respectively. Simulations with a cycle-accurate SystemC NoC simulator show that, with a 15% network use rate, the communication delay of high-critical flows is reduced by 80% while communication delay of low-critical flow is increased by 18% compared to usual solutions based on routers with multiple virtual channels. Mourad Dridi, Stéphane Rubini, Mounir Lallali, Martha Johanna Sepúlveda, Frank Singhoff, Jean-Philippe Diguet |
ICCD | 2 |
| 2017 | Scheduling analysis of tasks constrained by TDMA: Application to software radio protocols
Shuai Li 0007, Frank Singhoff, Stéphane Rubini, Michel Bourdellès |
J. Syst. Archit. | 3 |
| 2016 | A Cost Model for Virtual Machine Storage in Cloud IaaS ContextabstractThis paper proposes a storage system cost model for Infrastructure as a Service (IaaS) Cloud. The proposed cost model takes into account the virtualization environment, the storage system characteristics in addition to energy and QoS related parameters (Service Level Agreement and penalties). We show that those parameters are relevant and allow us to predict an accurate estimation of the overall cost of the IaaS infrastructure. We validate this cost model against real measures and we show less than 10% of error in most cases. Designers and administrators can use this cost model to perform optimization, load balancing, configuration and pricing of the Cloud infrastructure. Hamza Ouarnoughi, Jalil Boukhobza, Frank Singhoff, Stéphane Rubini |
PDP | 4 |
| 2015 | Addressing cache related preemption delay in fixed priority assignmentabstractHandling cache related preemption delay (CRPD) in preemptive scheduling context for real-time embedded systems still stays an open issue despite of its practical importance. Indeed, classical priority assignment algorithms are only optimal when preemption costs are neglected. For example, with Audsley's Optimal Priority Assignment (OPA), as the original algorithm does not take CRPD into account, it fails frequently in identifying the schedulable task sets as it happens that the algorithm qualifies a task set to be schedulable, while it is practically not because of CRPD. In this article, we propose an approach to adapt fixed priority assignment algorithms to real-time embedded systems with cache memory. For such a purpose, we propose three extensions of the original OPA algorithm that have different degrees of pessimism, different complexities, and give different results in terms of schedulable task sets coverage. Exhaustive experimentations were achieved to evaluate the proposed approaches in terms of complexity and efficiency. The result shows that our approach provides a mean to guarantee the schedulability of the real-time embedded system while taking into account CRPD. Hai Nam Tran, Frank Singhoff, Stéphane Rubini, Jalil Boukhobza |
ETFA | 3 |
| 2015 | MaCACH: An adaptive cache-aware hybrid FTL mapping scheme using feedback control for efficient page-mapped space management
Jalil Boukhobza, Pierre Olivier, Stéphane Rubini, Laurent Lemarchand, Yassine Hadjadj-Aoul, Arezki Laga |
J. Syst. Archit. | 3 |
| 2014 | Extending schedulability tests of tree-shaped transactions for TDMA radio protocolsabstractIn this paper, a schedulability test is proposed for tree-shaped transactions with non-immediate tasks. A tree-shaped transaction is a group of precedence dependent tasks, partitioned on different processors, which may release several other tasks upon completion. When there are non-immediate tasks, tasks are not necessarily released immediately upon their predecessor's completion. The schedulability test we propose is based on an existing test that does not handle non-immediate tasks directly. Simulation results show that tighter response time upper-bounds can be accessed when effects of non-immediateness are considered. Our schedulability test is motivated by real industrial TDMA systems developed at Thales, and experimental results show it provides less pessimistic schedulability results compared to current methods used by Thales system engineers. Shuai Li 0007, Frank Singhoff, Stéphane Rubini, Michel Bourdellès |
ETFA | 3 |
| 2014 | Instruction Cache in Hard Real-Time Systems: Modeling and Integration in Scheduling Analysis Tools with AADLabstractCache prediction for real-time systems in a preemptive scheduling context is still an open issue despite its practical importance. In this paper, we propose a modeling approach for taking into account the cache memory in realtime scheduling analysis. The goal is to have a simple but practical implementation to handle the cache memory with a real-time scheduling analyzer. The proposed contribution consists of three main parts: (1) modeling the targeted system with the Architecture Analysis and Design Language (AADL), (2) applying the cache analysis methods in a real time scheduling analysis tool and (3) performing scheduling simulation to access schedulability. For such a purpose, we present an extension of both the scheduling analysis tool Cheddar and of the AADL modeling language in order to integrate the cache modeling and analysis methodology we proposed. Experiments are presented to illustrate our propositions. They provide results on analysis that show examples of the timing impact of task preemption as well as the increase in overall responses time of the task set. This impact is important and the developed tool provides means to precisely assess it. Hai Nam Tran, Frank Singhoff, Stéphane Rubini, Jalil Boukhobza |
EUC | 3 |
| 2013 | CACH-FTL: A Cache-Aware Configurable Hybrid Flash Translation LayerabstractMany hybrid Flash Translation Layer (FTL) schemes have been proposed to leverage the erase-before-write and limited lifetime constraints of flash memories. Those schemes try to approach page mapping performance and flexibility while seeking block mapping memory usage. Furthermore, flash-specific cache systems were designed (1) to maximize lifetime by absorbing some erase operations, and (2) to reveal sequentiality from random write operations. Indeed, random writes represent the Achilles' heel of flash memories. Both cache systems and FTL schemes were designed independently from each other. This paper presents a scalable (in terms of mapping table size) and flexible (in terms of I/O workload support) Cache-Aware Configurable Hybrid (CACH) FTL. CACH-FTL uses a common feature of flash-specific cache systems that is flushing groups of pages from the same block. CACH-FTL partitions the flash memory space into two regions: (1) a data Block Mapped Region (BMR) collecting large groups of pages from the above cache (sequential I/Os), and (2) a small Page Mapped over-provisioning Region (PMR) which purpose is to collect/buffer small groups of pages coming from the cache (random I/Os) before moving them to BMR. CACH-FTL is flexible as it offers many configuration possibilities and can be adapted according to the I/O workload. CACH-FTL approaches the ideal page mapping FTL performance as it gives less than 15% performance difference in most cases. Jalil Boukhobza, Pierre Olivier, Stéphane Rubini |
PDP | 3 |
| 2011 | Modeling and Verification of Memory Architectures with AADL and REALabstractReal-Time Embedded systems must respect a wide range of non-functional properties, including safety, respect of deadlines, power or memory consumption. We note that correct hardware resource dimensioning requires taking into account the impact of the whole software, both the user code and the underlying run time environment. AADL allows one to precisely capture all of them. In this article, we evaluate the AADL modeling to define memory architectures, and then verification rules to assess that the memory is correctly dimensioned. We use the REAL domain-specific language to express memory requirements (such as layout or size) and then validate them on a case-study using the VxWorks real-time kernel. Stéphane Rubini, Frank Singhoff, Jérôme Hugues |
ICECCS | 1 |
| 2005 | Cluster of re-configurable nodes for scanning large genomic banks
Stéphane Guyetant, Mathieu Giraud, Ludovic L'Hours, Steven Derrien, Stéphane Rubini, Dominique Lavenier, Frédéric Raimbault |
Parallel Comput. | 5 |