EDBT 2026 Demo / reviewers in the wild / expert
Jalil Boukhobza
dblp:94/2593
· DBLP profile ↗
57ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0002-2194-4006ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 39 · 4 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | I/O patterns modeling of HPC applications with call stacks for predictive prefetchabstractModern high-performance computing (HPC) storage systems use heterogeneous storage technologies organized in tiers to find a compromise between capacity, performance, and cost. In these systems, prefetching is a common technique used to move the right data at the right moment from a slow to a fast tier to improve overall performance while using the costly high-performance tier only when needed. Effective prefetching requires precise knowledge of the application I/O patterns. This knowledge can be extracted through the source code, I/O tracing tools or I/O functions call stacks. State-of-the-art solutions based on the latter approach mainly focus on applications with regular I/O profiles to avoid scalability issues due to the grammar-based techniques used. In this paper, we present an approach based on I/O call stacks that models POSIX and STDIO I/O patterns for both regular and irregular applications, thanks to the use of directed graphs. We present different models usable for prefetching. Our models were used to predict the next I/O call stack on five real HPC applications with a prediction accuracy of up to 98%. Compared to the state-of-the-art Omnisc’IO, they incurred up to 120x lower model overhead (334 ns vs. 45 μ s on LAMMPS) and had a model size 10x to 15x smaller (463 B vs. 7 kB on LQCD). Louis-Marie Nicolas, Salim Mimouni, Philippe Couvée, Jalil Boukhobza |
Future Gener. Comput. Syst. | 4 |
| 2025 | Design and Optimization for AI/ML Acceleration on Resource-constrained and Edge SystemsabstractThe rapid advancement of AI (from foundational machine learning to Large Language Models) and edge computing has placed unprecedented demands on computation, memory, and storage on resource-constrained edge devices. As AI models scale, the ability to efficiently manage computing resources, utilize memory and storage, and reduce energy consumption has become critical. This paper introduces contributions on 4 topics related to deploying AI on resource-constrained edge devices: 1) unlocking training of foundational machine learning algorithms on the edge, 2) exploring hardware-aware DNN architecture and mapping co-optimization for inference on heterogeneous systems, 3) scaling RAG by leveraging advanced memory, storage, and energy-efficient designs, and 4) investigating cost-effective and high-performance large-scale graph processing. Jalil Boukhobza, Alessio Burrello, Yuan-Hao Chang 0001, Yawei Li 0001, Daniele Jahier Pagliari, Chun-Feng Wu, Ming-Chang Yang, Tsun-Yu Yang |
CASES | 1 |
| 2025 | Cache Less to Save More: A Cost-Based Distributed Caching Strategy for ICNabstractThe rapid growth of global data traffic has exposed limitations in traditional content delivery architectures. Information-Centric Networking (ICN) addresses these challenges by leveraging in-network caching to enhance scalability, reduce latency, and improve overall performance. However, existing caching strategies either optimize single-node cache management without considering network-wide costs, or address distribution without hardware-aware cost modeling. We propose a unified, cost-aware distributed caching strategy that integrates multi-tier caching at each node with network-wide replication, guided by a comprehensive cost model including resource depreciation, bandwidth, energy, and Service Level Agreement compliance. Our approach minimizes redundant replication on the network while maximizing cache hit rates and reducing latency. Experiments show on average 19.15 %, and up to 45.19 %, cost reduction, 8.11 %, and up to 32.15 %, cache hit ratio increase, and 9.01 %, and up to 27.21 %, latency improvement over other methods, offering a cost-effective solution for next-generation ICN systems. Lydia Ait-Oucheggou, Stéphane Rubini, Abdella Battou, Jalil Boukhobza |
CLUSTER | 4 |
| 2025 | Practicalizing Tree-Based Model Acceleration with CAM through Model Pruning and Data Placement OptimizationabstractTree-based model remains state-of-the-art for many tasks involving tabular data. While these models are favored in resource-constrained environments, the inherent characteristics result in inefficiency during inference, posing significant challenges for conventional accelerators. Recent research has achieved unprecedented acceleration with content-addressable memory (CAM), yet at the cost of overwhelming memory consumption with low utilization, which is impractical for numerous real-world applications. This work addresses these issues by introducing an end-to-end framework RETENTION. RETENTION incorporates (1) a pruning algorithm to minimize model complexity under a user-specified accuracy loss tolerance, and (2) two data placement strategies to enhance memory utilization and further reduce capacity requirement. Experiment results show that space efficiency can be improved from 4.35× to 207.12× with less than 3% accuracy loss. Yi-Chun Liao 0001, Chieh-Lin Tsai, Yuan-Hao Chang 0001, Camélia Slimani, Jalil Boukhobza, Tei-Wei Kuo |
CODES+ISSS | 5 |
| 2025 | DisPEED: Distributing Packet flow analyses in a swarm of heterogeneous EmbEddeD platformsabstractSecurity is a major challenge in swarm of drones. Network intrusion detection systems (IDS) are deployed to analyze and detect suspicious packet flows. Traditionally, they are implemented independently on each drone. However, due to heterogeneity and resource limitations of drones, IDS algorithms can fall short in satisfying Quality of Service (Qo$S$) metrics, such as latency and accuracy. We argue that a drone can make profit from the swarm by delegating part of the analysis of their packet flows to neighbor drones that have more processing power to enforce security. In this paper, we propose two solving methods to distribute the packet flows to analyze among drones in a way to ensure that it is processed with a minimum communication overhead to limit the attack surface, while ensuring Qo$S$metrics imposed by the drone mission. First, we propose a formulation of the distribution problem using both an Integer Linear Programming (ILP) and a Maximum-Flow Minimum-Cost (MFMC). Furthermore, we propose two specific solving methods for the distribution problem: (1) a Greedy Heuristic (GH), a non-exact solving method, but with small time overhead, and (2) an Adapted Edmonds-Karp (AEK) algorithm, an exact method, but with a higher time overhead. GH proved to be a very fast solution (up to more than 2000x faster than ILP with Branch and Bound), while AEK solution proved to find the exact solution even when the problem is very difficult. Louis Morge-Rollet, Camélia Slimani, Laurent Lemarchand, Frédéric Le Roy, David Espes, Jalil Boukhobza |
DATE | 6 |
| 2025 | Investigating the Use of File Advisory Hints on Lustre and GPFSabstractInternational audience Louis-Marie Nicolas, Salim Mimouni, Philippe Couvée, Jalil Boukhobza |
SSDBM | 4 |
| 2025 | QM-ARC: QoS-aware Multi-tier Adaptive Cache Replacement Strategy
Lydia Ait-Oucheggou, Stéphane Rubini, Abdella Battou, Jalil Boukhobza |
Future Gener. Comput. Syst. | 4 |
| 2025 | A study on characterizing energy, latency and security for Intrusion Detection Systems on heterogeneous embedded platforms
Camélia Slimani, Louis Morge-Rollet, Laurent Lemarchand, David Espes, Frédéric Le Roy, Jalil Boukhobza |
Future Gener. Comput. Syst. | 6 |
| 2025 | A Survey on Flash-Memory Storage Systems: A Host-Side PerspectiveabstractNAND flash memory has become the dominant storage media choice in a vast majority of application scenarios. Compared to mechanical hard disks, flash offers better access performance, energy efficiency, and shock resistance. However, the unique hardware peculiarities of this technology require dedicated facilities to manage the flash space and data. The implementation of flash management facilities has alternatively been realized either at the device or host computer level. Managing flash on the device side eases integration/compatibility and increases performance in certain scenarios. However, the limited computing resources inherent to devices and the lack of higher-level file system/application information make these solutions suboptimal in many situations. Managing flash on the host allows leveraging its abundant resources, and host-side knowledge such as data access patterns can be exploited to optimize flash management, at the cost of increased host-side complexity. The pros and cons of each approach also led to the appearance of hybrid, cross-layer solutions, enabling the collaboration of different layers of the storage stack. Recently, the pressure on modern storage systems requires that an increasing amount of flash management responsibilities is offloaded to the host, and the development of application-specific cross-layer solutions: In that context, it is crucial to review these developments. In this article, we make a comprehensive survey of the host-side management technologies of flash memory, application-/system-level flash-friendly designs, and emergent applications based on flash memory. Jalil Boukhobza, Pierre Olivier, Wen Sheng Lim, Liang-Chi Chen, Yun-Shan Hsieh, Shin-Ting Wu, Chien-Chung Ho, Po-Chun Huang, Yuan-Hao Chang 0001 |
ACM Trans. Storage | 1 |
| 2024 | HeROcache: Storage-Aware Scheduling in Heterogeneous Serverless Edge - The Case of IDSabstractIntrusion Detection Systems (IDS) are time-sensitive applications that aim to classify potentially malicious network traffic. IDSs are part of a class of applications that rely on short-lived functions that can be run reactively and, as such, could be deployed on edge resources, to offload processing from energy-constrained battery-backed devices. The serverless service model could fit the needs of such applications, given that the platform allows adequate levels of Quality of Service (QoS) for a variety of users, since the criticality of IDS applications depends on several parameters. Deploying serverless functions on unreserved edge resources requires to pay particular attention to (1) initialization delays that could be significant on low resources platforms, (2) inter-function communication between edge nodes, and (3) heterogeneous devices. In this paper, we propose both a storage-aware allocation and scheduling policy that seek to minimize task placement costs for service providers on edge devices while optimizing QoS for IDS users. To do so, we propose a caching and consolidation strategy that minimizes cold starts and inter-function communication delays while satisfying QoS by leveraging heterogeneous edge resources. We evaluated our platform in a simulation environment using characterization data from real-world IDS tasks and execution platforms and compared it with a vanilla Knative orchestrator and a storage-agnostic policy. Our strategy achieves 18% fewer QoS penalties while consolidating applications across 80% fewer edge nodes. Vincent Lannurien, Camélia Slimani, Laurent d'Orazio, Olivier Barais, Stéphane Paquelet, Jalil Boukhobza |
CCGrid | 6 |
| 2024 | IDS-DEEP: a strategy for selecting the best IDS for Drones with heterogeneous EmbEdded PlatformsabstractDrone swarms are increasingly being used to perform critical missions, such as inspection of ports and industrial installations. Each drone can embed heterogeneous execution platforms to successfully perform various computing tasks. As security threats may disrupt the progression of the drone mission, network intrusion detection systems (IDSs) are used. They analyze network traffic to detect malicious behaviors, but generally rely on resource-hungry machine learning models. To adapt to the dynamic nature of the mission, it is necessary to embed several IDS implementations leveraging heterogeneous computing resources of the drone and presenting a trade-off between security, throughput, and energy consumption. To address this issue, we propose, in this paper, an end-to-end flow composed of an offline phase to choose the IDS implementations to embed on the drone platform and an online phase to select the best implementation online considering the mission conditions at a given time. We devised a MILP formulation for the offline phase that proved to provide a 89.41% better Inverted Generational Distance (IGD) than a random choice. For the online phase, we investigated several solutions and designed a novel optimized strategy that proved to be around 16.76 times faster than TOPSIS while having comparable QoS metrics. Louis Morge-Rollet, Camélia Slimani, Laurent Lemarchand, Frédéric Le Roy, David Espes, Jalil Boukhobza |
SBAC-PAD | 6 |
| 2023 | HeROfake: Heterogeneous Resources Orchestration in a Serverless Cloud - An Application to Deepfake DetectionabstractServerless is a trending service model for cloud computing. It shifts a lot of the complexity from customers to service providers. However, current serverless platforms mostly consider the provider's infrastructure as homogeneous, as well as the users' requests. This limits possibilities for the provider to leverage heterogeneity in their infrastructure to improve function response time and reduce energy consumption. We propose a heterogeneity-aware serverless orchestrator for private clouds that consists of two components: the autoscaler allocates heterogeneous hardware resources (CPUs, GPUs, FPGAs) for function replicas, while the scheduler maps function executions to these replicas. Our objective is to guarantee function response time, while enabling the provider to reduce resource usage and energy consumption. This work considers a case study for a deepfake detection application relying on CNN inference. We devised a simulation environment that implements our model and a baseline Knative orchestrator, and evaluated both policies with regard to consolidation of tasks, energy consumption and SLA penalties. Experimental results show that our platform yields substantial gains for all those metrics, with an average of 35% less energy consumed for function executions while consolidating tasks on less than 40% of the infrastructure's nodes, and more than 60% less SLA violations. Vincent Lannurien, Laurent d'Orazio, Olivier Barais, Esther Bernard, Olivier Weppe, Laurent Beaulieu, Amine Kacete, Stéphane Paquelet, Jalil Boukhobza |
CCGrid | 9 |
| 2023 | Characterizing Intrusion Detection Systems On Heterogeneous Embedded PlatformsabstractSwarms of drones are more and more used for critical missions and need to be protected against malicious users. Intrusion Detection Systems (IDS) are used to analyze network traffic in order to detect possible threats. Modern IDSs rely on machine learning models for such a sake. Because of the absence of central management in swarms of drones, IDSs constitute a good second-line protective measure. Investigating the execution of IDS (resource-hungry) algorithms on drone (resource-constrained) devices is crucial when it comes to optimizing energy, response time, memory footprint and algorithm precision. In addition, embedded platforms used in drones often incorporate heterogeneous computing platforms on which IDSs could be executed. In this paper, we present a methodology and results about characterizing the execution of different IDS models on various platform (CPUs, GPUs). In effect, as swarm of drones operate in different mission contexts (e.g. criticity level) and states (e.g. energy budget, memory footprint), it is important to explore which IDS model to run on which platforms for a given mission in a given context. For this sake, we evaluated several metrics on different platforms: energy and resource consumption, accuracy for malicious traffic detection and response time. The models tested (RF, CNN, DNN) have shown different performance according to the measured metrics and the chosen platform and proved to be relevant in different mission states. Camélia Slimani, Louis Morge-Rollet, Laurent Lemarchand, Frédéric Le Roy, David Espes, Jalil Boukhobza |
DSD | 6 |
| 2023 | Training K-Means on Embedded Devices: A Deadline-Aware and Energy Efficient DesignabstractWith the surge in data production, Machine Learning techniques are now commonly used to build intelligent models. Traditionally, powerful platforms process data collected from endpoint devices. However, to address security threats and minimize communication traffic, models can be learned near endpoint devices, despite their resource shortage. K-means clustering is among the most common machine learning tasks used for embedded applications. Because the system is running on scarce resources, the learning process needs to obey a certain time limit. Even if current implementations of K-means have been optimized for embedded devices, they do not consider running within a predefined time budget. In this paper, we propose a deadline-aware and energy-efficient version of K-means called Embedded K-means (EK-means)11The source code is available on https://github.com/HafsaKaraAchira/EK-means-Embedded-K-means-.git, that relies on two main ideas: (1) smartly select the right subset of data to train on to meet the deadline at the expense of the smallest clustering error possible; (2) by dropping part of the data, slack times are identified and exploited opportunistically to apply Dynamic Voltage and Frequency Scaling techniques (DVFS) so as to decrease the energy consumption of the learning task. EK-means has been built on top of an I/O optimized version of K - means for embedded devices to maintain a low I/O proportion regardless of memory constraints. EK-means allows to cluster data while meeting more than 98% of the deadlines with a loss of 1.43 % of clustering quality, and an energy reduction of up to 84.26%. Hafsa Kara Achira, Camélia Slimani, Jalil Boukhobza |
MASCOTS | 3 |
| 2023 | Investigating Multi-Tier and QoS-Aware Caching Based on ARCabstractMemory caching is a common practice to reduce application latencies by buffering relevant data in high speed memory. When the volume of data to cache is too large or a DRAM - based solution too expensive, several technologies such as NVM or high speed SSDs could complement DRAM to form a multi-tier cache. Additionally, most existing policies focus on categorizing the data based on factors like recency and frequency, setting aside the fact that applications/customers have varying Quality-of-Service requirements. This concept is well established in Cloud environment with Service Level Agreement (SLA). In this paper, by extending the Adaptive Replacement Cache (ARC), that uses recency and frequency lists, we propose a QoS-aware Multi-tier Adaptive Replacement Cache (QM-ARC) policy with the ability to take into account data applications/customers priorities through the concept of penalty borrowed from the Cloud. QM-ARC is generic, as it can be applied whatever the number of tiers and can accommodate different penalty functions. Using synthetic and real traces, our solution improved QoS as compared to state-of-the-art work. Lydia Ait-Oucheggou, Stéphane Rubini, Abdella Battou, Jalil Boukhobza |
MASCOTS | 4 |
| 2023 | Accelerating Random Forest on Memory-Constrained Devices Through Data Storage OptimizationabstractRandom forests is a widely used classification algorithm. It consists of a set of decision trees each of which is a classifier built on the basis of a random subset of the training data-set. In an environment where the memory work-space is low in comparison to the data-set size, when training a decision tree, a large proportion of the execution time is related to I/O operations. These are caused by data blocks transfers between the storage device and the memory work-space (in both directions). Our analysis of random forests training algorithms showed that there are two major issues :(1)Block Under-utilization: data blocks are poorly used when loaded into memory and have to be reloaded multiple times, meaning that the algorithm exhibits a poor spatial locality;(2)Data Over-read: the data-set is supposed to be fully loaded in memory whereas a large proportion of data are not effectively useful when building a decision tree. Our proposed solution is structured to address these two issues. First, we propose to reorganize the data-set in such a way to enhance spatial locality and second, to remove the assumption that the data-set is entirely loaded into memory and access data only when effectively needed. Our experiments show that this method made it possible to reduce random forest building time by 51 to 95% in comparison to a state-of-the-art method. Camélia Slimani, Chun-Feng Wu, Stéphane Rubini, Yuan-Hao Chang 0001, Jalil Boukhobza |
IEEE Trans. Computers | 5 |
| 2022 | The Lannion report on Big Data and Security Monitoring ResearchabstractDuring the last decade, big data management has attracted increasing interest from both the industrial and academic communities. In parallel, Cyber Security has become mandatory due to various and more intensive threats. In June 2022, a group of researchers has met to reflect on their community’s impacts on current research challenges. In particular, they have considered four dimensions: (1) dedicated systems being data processing and analytic platforms or time series management systems; (2) graphs analytics and distributed computation; (3) privacy; and (4) new hardware. Laurent d'Orazio, Jalil Boukhobza, Omer F. Rana, Juba Agoun, Le Gruenwald, Hervé Rannou, Elisa Bertino, Mohand-Said Hacid, Taofik Saïdi, Georges Bossert, Dimitri Tombroff, Makoto Onizuka |
IEEE Big Data | 2 |
| 2022 | When IoT Data Meets Streaming in the FogabstractIoT and video streaming are the main driving applications for digital data generation today. The traditional way of storing and processing data in the Cloud cannot satisfy many latency critical applications. This is why Fog computing emerged as a continuum infrastructure from the Cloud to end-user devices. Misplacing data in such an infrastructure results in high latency, and consequently increases the penalty for Internet Service Providers (ISPs) incurred by violating the service level agreement (SLA). In past studies, two issues have been investigated separately: the IoT data placement and the streaming cache placement. However, both placements rely on the same Fog distributed storage system. In this paper, we address those issues in a unique model with the aim to minimize the penalty for ISPs incurred by the SLA violation and maximize storage resources usage. We subdivided each Fog node storage space into a storage part and a cache part. First, our model consists in placing IoT data in the storage part of Fog nodes, and then placing streaming data in the cache part of these nodes. The novelty of our model is the flexibility it offers for managing the cache volume, which can, adaptively, spill on the free part dedicated to IoT data. Experiments show that using our model makes it possible to reduce the streaming data penalty of the ISP’s SLA violation by more than 47% on average. Lydia Ait-Oucheggou, Mohammed Islam Naas, Yassine Hadjadj-Aoul, Jalil Boukhobza |
ICFEC | 4 |
| 2022 | RISCLESS: A Reinforcement Learning Strategy to Guarantee SLA on Cloud Ephemeral and Stable ResourcesabstractIn this paper, we propose RISCLESS, a Reinforcement Learning strategy to exploit unused Cloud resources. Our approach consists in using a small proportion of stable on-demand resources alongside the ephemeral ones in order to guarantee customers SLA and reduce the overall costs. The approach decides when and how much stable resources to allocate in order to fulfill customers’ demands. RISCLESS improved the Cloud Providers (CPs)’ profits by an average of 15.9% compared to past strategies. It also reduced the SLA violation time by 36.7% while increasing the amount of used ephemeral resources by 19.5%. SidAhmed Yalles, Mohamed Handaoui, Jean-Emile Dartois, Olivier Barais, Laurent d'Orazio, Jalil Boukhobza |
PDP | 6 |
| 2022 | Introduction to the Special Issue on Memory and Storage Systems for Embedded and IoT ApplicationsabstractInternational audience Yuan-Hao Chang 0001, Jalil Boukhobza, Song Han 0002 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2022 | Introduction to the Special Issue on Memory and Storage Systems for Embedded and IoT Applications: Part 2abstractNo abstract available. Yuan-Hao Chang 0001, Jalil Boukhobza, Song Han 0002 |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2021 | IoT Data Replication and Consistency Management in Fog Computing
Mohammed Islam Naas, Laurent Lemarchand, Philippe Raipin Parvédy, Jalil Boukhobza |
J. Grid Comput. | 4 |
| 2021 | HEVC hardware vs software decoding: An objective energy consumption analysis and comparison
Mohammed Bey Ahmed Khernache, Yahia Benmoussa, Jalil Boukhobza, Daniel Ménard |
J. Syst. Archit. | 3 |
| 2021 | Feasibility interval and sustainable scheduling simulation with CRPD on uniprocessor platform
Hai Nam Tran, Stéphane Rubini, Jalil Boukhobza, Frank Singhoff |
J. Syst. Archit. | 3 |
| 2021 | Investigating Machine Learning Algorithms for Modeling SSD I/O Performance for Container-Based VirtualizationabstractOne of the cornerstones of the cloud provider business is to reduce hardware resources cost by maximizing their utilization. This is done through smartly sharing processor, memory, network and storage, while fully satisfying SLOs negotiated with customers. For the storage part, while SSDs are increasingly deployed in data centers mainly for their performance and energy efficiency, their internal mechanisms may cause a dramatic SLO violation. In effect, we measured that I/O interference may induce a 10x performance drop. We are building a framework based on autonomic computing which aims to achieve intelligent container placement on storage systems by preventing bad I/O interference scenarios. One prerequisite to such a framework is to design SSD performance models that take into account interactions between running processes/containers, the operating system and the SSD. These interactions are complex. In this paper, we investigate the use of machine learning for building such models in a container based Cloud environment. We have investigated five popular machine learning algorithms along with six different I/O intensive applications and benchmarks. We analyzed the prediction accuracy, the learning curve, the feature importance and the training time of the tested algorithms on four different SSD models. Beyond describing modeling component of our framework, this paper aims to provide insights for cloud providers to implement SLO compliant container placement algorithms on SSDs. Our machine learning-based framework succeeded in modeling I/O interference with a median Normalized Root-Mean-Square Error (NRMSE) of 2.5 percent. Jean-Emile Dartois, Jalil Boukhobza, Anas Knefati, Olivier Barais |
IEEE Trans. Cloud Comput. | 2 |
| 2021 | Multi-objective Optimization of Data Placement in a Storage-as-a-Service Federated CloudabstractCloud federation enables service providers to collaborate to provide better services to customers. For cloud storage services, optimizing customer object placement for a member of a federation is a real challenge. Storage, migration, and latency costs need to be considered. These costs are contradictory in some cases. In this article, we modeled object placement as a multi-objective optimization problem. The proposed model takes into account parameters related to the local infrastructure, the federated environment, customer workloads, and their SLAs. For resolving this problem, we propose CDP-NSGAII IR , a Constraint Data Placement matheuristic based on NSGAII with Injection and Repair functions. The injection function aims to enhance the solutions’ quality. It consists to calculate some solutions using an exact method then inject them into the initial population of NSGAII. The repair function ensures that the solutions obey the problem constraints and so prevents from exploring large sets of unfeasible solutions. It reduces drastically the execution time of NSGAII. Experimental results show that the injection function improves the HV of NSGAII and the exact method by up to 94% and 60%, respectively, while the repair function reduces the execution time by an average of 68%. Amina Chikhaoui, Laurent Lemarchand, Kamel Boukhalfa, Jalil Boukhobza |
ACM Trans. Storage | 4 |
| 2020 | Salamander: a Holistic Scheduling of MapReduce Jobs on Ephemeral Cloud ResourcesabstractMost cloud data centers are over-provisioned and underutilized, primarily to handle peak loads and sudden failures. This has motivated many researchers to reclaim the unused resources, which are by nature ephemeral, to run data-intensive applications at a lower cost. Hadoop MapReduce is one of those applications. However, it was designed on the assumption that resources are available as long as users pay for the service. In order to make it possible for Hadoop to run on unused (ephemeral) resources, we have designed a heterogeneity and volatility-aware holistic scheduler consisting of three different components: (1) A MapReduce task and job scheduler that relies on a global vision of resource utilization predictions, (2) a scheduler-based data placement strategy that improves the data locality, and (3) a reactive QoS controller that ensures customers' service-level agreement (SLA) and minimizes interference between co-located workloads. Our framework makes it possible to take advantage of ephemeral resources efficiently. Indeed, for a given set of jobs, it reduces the overall execution time by up to 47.6% and an average of 18.7% as compared to state-of-the-art strategies. Mohamed Handaoui, Jean-Emile Dartois, Laurent Lemarchand, Jalil Boukhobza |
CCGRID | 4 |
| 2020 | ReLeaSER: A Reinforcement Learning Strategy for Optimizing Utilization Of Ephemeral Cloud ResourcesabstractCloud data center capacities are over-provisioned to handle demand peaks and hardware failures which leads to low resources' utilization. One way to improve resource utilization and thus reduce the total cost of ownership is to offer unused resources (referred to as ephemeral resources) at a lower price. However, reselling resources needs to meet the expectations of its customers in terms of Quality of Service. The goal is so to maximize the amount of reclaimed resources while avoiding SLA penalties. To achieve that, cloud providers have to estimate their future utilization to provide availability guarantees. The prediction should consider a safety margin for resources to react to unpredictable workloads. The challenge is to find the safety margin that provides the best trade-off between the amount of resources to reclaim and the risk of SLA violations. Most state-of-the-art solutions consider a fixed safety margin for all types of metrics (e.g., CPU, RAM). However, a unique fixed margin does not consider various workloads variations over time which may lead to SLA violations or/and poor utilization. In order to tackle these challenges, we propose ReLeaSER, a Reinforcement Learning strategy for optimizing the ephemeral resources' utilization in the cloud. ReLeaSER dynamically tunes the safety margin at the host-level for each resource metric. The strategy learns from past prediction errors (that caused SLA violations). Our solution reduces significantly the SLA violation penalties on average by 2.7× and up to 3.4×. It also improves considerably the CPs' potential savings by 27.6% on average and up to 43.6%. Mohamed Handaoui, Jean-Emile Dartois, Jalil Boukhobza, Olivier Barais, Laurent d'Orazio |
CloudCom | 3 |
| 2020 | Editorial
André-Luc Beylot, Jalil Boukhobza |
Perform. Evaluation | 2 |
| 2019 | Cuckoo: Opportunistic MapReduce on Ephemeral and Heterogeneous Cloud ResourcesabstractCloud infrastructures are generally over-provisioned for handling load peaks and node failures. However, the drawback of this approach is that a large portion of data center resources remains unused. In this paper, we propose a framework that leverages unused resources of data centers, which are ephemeral by nature, to run MapReduce jobs. Our approach allows: i) to run efficiently Hadoop jobs on top of heterogeneous Cloud resources, thanks to our data placement strategy, ii) to predict accurately the volatility of ephemeral resources, thanks to the quantile regression method, and iii) for avoiding the interference between MapReduce jobs and co-resident workloads, thanks to our reactive QoS controller. We have extended Hadoop implementation with our framework and evaluated it with three different data center workloads. The experimental results show that our approach divides Hadoop job execution time by up to 7 when compared to the standard Hadoop implementation. Jean-Emile Dartois, Heverson B. Ribeiro, Jalil Boukhobza, Olivier Barais |
CLOUD | 3 |
| 2019 | Tracking Application Fingerprint in a Trustless Cloud Environment for Sabotage DetectionabstractCompanies are more and more inclined to use collaborative cloud resources when their maximum internal capacities are reached in order to minimize their TCO. The downside of using such a collaborative cloud, made of private clouds' unused resources, is that malicious resource providers may sabotage the correct execution of third-party-owned applications due to its uncontrolled nature. In this paper, we propose an approach that allows sabotage detection in a trustless environment. To do so, we designed a mechanism that (1) builds an application fingerprint considering a large set of resources usage (such as CPU, I/O, memory) in a trusted environment using random forest algorithm, and (2) an online remote fingerprint recognizer that monitors application execution and that makes it possible to detect unexpected application behavior. Our approach has been tested by building the fingerprint of 5 applications on trusted machines. When running these applications on untrusted machines (with either homogeneous, heterogeneous or unspecified hardware from the one that was used to build the model), the fingerprint recognizer was able to ascertain whether the execution of the application is correct or not with a median accuracy of about 98% for heterogeneous hardware and about 40% for the unspecified one. Jean-Emile Dartois, Jalil Boukhobza, Vincent Françoise, Olivier Barais |
MASCOTS | 2 |
| 2019 | Leveraging cloud unused resources for Big data application while achieving SLAabstractIn this demo paper, we present an architecture that leverages unused but volatile Cloud resources to run big data jobs. It is based on a learning algorithm that accurately predicts future availability of resources to automatically scale the ran jobs. We also designed a mechanism that avoids interference between the Big data jobs and co-resident workloads. Our solution is based on Open-Source components such as kubernetes and Apache Spark. Jean-Emile Dartois, Ivan Meriau, Mohamed Handaoui, Jalil Boukhobza, Olivier Barais |
MASCOTS | 4 |
| 2019 | K -MLIO: Enabling K -Means for Large Data-Sets and Memory Constrained Embedded SystemsabstractMachine Learning (ML) algorithms are increasingly used in embedded systems to perform different tasks such as clustering and pattern recognition. These algorithms are both compute and memory intensive whilst embedded devices offer lower hardware capabilities as compared to traditional ML platforms. K-means clustering is one of the widely used ML algorithms. In the case of large data-sets, our analysis showed that on average, more than 70% of the execution time is spent on I/Os. In this paper, we present a version of K-means that drastically reduces the number of I/Os by spanning the data-set only once as compared to the traditional version that reads it several times according to the number of iterations performed. Our evaluation showed that the proposed strategy reduces the overall execution time on large data-sets by 60% on average while lowering the number I/Os operations by 90% with a comparable precision to the traditional K-means implementation. Camélia Slimani, Stéphane Rubini, Jalil Boukhobza |
MASCOTS | 3 |
| 2019 | Optimizing the cost of DBaaS object placement in hybrid storage systems
Djillali Boukhelef, Jalil Boukhobza, Kamel Boukhalfa, Hamza Ouarnoughi, Laurent Lemarchand |
Future Gener. Comput. Syst. | 2 |
| 2019 | Preserving SSD lifetime in deep learning applications with delta snapshots
Jalil Boukhobza, Zili Shao |
J. Parallel Distributed Comput. | 2 |
| 2018 | Using Quantile Regression for Reclaiming Unused Cloud Resources While Achieving SLAabstractAlthough Cloud computing techniques have reduced the total cost of ownership thanks to virtualization, the average usage of resources (e.g., CPU, RAM, Network, I/O) remains low. To address such issue, one may sell unused resources. Such a solution requires the Cloud provider to determine the resources available and estimate their future use to provide availability guarantees. This paper proposes a technique that uses machine learning algorithms (Random Forest, Gradient Boosting Decision Tree, and Long Short Term Memory) to forecast 24-hour of available resources at the host level. Our technique relies on the use of quantile regression to provide a flexible trade-off between the potential amount of resources to reclaim and the risk of SLA violations. In addition, several metrics (e.g., CPU, RAM, disk, network) were predicted to provide exhaustive availability guarantees. Our methodology was evaluated by relying on four in production data center traces and our results show that quantile regression is relevant to reclaim unused resources. Our approach may increase the amount of savings up to 20% compared to traditional approaches. Jean-Emile Dartois, Anas Knefati, Jalil Boukhobza, Olivier Barais |
CloudCom | 3 |
| 2018 | A Cost Model for Hybrid Storage Systems in a Cloud FederationsabstractA cloud federation gives to cloud service providers (CSP) the opportunity to collaborate in order to offer a better QoS to customers at a lower cost.To do so, CSPs make some spare resources available to others at a reduced cost.One of the most critical resources is the storage system as it represents the main system bottleneck.From this point of view, how to efficiently place data in a federation of Clouds with heterogeneous storage systems is a real challenge.To address this issue, one needs to accurately estimate the data placement cost.In this paper, we propose a cost model for hybrid storage systems in a cloud federation for a Database as a Service (DBaaS) application.It takes into account the storage system characteristics, customers I/O workloads and SLA.The proposed cost model considers both 1) Internal customers data placement cost including local placement, outsourcing, back-migration and penalty costs, and 2) External customers data placement cost including insourcing and geo-migration costs.It can be used to help in the decision-making process which aims to enhance customers QoS and reduce CSPs costs in a federation.Simulation results showed the relevance of the considered costs.We have shown that mis-considering some sub-costs may lead to a 95% cost error for external customers data placement and 80% for outsourcing customers.This may cause significant financial loss. Amina Chikhaoui, Kamel Boukhalfa, Jalil Boukhobza |
FedCSIS | 3 |
| 2018 | An Extension to iFogSim to Enable the Design of Data Placement StrategiesabstractFog computing consists in extending Cloud services down to the network edge by using resources such as base stations, routers and switches. It presents a dense, heterogeneous and geo-distributed infrastructure which pushes to investigate how data are placed within this infrastructure in order to minimize service latency, network utilization and energy consumption. iFogSim is a Fog and IoT environments simulator dedicated to manage IoT services in a Fog infrastructure. In this paper, we present an extension to iFogSim to be able to model and simulate scenarios with strategies aiming to optimize data placement in Fog and IoT contexts. Data placement problem is NP-Hard due to the large number of Fog nodes and the high amount of data to be placed. Thus, we added a support to divide and conquer strategies to subdivide the issued infrastructure into several parts hence reducing the data placement computing time. Moreover, the extension involves a generic smart city scenario with different workloads making it possible for the users to investigate the behavior of their strategies using various workloads. In order to optimize the execution time of the simulations, we parallelized the Floyd-Warshall algorithm. This algorithm is used in iFogSim to compute all shortest paths between nodes in order to simulate data transmission. We have evaluated this extension using the proposed smart city scenario with various infrastructure configurations. The experiments show that our extension has a small overhead in terms of simulation time and memory utilization. Mohammed Islam Naas, Jalil Boukhobza, Philippe Raipin Parvédy, Laurent Lemarchand |
ICFEC | 2 |
| 2018 | Emerging NVM: A Survey on Architectural Integration and Research ChallengesabstractThere has been a surge of interest in Non-Volatile Memory (NVM) in recent years. With many advantages, such as density and power consumption, NVM is carving out a place in the memory hierarchy and may eventually change our view of computer architecture. Many NVMs have emerged, such as Magnetoresistive random access memory (MRAM), Phase Change random access memory (PCM), Resistive random access memory (ReRAM), and Ferroelectric random access memory (FeRAM), each with its own peculiar properties and specific challenges. The scientific community has carried out a substantial amount of work on integrating those technologies in the memory hierarchy. As many companies are announcing the imminent mass production of NVMs, we think that it is time to have a step back and discuss the body of literature related to NVM integration. This article surveys state-of-the-art work on integrating NVM into the memory hierarchy. Specially, we introduce the four types of NVM, namely, MRAM, PCM, ReRAM, and FeRAM, and investigate different ways of integrating them into the memory hierarchy from the horizontal or vertical perspectives. Here, horizontal integration means that the new memory is placed at the same level as an existing one, while vertical integration means that the new memory is interleaved between two existing levels. In addition, we describe challenges and opportunities with each NVM technique. Jalil Boukhobza, Stéphane Rubini, Renhai Chen, Zili Shao |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2018 | Joint DVFS and Parallelism for Energy Efficient and Low Latency Software Video DecodingabstractIn this paper, we aim to bridge the gap between the energy efficiency of software and hardware video decoders by combining both DVFS and parallelism. For this purpose, we, first, propose an adaptive DVFS algorithm for energy efficient mono-core decoding of H.264 videos. The proposed solution uses metadata (normalized by MPEG) providing information about the upcoming workload. These metadata are processed within an adaptive filter to build dynamically an accurate complexity model used to calculate the minimal processor frequencies for decoding video frames while guaranteeing real time constraints. Then, we generalize the proposed DVFS to slice-based multi-threaded parallel video decoders on multi-core platforms. Our performance evaluations showed that the proposed algorithm for mono-core decoding is able to converge to an accurate complexity model (4 percent) in less than 1 second. Moreover, it is simple to implement, induces very low overhead and achieves up to 46 percent energy saving as compared to the ondemand Linux DVFS governor. On the other hand, joint use of parallelism and DVFS allows 720p software video decoding with only 17 percent more energy consumption as compared to a hardware video decoder. Yahia Benmoussa, Eric Senn, Nicolas Derouineau, Nicolas Tizon, Jalil Boukhobza |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2017 | COPS: Cost Based Object Placement Strategies on Hybrid Storage System for DBaaS CloudabstractSolid State Drives (SSD) are integrated together with Hard Disk Drives (HDD) in Hybrid Storage Systems (HSS) for Cloud environment. When it comes to storing data, some placement strategies are used to find the best location (SSD or HDD). These strategies should minimize the cost of data placement while satisfying Service Level Objectives (SLO). This paper presents two Cost based Object Placement Strategies (COPS) for DBaaS objects in HSS: a Genetic based approach (G-COPS) and an ad-hoc Heuristic approach (H-COPS) based on incremental optimization. While G-COPS proved to be closer to the optimal solution in case of small instances, H-COPS showed a better scalability as it approached the exact solution even for large instances (by 10% in average). In addition, H-COPS showed small execution times (few seconds) even for large instances which makes it a good candidate to be used in runtime. Both H-COPS and G-COPS performed better than state-of-the-art solutions as they satisfied SLOs while reducing the overall cost by more than 40% for problems of small and large instances. Djillali Boukhelef, Kamel Boukhalfa, Jalil Boukhobza, Hamza Ouarnoughi, Laurent Lemarchand |
CCGrid | 3 |
| 2017 | iFogStor: An IoT Data Placement Strategy for Fog InfrastructureabstractInternet of Things (IoT) will be one of the driving application for digital data generation in the next years as more than 50 billions of objects will be connected by 2020. IoT data can be processed and used by different devices spread all over the network. The traditional way of centralizing data processing in the Cloud can hardly scale because it cannot satisfy many of the latency critical IoT applications. In addition, it generates a too high network traffic when the number of objects and services increase. Fog infrastructure provides a beginning of an answer to such an issue. In this paper, we present a data placement strategy for Fog infrastructures called iFogStor. The objective of iFogStor is to take profit of the heterogeneity and location of Fog nodes to reduce the overall latency of storing and retrieving data in a Fog. We formulated the data placement problem as a Generalized Assignment Problem (GAP) and proposed two ways to solve it: 1) an exact solution using integer programming and 2) a heuristic one based on geographical zoning to reduce the solving time. Both solutions proved very good performance as they reduced the latency by more than 86% as compared to a Cloud based solution and by 60% as compared to a naive Fog solution. Using geographical zoning heuristic can allow solving problems with large number of Fog nodes efficiently and in a couple of seconds making iFogStor feasible in runtime and scalable. Mohammed Islam Naas, Philippe Raipin Parvédy, Jalil Boukhobza, Laurent Lemarchand |
ICFEC | 3 |
| 2017 | MONTRES : Merge ON-the-Run External Sorting Algorithm for Large Data Volumes on SSD Based Storage SystemsabstractExternal sorting algorithms are commonly used by data-centric applications to sort quantities of data that are larger than the main-memory. Many external sorting algorithms were proposed in state-of-the-art studies to take advantage of SSD performance properties to accelerate the sorting process. In this paper, we demonstrate that unfortunately, many of those algorithms fail to scale when it comes to increasing the dataset size under memory pressure. In order to address this issue, we propose a new sorting algorithm named MONTRES. MONTRES relies on SSD performance model while decreasing the overall number of I/O operations. It does this by reducing the amount of temporary data generated during the sorting process by continuously evicting small values in the final sorted file. MONTRES scales well with growing datasets under memory pressure. We tested MONTRES using several data distributions, different amounts of main-memory workspace and three SSD models. Results showed that MONTRES outperforms state-of-the-art algorithms as it reduces the sorting execution time of TPC-H datasets by more than 30 percent when the file size to main-memory size ratio is high. Arezki Laga, Jalil Boukhobza, Frank Singhoff, Michel Koskas |
IEEE Trans. Computers | 2 |
| 2016 | A Cost Model for DBaaS Storage
Djillali Boukhelef, Jalil Boukhobza, Kamel Boukhalfa |
DEXA (1) | 2 |
| 2016 | A Cost Model for Virtual Machine Storage in Cloud IaaS ContextabstractThis paper proposes a storage system cost model for Infrastructure as a Service (IaaS) Cloud. The proposed cost model takes into account the virtualization environment, the storage system characteristics in addition to energy and QoS related parameters (Service Level Agreement and penalties). We show that those parameters are relevant and allow us to predict an accurate estimation of the overall cost of the IaaS infrastructure. We validate this cost model against real measures and we show less than 10% of error in most cases. Designers and administrators can use this cost model to perform optimization, load balancing, configuration and pricing of the Cloud infrastructure. Hamza Ouarnoughi, Jalil Boukhobza, Frank Singhoff, Stéphane Rubini |
PDP | 2 |
| 2016 | A Methodology for Estimating Performance and Power Consumption of Embedded Flash File SystemsabstractIn the embedded systems domain, obtaining performance and power consumption estimations is extremely valuable in numerous cases. This is particularly true during the design stage, as designers of complex embedded systems face an increasingly large design space. Secondary storage is a well-known performance bottleneck and has also been reported as an important factor of power consumption. Flash memory is the main secondary storage media in an embedded system and exhibits specific constraints in its usage. One popular way to manage these constraints is to use dedicated Flash File Systems (FFS). In this article, we propose a methodology to estimate the performance and power consumption of applicative I/Os on an FFS-based storage system within embedded Linux. The methodology is divided into three sequential steps. In the exploration phase, the main factors of an FFS storage system impacting performance and power consumption are identified. In the modeling phase, this impact is formalized into models. Finally, in the last phase, the models are implemented in a simulator named OpenFlash. OpenFlash allows obtaining performance and power consumption estimations for an applicative workload processed by the Linux FFS storage stack on an embedded platform. The simulator is validated against real measurements and the estimation error stays below 10%. Pierre Olivier, Jalil Boukhobza, Eric Senn, Hamza Ouarnoughi |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2015 | Addressing cache related preemption delay in fixed priority assignmentabstractHandling cache related preemption delay (CRPD) in preemptive scheduling context for real-time embedded systems still stays an open issue despite of its practical importance. Indeed, classical priority assignment algorithms are only optimal when preemption costs are neglected. For example, with Audsley's Optimal Priority Assignment (OPA), as the original algorithm does not take CRPD into account, it fails frequently in identifying the schedulable task sets as it happens that the algorithm qualifies a task set to be schedulable, while it is practically not because of CRPD. In this article, we propose an approach to adapt fixed priority assignment algorithms to real-time embedded systems with cache memory. For such a purpose, we propose three extensions of the original OPA algorithm that have different degrees of pessimism, different complexities, and give different results in terms of schedulable task sets coverage. Exhaustive experimentations were achieved to evaluate the proposed approaches in terms of complexity and efficiency. The result shows that our approach provides a mean to guarantee the schedulability of the real-time embedded system while taking into account CRPD. Hai Nam Tran, Frank Singhoff, Stéphane Rubini, Jalil Boukhobza |
ETFA | 4 |
| 2015 | A methodology for performance/energy consumption characterization and modeling of video decoding on heterogeneous SoC and its applications
Yahia Benmoussa, Jalil Boukhobza, Eric Senn, Yassine Hadjadj-Aoul, Djamel Benazzouz |
J. Syst. Archit. | 2 |
| 2015 | MaCACH: An adaptive cache-aware hybrid FTL mapping scheme using feedback control for efficient page-mapped space management
Jalil Boukhobza, Pierre Olivier, Stéphane Rubini, Laurent Lemarchand, Yassine Hadjadj-Aoul, Arezki Laga |
J. Syst. Archit. | 1 |
| 2014 | What can Emerging Hardware do for your DBMS Buffer?abstractThe spectacular development of business intelligence applications (BIA), built around the data warehousing technology, increases the demand on query performance of DBMS hosting with its extremely high amount of data. In such a context a high interaction among queries exists since they share a large number of intermediate results. This is due to the fact that BIA use relational schemes such as a star schema in which each join passes through the fact table. The decision to cache these intermediate results in the traditional buffer becomes a critical issue since it depends on the size of the buffer and the number of intermediate results candidate for caching. As flash memory is more and more adopted in mass storage systems, we rely on it to buffer some intermediate results. In this paper, we first propose to couple the RAM and Solid State Drive, to respond to the problem combining buffer management and query scheduling sub problems. Secondly, a cost model for evaluating the quality of buffering data and scheduling queries is given. Based on this cost model, an algorithm is given to solve our joint problem. Simulations show that our proposal enhances the performance of SQL queries up to 86%. Salmi Cheikh, Abdelhakim Nacef, Ladjel Bellatreche, Jalil Boukhobza |
DOLAP | 4 |
| 2014 | Instruction Cache in Hard Real-Time Systems: Modeling and Integration in Scheduling Analysis Tools with AADLabstractCache prediction for real-time systems in a preemptive scheduling context is still an open issue despite its practical importance. In this paper, we propose a modeling approach for taking into account the cache memory in realtime scheduling analysis. The goal is to have a simple but practical implementation to handle the cache memory with a real-time scheduling analyzer. The proposed contribution consists of three main parts: (1) modeling the targeted system with the Architecture Analysis and Design Language (AADL), (2) applying the cache analysis methods in a real time scheduling analysis tool and (3) performing scheduling simulation to access schedulability. For such a purpose, we present an extension of both the scheduling analysis tool Cheddar and of the AADL modeling language in order to integrate the cache modeling and analysis methodology we proposed. Experiments are presented to illustrate our propositions. They provide results on analysis that show examples of the timing impact of task preemption as well as the increase in overall responses time of the task set. This impact is important and the developed tool provides means to precisely assess it. Hai Nam Tran, Frank Singhoff, Stéphane Rubini, Jalil Boukhobza |
EUC | 4 |
| 2014 | Open-PEOPLE, A Collaborative Platform for Remote & Accurate Measurement and Evaluation of Embedded Systems Power ConsumptionabstractThis paper presents Open-PEOPLE, a hardware and software platform which aims to widen access to accurate power consumption measurement to the scientific community. The idea behind this platform is to centralize and abstract the instrumentation effort and the investment cost then allow geographically remote users to make power measurement remotely without requiring any specific costly hardware and/or software. Yahia Benmoussa, Eric Senn, Jalil Boukhobza, Mickael Lanoe, Djamel Benazzouz |
MASCOTS | 3 |
| 2013 | Energy Consumption Modeling of H.264/AVC Video Decoding for GPP and DSPabstractMobile devices such as smart-phones and tablets are becoming the most important channel for delivering end-user Internet traffic especially multimedia content. One of the most popular multimedia application is video streaming. The video decoding process of this application is compute-intensive and is responsible of the consumption of a considerable part of the energy budget. Those mobile devices contain heterogeneous processing elements among-which we find Digital Signal Processors (DSP) and General Purpose Processors (GPP). In this context, the performance and energy estimation of those complex platforms is a difficult and time consuming task especially when considering both hardware and applicative parameters. In this paper, we propose a methodology for developing a unified high level video decoding performance and energy consumption analytical model for embedded heterogeneous platforms. This methodology is based on experimental measurements conducted on an embedded low-power platform. The developed model describes the performance and the energy consumption of H.264/AVC video decoding on both GPP and DSP in terms of video bit-rate, clock frequency and a set of comprehensive hardware and video related coefficients. It achieves a balance between a too abstract high level model and a detailed lower level one while guaranteeing a very good prediction properties (R-squared = 97%) for the tested videos. As a use case, we show that our model allows to accurately determine the bit-rate values for which video decoding on GPP is more energy-efficient than on DSP for a given platform. Yahia Benmoussa, Jalil Boukhobza, Eric Senn, Djamel Benazzouz |
DSD | 2 |
| 2013 | How to exploit the device diversity and database interaction to propose a generic cost model?abstractCost models have been following the life cycle of databases. In the first generation, they have been used by query optimizers, where the cost-based optimization paradigm has been developed and supported by most of important optimizers. The spectacular development of complex decision queries amplifies the interest of the physical design phase (PhD), where cost models are used to select the relevant optimization techniques such as indexes, materialized views, etc. Most of these cost models are usually developed for one storage device (usually disk) with a well identified storage model and ignore the interaction between the different components of databases: interaction between optimization techniques, interaction between queries, interaction between devices, etc. In this paper, we propose a generic cost model for the physical design that can be instantiated for each need. We contribute an ontology describing storage devices. Furthermore, we provide an instantiation of our meta model for two interdependent problems: query scheduling and buffer management. The evaluation results show the applicability of our model as well as its effectiveness. Ladjel Bellatreche, Salmi Cheikh, Sebastian Breß, Amira Kerkad, Ahcène Boukorca, Jalil Boukhobza |
IDEAS | 6 |
| 2013 | GPP vs DSP: A Performance/Energy Characterization and Evaluation of Video DecodingabstractMobile devices such as smart-phones and tablets are increasingly becoming the most important channel for delivering end-user Internet traffic especially multimedia content. One of the most popular use of these terminals is video streaming. In this type of application, video decoding is considered as the most compute and energy intensive part. Some specific processing units, such as dedicated Digital Signal Processors (DSPs), are added to those devices in order to optimize the performance and energy consumption. In this context, the objective of this paper is to give a comprehensive and comparative study of the performance and energy consumption of video decoding application on embedded heterogeneous platforms containing a GPP and a DSP. To achieve this goal, a performance and energy characterization methodology for H.264/AVC video decoding is proposed. This methodology considers a large set of video coding parameters and operating clock frequencies to reflect different execution scenarios ranging from low-quality video decoding on low-end mobile phones to high-quality video decoding on tablets. The obtained results revealed that the best performance-energy trade-off highly depends on the required video bit-rate and resolution. For instance, the GPP can be the best choice in many cases due to a significant overhead in DSP decoding which may represent 30% of the total decoding energy in some cases. Some explanations about the obtained performance and overheads are given. Finally, guidelines on which processing element to choose according to video properties are also proposed. Yahia Benmoussa, Jalil Boukhobza, Eric Senn, Djamel Benazzouz |
MASCOTS | 2 |
| 2013 | CACH-FTL: A Cache-Aware Configurable Hybrid Flash Translation LayerabstractMany hybrid Flash Translation Layer (FTL) schemes have been proposed to leverage the erase-before-write and limited lifetime constraints of flash memories. Those schemes try to approach page mapping performance and flexibility while seeking block mapping memory usage. Furthermore, flash-specific cache systems were designed (1) to maximize lifetime by absorbing some erase operations, and (2) to reveal sequentiality from random write operations. Indeed, random writes represent the Achilles' heel of flash memories. Both cache systems and FTL schemes were designed independently from each other. This paper presents a scalable (in terms of mapping table size) and flexible (in terms of I/O workload support) Cache-Aware Configurable Hybrid (CACH) FTL. CACH-FTL uses a common feature of flash-specific cache systems that is flushing groups of pages from the same block. CACH-FTL partitions the flash memory space into two regions: (1) a data Block Mapped Region (BMR) collecting large groups of pages from the above cache (sequential I/Os), and (2) a small Page Mapped over-provisioning Region (PMR) which purpose is to collect/buffer small groups of pages coming from the cache (random I/Os) before moving them to BMR. CACH-FTL is flexible as it offers many configuration possibilities and can be adapted according to the I/O workload. CACH-FTL approaches the ideal page mapping FTL performance as it gives less than 15% performance difference in most cases. Jalil Boukhobza, Pierre Olivier, Stéphane Rubini |
PDP | 1 |
| 2011 | Characterization of OLTP I/O Workloads for Dimensioning Embedded Write Cache for Flash Memories: A Case Study
Jalil Boukhobza, Ilyes Khetib, Pierre Olivier |
MEDI | 1 |