EDBT 2026 Demo / reviewers in the wild / expert
Hamza Ouarnoughi
dblp:157/3551
· DBLP profile ↗
27ranked-venue papers
1as first author
21since 2021 · last 2026
0000-0002-7490-5350ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 12 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SONATA: Self-adaptive Evolutionary Framework for Hardware-aware Neural Architecture SearchabstractInternational audience Halima Bouzidi, Smaïl Niar, Hamza Ouarnoughi, El-Ghazali Talbi |
GECCO | 3 |
| 2026 | Deep learning for anomaly detection in railway systems: A structured survey
Ammar Bouketta, Smaïl Niar, Hamza Ouarnoughi |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | Neural Architecture Search and Automatic Code Optimization: Techniques, trends, and challengesabstractDeep Learning models have experienced exponential growth in complexity and resource demands in recent years. Accelerating these models for efficient execution on resource-constrained devices has become more crucial than ever. Two notable techniques used to achieve this goal are Hardware-Aware Neural Architecture Search (HW-NAS) and Automatic Code Optimization (ACO). HW-NAS automatically designs accurate yet hardware-friendly neural networks, while ACO involves searching for the best code optimizations to apply on neural networks for efficient mapping and inference on the target hardware. This review explores recent work that combines these two techniques within a single framework. We present the fundamental principles of both domains and demonstrate their suboptimality when performed independently. We then investigate their integration into a joint optimization process that we call Hardware Aware- N eural A rchitecture and C ompiler O ptimizations co- S earch (NACOS). Inas Bachiri, Smaïl Niar, Riyadh Baghdadi, Hamza Ouarnoughi |
J. Syst. Archit. | 4 |
| 2025 | Machine Learning Vulnerabilities in 6G: Adversarial Attacks and Their Impact on Channel Gain Prediction and Resource Allocation in UC-CFmMIMO
Mahmoud Ghorbel, Selina Cheggour, Valeria Loscrì, Youcef Imine, Hamza Ouarnoughi, Smaïl Niar |
ESORICS (1) | 5 |
| 2024 | NeuraSearchLib: A Modular and Extensible Library for Neural Architecture Search with Configurable Search SpacesabstractWhile Deep Neural Networks (DNN) have driven technological innovation, designing new architectures remains labor-intensive, requiring human expertise and numerous trial-and-error iterations. Neural Architecture Search (NAS) aims to automate this process for more efficient exploration of DNN architectures. However, current NAS libraries lack flexibility, particularly in DNN search space customization, and lack universal, ready-to-use templates for state-of-the-art NAS search spaces. This paper introduces NeuraSearchLib, a modular and extensible NAS library that allows users to create and customize search spaces with high flexibility using simple high-level specifications. NeuraSearchLib1can reproduce existing NAS search spaces, create new ones, and conduct experiments with various search strategies, training methods, and performance evaluations, enabling researchers to explore new DNN architectures more efficiently. Kaouthar Essaheli, Selsabil Roubi, Halima Bouzidi, Hamza Ouarnoughi, Smaïl Niar |
BDCAT | 4 |
| 2024 | Analysis of Parallel Graph ApplicationsabstractDespite the increasing computing power of shared memory systems with high core counts, parallel graph processing frameworks cannot exploit it effectively. The reason behind this is the inherent challenges in parallel graph algorithms, which are efficient management of dynamically created tasks and irregular data access patterns. In this paper, we categorize several popular design choices into three design dimensions: (i) execution mode, (ii) data access pattern, and (iii) work activation. We provide their high-level parallel implementations and analyze various implementations of three representative iterative graph algorithms by considering these design dimensions. To gain a better understanding of design choices, we examine their impacts on performance, communication, scalability, and work efficiency. We also investigate the communication characteristics of the design choices on two state-of-the-art shared-memory platforms by performing micro-architectural analysis. Our microarchitectural analysis reveals that a topology-driven, pull-based model gives up to $20 x$ better performance. Funda Atik, Serif Yesil, Hamza Ouarnoughi, Smaïl Niar, Ozcan Ozturk 0001 |
ICPADS | 3 |
| 2024 | DfuseNAS: A Diffusion-Based Neural Architecture SearchabstractDeep Learning (DL) has revolutionized numerous domains by crafting highly effective Neural Network (NN) architectures. However, manual engineering approaches for NN design often yield sub-optimal solutions. Neural Architecture Search (NAS) addresses this issue by automating the design process and discovering state-of-the-art architectures. Nevertheless, the exploration of vast search spaces in NAS remains a challenging endeavor. Diffusion models have proven effective in traversing expansive search spaces encountered in generative image tasks. Their innate ability to compress and explore these spaces has shown a lot of promise. Building upon this inspiration, we introduce DfuseNAS, a novel NAS methodology rooted in diffusion processes. DfuseNAS brings substantial improvements in both NAS search efficiency and the quality of the generated neural network architectures. To the best of our knowledge, our work marks a pioneering effort in applying diffusion algorithms to enhance the search space exploration with NAS. Our experimental results, conducted on the widely-used NAS-Bench-101, showcase the remarkable capabilities of DfuseNAS. We achieved the highest average accuracy, outperforming other state-of-the-art methods, while completing the search process at least 2 times faster. Moreover, when provided with a specific architecture and a given task, the application of DfuseNAS consistently led to the generation of more accurate architectures in 98% of the times. Lotfi Abdelkrim Mecharbat, Hadjer Benmeziane, Hamza Ouarnoughi, Smaïl Niar, Kaoutar El Maghraoui |
IJCNN | 3 |
| 2024 | Grassroots operator search for model edge adaptation using mathematical search space
Hadjer Benmeziane, Kaoutar El Maghraoui, Hamza Ouarnoughi, Smaïl Niar |
Future Gener. Comput. Syst. | 3 |
| 2023 | Harmonic-NAS: Hardware-Aware Multimodal Neural Architecture Search on Resource-constrained Devices
Mohamed Imed Eddine Ghebriout, Halima Bouzidi, Smaïl Niar, Hamza Ouarnoughi |
ACML | 4 |
| 2023 | Map-and-Conquer: Energy-Efficient Mapping of Dynamic Neural Nets onto Heterogeneous MPSoCsabstractHeterogeneous MPSoCs comprise diverse processing units of varying compute capabilities. To date, the mapping strategies of neural networks (NNs) onto such systems are yet to exploit the full potential of processing parallelism, made possible through both the intrinsic NNs’ structure and underlying hardware composition. In this paper, we propose a novel framework to effectively map NNs onto heterogeneous MPSoCs in a manner that enables them to leverage the underlying processing concurrency. Specifically, our approach identifies an optimal partitioning scheme of the NN along its ‘width’ dimension, which facilitates deployment of concurrent NN blocks onto different hardware computing units. Additionally, our approach contributes a novel scheme to deploy partitioned NNs onto the MPSoC as dynamic multi-exit networks for additional performance gains. Our experiments on a standard MPSoC platform have yielded dynamic mapping configurations that are 2.1x more energy-efficient than the GPU-only mapping while incurring 1.7x less latency than DLA-only mapping. Halima Bouzidi, Mohanad Odema, Hamza Ouarnoughi, Smaïl Niar, Mohammad Abdullah Al Faruque |
DAC | 3 |
| 2023 | HADAS: Hardware-Aware Dynamic Neural Architecture Search for Edge Performance ScalingabstractDynamic neural networks (DyNNs) have become viable techniques to enable intelligence on resource-constrained edge devices while maintaining computational efficiency. In many cases, the implementation of DyNNs can be sub-optimal due to its underlying backbone architecture being developed at the design stage independent of both: (i) potential support for dynamic computing, e.g. early exiting, and (ii) resource efficiency features of the underlying hardware, e.g., dynamic voltage and frequency scaling (DVFS). Addressing this, we present HADAS, a novel Hardware-Aware Dynamic Neural Architecture Search framework that realizes DyNN architectures whose backbone, early exiting features, and DVFS settings have been jointly optimized to maximize performance and resource efficiency. Our experiments using the CIFAR-100 dataset and a diverse set of edge computing platforms have shown that HADAS can elevate dynamic models' energy efficiency by up to 57% for the same level of accuracy scores. Our code is available at https://github.com/HalimaBouzidi/HADAS Halima Bouzidi, Mohanad Odema, Hamza Ouarnoughi, Mohammad Abdullah Al Faruque, Smaïl Niar |
DATE | 3 |
| 2023 | Pareto Rank-Preserving Supernetwork for Hardware-Aware Neural Architecture SearchabstractIn neural architecture search (NAS), training every sampled architecture is very time-consuming and should be avoided. Weight-sharing is a promising solution to speed up the evaluation process. However, training the supernetwork incurs many discrepancies between the actual ranking and the predicted one. Additionally, efficient deep-learning engineering processes require incorporating realistic hardware-performance metrics into the NAS evaluation process, also known as hardware-aware NAS (HW-NAS). In HW-NAS, estimating task-specific performance and hardware efficiency are both required. This paper proposes a supernetwork training methodology that preserves the Pareto ranking between its different subnetworks resulting in more efficient and accurate neural networks for a variety of hardware platforms. The results show a 97% near Pareto front approximation in less than 2 GPU days of search, which provides 2x speed up compared to state-of-the-art methods. We validate our methodology on NAS-Bench-201, DARTS, and ImageNet. Our optimal model achieves 77.2% accuracy (+1.7% compared to baseline) with an inference time of 3.68ms on Edge GPU for ImageNet, which yields a 2.3x speedup. Training implementation can be found: https://github.com/IHIaadj/PRP-NAS. Hadjer Benmeziane, Kaoutar El Maghraoui, Hamza Ouarnoughi, Smaïl Niar |
ECAI | 3 |
| 2023 | FLASH-RL: Federated Learning Addressing System and Static Heterogeneity using Reinforcement LearningabstractWe propose FLASH-RL, a framework utilizing Double Deep Q-Learning (DDQL) to address system and static heterogeneity in Federated Learning (FL). FLASH-RL introduces a new reputation-based utility function to evaluate client contributions based on their current and past performances. Additionally, an adapted DDQL algorithm is proposed to expedite the learning process. Experimental results on MNIST and CIFAR-10 datasets demonstrate that FLASH-RL strikes a balance between model performance and end-to-end latency, reducing latency by up to 24.83% compared to FedAVG and 24.67% compared to FAVOR. It also reduces training rounds by up to 60.44% compared to FedAVG and 76% compared to FAVOR. Similar improvements are observed on the MobiAct Dataset for fall detection, underscoring the real-world applicability of our approach. Sofiane Bouaziz, Hadjer Benmeziane, Youcef Imine, Leila Hamdad, Smaïl Niar, Hamza Ouarnoughi |
ICCD | 6 |
| 2023 | Multi-objective Hardware-aware Neural Architecture Search with Pareto Rank-preserving Surrogate ModelsabstractDeep learning (DL) models such as convolutional neural networks (ConvNets) are being deployed to solve various computer vision and natural language processing tasks at the edge. It is a challenge to find the right DL architecture that simultaneously meets the accuracy, power, and performance budgets of such resource-constrained devices. Hardware-aware Neural Architecture Search (HW-NAS) has recently gained steam by automating the design of efficient DL models for a variety of target hardware platforms. However, such algorithms require excessive computational resources. Thousands of GPU days are required to evaluate and explore an architecture search space such as FBNet [ 45 ]. State-of-the-art approaches propose using surrogate models to predict architecture accuracy and hardware performance to speed up HW-NAS. Existing approaches use independent surrogate models to estimate each objective, resulting in non-optimal Pareto fronts. In this article, HW-PR-NAS, 1 a novel Pareto rank-preserving surrogate model for edge computing platforms, is presented. Our model integrates a new loss function that ranks the architectures according to their Pareto rank, regardless of the actual values of the various objectives. We employ a simple yet effective surrogate model architecture that can be generalized to any standard DL model. We then present an optimized evolutionary algorithm that uses and validates our surrogate model. Our approach has been evaluated on seven edge hardware platforms from various classes, including ASIC, FPGA, GPU, and multi-core CPU. The evaluation results show that HW-PR-NAS achieves up to 2.5× speedup compared to state-of-the-art methods while achieving 98% near the actual Pareto front. Hadjer Benmeziane, Hamza Ouarnoughi, Kaoutar El Maghraoui, Smaïl Niar |
ACM Trans. Archit. Code Optim. | 2 |
| 2023 | MaGNAS: A Mapping-Aware Graph Neural Architecture Search Framework for Heterogeneous MPSoC DeploymentabstractGraph Neural Networks (GNNs) are becoming increasingly popular for vision-based applications due to their intrinsic capacity in modeling structural and contextual relations between various parts of an image frame. On another front, the rising popularity of deep vision-based applications at the edge has been facilitated by the recent advancements in heterogeneous multi-processor Systems on Chips (MPSoCs) that enable inference under real-time, stringent execution requirements. By extension, GNNs employed for vision-based applications must adhere to the same execution requirements. Yet contrary to typical deep neural networks, the irregular flow of graph learning operations poses a challenge to running GNNs on such heterogeneous MPSoC platforms. In this paper, we propose a novel unified design-mapping approach for efficient processing of vision GNN workloads on heterogeneous MPSoC platforms. Particularly, we develop MaGNAS, a mapping-aware Graph Neural Architecture Search framework. MaGNAS proposes a GNN architectural design space coupled with prospective mapping options on a heterogeneous SoC to identify model architectures that maximize on-device resource efficiency. To achieve this, MaGNAS employs a two-tier evolutionary search to identify optimal GNNs and mapping pairings that yield the best performance trade-offs. Through designing a supernet derived from the recent Vision GNN (ViG) architecture, we conducted experiments on four (04) state-of-the-art vision datasets using both ( i ) a real hardware SoC platform (NVIDIA Xavier AGX) and ( ii ) a performance/cost model simulator for DNN accelerators. Our experimental results demonstrate that MaGNAS is able to provide 1.57 × latency speedup and is 3.38 × more energy-efficient for several vision datasets executed on the Xavier MPSoC vs. the GPU-only deployment while sustaining an average 0.11% accuracy reduction from the baseline. Mohanad Odema, Halima Bouzidi, Hamza Ouarnoughi, Smaïl Niar, Mohammad Abdullah Al Faruque |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2022 | CaW-NAS: Compression Aware Neural Architecture SearchabstractWith the ever-growing demand for deep learning (DL) at the edge, building small and efficient DL architectures has become a significant challenge. Optimization techniques such as quantization, pruning or hardware-aware neural architecture search (HW-NAS) have been proposed. In this paper, we present an efficient HW-NAS; Compression-Aware Neural Architecture search (CaW-NAS), that combines the search for the architecture and its quantization policy. While former works search over a fully quantized search space, we define our search space with quantized and non-quantized architectures. Our search strategy finds the best trade-off between accuracy and latency according to the target hardware. Experimental results on a mobile platform show that, our method allows to obtain more efficient networks in terms of accuracy, execution time and energy consumption when compared to the state of the art. Hadjer Benmeziane, Hamza Ouarnoughi, Smaïl Niar, Kaoutar El Maghraoui |
DSD | 2 |
| 2022 | Co-Optimization of DNN and Hardware Configurations on Edge GPUsabstractThe ever-increasing complexity of both Deep Neural Networks (DNN) and hardware accelerators has made the co-optimization of these domains extremely complex. Previous works typically focus on optimizing DNNs given a fixed hardware configuration or optimizing a specific hardware architecture given a fixed DNN model. Recently, the importance of the joint exploration of the two spaces drew more and more attention. Our work targets the co-optimization of DNN and hardware configurations on edge GPU accelerators. We propose an evolutionary-based co-optimization strategy by considering three metrics: DNN accuracy, execution latency, and power consumption. By combining the two search spaces, a larger number of configurations can be explored in a short time interval. In addition, a better tradeoff between DNN accuracy and hardware efficiency can be obtained. Experimental results show that the co-optimization outperforms the optimization of DNN for fixed hardware configuration with up to 53% hardware efficiency gains with the same accuracy and inference time. Halima Bouzidi, Hamza Ouarnoughi, Smaïl Niar, El-Ghazali Talbi, Abdessamad Ait El Cadi |
DSD | 2 |
| 2022 | Pareto Rank Surrogate Model for Hardware-aware Neural Architecture SearchabstractHardware-aware Neural Architecture Search (HWNAS) has recently gained much attention by automating the design of efficient deep learning models with tiny resources and reduced inference time requirements. However, HW-NAS inherits and exacerbates the expensive computational complexity of general NAS due to its significantly increased search spaces and more complex NAS evaluation component. To speed up HWNAS, existing efforts use surrogate models to predict a neural architecture’s accuracy and hardware performance on a specific platform. Thereby reducing the expensive training process and significantly reducing search time. We show that using multiple surrogate models to estimate the different objectives does not achieve the true Pareto front. Therefore, we propose HW-PRNAS, a novel Pareto Rank-preserving surrogate model. HWPR-NAS training is based on a new loss function that ranks the architectures according to their Pareto front. We evaluate our approach on seven different hardware platforms, including ASIC, FPGA, GPU and multi-cores. Our results show that we can achieve up to 2. 5x speedup while achieving better Pareto-front results than state of the art surrogate models. Hadjer Benmeziane, Smaïl Niar, Hamza Ouarnoughi, Kaoutar El Maghraoui |
ISPASS | 3 |
| 2022 | Performance Modeling of Computer Vision-based CNN on Edge GPUsabstractConvolutional Neural Networks (CNNs) are currently widely used in various fields, particularly for computer vision applications. Edge platforms have drawn tremendous attention from academia and industry due to their ability to improve execution time and preserve privacy. However, edge platforms struggle to satisfy CNNs’ needs due to their computation and energy constraints. Thus, it is challenging to find the most efficient CNN that respects accuracy, time, energy, and memory footprint constraints for a target edge platform. Furthermore, given the size of the design space of CNNs and hardware platforms, performance evaluation of CNNs entails several efforts. Consequently, designers need tools to quickly explore large design space and select the CNN that offers the best performance trade-off for a set of hardware platforms. This article proposes a Machine Learning (ML)–based modeling approach for CNN performances on edge GPU-based platforms for vision applications. We implement and compare five of the most successful ML algorithms for accurate and rapid CNN performance predictions on three different edge GPUs in image classification. Experimental results demonstrate the robustness and usefulness of our proposed methodology. For three of the five ML algorithms — XGBoost, Random Forest, and Ridge Polynomial regression — average errors of 11%, 6%, and 8% have been obtained for CNN inference execution time, power consumption, and memory usage, respectively. Halima Bouzidi, Hamza Ouarnoughi, Smaïl Niar, Abdessamad Ait El Cadi |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2021 | Performance prediction for convolutional neural networks on edge GPUsabstractEdge computing is increasingly used for Artificial Intelligence (AI) purposes to meet latency, privacy, and energy challenges. Convolutional Neural networks (CNN) are more frequently deployed on Edge devices for several applications. However, due to their constrained computing resources and energy budget, Edge devices struggle to meet CNN's latency requirements while maintaining good accuracy. It is, therefore, crucial to choose the CNN with the best accuracy and latency trade-off while respecting hardware constraints. This paper presents and compares five of the widely used Machine Learning (ML) based approaches to predict CNN's inference execution time on Edge GPUs. For these 5 methods, in addition to their prediction accuracy, we also explore the time needed for their training and their hyperparameters' tuning. Finally, we compare times to run the prediction models on different platforms. The use of these methods will highly facilitate design space exploration by quickly providing the best CNN on a target Edge GPU. Experimental results show that XGBoost provides an interesting average prediction error even for unexplored and unseen CNN architectures. Random Forest depicts comparable accuracy but needs more effort and time to be trained. The other 3 approaches (OLS, MLP, and SVR) are less accurate for CNN performance estimation. Halima Bouzidi, Hamza Ouarnoughi, Smaïl Niar, Abdessamad Ait El Cadi |
CF | 2 |
| 2021 | Hardware-Aware Neural Architecture Search: Survey and TaxonomyabstractThere is no doubt that making AI mainstream by bringing powerful, yet power hungry deep neural networks (DNNs) to resource-constrained devices would required an efficient co-design of algorithms, hardware and software. The increased popularity of DNN applications deployed on a wide variety of platforms, from tiny microcontrollers to data-centers, have resulted in multiple questions and challenges related to constraints introduced by the hardware. In this survey on hardware-aware neural architecture search (HW-NAS), we present some of the existing answers proposed in the literature for the following questions: "Is it possible to build an efficient DL model that meets the latency and energy constraints of tiny edge devices?", "How can we reduce the trade-off between the accuracy of a DL model and its ability to be deployed in a variety of platforms?". The survey provides a new taxonomy of HW-NAS and assesses the hardware cost estimation strategies. We also highlight the challenges and limitations of existing approaches and potential future directions. We hope that this survey will help to fuel the research towards efficient deep learning. Hadjer Benmeziane, Kaoutar El Maghraoui, Hamza Ouarnoughi, Smaïl Niar, Martin Wistuba, Naigang Wang |
IJCAI | 3 |
| 2020 | A Multi-Agent Approach for Vehicle-to-Fog Fair Computation OffloadingabstractFuture autonomous driving (AD) will require more data processing and storage capacities, exceeding the in-board capacities, especially with AI application requirements. The cloud or fog resources can provide off-load services. But to obtain an optimized task distribution between the on-board edge computing platform and cloud or Fog resources, communication delays and AD real-time constraints must be taken into account. The equity in resource allocation is rarely studied in this context where mobility adds a specific difficulty. Our approach of offloading mechanism is to leverage intelligent agents able to make decisions on task delegation based on past decisions. Agents at edge and fog levels communicate and exchange their knowledge and past decisions. A first scenario illustrates the proposition with simulated data. Emmanuelle Grislin, Hamza Ouarnoughi, Smaïl Niar |
AICCSA | 2 |
| 2020 | A GPU enhanced LIDAR Perception System for Autonomous VehiclesabstractEnvironment vision and understanding is a crucial task in Autonomous Driving (AD) context. This mainly needs image processing approaches such as Convolutional Neural Networks (CNN). Nevertheless, cameras have shown their limits for such a task, especially in dealing with difficult light conditions. LIDAR is a powerful and widely used sensor for AD. Indeed, LIDAR can then cope with the lack of information gathered from cameras. For AD, data processing from the sensors is the key function to obtain a high quality perception. For this, Graphics Processing Unit (GPU) platforms show great performances and outperform other processing platforms such as FPGA and Multi-cores. This work presents a new approach to produce multiple 2D representation from 3D points cloud coming from LIDAR. The 2D representation can therefore be used by any efficient image processing applications. Our approach uses only LIDAR sensor and exploits the high GPU parallelism for its implementation. The resulting 2D representations are then used by CNN for AD applications such as image classification and segmentation. Finally, our contributions have been evaluated using the KITTI road benchmark and showed encouraging results. Abderrahim Haneche, Mohammed Yazid Lachachi, Smaïl Niar, Hamza Ouarnoughi |
PDP | 4 |
| 2019 | Optimizing the cost of DBaaS object placement in hybrid storage systems
Djillali Boukhelef, Jalil Boukhobza, Kamel Boukhalfa, Hamza Ouarnoughi, Laurent Lemarchand |
Future Gener. Comput. Syst. | 4 |
| 2017 | COPS: Cost Based Object Placement Strategies on Hybrid Storage System for DBaaS CloudabstractSolid State Drives (SSD) are integrated together with Hard Disk Drives (HDD) in Hybrid Storage Systems (HSS) for Cloud environment. When it comes to storing data, some placement strategies are used to find the best location (SSD or HDD). These strategies should minimize the cost of data placement while satisfying Service Level Objectives (SLO). This paper presents two Cost based Object Placement Strategies (COPS) for DBaaS objects in HSS: a Genetic based approach (G-COPS) and an ad-hoc Heuristic approach (H-COPS) based on incremental optimization. While G-COPS proved to be closer to the optimal solution in case of small instances, H-COPS showed a better scalability as it approached the exact solution even for large instances (by 10% in average). In addition, H-COPS showed small execution times (few seconds) even for large instances which makes it a good candidate to be used in runtime. Both H-COPS and G-COPS performed better than state-of-the-art solutions as they satisfied SLOs while reducing the overall cost by more than 40% for problems of small and large instances. Djillali Boukhelef, Kamel Boukhalfa, Jalil Boukhobza, Hamza Ouarnoughi, Laurent Lemarchand |
CCGrid | 4 |
| 2016 | A Cost Model for Virtual Machine Storage in Cloud IaaS ContextabstractThis paper proposes a storage system cost model for Infrastructure as a Service (IaaS) Cloud. The proposed cost model takes into account the virtualization environment, the storage system characteristics in addition to energy and QoS related parameters (Service Level Agreement and penalties). We show that those parameters are relevant and allow us to predict an accurate estimation of the overall cost of the IaaS infrastructure. We validate this cost model against real measures and we show less than 10% of error in most cases. Designers and administrators can use this cost model to perform optimization, load balancing, configuration and pricing of the Cloud infrastructure. Hamza Ouarnoughi, Jalil Boukhobza, Frank Singhoff, Stéphane Rubini |
PDP | 1 |
| 2016 | A Methodology for Estimating Performance and Power Consumption of Embedded Flash File SystemsabstractIn the embedded systems domain, obtaining performance and power consumption estimations is extremely valuable in numerous cases. This is particularly true during the design stage, as designers of complex embedded systems face an increasingly large design space. Secondary storage is a well-known performance bottleneck and has also been reported as an important factor of power consumption. Flash memory is the main secondary storage media in an embedded system and exhibits specific constraints in its usage. One popular way to manage these constraints is to use dedicated Flash File Systems (FFS). In this article, we propose a methodology to estimate the performance and power consumption of applicative I/Os on an FFS-based storage system within embedded Linux. The methodology is divided into three sequential steps. In the exploration phase, the main factors of an FFS storage system impacting performance and power consumption are identified. In the modeling phase, this impact is formalized into models. Finally, in the last phase, the models are implemented in a simulator named OpenFlash. OpenFlash allows obtaining performance and power consumption estimations for an applicative workload processed by the Linux FFS storage stack on an embedded platform. The simulator is validated against real measurements and the estimation error stays below 10%. Pierre Olivier, Jalil Boukhobza, Eric Senn, Hamza Ouarnoughi |
ACM Trans. Embed. Comput. Syst. | 4 |