VLDB 2026 Research / reviewers in the wild / expert
Smaïl Niar
dblp:69/1858
· DBLP profile ↗
77ranked-venue papers
2as first author
27since 2021 · last 2026
0000-0002-7550-484XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 53 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 12 · 10 since 2021Software engineering, systems software and programming languages · 11 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SONATA: Self-adaptive Evolutionary Framework for Hardware-aware Neural Architecture SearchabstractInternational audience Halima Bouzidi, Smaïl Niar, Hamza Ouarnoughi, El-Ghazali Talbi |
GECCO | 2 |
| 2026 | OpenCL-based Deeply Pipelined HLS Implementation for Iterative Graph ApplicationsabstractGraph applications play a central role in different domains such as machine learning, data analytics, natural language processing, and fraud detection. Generating high‑performance kernels tailored to custom accelerator architectures for such workloads remains challenging due to their irregular memory access patterns and data‑dependent control flow. We propose an OpenCL‑based framework for the automated generation of deeply pipelined High Level Synthesis (HLS) implementations of iterative graph algorithms. The framework integrates a set of optimization techniques that restructure iterative graph algorithms to maximize pipeline utilization, reduce memory bottlenecks, and enable aggressive HLS optimizations. Although broadly applicable to a wide class of iterative graph workloads, we demonstrate the approach using PageRank as a representative case study. The experiment results demonstrate that efficient synthesis and high throughput can be obtained. The generated deeply pipelined HLS kernels can deliver substantial performance benefits for graph‑centric FPGA acceleration. Kenan Cagri Hirlak, Smaïl Niar, Ozcan Ozturk 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2026 | Deep learning for anomaly detection in railway systems: A structured survey
Ammar Bouketta, Smaïl Niar, Hamza Ouarnoughi |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Neural Architecture Search and Automatic Code Optimization: Techniques, trends, and challengesabstractDeep Learning models have experienced exponential growth in complexity and resource demands in recent years. Accelerating these models for efficient execution on resource-constrained devices has become more crucial than ever. Two notable techniques used to achieve this goal are Hardware-Aware Neural Architecture Search (HW-NAS) and Automatic Code Optimization (ACO). HW-NAS automatically designs accurate yet hardware-friendly neural networks, while ACO involves searching for the best code optimizations to apply on neural networks for efficient mapping and inference on the target hardware. This review explores recent work that combines these two techniques within a single framework. We present the fundamental principles of both domains and demonstrate their suboptimality when performed independently. We then investigate their integration into a joint optimization process that we call Hardware Aware- N eural A rchitecture and C ompiler O ptimizations co- S earch (NACOS). Inas Bachiri, Smaïl Niar, Riyadh Baghdadi, Hamza Ouarnoughi |
J. Syst. Archit. | 2 |
| 2025 | Zero-Shot Vision-Language Model for Event Detection in Smart SurveillanceabstractVision-Language Models (VLMs) have shown remarkable capabilities in processing visual data through natural language, yet their potential in smart surveillance for abnormal event detection remains underexplored. Traditional anomaly detection systems rely on predefined event classes, limiting their ability to identify novel or ambiguous anomalies in diverse scenarios. To address this, we propose VEZA (Vision Encoding for Zero-shot Anomalies), a novel architecture lever-aging pre-trained VLMs for zero-shot anomaly detection. With VLMs, we integrate vision and language modalities to encode and caption video streams, enabling flexible and context-rich anomaly detection without surveillance-specific training. This dual-modality approach enhances robustness across varied real-world settings. Evaluated on the UCF -Crime dataset, VEZA achieved up to 89% in open-vocabulary anomaly detection and video-text retrieval tasks, with 45% R @ 1 on the Ubnormal dataset as an additional evaluation. These results demonstrate VEZA's ability to identify and retrieve a wide variety of abnormal events without the need for prior training on surveillance-specific scenarios. Younes Kebour, Smaïl Niar, Nacim Ihaddadene, Abdelghani Bekrar, Hammouda Elbez |
CBMI | 2 |
| 2025 | Machine Learning Vulnerabilities in 6G: Adversarial Attacks and Their Impact on Channel Gain Prediction and Resource Allocation in UC-CFmMIMO
Mahmoud Ghorbel, Selina Cheggour, Valeria Loscrì, Youcef Imine, Hamza Ouarnoughi, Smaïl Niar |
ESORICS (1) | 6 |
| 2024 | NeuraSearchLib: A Modular and Extensible Library for Neural Architecture Search with Configurable Search SpacesabstractWhile Deep Neural Networks (DNN) have driven technological innovation, designing new architectures remains labor-intensive, requiring human expertise and numerous trial-and-error iterations. Neural Architecture Search (NAS) aims to automate this process for more efficient exploration of DNN architectures. However, current NAS libraries lack flexibility, particularly in DNN search space customization, and lack universal, ready-to-use templates for state-of-the-art NAS search spaces. This paper introduces NeuraSearchLib, a modular and extensible NAS library that allows users to create and customize search spaces with high flexibility using simple high-level specifications. NeuraSearchLib1can reproduce existing NAS search spaces, create new ones, and conduct experiments with various search strategies, training methods, and performance evaluations, enabling researchers to explore new DNN architectures more efficiently. Kaouthar Essaheli, Selsabil Roubi, Halima Bouzidi, Hamza Ouarnoughi, Smaïl Niar |
BDCAT | 5 |
| 2024 | Hardware Acceleration of Capsule Networks for Real-Time ApplicationsabstractCapsule networks (CapsNet) are a category of deep learning neural networks (DNN) that address one of the main issues and deficiencies of convolutional neural networks (CNN); loss of spatial information in pooling layers. However, the main concern with CapsNets is their compute-intensive nature, which is mainly related to the vector-based calculations during dynamic routing and acts as a barrier for their deployment in real-time applications. To address the specific computing requirements of dynamic routing in capsule layers of CapsN ets, we develop a hardware accelerator for dynamic routing using Vitis HLS. In this paper, we present a hardware acceleration solution for capsule networks by integrating AMD Xilinx deep processing unit (DPU) and a custom accelerator for the capsule layer. Our results show significant improvement in throughput compared with the baseline implementation when CapsNet is implemented on Zynq UltraScale+ MPSoC ZCU102 usinz Vitis AI DPUs. Maryam Hemmati, Earlene Starling Babette, Julia Shan, Morteza Biglari-Abhari, Smaïl Niar |
DSD | 5 |
| 2024 | Accelerated NAS via Pretrained Ensembles and Multi-fidelity Bayesian Optimization
Houssem Ouertatani, Cristian Maxim, Smaïl Niar, El-Ghazali Talbi |
ICANN (1) | 3 |
| 2024 | Analysis of Parallel Graph ApplicationsabstractDespite the increasing computing power of shared memory systems with high core counts, parallel graph processing frameworks cannot exploit it effectively. The reason behind this is the inherent challenges in parallel graph algorithms, which are efficient management of dynamically created tasks and irregular data access patterns. In this paper, we categorize several popular design choices into three design dimensions: (i) execution mode, (ii) data access pattern, and (iii) work activation. We provide their high-level parallel implementations and analyze various implementations of three representative iterative graph algorithms by considering these design dimensions. To gain a better understanding of design choices, we examine their impacts on performance, communication, scalability, and work efficiency. We also investigate the communication characteristics of the design choices on two state-of-the-art shared-memory platforms by performing micro-architectural analysis. Our microarchitectural analysis reveals that a topology-driven, pull-based model gives up to $20 x$ better performance. Funda Atik, Serif Yesil, Hamza Ouarnoughi, Smaïl Niar, Ozcan Ozturk 0001 |
ICPADS | 4 |
| 2024 | DfuseNAS: A Diffusion-Based Neural Architecture SearchabstractDeep Learning (DL) has revolutionized numerous domains by crafting highly effective Neural Network (NN) architectures. However, manual engineering approaches for NN design often yield sub-optimal solutions. Neural Architecture Search (NAS) addresses this issue by automating the design process and discovering state-of-the-art architectures. Nevertheless, the exploration of vast search spaces in NAS remains a challenging endeavor. Diffusion models have proven effective in traversing expansive search spaces encountered in generative image tasks. Their innate ability to compress and explore these spaces has shown a lot of promise. Building upon this inspiration, we introduce DfuseNAS, a novel NAS methodology rooted in diffusion processes. DfuseNAS brings substantial improvements in both NAS search efficiency and the quality of the generated neural network architectures. To the best of our knowledge, our work marks a pioneering effort in applying diffusion algorithms to enhance the search space exploration with NAS. Our experimental results, conducted on the widely-used NAS-Bench-101, showcase the remarkable capabilities of DfuseNAS. We achieved the highest average accuracy, outperforming other state-of-the-art methods, while completing the search process at least 2 times faster. Moreover, when provided with a specific architecture and a given task, the application of DfuseNAS consistently led to the generation of more accurate architectures in 98% of the times. Lotfi Abdelkrim Mecharbat, Hadjer Benmeziane, Hamza Ouarnoughi, Smaïl Niar, Kaoutar El Maghraoui |
IJCNN | 4 |
| 2024 | Grassroots operator search for model edge adaptation using mathematical search space
Hadjer Benmeziane, Kaoutar El Maghraoui, Hamza Ouarnoughi, Smaïl Niar |
Future Gener. Comput. Syst. | 4 |
| 2023 | Harmonic-NAS: Hardware-Aware Multimodal Neural Architecture Search on Resource-constrained Devices
Mohamed Imed Eddine Ghebriout, Halima Bouzidi, Smaïl Niar, Hamza Ouarnoughi |
ACML | 3 |
| 2023 | Map-and-Conquer: Energy-Efficient Mapping of Dynamic Neural Nets onto Heterogeneous MPSoCsabstractHeterogeneous MPSoCs comprise diverse processing units of varying compute capabilities. To date, the mapping strategies of neural networks (NNs) onto such systems are yet to exploit the full potential of processing parallelism, made possible through both the intrinsic NNs’ structure and underlying hardware composition. In this paper, we propose a novel framework to effectively map NNs onto heterogeneous MPSoCs in a manner that enables them to leverage the underlying processing concurrency. Specifically, our approach identifies an optimal partitioning scheme of the NN along its ‘width’ dimension, which facilitates deployment of concurrent NN blocks onto different hardware computing units. Additionally, our approach contributes a novel scheme to deploy partitioned NNs onto the MPSoC as dynamic multi-exit networks for additional performance gains. Our experiments on a standard MPSoC platform have yielded dynamic mapping configurations that are 2.1x more energy-efficient than the GPU-only mapping while incurring 1.7x less latency than DLA-only mapping. Halima Bouzidi, Mohanad Odema, Hamza Ouarnoughi, Smaïl Niar, Mohammad Abdullah Al Faruque |
DAC | 4 |
| 2023 | HADAS: Hardware-Aware Dynamic Neural Architecture Search for Edge Performance ScalingabstractDynamic neural networks (DyNNs) have become viable techniques to enable intelligence on resource-constrained edge devices while maintaining computational efficiency. In many cases, the implementation of DyNNs can be sub-optimal due to its underlying backbone architecture being developed at the design stage independent of both: (i) potential support for dynamic computing, e.g. early exiting, and (ii) resource efficiency features of the underlying hardware, e.g., dynamic voltage and frequency scaling (DVFS). Addressing this, we present HADAS, a novel Hardware-Aware Dynamic Neural Architecture Search framework that realizes DyNN architectures whose backbone, early exiting features, and DVFS settings have been jointly optimized to maximize performance and resource efficiency. Our experiments using the CIFAR-100 dataset and a diverse set of edge computing platforms have shown that HADAS can elevate dynamic models' energy efficiency by up to 57% for the same level of accuracy scores. Our code is available at https://github.com/HalimaBouzidi/HADAS Halima Bouzidi, Mohanad Odema, Hamza Ouarnoughi, Mohammad Abdullah Al Faruque, Smaïl Niar |
DATE | 5 |
| 2023 | Pareto Rank-Preserving Supernetwork for Hardware-Aware Neural Architecture SearchabstractIn neural architecture search (NAS), training every sampled architecture is very time-consuming and should be avoided. Weight-sharing is a promising solution to speed up the evaluation process. However, training the supernetwork incurs many discrepancies between the actual ranking and the predicted one. Additionally, efficient deep-learning engineering processes require incorporating realistic hardware-performance metrics into the NAS evaluation process, also known as hardware-aware NAS (HW-NAS). In HW-NAS, estimating task-specific performance and hardware efficiency are both required. This paper proposes a supernetwork training methodology that preserves the Pareto ranking between its different subnetworks resulting in more efficient and accurate neural networks for a variety of hardware platforms. The results show a 97% near Pareto front approximation in less than 2 GPU days of search, which provides 2x speed up compared to state-of-the-art methods. We validate our methodology on NAS-Bench-201, DARTS, and ImageNet. Our optimal model achieves 77.2% accuracy (+1.7% compared to baseline) with an inference time of 3.68ms on Edge GPU for ImageNet, which yields a 2.3x speedup. Training implementation can be found: https://github.com/IHIaadj/PRP-NAS. Hadjer Benmeziane, Kaoutar El Maghraoui, Hamza Ouarnoughi, Smaïl Niar |
ECAI | 4 |
| 2023 | FLASH-RL: Federated Learning Addressing System and Static Heterogeneity using Reinforcement LearningabstractWe propose FLASH-RL, a framework utilizing Double Deep Q-Learning (DDQL) to address system and static heterogeneity in Federated Learning (FL). FLASH-RL introduces a new reputation-based utility function to evaluate client contributions based on their current and past performances. Additionally, an adapted DDQL algorithm is proposed to expedite the learning process. Experimental results on MNIST and CIFAR-10 datasets demonstrate that FLASH-RL strikes a balance between model performance and end-to-end latency, reducing latency by up to 24.83% compared to FedAVG and 24.67% compared to FAVOR. It also reduces training rounds by up to 60.44% compared to FedAVG and 76% compared to FAVOR. Similar improvements are observed on the MobiAct Dataset for fall detection, underscoring the real-world applicability of our approach. Sofiane Bouaziz, Hadjer Benmeziane, Youcef Imine, Leila Hamdad, Smaïl Niar, Hamza Ouarnoughi |
ICCD | 5 |
| 2023 | Multi-objective Hardware-aware Neural Architecture Search with Pareto Rank-preserving Surrogate ModelsabstractDeep learning (DL) models such as convolutional neural networks (ConvNets) are being deployed to solve various computer vision and natural language processing tasks at the edge. It is a challenge to find the right DL architecture that simultaneously meets the accuracy, power, and performance budgets of such resource-constrained devices. Hardware-aware Neural Architecture Search (HW-NAS) has recently gained steam by automating the design of efficient DL models for a variety of target hardware platforms. However, such algorithms require excessive computational resources. Thousands of GPU days are required to evaluate and explore an architecture search space such as FBNet [ 45 ]. State-of-the-art approaches propose using surrogate models to predict architecture accuracy and hardware performance to speed up HW-NAS. Existing approaches use independent surrogate models to estimate each objective, resulting in non-optimal Pareto fronts. In this article, HW-PR-NAS, 1 a novel Pareto rank-preserving surrogate model for edge computing platforms, is presented. Our model integrates a new loss function that ranks the architectures according to their Pareto rank, regardless of the actual values of the various objectives. We employ a simple yet effective surrogate model architecture that can be generalized to any standard DL model. We then present an optimized evolutionary algorithm that uses and validates our surrogate model. Our approach has been evaluated on seven edge hardware platforms from various classes, including ASIC, FPGA, GPU, and multi-core CPU. The evaluation results show that HW-PR-NAS achieves up to 2.5× speedup compared to state-of-the-art methods while achieving 98% near the actual Pareto front. Hadjer Benmeziane, Hamza Ouarnoughi, Kaoutar El Maghraoui, Smaïl Niar |
ACM Trans. Archit. Code Optim. | 4 |
| 2023 | MaGNAS: A Mapping-Aware Graph Neural Architecture Search Framework for Heterogeneous MPSoC DeploymentabstractGraph Neural Networks (GNNs) are becoming increasingly popular for vision-based applications due to their intrinsic capacity in modeling structural and contextual relations between various parts of an image frame. On another front, the rising popularity of deep vision-based applications at the edge has been facilitated by the recent advancements in heterogeneous multi-processor Systems on Chips (MPSoCs) that enable inference under real-time, stringent execution requirements. By extension, GNNs employed for vision-based applications must adhere to the same execution requirements. Yet contrary to typical deep neural networks, the irregular flow of graph learning operations poses a challenge to running GNNs on such heterogeneous MPSoC platforms. In this paper, we propose a novel unified design-mapping approach for efficient processing of vision GNN workloads on heterogeneous MPSoC platforms. Particularly, we develop MaGNAS, a mapping-aware Graph Neural Architecture Search framework. MaGNAS proposes a GNN architectural design space coupled with prospective mapping options on a heterogeneous SoC to identify model architectures that maximize on-device resource efficiency. To achieve this, MaGNAS employs a two-tier evolutionary search to identify optimal GNNs and mapping pairings that yield the best performance trade-offs. Through designing a supernet derived from the recent Vision GNN (ViG) architecture, we conducted experiments on four (04) state-of-the-art vision datasets using both ( i ) a real hardware SoC platform (NVIDIA Xavier AGX) and ( ii ) a performance/cost model simulator for DNN accelerators. Our experimental results demonstrate that MaGNAS is able to provide 1.57 × latency speedup and is 3.38 × more energy-efficient for several vision datasets executed on the Xavier MPSoC vs. the GPU-only deployment while sustaining an average 0.11% accuracy reduction from the baseline. Mohanad Odema, Halima Bouzidi, Hamza Ouarnoughi, Smaïl Niar, Mohammad Abdullah Al Faruque |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2022 | CaW-NAS: Compression Aware Neural Architecture SearchabstractWith the ever-growing demand for deep learning (DL) at the edge, building small and efficient DL architectures has become a significant challenge. Optimization techniques such as quantization, pruning or hardware-aware neural architecture search (HW-NAS) have been proposed. In this paper, we present an efficient HW-NAS; Compression-Aware Neural Architecture search (CaW-NAS), that combines the search for the architecture and its quantization policy. While former works search over a fully quantized search space, we define our search space with quantized and non-quantized architectures. Our search strategy finds the best trade-off between accuracy and latency according to the target hardware. Experimental results on a mobile platform show that, our method allows to obtain more efficient networks in terms of accuracy, execution time and energy consumption when compared to the state of the art. Hadjer Benmeziane, Hamza Ouarnoughi, Smaïl Niar, Kaoutar El Maghraoui |
DSD | 3 |
| 2022 | Co-Optimization of DNN and Hardware Configurations on Edge GPUsabstractThe ever-increasing complexity of both Deep Neural Networks (DNN) and hardware accelerators has made the co-optimization of these domains extremely complex. Previous works typically focus on optimizing DNNs given a fixed hardware configuration or optimizing a specific hardware architecture given a fixed DNN model. Recently, the importance of the joint exploration of the two spaces drew more and more attention. Our work targets the co-optimization of DNN and hardware configurations on edge GPU accelerators. We propose an evolutionary-based co-optimization strategy by considering three metrics: DNN accuracy, execution latency, and power consumption. By combining the two search spaces, a larger number of configurations can be explored in a short time interval. In addition, a better tradeoff between DNN accuracy and hardware efficiency can be obtained. Experimental results show that the co-optimization outperforms the optimization of DNN for fixed hardware configuration with up to 53% hardware efficiency gains with the same accuracy and inference time. Halima Bouzidi, Hamza Ouarnoughi, Smaïl Niar, El-Ghazali Talbi, Abdessamad Ait El Cadi |
DSD | 3 |
| 2022 | Pareto Rank Surrogate Model for Hardware-aware Neural Architecture SearchabstractHardware-aware Neural Architecture Search (HWNAS) has recently gained much attention by automating the design of efficient deep learning models with tiny resources and reduced inference time requirements. However, HW-NAS inherits and exacerbates the expensive computational complexity of general NAS due to its significantly increased search spaces and more complex NAS evaluation component. To speed up HWNAS, existing efforts use surrogate models to predict a neural architecture’s accuracy and hardware performance on a specific platform. Thereby reducing the expensive training process and significantly reducing search time. We show that using multiple surrogate models to estimate the different objectives does not achieve the true Pareto front. Therefore, we propose HW-PRNAS, a novel Pareto Rank-preserving surrogate model. HWPR-NAS training is based on a new loss function that ranks the architectures according to their Pareto front. We evaluate our approach on seven different hardware platforms, including ASIC, FPGA, GPU and multi-cores. Our results show that we can achieve up to 2. 5x speedup while achieving better Pareto-front results than state of the art surrogate models. Hadjer Benmeziane, Smaïl Niar, Hamza Ouarnoughi, Kaoutar El Maghraoui |
ISPASS | 2 |
| 2022 | Reducing the fault vulnerability of hard real-time systems
Fabien Bouquillon, Smaïl Niar, Giuseppe Lipari |
J. Syst. Archit. | 2 |
| 2022 | Performance Modeling of Computer Vision-based CNN on Edge GPUsabstractConvolutional Neural Networks (CNNs) are currently widely used in various fields, particularly for computer vision applications. Edge platforms have drawn tremendous attention from academia and industry due to their ability to improve execution time and preserve privacy. However, edge platforms struggle to satisfy CNNs’ needs due to their computation and energy constraints. Thus, it is challenging to find the most efficient CNN that respects accuracy, time, energy, and memory footprint constraints for a target edge platform. Furthermore, given the size of the design space of CNNs and hardware platforms, performance evaluation of CNNs entails several efforts. Consequently, designers need tools to quickly explore large design space and select the CNN that offers the best performance trade-off for a set of hardware platforms. This article proposes a Machine Learning (ML)–based modeling approach for CNN performances on edge GPU-based platforms for vision applications. We implement and compare five of the most successful ML algorithms for accurate and rapid CNN performance predictions on three different edge GPUs in image classification. Experimental results demonstrate the robustness and usefulness of our proposed methodology. For three of the five ML algorithms — XGBoost, Random Forest, and Ridge Polynomial regression — average errors of 11%, 6%, and 8% have been obtained for CNN inference execution time, power consumption, and memory usage, respectively. Halima Bouzidi, Hamza Ouarnoughi, Smaïl Niar, Abdessamad Ait El Cadi |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2021 | Performance prediction for convolutional neural networks on edge GPUsabstractEdge computing is increasingly used for Artificial Intelligence (AI) purposes to meet latency, privacy, and energy challenges. Convolutional Neural networks (CNN) are more frequently deployed on Edge devices for several applications. However, due to their constrained computing resources and energy budget, Edge devices struggle to meet CNN's latency requirements while maintaining good accuracy. It is, therefore, crucial to choose the CNN with the best accuracy and latency trade-off while respecting hardware constraints. This paper presents and compares five of the widely used Machine Learning (ML) based approaches to predict CNN's inference execution time on Edge GPUs. For these 5 methods, in addition to their prediction accuracy, we also explore the time needed for their training and their hyperparameters' tuning. Finally, we compare times to run the prediction models on different platforms. The use of these methods will highly facilitate design space exploration by quickly providing the best CNN on a target Edge GPU. Experimental results show that XGBoost provides an interesting average prediction error even for unexplored and unseen CNN architectures. Random Forest depicts comparable accuracy but needs more effort and time to be trained. The other 3 approaches (OLS, MLP, and SVR) are less accurate for CNN performance estimation. Halima Bouzidi, Hamza Ouarnoughi, Smaïl Niar, Abdessamad Ait El Cadi |
CF | 3 |
| 2021 | Hardware-Aware Neural Architecture Search: Survey and TaxonomyabstractThere is no doubt that making AI mainstream by bringing powerful, yet power hungry deep neural networks (DNNs) to resource-constrained devices would required an efficient co-design of algorithms, hardware and software. The increased popularity of DNN applications deployed on a wide variety of platforms, from tiny microcontrollers to data-centers, have resulted in multiple questions and challenges related to constraints introduced by the hardware. In this survey on hardware-aware neural architecture search (HW-NAS), we present some of the existing answers proposed in the literature for the following questions: "Is it possible to build an efficient DL model that meets the latency and energy constraints of tiny edge devices?", "How can we reduce the trade-off between the accuracy of a DL model and its ability to be deployed in a variety of platforms?". The survey provides a new taxonomy of HW-NAS and assesses the hardware cost estimation strategies. We also highlight the challenges and limitations of existing approaches and potential future directions. We hope that this survey will help to fuel the research towards efficient deep learning. Hadjer Benmeziane, Kaoutar El Maghraoui, Hamza Ouarnoughi, Smaïl Niar, Martin Wistuba, Naigang Wang |
IJCAI | 4 |
| 2021 | Railway Obstacle Detection Using Unsupervised Learning: An Exploratory StudyabstractAutonomous Driving (AD) systems are heavily reliant on supervised models. In these approaches, a model is trained to detect only a predefined number of obstacles. However, for applications like railway obstacle detection, the training dataset is limited and not all possible obstacle classes are known beforehand. For such safety-critical applications, this situation is problematic and could limit the performance of obstacle detection in autonomous trains. In this paper, we propose an exploratory study using unsupervised models based on a large set of generated convolutional autoencoder models to detect obstacles on railway's track level. The study was conducted based on three components: loss functions, activations and optimizers. Existing works rely on fixing thresholds to judge the performance of the model. We propose instead a methodology based on Multi-Criteria Decision Making (MCDM) to evaluate the performance of all models. Furthermore, we introduce the notion of gap-score to evaluate each model by calculating the average difference between the reconstruction score on images with and without obstacles. The aim is to find models maximizing the average of gap-scores and rank them according to their performances. Experimental results show that the evaluated models can provide up to 68 % average gap-score. Amine Boussik, Waël Ben-Messaoud, Smaïl Niar, Abdelmalik Taleb-Ahmed |
IV | 3 |
| 2020 | A Multi-Agent Approach for Vehicle-to-Fog Fair Computation OffloadingabstractFuture autonomous driving (AD) will require more data processing and storage capacities, exceeding the in-board capacities, especially with AI application requirements. The cloud or fog resources can provide off-load services. But to obtain an optimized task distribution between the on-board edge computing platform and cloud or Fog resources, communication delays and AD real-time constraints must be taken into account. The equity in resource allocation is rarely studied in this context where mobility adds a specific difficulty. Our approach of offloading mechanism is to leverage intelligent agents able to make decisions on task delegation based on past decisions. Agents at edge and fog levels communicate and exchange their knowledge and past decisions. A first scenario illustrates the proposition with simulated data. Emmanuelle Grislin, Hamza Ouarnoughi, Smaïl Niar |
AICCSA | 3 |
| 2020 | Pedestrian Detection and Classification for Autonomous TrainabstractIn this paper, we present a combined approach for human localization and classification in Autonomous Train application. Our contribution is threefold. (a) The creation of a new dataset for workers wearing orange vests in a railway environment context. (b) A deep learning supervised YOLO object detector for persons detection combined with a linear SVM (Support Vector Machine) classifier for persons classification into workers wearing orange vests or travelers. (c) A realtime vision-based technique for the environment monitoring in a driverless train application. Experimental results evaluate the parameters of our two stages detection approach and show that our algorithm is robust in detecting and classifying railway workers for a real-time implementation on an embedded system. Our implementation on an embedded system allows a detection with a correct classification rate of 98.5 % of accuracy and a classification time of 1 ms per frame. Ankur Mahtani, Waël Ben-Messaoud, Abdelmalik Taleb-Ahmed, Smaïl Niar, Clément Strauss |
IPAS | 4 |
| 2020 | A GPU enhanced LIDAR Perception System for Autonomous VehiclesabstractEnvironment vision and understanding is a crucial task in Autonomous Driving (AD) context. This mainly needs image processing approaches such as Convolutional Neural Networks (CNN). Nevertheless, cameras have shown their limits for such a task, especially in dealing with difficult light conditions. LIDAR is a powerful and widely used sensor for AD. Indeed, LIDAR can then cope with the lack of information gathered from cameras. For AD, data processing from the sensors is the key function to obtain a high quality perception. For this, Graphics Processing Unit (GPU) platforms show great performances and outperform other processing platforms such as FPGA and Multi-cores. This work presents a new approach to produce multiple 2D representation from 3D points cloud coming from LIDAR. The 2D representation can therefore be used by any efficient image processing applications. Our approach uses only LIDAR sensor and exploits the high GPU parallelism for its implementation. The resulting 2D representations are then used by CNN for AD applications such as image classification and segmentation. Finally, our contributions have been evaluated using the KITTI road benchmark and showed encouraging results. Abderrahim Haneche, Mohammed Yazid Lachachi, Smaïl Niar, Hamza Ouarnoughi |
PDP | 3 |
| 2020 | Preemption-Aware Allocation, Deadline Assignment for Conditional DAGs on Partitioned EDFabstractHeterogeneous hardware platforms are often used for implementing complex critical real-time applications, like Advanced driver-assistance systems (ADAS) and autonomous driving. Typically, they are composed of CPU hosts and a set of accelerators. To better support real-time workloads, several hardware accelerators have evolved to allow preemption for computationally intensive tasks, such as GPUs. However, their preemption costs can be very high compared to classical CPU preemption, and therefore must be taken into account at design time and in the scheduling analysis. In this paper, we address mainly two tightly correlated problems: (i) task allocation for a set of real-time tasks, modeled by conditional directed acyclic graphs (C-DAG), onto multiprocessor platforms under partitioned preemptive Earliest Deadline First scheduling, assuming a non-negligible cost of preemption, and (ii) intermediate deadlines and offsets assignments to real-time C-DAGs, so to remove unnecessary preemption and reduce the total preemption overhead. The effectiveness of the proposed technique is evaluated using a large set of synthetic tasks sets. Houssam-Eddine Zahaf, Giuseppe Lipari, Smaïl Niar, Abou El Hassan Benyamina |
RTCSA | 3 |
| 2019 | A new memory reliability technique for multiple bit upsets mitigationabstractTechnological advances make it possible to produce increasingly complex electronic components. Nevertheless, these advances are convoyed by an increasing sensitivity to operating conditions and an accelerated aging process. In safety critical applications, it is vital to provide solutions to avoid these limitations and to guarantee a high level of reliability. In most of the existing methods in the literature only Single Event Upsets (SEU) are assumed. The next generations of embedded systems must on one side support Multiple-Bit Upsets (MBU) and avoid to induce a significant memory and processing overheads on the other side. This paper proposes a new method to increase the reliability of SRAM, without dramatically increasing costs in memory space and processing time. Our method, named DPSR for Double Parity Single Redundancy, offers a high level of reliability and takes into fault patterns occurring in real conditions. Alexandre Chabot, Ihsen Alouani, Smaïl Niar, Réda Nouacer |
CF | 3 |
| 2019 | Adaptive Vehicle Detection for Real-time Autonomous Driving SystemabstractModern cars are being equipped with powerful computational resources for autonomous driving systems (ADS) as one of their major parts to provide safer travels on roads. High accuracy and real-time requirements of ADS are addressed by HW/SW co-design methodology which helps in offloading the computationally intensive tasks to the hardware part. However, the limited hardware resources could be a limiting factor in complicated systems. This paper presents a dynamically reconfigurable system for ADS which is capable of real-time vehicle and pedestrian detection. Our approach employs different methods of vehicle detection in different lighting conditions to achieve better results. A novel deep learning method is presented for detection of vehicles in the dark condition where the road light is very limited or unavailable. We present a partial reconfiguration (PR) controller which accelerates the reconfiguration process on Zynq SoC for seamless detection in real-time applications. By partially reconfiguring the vehicle detection block on Zynq SoC, resource requirements is maintained low enough to allow for the existence of other functionalities of ADS on hardware which could complete their tasks without any interruption. Our presented system is capable of detecting pedestrian and vehicles in different lighting conditions at the rate of 50fps (frames per second) for HDTV (1080x1920) frame. Maryam Hemmati, Morteza Biglari-Abhari, Smaïl Niar |
DATE | 3 |
| 2019 | ENOrMOUS: ENergy Optimization for MObile plateform using User needS
Ismat Chaib Draa, Smaïl Niar, Emmanuelle Grislin, Morteza Biglari-Abhari, Jamel Tayeb |
J. Syst. Archit. | 2 |
| 2018 | Rapid in-memory matrix multiplication using associative processorabstractMemory hierarchy latency is one of the main problems that prevents processors from achieving high performance. To eliminate the need of loading/storing large sets of data, Resistive Associative Processors (ReAP) have been proposed as a solution to the von Neumann bottleneck. In ReAPs, logic and memory structures are combined together to allow inmemory computations. In this paper, we propose a new algorithm to compute the matrix multiplication inside the memory that exploits the benefits of ReAP. The proposed approach is based on the Cannon algorithm and uses a series of rotations without duplicating the data. It runs in O(n), where n is the dimension of the matrix. The method also applies to a large set of row by column matrix-based applications. Experimental results show several orders of magnitude increase in performance and reduction in energy and area when compared to the latest FPGA and CPU implementations. Mohamed A. Neggaz, Hasan Erdem Yantir, Smaïl Niar, Ahmed M. Eltawil, Fadi J. Kurdahi |
DATE | 3 |
| 2018 | A Reliability Study on CNNs for Critical Embedded SystemsabstractDeep learning systems such as Convolutional Neural Networks (CNNs) have shown remarkable efficiency in dealing with a variety of complex real life problems. To accelerate the execution of these heavy algorithms, a plethora of software implementations and hardware accelerators have been proposed. In a context of shrinking devices dimensions, reliability issues of CNN-hosting systems are under-explored. In this paper, we experimentally evaluate the inherent fault tolerance of CNNs by injecting errors within network modules, namely processing elements and memories. Our experiments demonstrate a non uniform sensitivity between different parts of the system. While CNNs are relatively resilient to errors occurring in processing elements, transient faults hitting memories lead to catastrophic degradation of accuracy. Mohamed A. Neggaz, Ihsen Alouani, Pablo R. Lorenzo, Smaïl Niar |
ICCD | 4 |
| 2018 | A Comprehensive Fault Injection Strategy for Embedded Systems Reliability AssessmentabstractThe embedded systems industry is moving towards the integration of higher performance, yet less reliable electronic components into new product generations. Technology and voltage scaling increased dramatically the susceptibility of new devices not only to Single Bit Upsets (SBU), but also to Multiple Bit Upsets (MBU). However, the system reliability assessment at the design phase of fault-tolerant computer systems is a complex and critical task. In this context, it is mandatory to enhance reliability analysis and evaluation techniques at early-stage of the system development. In this paper, we present a technique for reliability evaluation of embedded systems at early-stage by taking into account the application behavior and SBU/MBU phenomena. Instead of using the random fault injection, our approach models the architecture behavior under real working conditions. Our results demonstrate the efficiency of the proposed fault injection simulation platform for early-stage reliability studies. Alexandre Chabot, Ihsen Alouani, Smaïl Niar, Réda Nouacer |
RSP | 3 |
| 2018 | Power optimization techniques for associative processors
Hasan Erdem Yantir, Ahmed M. Eltawil, Smaïl Niar, Fadi J. Kurdahi |
J. Syst. Archit. | 3 |
| 2017 | Performance Exploration of AMBA AXI4 Bus Protocols for Wireless Sensor NetworksabstractModern System-on-Chip (SoC) designs are faced with many challenges among which efficient communication managing is one of the most important. On-chip communication architectures can have a strong impact on the performance of SoC designs. The traditional SoC interconnects, cannot keep up with the high demands of today's SoC. To address this problem, SoC makers propose new protocols to implement high performance data transfer. AMBA AXI4, is one of the widely used protocols as on-chip bus in recent Wireless Sensor Network SoCs. It includes three distinct interconnect protocols: stream, burst and lite. In this paper, we highlight the various parameters that must be taken into consideration to select the adequate interface for a given application. We analyze and compare the on-chip interfaces for hardware/software (HW/SW) communication when various code transformations are applied on the Zynq-7000 Xilinx platform. Many experiments have been conducted to evaluate the communication between the main processor and the reconfigurable hardware accelerators. Mariem Makni, Mouna Baklouti, Smaïl Niar, Mohamed Abid |
AICCSA | 3 |
| 2017 | Real-Time Multi-Scale Pedestrian Detection for Driver Assistance SystemsabstractPedestrian detection is one of the most challenging and vital tasks of driver assistance systems (DAS). Among several algorithms developed for human detection, histogram of oriented gradients (HOG) followed by support vector machine (SVM) has shown the most promising results. This paper presents a hardware accelerator for real-time pedestrian detection at different scales to fulfill the real-time requirements of DAS. It proposes an algorithmic modification to the conventional multi-scale object detection by means of HOG+SVM to increase the throughput and maintain the accuracy reasonably high. Our hardware accelerator detects pedestrians at the rate of 60 fps for HDTV (1080x1920) frame. Maryam Hemmati, Morteza Biglari-Abhari, Smaïl Niar, Stevan M. Berber |
DAC | 3 |
| 2017 | Design Space exploration of FPGA-based accelerators with multi-level parallelismabstractApplications containing compute-intensive kernels with nested loops can effectively leverage FPGAs to exploit fine-and coarse-grained parallelism. HLS tools used to translate these kernels from high-level languages (e.g., C/C--), however, are inefficient in exploiting multiple levels of parallelism automatically, thereby producing sub-optimal accelerators. Moreover, the large design space resulting from the various combinations of fineand coarse-grained parallelism options makes exhaustive design space exploration prohibitively time-consuming with HLS tools. Hence, we propose a rapid estimation framework, MPSeeker, to evaluate performance/area metrics of various accelerator options for an application at an early design phase. Experimental results show that MPSeeker can rapidly (in minutes) explore the complex design space and accurately estimate performance/area of various design points to identify the near-optimal (95.7% performance of the optimal on average) combination of parallelism options. Guanwen Zhong, Alok Prakash, Yun Liang 0001, Tulika Mitra, Smaïl Niar |
DATE | 6 |
| 2017 | Adaptive Reliability for Fault Tolerant Multicore SystemsabstractIn an era of continuously shrinking technology and escalating power density, Multiprocessor System on Chips (MPSoCs) suffer from a growing prominence of device defects and increase of dependability-related issues. This paper tackles the dependability challenge by suggesting an adaptive reliability enhancement strategy for multicore systems. We dynamically adapt the reliability enhancement to the actual tasks requirements as well as cores runtime operating conditions. As reliability improvement may adversely affect the parameters of embedded systems, we suggest a runtime recovery method. In fact, we implement a 3-mode mapping technique to limit redundancy overheads through judicious task migrating and dropping. Our experiments show promising results in terms of error mitigation with controllable power and thermal overheads. Ihsen Alouani, Thomas Wild, Andreas Herkersdorf, Smaïl Niar |
DSD | 4 |
| 2017 | An Energy-Aware Learning Agent for Power Management in Mobile Devices
Ismat Chaib Draa, Emmanuelle Grislin, Smaïl Niar |
IEA/AIE (1) | 3 |
| 2017 | A Rapid Data Communication Exploration Tool for Hybrid CPU-FPGA ArchitecturesabstractModern System-on-Chip (SoC) designs face many challenges. Choosing the best communication protocol among the different processing nodes is one of the most important design decisions. On-chip communication architectures can have a significant impact on the performance of SoC designs. However, in most of the existing design tools, only the computation cost is accurately estimated. To address this challenge, we present a high-level analytical tool to estimate the data communication cost for hybrid CPU-FPGA architectures. The proposed model allows to estimate, rapidly and accurately, both computation and communication cost of applications containing multiple nested loops. This paper also explores the benefits of applying various optimization pragmas including dataflow and loop pipelining, at the compilation phase. Experimental results show that the proposed model provides accurate data communication estimation for hybrid CPUFPGA architectures. Mariem Makni, Smaïl Niar, Mouna Baklouti, Guanwen Zhong, Tulika Mitra, Mohamed Abid |
PDP | 2 |
| 2016 | Lin-analyzer: a high-level performance analysis tool for FPGA-based acceleratorsabstractThe increasing complexity of FPGA-based accelerators, coupled with time-to-market pressure, makes high-level synthesis (HLS) an attractive solution to improve designer productivity by abstracting the programming effort above register-transfer level (RTL). HLS offers various architectural design options with different trade-offs via pragmas (loop unrolling, loop pipelining, array partitioning). However, non-negligible HLS runtime renders manual or automated HLS-based exhaustive architectural exploration practically infeasible. To address this challenge, we present Lin-Analyzer, a high-level accurate performance analysis tool that enables rapid design space exploration with various pragmas for FPGA-based accelerators without requiring RTL implementations. Guanwen Zhong, Alok Prakash, Yun Liang 0001, Tulika Mitra, Smaïl Niar |
DAC | 5 |
| 2016 | NS-SRAM: Neighborhood Solidarity SRAM for Reliability Enhancement of SRAM MemoriesabstractTechnology shift and voltage scaling increased the susceptibility of Static Random Access Memories (SRAMs) to errors dramatically. In this paper, we present NS-SRAM, for Neighborhood Solidarity SRAM, a new technique to enhance error resilience of SRAMs by exploiting the adjacent memory bit data. Bit cells of a memory line are paired together in circuit level to mutually increase the static noise margin and critical charge of a cell. Unlike existing techniques, NS-SRAM aims to enhance both Bit Error Rate (BER) and Soft Error rate (SER) at the same time. Due to auto-adaptive joiners, each of the adjacent cells' nodes is connected to its counterpart in the neighbor bit. NS-SRAM enhances read-stability by increasing critical Read Static Noise Margin (RSNM), thereby decreasing faults when circuit operates under voltage scaling. It also increases hold-stability and critical charge to mitigate soft-errors. By the proposed technique, reliability of SRAM based structures such as cache memories and register files can drastically be improved with comparable area overhead to existing hardening techniques. Moreover it does not require any extra-memory, does not impact the memory effective size, and has no negative impact on performance. Ihsen Alouani, Hamzeh Ahangari, Ozcan Ozturk 0001, Smaïl Niar |
DSD | 4 |
| 2016 | Device Context Classification for Mobile Power Consumption ReductionabstractThe diverse range of wireless interfaces, sensors, processing components added to the increasing popularity of power-hungry applications reduce the battery life of mobile devices. This paper proposes a tool for identifying the device context, understanding the user habits and preferences in order to adjust available resources and find trade-off between the power consumption and the user satisfaction. We use Machine Learning (ML) methods to identify and classify user/device contexts. On this basis, a software is developed to control at run-time system component activities. When applied only for the screen brightness level knob, the proposed solution can lower the power consumption by up to 20% vs. the out-of-the-box OS brightness manager with a negligible energy overhead. Ismat Chaib Draa, Maroua Nouiri, Smaïl Niar, Abdelghani Bekrar |
DSD | 3 |
| 2015 | Enhanced Quality Using Intensive Test and Analysis on SimulatorsabstractEmbedded systems are becoming ubiquitous and are subject to demanding standards in both safety and reliability. Modern vehicles, which must respect ISO 26262 standards, use up to 100 Electronic Control Unit (ECUs). Advances in microelectronics enable integration of more functions in the ECU, but at the cost of greater unreliability in hostile operating environments, such as electromagnetic fields, temperature, and humidity. Their software mainly drives embedded system flexibility and smartness. However due to lack of automation, its validation and verification (V&V) takes place throughout the design process and tends to swallow up 40% to 50% of the total development cost. The “Enhanced Quality Using Intensive Test Analysis on Simulators” (EQUITAS) project intends to limit the impact of software V&V on embedded systems cost and time-to-market while improving reliability and functional safety. Project activities include: development of a continuous tool-chain to automate the V&V process of embedded computers, improving the relevance of the test campaigns by detecting the redundant tests using equivalence classes, providing assistance for hardware failure effect analysis (FMEA), and finally assessing the tool-chain under the ISO 26262 requirements. Réda Nouacer, Manel Djemal, Smaïl Niar, Gilles Mouchard, Nicolas Rapin, Jean-Pierre Gallois, Philippe Fiani, Francois Chastrette, Toni Adriano, Bryan MacEachen |
DSD | 3 |
| 2015 | A multi-objective approach for software/hardware partitioning in a multi-target tracking systemabstractHeterogeneous Multiprocessor System-on-Chips (MPSoCs) are getting increasingly used to cope with new embedded applications performance requirements. In such systems, the promising cohabitation of processing elements (PEs) having different aspects allows designers to better exploit the synergy of hardware and software cores. Software/Hardware partitioning investigates the design of MPSOCs to take advantage of software flexibility and hardware high performance with the lowest possible costs. Signal-processing-oriented systems handle huge amounts of data and consequently demand highly performant architecture. In this paper, we propose a Software/Hardware partitioning approach for high speed reconfigurable DSP-oriented embedded systems. We present two multi-objective techniques aiming at exploring the partitioning configurations that minimize execution time, resource utilization and time to market of the MPSoC. Ihsen Alouani, Braham Lotfi Mediouni, Smaïl Niar |
RSP | 3 |
| 2014 | Fast System Level Benchmarks for Multicore ArchitecturesabstractWe present a framework that automatically generates system level synthetic benchmarks from traditional benchmarks. Synthetic benchmarks have similar performance behavior as the original benchmarks that they are generated from and they can run faster. Synthetics can also be used as proxies where original applications are not available in source form. In experiments we observe that not only are our system level benchmarks much smaller than the real benchmarks that they are generated from but they are also much faster. For example, when we generate synthetic benchmarks from the well-known multicore benchmark suite, PARSEC, our benchmarks have an average speedup of 149x over PARSEC benchmarks. We also observe that the performance behavior of synthetics have more than 85% similarity to the real benchmarks. Alper Sen 0001, Gökçehan Kara, Etem Deniz, Smaïl Niar |
DSD | 4 |
| 2014 | Design Space Exploration for Customized Asymmetric Heterogeneous MPSoCabstractModern FPGA allows the design of very complex System-on-Chips (SoC). To fulfil modern application requirements, in terms of performance/energy consumption ratio, Heterogeneous Multiprocessor System-on-Chip (Ht- MPSoC) architectures represent a promising solution. In such systems, the processor instruction set is enhanced by application-specific custom instructions implemented on reconfigurable fabrics, namely FPGA. To increase area utilization and guarantee application constraint respect, we propose a new Ht-MPSoC architecture where hardware accelerators (HW accelerators) are shared among different processors in an intelligent manner. In this paper, we extend existing Ht-MPSoC architectures by considering asymmetric (AHt-MPSoC). In these architectures, cores have different resources that may share in different manners. Depending on the running applications and their needs in processing, private and shared HW accelerators are attached to the different cores. On a 8-core AHt-MPSoC we obtained a speed of 2.6 with a reduced number of HW accelerators for our benchmarks. Bouthaina Damak, Rachid Benmansour, Mouna Baklouti, Smaïl Niar, Mohamed Abid |
DSD | 4 |
| 2014 | HOG Feature Extractor Hardware Accelerator for Real-Time Pedestrian DetectionabstractHistogram of oriented gradients (HOG) is considered as the most promising algorithm in human detection, however its complexity and intensive computational load is an issue for real-time detection in embedded systems. This paper presents a hardware accelerator for HOG feature extractor to fulfill the requirements of real-time pedestrian detection in driver assistance systems. Parallel and deep pipelined hardware architecture with special defined memory access pattern is employed to improve the throughput while maintaining the accuracy of the original algorithm reasonably high. Adoption of efficient memory access pattern, which provides simultaneous access to the required memory area for different functional blocks, avoids repetitive calculation at different stages of computation, resulting in both higher throughput and lower power. It does not impose any further resource requirements with regard to memory utilization. Our presented hardware accelerator is capable of extracting HOG features for 60 fps (frame per second) of HDTV (1080x1920) frame and could be employed with several instances of support vector machine (SVM) classifier in order to provide multiple object detection. Maryam Hemmati, Morteza Biglari-Abhari, Stevan M. Berber, Smaïl Niar |
DSD | 4 |
| 2014 | A mixed integer linear programming approach for design space exploration in FPGA-based MPSoCabstractHeterogeneous Multiprocessor System-on-Chip (Ht-MPSoC) architectures represent a promising approach as they allow a higher performance/energy consumption trade-off. In such systems, the processor instruction set is enhanced by application-specific custom instructions implemented on reconfigurable fabrics, namely FPGA. To increase area utilization and guarantee application constraint respect, we propose a new architecture where Ht-MPSoC hardware accelerators are shared among different processors in an intelligent manner. In this paper, a Mixed Integer Linear Programming (MILP) model is proposed to systematically explore the complex design space of the different configurations. Bouthaina Damak, Rachid Benmansour, Smaïl Niar, Mouna Baklouti, Mohamed Abid |
FPL | 3 |
| 2014 | Application specific multi-port memory customization in FPGAsabstractFPGA block RAMs (BRAMs) offer speed advantages compared to LUT-based memory designs but a BRAM has only one read and one write port. Designers need to use multiple BRAMs in order to create multi-port memory structures which are more difficult than designing with LUT-based multiport memories. Multi-port memory designs increase overall performance but comes with area cost. In this paper, we present a fully automated methodology that tailors our multi-port memory from a given application. We present our performance improvements and area tradeoffs on state-of-the-art string matching algorithms. Gorker Alp Malazgirt, Hasan Erdem Yantir, Arda Yurdakul, Smaïl Niar |
FPL | 4 |
| 2014 | Design space exploration of multiple loops on FPGAs using high level synthesisabstractReal-world applications such as image processing, signal processing, and others often contain a sequence of computation intensive kernels, each represented in the form of a nested loop. High-level synthesis (HLS) enables efficient hardware implementation of these loops using high-level programming languages. HLS tools also allow the designers to evaluate design choices with different trade-offs through pragmas/directives. Prior design space exploration techniques for HLS primarily focus on either single nested loop or multiple loops without consideration to the data dependencies among them. In this paper, we propose efficient design space exploration techniques for applications that consist of multiple nested loops with or without data dependencies. In particular, we develop an algorithm to derive the Pareto-optimal curve (performance versus area) of the application when mapped onto FPGAs using HLS. Our algorithm is efficient as it effectively prunes the dominated points in the design space. We also develop accurate performance and area models to assist the design space exploration process. Experiments on various scientific kernels and real-world applications demonstrate that our design space exploration technique is accurate and efficient. Guanwen Zhong, Vanchinathan Venkataramani, Yun Liang 0001, Tulika Mitra, Smaïl Niar |
ICCD | 5 |
| 2013 | Radar signature in multiple target tracking system for driver assistant applicationabstractThis paper presents a new Driver Assistant System (DAS) using radar signatures. The new system is able in one hand to track multiple obstacles and on the other hand to identify obstacles during vehicle movements. The combination of these two functions on the same DAS gives the benefits of avoiding false alarms. Also, it makes possible to generate alarms that take into account the identification of the obstacles. The obstacle tracking process is simplified thanks to the identification stage. Hence, our low cost FPGA-based System-on-Chip is able to detect, recognize and track a large number of obstacles in a relatively reduced time period. Our experimental result proves that a speed up of 32% can be obtained compared to the standard system. Haisheng Liu, Smaïl Niar |
DATE | 2 |
| 2013 | Two-level caches tuning technique for energy consumption in reconfigurable embedded MPSoC
A. Bengueddach, B. Senouci, Smaïl Niar, Bouziane Beldjilali |
J. Syst. Archit. | 3 |
| 2012 | H.264 Macroblock Line Level Parallel Video Decoding on Embedded Multicore ProcessorsabstractThe adaptation of intensive calculation algorithms made the new emerging H.264 an efficient video codec. On the other hand, embedded processors are equipped with multicore processors, thus offering additional processing power. The H.264 codec cannot benefit from this processing power in its current state. One solution is to execute the codec on different cores concurrently. H.264 codec is a complex video compression standard that is widely used in multimedia applications. In this paper, a new parallelization technique for the H.264 decoder is proposed based on Macroblock (MB) lines distribution of a video frame on a multicore architecture. A pipeline for the Entropy Decoder (ED) at the slice level is also applied in order to speed up the processing time. Simulations conducted with High Definition (HD) resolutions show an upper limit speedup of 4.7 using the Baseline profile and 3.2 using the Main profile on a 16-core embedded processor. Elias Baaklini, Hassan Sbeity, Smaïl Niar |
DSD | 3 |
| 2012 | An efficient power estimation methodology for complex RISC processor-based platformsabstractIn this contribution, we propose an efficient power estimation methodology for complex RISC processor-based platforms. In this methodology, the Functional Level Power Analysis (FLPA) is used to set up generic power models for the different parts of the system. Then, a simulation framework based on virtual platform is developed to evaluate accurately the activities used in the related power models. The combination of the two parts above leads to a heterogeneous power estimation that gives a better trade-off between accuracy and speed. The usefulness and effectiveness of our proposed methodology is validated through ARM9 and ARM CortexA8 processor designed respectively around the OMAP5912 and OMAP3530 boards. This efficiency and the accuracy of our proposed methodology is evaluated by using a variety of basic programs to complete media benchmarks. Estimated power values are compared to real board measurements for the both ARM940T and ARM CortexA8 architectures. Our obtained power estimation results provide less than 3% of error for ARM940T processor, 3.5% for ARM CortexA8 processor-based system and 1x faster compared to the state-of-the-art power estimation tools. Santhosh Kumar Rethinagiri, Rabie Ben Atitallah, Jean-Luc Dekeyser, Eric Senn, Smaïl Niar |
ACM Great Lakes Symposium on VLSI | 5 |
| 2012 | Parity-based mono-Copy Cache for low power consumption and high reliabilityabstractThe power consumption is one of the most important preoccupations of the chip designers. However, reducing power consumption has its negative impact on the circuit. For example, reducing the supply voltage of a microprocessor implies an increase in the probability of process-variation-induced failures. Fault tolerant architectures propose a trade-off by boosting the reliability while reducing power consumption. Since a large part of the microprocessor power is consumed by the cache memory, we propose in this paper the Parity-based mono-Copy Cache (PmC2) that maintains cache reliability under aggressive voltage scaling. PmC2results in reducing energy consumption considerably with very low performance penalty. PmC2uses a parity check mechanism in error detection and only one cache block redundancy for error correction. Our experimental results demonstrate that reducing the supply voltage with roughly 25% of nominal Vdd achieves more than 62% reduction in cache power consumption with a negligible IPC loss that does not exceed 0.15%. Ihsen Alouani, Smaïl Niar, Fadi J. Kurdahi, Mohamed Abid |
RSP | 2 |
| 2011 | Hybrid system level power consumption estimation for FPGA-based MPSoCabstractThis paper proposes an efficient Hybrid System Level (HSL) power estimation methodology for FPGA-based MPSoC. Within this methodology, the Functional Level Power Analysis (FLPA) is extended to set up generic power models for the different parts of the system. Then, a simulation framework is developed at the transactional level to evaluate accurately the activities used in the related power models. The combination of the above two parts lead to a hybrid power estimation that gives a better trade-off between accuracy and speed. The proposed methodology has several benefits: it considers the power consumption of the embedded system in its entirety and leads to accurate estimates without a costly and complex material. The proposed methodology is also scalable for exploring complex embedded architectures. The usefulness and effectiveness of our HSL methodology is validated through a typical mono-processor and multiprocessor embedded system designed around the Xilinx Virtex II Pro FPGA board. Our experiments performed on an explicit embedded platform show that the obtained power estimation results are less than 1.2% of error when compared to the real board measurements and faster compared to other power estimation tools. Santhosh Kumar Rethinagiri, Rabie Ben Atitallah, Smaïl Niar, Eric Senn, Jean-Luc Dekeyser |
ICCD | 3 |
| 2010 | H.264 Color Components Video Decoding Parallelization on Multi-core ProcessorsabstractMultiprocessor-system-on-a-chip will be the dominating architecture in embedded systems as it provides an increase in concurrency improving the performance of the system rather than increasing the clock speed which affects the power consumption of the system. However, concurrency needs to be exploited in order to improve the system performance in the different applications'environments. The new emerging H.264/AVC coding standard is designed to cover a wide range of applications (real-time conversational services such as videoconferencing, video phone, etc.). It has many new features that require complex computations compared to previous video coding standards. This coding standard will be a challenging workload for future MPSoC embedded systems. Exploiting the different levels of parallelism for video codec applications can be done at the data level, the functional level, or both simultaneously. Our intention in this paper is to explore the natural existent parallelism in the H.264 decoder software itself without any modification to the encoder phase, rather than forcing parallelization techniques. Our novel idea is based on the fact that the H.264 decoder decodes the luminance and chrominance signals separately, but the decoder is implemented in a way to decode them in series. Our approach is to execute the different decoding phases of the luminance signals in parallel to the chrominance signals. Using two cores to decode the luma and the chroma signals in parallel gives a gain of 15-20% of the decoding processing time and combining them the functional pipelined implementation over four cores or more, the gain can reach 60% compared to the current sequential execution. Elias Baaklini, Hassan Sbeity, Smaïl Niar, Nouhad Amaneddine |
DSD | 3 |
| 2010 | An Improved Automotive Multiple Target Tracking System DesignabstractMultiple Target Tracking (MTT) algorithms are widely used in various military and civilian applications but its use in automotive safety has little been investigated. In MTT algorithms, implemented in embedded systems, it is important to use the minimum required resources to allow the entire DAS system to be integrated on the same chip (data acquisition, MTT and alarm restitution). This allows the reduction of the System on Chip (SoC) complexity and cost. This paper presents an efficient Driver Assistance System (DAS) based on MTT application. To do so, we first identified the performance bottlenecks in the application. In this application, a set of optimizations were applied to reduce the MTT algorithm's complexity. Tuning in conjunction the hardware and the software yielded to optimize the final system and to meet the functional requirements. The result is a complete embedded MTT application running on an embedded system that fits in a contemporary medium sized FPGA device. Naim Harb, Haisheng Liu, Smaïl Niar, Rabie Ben Atitallah |
DSD | 4 |
| 2009 | A Dynamic Hybrid Cache Coherency Protocol for Shared-Memory MPSoCabstractIn Multi-Processor System-on-Chip (MPSoC) architectures equipped with shared-memory, caches have significant impact on performance and energy consumption. Indeed, if the executed application depicts a high degree of reference locality, caches may reduce the amount of shared-memory accesses and data transfers on the interconnection network. Hence, execution time and energy consumption can be greatly optimized. However, caches in MPSoC architectures put forward the data coherency problem. In this context, most of the existing solutions are based either on data invalidation or data update protocols. These protocols do not consider the change in the application behavior. This paper presents a new hybrid cache-coherency protocol that is able to dynamically adapt its functioning mode according to the application needs. An original architecture which facilitates this protocol's implementation in Network-On-Chip based MPSoC architectures is also proposed. Performances, in terms of speed up factor and energy reduction gain of the proposed protocol, have been evaluated using a Cycle Accurate Bit Accurate (CABA) simulation platform. Experimental results in comparison with other existing solutions show that this protocol may give significant reductions in execution time and energy consumption can be achieved. Hajer Chtioui, Rabie Ben Atitallah, Smaïl Niar, Jean-Luc Dekeyser, Mohamed Abid |
DSD | 3 |
| 2008 | An MPSoC architecture for the Multiple Target Tracking application in driver assistant systemabstractThis article discusses the design of an application specific MPSoC architecture dedicated to multiple target tracking (MTT). This application has its utility in driver assistant systems, more precisely in collision avoidance and warning systems. An automotive-radar is used as the front end sensor in our application. The article examines the tradeoffs that must be taken into consideration in the realization of the entire MTT application in an embedded system. In our implementation of MTT, several independent parallel tasks have been identified and mapped onto a multiprocessor architecture to ensure the deadlines imposed by the application. Our study demonstrates that the joint utilization of reconfigurable circuits (namely FPGA) and MPSoC, facilitates the development of a flexible and efficient MTT system. Jehangir Khan, Smaïl Niar, Atika Rivenq, Yassin Elhillali, Jean-Luc Dekeyser |
ASAP | 2 |
| 2008 | Multi-granularity sampling for simulating concurrent heterogeneous applicationsabstractDetailed or cycle-accurate/bit-accurate (CABA) simulation is a critical phase in the design flow of embedded systems. However, with increasing system complexity, full detailed simulation is prohibitively slower than the hardware being simulated. In this paper, we present an approach that uses the sampling technique to speed up the design flow of Multiprocessor System-on-Chip (MPSoC) systems. Based on the dynamic behavior of the applications running concurrently, our method dynamically chooses between multiple granularities of the sampling phase. The similarities of the execution phases for all possible granularities are first analyzed, then transitions between phase overlaps are discretized. To facilitate the detection of repetitions, one phase, with an appropriate granularity, is chosen per process. Unlike most other proposals, the associated performance is usually accurate enough not to need repeated resampling. The use of checkpointing in conjunction with our approach is simplified because the amount of the needed disk space is significantly reduced. Experimental results show that the simulation of concurrent heterogeneous applications can be accelerated by a factor of up to 60x, while maintaining an average performance estimation error lower than 5%. Melhem Tawk, Khaled Z. Ibrahim, Smaïl Niar |
CASES | 3 |
| 2007 | Adaptive Sampling for Efficient MPSoC Architecture SimulationabstractModern micro-architecture simulators are many orders of magnitude slower than the hardware they simulate. The use of multiprocessor architectures for supporting future mobile and embedded applications will exacerbate this slowness. In this paper, we focus on the usage of the sampling technique for simulation acceleration, in the case of design space exploration (DSE), considering MPSoC. Among the addressed issue is the formation of sampling intervals that are executed simultaneously by the different processors. We propose a technique that dynamically adjusts the size for simulation samples for multiprocessor activities overlaps. Experimental results show that with our method, the simulation can be speedup by a factor of up to 800 with a relatively small estimation error. Melhem Tawk, Khaled Z. Ibrahim, Smaïl Niar |
MASCOTS | 3 |
| 2007 | An MPSoC Performance Estimation Framework Using Transaction Level ModelingabstractTo use the tremendous hardware resources available in next generation multiprocessor systems-on-chip (MPSoC) efficiently, rapid and accurate design space exploration (DSE) methods are needed to evaluate the different design alternatives. In this paper, we present a framework that makes fast simulation and performance evaluation of MPSoC possible early in the design flow, thus reducing the time-to-market. In this framework and within the transaction level modeling (TLM) approach, we present a new definition of the timed programmer's view (PVT) level by introducing two complementary modeling sublevels. The first one, PVT transaction accurate (PVT-TA), offers a high simulation speedup factor over the cycle accurate bit accurate (CABA) level modeling. The second one, PVT event accurate (PVT-EA), provides a better accuracy with a still acceptable speedup factor. An MPSoC platform has been developed using these two sublevels including performance estimation models. Simulation results show that the combination of these two sublevels gives a high simulation speedup factor of up to 18 with a negligible performance estimation error margin. Rabie Ben Atitallah, Smaïl Niar, Samy Meftali, Jean-Luc Dekeyser |
RTCSA | 2 |
| 2006 | Adapting EPIC Architecture's Register Stack for Virtual Stack MachinesabstractThe register stack (RS) is a major component of the explicit parallel instruction computer (EPIC) architecture. In this paper, our objective is to close the theoretical performance gap between EPIC and stack processors running virtual stack machines - using forth, a simple and canonical stack machine. For this purpose, we first introduce a new calling mechanism using the RS to implement a software-only virtual stack machine. Based upon our performance measurements, we show that the new calling mechanism is a promising technique to improve the performance of stack-based interpretative virtual machines. But limitation in EPIC makes the need for hardware support to reach optimal performance. As a second step, we define an addition to Itanium 2 processor's instruction set to accommodate the new calling mechanism. As our third and last step, we describe a conservative architectural implementation of the extended instruction set Jamel Tayeb, Smaïl Niar |
DSD | 2 |
| 2006 | A Real Time Signal Processing for an Anticollision Road Radar SystemabstractThis paper describes the real time processing unit used for an anticollision road radar system. This radar based on a numerical correlation between the transmitted signal and the received signal is under development. The signal uses orthogonal codes to ensure a multiple access communication between all vehicles in near area. In this paper, a real time processing unit associated to an original anticollision radar is presented. The studied radar is based on spreading spectrum coded radar waveforms at 76-77 GHz and a numerical correlation receiver. This sensor associated to other sensors like Lidar and Camera will be used on-board vehicles for more safety on road. The studied receiver computes the numerical crosscorrelation between the received signal and a replica of the transmitted code to allow an optimal detection. The appropriate coding and processing have been used to implement a laboratory radar mock-up. The real time processing is tested in order to show their performances and disadvantages when applied to obstacles detection. The main idea is to achieve an efficient real time detection using a simple and low cost system. Laila Sakkila, Pascal Deloof, Yassin Elhillali, Atika Rivenq, Smaïl Niar |
VTC Fall | 5 |
| 2006 | Pattern-driven prefetching for multimedia applications on embedded processors
Hassan Sbeyti, Smaïl Niar, Lieven Eeckhout |
J. Syst. Archit. | 2 |
| 2005 | Optimal sample length for efficient cache simulation
Lieven Eeckhout, Smaïl Niar, Koen De Bosschere |
J. Syst. Archit. | 2 |
| 2004 | Adaptive Prefetching for Multimedia Applications in Embedded SystemsabstractThis paper presents a new and simple prefetching mechanism to improve the memory performance of multimedia applications. This method adapts the memory access mechanism to the access patterns as observed in the application. By doing so, performance is increased, the available resources are better utilized and energy consumption is reduced. Using our prefetch method, we are able to get up to 5.5% IPC improvement, more than 50% cache miss reduction, and up to 4.5% energy reduction. Our mechanism results in better performance for a 2KB data cache than is achievable with an 8KB data cache (without prefetching) for StrongArm SA1110 and Xscale-like processor configurations. This mechanism requires limited hardware resources and generates little additional external bus transfers. This makes this adaptive prefetching well suited for embedded microprocessor systems. Hassan Sbeyti, Smaïl Niar, Lieven Eeckhout |
DATE | 2 |
| 2004 | An automatic communication synthesis for high level SOC desing using transaction level modelling (poster)
E. Turbatu, Samy Meftali, Smaïl Niar, Jean-Luc Dekeyser |
FDL | 3 |
| 2001 | Performances of a Dynamic Threads Scheduler
Smaïl Niar, Mahamed Adda |
Euro-Par | 1 |
| 1990 | The evaluation of the N-arch emulator on a transputer network
Smaïl Niar, Gilles Goncalves, Bernard Toursel, Marie-Paule Lecouffe |
Microprocessing and Microprogramming | 1 |
| 1988 | A network of transputers to emulate a parallel symbolic processor
Gilles Goncalves, Marie-Paule Lecouffe, Bernard Toursel, Smaïl Niar |
Microprocess. Microprogramming | 4 |