EDBT 2026 Demo / reviewers in the wild / expert
Hadjer Benmeziane
dblp:284/0784
· DBLP profile ↗
15ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-5259-0749ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | STARC: Selective Token Access with Remapping and Clustering for Efficient LLM Decoding on PIM SystemsabstractServing large language models (LLMs) places significant pressure on memory systems due to frequent accesses and growing key–value (KV) caches as context lengths increase. Processing-in-memory (PIM) architectures offer high internal bandwidth and near-data compute parallelism, but current designs target dense attention and perform poorly under the irregular access patterns of dynamic KV cache sparsity. To mitigate this limitation, we propose STARC, a sparsity-optimized data mapping scheme for efficient LLM decoding on PIM. STARC clusters semantically similar KV pairs and co-locates them contiguously within PIM banks, enabling retrieval at cluster granularity by matching queries against precomputed centroids. This bridges the gap between fine-grained sparse attention and row-level PIM operations, improving utilization while minimizing overhead. On a simulated HBM-PIM system, under constrained KV budgets, STARC achieves up to 78% and 65% reductions in attention-layer latency and energy over token-wise sparsity methods, and up to 93% and 92% reductions relative to full attention, while preserving model accuracy. Zehao Fan, Yunzhen Liu, Garrett Gagnon, Yayue Hou, Hadjer Benmeziane, Kaoutar El Maghraoui, Liu Liu 0017 |
ASPLOS (2) | 6 |
| 2026 | HILAL: Hessian-Informed Layer Allocation for Heterogeneous Analog-Digital InferenceabstractHeterogeneous AI accelerators that combine high-precision digital cores with energy-efficient analog in-memory computing (AIMC) units offer a promising path to overcome the energy and scalability limits of deep learning. A key challenge, however, is to determine which neural network layers can be executed on noisy analog units without compromising accuracy. Existing mapping strategies rely on ad-hoc heuristics and lack principled noise-sensitivity estimation. We propose HILAL (Hessian-Informed Layer Allocation), a framework that systematically quantifies layer robustness to analog noise using two complementary metrics: Hessian-based Noise Impact and Spectral Concentration Ratio. Layers are partitioned into robust and sensitive groups via clustering, enabling threshold-free mapping to analog or digital units. To further mitigate accuracy loss, we gradually offload layers to AIMC while retraining with noise-injection. Experiments on convolutional networks and transformers across CIFAR-10/100, ImageNet and SQuAD show that HILAL is on average 3.09x faster in search and mapping runtime than SOTA methods while achieving less accuracy degradation and maximizing analog utilization. Aniss Bessalah, Hatem Mohamed Abdelmoumen, Karima Benatchba, Hadjer Benmeziane |
DATE | 4 |
| 2026 | HARP: Heterogeneous Analog-Digital Resource-Aware Performance and Scheduling Framework for Transformer Acceleration
Elena Ferro, Hadjer Benmeziane, Irem Boybat |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | Hardware-Aware Compilation and Simulation for In-Memory ComputingabstractThis brief presents an overview of recent tools and research efforts aimed at enhancing the programmability and reliability of In-Memory Computing (IMC)-based systems. We discuss hardware-aware training techniques that improve model resilience to analog device imperfections, and explore mapping strategies that balance accuracy and performance for heterogeneous IMC-based accelerators. Additionally, we examine a compiler framework that abstracts hardware complexities and enables seamless integration of these accelerators into existing deployment pipelines. By combining these approaches with advanced simulation tools, we propose an end-to-end workflow that facilitates the practical deployment and optimization of IMC technologies across diverse memory types and architectural designs. Asif Ali Khan, Hadjer Benmeziane, Hamid Farzaneh, João Paulo C. de Lima, William Andrew Simon, Yiyu Shi 0001, Zheyu Yan, Abu Sebastian, Xiaobo Sharon Hu, Jerónimo Castrillón, Corey Lammie |
CASES | 2 |
| 2025 | A framework for analog-digital mixed-precision neural network training and inferenceabstractRecent advancements in AI hardware highlight the potential of mixed-signal accelerators, which integrate analog computation for matrix multiplications with reduced-precision digital operations, to achieve superior performance and energy efficiency. In this paper, we present a framework designed to perform hardware-aware training and inference evaluation of neural networks (NNs) on such accelerators. This framework extends an existing toolkit, the IBM Analog Hardware Acceleration Kit (AIHWKit), using a quantization library, enabling flexible layer-wise deployment in either analog or digital units, the latter with configurable precision and quantization options. Our combined framework supports simultaneous quantization-and analog-aware training as well as post-training calibration routines. It can also evaluate the accuracy of NNs when deployed on mixed-signal accelerators. We demonstrate the need of such a framework through ablation studies on a ResNet-based vision model and a BERT-based language model, highlighting the importance of its functionality for maximizing accuracy during deployment. Our contribution is open-sourced as part of the core code of AIHWKit [1]. Athanasios Vasilopoulos, Emma Boulharts, Corey Lammie, Julian Büchel, Hadjer Benmeziane, Manuel Le Gallo, Abu Sebastian |
ISCAS | 5 |
| 2024 | Analog AI as a Service: A Cloud Platform for In-Memory ComputingabstractThis paper introduces the Analog AI Cloud Composer platform, a service that allows users to access Analog In-Memory Computing (AIMC) simulation and computing resources over the cloud. We introduce the concept of an Analog AI as a Service (AAaaS). AIMC offers a novel approach for decreasing both the latency and energy usage associated with Deep Neural Network (DNN) inference and training. This platform democratizes access to AIMC computing, making it available to a broader audience, including researchers, developers, and businesses. Emphasizing a user-friendly, no-code approach, AAaaS integrates the Analog Hardware Acceleration Kit (AIHWKit) simulation platform within a fully managed cloud environment. We discuss the architecture of the Analog AI Cloud Composer (AAICC), focusing on its key services such as inference, training, and AIMC hardware access. The platform's design, grounded in cloud services and guidelines, ensures a secure, data-centric user experience with robust control and validation mechanisms. Kaoutar El Maghraoui, Kim Tran, Kurtis Ruby, Borja Godoy, Jordan Murray, Manuel Le Gallo-Bourdeau, Todd Deshane, Pablo Gonzalez, Diego Moreda, Hadjer Benmeziane, Corey Lammie, Julian Büchel, Malte J. Rasch, Abu Sebastian, Vijay Narayanan |
SSE | 10 |
| 2024 | Medical Neural Architecture Search: Survey and Taxonomy
Hadjer Benmeziane, Imane Hamzaoui, Zayneb Cherif, Kaoutar El Maghraoui |
IJCAI | 1 |
| 2024 | DfuseNAS: A Diffusion-Based Neural Architecture SearchabstractDeep Learning (DL) has revolutionized numerous domains by crafting highly effective Neural Network (NN) architectures. However, manual engineering approaches for NN design often yield sub-optimal solutions. Neural Architecture Search (NAS) addresses this issue by automating the design process and discovering state-of-the-art architectures. Nevertheless, the exploration of vast search spaces in NAS remains a challenging endeavor. Diffusion models have proven effective in traversing expansive search spaces encountered in generative image tasks. Their innate ability to compress and explore these spaces has shown a lot of promise. Building upon this inspiration, we introduce DfuseNAS, a novel NAS methodology rooted in diffusion processes. DfuseNAS brings substantial improvements in both NAS search efficiency and the quality of the generated neural network architectures. To the best of our knowledge, our work marks a pioneering effort in applying diffusion algorithms to enhance the search space exploration with NAS. Our experimental results, conducted on the widely-used NAS-Bench-101, showcase the remarkable capabilities of DfuseNAS. We achieved the highest average accuracy, outperforming other state-of-the-art methods, while completing the search process at least 2 times faster. Moreover, when provided with a specific architecture and a given task, the application of DfuseNAS consistently led to the generation of more accurate architectures in 98% of the times. Lotfi Abdelkrim Mecharbat, Hadjer Benmeziane, Hamza Ouarnoughi, Smaïl Niar, Kaoutar El Maghraoui |
IJCNN | 2 |
| 2024 | Grassroots operator search for model edge adaptation using mathematical search space
Hadjer Benmeziane, Kaoutar El Maghraoui, Hamza Ouarnoughi, Smaïl Niar |
Future Gener. Comput. Syst. | 1 |
| 2023 | Pareto Rank-Preserving Supernetwork for Hardware-Aware Neural Architecture SearchabstractIn neural architecture search (NAS), training every sampled architecture is very time-consuming and should be avoided. Weight-sharing is a promising solution to speed up the evaluation process. However, training the supernetwork incurs many discrepancies between the actual ranking and the predicted one. Additionally, efficient deep-learning engineering processes require incorporating realistic hardware-performance metrics into the NAS evaluation process, also known as hardware-aware NAS (HW-NAS). In HW-NAS, estimating task-specific performance and hardware efficiency are both required. This paper proposes a supernetwork training methodology that preserves the Pareto ranking between its different subnetworks resulting in more efficient and accurate neural networks for a variety of hardware platforms. The results show a 97% near Pareto front approximation in less than 2 GPU days of search, which provides 2x speed up compared to state-of-the-art methods. We validate our methodology on NAS-Bench-201, DARTS, and ImageNet. Our optimal model achieves 77.2% accuracy (+1.7% compared to baseline) with an inference time of 3.68ms on Edge GPU for ImageNet, which yields a 2.3x speedup. Training implementation can be found: https://github.com/IHIaadj/PRP-NAS. Hadjer Benmeziane, Kaoutar El Maghraoui, Hamza Ouarnoughi, Smaïl Niar |
ECAI | 1 |
| 2023 | FLASH-RL: Federated Learning Addressing System and Static Heterogeneity using Reinforcement LearningabstractWe propose FLASH-RL, a framework utilizing Double Deep Q-Learning (DDQL) to address system and static heterogeneity in Federated Learning (FL). FLASH-RL introduces a new reputation-based utility function to evaluate client contributions based on their current and past performances. Additionally, an adapted DDQL algorithm is proposed to expedite the learning process. Experimental results on MNIST and CIFAR-10 datasets demonstrate that FLASH-RL strikes a balance between model performance and end-to-end latency, reducing latency by up to 24.83% compared to FedAVG and 24.67% compared to FAVOR. It also reduces training rounds by up to 60.44% compared to FedAVG and 76% compared to FAVOR. Similar improvements are observed on the MobiAct Dataset for fall detection, underscoring the real-world applicability of our approach. Sofiane Bouaziz, Hadjer Benmeziane, Youcef Imine, Leila Hamdad, Smaïl Niar, Hamza Ouarnoughi |
ICCD | 2 |
| 2023 | Multi-objective Hardware-aware Neural Architecture Search with Pareto Rank-preserving Surrogate ModelsabstractDeep learning (DL) models such as convolutional neural networks (ConvNets) are being deployed to solve various computer vision and natural language processing tasks at the edge. It is a challenge to find the right DL architecture that simultaneously meets the accuracy, power, and performance budgets of such resource-constrained devices. Hardware-aware Neural Architecture Search (HW-NAS) has recently gained steam by automating the design of efficient DL models for a variety of target hardware platforms. However, such algorithms require excessive computational resources. Thousands of GPU days are required to evaluate and explore an architecture search space such as FBNet [ 45 ]. State-of-the-art approaches propose using surrogate models to predict architecture accuracy and hardware performance to speed up HW-NAS. Existing approaches use independent surrogate models to estimate each objective, resulting in non-optimal Pareto fronts. In this article, HW-PR-NAS, 1 a novel Pareto rank-preserving surrogate model for edge computing platforms, is presented. Our model integrates a new loss function that ranks the architectures according to their Pareto rank, regardless of the actual values of the various objectives. We employ a simple yet effective surrogate model architecture that can be generalized to any standard DL model. We then present an optimized evolutionary algorithm that uses and validates our surrogate model. Our approach has been evaluated on seven edge hardware platforms from various classes, including ASIC, FPGA, GPU, and multi-core CPU. The evaluation results show that HW-PR-NAS achieves up to 2.5× speedup compared to state-of-the-art methods while achieving 98% near the actual Pareto front. Hadjer Benmeziane, Hamza Ouarnoughi, Kaoutar El Maghraoui, Smaïl Niar |
ACM Trans. Archit. Code Optim. | 1 |
| 2022 | CaW-NAS: Compression Aware Neural Architecture SearchabstractWith the ever-growing demand for deep learning (DL) at the edge, building small and efficient DL architectures has become a significant challenge. Optimization techniques such as quantization, pruning or hardware-aware neural architecture search (HW-NAS) have been proposed. In this paper, we present an efficient HW-NAS; Compression-Aware Neural Architecture search (CaW-NAS), that combines the search for the architecture and its quantization policy. While former works search over a fully quantized search space, we define our search space with quantized and non-quantized architectures. Our search strategy finds the best trade-off between accuracy and latency according to the target hardware. Experimental results on a mobile platform show that, our method allows to obtain more efficient networks in terms of accuracy, execution time and energy consumption when compared to the state of the art. Hadjer Benmeziane, Hamza Ouarnoughi, Smaïl Niar, Kaoutar El Maghraoui |
DSD | 1 |
| 2022 | Pareto Rank Surrogate Model for Hardware-aware Neural Architecture SearchabstractHardware-aware Neural Architecture Search (HWNAS) has recently gained much attention by automating the design of efficient deep learning models with tiny resources and reduced inference time requirements. However, HW-NAS inherits and exacerbates the expensive computational complexity of general NAS due to its significantly increased search spaces and more complex NAS evaluation component. To speed up HWNAS, existing efforts use surrogate models to predict a neural architecture’s accuracy and hardware performance on a specific platform. Thereby reducing the expensive training process and significantly reducing search time. We show that using multiple surrogate models to estimate the different objectives does not achieve the true Pareto front. Therefore, we propose HW-PRNAS, a novel Pareto Rank-preserving surrogate model. HWPR-NAS training is based on a new loss function that ranks the architectures according to their Pareto front. We evaluate our approach on seven different hardware platforms, including ASIC, FPGA, GPU and multi-cores. Our results show that we can achieve up to 2. 5x speedup while achieving better Pareto-front results than state of the art surrogate models. Hadjer Benmeziane, Smaïl Niar, Hamza Ouarnoughi, Kaoutar El Maghraoui |
ISPASS | 1 |
| 2021 | Hardware-Aware Neural Architecture Search: Survey and TaxonomyabstractThere is no doubt that making AI mainstream by bringing powerful, yet power hungry deep neural networks (DNNs) to resource-constrained devices would required an efficient co-design of algorithms, hardware and software. The increased popularity of DNN applications deployed on a wide variety of platforms, from tiny microcontrollers to data-centers, have resulted in multiple questions and challenges related to constraints introduced by the hardware. In this survey on hardware-aware neural architecture search (HW-NAS), we present some of the existing answers proposed in the literature for the following questions: "Is it possible to build an efficient DL model that meets the latency and energy constraints of tiny edge devices?", "How can we reduce the trade-off between the accuracy of a DL model and its ability to be deployed in a variety of platforms?". The survey provides a new taxonomy of HW-NAS and assesses the hardware cost estimation strategies. We also highlight the challenges and limitations of existing approaches and potential future directions. We hope that this survey will help to fuel the research towards efficient deep learning. Hadjer Benmeziane, Kaoutar El Maghraoui, Hamza Ouarnoughi, Smaïl Niar, Martin Wistuba, Naigang Wang |
IJCAI | 1 |