El Houcine Bergou

dblp:152/3495 · DBLP profile ↗
← Back
20ranked-venue papers
0as first author
17since 2021 · last 2026
0000-0001-8685-6974ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Stabilizing Policy Gradient Methods via Reward Profiling
abstract
Policy gradient methods, which have been extensively studied in the last decade, offer an effective and efficient framework for reinforcement learning problems. However, their performances can often be unsatisfactory, suffering from unreliable reward improvements and slow convergence, due to high variance in gradient estimations. In this paper, we propose a universal reward profiling framework that can be seamlessly integrated with any policy gradient algorithm, where we selectively update the policy based on high-confidence performance estimations. We theoretically justify that our technique will not slow down the convergence of the baseline policy gradient methods, but with high probability, will result in stable and monotonic improvements of their performance. Empirically, on eight continuous‐control benchmarks (Box2D and MuJoCo/PyBullet), our profiling yields up to 1.5x faster convergence to near‐optimal returns, up to 1.75x reduction in return variance on some setups. Our profiling approach offers a general, theoretically grounded path to more reliable and efficient policy learning in complex environments.
Shihab Ahmed, El Houcine Bergou, Yue Wang 0068, Aritra Dutta
AAAI2
2026 Just Few States Are Enough: Randomized Sparse Feedback for Stability of Dynamical Systems
abstract
While classical control theory assumes that the controller has access to measurements of the entire state (or output) at every time instant, this paper investigates a setting where the feedback controller can only access a randomly selected subset of the state vector at each time step. Due to the random sparsification that selects only a subset of the state components at each step, we analyze the stability of the closed-loop system in terms of Asymptotic Mean-Square Stability (AMSS), which ensures that the system state converges to zero in the mean-square sense. We consider the problem of designing both a feedback gain matrix and a measurement sparsification strategy that minimizes the number of state components required for feedback, while ensuring AMSS of the closed-loop system. Interestingly, (1) we provide conditions on the dynamics of the system under which it is possible to find a sparsification strategy, and (2) we propose a Linear Matrix Inequality (LMI) based algorithm that jointly computes a stabilizing gain matrix, and a randomized sparsification strategy that minimizes the expected number of measured state coordinates while preserving the AMSS. Our approach is then extended to the case where the sparsification probabilities vary across the state components. Based on these theoretical findings, we propose an algorithmic procedure to compute the vector of sparsification parameters, along with the corresponding feedback gain matrix. To the best of our knowledge, this is the first study to investigate the stability properties of control systems that rely solely on randomly selected state measurements. Numerical simulations demonstrate that, in some settings, the system achieves comparable performance to full-state feedback while requiring measurements from only 0.3 percent of the state coordinates.
Zaid Hadach, Hajar Elhammouti, El Houcine Bergou, Adnane Saoud
AAAI3
2025 On the Fairness of Ensemble Learning Methods in Student Dropout Prediction
Abdelghafour Aboukacem, Loubna Mekouar, El Houcine Bergou, Youssef Iraqi, Ismail Berrada
AIED (5)3
2024 Minibatch Stochastic Three Points Method for Unconstrained Smooth Minimization
abstract
We present a new zero-order optimization method called Minibatch Stochastic Three Points (MiSTP), specifically designed to solve stochastic unconstrained minimization problems when only an approximate evaluation of the objective function is possible. MiSTP is an extension of the Stochastic Three Point Method (STP). The key innovation of MiSTP is that it selects the next point solely based on the objective function approximation, without relying on its exact evaluation. At each iteration, MiSTP generates a random search direction and compares the approximations of the objective function at the current point, the randomly generated direction and its opposite. The best of these three points is chosen as the next iterate. We analyze the worst-case complexity of MiSTP in the convex and non-convex cases and demonstrate that it matches the most accurate complexity bounds known in the literature for zero-order optimization methods. We perform extensive numerical evaluations to assess the computational efficiency of MiSTP and compare its performance to other state-of-the-art methods by testing it on several machine learning tasks. The results show that MiSTP outperforms or has comparable performance against state-of-the-art methods indicating its potential for a wide range of practical applications.
Soumia Boucherouite, Grigory Malinovsky, Peter Richtárik, El Houcine Bergou
AAAI4
2024 Investigating the Predictive Potential of Large Language Models in Student Dropout Prediction
Abdelghafour Aboukacem, Ismail Berrada, El Houcine Bergou, Youssef Iraqi, Loubna Mekouar
AIED (2)3
2024 Student At-Risk Identification and Classification Through Multitask Learning: A Case Study on the Moroccan Education System
Ismail Elbouknify, Ismail Berrada, Loubna Mekouar, Youssef Iraqi, El Houcine Bergou, Hind Belhabib, Younes Nail, Souhail Wardi
AIED (2)5
2024 Tolerating Outliers: Gradient-Based Penalties for Byzantine Robustness and Inclusion
Latifa Errami, El Houcine Bergou
IJCAI2
2024 If You Want to Be Robust, Be Wary of Initialization
abstract
Graph Neural Networks (GNNs) have demonstrated remarkable performance across a spectrum of graph-related tasks, however concerns persist regarding their vulnerability to adversarial perturbations. While prevailing defense strategies focus primarily on pre-processing techniques and adaptive message-passing schemes, this study delves into an under-explored dimension: the impact of weight initialization and associated hyper-parameters, such as training epochs, on a model’s robustness. We introduce a theoretical framework bridging the connection between initialization strategies and a network's resilience to adversarial perturbations. Our analysis reveals a direct relationship between initial weights, number of training epochs and the model’s vulnerability, offering new insights into adversarial robustness beyond conventional defense mechanisms. While our primary focus is on GNNs, we extend our theoretical framework, providing a general upper-bound applicable to Deep Neural Networks. Extensive experiments, spanning diverse models and real-world datasets subjected to various adversarial attacks, validate our findings. We illustrate that selecting appropriate initialization not only ensures performance on clean datasets but also enhances model robustness against adversarial perturbations, with observed gaps of up to 50\% compared to alternative initialization approaches.
Sofiane Ennadir, Johannes F. Lutzeyer, Michalis Vazirgiannis, El Houcine Bergou
NeurIPS4
2024 EUREKHA: Enhancing User Representation for Key Hackers Identification in Underground Forums
abstract
Underground forums serve as hubs for cybercriminal activities, offering a space for anonymity and evasion of conventional online oversight. In these hidden communities, malicious actors collaborate to exchange illicit knowledge, tools, and tactics, driving a range of cyber threats—from hacking techniques to the sale of stolen data, malware, and zero-day exploits. Identifying the key instigators (i.e., key hackers), behind these operations is essential but remains a complex challenge. This paper presents a novel method called EUREKHA (Enhancing User Representation for Key Hacker Identification in Underground Forums), designed to identify these key hackers by modeling each user as a textual sequence. This sequence is processed through a large language model (LLM) for domain-specific adaptation, with LLMs acting as feature extractors. These extracted features are then fed into a Graph Neural Network (GNN) to model user structural relationships, significantly improving identification accuracy. Furthermore, we employ BERTopic (Bidirectional Encoder Representations from Transformers Topic Modeling) to extract personalized topics from user-generated content, enabling multiple textual representations per user and optimizing the selection of the most representative sequence. Our study demonstrates that fine-tuned LLMs outperform state-of-the-art methods in identifying key hackers. Additionally, when combined with GNNs, our model achieves significant improvements, resulting in approximately 6% and 10% increases in accuracy and F1-score, respectively, over existing methods. EUREKHA was tested on the Hack-Forums1dataset, and we provide open-source access to our code2.
Abdoul Nasser Hassane Amadou, Anas Motii, Saida Elouardi, El Houcine Bergou
TrustCom4
2024 Energy Efficient Aerial RIS: Phase Shift Optimization and Trajectory Design
abstract
Reconfigurable Intelligent Surface (RIS) technology has gained significant attention due to its ability to enhance the performance of wireless communication systems. The main advantage of RIS is that it can be strategically placed in the environment to control wireless signals, enabling improvements in coverage, capacity, and energy efficiency. In this paper, we investigate a scenario in which a drone, equipped with a RIS, travels from an initial point to a target destination. In this scenario, the aerial RIS (ARIS) is deployed to establish a direct link between the base station and obstructed users. Our objective is to maximize the energy efficiency of the ARIS while taking into account its dynamic model including its velocity and acceleration along with the phase shift of the RIS. To this end, we formulate the energy efficiency problem under the constraints of the dynamic model of the drone. The studied problem is challenging to solve. To address this, we proceed as follows. First, we introduce an efficient solution that involves decoupling the phase shift optimization and the trajectory design. Specifically, the closed-form expression of the phase-shift is obtained using a convex approximation, which is subsequently integrated into the trajectory design problem. We then employ tools inspired by economic model predictive control (EMPC) to solve the resulting trajectory optimization. Our simulation results show a significant improvement in energy efficiency against the scenario where the dynamic model of the UAV is ignored.
Hajar Elhammouti, Adnane Saoud, Asma Ennahkami, El Houcine Bergou
VTC Spring4
2024 Latency Minimization in Heterogeneous Federated Learning through Joint Compression and Resource Allocation
abstract
Federated Learning (FL) has emerged as a promising decentralized machine learning (ML) paradigm where distributed clients collaboratively train models without sharing their private data. However, the heterogeneous properties of the clients, combined with the high dimensions of ML models considerably slow down the wall-clock convergence time. To address these challenges, we propose FedHC, a framework that jointly optimizes resource allocations and uplink compression levels of the clients to minimize the overall latency while respecting the energy budget and convergence guarantees. To solve the formulated optimization problem, we first derive the required number of global training rounds, to achieve the target accuracy. Then we propose an iterative algorithm, where at each step optimal CPU levels and bandwidth along with compression levels are derived. Our numerical results show the performance -with time reduction up to 4×- and robustness to non-IID data of our approach, compared to the benchmarks.
Ouiame Marnissi, Hajar Elhammouti, El Houcine Bergou
VTC Fall3
2024 Semantic-Aware Resource Allocation in Constrained Networks with Limited User Participation
abstract
Semantic communication has gained attention as a key enabler for intelligent and context-aware communication. However, one of the key challenges of semantic communications is the need to tailor the resource allocation to meet the specific requirements of semantic transmission. In this paper, we focus on networks with limited resources where devices are constrained to transmit with limited bandwidth and power over large distance. Specifically, we devise an efficient strategy to select the most pertinent semantic features and participating users, taking into account the channel quality, the transmission time, and the recovery accuracy. To this end, we formulate an optimization problem with the goal of selecting the most relevant and accurate semantic features over devices while satisfying constraints on transmission time and quality of the channel. This involves optimizing communication resources, identifying participating users, and choosing specific semantic information for transmission. The underlying problem is inherently complex due to its non-convex nature and combinatorial constraints. To overcome this challenge, we efficiently approximate the optimal solution by solving a series of integer linear programming problems. Our numerical findings illustrate the effectiveness and efficiency of our approach in managing semantic communications in networks with limited resources.
Ouiame Marnissi, Hajar Elhammouti, El Houcine Bergou
WCNC3
2024 Age-of-Information in UAV-assisted Networks: a Decentralized Multi-Agent Optimization
abstract
Unmanned aerial vehicles (UAVs) are a highly promising technology with diverse applications in wireless networks. One of their primary uses is the collection of time-sensitive data from Internet of Things (IoT) devices. In UAV-assisted networks, the Age-of-Information (AoI) serves as a fundamental metric for quantifying data timeliness and freshness. In this work, we are interested in a generalized AoI formulation, where each packet's age is weighted based on its generation time. Our objective is to find the optimal UAVs' trajectories and the subsets of selected devices such that the weighted AoI is minimized. To address this challenge, we formulate the problem as a Mixed-Integer Nonlinear Programming (MINLP), incorporating time and quality of service constraints. To efficiently tackle this complex problem and minimize communication overhead among UAVs, we propose a distributed approach. This approach enables drones to make independent decisions based on locally acquired data. Specifically, we reformulate our problem such that our objective function is easily decomposed into individual rewards. The reformulated problem is solved using a distributed implementation of Multi-Agent Reinforcement Learning (MARL). Our empirical results show that the proposed decentralized approach achieves results that are nearly equivalent to a centralized implementation with a notable reduction in communication overhead.
Mouhamed Naby Ndiaye, El Houcine Bergou, Hajar Elhammouti
WCNC2
2023 Ensemble DNN for Age-of-Information Minimization in UAV-assisted Networks
abstract
This paper addresses the problem of Age-of-Information (AoI) in UAV-assisted networks. Our objective is to minimize the expected AoI across devices by optimizing UAVs’ stopping locations and device selection probabilities. To tackle this problem, we first derive a closed-form expression of the expected AoI that involves the probabilities of selection of devices. Then, we formulate the problem as a non-convex minimization subject to quality of service constraints. Since the problem is challenging to solve, we propose an Ensemble Deep Neural Network (EDNN) based approach which takes advantage of the dual formulation of the studied problem. Specifically, the Deep Neural Networks (DNNs) in the ensemble are trained in an unsupervised manner using the Lagrangian function of the studied problem. Our experiments show that the proposed EDNN method outperforms traditional DNNs in reducing the expected AoI, achieving a remarkable reduction of 29.5%.
Mouhamed Naby Ndiaye, El Houcine Bergou, Hajar Elhammouti
VTC Fall2
2022 Age-of-Updates Optimization for UAV-assisted Networks
abstract
Unmanned aerial vehicles (UAVs) have been proposed as a promising technology to collect data from IoT devices and relay it to the network. In this work, we are interested in scenarios where the data is updated periodically, and the collected updates are time-sensitive. In particular, the data updates may lose their value if they are not collected and analyzed timely. To maximize the data freshness, we optimize a new performance metric, namely the Age-of-Updates (AoU). Our objective is to carefully schedule the UAVs hovering positions and the users' association so that the AoU is minimized. Unlike existing works where the association parameters are considered as binary variables, we assume that devices send their updates according to a probability distribution. As a consequence, instead of optimizing a deterministic objective function, the objective function is replaced by an expectation over the probability distribution. The expected AoU is therefore optimized under quality of service and energy constraints. The original problem being non-convex, we propose an equivalent convex optimization that we solve using an interior-point method. Our simulation results show the performance of the proposed approach against a binary association.
Mouhamed Naby Ndiaye, El Houcine Bergou, Mounir Ghogho, Hajar Elhammouti
GLOBECOM2
2021 Outsmarting the Atmospheric Turbulence for Ground-Based Telescopes Using the Stochastic Levenberg-Marquardt Method
Yuxi Hong 0001, El Houcine Bergou, Nicolas Doucet, Jesse Cranney, Hatem Ltaief, Damien Gratadour, François Rigaut, David E. Keyes
Euro-Par2
2021 GRACE: A Compressed Communication Framework for Distributed Machine Learning
abstract
Powerful computer clusters are used nowadays to train complex deep neural networks (DNN) on large datasets. Distributed training increasingly becomes communication bound. For this reason, many lossy compression techniques have been proposed to reduce the volume of transferred data. Unfortunately, it is difficult to argue about the behavior of compression methods, because existing work relies on inconsistent evaluation testbeds and largely ignores the performance impact of practical system configurations. In this paper, we present a comprehensive survey of the most influential compressed communication methods for DNN training, together with an intuitive classification (i.e., quantization, sparsification, hybrid and low-rank). Next, we propose GRACE, a unified framework and API that allows for consistent and easy implementation of compressed communication on popular machine learning toolkits. We instantiate GRACE on TensorFlow and PyTorch, and implement 16 such methods. Finally, we present a thorough quantitative evaluation with a variety of DNNs (convolutional and recurrent), datasets and system configurations. We show that the DNN architecture affects the relative performance among methods. Interestingly, depending on the underlying communication library and computational cost of compression / decompression, we demonstrate that some methods may be impractical. GRACE and the entire benchmarking suite are available as open-source.
Chen-Yu Ho 0001, Ahmed M. Abdelmoniem, Aritra Dutta, El Houcine Bergou, Konstantinos Karatsenidis, Marco Canini, Panos Kalnis
ICDCS5
2020 A Stochastic Derivative-Free Optimization Method with Importance Sampling: Theory and Learning to Control
abstract
We consider the problem of unconstrained minimization of a smooth objective function in ℝn in a setting where only function evaluations are possible. While importance sampling is one of the most popular techniques used by machine learning practitioners to accelerate the convergence of their models when applicable, there is not much existing theory for this acceleration in the derivative-free setting. In this paper, we propose the first derivative free optimization method with importance sampling and derive new improved complexity results on non-convex, convex and strongly convex functions. We conduct extensive experiments on various synthetic and real LIBSVM datasets confirming our theoretical results. We test our method on a collection of continuous control tasks on MuJoCo environments with varying difficulty. Experiments show that our algorithm is practical for high dimensional continuous control problems where importance sampling results in a significant sample complexity improvement.
Adel Bibi, El Houcine Bergou, Ozan Sener, Bernard Ghanem, Peter Richtárik
AAAI2
2020 On the Discrepancy between the Theoretical Analysis and Practical Implementations of Compressed Communication for Distributed Deep Learning
abstract
Compressed communication, in the form of sparsification or quantization of stochastic gradients, is employed to reduce communication costs in distributed data-parallel training of deep neural networks. However, there exists a discrepancy between theory and practice: while theoretical analysis of most existing compression methods assumes compression is applied to the gradients of the entire model, many practical implementations operate individually on the gradients of each layer of the model.In this paper, we prove that layer-wise compression is, in theory, better, because the convergence rate is upper bounded by that of entire-model compression for a wide range of biased and unbiased compression methods. However, despite the theoretical bound, our experimental study of six well-known methods shows that convergence, in practice, may or may not be better, depending on the actual trained model and compression ratio. Our findings suggest that it would be advantageous for deep learning frameworks to include support for both layer-wise and entire-model compression.
Aritra Dutta, El Houcine Bergou, Ahmed M. Abdelmoniem, Chen-Yu Ho 0001, Atal Narayan Sahu, Marco Canini, Panos Kalnis
AAAI2
2020 A Stochastic Derivative Free Optimization Method with Momentum
Eduard Gorbunov, Adel Bibi, Ozan Sener, El Houcine Bergou, Peter Richtárik
ICLR4