Youcef Djenouri

dblp:130/8288 · DBLP profile ↗
← Back
109ranked-venue papers
56as first author
74since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 51 · 22 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 11 first-author · 22 since 2021Databases, data management, data science and information retrieval · 19 · 12 first-author · 7 since 2021Computer networks · 17 · 6 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 4 since 2021Systems, architecture and hardware · 5 · 4 first-author · 3 since 2021Theory of computation · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Uncertainty-Aware Drone Swarm Disruption for Threat Mitigation
abstract
This study tackles the challenge of neutralizing a malicious drone swarm targeting critical infrastructure. Given the impracticality of destroying all drones, we propose an uncertainty-aware threat assessment method, modeling payload size and positional uncertainties probabilistically. We studynode removalas the core decision problem to fragment communication and reduce attack power. The threat of each drone is quantified by its expected payload and distance-based probabilistic model. We also assess the probabilistic connectivity between drones to evaluate the swarm’s robustness. A novel evaluation function integrates these factors, and we introduce a probabilistic greedy search algorithm to minimize the swarm’s threat. We evaluate against centrality-based and weight-augmented dismantling heuristics, a recent DQN policy, and a small-n brute-force oracle. Our approach achieves a 70% reduction in attack power and near-optimal performance, deviating by less than 5% from the best possible outcomes under uncertainty.
Noor Ullah, Youcef Djenouri, Tomasz P. Michalak, Ahmed Nabil Belbachir, Gautam Srivastava 0001
IEEE Internet Things J.2
2025 Drone Swarm Sensitivity Estimation Using Bayesian Theory for Search and Rescue Operations
abstract
Search and rescue operations demand a fast and reliable process for locating victims. The use of unmanned aerial vehicles (UAV s) has become a prominent research topic due to their lower cost and operational flexibility compared to manned air-craft. This study investigates the use of Bayesian probabilistic methods to estimate the sensitivity of a drone swarm in maritime search and rescue missions. An initial sensitivity distribution is defined based on both prior exact and approximate knowledge, and is subsequently updated using observational data: the number of drones (out of 10) that successfully identified a survivor while flying over the same area. The model is iteratively updated and evaluated. Results show that even after just four missions, there is a consistent convergence of the mean and a reduction in the standard deviation from 0.1307 to 0.0449. These findings demonstrate a reusable framework that improves swarm sensitivity estimation even with limited data. The full code is available at https://github.com/lima0luciano/bayesian-drone-sar.
Luciano Netto de Lima, Ali Ghaderi, Fabio A. A. Andrade, Carlos Pfeiffer, Youcef Djenouri, Marcos Moura
DSAA5
2025 Knowledge-Driven Bayesian Uncertainty Quantification for Reliable Fake News Detection
abstract
The pervasive dissemination of fake news presents significant challenges to societal well-being and informed decision-making, necessitating robust detection mechanisms with calibrated uncertainty measures. This paper proposes a novel hybrid framework for fake news detection, integrating uncertainty quantification with a domain-specific Knowledge Base approach. The BANED knowledge base models word-level probabilistic significance, leveraging statistical support metrics to assess prediction uncertainty. By incorporating these metrics into a Bayesian framework, our method provides well-calibrated predictive distributions, offering enhanced interpretability and robustness in the presence of ambiguous or conflicting news data. The proposed approach is evaluated on the FakeNewsNet and ISOT Fake News datasets, demonstrating competitive accuracy and superior reliability compared to state-of-the-art Bayesian inference techniques. Combining word-level probabilistic significance with Monte Carlo Dropout decreases mean calibration error and narrows the interquartile range of predictions. Full code and supplementary materials of BANED might be found at https://github.com/micbizon/BANED.
Julia Puczynska, Youcef Djenouri, Michal Bizon, Tomasz P. Michalak, Piotr Sankowski
ECAI2
2025 Optimizing Object Detection for Maritime Search and Rescue: Progressive Fine-Tuning of YOLOv9 with Real and Synthetic Data
Luciano Netto de Lima, Fabio A. A. Andrade, Youcef Djenouri, Carlos Pfeiffer, Marcos Moura
ICAART (3)3
2025 Learning Graph Representation of Agent Diffusers
Youcef Djenouri, Nassim Belmecheri, Tomasz P. Michalak, Jan Dubinski, Ahmed Nabil Belbachir, Anis Yazidi
AAMAS1
2025 Hybrid Visibility Graph and Long Short Term Memory for Schizophrenia Detection
abstract
Schizophrenia is a complex neuropsychiatric disorder that affects cognitive function and brain activity. Electroencephalography (EEG) has emerged as a valuable tool for detecting schizophrenia-related neural patterns, but accurate classification remains a challenge due to the intricate nature of EEG signals. In this study, we propose a Hybrid Visibility Graph and Long Short-Term Memory (VG-LSTM) framework for schizophrenia detection. Our approach transforms EEG time series into graph structures using visibility graphs (VGs) to capture the underlying topological properties of brain activity. We then employ LSTM networks to model sequential dependencies, effectively integrating both structural and sequential information for robust classification. Experimental results on publicly available schizophrenia EEG datasets demonstrate that our VG-LSTM framework achieves superior performance compared to conventional deep learning approaches. The results highlight the potential of combining graph-theoretic and deep sequential modeling techniques for EEG-based neuropsychiatric disorder detection. The full code of this research work is available on https://github.com/YousIA/VG-LSTM.
Asma Belhadi, Youcef Djenouri, Pedro G. Lind, Anis Yazidi
IJCNN2
2025 Shared Knowledge Base for Multi Deep Learning in Defect Detection
abstract
In recent years, there has been growing interest in applying deep learning techniques for visual anomaly detection, particularly in the manufacturing sector. Various models have been developed to identify defects in manufacturing data, yet selecting and optimizing these models for anomaly detection in intelligent manufacturing environments remains a significant challenge. This research focuses on general-purpose visual anomaly detection, aiming to reduce dependence on domain-specific knowledge and create flexible, generic models. We propose a novel deep learning framework in which multiple models are trained for each image. The visual features and loss values from these models are computed and stored during training. During the testing phase, this stored information is used to select the most appropriate model for each new image using a k-Nearest Neighbors (kNN) approach. The proposed method, KGDL-VAD (Knowledge-Guided Deep Learning for Visual Anomaly Detection), was evaluated on the MVTec AD, and standard aerospace defect detection datasets, achieving an area under the curve (AUC) score of 0.96, outperforming baseline methods. In addition, KGDL-VAD surpasses ensemble learning approaches across multiple domain-independent datasets with varying numbers of trained classes.
Youcef Djenouri, Asma Belhadi, Gautam Srivastava 0001, Ahmed Nabil Belbachir, Alberto Cano 0001
IJCNN1
2025 Shapley Consensus Deep Learning for Ensemble Pruning
abstract
This paper targets a new foundation for designing general-purpose learning systems, by establishing a consensus method that facilitates self-adaptation and flexibility to deal with different learning tasks and different data distribution. We present the Shapely Consensus Deep Learning (SCDL) as a consensus method for general-purpose solutions that do not require the help of domain experts. SCDL is two-level based learning process. In the first level, several deep learning models are trained and the Shapley Value is used to determine the contribution of each subset of models in the training. The models are pruned according to their contribution in the learning process. In the second level, the loss information of each data distribution is saved in the knowledge base. Both levels are explored to prune the models for each new observation. We present the evaluation of the generality of SCDL using different datasets with different shapes, and complexities. The results reveal the effectiveness of SCDL for weakly classification. Concretely, SCDL achieved 90% of AUC with less than 86% for the baseline solutions.
Youcef Djenouri, Ahmed Nabil Belbachir, Asma Belhadi, Nassim Belmecheri, Tomasz P. Michalak
WACV1
2025 EEG Data Classification: Review and Taxonomy
abstract
EEG data classification plays a pivotal role in understanding brain activity and its applications in various domains. Deep learning has emerged as a powerful paradigm for automatically learning complex patterns from raw data, eliminating the need for manual feature extraction. However, in the context of medical data, and in particular for EEG analysis, the use of deep learning approaching while having been very successful is not being included in medical diagnosis routines, yet. The aim of this survey is twofold. On one side, it provides a comprehensive overview of the current state-of-the-art in EEG data classification, with a specific focus on the use of deep learning techniques. On the other side, it also addresses the clinician community, explaining the power and trustfulness of such new approaches. The survey begins with an introduction highlighting the limitations of traditional model-based approaches and the potential of deep learning in EEG data classification. The fundamental principles and architectures of deep learning models are presented, including convolutional neural networks (CNNs), recurrent neural networks (RNNs), and graph convolution neural networks (GCNNs) that have been successfully applied to EEG data classification tasks. A detailed review and analysis of existing literature on deep learning-based EEG data classification are provided, categorizing the studies based on the type of the input data, e.g., sequences, images, graphs, or multi-modalities. We also discuss about the existing tools and technologies for EEG data classification and highlights the challenges and limitations associated with deep learning in EEG data classification, including limited data availability, interpretability of deep models, and bias mitigation. Potential solutions and ongoing research efforts to overcome these challenges are explored, providing insights into the future directions of this field. This survey serves as a valuable resource for researchers, practitioners, and healthcare professionals involved in EEG data classification. It provides an extensive understanding of the advancements, challenges, and potential applications of deep learning techniques in this domain, guiding further research and development of accurate and interpretable approaches for EEG data analysis and interpretation.
Asma Belhadi, Anis Yazidi, Pedro G. Lind, Youcef Djenouri
ACM Trans. Comput. Heal.4
2025 Next-Gen Metaverse Security Through Intrusion Detection Enhanced by Transformers and GANs
abstract
As the metaverse grows in popularity and complexity, securing its virtual environment is critical. Metaverse intrusion detection involves identifying and preventing unauthorized access, malicious activities, and potential threats. To address these challenges, we propose a novel Metaverse intrusion detection system (MIDS) that combines generative adversarial networks (GAN) and Transformer-based classifiers. The system operates in three stages: 1) generating diverse and realistic network traffic using GAN; 2) detecting intrusions with a Transformer-based classifier; and 3) ensuring data privacy through federated learning and a trusted authority mechanism. Unlike traditional methods, our approach employs dual aggregation, generating both global and local models tailored to users’ needs. Tested on public datasets, the method achieves state-of-the-art performance with an F1-score of 0.9984, demonstrating its effectiveness in generating realistic training data and improving MIDS performance. This approach can extend to other security domains requiring diverse data for training.
Youcef Djenouri, Ahmed Nabil Belbachir, Asma Belhadi, Tomasz P. Michalak, Gautam Srivastava 0001
IEEE Internet Things J.1
2025 Detection for User Impersonation Attacks in Mobile Social Networks Based on High-order Markov Chains
abstract
Abstract In security defense of MSN (MSN), attackers often impersonate themselves as other users, making it difficult to detect network user attacks based on user behavior. Multi-order Markov chains can consider the front-to-back correlation of user behavior, thereby more accurately identifying disguised users. Therefore, this paper proposes a user impersonation attack detection method based on multi-order Markov chains. First, the relevance coefficient method is used to determine the order of the multi-order Markov chain, and by defining appropriate multi-order Markov chain states to capture key features in user behavior, a multi-order Markov chain is established. Then, through the multi-order Markov chain combined with Shell commands, the normal behavior profile of legitimate users is established, and based on this, the probability of occurrence of the state sequence is calculated to complete the detection of userimpersonation attacks. The experimental results show that the similarity between the results of the proposed method and the actual situation in detecting impersonation attacks is more than 97%, indicating that this method can detect MSN user impersonation attacks with high accuracy.
Wenhui Gong, Youcef Djenouri, Syed Atif Moqurrab
Mob. Networks Appl.3
2025 Grouping Interesting Patterns for Understanding Customer Behaviors
abstract
This article presents a highly efficient technique for pattern mining in the realm of customer behavior analysis, termed hybrid clustering patterns for customer behavior analysis (HCP-CBA). It leverages decomposition techniques to uncover relevant patterns by examining correlations among customer transactions within the dataset. Initially, the transaction dataset undergoes decomposition, grouping together transactions exhibiting high correlations. Subsequently, relevant patterns are extracted by applying a pattern mining algorithm represented by Apriori to each group. It incorporates both groups of transactions and shared items between groups. To assess the effectiveness of the HCP-CBA framework, extensive experiments are conducted across customer behavior dataset. The experimental results demonstrate notable reductions in both runtime and scalability. The full code of this research work is available onhttps://github.com/YousIA/ConsumerAnalytics.
Kristian Brathovde, Youcef Djenouri, Anis Yazidi, Gautam Srivastava 0001
IEEE Trans. Comput. Soc. Syst.2
2025 Knowledge Guided Visual Transformers for Intelligent Transportation Systems
abstract
We present a novel approach for addressing computer vision tasks in intelligent transportation systems, with a strong focus on data security during training through federated learning. Our method leverages visual transformers, training multiple models for each image. By calculating and storing visual image features as well as loss values, we propose a novel Shapley value model based on model performance consistency to select the most appropriate models during testing. To enhance security, we introduce an intelligent federated learning strategy, where users are grouped into clusters based on constrastive clustering for creating a global model as well as customized local models. Users receive both global as well as local models, enabling tailored computer vision applications. We evaluated KGVT-ITS (Knowledge Guided Visual Transformers for Intelligent Transportation Systems) on various ITS challenges, including pedestrian detection, abnormal event detection, as well as near-crash detection. The results demonstrate the superiority of KGVT-ITS over baseline solutions, showcasing its effectiveness and robustness in intelligent transportation scenarios. More particularly, KGVT-ITS achieves significant improvements of about 8% against the existing ITS methods.
Asma Belhadi, Youcef Djenouri, Ahmed Nabil Belbachir, Tomasz P. Michalak, Gautam Srivastava 0001
IEEE Trans. Intell. Transp. Syst.2
2024 Vision-based Spatiotemporal Learning for Human Activity Recognition
abstract
This paper introduces a novel concept for Human Activity Recognition (HAR) that allows robust analysis, classification, and understanding of human movements in various environments. It can be applied in various applications such as Health monitoring and analysis, fitness/dance training and performance analysis, interactive gaming, smart homes, and wearable devices. The novel method, coined as STL-HAR (SpatioTemporal Learning for HAR), learns from sensor data jointly represented in space and time to robustify the HAR process. In the new concept, we propose a hybrid model based on GNN (Graph Neural Network), and LSTM (Long Short-Term Memory). GNN first learns the spatial features from different sensor data locations. The learned features will then be injected to LSTM where the temporal information is captured by observing sensor status at different timestamps. We evaluate and analyze the performance of STL-HAR in real use case scenarios of HAR data compared with baseline HAR-based solutions. STL-HAR has achieved a recognition rate of 92% under different scenarios.
Youcef Djenouri, Ahmed Nabil Belbachir, Gautam Srivastava 0001, Alberto Cano 0001
IJCNN1
2024 POPD: Partial Occluded Pedestrian Detection Using A Multimodal Deep Learning Approach
abstract
With the enhancement in the growth of the automotive industry, pedestrian detection is one of the major issues to prevent collisions on roads. This also plays a significant role in various applications ranging from road safety to urban planning. Hence, in this paper, a multimodal deep learning scheme based Partial Occluded Pedestrian Detection (POPD) approach has been proposed. Primarily, data fusion has been performed by considering multimodal data from different sources for effective implementation of pedestrian detection. In the next step, a key-point detection algorithm was applied to this aggregated data, and Mask-RCNN was deployed to classify different occlusion profiles. To evaluate results, extensive simulations have been performed and exhibited results show that the proposed scheme has provided satisfactory results when compared with benchmark schemes. The findings of this study show a considerable increase in terms of detection accuracy, particularly in situations with a lot of occlusions.
Deepanshu Garg, Sivaraman Eswaran, Youcef Djenouri, Gautam Srivastava 0001
IJCNN4
2024 Enhancing smart road safety with federated learning for Near Crash Detection to advance the development of the Internet of Vehicles
abstract
We introduce an innovative methodology for the identification of vehicular collisions within Internet of Vehicles (IoV) applications. This approach combines a knowledge base system with deep learning for model selection in an ensemble learning setting. It is designed to provide a general near-crash detection capability without relying on domain-specific knowledge, enabling the development of generic deep learning models. Our proposed methodology employs a novel deep learning approach, wherein multiple learning models are individually trained for each image. Subsequently, visual features are computed and stored for each trained image, along with the associated loss values from the training phase. This stored information is utilized to select the most suitable models for processing new image data during the testing phase. To facilitate efficient model selection, we employ a kNN (k Nearest Neighbors) strategy. To enhance both data and model security in IoV environments, we implement an intelligent federated learning (FL) strategy. Users are organized into clusters, and we employ two distinct aggregation methods, departing from conventional federated learning approaches. In the initial stage, we aggregate model data from all users to create a global model representing collective knowledge. In the subsequent stage, we aggregate models from each cluster to generate customized local models. Users are provided with both global and local models, allowing them to select the most suitable model for their specific crash detection needs. We test our approach, that we call Knowledge Guided Deep Learning for Near Crash Detection (KGDL-NCD), on well-known NCD benchmarks. The results demonstrate that KGDL-NCD surpasses baseline solutions, achieving an AUC (Area Under Curve) metric of 0.95.
Youcef Djenouri, Ahmed Nabil Belbachir, Tomasz P. Michalak, Asma Belhadi, Gautam Srivastava 0001
Eng. Appl. Artif. Intell.1
2024 Artificial intelligence of medical things for disease detection using ensemble deep learning and attention mechanism
abstract
Abstract In this paper, we present a novel paradigm for disease detection. We build an artificial intelligence based system where various biomedical data are retrieved from distributed and homogeneous sensors. We use different deep learning architectures (VGG16, RESNET, and DenseNet) with ensemble learning and attention mechanisms to study the interactions between different biomedical data to detect and diagnose diseases. We conduct extensive testing on biomedical data. The results show the benefits of using deep learning technologies in the field of artificial intelligence of medical things to diagnose diseases in the healthcare decision‐making process. For example, the disease detection rate using the proposed methodology achieves 92%, which is greatly improved compared to the higher‐level disease detection models.
Youcef Djenouri, Asma Belhadi, Anis Yazidi, Gautam Srivastava 0001, Jerry Chun-Wei Lin
Expert Syst. J. Knowl. Eng.1
2024 A Federated Convolution Transformer for Fake News Detection
abstract
We present a novel approach to detecting fake news in Internet of Things (IoT) applications. By investigating federated learning and trusted authority methods, we address the issue of data security during training. Simultaneously, by investigating convolution transformers and user clustering, we deal with multi-modality in fake news data. Firstly, we use dense embedding and the k-means algorithm to cluster users into groups that are similar to one another. We then develop a local model for each user using their local data. The server then receives the local models of the users along with the clustering information, and a trusted authority verifies their integrity there. We use two different types of aggregation in place of conventional federated learning systems. The initial step is to combine all the users' models to create a single global model. The second step entails compiling each user's model into a local model of comparable users. Both models are supplied to the users, who then select the most suitable model for identifying fake news. By conducting extensive experiments using Twitter data, we demonstrate that the proposed method outperforms various baselines, where it achieves an average accuracy of 0.85 in comparison to others that do not exceed 0.81.
Youcef Djenouri, Ahmed Nabil Belbachir, Tomasz P. Michalak, Gautam Srivastava 0001
IEEE Trans. Big Data1
2024 A Secure Parallel Pattern Mining System for Medical Internet of Things
abstract
In this paper, a new generic parallel pattern mining framework called multi-objective Decomposition for Parallel Pattern-Mining (MD-PPM) is developed to solve challenges in the Internet of Medical Things through big data exploration. MD-PPM discovers important patterns by using decomposition and parallel mining methods to explore connectivity between medical data. First, a new technique, the multi-objective k-means algorithm, is used to aggregate medical data. A parallel pattern mining approach based on GPU and MapReduce architectures is also used to create useful patterns. To ensure complete privacy and security of the medical data, blockchain technology has been integrated throughout the system. Several tests were conducted to demonstrate the high performance of two sequential and graph pattern mining problems on large medical data and to evaluate the developed MD-PPM framework. From our results, our proposed MD-PPM has achieved strong results in terms of memory usage and computation time in terms of efficiency. Moreover, MD-PPM performs well in terms of accuracy and feasibility compared to existing models.
Youcef Djenouri, Asma Belhadi, Gautam Srivastava 0001, Jerry Chun-Wei Lin
IEEE Trans. Comput. Biol. Bioinform.1
2024 Social Web in IoT: Can Evolutionary Computation and Clustering Improve Ontology Matching for Social Web of Things?
abstract
Many Internet of Things (IoT) applications can benefit from Social Web of Things (S-WoT) methods that enable knowledge discovery and help solving interoperability problems. The semantic modeling of S-WoT is the main emphasis of this work where we suggest a novel solution, evolutionary clustering for ontology matching (ECOM), to explore correlations between S-WoT data using clustering and evolutionary computation methodologies. The ECOM approach uses a variety of clustering techniques to aggregate S-WoT data's strongly related ontologies into comparable categories. The principle is to match concepts of similar groups rather than full concepts of two ontologies, which necessitates splitting examples of each ontology into similar groups. We design two clustering algorithms for ontology matching using conventional methods, as well as sophisticated clustering techniques. Moreover, we develop an intelligent matching algorithm that uses evolutionary computation to quickly converge to (or ideally identify) optimal matches. Numerous simulations have been conducted using various ontology databases to demonstrate the application and precision of ECOM. Our findings clearly show that ECOM has better results when compared to cutting-edge ontology matching methods. The F-measure of ECOM exceeds 95% whereas it does not reach 90% for all baseline methods. The results also confirm that ECOM scales with big data in S-WoT environments.
Asma Belhadi, Djamel Djenouri, Youcef Djenouri, Ahmed Nabil Belbachir, Gautam Srivastava 0001
IEEE Trans. Comput. Soc. Syst.3
2024 DRL-Based URLLC-Constraint and Energy-Efficient Task Offloading for Internet of Health Things
abstract
Internet of Health Things (IoHT) is a promising e-Health paradigm that involves offloading numerous computational-intensive and delay-sensitive tasks from locally limited IoHT points to edge servers (ESs) with abundant computational resources in close proximity. However, existing computation offloading techniques struggle to meet the burgeoning health demands in ultra-reliable and low-latency communication (URLLC), one of the 5G application scenarios. This article proposes a Multi-Agent Soft-Actor-Critic-discrete based URLLC-constrained task offloading and resource allocation (MASACDUA) scheme to maximize throughput while minimizing power consumption on the remote side, considering the long-term URLLC constraints. The URLLC constraint conditions are formulated using extreme value theory, and Lyapunov optimization is employed to divide the problem into task offloading and computation resource allocation. MASAC-discrete and a queue backlog-aware algorithm are utilized to approach task offloading and computation resource allocation, respectively. Extensive simulation results demonstrate that MASACDUA outperforms traditional DRL algorithms under different IoHT points and data arrival rate intervals and achieves superior performance in delay, bound violation probability, and other characteristics related to URLLC.
Yixiao Wang 0002, Huaming Wu, Rutvij H. Jhaveri, Youcef Djenouri
IEEE J. Biomed. Health Informatics4
2024 An Efficient and Accurate GPU-based Deep Learning Model for Multimedia Recommendation
abstract
This article proposes the use of deep learning in human-computer interaction and presents a new explainable hybrid framework for recommending relevant hashtags on a set of orpheline tweets, which are tweets with hashtags. The approach starts by determining the set of batches used in the convolution neural network based on frequent pattern mining solutions. The convolutional neural network is then applied to the set of batches of tweets to learn the hashtags of the tweets. An optimization strategy has been proposed to accurately perform the learning process by reducing the number of frequent patterns. Moreover, eXplainable AI is introduced for hashtag recommendations by analyzing the user preferences and understanding the different weights of the deep learning model used in the learning process. This is performed by learning the hyper-parameters of the deep architecture using the genetic algorithm. GPU computing is also investigated to achieve high speed and enable the execution of the overall framework in real time. Extensive experimental analysis has been performed to show that our methodology is useful on different collections of tweets. The experimental results clearly show the efficiency of our proposed approach compared to baseline approaches in terms of both runtime and accuracy. Thus, the proposed solution achieves an accuracy of 90% when analyzing complex Wikipedia data while the other algorithms did not achieve 85% when processing the same amount of data.
Youcef Djenouri, Asma Belhadi, Gautam Srivastava 0001, Jerry Chun-Wei Lin
ACM Trans. Multim. Comput. Commun. Appl.1
2024 TG-SPRED: Temporal Graph for Sensorial Data PREDiction
abstract
This study introduces an innovative method aimed at reducing energy consumption in sensor networks by predicting sensor data, thereby extending the network’s operational lifespan. Our model, Temporal Graph Sensor Prediction (TG-SPRED), predicts readings for a subset of sensors designated to enter sleep mode in each time slot, based on a non-scheduling-dependent approach. This flexibility allows for extended sensor inactivity periods without compromising data accuracy. TG-SPRED addresses the complexities of event-based sensing—a domain that has been somewhat overlooked in existing literature—by recognizing and leveraging the inherent temporal and spatial correlations among events. It combines the strengths of Gated Recurrent Units and Graph Convolutional Networks to analyze temporal data and spatial relationships within the sensor network graph, where connections are defined by sensor proximities. An adversarial training mechanism, featuring a critic network employing the Wasserstein distance for performance measurement, further refines the predictive accuracy. Comparative analysis against six leading solutions using four critical metrics—F-score, energy consumption, network lifetime, and computational efficiency—showcases our approach’s superior performance in both accuracy and energy efficiency.
Roufaida Laidi, Djamel Djenouri, Youcef Djenouri, Jerry Chun-Wei Lin
ACM Trans. Sens. Networks3
2023 Empowering Search and Rescue Operations with Big Data Technology: A Comprehensive Study of YOLOv8 Transfer Learning for Transportation Safety
abstract
In this research work, we demonstrate the important role of object detection technology and how to optimize it for elevating the efficiency of accident rescue missions in the maritime transportation industry. In this context, the use of unmanned aerial vehicles in search and rescue missions is a promising research topic. However, the processing power limitations and the lack of data focused on sea operations present some challenges in this area. The mix of synthetic and real data during the training and even the total replacement of real data with virtual generated ones can lead to a good and flexible solution for the dataset challenges. Another strategy for dealing with the lack of data is the use of transfer learning for leveraging the knowledge in a domain with an abundance of data when compared to a new domain of interest. In this work, the use of transfer learning and synthetic and real mixed datasets is explored for the field of search and rescue. The YOLOv8 is trained in different configurations of regular learning and transfer learning, with fine-tuning and 4 and 7 frozen layers, using both synthetic and real data. Finally, the models and the set of data are evaluated based on mAP50-95 showing some possible reasons for a performance difference between real and synthetic data in the training process.
Luciano Netto de Lima, Fabio A. A. Andrade, Youcef Djenouri, Carlos Pfeiffer, Marcos Moura
IEEE Big Data3
2023 Knowledge Guided Deep Learning for General-Purpose Computer Vision Applications
Youcef Djenouri, Ahmed Nabil Belbachir, Rutvij H. Jhaveri, Djamel Djenouri
CAIP (1)1
2023 Empowering Urban Connectivity in Smart Cities using Federated Intrusion Detection
abstract
The advent of transformative technologies such as the Internet of Things (IoT) has brought forth significant advancements in various sectors like smart cities, fintech, learning, and healthcare, as well as revolutionized online activities. The IoT has facilitated widespread connectivity by interconnecting numerous objects and services, but it has also made IoT and cloud infrastructures susceptible to cyberattacks, making cybersecurity a paramount concern, particularly for the development of reliable IoT systems, especially those powering smart city networks. In this research endeavor, we embark on exploring a cutting-edge pipeline that amalgamates federated deep learning with a trusted authority approach to tackle the intricate challenges associated with intrusion detection in smart city networks. To identify anomalies and intrusions effectively within the network, we devise an improved LSTM (Long Short-Term Memory) model. Additionally, we propose an intelligent swarm optimization solution to address dimensionality reduction concerns. Thorough evaluations of our federated learning-based approach are conducted, and these are juxtaposed with several basic approaches, utilizing the renowned NSL-KDD dataset. Encouragingly, our findings reveal that the proposed framework remarkably outperforms the baseline solutions, particularly when dealing with datasets containing a substantial volume of transactions. Furthermore, our method ensures robust data security for the model, as it becomes the pioneering endeavor to incorporate the principle of trusted authority into the realm of federated learning for the management of smart city networks.
Youcef Djenouri, Ahmed Nabil Belbachir
DSAA1
2023 Hybrid Genetic U-Net Algorithm for Medical Segmentation
Jon-Olav Holland, Youcef Djenouri, Roufaida Laidi, Anis Yazidi
ICAART (3)2
2023 Generating Event Sensor Readings Using Spatial Correlations and a Graph Sensor Adversarial Model for Energy Saving in IoT: GSAVES
abstract
This work targets a comprehensive model enabling energy-constrained IoT (Internet of Things) sensor devices to be inactive for extended periods while estimating their readings of real-time events. Although events seem semantically uncoupled, they are usually spatially and temporally related. We propose GSAVES (Graph Sensor AdVersarial for Energy Saving), which uses readings from active devices and spatial correlations to generate the missing data due to sensor inactivity. The missing readings are generated with Graph Convolutional Network (GCN) that learns embeddings from data and the graph structure. GSAVES is evaluated against four state-of-the-art solutions using three network sizes and four performance metrics. The results demonstrate the efficiency of GSAVES for providing the best balance between the considered metrics, outperforming all the solutions in reducing energy consumption and improving accuracy.
Roufaida Laidi, Djamel Djenouri, Miloud Bagaa, Lyes Khelladi, Youcef Djenouri
PIMRC5
2023 Interpretable intrusion detection for next generation of Internet of Things
abstract
This paper presents a new framework for intrusion detection in the next-generation Internet of Things. MinMax normalization strategy is used to collect and preprocess data. The Marine Predator algorithm is then used to select relevant features to be used in the learning process. The selected features are then trained with an advanced and state-of-the-art recurrent neural network that includes an attention mechanism. Finally, Shapely values are calculated to determine how much each feature contributes to the final output. The dataset NSL-KDD was used for intensive simulations. The results show the advantages of the proposed system as well as its superiority over state-of-the-art methods. In fact, the proposed solution achieved a rate of more than 94% for both true negative and true position, while the rates of the existing solutions are below 90% for the challenging NSL-KDD datasets.
Youcef Djenouri, Asma Belhadi, Gautam Srivastava 0001, Jerry Chun-Wei Lin, Anis Yazidi
Comput. Commun.1
2023 Hybrid graph convolution neural network and branch-and-bound optimization for traffic flow forecasting
abstract
In this study, we combine graph optimization and prediction in a single pipeline to investigate an innovative convolutional graph-based neural network for urban traffic flow prediction in an edge IoT environment. Pre-processing of the linked graph is first performed to remove noise from the set of original road networks of urban traffic data. Outlier detection strategy is used to efficiently explore the road network and remove irrelevant patterns and noise. The resulting graph is then implemented to train an extended graph convolutional neural network to estimate the traffic flow in the city. To accurately tune the hyperparameter values of the proposed framework, a new optimization technique is developed based on branch and bound. For comparison, an intensive evaluation is conducted with multiple datasets and baseline methods. The results show that the proposed framework outperforms the baseline solutions, especially when the number of nodes in the graph is large.
Youcef Djenouri, Asma Belhadi, Gautam Srivastava 0001, Jerry Chun-Wei Lin
Future Gener. Comput. Syst.1
2023 Federated deep learning for smart city edge-based applications
Youcef Djenouri, Tomasz P. Michalak, Jerry Chun-Wei Lin
Future Gener. Comput. Syst.1
2023 Fast and Accurate Deep Learning Framework for Secure Fault Diagnosis in the Industrial Internet of Things
abstract
This article introduced a new deep learning framework for fault diagnosis in electrical power systems. The framework integrates the convolution neural network and different regression models to visually identify which faults have occurred in electric power systems. The approach includes three main steps: 1) data preparation; 2) object detection; and 3) hyperparameter optimization. Inspired by deep learning and evolutionary computation (EC) techniques, different strategies have been proposed in each step of the process. In addition, we propose a new hyperparameters optimization model based on EC that can be used to tune parameters of our deep learning framework. In the validation of the framework’s usefulness, experimental evaluation is executed using the well known and challenging VOC 2012, the COCO data sets, and the large NESTA 162-bus system. The results show that our proposed approach significantly outperforms most of the existing solutions in terms of runtime and accuracy.
Youcef Djenouri, Asma Belhadi, Gautam Srivastava 0001, Uttam Ghosh, Pushpita Chatterjee, Jerry Chun-Wei Lin
IEEE Internet Things J.1
2023 Emergent Deep Learning for Anomaly Detection in Internet of Everything
abstract
This research presents a new generic deep learning (DL) framework for anomaly detection in the Internet of Everything (IoE). It combines decomposition methods, deep neural networks, and evolutionary computation to better detect outliers in IoE environments. The data set is first decomposed into clusters, while similar observations in the same cluster are grouped. Five clustering algorithms were used for this purpose. The generated clusters are then trained using DL architectures. In this context, we propose a new recurrent neural network for training time-series data. Two evolutionary computational algorithms are also proposed: 1) the genetic and 2) the bee swarm, to fine-tune the training step. These algorithms consider the hyperparameters of the trained models and try to find the optimal values. The proposed solutions have been experimentally evaluated for two use cases: 1) road traffic outlier detection and 2) network intrusion detection. The results show the advantages of the proposed solutions and a clear superiority compared to state-of-the-art approaches.
Youcef Djenouri, Djamel Djenouri, Asma Belhadi, Gautam Srivastava 0001, Jerry Chun-Wei Lin
IEEE Internet Things J.1
2023 High-Precision Matching Algorithm for Multi-Image Segmentation of Micro Animation Videos in Mobile Network Environment
abstract
Abstract In the mobile network environment, the accuracy of related image matching algorithms is affected by factors such as bandwidth uncertainty and channel interference, resulting in significant limitations in image feature matching. This article designs a high-precision matching algorithm for multi-image segmentation of micro animation videos in mobile network environments. Fully denoise micro animation video images using 2D High Density Discrete Wavelet Transform (HD-DWT), and apply fixed block count segmentation to process micro animation video images; Using Harris algorithm to complete image corner detection and obtain corner features of sub images; In the K-means clustering algorithm, SIFT feature vectors are divided into clusters and paired with the nearest neighbor cluster in another sub image to form a sub image matching pair, completing block based sub image matching; Combine all sub image matching results to obtain video image matching results, and use the Improved Random Sampling Consistency (RANCAS) algorithm to remove incorrect matching during the matching process, improving matching accuracy. The experimental results show that the designed algorithm can effectively reduce image noise, improve image quality, and generate a large number of matching pairs in mobile network environments. After the application of the designed algorithm, the production effect of micro animated videos in mobile networks can be significantly improved.
Yehui Su, Youcef Djenouri
Mob. Networks Appl.2
2023 Fast and Accurate Framework for Ontology Matching in Web of Things
abstract
The Web of Things (WoT) can help with knowledge discovery and interoperability issues in many Internet of Things (IoT) applications. This article focuses on semantic modeling of WoT and proposes a new approach called Decomposition for Ontology Matching (DOM) to discover relevant knowledge by exploring correlations between WoT data using decomposition strategies. The DOM technique adopts several decomposition techniques to order highly linked ontologies of WoT data into similar groups. The main idea is to decompose the instances of each ontology into similar groups and then match instances of similar groups instead of entire instances of two ontologies. Three main algorithms for decomposition have been developed. The first algorithm is based on radar scanning, which determines the distribution of distances between each instance and all other instances to determine the cluster centroid. The second algorithm is based on adaptive grid clustering, where it focuses on distribution information and the construction of spanning trees. The third algorithm is based on split index clustering, where instances are divided into groups of cells from which noise is removed during the merging process. Several studies were conducted with different ontology databases to illustrate the use of the DOM technique. The results show that DOM outperforms state-of-the-art ontology matching models in terms of computational cost while maintaining the quality of the matching. Moreover, these results demonstrate that DOM is capable of handling various large datasets in WoT contexts.
Asma Belhadi, Youcef Djenouri, Gautam Srivastava 0001, Jerry Chun-Wei Lin
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2023 Advanced Pattern-Mining System for Fake News Analysis
abstract
Decomposition MapReduce mining for fake news analysis (DMRM-FNA), a novel generic parallel pattern-mining framework, is developed in this article to solve difficulties in social network analysis using big data exploration. The first difficulty faced by existing techniques is the inability to retrieve actionable insights into the structure of fake news data. This can be solved by extracting patterns from fake news and matching them with real news data. The second difficulty is the computational time of existing pattern-mining solutions. This might be solved by combining both decomposition and MapReduce mining techniques to extract relevant patterns from fake news data. The multiobjective${k}$-means algorithm is used to first aggregate fake news data. To generate useful patterns, a parallel pattern-mining method based on MapReduce structures is applied. To evaluate the created DMRM-FNA framework (DFAST) and demonstrate the high performance of sequential pattern-mining challenges on massive social network data, several tests were conducted. Our results show that the proposed DMRM-FNA performs well in terms of memory usage and efficiency. Moreover, DMRM-FNA outperforms existing models in terms of accuracy and feasibility.
Youcef Djenouri, Asma Belhadi, Gautam Srivastava 0001, Jerry Chun-Wei Lin
IEEE Trans. Comput. Soc. Syst.1
2023 Intelligent Deep Fusion Network for Anomaly Identification in Maritime Transportation Systems
abstract
This paper introduces a novel deep learning architecture for identifying outliers in the context of intelligent transportation systems. The use of a convolutional neural network with decomposition is explored to find abnormal behavior in maritime data. The set of maritime data is first decomposed into similar clusters containing homogeneous data, and then a convolutional neural network is used for each data cluster. Different models are trained (one per cluster), and each model is learned from highly correlated data. Finally, the results of the models are merged using a simple but efficient fusion strategy. To verify the performance of the proposed framework, intensive experiments were conducted on marine data. The results show the superiority of the proposed framework compared to the baseline solutions in terms of several accuracy metrics.
Youcef Djenouri, Asma Belhadi, Djamel Djenouri, Gautam Srivastava 0001, Jerry Chun-Wei Lin
IEEE Trans. Intell. Transp. Syst.1
2023 A Secure Intelligent System for Internet of Vehicles: Case Study on Traffic Forecasting
abstract
Significant efforts have been made for vehicle-to-vehicle communications that now enable the Internet of Vehicles (IoV). However, current IoV solutions are unable to capture traffic data both accurately and securely. Another drawback of current IoV models that are based on deep learning is that the methods used do not tune hyperparameters efficiently. In this paper, a new system known as Secure and Intelligent System for the Internet of Vehicles (SISIV) is developed. A deep learning architecture based on graph convolutional networks and an attention mechanism are implemented. In addition, blockchain technology is used to protect data transmission between nodes in the IoV system. Moreover, the hyperparameters of the generated deep learning model are intelligently selected using a branch-and-bound technique. To validate SISIV, experiments were conducted on four networked vehicle databases dealing with prediction problems. In terms of forecasting rate ($>$90%), F-measure ($>$80%), and attack detection (< 75%), the results clearly show the superiority of SISIV over baseline systems. Moreover, compared to state-of-the-art solutions based on traffic prediction, SISIV enables efficient and reliable prediction of traffic flow in an IoV context.
Youcef Djenouri, Asma Belhadi, Djamel Djenouri, Gautam Srivastava 0001, Jerry Chun-Wei Lin
IEEE Trans. Intell. Transp. Syst.1
2023 Intelligent Graph Convolutional Neural Network for Road Crack Detection
abstract
This paper presents a novel intelligent system based on graph convolutional neural networks to study road crack detection in intelligent transportation systems. The visual features of the input images are first computed using the well-known Scale-Invariant Feature Transform (SIFT) extraction algorithm. Then, a correlation between SIFT features of similar images is analyzed and a series of graphs are generated. The graphs are trained on a graph convolutional neural network, and a hyper-optimization algorithm is developed to supervise the training process. A case study of road crack detection data is analyzed. The results show a clear superiority of the proposed framework over state-of-the-art solutions. In fact, the precision of the proposed solution exceeds 70%, while the precision of the baseline methods does not exceed 60%.
Youcef Djenouri, Asma Belhadi, Essam H. Houssein, Gautam Srivastava 0001, Jerry Chun-Wei Lin
IEEE Trans. Intell. Transp. Syst.1
2023 Towards an Advanced Deep Learning for the Internet of Behaviors: Application to Connected Vehicles
abstract
In recent years, intensive research has been conducted to enable people to live more comfortably. Developments in the Internet of Things (IoT) , big data, and artificial intelligence have taken this type of research to a new level and led to the emergence of the Internet of Behaviors (IoB) , which analyzes behavioral patterns. However, current IoB technologies are not capable of handling heterogeneous data. While it is quite common to have different formats of sensor data for the same behavioral observation, the use of these different data formats can significantly help to obtain a more accurate classification of the observation. Another limitation is that existing IoB deep learning models rely on inefficient hyperparameter tuning strategies. In this paper, we present an Advanced Deep Learning framework for IoB (ADLIoB) applied to connected vehicles. Several deep learning architectures are employed in this framework: CNN, Graph CNN (GCNN), and LSTM are used to train sensor data of different formats. In addition, a branch-and-bound technique is used to intelligently select hyperparameters. To validate ADLIoB, experiments were conducted on four databases for connected vehicles. The results clearly show that ADLIoB is superior to the baseline solutions in terms of both accuracy and runtime.
Tinhinane Mezair, Youcef Djenouri, Asma Belhadi, Gautam Srivastava 0001, Jerry Chun-Wei Lin
ACM Trans. Sens. Networks2
2022 How Image Retrieval and Matching Can Improve Object Localisation on Offshore Platforms
Youcef Djenouri, Jon Mikkelsen Hjelmervik, Elias Bjørne, Milad Mobarhan
IDEAL1
2022 Deep learning based decomposition for visual navigation in industrial platforms
abstract
Abstract In the heavy asset industry, such as oil & gas, offshore personnel need to locate various equipment on the installation on a daily basis for inspection and maintenance purposes. However, locating equipment in such GPS denied environments is very time consuming due to the complexity of the environment and the large amount of equipment. To address this challenge we investigate an alternative approach to study the navigation problem based on visual imagery data instead of current ad-hoc methods where engineering drawings or large CAD models are used to find equipment. In particular, this paper investigates the combination of deep learning and decomposition for the image retrieval problem which is central for visual navigation. A convolutional neural network is first used to extract relevant features from the image database. The database is then decomposed into clusters of visually similar images, where several algorithms have been explored in order to make the clusters as independent as possible. The Bag-of-Words (BoW) approach is then applied on each cluster to build a vocabulary forest. During the searching process the vocabulary forest is exploited to find the most relevant images to the query image. To validate the usefulness of the proposed framework, intensive experiments have been carried out using both standard datasets and images from industrial environments. We show that the suggested approach outperforms the BoW-based image retrieval solutions, both in terms of computing time and accuracy. We also show the applicability of this approach on real industrial scenarios by applying the model on imagery data from offshore oil platforms.
Youcef Djenouri, Johan Hatleskog, Jon Mikkelsen Hjelmervik, Elias Bjørne, Trygve Utstumo, Milad Mobarhan
Appl. Intell.1
2022 An edge-driven multi-agent optimization model for infectious disease detection
abstract
This research work introduces a new intelligent framework for infectious disease detection by exploring various emerging and intelligent paradigms. We propose new deep learning architectures such as entity embedding networks, long-short term memory, and convolution neural networks, for accurately learning heterogeneous medical data in identifying disease infection. The multi-agent system is also consolidated for increasing the autonomy behaviours of the proposed framework, where each agent can easily share the derived learning outputs with the other agents in the system. Furthermore, evolutionary computation algorithms, such as memetic algorithms, and bee swarm optimization controlled the exploration of the hyper-optimization parameter space of the proposed framework. Intensive experimentation has been established on medical data. Strong results obtained confirm the superiority of our framework against the solutions that are state of the art, in both detection rate, and runtime performance, where the detection rate reaches 98% for handling real use cases.
Youcef Djenouri, Gautam Srivastava 0001, Anis Yazidi, Jerry Chun-Wei Lin
Appl. Intell.1
2022 Efficient evolutionary computation model of closed high-utility itemset mining
Jerry Chun-Wei Lin, Youcef Djenouri, Gautam Srivastava 0001, Philippe Fournier-Viger
Appl. Intell.2
2022 An ontology matching approach for semantic modeling: A case study in smart cities
abstract
Abstract This paper investigates the semantic modeling of smart cities and proposes two ontology matching frameworks, called Clustering for Ontology Matching‐based Instances (COMI) and Pattern mining for Ontology Matching‐based Instances (POMI). The goal is to discover the relevant knowledge by investigating the correlations among smart city data based on clustering and pattern mining approaches. The COMI method first groups the highly correlated ontologies of smart‐city data into similar clusters using the generic k‐means algorithm. The key idea of this method is that it clusters the instances of each ontology and then matches two ontologies by matching their clusters and the corresponding instances within the clusters. The POMI method studies the correlations among the data properties and selects the most relevant properties for the ontology matching process. To demonstrate the usefulness and accuracy of the COMI and POMI frameworks, several experiments on the DBpedia, Ontology Alignment Evaluation Initiative, and NOAA ontology databases were conducted. The results show that COMI and POMI outperform the state‐of‐the‐art ontology matching models regarding computational cost without losing the quality during the matching process. Furthermore, these results confirm the ability of COMI and POMI to deal with heterogeneous large‐scale data in smart‐city environments.
Youcef Djenouri, Hiba Belhadi, Karima Akli-Astouati, Alberto Cano 0001, Jerry Chun-Wei Lin
Comput. Intell.1
2022 Intelligent deep fusion network for urban traffic flow anomaly identification
abstract
This paper presents a novel deep learning architecture for identifying outliers in the context of intelligent transportation systems. The use of a convolutional neural network with an efficient decomposition strategy is explored to find the anomalous behavior of urban traffic flow data. The urban traffic flow data set is decomposed into similar clusters, each containing homogeneous data. The convolutional neural network is used for each data cluster. In this way, different models are trained, each learned from highly correlated data. A merging strategy is finally used to fuse the results of the obtained models. To validate the performance of the proposed framework, intensive experiments were conducted on urban traffic flow data. The results show that our system outperforms the competition on several accuracy criteria.
Youcef Djenouri, Asma Belhadi, Hsing-Chung Chen, Jerry Chun-Wei Lin
Comput. Commun.1
2022 A sustainable deep learning framework for fault detection in 6G Industry 4.0 heterogeneous data environments
abstract
The integration of 5G and Beyond 5G (B5G)/6G in Machine-to-Machine (M2M) communications, is making Industry 4.0 smarter. However, the goal of having a sustainable self-monitored industry has not been reached yet. State-of-the-art deep learning-based Fault Detection algorithms cannot handle heterogeneous data, meaning that more than one fault detection computational device has to be used for each data format, in addition to the inability to take advantage of the combination of all the information available in different formats to derive more accurate conclusions. Moreover, these algorithms rely on inefficient hyper-parameters tuning strategies. In this paper, we propose an Advanced Deep Learning framework for Fault Diagnosis in Industry 4.0 (ADL-FDI4), which combines Long Short Term Memory (LSTM), Convolutional Neural Networks (CNN) and graph CNN (GNN), to handle heterogeneous data. Furthermore, our novel framework uses a Branch-and-Bound procedure to guide the learning process. Our experimental results show that ADL-FDI4 outperforms the state-of-the-art solutions in terms of detection rate and running time, and for that, it consumes less energy. In addition to handling heterogeneous data, which implies that one computational device is sufficient to handle all data formats.
Tinhinane Mezair, Youcef Djenouri, Asma Belhadi, Gautam Srivastava 0001, Jerry Chun-Wei Lin
Comput. Commun.2
2022 Hybrid intelligent framework for automated medical learning
abstract
Abstract This paper investigates the automated medical learning and proposes hybrid intelligent framework, called Hybrid Automated Medical Learning (HAML). The goal is the efficient combination of several intelligent components in order to automatically learn the medical data. Multi agents system is proposed by using distributed deep learning, and knowledge graph for learning medical data. The distributed deep learning is used for efficient learning of the different agents in the system, where the knowledge graph is used for dealing with heterogeneous medical data. To demonstrate the usefulness and accuracy of the HAML framework, intensive simulations on medical data were conducted. A wide range of experiments were conducted to verify the efficiency of the proposed system. Three case studies are discussed in this research, the first case study is related to process mining, and more precisely on the ability of HAML to detect relevant patterns from event medical data. The second case study is related to smart building, and the ability of HAML to recognize the different activities of the patients. The third one is related to medical image retrieval, and the ability of HAML to find the most relevant medical images according to the image query. The results show that the developed HAML achieves good performance compared to the most up‐to‐date medical learning models regarding both the computational and cost the quality of returned solutions.
Asma Belhadi, Youcef Djenouri, Vicente García-Díaz, Essam H. Houssein, Jerry Chun-Wei Lin
Expert Syst. J. Knowl. Eng.2
2022 Sensor data fusion for the industrial artificial intelligence of things
abstract
Abstract The emergence of smart sensors, artificial intelligence, and deep learning technologies yield artificial intelligence of things, also known as the AIoT. Sophisticated cooperation of these technologies is vital for the effective processing of industrial sensor data. This paper introduces a new framework for addressing the different challenges of the AIoT applications. The proposed framework is an intelligent combination of multi‐agent systems, knowledge graphs and deep learning. Deep learning architectures are used to create models from different sensor‐based data. Multi‐agent systems can be used for simulating the collective behaviours of the smart sensors using IoT settings. The communication among different agents is realized by integrating knowledge graphs. Different optimizations based on constraint satisfaction as well as evolutionary computation are also investigated. Experimental analysis is undertaken to compare the methodology presented to state‐of‐the‐art AIoT technologies. We show through experimentation that our designed framework achieves good performance compared to baseline solutions.
Youcef Djenouri, Asma Belhadi, Gautam Srivastava 0001, Essam H. Houssein, Jerry Chun-Wei Lin
Expert Syst. J. Knowl. Eng.1
2022 Deep learning based hashtag recommendation system for multimedia data
abstract
This work aims to provide a novel hybrid architecture to suggest appropriate hashtags to a collection of orpheline tweets. The methodology starts with defining the collection of batches used in the convolutional neural network. This methodology is based on frequent pattern extraction methods. The hashtags of the tweets are then learned using the convolution neural network that was applied to the collection of batches of tweets. In addition, a pruning approach should ensure that the learning process proceeds properly by reducing the number of common patterns. Besides, the evolutionary algorithm is involved to extract the optimal parameters of the deep learning model used in the learning process. This is achieved by using a genetic algorithm that learns the hyper-parameters of the deep architecture. The effectiveness of our methodology has been demonstrated in a series of detailed experiments on a set of Twitter archives. From the results of the experiments, it is clear that the proposed method is superior to the baseline methods in terms of efficiency.
Youcef Djenouri, Asma Belhadi, Gautam Srivastava 0001, Jerry Chun-Wei Lin
Inf. Sci.1
2022 Vehicle detection using improved region convolution neural network for accident prevention in smart roads
abstract
This paper explores the vehicle detection problem and introduces an improved regional convolution neural network. The vehicle data (set of images) is first collected, from which the noise (set of outlier images) is removed using the SIFT extractor. The region convolution neural network is then used to detect the vehicles. We propose a new hyper-parameters optimization model based on evolutionary computation that can be used to tune parameters of the deep learning framework. The proposed solution was tested using the well-known boxy vehicle detection data, which contains more than 200,000 vehicle images and 1,990,000 annotated vehicles. The results are very promising and show superiority over many current state-of-the-art solutions in terms of runtime and accuracy performances.
Youcef Djenouri, Asma Belhadi, Gautam Srivastava 0001, Djamel Djenouri, Jerry Chun-Wei Lin
Pattern Recognit. Lett.1
2022 Toward a Cognitive-Inspired Hashtag Recommendation for Twitter Data Analysis
abstract
This research investigates hashtag suggestions in a heterogeneous and huge social network, as well as a cognitive-based deep learning solution based on distributed knowledge graphs. Community detection is first performed to find the connected communities in a vast and heterogeneous social network. The knowledge graph is subsequently generated for each discovered community, with an emphasis on expressing the semantic relationships among the Twitter platform’s user communities. Each community is trained with the embedded deep learning model. To recommend hashtags for the new user in the social network, the correlation between the tweets of such user and the knowledge graph of each community is explored to set the relevant communities of such user. The models of the relevant communities are used to infer the hashtags of the tweets of such users. We conducted extensive testing to demonstrate the usefulness of our methods on a variety of tweet collections. Experimental results show that the proposed approach is more efficient than the baseline approaches in terms of both runtime and accuracy.
Youcef Djenouri, Asma Belhadi, Gautam Srivastava 0001, Jerry Chun-Wei Lin
IEEE Trans. Comput. Soc. Syst.1
2022 Deep Learning Versus Traditional Solutions for Group Trajectory Outliers
abstract
This article introduces a new model to identify a group of trajectory outliers from a large trajectory database and proposes several algorithms. These can be split into three categories: 1) algorithms based on data mining and knowledge discovery, which study the different correlations among the trajectory data and identify the group of abnormal trajectories from the knowledge extracted; 2) algorithms based on machine learning and computational intelligence methods, which use the ensemble learning and metaheuristics to find the group of trajectory outliers; and 3) an algorithm exploring the convolution deep neural network that learns the different features of historical data to determine the group of trajectory outliers. Experiments on different trajectory databases have been carried out to investigate the proposed algorithms. The results show that the deep learning solution outperforms data mining, machine learning, and computational intelligence solutions, as well as state-of-the-art solutions in terms of runtime and accuracy performance.
Asma Belhadi, Youcef Djenouri, Djamel Djenouri, Tomasz P. Michalak, Jerry Chun-Wei Lin
IEEE Trans. Cybern.2
2022 Secure Collaborative Augmented Reality Framework for Biomedical Informatics
abstract
Augmented reality is currently of interest in biomedical health informatics. At the same time, several challenges have appeared, in particular with the rapid progress of smart sensor technologies, and medical artificial intelligence. This yields the necessity of new needs in biomedical health informatics. Collaborative learning and privacy are just some of the challenges of augmented reality technology in biomedical health informatics. This paper introduces a novel secure collaborative augmented reality framework for biomedical health informatics-based applications. Distributed deep learning is performed across a multi-agent system platform. The privacy strategy is then developed for ensuring better communications of the different intelligent agents in the system. In this research work, a system of multiple agents is created for the simulation of the collective behaviours of the smart components of biomedical health informatics. Augmented reality is also incorporated for better visualization of medical patterns. A novel privacy strategy based on blockchain is investigated for ensuring the confidentiality of the learning process. Experiments are conducted on real use cases of the biomedical segmentation process. Our strong experimental analysis reveals the strength of the proposed framework when directly compared to state-of-the-art biomedical health informatics solutions.
Youcef Djenouri, Asma Belhadi, Gautam Srivastava 0001, Jerry Chun-Wei Lin
IEEE J. Biomed. Health Informatics1
2022 Deviation Point Curriculum Learning for Trajectory Outlier Detection in Cooperative Intelligent Transport Systems
abstract
Cooperative Intelligent Transport Systems (C-ITS) are emerging in the field of transportation systems, which can be used to provide safety, sustainability, efficiency, communication and cooperation between vehicles, roadside units, and traffic command centres. With improved network structure and traffic mobility, a large amount of trajectory-based data is generated. Trajectory-based knowledge graphs help to give semantic and interconnection capabilities for intelligent transport systems. Prior works consider trajectory as the single point of deviation for the individual outliers. However, in real-world transportation systems, trajectory outliers can be seen in the groups, e.g., a group of vehicles that deviates from a single point based on the maintenance of streets in the vicinity of the intelligent transportation system. In this paper, we propose a trajectory deviation point embedding and deep clustering method for outlier detection. We first initiate network structure and nodes’ neighbours to construct a structural embedding by preserving nodes relationships. We then implement a method to learn the latent representation of deviation points in road network structures. A hierarchy multilayer graph is designed with a biased random walk to generate a set of sequences. This sequence is implemented to tune the node embeddings. After that, embedding values of the node were averaged to get the trip embedding. Finally, LSTM-based pairwise classification method is initiated to cluster the embedding with similarity-based measures. The results obtained from the experiments indicate that the proposed learning trajectory embedding captured structural identity and increasedF-measureby 5.06% and 2.4% while compared with genericNode2VecandStruct2Vecmethods.
Usman Ahmed, Gautam Srivastava 0001, Youcef Djenouri, Jerry Chun-Wei Lin
IEEE Trans. Intell. Transp. Syst.3
2022 Hybrid Group Anomaly Detection for Sequence Data: Application to Trajectory Data Analytics
abstract
Many research areas depend on group anomaly detection. The use of group anomaly detection can maintain and provide security and privacy to the data involved. This research attempts to solve the deficiency of the existing literature in outlier detection thus a novel hybrid framework to identify group anomaly detection from sequence data is proposed in this paper. It proposes two approaches for efficiently solving this problem: i)Hybrid Data Mining-based algorithm, consists of three main phases: first, the clustering algorithm is applied to derive the micro-clusters. Second, the$kNN$algorithm is applied to each micro-cluster to calculate the candidates of the group’s outliers. Third, a pattern mining framework gets applied to the candidates of the group’s outliers as a pruning strategy, to generate the groups of outliers, and ii) aGPU-basedapproach is presented, which benefits from the massively GPU computing to boost the runtime of the hybrid data mining-based algorithm. Extensive experiments were conducted to show the advantages of different sequence databases of our proposed model. Results clearly show the efficiency of a GPU direction when directly compared to a sequential approach by reaching a speedup of451. In addition, both approaches outperform the baseline methods for group detection.
Asma Belhadi, Youcef Djenouri, Gautam Srivastava 0001, Alberto Cano 0001, Jerry Chun-Wei Lin
IEEE Trans. Intell. Transp. Syst.2
2022 Hybrid RESNET and Regional Convolution Neural Network Framework for Accident Estimation in Smart Roads
abstract
Road safety is tackled and an intelligent deep learning framework is proposed in this work, which includes outlier detection, vehicle detection, and accident estimation. The road state is first collected, while an intelligent filter, based on SIFT extractor and a Chinese restaurant process is used to remove noise. The extended region-based convolution neural network is then applied to identify the closest vehicles to the given driver. The residual network will benefit from the vehicle detection process to make a binary classification on whether the current road state might cause an accident or not. Finally, we propose a novel optimization model for optimizing hyper-parameters in deep learning methodologies by using evolutionary computation. The proposed solution has been tested using benchmark vehicle detection and accident estimation datasets. The results are very promising and show superiority over many current state-of-the-art solutions in terms of runtime and accuracy, where the proposed solution has more than 5% of improved accident estimation rate compared to the conventional methods.
Youcef Djenouri, Gautam Srivastava 0001, Djamel Djenouri, Asma Belhadi, Jerry Chun-Wei Lin
IEEE Trans. Intell. Transp. Syst.1
2022 Scalable Mining of High-Utility Sequential Patterns With Three-Tier MapReduce Model
abstract
High-utility sequential pattern mining (HUSPM) is a hot research topic in recent decades since it combines both sequential and utility properties to reveal more information and knowledge rather than the traditional frequent itemset mining or sequential pattern mining. Several works of HUSPM have been presented but most of them are based on main memory to speed up mining performance. However, this assumption is not realistic and not suitable in large-scale environments since in real industry, the size of the collected data is very huge and it is impossible to fit the data into the main memory of a single machine. In this article, we first develop a parallel and distributed three-stage MapReduce model for mining high-utility sequential patterns based on large-scale databases. Two properties are then developed to hold the correctness and completeness of the discovered patterns in the developed framework. In addition, two data structures called sidset and utility-linked list are utilized in the developed framework to accelerate the computation for mining the required patterns. From the results, we can observe that the designed model has good performance in large-scale datasets in terms of runtime, memory, efficiency of the number of distributed nodes, and scalability compared to the serial HUSP-Span approach.
Jerry Chun-Wei Lin, Youcef Djenouri, Gautam Srivastava 0001, Yuanfa Li, Philip S. Yu
ACM Trans. Knowl. Discov. Data2
2021 Detection of Trajectory Outliers in Intelligent Transportation Systems
abstract
In this paper, we provide a technique for identifying outliers based on embedding trajectory deviation points and deep clustering. We begin by constructing the network topology and the neighbors of the nodes to create a structural embedding while capturing the interactions of the nodes. We then develop a strategy to determine the hidden representation of distraction points in the road network topology. To create a collection of sequences from a hierarchical multilayer network, a biased random walk is used. This sequence is used to fine tune the embedding of the nodes. The trip embedding was then determined by averaging the node embedding values. Finally, the embeddings are clustered using an LSTM-based pairwise classification strategy based on similarity metrics. The experimental results show that compared to the generic techniques Node2Vec and Struct2Vec, the proposed embedding learning trajectory captures the structural identity and improves the F-measure by 5.06% and 2.4%, respectively.
Usman Ahmed, Jerry Chun-Wei Lin, Gautam Srivastava 0001, Youcef Djenouri, Jimmy Ming-Tai Wu
IEEE BigData4
2021 LSTM for Periodic Broadcasting in Green IoT Applications over Energy Harvesting Enabled Wireless Networks: Case Study on ADAPCAST
abstract
The present paper considers emerging Internet of Things (IoT) applications and proposes a Long Short Term Memory (LSTM) based neural network for predicting the end of the broadcasting period under slotted CSMA (Carrier Sense Multiple Access) based MAC protocol and Energy Harvesting enabled Wireless Networks (EHWNs). The goal is to explore LSTM for minimizing the number of missed nodes and the number of broadcasting time-slots required to reach all the nodes under periodic broadcast operations. The proposed LSTM model predicts the end of the current broadcast period relying on the Root Mean Square Error (RMSE) values generated by its output, which (the RMSE) is used as an indicator for the divergence of the model. As a case study, we enhance our already developed broadcast policy, ADAPCAST by applying the proposed LSTM. This allows to dynamically adjust the end of the broadcast periods, instead of statically fixing it beforehand. An artificial data-set of the historical data is used to feed the proposed LSTM with information about the amounts of incoming, consumed, and effective energy per time-slot, and the radio activity besides the average number of missed nodes per frame. The obtained results prove the efficiency of the proposed LSTM model in terms of minimizing both the number of missed nodes and the number of time-slots required for completing broadcast operations.
Mustapha Khiati, Djamel Djenouri, Jianguo Ding, Youcef Djenouri
MSN4
2021 Privacy reinforcement learning for faults detection in the smart grid
abstract
Recent anticipated advancements in ad hoc Wireless Mesh Networks (WMN) have made them strong natural candidates for Smart Grid’s Neighborhood Area Network (NAN) and the ongoing work on Advanced Metering Infrastructure (AMI). Fault detection in these types of energy systems has recently shown lots of interest in the data science community, where anomalous behavior from energy platforms is identified. This paper develops a new framework based on privacy reinforcement learning to accurately identify anomalous patterns in a distributed and heterogeneous energy environment. The local outlier factor is first performed to derive the local simple anomalous patterns in each site of the distributed energy platform. A reinforcement privacy learning is then established using blockchain technology to merge the local anomalous patterns into global complex anomalous patterns. Besides, different optimization strategies are suggested to improve the whole outlier detection process. To demonstrate the applicability of the proposed framework, intensive experiments have been carried out on well-known CASAS (Center of Advanced Studies in Adaptive Systems) platform. Our results show that our proposed framework outperforms the baseline fault detection solutions.
Asma Belhadi, Youcef Djenouri, Gautam Srivastava 0001, Alireza Jolfaei, Jerry Chun-Wei Lin
Ad Hoc Networks2
2021 Cluster-based information retrieval using pattern mining
abstract
Abstract This paper addresses the problem of responding to user queries by fetching the most relevant object from a clustered set of objects. It addresses the common drawbacks of cluster-based approaches and targets fast, high-quality information retrieval. For this purpose, a novel cluster-based information retrieval approach is proposed, named Cluster-based Retrieval using Pattern Mining (CRPM). This approach integrates various clustering and pattern mining algorithms. First, it generates clusters of objects that contain similar objects. Three clustering algorithms based on k-means, DBSCAN (Density-based spatial clustering of applications with noise), and Spectral are suggested to minimize the number of shared terms among the clusters of objects. Second, frequent and high-utility pattern mining algorithms are performed on each cluster to extract the pattern bases. Third, the clusters of objects are ranked for every query. In this context, two ranking strategies are proposed: i) Score Pattern Computing (SPC), which calculates a score representing the similarity between a user query and a cluster; and ii) Weighted Terms in Clusters (WTC), which calculates a weight for every term and uses the relevant terms to compute the score between a user query and each cluster. Irrelevant information derived from the pattern bases is also used to deal with unexpected user queries. To evaluate the proposed approach, extensive experiments were carried out on two use cases: the documents and tweets corpus. The results showed that the designed approach outperformed traditional and cluster-based information retrieval approaches in terms of the quality of the returned objects while being very competitive in terms of runtime.
Youcef Djenouri, Asma Belhadi, Djamel Djenouri, Jerry Chun-Wei Lin
Appl. Intell.1
2021 Linguistic frequent pattern mining using a compressed structure
Jerry Chun-Wei Lin, Usman Ahmed, Gautam Srivastava 0001, Jimmy Ming-Tai Wu, Tzung-Pei Hong, Youcef Djenouri
Appl. Intell.6
2021 Reinforcement learning multi-agent system for faults diagnosis of mircoservices in industrial settings
Asma Belhadi, Youcef Djenouri, Gautam Srivastava 0001, Jerry Chun-Wei Lin
Comput. Commun.2
2021 Self-adjusting k nearest neighbors for continual learning from multi-label drifting data streams
Martha I. Roseberry, Bartosz Krawczyk, Youcef Djenouri, Alberto Cano 0001
Neurocomputing3
2021 Privacy-Preserving Multiobjective Sanitization Model in 6G IoT Environments
abstract
The next revolution of the smart industry relies on the emergence of the Industrial Internet of Things (IoT) and 5G/6G technology. The properties of such sophisticated communication technologies will change our perspective of information and communication by enabling seamless connectivity and bring closer entities, data, and “things.” Terahertz-based 6G networks promise the best speed and reliability, but they will face new man-in-the-middle attacks. In such critical and high-sensitive environments, the security of data and privacy of information still a big challenge. Without privacy-preserving considerations, the configuration state may be attacked or modified, thus causing security problems and damage to data. In this article, motivated by the need to secure 6G IoT networks, an ant colony optimization (ACO) approach is presented by adopting multiple objectives as well as using transaction deletion to secure confidential and sensitive information. Each ant in the population is represented as a set of possible deletion transactions for hiding sensitive information. We utilize the use of a prelarge concept to assist in the reduction of multiple database scans in the evaluation progress. We then also adopt external solutions to maintain discovered Pareto solutions, thus improving effectiveness to find optimized solutions. Experiments are conducted comparing our methodology to state-of-the-art bioinspired particle swarm optimization (PSO) as well as genetic algorithm (GA). Our strong results clearly show that the designed approach achieves fewer side effects while maintaining low computational cost overall (Chen et al., 2020).
Jerry Chun-Wei Lin, Gautam Srivastava 0001, Yuyu Zhang, Youcef Djenouri, Moayad Aloqaily
IEEE Internet Things J.4
2021 An improved opposition-based marine predators algorithm for global optimization and multilevel thresholding image segmentation
Essam H. Houssein, Kashif Hussain 0001, Laith Mohammad Abualigah, Mohamed E. Abd Elaziz, Waleed Alomoush, Gaurav Dhiman 0001, Youcef Djenouri, Erik Valdemar Cuevas Jiménez
Knowl. Based Syst.7
2021 ASRNN: A recurrent neural network with an attention model for sequence labeling
Jerry Chun-Wei Lin, Yinan Shao, Youcef Djenouri, Unil Yun
Knowl. Based Syst.3
2021 Mining of High-Utility Patterns in Big IoT-based Databases
Jimmy Ming-Tai Wu, Gautam Srivastava 0001, Jerry Chun-Wei Lin, Youcef Djenouri, Reza M. Parizi, Mohammad S. Khan
Mob. Networks Appl.4
2021 Fast and Accurate Convolution Neural Network for Detecting Manufacturing Data
abstract
This article introduces a technique known as clustering with particle for object detection (CPOD) for use in smart factories. CPOD builds on regional-based methods to identify smart object data using outlier detection, clustering, particle swarm optimization (PSO), and deep convolutional networks. The process starts by removing noise and errors from the images database by the local outlier factor (LOF) algorithm. Next, the algorithm studies different correlations from the set of images in the database. This creates homogeneous, and similar clusters using the well-known k-means algorithm, and the FastRCNN (fast region convolutional neural network) uses these clusters to design efficient and more focused models. PSO is used to optimize the different parameters including, the number of neighbors of LOF, the number of clusters of k-means, the number of epochs, and the error learning rate for FastRCNN. The inference process benefits from the knowledge provided by training. Instead of considering a complex single model of the whole images database, we consider a simple homogeneous model. To demonstrate the usefulness of our approach, intensive experiments have been carried out on standard images database, and real smart manufacturer data. Our results show that CPOD when compared to baseline object detection solutions is superior in terms of runtime and accuracy.
Youcef Djenouri, Gautam Srivastava 0001, Jerry Chun-Wei Lin
IEEE Trans. Ind. Informatics1
2021 A Two-Phase Anomaly Detection Model for Secure Intelligent Transportation Ride-Hailing Trajectories
abstract
This paper addresses the taxi fraud problem and introduces a new solution to identify trajectory outliers. The approach as presented allows to identify both individual and group outliers and is based on a two phase-based algorithm. The first phase determines the individual trajectory outliers by computing the distance of each point in each trajectory, whereas the second identifies the group trajectory outliers by exploring the individual trajectory outliers using both feature selection and sliding windows strategies. A parallel version of the algorithm is also proposed using a sliding window-based GPU approach to boost the runtime performance. Extensive experiments have been carried out to thoroughly demonstrate the usefulness of our methodology on both synthetic and real trajectory databases. The results show that the GPU approach enables reaching a speed-up of 341 over the sequential algorithm on large synthetic databases. The efficiency of the proposed method to detect both individual and group trajectory outliers on a real-world taxi trajectory database is also demonstrated in comparison with baseline trajectory outlier and group detection algorithms. The results are very promising and show superiority of the proposed method both in reducing computational time and enhancing the quality of returned outliers. Finally, we prime our methodology and results for future refinement using deep learning methodologies.
Asma Belhadi, Youcef Djenouri, Gautam Srivastava 0001, Djamel Djenouri, Alberto Cano 0001, Jerry Chun-Wei Lin
IEEE Trans. Intell. Transp. Syst.2
2021 Uncertain-Driven Analytics of Sequence Data in IoCV Environments
abstract
As the increasing availability and use of dynamic mobile communications, information from an Internet of Things (IoT) subset of devices, known as Internet of Connected Vehicles (IoCV), is collected with a level of uncertainty. To bridge this gap of data analytics, some studies take two factors individually to mine knowledge or information, such as uncertainty and utility as two exemplary factors. However, this approach may cause actual loss of knowledge integrity. In this work, our first result is a knowledge called High Expected Utility Sequential Patterns (HEUSPs) that is both novel and also provides an alternative option for knowledge discovery regarding utility and uncertainty factors by a single threshold in IoCV environments. Furthermore, two PUL-Chain and EUL-Chain structures with six pruning methodologies are respectively developed to maintain information that is necessary and reduce the search space for improving mining performance. Our experimental results show both efficiency and strength of the designed algorithm compared to HUS-Span which is considered to be the current standard in utility-oriented sequential pattern mining.
Gautam Srivastava 0001, Jerry Chun-Wei Lin, Alireza Jolfaei, Yuanfa Li, Youcef Djenouri
IEEE Trans. Intell. Transp. Syst.5
2021 SS-ITS: secure scalable intelligent transportation systems
Asma Belhadi, Youcef Djenouri, Gautam Srivastava 0001, Jerry Chun-Wei Lin
J. Supercomput.2
2021 Trajectory Outlier Detection: New Problems and Solutions for Smart Cities
abstract
This article introduces two new problems related to trajectory outlier detection: (1) group trajectory outlier (GTO) detection and (2) deviation point detection for both individual and group of trajectory outliers. Five algorithms are proposed for the first problem by adapting DBSCAN , k nearest neighbors (kNN) , and feature selection (FS) . DBSCAN-GTO first applies DBSCAN to derive the micro clusters , which are considered as potential candidates. A pruning strategy based on density computation measure is then suggested to find the group of trajectory outliers. kNN-GTO recursively derives the trajectory candidates from the individual trajectory outliers and prunes them based on their density. The overall process is repeated for all individual trajectory outliers. FS-GTO considers the set of individual trajectory outliers as the set of all features, while the FS process is used to retrieve the group of trajectory outliers. The proposed algorithms are improved by incorporating ensemble learning and high-performance computing during the detection process. Moreover, we propose a general two-phase-based algorithm for detecting the deviation points, as well as a version for graphic processing units implementation using sliding windows. Experiments on a real trajectory dataset have been carried out to demonstrate the performance of the proposed approaches. The results show that they can efficiently identify useful patterns represented by group of trajectory outliers, deviation points, and that they outperform the baseline group detection algorithms.
Youcef Djenouri, Djamel Djenouri, Jerry Chun-Wei Lin
ACM Trans. Knowl. Discov. Data1
2020 Mining Multiple Fuzzy Frequent Patterns with Compressed List Structures
abstract
Fuzzy-set theory was invented to represent more meaningful representations of knowledge for human reasoning, which can also be applied and utilized for handling the quantitative database. In this paper, an efficient fuzzy mining (EFM) algorithm is presented to fast discover the multiple fuzzy frequent patterns from quantitative databases under type-2 fuzzy-set theory. A compressed fuzzy-list (CFL)-structure is developed to maintain complete information for rule generation. Two pruning techniques are developed to reduce the search space and speed up mining progress. Several experiments are carried out for the purpose of verifying the efficiency and effectiveness of the designed approach in terms of runtime and the number of examined nodes under different minimum support thresholds and the results indicated the designed EFM achieves the best performance compared to the existing models.
Jerry Chun-Wei Lin, Jimmy Ming-Tai Wu, Youcef Djenouri, Gautam Srivastava 0001, Tzung-Pei Hong
FUZZ-IEEE3
2020 Hybrid Decomposition Convolution Neural Network and Vocabulary Forest for Image Retrieval
abstract
This paper introduces a highly efficient image retrieval technique called DCNN-vForest (Decomposition Convolution Neural Network and vocabulary Forest), which aims to retrieve the relevant images to the given image query by studying the correlation between images in the image database based on decomposition. The regional and global features of the image database are first extracted using the convolution neural network, and then divided into clusters of similar images using the Kmeans algorithm. We propose a new structure called vForest (vocabulary Forest), by calculating the vocabulary tree on each cluster of images. The retrieval process benefits from the knowledge provided by the vForest, and instead of considering the whole image database, only the most similar cluster to the image query is explored. To demonstrate the usefulness of our approach, intensive experiments have been carried out on ground-truth image databases, the results reveal the superiority of DCNN- vForest against the baseline image retrieval solutions, in terms of runtime and accuracy.
Youcef Djenouri, Jon Mikkelsen Hjelmervik
ICPR1
2020 Efficient Mining of Pareto-Front High Expected Utility Patterns
Usman Ahmed, Jerry Chun-Wei Lin, Jimmy Ming-Tai Wu, Youcef Djenouri, Gautam Srivastava 0001, Suresh Kumar Mukhiya
IEA/AIE4
2020 Data mining-based approach for ontology matching problem
Hiba Belhadi, Karima Akli-Astouati, Youcef Djenouri, Jerry Chun-Wei Lin
Appl. Intell.3
2020 A recurrent neural network for urban long-term traffic flow forecasting
abstract
Abstract This paper investigates the use of recurrent neural network to predict urban long-term traffic flows. A representation of the long-term flows with related weather and contextual information is first introduced. A recurrent neural network approach, named RNN-LF, is then proposed to predict the long-term of flows from multiple data sources. Moreover, a parallel implementation on GPU of the proposed solution is developed (GRNN-LF), which allows to boost the performance of RNN-LF. Several experiments have been carried out on real traffic flow including a small city (Odense, Denmark) and a very big city (Beijing). The results reveal that the sequential version (RNN-LF) is capable of dealing effectively with traffic of small cities. They also confirm the scalability of GRNN-LF compared to the most competitive GPU-based software tools when dealing with big traffic flow such as Beijing urban data.
Asma Belhadi, Youcef Djenouri, Djamel Djenouri, Jerry Chun-Wei Lin
Appl. Intell.2
2020 A general-purpose distributed pattern mining system
abstract
Abstract This paper explores five pattern mining problems and proposes a new distributed framework called DT-DPM: Decomposition Transaction for Distributed Pattern Mining. DT-DPM addresses the limitations of the existing pattern mining problems by reducing the enumeration search space. Thus, it derives the relevant patterns by studying the different correlation among the transactions. It first decomposes the set of transactions into several clusters of different sizes, and then explores heterogeneous architectures, including MapReduce, single CPU, and multi CPU, based on the densities of each subset of transactions. To evaluate the DT-DPM framework, extensive experiments were carried out by solving five pattern mining problems (FIM: Frequent Itemset Mining, WIM: Weighted Itemset Mining, UIM: Uncertain Itemset Mining, HUIM: High Utility Itemset Mining, and SPM: Sequential Pattern Mining). Experimental results reveal that by using DT-DPM, the scalability of the pattern mining algorithms was improved on large databases. Results also reveal that DT-DPM outperforms the baseline parallel pattern mining algorithms on big databases.
Asma Belhadi, Youcef Djenouri, Jerry Chun-Wei Lin, Alberto Cano 0001
Appl. Intell.2
2020 Incrementally updating the high average-utility patterns with pre-large concept
abstract
Abstract High-utility itemset mining (HUIM) is considered as an emerging approach to detect the high-utility patterns from databases. Most existing algorithms of HUIM only consider the itemset utility regardless of the length. This limitation raises the utility as a result of a growing itemset size. High average-utility itemset mining (HAUIM) considers the size of the itemset, thus providing a more balanced scale to measure the average-utility for decision-making. Several algorithms were presented to efficiently mine the set of high average-utility itemsets (HAUIs) but most of them focus on handling static databases. In the past, a fast-updated (FUP)-based algorithm was developed to efficiently handle the incremental problem but it still has to re-scan the database when the itemset in the original database is small but there is a high average-utility upper-bound itemset (HAUUBI) in the newly inserted transactions. In this paper, an efficient framework called PRE-HAUIMI for transaction insertion in dynamic databases is developed, which relies on the average-utility-list (AUL) structures. Moreover, we apply the pre-large concept on HAUIM. A pre-large concept is used to speed up the mining performance, which can ensure that if the total utility in the newly inserted transaction is within the safety bound, the small itemsets in the original database could not be the large ones after the database is updated. This, in turn, reduces the recurring database scans and obtains the correct HAUIs. Experiments demonstrate that the PRE-HAUIMI outperforms the state-of-the-art batch mode HAUI-Miner, and the state-of-the-art incremental IHAUPM and FUP-based algorithms in terms of runtime, memory, number of assessed patterns and scalability.
Jerry Chun-Wei Lin, Matin Pirouz, Youcef Djenouri, Chien-Fu Cheng, Usman Ahmed
Appl. Intell.3
2020 Space-time series clustering: Algorithms, taxonomy, and case study on urban smart cities
abstract
This paper provides a short overview of space–time series clustering, which can be generally grouped into three main categories such as: hierarchical, partitioning-based, and overlapping clustering. The first hierarchical category is to identify hierarchies in space–time series data. The second partitioning-based category focuses on determining disjoint partitions among the space–time series data, whereas the third overlapping category explores fuzzy logic to determine the different correlations between the space–time series clusters. We also further describe solutions for each category in this paper. Furthermore, we show the applications of these solutions in an urban traffic data captured on two urban smart cities (e.g., Odense in Denmark and Beijing in China). The perspectives on open questions and research challenges are also mentioned and discussed that allow to obtain a better understanding of the intuition, limitations, and benefits for the various space–time series clustering methods. This work can thus provide the guidances to practitioners for selecting the most suitable methods for their used cases, domains, and applications.
Asma Belhadi, Youcef Djenouri, Kjetil Nørvåg, Heri Ramampiaro, Florent Masseglia, Jerry Chun-Wei Lin
Eng. Appl. Artif. Intell.2
2019 Mining High-Utility Sequential Patterns from Big Datasets
abstract
High-Utility Sequential Pattern Mining (HUSPM) has become an emerging issue in recent decades since it reveals more information such as the utility and sequence factors for knowledge discovery. For the previous works, many algorithms were presented to speed up the mining performance regarding a single machine with small datasets. In real-world applications, the size of dataset can be collected from many places or devices, such as PC, Internet of Things (IoT), mobile devices, and shopping malls, among others. It is necessary to build an efficient model to handle the big dataset for HUSPM. In this paper, we present a four-stages MapReduce framework based on the Spark platform for mining the high-utility sequential patterns from a very large database. From the experimental results, we then can observe that the designed model outperforms the state-of-the-art approaches for handling the very big dataset.
Jerry Chun-Wei Lin, Yuanfa Li, Philippe Fournier-Viger, Youcef Djenouri, Shyue-Liang Wang
IEEE BigData4
2019 A Novel Parallel Framework for Metaheuristic-based Frequent Itemset Mining
abstract
Frequent Itemset Mining (FIM) is an important but very time-consuming data mining task. As a result, traditional FIM algorithms are often not scalable to large databases. To address this issue, several metaheuristics have been developed in recent years to find good approximate solutions to the FIM problem. It was shown that such approaches can be much more efficient than exact algorithms. However, metaheuristics often have long runtimes on massive datasets and the quality of their solutions can be improved. To address this issue, this paper proposes a parallel framework called CFIM (Cluster for Frequent Itemset Mining) for metaheuristic-based FIM. It accelerates FIM by using multiple cluster workers. The proposed approach partitions a transactional database and the set of all items at the level of cluster workers. The itemset generation process is performed by each worker, which then send results to a master node. This latter performs a merging step to only keep high quality itemsets by considering their frequency and diversification. Three metaheuristics (GA, PSO and BSO) are integrated in this framework to yield three novel metaheuristics (CGA, CPSO and CBSO). Extensive experiments show that CPSO outperforms CGA, CBSO, and state-of-the-art high performance computing FIM approaches.
Youcef Djenouri, Djamel Djenouri, Asma Belhadi, Jerry Chun-Wei Lin, Ahcène Bendjoudi, Philippe Fournier-Viger
CEC1
2019 Single Scan Polynomial Algorithms for Frequent Itemset Mining in Big Databases
abstract
This paper considers frequent itemset mining in big transactional databases. It first introduces a novel approach (Bio-SS) that combines the bio-inspired algorithms with the single scan algorithm (SSFIM). The proposed approach addresses the limitations of SSFIM by utilizing the bio-inspired operators in the generation process. This reduces the time complexity of SSFIM from exponential to polynomial, while taking advantage of the capacity to derive the frequent itemsets by performing a single database scan, independently from the minimum support value. This allows to considerably accelerate the scan procedure compared to existing approaches, especially when dealing with large scale databases. The numerical results show that the designed Bio-SS outperforms both accurate and metaheuristics baseline FIM approaches when dealing with big databases.
Youcef Djenouri, Djamel Djenouri, Jerry Chun-Wei Lin, Asma Belhadi
CEC1
2019 A Swarm-based Data Sanitization Algorithm in Privacy-Preserving Data Mining
abstract
In recent decades, data protection (PPDM), which not only hides information, but also provides information that is useful to make decisions, has become a critical concern. We present a sanitization algorithm with the consideration of four side effects based on multi-objective PSO and hierarchical clustering methods to find optimized solutions for PPDM. Experiments showed that compared to existing approaches, the designed sanitization algorithm based on the hierarchical clustering method achieves satisfactory performance in terms of hiding failure, missing cost, and artificial cost.
Jimmy Ming-Tai Wu, Jerry Chun-Wei Lin, Youcef Djenouri, Philippe Fournier-Viger, Yuyu Zhang
CEC3
2019 Highly Efficient Pattern Mining Based on Transaction Decomposition
abstract
This paper introduces a highly efficient pattern mining technique called Clustering-Based Pattern Mining (CBPM). This technique discovers relevant patterns by studying the correlation between transactions in transaction databases using clustering techniques. The set of transactions are first clus-tered using the k-means algorithm, where highly correlated transactions are grouped together. Next, the relevant patterns are derived by applying a pattern mining algorithm to each cluster. We present two different pattern mining algorithms, one approximate and one exact. We demonstrate the efficiency and effectiveness of CBPM through a thorough experimental evaluation.
Youcef Djenouri, Jerry Chun-Wei Lin, Kjetil Nørvåg, Heri Ramampiaro
ICDE1
2019 GPU-based swarm intelligence for Association Rule Mining in big databases
abstract
Association Rule Mining (ARM) is a fundamental data mining task that is time-consuming on big datasets. Thus, developing new scalable algorithms for this problem is desirable. Recently, Bee Swarm Optimization (BSO)-based meta-heuristics were shown effective to reduce the time required for ARM. But these approaches were applied only on small or medium scale databases. To perform ARM on big databases, a promising approach is to design parallel algorithms using the massively parallel threads of a GPU processor. While some GPU-based ARM algorithms have been developed, they only benefit from GPU parallelism during the evaluation step of solutions obtained by the BSO-metaheuristics. This paper improves this approach by parallelizing the other steps of the BSO process (diversification and intensification). Based on these novel ideas, three novel algorithms are presented, i) DRGPU (Determination of Regions on GPU), ii) SAGPU (Search Area on GPU, and, iii) ALLGPU (All steps on GPU). These solutions are analyzed and empirically compared on benchmark datasets. Experimental results show that ALLGPU outperforms the three other approaches in terms of speed up. Moreover, results confirm that ALLGPU outperforms the state-of-the-art GPU-based ARM approaches on big ARM databases such as the Webdocs dataset. Furthermore, ALLGPU is extended to mine big frequent graphs and results demonstrate its superiority over the state-of-the-art D-Mine algorithm for frequent graph mining on the large Pokec social network dataset.
Youcef Djenouri, Philippe Fournier-Viger, Jerry Chun-Wei Lin, Djamel Djenouri, Asma Belhadi
Intell. Data Anal.1
2019 Exploiting GPU and cluster parallelism in single scan frequent itemset mining
Youcef Djenouri, Djamel Djenouri, Asma Belhadi, Alberto Cano 0001
Inf. Sci.1
2019 Exploiting GPU parallelism in improving bees swarm optimization for mining big transactional databases
Youcef Djenouri, Djamel Djenouri, Asma Belhadi, Philippe Fournier-Viger, Jerry Chun-Wei Lin, Ahcène Bendjoudi
Inf. Sci.1
2019 Bee swarm optimization for solving the MAXSAT problem using prior knowledge
Youcef Djenouri, Zineb Habbas, Djamel Djenouri, Philippe Fournier-Viger
Soft Comput.1
2019 Hiding sensitive itemsets with multiple objective optimization
Jerry Chun-Wei Lin, Yuyu Zhang, Philippe Fournier-Viger, Youcef Djenouri
Soft Comput.5
2018 Anonymization of Multiple and Personalized Sensitive Attributes
Jerry Chun-Wei Lin, Qiankun Liu 0002, Philippe Fournier-Viger, Youcef Djenouri, Ji Zhang 0001
DaWaK4
2018 A Metaheuristic Algorithm for Hiding Sensitive Itemsets
Jerry Chun-Wei Lin, Yuyu Zhang, Philippe Fournier-Viger, Youcef Djenouri, Ji Zhang 0001
DEXA (2)4
2018 Outlier Detection in Urban Traffic Flow Distributions
abstract
Urban traffic data consists of observations like number and speed of cars or other vehicles at certain locations as measured by deployed sensors. These numbers can be interpreted as traffic flow which in turn relates to the capacity of streets and the demand of the traffic system. City planners are interested in studying the impact of various conditions on the traffic flow, leading to unusual patterns, i.e., outliers. Existing approaches to outlier detection in urban traffic data take into account only individual flow values (i.e., an individual observation). This can be interesting for real time detection of sudden changes. Here, we face a different scenario: The city planners want to learn from historical data, how special circumstances (e.g., events or festivals) relate to unusual patterns in the traffic flow, in order to support improved planing of both, events and the layout of the traffic system. Therefore, we propose to consider the sequence of traffic flow values observed within some time interval. Such flow sequences can be modeled as probability distributions of flows. We adapt an established outlier detection method, the local outlier factor (LOF), to handling flow distributions rather than individual observations. We apply the outlier detection online to extend the database with new flow distributions that are considered inliers. For the validation we consider a special case of our framework for comparison with state-of-the-art outlier detection on flows. In addition, a real case study on urban traffic flow data showcases that our method finds meaningful outliers in the traffic flow data.
Youcef Djenouri, Arthur Zimek, Marco Chiarandini
ICDM1
2018 A new framework for metaheuristic-based frequent itemset mining
Youcef Djenouri, Djamel Djenouri, Asma Belhadi, Philippe Fournier-Viger, Jerry Chun-Wei Lin
Appl. Intell.1
2018 Maintenance algorithm for high average-utility itemsets with transaction deletion
Jerry Chun-Wei Lin, Yinan Shao, Philippe Fournier-Viger, Youcef Djenouri, Xiangmin Guo
Appl. Intell.4
2018 How to exploit high performance computing in population-based metaheuristics for solving association rule mining problem
Youcef Djenouri, Djamel Djenouri, Zineb Habbas, Asma Belhadi
Distributed Parallel Databases1
2018 Bees swarm optimization guided by data mining techniques for document information retrieval
Youcef Djenouri, Asma Belhadi, Riadh Belkebir
Expert Syst. Appl.1
2018 Mining diversified association rules in big datasets: A cluster/GPU/genetic approach
Youcef Djenouri, Asma Belhadi, Philippe Fournier-Viger, Hamido Fujita
Inf. Sci.1
2018 Fast and effective cluster-based information retrieval using frequent closed itemsets
Youcef Djenouri, Asma Belhadi, Philippe Fournier-Viger, Jerry Chun-Wei Lin
Inf. Sci.1
2018 Extracting useful knowledge from event logs: A frequent itemset mining approach
Youcef Djenouri, Asma Belhadi, Philippe Fournier-Viger
Knowl. Based Syst.1
2017 SS-FIM: Single Scan for Frequent Itemsets Mining in Transactional Databases
Youcef Djenouri, Marco Comuzzi, Djamel Djenouri
PAKDD (2)1
2017 GPU-based Bio-inspired Model for Solving Association Rules Mining Problem
abstract
We explore in this paper the application of bioinspired approaches to the association rules mining (ARM) problem for the purpose of accelerating the process of extracting the correlations between items in sizeable data instances. A new bio-inspired GPU-based model is proposed, which benefits from the massively GPU threading by evaluating multiple rules in parallel on GPU. To validate the proposed model, the most used bio-inspired approaches (GA, PSO, and BSO) have been executed on GPU to solve wellknown large ARM instances. Real experiments have been carried out on an Intel Xeon 64 bit quad-core processor E5520 coupled to an Nvidia Tesla C2075 GPU device. The results show that the genetic algorithm outperforms PSO and BSO. Moreover, it outperforms the state-of-the-art GPU-based ARM approaches when dealing with the challenging Webdocs instance.
Youcef Djenouri, Ahcène Bendjoudi, Djamel Djenouri, Marco Comuzzi
PDP1
2017 Reducing thread divergence in GPU-based bees swarm optimization applied to association rule mining
abstract
Summary The association rules mining (ARM) problem is one of the most important problems in the area of data mining. It aims at finding all relevant association rules from transactional databases. It is CPU time intensive and requires a huge computing power when dealing with large transactional databases. To deal with this issue, Graphics Processing Units (GPUs) are a powerful tool to speed up the search process. However, their performance is closely subject to thread/branch divergence resulting from the single instruction multiple data parallel model of GPUs. In this paper, we propose three approaches based on database reorganization, aiming to reduce thread divergence in GPU‐based bees swarm optimization metaheuristic for ARM, respectively, named block‐based reordering, transactions‐based reordering, and transactions‐based reordering with median value. Theoretical and experimental studies have been carried out using well‐known large ARM instances. The experiments have been performed on an Intel Xeon 64 bit quad‐core processor E5520 coupled to Nvidia Tesla C2075 448 cores. The results show that the proposed approaches minimize considerably the number of thread divergence and improve the overall execution time. Indeed, the number of thread divergence occurrences has been reduced by up to eight times making the execution much faster. Copyright © 2016 John Wiley & Sons, Ltd.
Youcef Djenouri, Ahcène Bendjoudi, Zineb Habbas, Malika Mehdi, Djamel Djenouri
Concurr. Comput. Pract. Exp.1
2017 Combining Apriori heuristic and bio-inspired algorithms for solving the frequent itemsets mining problem
Youcef Djenouri, Marco Comuzzi
Inf. Sci.1
2016 Bees Swarm Optimization Metaheuristic Guided by Decomposition for Solving MAX-SAT
abstract
Decomposition methods aim to split a problem into a collection a collection of smaller interconnected sub-problems. Several research works have explored decomposition methods for solving large optimization problems. Due to its theroretical properties, Tree decomposition has been especially the subject of numerous successfull studies in the context of exact optimization solvers. More recently, Tree decomposition has been successfully used to guide the Variable Neighbor Search (VNS) local search method. Our present contribution follows this last direction and proposes two approaches called BSOGD1 and BSOGD2 for guiding the Bees Swarm Optimization (BSO) metaheuristic by using a decomposition method. More pragmatically, this paper deals with the MAX-SAT problem and uses the Kmeans algorithm as a decomposition method. Several experimental results conducted on DIMACS benchmarks and some other hard SAT instances lead to promising results in terms of the quality of the solutions. Moreover, these experiments highlight a good stability of the two approaches, more especially, when dealing with hard instances like the Parity8 family from DIMACS. Beyond these first promising results, note that this approach can be easily applied to many other optimization problems such as the Weighted MAX-SAT, the MAX-CSP or the coloring problem and can be used with other decomposition methods as well as other metaheuristics.
Youcef Djenouri, Zineb Habbas, Wassila Aggoune-Mtalaa
ICAART (2)1
2015 GPU-based bees swarm optimization for association rules mining
Youcef Djenouri, Ahcène Bendjoudi, Malika Mehdi, Nadia Nouali-Taboudjemat, Zineb Habbas
J. Supercomput.1
2013 Multilevel Clustering of Induction Rules for Web Meta-knowledge
Amine Chemchem, Habiba Drias, Youcef Djenouri
WorldCIST3