EDBT 2026 Demo / reviewers in the wild / expert
Mounim A. El-Yacoubi
dblp:54/3370 · also Abdenaim El Yacoubi, Mounîm A. El-Yacoubi
· DBLP profile ↗
67ranked-venue papers
6as first author
30since 2021 · last 2026
0000-0002-7383-0588ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 9 · 9 since 2021Security and privacy · 8 · 7 since 2021Computer networks · 5 · 2 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Facial digital markers For hypomimia detection in Parkinson's disease: A systematic review
Anas Filali Razzouki, Laetitia Jeancolas, Dijana Petrovska-Delacrétaz, Mounim A. El-Yacoubi |
Pattern Recognit. | 4 |
| 2026 | Before memories fade: Large language models for language-based dementia detection, a systematic reviewabstractLarge language models (LLMs) have emerged as tools for automatic dementia detection from language. This systematic review synthesizes 87 studies published between 1 January 2020 and 15 October 2025 that apply LLMs to spoken or written patient language for detecting Alzheimer’s disease, mild cognitive impairment, and related dementia. We group existing methods into three main families: finetuning, frozen LLMs used as embedding extractors, and prompting-based approaches, and compare them across datasets, languages, and task formulations. The analysis reveals a clear temporal shift from early dominance of finetuning and embedding-based pipelines to a recent surge in prompting methods, alongside heavy reliance on a small set of English speech corpora and limited work on handwriting as a linguistic modality. We also identify recurring methodological limitations, including scarce cross-dataset validation and limited attention to explainability. Taken together, the review outlines current capabilities and gaps, and proposes concrete directions for developing more robust, multimodal, and clinically interpretable LLM-based systems for dementia detection. Jana Sweidan, Mounim A. El-Yacoubi, Nasredine Semmar |
Pattern Recognit. | 2 |
| 2026 | MsMemoryGAN: A Multiscale Memory GAN for Palm-Vein Adversarial PurificationabstractDeep neural networks have recently achieved promising performance in the vein recognition task and have shown an increasing application trend. However, they are prone to adversarial attacks by adding imperceptible perturbations to the input, resulting in incorrect recognition. To address this issue, we propose a novel defense model named MsMemoryGAN, which aims to filter the perturbations from adversarial samples before recognition. First, we design a multiscale memory autoencoder (MsMemoryAE) to achieve high-quality reconstruction, where the memory module (MM) within it is capable of learning the detailed patterns of normal samples at different scales. Second, to overcome the limitations of handcrafted similarity metrics, we propose an MM with learnable similarity (LSMM), which retrieves the most relevant memory items to purify the input feature. Finally, the perceptual loss and adversarial loss are integrated with the pixel loss to further enhance the quality of the reconstructed image. During the training phase, the MsMemoryGAN learns to reconstruct the input by merely using fewer prototypical elements of the normal patterns recorded in the memory. At the testing stage, given an adversarial sample, the MsMemoryGAN retrieves its most relevant normal patterns in MMs for reconstruction. Perturbations in the adversarial sample are usually not reconstructed well, resulting in adversarial purification. We conduct extensive experiments on two public vein datasets under different adversarial attack methods to evaluate the performance of the proposed approach. The experimental results show that our approach removes a wide variety of adversarial perturbations, allowing vein classifiers to achieve the highest recognition accuracy. Huafeng Qin, Yuming Fu 0001, Huiyan Zhang 0001, Mounim A. El-Yacoubi, Xinbo Gao 0001, Qun Song 0007, Jun Wang 0071 |
IEEE Trans. Cybern. | 4 |
| 2026 | Neural Architecture Search-Based Global-Local Vision Mamba for Palm-Vein RecognitionabstractOwing to its inherent attributes of high security, privacy preservation, and liveness detection, vein recognition has garnered significant attention, with deep learning (DL) models prevailing in the field. In particular, Mamba, a recent DL architecture showing robust feature representation with linear computational complexity, has been applied successfully for visual tasks. However, Vision Mamba captures long-distance feature dependencies but deteriorates local feature details. Besides, manually designing Mamba architecture based on human prior knowledge is very time-consuming and error-prone. To address these limitations, we propose a hybrid network structure named Global-local Vision Mamba (GLVM) to learn both local correlations and global dependencies within images for comprehensive vein feature representation. Second, we design a Multi-head Mamba to learn the dependencies along different directions, so as to improve the feature representation of Vision Mamba. Third, to learn complementary features, we propose a ConvMamba block consisting of three branches: Multi-head Mamba branch (MHMamba), Feature Iteration Unit branch (FIU), and Convolutional Neural Network (CNN) branch, with FIU aiming to fuse convolutional local features with Mamba global representations. Finally, we propose a Global-local Alternate Neural Architecture Search (GLNAS) method, which alternately searches for the optimal architecture of GLVM through weight entanglement strategy and evolutionary algorithm. We have carried out rigorous experiments on five public vein datasets to assess performance. Our approach achieves the highest 96.84%, 99.63%, 95.73%, 99.72%, 99.14% accuracies and the lowest 0.27%, 0.07%, 0.48%, 0.07%, 0.12% EER among all existing approaches on five public vein datasets, which demonstrates that our approach is capable of learning more complete features than existing approaches. In addition, the visual assessment experiments also show that our approach extracts more global vein architecture and local vein detail for recognition. Huafeng Qin, Yuming Fu 0001, Jing Chen 0050, Mounim A. El-Yacoubi, Xinbo Gao 0001, Feng Xi |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Exploring Emotion Expression Recognition in Older Adults Interacting With a Virtual CoachabstractThe EMPATHIC project aimed to design an emotionally expressive virtual coach capable of engaging healthy seniors to improve well-being and promote independent aging. In particular, the system's human sensing capabilities allow for the perception of emotional states to provide a personalized experience. This paper outlines the development of the emotion expression recognition module of the virtual coach, encompassing data collection, annotation design, and a first methodological approach, all tailored to the project requirements. With the latter, we investigate the role of various modalities, individually and combined, for discrete emotion expression recognition in this context: speech from audio, and facial expressions, gaze, and head dynamics from video. The collected corpus includes users from Spain, France, and Norway, and was annotated separately for the audio and video channels with distinct emotional labels, allowing for a performance comparison across cultures and label types. Results confirm the informative power of the modalities studied for the emotional categories considered, with multimodal methods generally outperforming others (around 68% accuracy with audio labels and 72-74% with video labels). The findings are expected to contribute to the limited literature on emotion recognition applied to older adults in conversational human-machine interaction, and guide the development of future systems. Cristina Palmero, Mikel de Velasco-Vázquez, Mohamed Amine Hmani, Aymen Mtibaa, Leila Ben Letaifa, Pau Buch-Cardona, Raquel Justo, Terry Amorese, Eduardo Gonzalez-Fraile, Begoña Fernández-Ruanova, Jofre Tenorio-Laranga, Anna Torp Johansen, Micaela Rodrigues da Silva, L. J. Martinussen, Maria Stylianou Korsnes, Gennaro Cordasco, Anna Esposito, Mounim A. El-Yacoubi, Dijana Petrovska-Delacrétaz, M. Inés Torres, Sergio Escalera |
IEEE Trans. Affect. Comput. | 18 |
| 2025 | AdVeinSAM: Adversarial Learning-Based Large Model for Palm-Vein Feature SegmentationabstractPalm-vein recognition is gaining significant attention as a high-security biometric recognition technology. However, the vein image acquisition process is easily affected by several factors, making vein texture segmentation a challenging task. Recently, foundation models such as Segment Anything Model (SAM) have shown remarkable potential in image segmentation without requiring prior retraining. Nevertheless, due to the large domain discrepancy between the resource and target domains, as well as limited datasets, existing solutions that rely heavily on abundant training images often struggle to extract robust vein texture patterns. To address this challenge, we propose AdVeinSAM, an adversarial learning-based large model for palm-vein texture extraction, which leverages rich knowledge of large models to enhance vein pattern segmentation. Specifically, by alternately optimizing the vein segmentation model and the image generator, AdVeinSAM generates diverse training samples, effectively transferring knowledge from the large model to enhance feature extraction robustness. First, we incorporate the wavelet transform into xLSTM-UNet to design Wavelet-xLSTM-UNet, which generates diverse and realistic vein images for data augmentation. Then, we improve the NOLA model to fine-tune the segmentation anything model (SAM) and develop a specialized vein segmentation model (VeinSAM), which effectively extracts palm-vein texture features. Finally, the image generator (Wavelet-xLSTM-UNet) and the vein segmentation model (VeinSAM) are combined to form AdVeinSAM, where the generator and the VeinSAM are alternatively updated through adversarial training. Concretely, the image generator generates challenging samples to increase the segmentation difficulty for VeinSAM, while VeinSAM learns more robust feature representations from these challenging samples to improve the generalization and segmentation accuracy. We conduct extensive experiments on three public palm-vein databases and experimental results demonstrate that the proposed AdVeinSAM model outperforms state-of-the-art solutions, achieving the lowest equal error rates (EERs) of 1.48%, 4.76%, and 0.72%, respectively. These results confirm the effectiveness and robustness of AdVeinSAM in palm-vein texture extraction. Huafeng Qin, Hulei Deng, Yantao Li 0001, Mounim A. El-Yacoubi |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | WTxGRN: Wavelet Transform-Based Extended Gated Recurrent Network for Palm Vein RecognitionabstractVein recognition technology offers high security and privacy as an advanced biometric identification method. While deep learning techniques have achieved state-of-the-art performance in vein recognition due to their powerful pattern recognition capabilities, the Gated Recurrent Unit (GRU), a simplified version of LSTM, still faces limitations: 1) inability to process sequence information in parallel, leading to inefficient training; 2) loss of sensitivity to local features crucial for pattern recognition, despite excelling at modeling long-distance dependencies. To address these issues, we propose WTxGRN, a Wavelet Transform-based extended Gated Recurrent Network, which simultaneously extracts global and local features and supports parallel sequence processing. Specifically, we modify the GRU memory structure to enable parallel training and enhance feature representation through exponential gating and stabilization techniques, resulting in an extended GRU architecture called xGRU. We integrate xGRU into a wavelet transform-based residual backbone to form the xGRU Block. By incorporating a wavelet convolution branch and two Mixer Modules, we facilitate multi-scale feature extraction and fusion, enhancing vein recognition robustness and yielding the WTxGRU Block. Stacking these blocks constructs the WTxGRN. Furthermore, we present Spiking WTxGRN, an energy-efficient spiking version of WTxGRN, pioneering the application of spiking neural networks in vein recognition. Spiking WTxGRN offers high energy efficiency while maintaining excellent recognition performance, making it suitable for real-time vein recognition tasks. Extensive experiments on three public palm vein datasets demonstrate that our methods outperform state-of-the-art models across multiple benchmarks, achieving superior performance. Huafeng Qin, Yuming Fu 0001, Jing Chen 0050, Qun Song 0007, Yantao Li 0001, Mounim A. El-Yacoubi, Dexing Zhong |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Detection of Early Parkinson's Disease by Leveraging Speech Foundation ModelsabstractParkinson's disease (PD) is a progressive neurodegenerative disorder affecting millions worldwide, characterized by a wide range of motor and non-motor symptoms. Among these symptoms, alterations in speech and voice quality stand out as early and prominent indicators of the disease. Recently, the emergence of speech foundation models has revolutionized the field by providing powerful tools for speech processing and feature extraction. In this article, we investigate the capabilities of three state-of the art speech foundation models, wav2vec2.0, Whisper and SeamlessM4T, to develop robust and accurate methods for PD detection from voice recordings. We experiment with both direct feature extraction and finetuning of the foundation models for the PD classification task, and validate the results against clinical and neuroimaging data. We achieve promising results using both pretrained features and models' finetuning, with finetuning providing stronger performance, up to 91.35% for AUC, which is the new state of the art on the ICEBERG dataset. The predictions of our models also show good correlation with clinical as well as DaTSCAN scores, proving the feasibility to apply speech foundation models for detection of early PD. Quang Dao, Laetitia Jeancolas, Graziella Mangone, Sara Sambin, Alizé Chalançon, Manon Gomes, Stéphane Lehéricy, Jean-Christophe Corvol, Marie Vidailhet, Isabelle Arnulf, Dijana Petrovska-Delacrétaz, Mounim A. El-Yacoubi |
IEEE J. Biomed. Health Informatics | 12 |
| 2025 | Hybrid Transformer for Early Alzheimer's Detection: Integration of Handwriting-Based 2D Images and 1D Signal FeaturesabstractAlzheimer's Disease (AD) is a prevalent neurodegenerative condition where early detection is vital. Handwriting, often affected early in AD, offers a non-invasive and cost-effective way to capture subtle motor changes. State-of-the-art research on handwriting, mostly online, based AD detection has predominantly relied on manually extracted features, fed as input to shallow machine learning models. Some recent works have proposed deep learning (DL)-based models, either 1D-CNN or 2D-CNN architectures, with performance comparing favorably to handcrafted schemes. These approaches, however, overlook the intrinsic relationship between the 2D spatial patterns of handwriting strokes and their 1D dynamic characteristics, thus limiting their capacity to capture the multimodal nature of handwriting data. Moreover, the application of Transformer models remains basically unexplored. To address these limitations, we propose a novel approach for AD detection, consisting of a learnable multimodal hybrid attention model that integrates simultaneously 2D handwriting images with 1D dynamic handwriting signals. Our model leverages a gated mechanism to combine similarity and difference attention, blending the two modalities and learning robust features by incorporating information at different scales. Our model achieved state-of-the-art performance on the DARWIN dataset, with an F1-score of 90.32% and accuracy of 90.91% in Task 8 ('L' writing), surpassing the previous best by 4.61% and 6.06% respectively. Huafeng Qin, Mounim A. El-Yacoubi |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | SUMix: Mixup with Semantic and Uncertain Information
Huafeng Qin, Xin Jin 0009, Hongchao Liao, Mounim A. El-Yacoubi, Xinbo Gao 0001 |
ECCV (87) | 5 |
| 2024 | Exploring the Efficacy of Text Embeddings in Early Dementia Diagnosis from SpeechabstractLanguage impairment is a key biomarker for neurodegenerative diseases such as Alzheimer's dis-ease (AD). With the rapid growth of Large Language Models, natural language processing (NLP) has be-come a preferred modality for the early prediction of AD from speech. In this work, we propose a two-stage process for early detection of AD from transcriptions of speech. The first step involves extracting a discriminative text embedding representation using public models from OpenAI. This embedding serves as input for a machine learning classifier in the second stage. In this paper, we investigate three text embedding models and eight machine learning classifiers, both deep learning (DL) based and non-DL based. The evaluation was conducted using the public ADReSSo dataset of 237 patients. The results show that models “ada-002” and “3-small” produce discriminative embeddings that lead to good performance when combined with a Deep Neural Network in classification, achieving accuracy rates of 83.10% and 84.51 %, respectively. Khaoula Ajroudi, Mohamed Ibn Khedher, Olfa Jemai, Mounim A. El-Yacoubi |
HSI | 4 |
| 2024 | Soft-Attention Based Person Re-Identification in Real-world Settings using Variational AutoEncodersabstractPerson re-identification is still an open challenging task in various fields due to numerous factors, including illumination changes, background clutter, pose state variations and cloth changes. Several approaches have been suggested to address this problem in the context of deep learning. Generative models, particularly Variational Autoencoders (VAEs), have emerged as promising tools to address these challenges by learning discriminative feature representations of individual images. In this paper, we present Soft-Attention based Person Re-Identification (SAPRI), a novel approach that combines VAEs with a supervised ReID method to enhance the resilience and efficacy of ReID systems. The proposed approach focuses on data reconstruction based on soft attention. Variational autoen-coders encode principally person data, while ignoring irrelevant information. By incorporating supervised ReID, the model learns to appropriately classify persons in real world environments. Our SAPRI proposed method has been evaluated on well-known benchmarks, DukeMTMC-reID and CUHK03, demonstrating superior performance compared to existing state-of-the-art techniques in terms of the mean Average Precision evaluation metric (mAP). Additionally, qualitative results show the effectiveness of the VAE in generating discriminative representations of person images. Emna Ben Baoues, Imen Jegham, Mounim A. El-Yacoubi, Anouar Ben Khalifa |
HSI | 3 |
| 2024 | Human Pose Estimation Based Biomechanical Feature Extraction for Long JumpsabstractBiomechanical features describing movements and poses of athletes have been proposed by experts to help study athletic performances, but the traditional way of measuring those features are high-cost, time-consuming and intrusive. In this paper, we propose a deep learning-based method that can estimate athletic biomechanical features from typical broadcast competition videos, i.e. single-camera-shot moving videos. This method involves state-of-the-art human pose estimation models and a biomechanical analysis to reconstruct the trajectory. We then leverage the reconstructed trajectory to estimate the target features. To evaluate the method, we gathered a dataset from the long jump World Championships of 2017 and 2018, comprising 22 expert-proposed long-jump biomechanical features about the trajectories, taking-off and landing characteristics. Our experiments show the effectiveness of the pipeline in automatically estimating the biomechanical features. By analysing the results, we identify the challenges towards high-accuracy athletes' feature estimations from monocular broadcast competition videos. Code is available at https://github.com/QGAN2019/Long_Jump_Feature_Estimation. Qi Gan, Sao Mai Nguyen, Mounim A. El-Yacoubi, Eric Fenaux, Stéphan Clémençon |
HSI | 3 |
| 2024 | Assessing the Interpretability of Machine Learning Models in Early Detection of Alzheimer's DiseaseabstractAlzheimer's disease (AD) is a chronic and irreversible neurological disorder, making early detection essential for managing its progression. This study investigates the coherence of SHAP values with medical scientific truth. It examines three types of features: clinical, demographic, and FreeSurfer extracted from MRI scans. A set of six ML classifiers are investigated for their interpretability levels. This study is validated on the OASIS-3 dataset with binary classification. The results show that clinical data outperforms the others, with a margin of 14% over FreeSurfer features, the second-best features. In the case of clinical features, the explanations provided by the tree-based classifiers consistently align with medical insights. This comparison was calculated using the Kendall Tau distance. Karim Haddada, Mohamed Ibn Khedher, Olfa Jemai, Sarra Iben Khedher, Mounim A. El-Yacoubi |
HSI | 5 |
| 2024 | GT&I GAN: A Generative Adversarial Network for Data Augmentation in Regression and Segmentation TasksabstractFor data augmentation (DA), Generative Adversarial Networks (GANs) are typically integrated with CNNs or MLPs to generate samples in classification and segmentation tasks. For classification, categorical ground truth is leveraged in conditional GANs to generate samples for each class. For regression, data generation becomes complex as the aim now is to generate both the samples (images) and their continuous ground truth vectors. GANs for classification can no longer, therefore, be leveraged for DA on regression. To address this issue, we propose GT&I_GAN, a novel GAN-based DA model that generates jointly image samples and their ground truth continuous vectors by learning their conjoint distribution. The main idea behind GT$\mathbf{\& I-GAN}$is to add, to the RGB sample image, an additional (fourth) channel associated with the ground vector. GT&I_GAN offers the great advantage of generating conjointly the samples and their ground truths by a single model without needing an additional network. We assess our approach on an image dataset where the ground truth consists of a high dimensional vector of continuous values. The results show that the synthetic data consisting of the image & ground truth vector pairs are realistic and allow improving the CNN regressor performance. Moreover, we show that our GT&I_GAN can be leveraged seamlessly for segmentation tasks by adding, in a similar way, the ground truth segmentation mask as an additional channel to the input RGB image. Hajar Hammouch, Sambit Mohapatra, Mounim A. El-Yacoubi, Huafeng Qin, Hassan Berbia |
HSI | 3 |
| 2024 | GAN et: Gabor Attention Aggregation Network for Palmvein IdentificationabstractPalm vein recognition has attracted recently wide attention thanks to its robust feature representation and high accuracy. Despite advancements in the literature, however, existing solutions suffer from the following issues: 1) Insufficient large-scale data for deep learning-based recognition of vein biometrics, resulting in decreased generalization performance and model accuracy. 2) Lack of methods based on machine learning convolutional neural networks capable of capturing the global receptive field for vein biometric recognition. In addressing these issues, this paper proposes a method to acquire the global receptive field, termed G AN et, which extracts vein features using Gabor filters and computes an attention mechanism to capture the global receptive field for downstream palm vein recognition models. Initially, vein features are extracted using multi-scale fixed Gabor filters and multi-scale adaptive Gabor filters. Subsequently, self-attention mechanisms are employed to compute relationships between blocks to obtain the global receptive field. To perform recognition, the Euclidean distance between feature vectors is then computed. Our experiments on three datasets show that our approach outperforms existing palm vein recognition methods. Hongchao Liao, Xin Jin 0009, Yuming Fu 0001, Mounim A. El-Yacoubi, Huafeng Qin |
HSI | 5 |
| 2024 | Early-Stage Parkinson's Disease Detection Based on Optical Flow and Video Vision TransformerabstractHypomimia, a symptom of Parkinson's disease (PD), is marked by reduced facial movements and loss of face emotional expressions. This study focuses on identifying hypomimia in individuals with early-stage PD using optical-flow-based video vision transformer. Our study included video recordings from 109 PD and 45 healthy control (HC) subjects with an average of two videos per person (294 videos in total). The participants asked to speak freely while being recorded. To extract typical facial muscle movements from subjects, we computed the optical flow (OF) from the videos. Video vision transformer is then used to infer feature representations from OF and RGB modalities, input to a Random Forest (RF) classifier to classify PD vs. HC. We obtained classification scores up to 83% in terms of balanced accuracy (BA) and an area under the curve (AUC) of 84% at subject level. The results are promising for identifying hypomimia in the early stages of PD, and this research could lead to the possibility of continuous monitoring of hypomimia outside of hospital settings via telemedicine. Anas Filali Razzouki, Laetitia Jeancolas, Graziella Mangone, Sara Sambin, Alizé Chalançon, Manon Gomes, Stéphane Lehéricy, Jean-Christophe Corvol, Marie Vidailhet, Isabelle Arnulf, Mounim A. El-Yacoubi, Dijana Petrovska-Delacrétaz |
HSI | 11 |
| 2024 | PredictStr: A Balanced Benchmark Dataset for Improve Stroke PredictionabstractPredicting strokes is essential for improving healthcare outcomes and saving lives. This paper introduces a benchmarking dataset, PredictStr, specifically developed to enhance stroke prediction. This dataset improves upon a previously unique dataset identified in the literature. Our methodology comprises two main steps: firstly, we outline a series of preprocessing and cleaning measures to enhance data quality. Secondly, we present a novel algorithm, the Dynamic Hybrid Balancing Algorithm, which builds upon the ADSYSN algorithm by integrating consistency constraints to address class imbalances. Our contribution extends to the application of sophisticated analysis techniques, including histogram and boxplot analyses, feature distribution assessments, statistical explorations, correlation evaluations, feature importance rankings, and Individual Conditional Expectation (ICE) plots. These methodologies are designed to provide valuable insights into feature significance, thereby assisting researchers in identifying the most critical attributes for effective stroke detection. Taissir Fekih Romdhane, Mohamed Ibn Khedher, Mounim A. El-Yacoubi |
HSI | 3 |
| 2024 | Improving Alzheimer's Diagnosis Using Vision Transformers and Transfer LearningabstractAlzheimer's disease is a neurodegenerative disorder defined by memory loss and primarily affects older individuals. Currently, there is no definitive cure available. Although medications are accessible, they only serve to slow the progression of the disease. In this paper, we propose the use of Vision Transformers and Transfer Learning for Alzheimer's classification. Our approach leverages the temporal aspect of the transformer to model the correlation between different image patches. Transfer learning enables us to mitigate the issue of insufficient available data. Our method has been validated on the OASIS dataset, which consists of 250 brain scans. The results demonstrate that transfer learning with Transformer models surpasses the performance of transfer learning with CNN models by 4% and exceeds traditional CNN models without transfer learning by 8%. Two types of Transformers were tested: ViT-B16 and ViT-B32. The results are comparable, with ViT-B32 outperforming ViT-B16 by 1%. Marwa Zaabi, Mohamed Ibn Khedher, Mounim A. El-Yacoubi |
HSI | 3 |
| 2024 | Adversarial AutoMixupabstractData mixing augmentation has been widely applied to improve the generalization ability of deep neural networks. Recently, offline data mixing augmentation, e.g. handcrafted and saliency information-based mixup, has been gradually replaced by automatic mixing approaches. Through minimizing two sub-tasks, namely, mixed sample generation and mixup classification in an end-to-end way, AutoMix significantly improves accuracy on image classification tasks. However, as the optimization objective is consistent for the two sub-tasks, this approach is prone to generating consistent instead of diverse mixed samples, which results in overfitting for target task training. In this paper, we propose AdAutomixup, an adversarial automatic mixup augmentation approach that generates challenging samples to train a robust classifier for image classification, by alternatively optimizing the classifier and the mixup sample generator. AdAutomixup comprises two modules, a mixed example generator, and a target classifier. The mixed sample generator aims to produce hard mixed examples to challenge the target classifier, while the target classifier's aim is to learn robust features from hard mixed examples to improve generalization. To prevent the collapse of the inherent meanings of images, we further introduce an exponential moving average (EMA) teacher and cosine similarity to train AdAutomixup in an end-to-end way. Extensive experiments on seven image benchmarks consistently prove that our approach outperforms the state of the art in various classification scenarios. The source code is available at
https://github.com/JinXins/Adversarial-AutoMixup. Huafeng Qin, Xin Jin 0009, Mounim A. El-Yacoubi, Xinbo Gao 0001 |
ICLR | 4 |
| 2024 | Special section: Best papers of the international conference on pattern recognition and artificial intelligence (ICPRAI) 2022
Mounim A. El-Yacoubi, Umapada Pal 0001, Eric Granger, Pong C. Yuen |
Pattern Recognit. Lett. | 1 |
| 2024 | Adversarial Learning-Based Data Augmentation for Palm-Vein IdentificationabstractPalm-vein identification is a highly secure pattern biometrics that has become an active research area in recent years. Despite the recent progress in deep neural networks (DNNs) for vein identification, existing solutions for feature representation continue to lack robustness due to the limited training samples. To address this limitation, data augmentation approaches, including Generative Adversarial Networks (GANs), have been investigated, but these schemes suffer from the following issues. First, it is practically unfeasible to use all the generated samples for classifier training due to the limited storage space and computation resources. Further, some of these generated samples may be non-representative or ineffective, seriously compromising models’ generalization capabilities. Second, the augmented dataset is fed to the target classifier repeatedly, resulting in overfitting after substantial training epochs. To tackle the above problems, we propose AdveinAU, an Adversarial vein AUtomatic AUgmentation approach that generates challenging samples to train a more robust vein classifier for palm-vein identification by alternatively optimizing the vein classifier and a set of latent variables. First, we consider a conditional deep convolution generative adversarial net (cDCGAN) to learn the distribution of real data and the generated data, and then a latent variable from the latent variable space is mapped to the sample space. Second, we combine the trained generator with the vein classifier to constitute AdveinAU, where the input sets of the generator and the classifier are alternatively updated by adversarial training. Specifically, a latent variable set is learned to increase the training loss of a target network through generating adversarial samples, while the classifier learns more robust features from harder examples to improve the generalization. To avoid collapsing inherent meanings of images, an exponential moving average (EMA) teacher andcosinesimilarity are employed for regularization to reduce the search space. Unlike previous works where GANs synthesize new realistic images, our model aims to search a latent variable set, based on which the generator can produce challenging samples along with the training process to improve the classifier’s performance. Finally, we conduct extensive experiments on three public palm-vein datasets to evaluate the performance of AdveinAU, and the experimental results demonstrate that the proposed AdveinAU is capable of generating harder samples to improve the performance of the vein classifier. Huafeng Qin, Haofei Xi, Yantao Li 0001, Mounim A. El-Yacoubi, Jun Wang 0071, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | AG-NAS: An Attention GRU-Based Neural Architecture Search for Finger-Vein RecognitionabstractFinger-vein recognition has attracted extensive attention due to its exceptional level of security and privacy. Recently, deep neural networks (DNNs), such as convolutional neural networks (CNNs) showing robust capacity for feature representation, have been proposed for vein recognition. The architectures of these DNNs, however, have primarily been manually designed based on human prior knowledge, which is both time-consuming and error-prone. To overcome these problems, we propose AG-NAS, an Attention Gated recurrent unit-based Neural Architecture Search to automatically search for the optimal network architecture, thereby improving the recognition performance for different finger-vein recognition tasks. First, we combine the self-attention mechanism and gated recurrent unit (GRU) to propose an attention GRU module employed as a controller to generate the architectural hyperparameters of candidate neural networks automatically. Second, we investigate a parameter-sharing supernet policy to reduce the search space, computation, and time costs. Finally, we conduct rigorous experiments on our finger-vein database and two public finger-vein databases. The experimental results demonstrate that the proposed AG-NAS outperforms the representative approaches and achieves state-of-the-art recognition accuracy. Huafeng Qin, Shaojiang Deng, Yantao Li 0001, Mounim A. El-Yacoubi, Gang Zhou 0002 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Attention BLSTM-Based Temporal-Spatial Vein Transformer for Multi-View Finger-Vein RecognitionabstractFinger-vein biometrics has recently gained significant attention due to its robust privacy and high security features. Despite notable advancements, most existing methods focus on extracting features from a 2-dimensional (2D) image projected from 3D vein vessels with a single view. However, recognition based on a single view is prone to errors due to variations in finger positioning, especially those caused by finger roll movements, which can degrade recognition performance. To address this challenge, we propose ABLSTM-TSVT, an Attention Bidirectional LSTM-based Temporal-Spatial Vein Transformer for multi-view finger-vein recognition. First, we enhance LSTM with an attention mechanism to create an attention LSTM for extracting temporal features. We further improve this by introducing a local attention module, which learns temporal dependencies between a patch (token) and its adjacent patches across multiple views, integrating it with the attention LSTM to form a temporal attention module. Second, we develop a spatial attention module that captures the spatial dependencies of patches within an image. Finally, merging the temporal and the spatial attention modules, we create our temporal-spatial transformer model, which effectively represents features from multi-view images. Experimental results on two multi-view datasets demonstrate that our approach outperforms state-of-the-art approaches in enhancing identification accuracy and reducing verification errors in vein classifiers. Huafeng Qin, Zhipeng Xiong, Yantao Li 0001, Mounim A. El-Yacoubi, Jun Wang 0071 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Memory-Augmented Autoencoder Based Continuous Authentication on Smartphones With Conditional Transformer GANsabstractOver the last years, sensor-based continuous authentication on mobile devices has achieved great success on personal information protection. These proposed mechanisms, however, require both legal and illegal users’ data for authentication model training, which takes time and is impractical. In this paper, we present MAuGANs, a lightweight and practical Memory-Augmented Autoencoder-based continuous Authentication system on smartphones with conditional transformer Generative Adversarial Networks (GANs), where the conditional transformer GANs (CTGANs) are used for data augmentation and the memory-augmented autoencoder (MAu) is utilized to identify users. Specifically, MAuGANs exploits the smartphone built-in accelerometer and gyroscope sensors to implicitly collect users’ behavioral patterns. With the normalized legitimate user's sensor data, MAuGANs uses a CTGAN composed of a conditional transformer-based generator and a conditional transformer-based discriminator to create additional training data for the MAu. Then, the MAu is trained on the augmented legitimate user's data. The trained MAu reconstructs the current user data and then calculates the reconstruction error between the reconstructed data and current user data. To carry out user authentication, MAuGANs compares the reconstruction error with a predefined authentication threshold. We evaluate the performance of MAuGANs on our dataset, where our extensive experiments demonstrate that MAuGANs reaches the best authentication performance, when comparing with the representative state-of-the-art methods, by 0.33% EER and 99.65% accuracy on 10 unseen users. Yantao Li 0001, Shaojiang Deng, Huafeng Qin, Mounim A. El-Yacoubi, Gang Zhou 0002 |
IEEE Trans. Mob. Comput. | 5 |
| 2023 | Local Attention Transformer-Based Full-View Finger-Vein IdentificationabstractMulti-view finger-vein recognition technology has attracted increasing attentions in recent years. Despite recent advances in the multi-view finger-vein identification, existing solutions employ multiple monocular cameras from different views to record two-dimensional (2D) projections of 3D vein vessels, which causes the following problems: 1) 2D images collected from limited views (two or three views) are insufficient for robust 3D vein vessel feature representation. Furthermore, image sequences of the same finger acquired from different views usually show significant differences. As a result, the existing works are still sensitive to positional variations of the fingers, specifically those caused by finger roll movements. 2) Using multiple cameras can lead to increased costs. Moreover, it is impossible to employ several cameras to acquire full-view images because of the limited space on capturing devices. To address the above issues, we present$\mathbb {FV}$-LT, a Full-View Finger-Vein identification system based on a Local attention Transformer, by implementing an image acquisition device with a single camera. First, we design and implement a finger-vein acquisition prototype device that utilizes a single camera and a LED group to rotate along a finger for full-view image collection. This allows capturing all vein patterns concealed beneath human skin to form a complete representation of finger features. Second, given the full-view vein images, we propose a local attention transformer-based approach to extract dependency features of a token (a patch or an image) on its neighborhood’s tokens among image patches and among full-view images, respectively. These dependency features are shown to be robust to positional variations induced by finger rolls. Based on the public database of full-view finger-vein images captured by our designed device and a single-view database, we verify the performance of the proposed$\mathbb {FV}$-LT. The experimental results show that$\mathbb {FV}$-LT significantly outperforms existing 2D/multi-view based approaches with respect to improving the tolerance against finger roll and achieving the state-of-the-art identification accuracy. Huafeng Qin, Rongshan Hu, Mounim A. El-Yacoubi, Yantao Li 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Transformer Based Defense GAN Against Palm-Vein Adversarial AttacksabstractVein biometrics is a high security and privacy preserving identification technology that has attracted increasing attention over the last decade. Deep neural networks (DNNs), such as convolutional neural networks (CNN), have shown strong capabilities for robust feature representation, and have achieved, as a result, state-of-the-art performance on various vision tasks. Inspired by their success, deep learning models have been widely investigated for vein recognition and have shown significant improvement of identification accuracy compared to handcrafted models. Existing deep learning models, however, are vulnerable to adversarial perturbation attacks, where thoughtfully crafted small perturbations can cause misclassification of legitimate images, degrading, thereby, the efficiency of vein recognition systems. To address this problem, we propose, in this paper, VeinGuard, a novel defense framework to defend deep learning classifiers against adversarial palm-vein image attacks, composed of a local transformer-based GAN and a purifier. VeinGuard comprises two components: a local transformer-based GAN (LTGAN) that learns the distribution of unperturbed vein images and generates high-quality palm-vein images, and a purifier consisting of a trainable residual network and of a pre-trained generator from LTGAN that automatically removes a wide variety of adversarial perturbations. The resulting clean images are fed to vein classifiers for identification, thereby avoiding adversarial attacks. We evaluate VeinGuard on three public vein datasets in terms of white-box attacks, black-box attacks, ablation experiments, and computation time. The experimental results show that VeinGuard allows filtering the perturbations and enables the classifiers to achieve state-of-the-art recognition results for different adversarial attacks. Yantao Li 0001, Song Ruan, Huafeng Qin, Shaojiang Deng, Mounim A. El-Yacoubi |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | Adaptive Deep Feature Fusion for Continuous Authentication With Data AugmentationabstractMobile devices are becoming increasingly popular and are playing significant roles in our daily lives. Insufficient security and weak protection mechanisms, however, cause serious privacy leakage of the unattended devices. To fully protect mobile device privacy, we propose ADFFDA, a novel mobile continuous authentication system using an Adaptive Deep Feature Fusion scheme for effective feature representation, and a transformer-based GAN for Data Augmentation, by leveraging smartphone built-in sensors of the accelerometer, gyroscope and magnetometer. Given the normalized sensor data, ADFFDA utilizes the transformer-based GAN consisting of a transformer-based generator and a CNN-based discriminator to augment the training data for CNN training. With the augmented data and the especially-designed CNN based on the ghost module and ghost bottleneck, ADFFDA extracts deep features from the three sensors by the trained CNN, and exploits an adaptive-weighted concatenation method to adaptively fuse the CNN-extracted features. Based on the fused features, ADFFDA authenticates users by using the one-class SVM (OC-SVM) classifier. We evaluate the authentication performance of ADFFDA in terms of the efficiency of the transformer-based GAN, GAN-based data augmentation, CNN architecture, adaptive-weighted feature fusion, OC-SVM classifier, and security analysis. The experimental results show that ADFFDA obtains the best authentication performance w.r.t representative approaches, by achieving a mean equal error rate of 0.01%. Yantao Li 0001, Huafeng Qin, Shaojiang Deng, Mounim A. El-Yacoubi, Gang Zhou 0002 |
IEEE Trans. Mob. Comput. | 5 |
| 2021 | Enhancing the Interpretability of Deep Models in Healthcare Through Attention: Application to Glucose Forecasting for Diabetic PeopleabstractThe adoption of deep learning in healthcare is hindered by their “black box” nature. In this paper, we explore the RETAIN architecture for the task of glucose forecasting for diabetic people. By using a two-level attention mechanism, the recurrent-neural-network-based RETAIN model is interpretable. We evaluate the RETAIN model on the type-2 IDIAB and the type-1 OhioT1DM datasets by comparing its statistical and clinical performances against two deep models and three models based on decision trees. We show that the RETAIN model offers a very good compromise between accuracy and interpretability, being almost as accurate as the LSTM and FCN models while remaining interpretable. We show the usefulness of its interpretable nature by analyzing the contribution of each variable to the final prediction. It revealed that signal values older than 1[Formula: see text]h are not used by the RETAIN model for 30[Formula: see text]min ahead of time prediction of glucose. Also, we show how the RETAIN model changes its behavior upon the arrival of an event such as carbohydrate intakes or insulin infusions. In particular, it showed that the patient’s state before the event is particularly important for the prediction. Overall the RETAIN model, thanks to its interpretability, seems to be a very promising model for regression or classification tasks in healthcare. Maxime De Bois, Mounim A. El-Yacoubi, Mehdi Ammi |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2021 | Multi-Scale and Multi-Direction GAN for CNN-Based Single Palm-Vein IdentificationabstractDespite recent advances of deep neural networks in hand vein identification, the existing solutions assume the availability of a large and rich set of training image samples. These solutions, therefore, still lack the capability to extract robust and discriminative hand-vein features from a single training image sample. To overcome this problem, we propose a single-sample-per-person (SSPP) palm-vein identification approach, where only a single sample per class is enrolled in the gallery set for training. Our approach, named MSMDGAN + CNN, consists of a multi-scale and multi-direction generative adversarial network (MSMDGAN) for data augmentation and a convolutional neural network (CNN) for palm-vein identification. First, a novel data augmentation approach, MSMDGAN, is developed to learn the internal distribution of patches in a single image. The proposed MSMDGAN consists of multiple fully convolutional GANs, each of which is responsible for learning the patch distribution within an image at a different scale and at a different direction. Second, given the resulting augmented data by MSMDGAN, we design a CNN for single sample palm-vein recognition. The experimental results on two public hand-vein databases demonstrate that MSMDGAN is able to generate realistic and diverse samples, which, in turn, improves the stability of the CNN. In terms of accuracy, MSMDGAN + CNN outperforms other representative approaches and achieves state-of-the-art recognition results. Huafeng Qin, Mounim A. El-Yacoubi, Yantao Li 0001, Chong-Wen Liu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | Automatic processing of Historical Arabic Documents: A comprehensive Survey
Mohamed Ibn Khedher, Houda Jmila, Mounim A. El-Yacoubi |
Pattern Recognit. | 3 |
| 2019 | Model Fusion to Enhance the Clinical Acceptability of Long-Term Glucose PredictionsabstractThis paper presents the Derivatives Combination Predictor (DCP), a novel model fusion algorithm for making long-term glucose predictions for diabetic people. First, using the history of glucose predictions made by several models, the future glucose variation at a given horizon is predicted. Then, by accumulating the past predicted variations starting from a known glucose value, the fused glucose prediction is computed. A new loss function is introduced to make the DCP model learn to react faster to changes in glucose variations. The algorithm has been tested on 10 in-silico type-1 diabetic children from the T1DMS software. Three initial predictors have been used: a Gaussian process regressor, a feed-forward neural network and an extreme learning machine model. The DCP and two other fusion algorithms have been evaluated at a prediction horizon of 120 minutes with the root-mean-squared error of the prediction, the root-mean-squared error of the predicted variation, and the continuous glucose-error grid analysis. By making a successful trade-off between prediction accuracy and predicted-variation accuracy, the DCP, alongside with its specifically designed loss function, improves the clinical acceptability of the predictions, and therefore the safety of the model for diabetic people. Maxime De Bois, Mehdi Ammi, Mounim A. El-Yacoubi |
BIBE | 3 |
| 2019 | Prediction-Coherent LSTM-Based Recurrent Neural Network for Safer Glucose Predictions in Diabetic People
Maxime De Bois, Mounim A. El-Yacoubi, Mehdi Ammi |
ICONIP (3) | 2 |
| 2019 | Siamese Network Based Feature Learning for Improved Intrusion Detection
Houda Jmila, Mohamed Ibn Khedher, Gregory Blanc, Mounim A. El-Yacoubi |
ICONIP (1) | 4 |
| 2019 | Study of Short-Term Personalized Glucose Predictive Models on Type-1 Diabetic ChildrenabstractResearch in diabetes, especially when it comes to building data-driven models to forecast future glucose values, is hindered by the sensitive nature of the data. Because researchers do not share the same data between studies, progress is hard to assess. This paper aims at comparing the most promising algorithms in the field, namely Feedforward Neural Networks (FFNN), Long Short-Term Memory (LSTM) Recurrent Neural Networks, Extreme Learning Machines (ELM), Support Vector Regression (SVR) and Gaussian Processes (GP). They are personalized and trained on a population of 10 virtual children from the Type 1 Diabetes Metabolic Simulator software to predict future glucose values at a prediction horizon of 30 minutes. The performances of the models are evaluated using the Root Mean Squared Error (RMSE) and the Continuous Glucose-Error Grid Analysis (CG-EGA). While most of the models end up having low RMSE, the GP model with a Dot-Product kernel (GP-DP), a novel usage in the context of glucose prediction, has the lowest. Despite having good RMSE values, we show that the models do not necessarily exhibit a good clinical acceptability, measured by the CG-EGA. Only the LSTM, SVR and GP-DP models have overall acceptable results, each of them performing best in one of the glycemia regions. Maxime De Bois, Mounim A. El-Yacoubi, Mehdi Ammi |
IJCNN | 2 |
| 2019 | Unsupervised deep neuron-per-neuron hashing
Sanaa Chafik, Mounim A. El-Yacoubi, Imane Daoudi, Hamid El Ouardi |
Appl. Intell. | 2 |
| 2019 | Finger-Vein Quality Assessment Based on Deep Features From Grayscale and Binary ImagesabstractFinger-vein verification is a highly secure biometric authentication that has been widely investigated over the last years. One of its challenges, however, is the possible degradation of image quality, that results in spurious and missing vein patterns, which increases the verification error. Despite recent advances in finger-vein quality assessment, the proposed solutions are limited as they depend on human expertise and domain knowledge to extract handcrafted features for assessing quality. We have proposed, recently, the first deep neural network (DNN) framework for assessing finger-vein quality, that does not require manual labeling of high and low quality images, as is the case for state of the art methods, but infers such annotations automatically based on an objective indicator, the biometric verification decision. This framework has significantly outperformed the existing methods, whether the input image is in grayscale or is binary. Motivated by these performances, we propose, in this work, a representation learning of finger vein image quality, where a DNN takes as input conjointly the grayscale and binary versions of the input image to predict vein quality. Our model allows to learn the joint representation from grayscale and binary images, for quality assessment. The experimental results, obtained on a large public dataset, demonstrates that our proposed method accurately identifies high and low quality images, and outperforms other techniques in terms of equal error rate (EER) minimization, including our previous DNN models, based either on grayscale or binary input. Huafeng Qin, Mounim A. El-Yacoubi |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2019 | From aging to early-stage Alzheimer's: Uncovering handwriting multimodal behaviors by semi-supervised learning and sequential representation learning
Mounim A. El-Yacoubi, Sonia Garcia-Salicetti, Christian Kahindo, Anne-Sophie Rigaud, Victoria Cristancho-Lacroix |
Pattern Recognit. | 1 |
| 2019 | Estimation of Static and Dynamic Urban Populations with Mobile Network MetadataabstractCommunication-enabled devices routinely carried by individuals have become pervasive, opening unprecedented opportunities for collecting digital metadata about the mobility of large populations. In this paper, we propose a novel methodology for the estimation of people density at metropolitan scales, using subscriber presence metadata collected by a mobile operator. Our approach suits the estimation of static population densities, i.e., of the distribution of dwelling units per urban area contained in traditional censuses. More importantly, it enables the estimation of dynamic population densities, i.e., the time-varying distributions of people in a conurbation. By leveraging substantial real-world mobile network metadata and ground-truth information, we demonstrate that the accuracy of our solution is superior to that granted by state-of-the-art methods in practical heterogeneous urban scenarios. Ghazaleh Khodabandelou, Vincent Gauthier, Marco Fiore 0001, Mounim A. El-Yacoubi |
IEEE Trans. Mob. Comput. | 4 |
| 2018 | Fusion of Interest Point/Image based descriptors for efficient person re-identificationabstractThe paper proposes a novel video-based person re-identification system that consists of describing a person using both Interest Points (IP) and Image-based features. The Image-based descriptor extracts global image representation that includes the silhouette but also possibly extra objects (i.e animal, stroller, etc) while the IP-based descriptor extracts salient points associated each with a local region of one of the objects. Two reidentification systems are proposed: an IP-based system using SURF interest points matched via sparse representation, and Image-based system using a Convolutional Neural Network. To harness both representations, we propose a fusing strategy based on the scores product rule, the scores being vote vectors associated with each descriptor for each person. Our proposal is evaluated on the large public dataset PRID-2011 and the results show its effectiveness compared to the state of the art. Mohamed Ibn Khedher, Houda Jmila, Mounim A. El-Yacoubi |
IJCNN | 3 |
| 2018 | Combining Bayesian Inference and Clustering for Transport Mode Detection from Sparse and Noisy Geolocation Data
Danya Bachir, Ghazaleh Khodabandelou, Vincent Gauthier, Mounim A. El-Yacoubi, Eric Vachon |
ECML/PKDD (3) | 4 |
| 2018 | Characterizing Early-Stage Alzheimer Through Spatiotemporal Dynamics of HandwritingabstractWe propose an original approach for characterizing early Alzheimer, based on the analysis of online handwritten cursive loops. Unlike the literature, we model the loop velocity trajectory (full dynamics) in an unsupervised way. Through a temporal clustering based on K-medoids, with dynamic time warping as dissimilarity measure, we uncover clusters that give new insights on the problem. For classification, we consider a Bayesian formalism that aggregates the contributions of the clusters, by probabilistically combining the discriminative power of each. On a dataset consisting of two cognitive profiles, early-stage Alzheimer disease and healthy persons, each comprising 27 persons collected at Broca Hospital in Paris, our classification performance significantly outperforms the state-of-the-art, based on global kinematic features. Christian Kahindo, Mounim A. El-Yacoubi, Sonia Garcia-Salicetti, Anne-Sophie Rigaud, Victoria Cristancho-Lacroix |
IEEE Signal Process. Lett. | 2 |
| 2018 | Deep Representation for Finger-Vein Image-Quality AssessmentabstractFinger-vein biometrics has been extensively investigated for personal authentication. One of the open issues in finger-vein verification is the lack of robustness against image-quality degradation. Spurious and missing features in poor-quality images may degrade the system's performance. Despite recent advances in finger-vein quality assessment, current solutions depend on domain knowledge. In this paper, we propose a deep neural network (DNN) for representation learning to predict image quality using very limited knowledge. Driven by the primary target of biometric quality assessment, i.e., verification error minimization, we assume that low-quality images are falsely rejected in a verification system. Based on this assumption, the low- and high-quality images are labeled automatically. We then train a DNN on the resulting data set to predict the image quality. To further improve the DNN's robustness, the finger-vein image is divided into various patches, on which a patch-based DNN is trained. The deepest layers associated with the patches form together a complementary and an over-complete representation. Subsequently, the quality of each patch from a testing image is estimated and the quality scores from the image patches are conjointly input to probabilistic support vector machines (P-SVM) to boost quality-assessment performance. To the best of our knowledge, this is the first proposed work of deep learning-based quality assessment, not only for finger-vein biometrics, but also for other biometrics in general. The experimental results on two public finger-vein databases show that the proposed scheme accurately identifies high- and low-quality images and significantly outperforms existing approaches in terms of the impact on equal error-rate decrease. Huafeng Qin, Mounim A. El-Yacoubi |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Refining Visual Activity Recognition with Semantic ReasoningabstractAs elderly care is getting more and more important, monitoring of activity of daily living (ADL) has become an active research topic. Both robotic and pervasive computing domains, through smart homes, are creating opportunities to move forward in ADL field. Multiple techniques were proposed to identify activities, each with their features, advantages and limits. However, it is a very challenging issue and none of the existing methods provides robust results, in particular in real daily living scenarios. This is particularly true for visionbased approaches used by robots. In this paper, we propose to refine a robot's visual activity recognition process by relying on smart home sensors. We assert that the consideration of further sensors and the knowledge about the target user together with the semantic by means of an ontology and a reasoning layer in the recognition process, has improved the existing works results. We experimented through multiple activity recognition scenarios with and without refinement to assess the relevance of such a combination. Although our tests reveal positive results, they also point out limits and challenges that we discuss in this paper. Nathan Ramoly, Vincent Vassout, Amel Bouzeghoub, Mounim A. El-Yacoubi, Mossaab Hariz |
AINA | 4 |
| 2017 | Comparing Hybrid NN-HMM and RNN for Temporal Modeling in Gesture Recognition
Nicolas Granger 0001, Mounim A. El-Yacoubi |
ICONIP (2) | 2 |
| 2017 | Estimating VNF Resource Requirements Using Machine Learning Techniques
Houda Jmila, Mohamed Ibn Khedher, Mounim A. El-Yacoubi |
ICONIP (1) | 3 |
| 2017 | Fusion of appearance and motion-based sparse representations for multi-shot person re-identification
Mohamed Ibn Khedher, Mounim A. El-Yacoubi, Bernadette Dorizzi |
Neurocomputing | 2 |
| 2017 | Deep Representation-Based Feature Extraction and Recovering for Finger-Vein VerificationabstractFinger-vein biometrics has been extensively investigated for personal verification. Despite recent advances in finger-vein verification, current solutions completely depend on domain knowledge and still lack the robustness to extract finger-vein features from raw images. This paper proposes a deep learning model to extract and recover vein features using limited a priori knowledge. First, based on a combination of the known state-of-the-art handcrafted finger-vein image segmentation techniques, we automatically identify two regions: a clear region with high separability between finger-vein patterns and background, and an ambiguous region with low separability between them. The first is associated with pixels on which all the above-mentioned segmentation techniques assign the same segmentation label (either foreground or background), while the second corresponds to all the remaining pixels. This scheme is used to automatically discard the ambiguous region and to label the pixels of the clear region as foreground or background. A training data set is constructed based on the patches centered on the labeled pixels. Second, a convolutional neural network (CNN) is trained on the resulting data set to predict the probability of each pixel of being foreground (i.e., vein pixel), given a patch centered on it. The CNN learns what a finger-vein pattern is by learning the difference between vein patterns and background ones. The pixels in any region of a test image can then be classified effectively. Third, we propose another new and original contribution by developing and investigating a fully convolutional network to recover missing finger-vein patterns in the segmented image. The experimental results on two public finger-vein databases show a significant improvement in terms of finger-vein verification accuracy. Huafeng Qin, Mounim A. El-Yacoubi |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2016 | Multimodal Sequential Modeling and Recognition of Human Activities
Mouna Selmi, Mounim A. El-Yacoubi |
ICCHP (2) | 2 |
| 2016 | Population estimation from mobile network traffic metadataabstractSmartphones and other mobile devices are today pervasive across the globe. As an interesting side effect of the surge in mobile communications, mobile network operators can now easily collect a wealth of high-resolution data on the habits of large user populations. The information extracted from mobile network traffic data is very relevant in the context of population mapping: it provides a tool for the automatic and live estimation of population densities, overcoming the limitations of traditional data sources such as censuses and surveys. In this paper, we propose a new approach to infer population densities at urban scales, based on aggregated mobile network traffic metadata. Our approach allows estimating both static and dynamic populations, achieves a significant improvement in terms of accuracy with respect to state-of-the-art solutions in the literature, and is validated on different city scenarios. Ghazaleh Khodabandelou, Vincent Gauthier, Mounim A. El-Yacoubi, Marco Fiore 0001 |
WoWMoM | 3 |
| 2016 | CT-Mapper: Mapping sparse multimodal cellular trajectories using a multilayer transportation networkabstractMobile phone data have recently become an attractive source of information about mobility behavior. Since cell phone data can be captured in a passive way for a large user population, they can be harnessed to collect well-sampled mobility information. In this paper, we propose CT-Mapper , an unsupervised algorithm that enables the mapping of mobile phone traces over a multimodal transport network. One of the main strengths of CT-Mapper is its capability to map noisy sparse cellular multimodal trajectories over a multilayer transportation network where the layers have different physical properties and not only to map trajectories associated with a single layer. Such a network is modeled by a large multilayer graph in which the nodes correspond to metro/train stations or road intersections and edges correspond to connections between them. The mapping problem is modeled by an unsupervised HMM where the observations correspond to sparse user mobile trajectories and the hidden states to the multilayer graph nodes. The HMM is unsupervised as the transition and emission probabilities are inferred using respectively the physical transportation properties and the information on the spatial coverage of antenna base stations. To evaluate CT-Mapper we collected cellular traces with their corresponding GPS trajectories for a group of volunteer users in Paris and vicinity (France). We show that CT-Mapper is able to accurately retrieve the real cell phone user paths despite the sparsity of the observed trace trajectories. Furthermore our transition probability model is up to 20% more accurate than other naive models. Fereshteh Asgari, Alexis Sultan, Haoyi Xiong, Vincent Gauthier, Mounim A. El-Yacoubi |
Comput. Commun. | 5 |
| 2016 | Two-layer discriminative model for human activity recognitionabstractMost of recent methods for action/activity recognition, usually based on static classifiers, have achieved improvements by integrating context of local interest point (IP) features such as spatiotemporal IPs by characterising their neighbourhood under different scales. In this study, the authors propose a new approach that explicitly models the sequential aspect of activities. First, a sliding window segmentation technique splits the video stream into overlapping short segments. Each window is characterised by a local bag of words of IPs encoded by motion information. A first‐layer support vector machine provides for each window a vector of conditional class probabilities that summarises all discriminant information that is relevant for sequence recognition. The sequence of these stochastic vectors is then fed to a hidden conditional random field for inference at the sequence level. They also show how their approach can be naturally extended to the problem of conjoint segmentation and recognition of a sequence of action classes within a continuous video stream. They have tested their model on various human action and activity datasets and the obtained results compare favourably with current state of the art. Mouna Selmi, Mounim A. El-Yacoubi, Bernadette Dorizzi |
IET Comput. Vis. | 2 |
| 2015 | Two-Stage Filtering Scheme for Sparse Representation Based Interest Point Matching for Person Re-identification
Mohamed Ibn Khedher, Mounim A. El-Yacoubi |
ACIVS | 2 |
| 2015 | Age and Gender Characterization Through a Two Layer Clustering of Online Handwriting
Gabriel Marzinotto, José C. Rosales 0002, Mounim A. El-Yacoubi, Sonia Garcia-Salicetti |
ACIVS | 3 |
| 2015 | Cluster-based data oriented hashingabstractMany multidimensional hashing schemes have been actively studied in recent years, providing efficient nearest neighbor search. Generally, we can distinguish several hashing families, such as learning based hashing, which provides better hash function selectivity by learning the dataset distribution. The spacial hashing family proposes a suitable partition of the multidimensional space, more adapted to data points distribution. In spite of the efficiency of multidimensional hashing techniques to solve the nearest neighbor search problem, these techniques suffer from scalabity issues. In this paper, we propose a novel hashing algorithm, named Cluster Based Data Oriented Hashing, that combines space hashing and learning based hashing techniques. The proposed approach applies first a clustering algorithm for structuring the multidimensional space into clusters. Then, in each cluster, a learning based hashing algorithm is applied by selecting an appropriate hash function that fits the data distribution. Experimental comparisons with standard Euclidean Locality Sensitive Hashing demonstrate the effectiveness of the proposed method for large datasets. Sanaa Chafik, Imane Daoudi, Mounim A. El-Yacoubi, Hamid El Ouardi |
DSAA | 3 |
| 2015 | Local Sparse Representation Based Interest Point Matching for Person Re-identification
Mohamed Ibn Khedher, Mounim A. El-Yacoubi |
ICONIP (3) | 2 |
| 2015 | Finger-Vein Quality Assessment by Representation Learning from Binary Images
Huafeng Qin, Mounim A. El-Yacoubi |
ICONIP (1) | 2 |
| 2013 | Multi-shot SURF-based person re-identification via sparse representationabstractWe present in this paper a multi-shot human reidentification system from video sequences based on SURF matching. Our contribution is about the matching step which is crucial. In this context, we propose a new method of SURF matching via sparse representation. Each SURF Interest Point in the test sequence is represented by a sparse representation of SURFs points in the reference dataset. For efficiency purposes, a dynamic dictionary is selected for each SURF from this dataset through KD-Tree Neighborhood search. Then a majority vote rule is applied to classify the test sequence. This approach is evaluated on two public datasets : PRID-2011 and CAVIAR4REID. The experimental results show that our approach compares favorably with and outperforms current state-of-the-art on the two datasets by 1% to 7%. Mohamed Ibn Khedher, Mounim A. El-Yacoubi, Bernadette Dorizzi |
AVSS | 2 |
| 2013 | A Combined SVM/HCRF Model for Activity Recognition based on STIPs Trajectories
Mouna Selmi, Mounim A. El-Yacoubi, Bernadette Dorizzi |
ICPRAM | 2 |
| 2012 | Human Action Recognition using Continuous HMMs and HOG/HOF Silhouette Representation
Mohamed Ibn Khedher, Mounim A. El-Yacoubi, Bernadette Dorizzi |
ICPRAM (2) | 2 |
| 2009 | Handwriting recognition research: Twenty years of achievement... and beyond
Mohamed Cheriet, Mounim A. El-Yacoubi, Hiromichi Fujisawa, Daniel P. Lopresti, Guy Lorette |
Pattern Recognit. | 2 |
| 2002 | A Statistical Approach for Phrase Location and Recognition within a Text Line: An Application to Street Name RecognitionabstractWe describe an approach to conjointly locate and recognize a street name within a street line. The system developed is based on a probabilistic framework that naturally integrates various knowledge sources to emit a final decision. At the handwriting signal level, hidden Markov models are extensively used to provide the needed matching scores. Several optimization techniques are employed to speed up the processing time. Experiments carried out on large data sets of street line images, automatically extracted from real French mail envelope images, show very promising results. Mounim A. El-Yacoubi, Michel Gilloux, Jean-Michel Bertille |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | Handwritten Month Word Recognition on Brazilian Bank ChecksabstractThis paper describes an off-line system under development to process unconstrained handwritten dates on Brazilian bank cheques in an omni-writer context. We show here some improvements on our previous work on isolated month word recognition using hidden Markov models (HMM). After preprocessing, a word image is explicitly segmented into characters or pseudo-characters and represented by two feature sequences of equal length, which are combined using HMM. The word models are generated from the concatenation of appropriate character models. In addition to the small date database, we also make use of the legal amount database to increase the frequency of characters in the training and the validation sets. Although this study deals with a limited lexicon, the many similarities among the word classes can affect the performance of the recognition. Experiments show an increase in the average recognition rate from 84% to 91%. Finally, we present our perspectives of future work. Marisa E. Morita, Robert Sabourin, Mounim A. El-Yacoubi, Flávio Bortolozzi, Ching Y. Suen |
ICDAR | 3 |
| 1999 | Influence of Word Length on Handwriting RecognitionabstractTwo strategies can be considered in handwriting recognition: phrase or word approaches. In this paper, we demonstrate the superiority of the phrase-based strategy, especially in city name recognition. The performance of an HMM-based off-line system using an analytic approach with explicit segmentation is evaluated on two databases: (i) city names in full, and (ii) city names in single words. A difference in performance is observed, principally caused by the dissimilarity of word lengths between the two databases. After generating other data sets and lexicons, experiments were performed yielding results which lead us to conclude that word length in the data set, as well as in lexicons, significantly influences recognition performance, and also that it is preferable to perform city name recognition based on the phrase approach rather than by word recognition. Frédéric Grandidier, Robert Sabourin, Mounim A. El-Yacoubi, Michel Gilloux, Ching Y. Suen |
ICDAR | 3 |
| 1999 | An HMM-Based Approach for Off-Line Unconstrained Handwritten Word Modeling and RecognitionabstractDescribes a hidden Markov model-based approach designed to recognize off-line unconstrained handwritten words for large vocabularies. After preprocessing, a word image is segmented into letters or pseudoletters and represented by two feature sequences of equal length, each consisting of an alternating sequence of shape-symbols and segmentation-symbols, which are both explicitly modeled. The word model is made up of the concatenation of appropriate letter models consisting of elementary HMMs and an HMM-based interpolation technique is used to optimally combine the two feature sets. Two rejection mechanisms are considered depending on whether or not the word image is guaranteed to belong to the lexicon. Experiments carried out on real-life data show that the proposed approach can be successfully used for handwritten word recognition. Mounim A. El-Yacoubi, Michel Gilloux, Robert Sabourin, Ching Y. Suen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1998 | Improved model architecture and training phase in an off-line HMM-based word recognition systemabstractDescribes the latest developments to enhance the performance of our HMM-based handwritten word recognition system. These methods only deal with the recognition phase and involve the improvement of the HMM architecture as well as the optimization of the training phase. Experiments carried out on real data show that the proposed approaches lead to significant improvements in the accuracy of the system. Mounim A. El-Yacoubi, Robert Sabourin, Michel Gilloux, Ching Y. Suen |
ICPR | 1 |
| 1995 | Conjoined location and recognition of street names within a postal address delivery lineabstractThis paper describes a global model designed to jointly detect and recognize a street name within a delivery line of an handwritten address block image. The model used is based on Hidden Markov Models (HMM). The lines are firstly preprocessed, then segmented and characterized by two types of features. We create a HMM for each street name by simply concatenating the corresponding letter models, elementary HMM learned on a large city name database. A street name may often be surrounded by other information on the left and on the right within the delivery line. These phenomena are roughly modelled by trigrams. The global model is then simply obtained by the concatenation of trigram models with the HMMs corresponding to street names. Mounim A. El-Yacoubi, Jean-Michel Bertille, Michel Gilloux |
ICDAR | 1 |