EDBT 2026 Demo / reviewers in the wild / expert
Wassim Bouachir
dblp:120/4430
· DBLP profile ↗
22ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0003-3896-7674ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From unaltered raw waveform to emotion: Synergizing convolutional and gated recurrent networks for holistic speech emotion analysis
Alaa Nfissi, Wassim Bouachir, Nizar Bouguila, Brian L. Mishara |
Appl. Intell. | 2 |
| 2025 | SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion RecognitionabstractIn the field of human-computer interaction and psychological assessment, speech emotion recognition (SER) plays an important role in deciphering emotional states from speech signals. Despite advancements, challenges persist due to system complexity, feature distinctiveness issues, and noise interference. This paper introduces a new end-to-end (E2E) deep learning multi-resolution framework for SER, addressing these limitations by extracting meaningful representations directly from raw waveform speech signals. By leveraging the properties of the fast discrete wavelet transform (FDWT), including the cascade algorithm, conjugate quadrature filter, and coefficient denoising, our approach introduces a learnable model for both wavelet bases and denoising through deep learning techniques. The framework incorporates an activation function for learnable asymmetric hard thresholding of wavelet coefficients. Our approach exploits the capabilities of wavelets for effective localization in both time and frequency domains. We then combine one-dimensional dilated convolutional neural networks (1D dilated CNN) with a spatial attention layer and bidirectional gated recurrent units (Bi-GRU) with a temporal attention layer to efficiently capture the nuanced spatial and temporal characteristics of emotional features. By handling variable-length speech without segmentation and eliminating the need for pre or post-processing, the proposed model outperformed state-of-the-art methods on IEMOCAP and EMO-DB datasets. Alaa Nfissi, Wassim Bouachir, Nizar Bouguila, Brian L. Mishara |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | ReL-SAR: Representation Learning for Skeleton Action Recognition with Convolutional Transformers and BYOLabstractTo extract robust and generalizable skeleton action recognition features, large amounts of well-curated data are typically required, which is a challenging task hindered by annotation and computation costs. Therefore, unsupervised representation learning is of prime importance to leverage unlabeled skeleton data. In this work, we investigate unsupervised representation learning for skeleton action recognition. For this purpose, we designed a lightweight convolutional transformer framework, named ReL-SAR, exploiting the complementarity of convolutional and attention layers for jointly modeling spatial and temporal cues in skeleton sequences. We also use a Selection-Permutation strategy for skeleton joints to ensure more infor-mative descriptions from skeletal data. Finally, we capitalize on Bootstrap Your Own Latent (BYOL) to learn robust representations from unlabeled skeleton sequence data. We achieved very competitive results on limited-size datasets: MCAD, IXMAS, JH-MDB, and NW-UCLA, showing the effectiveness of our proposed method against state-of-the-art methods in terms of both performance and computational efficiency. To ensure reproducibility and reusability, the source code including all implementation parameters is provided at https://github.com/SafwenNaimi. Safwen Naimi, Wassim Bouachir, Guillaume-Alexandre Bilodeau |
ICMLA | 2 |
| 2024 | Learnable Deep Wavelet Packet Transform for Speech Emotion Recognition in High-Risk Suicide CallsabstractIn human-computer interaction and psychological evaluation, speech emotion recognition (SER) is crucial for interpreting emotional states from spoken language. Although there have been advancements, challenges such as system complexity, issues with feature distinctiveness, and noise interference continue to persist. This paper presents a novel end-to-end (E2E) deep learning multi-resolution framework for SER, which tackles these limitations by deriving significant representations directly from raw speech waveform signals. By leveraging the properties of wavelet packet transform (WPT), our approach introduces a learnable model for both wavelet bases and denoising through deep learning techniques. Unlike discrete wavelet transform (DWT), WPT offers a more detailed analysis by decomposing both approximation and detail coefficients, providing a finer resolution in the time-frequency domain. This capability enhances feature extraction by capturing more nuanced signal characteristics across different frequency bands. The framework incorporates a learnable activation function for asymmetric hard thresholding of wavelet packet coefficients. Our approach exploits the capabilities of wavelet packets for effective localization in both time and frequency domains. We then combine one-dimensional dilated convolutional neural networks (1D dilated CNN) with a spatial attention layer and bidirectional gated recurrent units (Bi-GRU) with a temporal attention layer to efficiently capture emotional features' nuanced spatial and temporal characteristics. By handling variable-length speech without segmentation and eliminating the need for pre or post-processing, the proposed model outperforms state-of-the-art methods on our NSPL-CRISE dataset and on the public IEMOCAP dataset. The source code of this paper is shared on this repository: https://github.com/alaaNfissi/WPT-Deep-Learning-SER-for-Suicide-Monitoring. Alaa Nfissi, Wassim Bouachir, Nizar Bouguila, Brian L. Mishara |
ICMLA | 2 |
| 2024 | Unveiling hidden factors: explainable AI for feature boosting in speech emotion recognition
Alaa Nfissi, Wassim Bouachir, Nizar Bouguila, Brian L. Mishara |
Appl. Intell. | 2 |
| 2024 | 1D-convolutional transformer for Parkinson disease diagnosis from gait
Safwen Naimi, Wassim Bouachir, Guillaume-Alexandre Bilodeau |
Neural Comput. Appl. | 2 |
| 2023 | HCT: Hybrid Convnet-Transformer for Parkinson's Disease Detection and Severity Prediction from GaitabstractIn this paper, we propose a novel deep learning method based on a new Hybrid ConvNet-Transformer archi-tecture to detect and stage Parkinson's disease (PD) from gait data. We adopt a two-step approach by dividing the problem into two sub-problems. Our Hybrid ConvNet-Transformer model first distinguishes healthy versus parkinsonian patients. If the patient is parkinsonian, a multi-class Hybrid ConvNet-Transformer model determines the Hoehn and Yahr (H&Y) score to assess the PD severity stage. Our hybrid architecture exploits the strengths of both Convolutional Neural Networks (ConvNets) and Transformers to accurately detect PD and determine the severity stage. In particular, we take advantage of ConvNets to capture local patterns and correlations in the data, while we exploit Transformers for handling long-term dependencies in the input signal. We show that our hybrid method achieves superior performance when compared to other state-of-the-art methods, with a PD detection accuracy of 97% and a severity staging accuracy of 87%. Our source code is available at https://github.com/SafwenNaimi. Safwen Naimi, Wassim Bouachir, Guillaume-Alexandre Bilodeau |
ICMLA | 2 |
| 2023 | Automating Lichen Monitoring in Ecological Studies Using Instance Segmentation of Time-Lapse ImagesabstractLichens are symbiotic organisms composed of fungi, algae, and/or cyanobacteria that thrive in a variety of environments. They play important roles in carbon and nitrogen cycling, and contribute directly and indirectly to biodiversity. Ecologists typically monitor lichens by using them as indicators to assess air quality and habitat conditions. In particular, epiphytic lichens, which live on trees, are key markers of air quality and environmental health. A new method of monitoring epiphytic lichens involves using time-lapse cameras to gather images of lichen populations. These cameras are used by ecologists in Newfoundland and Labrador to subsequently analyze and manually segment the images to determine lichen thalli condition and change. These methods are time-consuming and susceptible to observer bias. In this work, we aim to automate the monitoring of lichens over extended periods and to estimate their biomass and condition to facilitate the task of ecologists. To accomplish this, our proposed framework uses semantic segmentation with an effective training approach to automate monitoring and biomass estimation of epiphytic lichens on time-lapse images. We show that our method has the potential to significantly improve the accuracy and efficiency of lichen population monitoring, making it a valuable tool for forest ecologists and environmental scientists to evaluate the impact of climate change on Canada's forests. To the best of our knowledge, this is the first time that such an approach has been used to assist ecologists in monitoring and analyzing epiphytic lichens. Safwen Naimi, Olfa Koubaa, Wassim Bouachir, Guillaume-Alexandre Bilodeau, Gregory Jeddore, Patricia Baines, David L. P. Correia, Andre Arsenault |
ICMLA | 3 |
| 2023 | Iterative Feature Boosting for Explainable Speech Emotion RecognitionabstractIn speech emotion recognition (SER), using predefined features without considering their practical importance may lead to high dimensional datasets, including redundant and irrelevant information. Consequently, high-dimensional learning often results in decreasing model accuracy while increasing computational complexity. Our work underlines the importance of carefully considering and analyzing features in order to build efficient SER systems. We present a new supervised SER method based on an efficient feature engineering approach. We pay particular attention to the explainability of results to evaluate feature relevance and refine feature sets. This is performed iteratively through feature evaluation loop, using Shapley values to boost feature selection and improve overall framework performance. Our approach allows thus to balance the benefits between model performance and transparency. The proposed method outperforms human-level performance (HLP) and state-of-the-art machine learning methods in emotion recognition on the TESS dataset. The source code of this paper is publicly available at Iterative-Feature-Boosting-for-Explainable-Speech-Emotion-Recoanition. Alaa Nfissi, Wassim Bouachir, Nizar Bouguila, Brian L. Mishara |
ICMLA | 2 |
| 2022 | Automatic counting of mounds on UAV images: combining instance segmentation and patch-level correctionabstractSite preparation by mounding is a commonly used silvicultural treatment that improves tree growth conditions by mechanically creating planting microsites called mounds. Following site preparation, the next critical step is to count the number of mounds, which provides forest managers with a precise estimate of the number of seedlings required for a given plantation block. Counting the number of mounds is generally conducted through manual field surveys by forestry workers, which is costly and prone to errors, especially for large areas. To address this issue, we present a novel framework exploiting advances in Unmanned Aerial Vehicle (UAV) imaging and computer vision to estimate the number of mounds on a planting block accurately. The proposed framework comprises two main components. First, we exploit a visual recognition method based on a deep learning algorithm for multiple object detection by segmentation. This enables a preliminary counting of visible mounds, as well as other frequently seen objects (e.g., trees, debris, accumulation of water), to be used to characterize the planting block. Second, since visual recognition could be limited by several perturbation factors (e.g., mound erosion, occlusion), we employ a machine learning estimation function to predict the final number of mounds based on the local block properties extracted in the first stage. We evaluate the proposed framework on a new UAV dataset representing numerous planting blocks with varying features. The proposed method outperformed manual counting methods in terms of relative counting precision, indicating that it has the potential to be advantageous and efficient under challenging situations. Majid Nikougoftar Nategh, Ahmed Zgaren, Wassim Bouachir, Nizar Bouguila |
ICMLA | 3 |
| 2022 | CNN-n-GRU: end-to-end speech emotion recognition from raw waveform signal using CNNs and gated recurrent unit networksabstractWe present CNN-n-GRU, a new end-to-end (E2E) architecture built of an n-layer convolutional neural network (CNN) followed sequentially by an n-layer Gated Recurrent Unit (GRU) for speech emotion recognition. CNNs and RNNs both exhibited promising outcomes when fed raw waveform voice inputs. This inspired our idea to combine them into a single model to maximise their potential. Instead of using handcrafted features or spectrograms, we train CNNs to recognise low-level speech representations from raw waveform, which allows the network to capture relevant narrow-band emotion characteristics. On the other hand, RNNs (GRUs in our case) can learn temporal characteristics, allowing the network to better capture the signal’s time-distributed features. Because a CNN can generate multiple levels of representation abstraction, we exploit early layers to extract high-level features, then to supply the appropriate input to subsequent RNN layers in order to aggregate long-term dependencies. By taking advantage of both CNNs and GRUs in a single model, the proposed architecture has important advantages over other models from the literature. The proposed model was evaluated using the TESS dataset and compared to state-of-the-art methods. Our experimental results demonstrate that the proposed model is more accurate than traditional classification approaches for speech emotion recognition. Alaa Nfissi, Wassim Bouachir, Nizar Bouguila, Brian L. Mishara |
ICMLA | 2 |
| 2022 | Transformers for 1D signals in Parkinson's disease detection from gaitabstractThis paper focuses on the detection of Parkinson’s disease based on the analysis of a patient’s gait. The growing popularity and success of Transformer networks in natural language processing and image recognition motivated us to develop a novel method for this problem based on an automatic features extraction via Transformers. The use of Transformers in 1D signal is not really widespread yet, but we show in this paper that they are effective in extracting relevant features from 1D signals. As Transformers require a lot of memory, we decoupled temporal and spatial information to make the model smaller. Our architecture used temporal Transformers, dimension reduction layers to reduce the dimension of the data, a spatial Transformer, two fully connected layers and an output layer for the final prediction. Our model outperforms the current state-of-the-art algorithm with 95.2% accuracy in distinguishing a Parkinsonian patient from a healthy one on the Physionet dataset. A key learning from this work is that Transformers allow for greater stability in results. The source code and pre-trained models are released in https://github.com/DucMinhDimitriNguyen1. Duc Minh Dimitri Nguyen, Mehdi Miah, Guillaume-Alexandre Bilodeau, Wassim Bouachir |
ICPR | 4 |
| 2021 | Multiple convolutional features in Siamese networks for object tracking
Zhenxi Li, Guillaume-Alexandre Bilodeau, Wassim Bouachir |
Mach. Vis. Appl. | 3 |
| 2021 | Efficient integration of generative topic models into discriminative classifiers using robust probabilistic kernels
Koffi Eddy Ihou, Nizar Bouguila, Wassim Bouachir |
Pattern Anal. Appl. | 3 |
| 2020 | MFST: Multi-Features Siamese TrackerabstractSiamese trackers have recently achieved interesting results due to their balance between accuracy and speed. This success is mainly due to the fact that deep similarity networks were specifically designed to address the image similarity problem. Therefore, they are inherently more appropriate than classical CNNs for the tracking task. However, Siamese trackers rely on the last convolutional layers for similarity analysis and target search, which restricts their performance. In this paper, we argue that using a single convolutional layer as feature representation is not the optimal choice within the deep similarity framework, as multiple convolutional layers provide several abstraction levels in characterizing an object. Starting from this motivation, we present the Multi-Features Siamese Tracker (MFST), a novel tracking algorithm exploiting several hierarchical feature maps for robust deep similarity tracking. MFST proceeds by fusing hierarchical features to ensure a richer and more efficient representation. Moreover, we handle appearance variation by calibrating deep features extracted from two different CNN models. Based on this advanced feature representation, our algorithm achieves high tracking accuracy, while outperforming several state-of-the-art trackers, including standard Siamese trackers. Zhenxi Li, Guillaume-Alexandre Bilodeau, Wassim Bouachir |
ICPR | 3 |
| 2020 | Deep 1D-Convnet for accurate Parkinson disease detection and severity prediction from gait
Imanne El Maachi, Guillaume-Alexandre Bilodeau, Wassim Bouachir |
Expert Syst. Appl. | 3 |
| 2018 | Intelligent video surveillance for real-time detection of suicide attempts
Wassim Bouachir, Rafik Gouiaa, Bo Li 0118, Rita Noumeir |
Pattern Recognit. Lett. | 1 |
| 2015 | Reproducible evaluation of Pan-Tilt-Zoom trackingabstractTracking with a Pan-Tilt-Zoom (PTZ) camera has been a research topic in computer vision for many years. However, it is difficult to assess the progress that has been made because there is no standard evaluation methodology. The difficulty in evaluating PTZ tracking algorithms arises from their dynamic nature. In contrast to other forms of tracking, PTZ tracking involves both locating the target in the image and controlling the motors of the camera to aim it so that the target stays in its field of view. This type of tracking can only be performed online. In this paper, we propose a new evaluation framework based on a virtual PTZ camera. With this framework, tracking scenarios do not change for each experiment and we are able to replicate the main principles of online PTZ camera control and behavior including camera positioning delays, tracker processing delays, and numerical zoom. We tested our evaluation framework with the Camshift tracker to show its viability and to establish baseline results. Gengjie Chen, Pierre-Luc St-Charles, Wassim Bouachir, Guillaume-Alexandre Bilodeau, Robert Bergevin |
ICIP | 3 |
| 2015 | Part-Based Tracking via Salient Collaborating FeaturesabstractWe present a novel part-based method for model-free tracking. In our model, key points are considered as elementary predictors, collaborating to localize the target. In order to differentiate reliable features from outliers and bad predictors, we define the notion of feature saliency including three factors: the persistence, the spatial consistency, and the predictive power of local features. Saliency information is learned during tracking to be used in several algorithmic steps: local predictions, global localization, feature removal, etc. By exploiting saliency information and key point structural properties, the proposed algorithm is able to track accurately generic objects, facing several difficulties such as occlusions, presence of distractors, and abrupt motion. The proposed tracker demonstrated a high robustness on challenging public datasets, outperforming significantly five recent state-of-the-art trackers. Wassim Bouachir, Guillaume-Alexandre Bilodeau |
WACV | 1 |
| 2015 | Collaborative part-based tracking using salient local predictors
Wassim Bouachir, Guillaume-Alexandre Bilodeau |
Comput. Vis. Image Underst. | 1 |
| 2015 | Exploiting structural constraints for visual object tracking
Wassim Bouachir, Guillaume-Alexandre Bilodeau |
Image Vis. Comput. | 1 |
| 2014 | Structure-aware keypoint tracking for partial occlusion handlingabstractThis paper introduces a novel keypoint-based method for visual object tracking. To represent the target, we use a new model combining color distribution with keypoints. The appearance model also incorporates the spatial layout of the keypoints, encoding the object structure learned during tracking. With this multi-feature appearance model, our Structure-Aware Tracker (SAT) estimates accurately the target location using three main steps. First, the search space is reduced to the most likely image regions with a probabilistic approach. Second, the target location is estimated in the reduced search space using deterministic keypoint matching. Finally, the location prediction is corrected by exploiting the keypoint structural model with a voting-based method. By applying our SAT on several tracking problems, we show that location correction based on structural constraints is a key technique to improve prediction in moderately crowded scenes, even if only a small part of the target is visible. We also conduct comparison with a number of state-of-the-art trackers and demonstrate the competitiveness of the proposed method. Wassim Bouachir, Guillaume-Alexandre Bilodeau |
WACV | 1 |