VLDB 2026 Research / reviewers in the wild / expert
Guillaume-Alexandre Bilodeau
dblp:90/2054
· DBLP profile ↗
58ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0003-3227-5060ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Continuous conditional video synthesis by neural processesabstractDifferent conditional video synthesis tasks, such as frame interpolation and future frame prediction, are typically addressed individually by task-specific models, despite their shared underlying characteristics. Additionally, most conditional video synthesis models are limited to discrete frame generation at specific integer time steps. This paper presents a unified model that tackles both challenges simultaneously. We demonstrate that conditional video synthesis can be formulated as a neural process, where input spatio-temporal coordinates are mapped to target pixel values by conditioning on context spatio-temporal coordinates and pixel values. Our approach leverages a Transformer-based non-autoregressive conditional video synthesis model that takes the implicit neural representation of coordinates and context pixel features as input. Our task-specific models outperform previous methods for future frame prediction and frame interpolation across multiple datasets. Importantly, our model enables temporal continuous video synthesis at arbitrary high frame rates, outperforming the previous state-of-the-art. The source code and video demos for our model are available at https://npvp.github.io . Xi Ye 0005, Guillaume-Alexandre Bilodeau |
Comput. Vis. Image Underst. | 2 |
| 2025 | Learning data association for multi-object tracking using only coordinatesabstractWe propose a novel Transformer-based module to address the data association problem for multi-object tracking. From detections obtained by a pretrained detector, this module uses only coordinates from bounding boxes to estimate an affinity score between pairs of tracks extracted from two distinct temporal windows. This module, named TWiX, is trained on sets of tracks with the objective of discriminating pairs of tracks coming from the same object from those which are not. Our module does not use the intersection over union measure, nor does it requires any motion priors or any camera motion compensation technique. By inserting TWiX within an online cascade matching pipeline, our tracker C-TWiX achieves state-of-the-art performance on the DanceTrack and KITTIMOT datasets, and gets competitive results on the MOT17 dataset. The code will be made available upon publication on the website https://mehdimiah.com/twix . • Our Transformer-based model, TWiX, can learn to associate objects using only coordinates. • We show that motion priors or intersection-over-union measure are not required for tracking. Using pairs of tracks is sufficient. • Tracking with TWiX gives competitive or state-of-the-results on several datasets. Mehdi Miah, Guillaume-Alexandre Bilodeau, Nicolas Saunier |
Pattern Recognit. | 2 |
| 2024 | STDiff: Spatio-Temporal Diffusion for Continuous Stochastic Video PredictionabstractPredicting future frames of a video is challenging because it is difficult to learn the uncertainty of the underlying factors influencing their contents. In this paper, we propose a novel video prediction model, which has infinite-dimensional latent variables over the spatio-temporal domain. Specifically, we first decompose the video motion and content information, then take a neural stochastic differential equation to predict the temporal motion information, and finally, an image diffusion model autoregressively generates the video frame by conditioning on the predicted motion feature and the previous frame. The better expressiveness and stronger stochasticity learning capability of our model lead to state-of-the-art video prediction performances. As well, our model is able to achieve temporal continuous prediction, i.e., predicting in an unsupervised way the future video frames with an arbitrarily high frame rate. Our code is available at https://github.com/XiYe20/STDiffProject. Xi Ye 0005, Guillaume-Alexandre Bilodeau |
AAAI | 2 |
| 2024 | ReL-SAR: Representation Learning for Skeleton Action Recognition with Convolutional Transformers and BYOLabstractTo extract robust and generalizable skeleton action recognition features, large amounts of well-curated data are typically required, which is a challenging task hindered by annotation and computation costs. Therefore, unsupervised representation learning is of prime importance to leverage unlabeled skeleton data. In this work, we investigate unsupervised representation learning for skeleton action recognition. For this purpose, we designed a lightweight convolutional transformer framework, named ReL-SAR, exploiting the complementarity of convolutional and attention layers for jointly modeling spatial and temporal cues in skeleton sequences. We also use a Selection-Permutation strategy for skeleton joints to ensure more infor-mative descriptions from skeletal data. Finally, we capitalize on Bootstrap Your Own Latent (BYOL) to learn robust representations from unlabeled skeleton sequence data. We achieved very competitive results on limited-size datasets: MCAD, IXMAS, JH-MDB, and NW-UCLA, showing the effectiveness of our proposed method against state-of-the-art methods in terms of both performance and computational efficiency. To ensure reproducibility and reusability, the source code including all implementation parameters is provided at https://github.com/SafwenNaimi. Safwen Naimi, Wassim Bouachir, Guillaume-Alexandre Bilodeau |
ICMLA | 3 |
| 2024 | 1D-convolutional transformer for Parkinson disease diagnosis from gait
Safwen Naimi, Wassim Bouachir, Guillaume-Alexandre Bilodeau |
Neural Comput. Appl. | 3 |
| 2023 | HCT: Hybrid Convnet-Transformer for Parkinson's Disease Detection and Severity Prediction from GaitabstractIn this paper, we propose a novel deep learning method based on a new Hybrid ConvNet-Transformer archi-tecture to detect and stage Parkinson's disease (PD) from gait data. We adopt a two-step approach by dividing the problem into two sub-problems. Our Hybrid ConvNet-Transformer model first distinguishes healthy versus parkinsonian patients. If the patient is parkinsonian, a multi-class Hybrid ConvNet-Transformer model determines the Hoehn and Yahr (H&Y) score to assess the PD severity stage. Our hybrid architecture exploits the strengths of both Convolutional Neural Networks (ConvNets) and Transformers to accurately detect PD and determine the severity stage. In particular, we take advantage of ConvNets to capture local patterns and correlations in the data, while we exploit Transformers for handling long-term dependencies in the input signal. We show that our hybrid method achieves superior performance when compared to other state-of-the-art methods, with a PD detection accuracy of 97% and a severity staging accuracy of 87%. Our source code is available at https://github.com/SafwenNaimi. Safwen Naimi, Wassim Bouachir, Guillaume-Alexandre Bilodeau |
ICMLA | 3 |
| 2023 | Automating Lichen Monitoring in Ecological Studies Using Instance Segmentation of Time-Lapse ImagesabstractLichens are symbiotic organisms composed of fungi, algae, and/or cyanobacteria that thrive in a variety of environments. They play important roles in carbon and nitrogen cycling, and contribute directly and indirectly to biodiversity. Ecologists typically monitor lichens by using them as indicators to assess air quality and habitat conditions. In particular, epiphytic lichens, which live on trees, are key markers of air quality and environmental health. A new method of monitoring epiphytic lichens involves using time-lapse cameras to gather images of lichen populations. These cameras are used by ecologists in Newfoundland and Labrador to subsequently analyze and manually segment the images to determine lichen thalli condition and change. These methods are time-consuming and susceptible to observer bias. In this work, we aim to automate the monitoring of lichens over extended periods and to estimate their biomass and condition to facilitate the task of ecologists. To accomplish this, our proposed framework uses semantic segmentation with an effective training approach to automate monitoring and biomass estimation of epiphytic lichens on time-lapse images. We show that our method has the potential to significantly improve the accuracy and efficiency of lichen population monitoring, making it a valuable tool for forest ecologists and environmental scientists to evaluate the impact of climate change on Canada's forests. To the best of our knowledge, this is the first time that such an approach has been used to assist ecologists in monitoring and analyzing epiphytic lichens. Safwen Naimi, Olfa Koubaa, Wassim Bouachir, Guillaume-Alexandre Bilodeau, Gregory Jeddore, Patricia Baines, David L. P. Correia, Andre Arsenault |
ICMLA | 4 |
| 2023 | Video prediction by efficient transformers
Xi Ye 0005, Guillaume-Alexandre Bilodeau |
Image Vis. Comput. | 2 |
| 2022 | Transformers for 1D signals in Parkinson's disease detection from gaitabstractThis paper focuses on the detection of Parkinson’s disease based on the analysis of a patient’s gait. The growing popularity and success of Transformer networks in natural language processing and image recognition motivated us to develop a novel method for this problem based on an automatic features extraction via Transformers. The use of Transformers in 1D signal is not really widespread yet, but we show in this paper that they are effective in extracting relevant features from 1D signals. As Transformers require a lot of memory, we decoupled temporal and spatial information to make the model smaller. Our architecture used temporal Transformers, dimension reduction layers to reduce the dimension of the data, a spatial Transformer, two fully connected layers and an output layer for the final prediction. Our model outperforms the current state-of-the-art algorithm with 95.2% accuracy in distinguishing a Parkinsonian patient from a healthy one on the Physionet dataset. A key learning from this work is that Transformers allow for greater stability in results. The source code and pre-trained models are released in https://github.com/DucMinhDimitriNguyen1. Duc Minh Dimitri Nguyen, Mehdi Miah, Guillaume-Alexandre Bilodeau, Wassim Bouachir |
ICPR | 3 |
| 2022 | VPTR: Efficient Transformers for Video PredictionabstractIn this paper, we propose a new Transformer block for video future frames prediction based on an efficient local spatial-temporal separation attention mechanism. Based on this new Transformer block, a fully autoregressive video future frames prediction Transformer is proposed. In addition, a non-autoregressive video prediction Transformer is also proposed to increase the inference speed and reduce the accumulated inference errors of its autoregressive counterpart. In order to avoid the prediction of very similar future frames, a contrastive feature loss is applied to maximize the mutual information between predicted and ground-truth future frame features. This work is the first that makes a formal comparison of the two types of attention-based video future frames prediction models over different scenarios. The proposed models reach a performance competitive with more complex state-of-the-art models. The source code is available at https://github.com/XiYe20/VPTR. Xi Ye 0005, Guillaume-Alexandre Bilodeau |
ICPR | 2 |
| 2021 | Multiple convolutional features in Siamese networks for object tracking
Zhenxi Li, Guillaume-Alexandre Bilodeau, Wassim Bouachir |
Mach. Vis. Appl. | 2 |
| 2021 | FFAVOD: Feature fusion architecture for video object detection
Hughes Perreault, Guillaume-Alexandre Bilodeau, Nicolas Saunier, Maguelonne Héritier |
Pattern Recognit. Lett. | 2 |
| 2020 | Domain Siamese CNNs for Sparse Multispectral Disparity EstimationabstractMultispectral disparity estimation is a difficult task for many reasons: it has all the same challenges as traditional visible-visible disparity estimation (occlusions, repetitive patterns, textureless surfaces), in addition of having very few common visual information between images (e.g. colour information vs. thermal information). In this paper, we propose a new CNN architecture able to do disparity estimation between images from different spectra, namely thermal and visible in our case. Our proposed model takes two patches as input and proceeds to do domain feature extraction for each of them. Features from both domains are then merged with two fusion operations, namely correlation and concatenation. These merged vectors are then forwarded to their respective classification heads, which are responsible for classifying the inputs as being same or not. Using two merging operations gives more robustness to our feature extraction process, which leads to more precise disparity estimation. Our method was tested using the publicly available LITIV 2014 and LITIV 2018 datasets, and showed best results when compared to other state-of-the-art methods. David-Alexandre Beaupré, Guillaume-Alexandre Bilodeau |
ICPR | 2 |
| 2020 | A Grid-based Representation for Human Action RecognitionabstractHuman action recognition (HAR) in videos is a fundamental research topic in computer vision. It consists mainly in understanding actions performed by humans based on a sequence of visual observations. In recent years, HAR have witnessed significant progress, especially with the emergence of deep learning models. However, most of existing approaches for action recognition rely on information that is not always relevant for this task, and are limited in the way they fuse the temporal information. In this paper, we propose a novel method for human action recognition that encodes efficiently the most discriminative appearance information of an action with explicit attention on representative pose features, into a new compact grid representation. Our GRAR (Grid-based Representation for Action Recognition) method is tested on several benchmark datasets demonstrating that our model can accurately recognize human actions, despite intra-class appearance variations and occlusion challenges. Soufiane Lamghari, Guillaume-Alexandre Bilodeau, Nicolas Saunier |
ICPR | 2 |
| 2020 | MFST: Multi-Features Siamese TrackerabstractSiamese trackers have recently achieved interesting results due to their balance between accuracy and speed. This success is mainly due to the fact that deep similarity networks were specifically designed to address the image similarity problem. Therefore, they are inherently more appropriate than classical CNNs for the tracking task. However, Siamese trackers rely on the last convolutional layers for similarity analysis and target search, which restricts their performance. In this paper, we argue that using a single convolutional layer as feature representation is not the optimal choice within the deep similarity framework, as multiple convolutional layers provide several abstraction levels in characterizing an object. Starting from this motivation, we present the Multi-Features Siamese Tracker (MFST), a novel tracking algorithm exploiting several hierarchical feature maps for robust deep similarity tracking. MFST proceeds by fusing hierarchical features to ensure a richer and more efficient representation. Moreover, we handle appearance variation by calibrating deep features extracted from two different CNN models. Based on this advanced feature representation, our algorithm achieves high tracking accuracy, while outperforming several state-of-the-art trackers, including standard Siamese trackers. Zhenxi Li, Guillaume-Alexandre Bilodeau, Wassim Bouachir |
ICPR | 2 |
| 2020 | An Empirical Analysis of Visual Features for Multiple Object Tracking in Urban ScenesabstractThis paper addresses the problem of selecting appearance features for multiple object tracking (MOT) in urban scenes. Over the years, a large number of features has been used for MOT. However, it is not clear whether some of them are better than others. Commonly used features are color histograms, histograms of oriented gradients, deep features from convolutional neural networks and re-identification (ReID) features. In this study, we assess how good these features are at discriminating objects enclosed by a bounding box in urban scene tracking scenarios. Several affinity measures, namely the L1, L2and the Bhattacharyya distances, Rank-1 counts and the cosine similarity, are also assessed for their impact on the discriminative power of the features. Results on several datasets show that features from ReID networks are the best for discriminating instances from one another regardless of the quality of the detector. If a ReID model is not available, color histograms may be selected if the detector has a good recall and there are few occlusions; otherwise, deep features are more robust to detectors with lower recall. Mehdi Miah, Justine Pepin, Nicolas Saunier, Guillaume-Alexandre Bilodeau |
ICPR | 4 |
| 2020 | Deep 1D-Convnet for accurate Parkinson disease detection and severity prediction from gait
Imanne El Maachi, Guillaume-Alexandre Bilodeau, Wassim Bouachir |
Expert Syst. Appl. | 2 |
| 2020 | Robust face tracking using multiple appearance models and graph relational learning
Tanushri Chakravorty, Guillaume-Alexandre Bilodeau, Eric Granger |
Mach. Vis. Appl. | 2 |
| 2019 | Online Mutual Foreground Segmentation for Multispectral Stereo Videos
Pierre-Luc St-Charles, Guillaume-Alexandre Bilodeau, Robert Bergevin |
Int. J. Comput. Vis. | 2 |
| 2019 | From superpixel to human shape modelling for carried object detection
Farnoosh Ghadiri, Robert Bergevin, Guillaume-Alexandre Bilodeau |
Pattern Recognit. | 3 |
| 2019 | Domain-Specific Face Synthesis for Video Face Recognition From a Single Sample Per PersonabstractIn video surveillance, face recognition (FR) systems are employed to detect individuals of interest appearing over a distributed network of cameras. The performance of still-to-video FR systems can decline significantly because faces captured in unconstrained operational domain (OD) over multiple video cameras have a different underlying data distribution compared to faces captured under controlled conditions in the enrollment domain with a still camera. This is particularly true when individuals are enrolled to the system using a single reference still. To improve the robustness of these systems, it is possible to augment the reference set by generating synthetic faces based on the original still. However, without the knowledge of the OD, many synthetic images must be generated to account for all possible capture conditions. FR systems may, therefore, require complex implementations and yield lower accuracy when training on many less relevant images. This paper introduces an algorithm for domain-specific face synthesis (DSFS) that exploits the representative intra-class variation information available from the OD. Prior to operation (during camera calibration), a compact set of faces from unknown persons appearing in the OD is selected through affinity propagation clustering in the captured condition space (defined by pose and illumination estimation). The domain-specific variations of these face images are then projected onto the reference still of each individual by integrating an image-based face relighting technique inside the 3-D reconstruction framework. A compact set of synthetic faces is generated that resemble individuals of interest under the capture conditions relevant to the OD. In a particular implementation based on sparse representation classification, the synthetic faces generated with the DSFS are employed to form a cross-domain dictionary that accounts for structured sparsity, where the dictionary blocks combine the original and synthetic faces of each individual. Experimental results obtained with videos from the Chokepoint and COX-S2V data sets reveal that augmenting the reference gallery set of still-to-video FR systems using the proposed DSFS approach can provide a significantly higher level of accuracy compared with the state-of-the-art approaches, with only a moderate increase in its computational complexity. Fania Mokhayeri, Eric Granger, Guillaume-Alexandre Bilodeau |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | Tracking using Numerous Anchor Points
Tanushri Chakravorty, Guillaume-Alexandre Bilodeau, Eric Granger |
Mach. Vis. Appl. | 2 |
| 2018 | SPiKeS: Superpixel-Keypoints structure for robust visual tracking
François-Xavier Derue, Guillaume-Alexandre Bilodeau, Robert Bergevin |
Mach. Vis. Appl. | 2 |
| 2017 | Spatio-Temporal Consistency to Detect and Segment Carried Objects
Farnoosh Ghadiri, Robert Bergevin, Guillaume-Alexandre Bilodeau |
BMVC | 3 |
| 2017 | Dynamic Selection of Exemplar-SVMs for Watch-list Screening through Domain Adaptation
Saman Bashbaghi, Eric Granger, Robert Sabourin, Guillaume-Alexandre Bilodeau |
ICPRAM | 4 |
| 2017 | Robust watch-list screening using dynamic ensembles of SVMs based on multiple face representations
Saman Bashbaghi, Eric Granger, Robert Sabourin, Guillaume-Alexandre Bilodeau |
Mach. Vis. Appl. | 4 |
| 2017 | Dynamic ensembles of exemplar-SVMs for still-to-video face recognition
Saman Bashbaghi, Eric Granger, Robert Sabourin, Guillaume-Alexandre Bilodeau |
Pattern Recognit. | 4 |
| 2016 | Carried Object Detection Based on an Ensemble of Contour Exemplars
Farnoosh Ghadiri, Robert Bergevin, Guillaume-Alexandre Bilodeau |
ECCV (7) | 3 |
| 2016 | Online multi-object tracking by detection based on generative appearance models
Dorra Riahi, Guillaume-Alexandre Bilodeau |
Comput. Vis. Image Underst. | 2 |
| 2016 | Universal Background Subtraction Using Word Consensus ModelsabstractBackground subtraction is often used as the first step in video analysis and smart surveillance applications. However, the issue of inconsistent performance across different scenarios due to a lack of flexibility remains a serious concern. To address this, we propose a novel non-parametric, pixel-level background modeling approach based on word dictionaries that draws from traditional codebooks and sample consensus approaches. In this new approach, the importance of each background sample (or word) is evaluated online based on their recurrence among all local observations. This helps build smaller pixel models that are better suited for long-term foreground detection. Combining these models with a frame-level dictionary and local feedback mechanisms leads us to our proposed background subtraction method, coined “PAWCS.” Experiments on the 2012 and 2014 versions of the ChangeDetection.net data set show that PAWCS outperforms 26 previously tested and published methods in terms of overall F-Measure as well as in most categories taken individually. Our results can be reproduced with a C++ implementation available online. Pierre-Luc St-Charles, Guillaume-Alexandre Bilodeau, Robert Bergevin |
IEEE Trans. Image Process. | 2 |
| 2016 | Tracking All Road Users at Multimodal Urban Traffic IntersectionsabstractBecause of the large variability of road user appearance in an urban setting, it is very challenging to track all of them with the purpose of obtaining precise and reliable trajectories. However, obtaining the trajectories of the various road users is very useful for many transportation applications. It is particularly essential for any task that requires higher level behavior interpretation, including new safety diagnosis methods that rely on the observation of road user interactions without a collision and therefore do not require waiting for collisions to happen. In this paper, we propose a tracking method that has been specifically designed to track the various road users that may be encountered in an urban environment. Since road users have very diverse shapes and appearances, our proposed method starts from background subtraction to extract the potential a priori unknown road users. Each of these road users is then tracked using a collection of keypoints inside the detected foreground regions, which allows the interpolation of object locations even during object merges or occlusions. A finite state machine handles fragmentation, splitting, and merging of the road users to correct and improve the resulting object trajectories. The proposed tracker was tested on several urban intersection videos and is shown to outperform an existing reference tracker used in transportation research. Jean-Philippe Jodoin, Guillaume-Alexandre Bilodeau, Nicolas Saunier |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2015 | Ensembles of exemplar-SVMs for video face recognition from a single sample per personabstractRecognizing the face of target individuals in a watch-list is among the most challenging applications in video surveillance, especially when enrollment is based on one reference still facial image. Besides the limited representativeness of facial models used for matching, the appearance of faces captured in videos varies due to changes in illumination, pose, scales, etc., and to camera inter-operability. A multi-classifier system is proposed in this paper for robust still-to-video face recognition (FR) based on multiple diverse face representations. An individual-specific ensemble of exemplar-SVMs (e-SVMs) classifiers is assigned to each target person, where each classifier is trained using a high-quality reference face still versus many lower-quality faces of non-target individuals captured in videos. Diverse face representations are generated from different patches isolated in facial images and face descriptors that are robust to various nuisance factors (e.g., illumination and pose) commonly encountered in surveillance environments. Discriminant feature subsets, training samples, and ensemble fusion functions are selected using faces of non-target individuals captured in videos of the scene. Experiments on videos from the Chokepoint dataset reveal that the proposed ensemble of e-SVMs outperforms state-of-the-art FR systems specialized for the single sample per person problem. Saman Bashbaghi, Eric Granger, Robert Sabourin, Guillaume-Alexandre Bilodeau |
AVSS | 4 |
| 2015 | Contextual object tracker with structure encodingabstractMotivated by the problem of object tracking in video sequences, this paper presents a new Contextual Object Tracker with Structural Encoding (CTSE). The novelty in our tracking approach lies in the application of contextual and structural information (that is specific to a target object) into a model-free tracker. This is first achieved by including features from a complementary region having correlated motion with the target object. Second, a local structure that represents a spatial constraint between features within the target object are included. SIFT keypoints are used as features to encode both these information. The tracking is done in three steps. Firstly, keypoints are detected and described to encode object structure. Secondly, they are matched in every frame. Finally, each matched keypoint votes for the target object location locally in a voting matrix by using the encoded object structure. The voting method gives more priority to the keypoints that have been matched more often and are closest to the target's center than the rest. The proposed tracker is competitive with state-of-the art trackers while being significantly faster. It ranks as first or second most accurate tracker in experiments with standard datasets. Tanushri Chakravorty, Guillaume-Alexandre Bilodeau, Eric Granger |
ICIP | 2 |
| 2015 | Reproducible evaluation of Pan-Tilt-Zoom trackingabstractTracking with a Pan-Tilt-Zoom (PTZ) camera has been a research topic in computer vision for many years. However, it is difficult to assess the progress that has been made because there is no standard evaluation methodology. The difficulty in evaluating PTZ tracking algorithms arises from their dynamic nature. In contrast to other forms of tracking, PTZ tracking involves both locating the target in the image and controlling the motors of the camera to aim it so that the target stays in its field of view. This type of tracking can only be performed online. In this paper, we propose a new evaluation framework based on a virtual PTZ camera. With this framework, tracking scenarios do not change for each experiment and we are able to replicate the main principles of online PTZ camera control and behavior including camera positioning delays, tracker processing delays, and numerical zoom. We tested our evaluation framework with the Camshift tracker to show its viability and to establish baseline results. Gengjie Chen, Pierre-Luc St-Charles, Wassim Bouachir, Guillaume-Alexandre Bilodeau, Robert Bergevin |
ICIP | 4 |
| 2015 | Synthetic face generation under various operational conditions in video surveillanceabstractIn still-to-video face recognition (FR), the faces captured with surveillance cameras are matched against reference stills of target individuals enrolled to the system. FR is a challenging problem in video surveillance due to uncontrolled capture conditions (variations in pose, expression, illumination, blur, scale, etc.), and the limited number of reference stills to model target individuals. This paper introduces a new approach to generate multiple synthetic face images per reference still based on camera-specific capture conditions to deal with illumination variations. For each reference still, a diverse set of faces from non-target individuals appearing in the camera viewpoint are selected based on luminance and contrast distortion. These face images are then decomposed into detail layer and large scale layer using an edge-preserving image decomposition to obtain their illumination dependent component. Finally, the large scale layers of these images are morphed with each reference still image to generate multiple synthetic reference stills that incorporate illumination and contrast conditions. Experimental results obtained with the ChokePoint dataset reveal that these synthetic faces produce an enhanced face model. As the number of synthetic faces grows, the proposed approach provides a higher level of accuracy and robustness across a range of capture conditions. Fania Mokhayeri, Eric Granger, Guillaume-Alexandre Bilodeau |
ICIP | 3 |
| 2015 | Multiple object tracking based on sparse generative appearance modelingabstractThis paper addresses multiple object tracking which still remains a challenging problem because of factors like frequent occlusions, unknown number of targets and similarity in objects' appearance. We propose a novel approach for multiple object tracking using a multiple feature framework. The main focus of the proposed method is to build a robust appearance model. The appearance model of an object is built using a color model, a sparse appearance model, a motion model and spatial information. We validated the proposed algorithm on four publicly available videos with comparisons with state-of-the-art approaches. We demonstrate that our algorithm achieves competitive results. Dorra Riahi, Guillaume-Alexandre Bilodeau |
ICIP | 2 |
| 2015 | Part-Based Tracking via Salient Collaborating FeaturesabstractWe present a novel part-based method for model-free tracking. In our model, key points are considered as elementary predictors, collaborating to localize the target. In order to differentiate reliable features from outliers and bad predictors, we define the notion of feature saliency including three factors: the persistence, the spatial consistency, and the predictive power of local features. Saliency information is learned during tracking to be used in several algorithmic steps: local predictions, global localization, feature removal, etc. By exploiting saliency information and key point structural properties, the proposed algorithm is able to track accurately generic objects, facing several difficulties such as occlusions, presence of distractors, and abrupt motion. The proposed tracker demonstrated a high robustness on challenging public datasets, outperforming significantly five recent state-of-the-art trackers. Wassim Bouachir, Guillaume-Alexandre Bilodeau |
WACV | 2 |
| 2015 | A Self-Adjusting Approach to Change Detection Based on Background Word ConsensusabstractAlthough there has long been interest in foreground background segmentation based on change detection for video surveillance applications, the issue of inconsistent performance across different scenarios remains a serious concern. To address this, we propose a new type of word based approach that regulates its own internal parameters using feedback mechanisms to withstand difficult conditions while keeping sensitivity intact in regular situations. Coined "PAWCS", this method's key advantages lie in its highly persistent and robust dictionary model based on color and local binary features as well as its ability to automatically adjust pixel-level segmentation behavior. Experiments using the 2012 Change Detection.net dataset show that it outranks numerous recently proposed solutions in terms of overall performance as well as in each category. A complete C++ implementation based on OpenCV is available online. Pierre-Luc St-Charles, Guillaume-Alexandre Bilodeau, Robert Bergevin |
WACV | 2 |
| 2015 | Collaborative part-based tracking using salient local predictors
Wassim Bouachir, Guillaume-Alexandre Bilodeau |
Comput. Vis. Image Underst. | 2 |
| 2015 | Exploiting structural constraints for visual object tracking
Wassim Bouachir, Guillaume-Alexandre Bilodeau |
Image Vis. Comput. | 2 |
| 2015 | SuBSENSE: A Universal Change Detection Method With Local Adaptive SensitivityabstractForeground/background segmentation via change detection in video sequences is often used as a stepping stone in high-level analytics and applications. Despite the wide variety of methods that have been proposed for this problem, none has been able to fully address the complex nature of dynamic scenes in real surveillance tasks. In this paper, we present a universal pixel-level segmentation method that relies on spatiotemporal binary features as well as color information to detect changes. This allows camouflaged foreground objects to be detected more easily while most illumination variations are ignored. Besides, instead of using manually set, frame-wide constants to dictate model sensitivity and adaptation speed, we use pixel-level feedback loops to dynamically adjust our method's internal parameters without user intervention. These adjustments are based on the continuous monitoring of model fidelity and local segmentation noise levels. This new approach enables us to outperform all 32 previously tested state-of-the-art methods on the 2012 and 2014 versions of the ChangeDetection.net dataset in terms of overall F-Measure. The use of local binary image descriptors for pixel-level modeling also facilitates high-speed parallel implementations: our own version, which used no low-level or architecture-specific instruction, reached real-time processing speed on a midlevel desktop CPU. A complete C++ implementation based on OpenCV is available online. Pierre-Luc St-Charles, Guillaume-Alexandre Bilodeau, Robert Bergevin |
IEEE Trans. Image Process. | 2 |
| 2014 | Watch-List Screening Using Ensembles Based on Multiple Face RepresentationsabstractStill-to-video face recognition (FR) is an important function in watch list screening, where faces captured over a network of video surveillance cameras are matched against reference stills of target individuals. Recognizing faces in a watch list is a challenging problem in semi -- and unconstrained surveillance environments due to the lack of control over capture and operational conditions, and to the limited number of reference stills. This paper provides a performance baseline and guidelines for ensemble-based systems using a single high-quality reference still per individual, as found in many watch list screening applications. In particular, modular systems are considered, where an ensemble of template matchers based on multiple face representations is assigned to each individual of interest. During enrollment, multiple feature extraction (FE) techniques are applied to patches isolated in the reference still to generate diverse face-part representations that are robust to various nuisance factors (e.g., illumination and pose) encountered in video surveillance. The selection of relevant feature subsets, decision thresholds, and fusion functions of ensembles are achieved using faces of non-target individuals selected from reference videos (forming a universal background model). During operations, a face tracker gradually regroups faces captured from different people appearing in a scene, while each user-specific ensemble generates a decision per face capture. This leads to robust spatio-temporal FR when accumulated ensemble predictions surpass a detection threshold. Simulation results obtained with the Chokepoint video dataset show a significant improvement to accuracy, (1) when performing score-level fusion of matchers, where patches-based and FE techniques generate ensemble diversity, (2) when defining feature subsets and decision thresholds for each individual matcher of an ensemble using non-target videos, and (3) when accumulating positive detections over multiple frames. Saman Bashbaghi, Eric Granger, Robert Sabourin, Guillaume-Alexandre Bilodeau |
ICPR | 4 |
| 2014 | Structure-aware keypoint tracking for partial occlusion handlingabstractThis paper introduces a novel keypoint-based method for visual object tracking. To represent the target, we use a new model combining color distribution with keypoints. The appearance model also incorporates the spatial layout of the keypoints, encoding the object structure learned during tracking. With this multi-feature appearance model, our Structure-Aware Tracker (SAT) estimates accurately the target location using three main steps. First, the search space is reduced to the most likely image regions with a probabilistic approach. Second, the target location is estimated in the reduced search space using deterministic keypoint matching. Finally, the location prediction is corrected by exploiting the keypoint structural model with a voting-based method. By applying our SAT on several tracking problems, we show that location correction based on structural constraints is a key technique to improve prediction in moderately crowded scenes, even if only a small part of the target is visible. We also conduct comparison with a number of state-of-the-art trackers and demonstrate the competitiveness of the proposed method. Wassim Bouachir, Guillaume-Alexandre Bilodeau |
WACV | 2 |
| 2014 | Urban Tracker: Multiple object tracking in urban mixed trafficabstractIn this paper, we study the problem of detecting and tracking multiple objects of various types in outdoor urban traffic scenes. This problem is especially challenging due to the large variation of road user appearances. To handle that variation, our system uses background subtraction to detect moving objects. In order to build the object tracks, an object model is built and updated through time inside a state machine using feature points and spatial information. When an occlusion occurs between multiple objects, the positions of feature points at previous observations are used to estimate the positions and sizes of the individual occluded objects. Our Urban Tracker algorithm is validated on four outdoor urban videos involving mixed traffic that includes pedestrians, cars, large vehicles, etc. Our method compares favorably to a current state of the art feature-based tracker for urban traffic scenes on pedestrians and mixed traffic. Jean-Philippe Jodoin, Guillaume-Alexandre Bilodeau, Nicolas Saunier |
WACV | 2 |
| 2014 | Improving background subtraction using Local Binary Similarity PatternsabstractMost of the recently published background subtraction methods can still be classified as pixel-based, as most of their analysis is still only done using pixel-by-pixel comparisons. Few others might be regarded as spatial-based (or even spatiotemporal-based) methods, as they take into account the neighborhood of each analyzed pixel. Although the latter types can be viewed as improvements in many cases, most of the methods that have been proposed so far suffer in complexity, processing speed, and/or versatility when compared to their simpler pixel-based counterparts. In this paper, we present an adaptive background subtraction method, derived from the low-cost and highly efficient ViBe method, which uses a spatiotemporal binary similarity descriptor instead of simply relying on pixel intensities as its core component. We then test this method on multiple video sequences and show that by only replacing the core component of a pixel-based method it is possible to dramatically improve its overall performance while keeping memory usage, complexity and speed at acceptable levels for online applications. Pierre-Luc St-Charles, Guillaume-Alexandre Bilodeau |
WACV | 2 |
| 2014 | A computationally efficient importance sampling tracking algorithm
Rana Farah, Qifeng Gan, J. M. Pierre Langlois, Guillaume-Alexandre Bilodeau, Yvon Savaria |
Mach. Vis. Appl. | 4 |
| 2013 | A LSS-based registration of stereo thermal-visible videos of multiple people using belief propagation
Atousa Torabi, Guillaume-Alexandre Bilodeau |
Comput. Vis. Image Underst. | 2 |
| 2013 | Local self-similarity-based registration of human ROIs in pairs of stereo thermal-visible videos
Atousa Torabi, Guillaume-Alexandre Bilodeau |
Pattern Recognit. | 2 |
| 2013 | Catching a Rat by Its EdgletsabstractComputer vision is a noninvasive method for monitoring laboratory animals. In this article, we propose a robust tracking method that is capable of extracting a rodent from a frame under uncontrolled normal laboratory conditions. The method consists of two steps. First, a sliding window combines three features to coarsely track the animal. Then, it uses the edglets of the rodent to adjust the tracked region to the animal's boundary. The method achieves an average tracking error that is smaller than a representative state-of-the-art method. Rana Farah, J. M. Pierre Langlois, Guillaume-Alexandre Bilodeau |
IEEE Trans. Image Process. | 3 |
| 2012 | An iterative integrated framework for thermal-visible image registration, sensor fusion, and people tracking for video surveillance applications
Atousa Torabi, Guillaume Massé, Guillaume-Alexandre Bilodeau |
Comput. Vis. Image Underst. | 3 |
| 2012 | Body temperature estimation of a moving subject from thermographic images
Guillaume-Alexandre Bilodeau, Atousa Torabi, Maxime Levesque, Charles Ouellet, J. M. Pierre Langlois, Pablo Lema, Lionel Carmant |
Mach. Vis. Appl. | 1 |
| 2011 | Comparative analysis of contrast enhancement algorithms in surveillance imagingabstractImage contrast enhancement methods play a key role in many image processing and vision applications. For surveillance applications, real-time contrast improvement over the whole image is required when videos are taken in poor lighting conditions. It is also necessary to highlight details in shadowed regions without introducing artifacts. In this paper, several state-of-the-art contrast enhancement methods are compared. Image quality is evaluated by means of objective metrics such as intensity contrast and brightness error, and by subjective assessment. Execution time is also measured. Experimental results show that a technique based on histogram modification presents a better trade-off considering both aspects. Diana Carolina Gil, Rana Farah, J. M. Pierre Langlois, Guillaume-Alexandre Bilodeau, Yvon Savaria |
ISCAS | 4 |
| 2011 | Visible and infrared image registration using trajectories and composite foreground images
Guillaume-Alexandre Bilodeau, Atousa Torabi, François Morin |
Image Vis. Comput. | 1 |
| 2011 | People tracking using a network-based PTZ camera
Parisa Darvish Zadeh Varcheie, Guillaume-Alexandre Bilodeau |
Mach. Vis. Appl. | 2 |
| 2009 | Fuzzy Feature-Based Upper Body Tracking with IP PTZ Camera Control
Parisa Darvish Zadeh Varcheie, Guillaume-Alexandre Bilodeau |
CIARP | 2 |
| 2007 | Qualitative part-based models in content-based image retrieval
Guillaume-Alexandre Bilodeau, Robert Bergevin |
Mach. Vis. Appl. | 1 |
| 2002 | Part segmentation of objects in real images
Guillaume-Alexandre Bilodeau, Robert Bergevin |
Pattern Recognit. | 1 |
| 2000 | Generic Modeling of 3D Objects from Single 2D ImagesabstractAddresses the problem of building generic 3D models of structured objects on the basis of single 2D intensity images. In the context of the paper, generic modeling refers to the situation where analysis of the image information is performed on the sole basis of generic knowledge. That is, no a priori knowledge about the specific quantitative shape properties of the objects of interest is ever assumed. Moreover, images of interest are realistic. For instance, they may contain complex foreground 3D objects with textures and shadows, and a cluttered background. Objects are modeled by their constituent parts and connections. Therefore, a partly occluded object could be recognized from its model. Part models are based on geons, which are a set of qualitative generalized cylinders. An overview of the architecture of the modeling system is presented, along with the functionality of each subsystem and processing results. Guillaume-Alexandre Bilodeau, Robert Bergevin |
ICPR | 1 |