Guillaume-Alexandre Bilodeau

dblp:90/2054 · DBLP profile ↗
← Back
58ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0003-3227-5060ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 41 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Continuous conditional video synthesis by neural processes
abstract
Different conditional video synthesis tasks, such as frame interpolation and future frame prediction, are typically addressed individually by task-specific models, despite their shared underlying characteristics. Additionally, most conditional video synthesis models are limited to discrete frame generation at specific integer time steps. This paper presents a unified model that tackles both challenges simultaneously. We demonstrate that conditional video synthesis can be formulated as a neural process, where input spatio-temporal coordinates are mapped to target pixel values by conditioning on context spatio-temporal coordinates and pixel values. Our approach leverages a Transformer-based non-autoregressive conditional video synthesis model that takes the implicit neural representation of coordinates and context pixel features as input. Our task-specific models outperform previous methods for future frame prediction and frame interpolation across multiple datasets. Importantly, our model enables temporal continuous video synthesis at arbitrary high frame rates, outperforming the previous state-of-the-art. The source code and video demos for our model are available at https://npvp.github.io .
Xi Ye 0005, Guillaume-Alexandre Bilodeau
Comput. Vis. Image Underst.2
2025 Learning data association for multi-object tracking using only coordinates
abstract
We propose a novel Transformer-based module to address the data association problem for multi-object tracking. From detections obtained by a pretrained detector, this module uses only coordinates from bounding boxes to estimate an affinity score between pairs of tracks extracted from two distinct temporal windows. This module, named TWiX, is trained on sets of tracks with the objective of discriminating pairs of tracks coming from the same object from those which are not. Our module does not use the intersection over union measure, nor does it requires any motion priors or any camera motion compensation technique. By inserting TWiX within an online cascade matching pipeline, our tracker C-TWiX achieves state-of-the-art performance on the DanceTrack and KITTIMOT datasets, and gets competitive results on the MOT17 dataset. The code will be made available upon publication on the website https://mehdimiah.com/twix . • Our Transformer-based model, TWiX, can learn to associate objects using only coordinates. • We show that motion priors or intersection-over-union measure are not required for tracking. Using pairs of tracks is sufficient. • Tracking with TWiX gives competitive or state-of-the-results on several datasets.
Mehdi Miah, Guillaume-Alexandre Bilodeau, Nicolas Saunier
Pattern Recognit.2
2024 STDiff: Spatio-Temporal Diffusion for Continuous Stochastic Video Prediction
abstract
Predicting future frames of a video is challenging because it is difficult to learn the uncertainty of the underlying factors influencing their contents. In this paper, we propose a novel video prediction model, which has infinite-dimensional latent variables over the spatio-temporal domain. Specifically, we first decompose the video motion and content information, then take a neural stochastic differential equation to predict the temporal motion information, and finally, an image diffusion model autoregressively generates the video frame by conditioning on the predicted motion feature and the previous frame. The better expressiveness and stronger stochasticity learning capability of our model lead to state-of-the-art video prediction performances. As well, our model is able to achieve temporal continuous prediction, i.e., predicting in an unsupervised way the future video frames with an arbitrarily high frame rate. Our code is available at https://github.com/XiYe20/STDiffProject.
Xi Ye 0005, Guillaume-Alexandre Bilodeau
AAAI2
2024 ReL-SAR: Representation Learning for Skeleton Action Recognition with Convolutional Transformers and BYOL
abstract
To extract robust and generalizable skeleton action recognition features, large amounts of well-curated data are typically required, which is a challenging task hindered by annotation and computation costs. Therefore, unsupervised representation learning is of prime importance to leverage unlabeled skeleton data. In this work, we investigate unsupervised representation learning for skeleton action recognition. For this purpose, we designed a lightweight convolutional transformer framework, named ReL-SAR, exploiting the complementarity of convolutional and attention layers for jointly modeling spatial and temporal cues in skeleton sequences. We also use a Selection-Permutation strategy for skeleton joints to ensure more infor-mative descriptions from skeletal data. Finally, we capitalize on Bootstrap Your Own Latent (BYOL) to learn robust representations from unlabeled skeleton sequence data. We achieved very competitive results on limited-size datasets: MCAD, IXMAS, JH-MDB, and NW-UCLA, showing the effectiveness of our proposed method against state-of-the-art methods in terms of both performance and computational efficiency. To ensure reproducibility and reusability, the source code including all implementation parameters is provided at https://github.com/SafwenNaimi.
Safwen Naimi, Wassim Bouachir, Guillaume-Alexandre Bilodeau
ICMLA3
2024 1D-convolutional transformer for Parkinson disease diagnosis from gait
Safwen Naimi, Wassim Bouachir, Guillaume-Alexandre Bilodeau
Neural Comput. Appl.3
2023 HCT: Hybrid Convnet-Transformer for Parkinson's Disease Detection and Severity Prediction from Gait
abstract
In this paper, we propose a novel deep learning method based on a new Hybrid ConvNet-Transformer archi-tecture to detect and stage Parkinson's disease (PD) from gait data. We adopt a two-step approach by dividing the problem into two sub-problems. Our Hybrid ConvNet-Transformer model first distinguishes healthy versus parkinsonian patients. If the patient is parkinsonian, a multi-class Hybrid ConvNet-Transformer model determines the Hoehn and Yahr (H&Y) score to assess the PD severity stage. Our hybrid architecture exploits the strengths of both Convolutional Neural Networks (ConvNets) and Transformers to accurately detect PD and determine the severity stage. In particular, we take advantage of ConvNets to capture local patterns and correlations in the data, while we exploit Transformers for handling long-term dependencies in the input signal. We show that our hybrid method achieves superior performance when compared to other state-of-the-art methods, with a PD detection accuracy of 97% and a severity staging accuracy of 87%. Our source code is available at https://github.com/SafwenNaimi.
Safwen Naimi, Wassim Bouachir, Guillaume-Alexandre Bilodeau
ICMLA3
2023 Automating Lichen Monitoring in Ecological Studies Using Instance Segmentation of Time-Lapse Images
abstract
Lichens are symbiotic organisms composed of fungi, algae, and/or cyanobacteria that thrive in a variety of environments. They play important roles in carbon and nitrogen cycling, and contribute directly and indirectly to biodiversity. Ecologists typically monitor lichens by using them as indicators to assess air quality and habitat conditions. In particular, epiphytic lichens, which live on trees, are key markers of air quality and environmental health. A new method of monitoring epiphytic lichens involves using time-lapse cameras to gather images of lichen populations. These cameras are used by ecologists in Newfoundland and Labrador to subsequently analyze and manually segment the images to determine lichen thalli condition and change. These methods are time-consuming and susceptible to observer bias. In this work, we aim to automate the monitoring of lichens over extended periods and to estimate their biomass and condition to facilitate the task of ecologists. To accomplish this, our proposed framework uses semantic segmentation with an effective training approach to automate monitoring and biomass estimation of epiphytic lichens on time-lapse images. We show that our method has the potential to significantly improve the accuracy and efficiency of lichen population monitoring, making it a valuable tool for forest ecologists and environmental scientists to evaluate the impact of climate change on Canada's forests. To the best of our knowledge, this is the first time that such an approach has been used to assist ecologists in monitoring and analyzing epiphytic lichens.
Safwen Naimi, Olfa Koubaa, Wassim Bouachir, Guillaume-Alexandre Bilodeau, Gregory Jeddore, Patricia Baines, David L. P. Correia, Andre Arsenault
ICMLA4
2023 Video prediction by efficient transformers
Xi Ye 0005, Guillaume-Alexandre Bilodeau
Image Vis. Comput.2
2022 Transformers for 1D signals in Parkinson's disease detection from gait
abstract
This paper focuses on the detection of Parkinson’s disease based on the analysis of a patient’s gait. The growing popularity and success of Transformer networks in natural language processing and image recognition motivated us to develop a novel method for this problem based on an automatic features extraction via Transformers. The use of Transformers in 1D signal is not really widespread yet, but we show in this paper that they are effective in extracting relevant features from 1D signals. As Transformers require a lot of memory, we decoupled temporal and spatial information to make the model smaller. Our architecture used temporal Transformers, dimension reduction layers to reduce the dimension of the data, a spatial Transformer, two fully connected layers and an output layer for the final prediction. Our model outperforms the current state-of-the-art algorithm with 95.2% accuracy in distinguishing a Parkinsonian patient from a healthy one on the Physionet dataset. A key learning from this work is that Transformers allow for greater stability in results. The source code and pre-trained models are released in https://github.com/DucMinhDimitriNguyen1.
Duc Minh Dimitri Nguyen, Mehdi Miah, Guillaume-Alexandre Bilodeau, Wassim Bouachir
ICPR3
2022 VPTR: Efficient Transformers for Video Prediction
abstract
In this paper, we propose a new Transformer block for video future frames prediction based on an efficient local spatial-temporal separation attention mechanism. Based on this new Transformer block, a fully autoregressive video future frames prediction Transformer is proposed. In addition, a non-autoregressive video prediction Transformer is also proposed to increase the inference speed and reduce the accumulated inference errors of its autoregressive counterpart. In order to avoid the prediction of very similar future frames, a contrastive feature loss is applied to maximize the mutual information between predicted and ground-truth future frame features. This work is the first that makes a formal comparison of the two types of attention-based video future frames prediction models over different scenarios. The proposed models reach a performance competitive with more complex state-of-the-art models. The source code is available at https://github.com/XiYe20/VPTR.
Xi Ye 0005, Guillaume-Alexandre Bilodeau
ICPR2
2021 Multiple convolutional features in Siamese networks for object tracking
Zhenxi Li, Guillaume-Alexandre Bilodeau, Wassim Bouachir
Mach. Vis. Appl.2
2021 FFAVOD: Feature fusion architecture for video object detection
Hughes Perreault, Guillaume-Alexandre Bilodeau, Nicolas Saunier, Maguelonne Héritier
Pattern Recognit. Lett.2
2020 Domain Siamese CNNs for Sparse Multispectral Disparity Estimation
abstract
Multispectral disparity estimation is a difficult task for many reasons: it has all the same challenges as traditional visible-visible disparity estimation (occlusions, repetitive patterns, textureless surfaces), in addition of having very few common visual information between images (e.g. colour information vs. thermal information). In this paper, we propose a new CNN architecture able to do disparity estimation between images from different spectra, namely thermal and visible in our case. Our proposed model takes two patches as input and proceeds to do domain feature extraction for each of them. Features from both domains are then merged with two fusion operations, namely correlation and concatenation. These merged vectors are then forwarded to their respective classification heads, which are responsible for classifying the inputs as being same or not. Using two merging operations gives more robustness to our feature extraction process, which leads to more precise disparity estimation. Our method was tested using the publicly available LITIV 2014 and LITIV 2018 datasets, and showed best results when compared to other state-of-the-art methods.
David-Alexandre Beaupré, Guillaume-Alexandre Bilodeau
ICPR2
2020 A Grid-based Representation for Human Action Recognition
abstract
Human action recognition (HAR) in videos is a fundamental research topic in computer vision. It consists mainly in understanding actions performed by humans based on a sequence of visual observations. In recent years, HAR have witnessed significant progress, especially with the emergence of deep learning models. However, most of existing approaches for action recognition rely on information that is not always relevant for this task, and are limited in the way they fuse the temporal information. In this paper, we propose a novel method for human action recognition that encodes efficiently the most discriminative appearance information of an action with explicit attention on representative pose features, into a new compact grid representation. Our GRAR (Grid-based Representation for Action Recognition) method is tested on several benchmark datasets demonstrating that our model can accurately recognize human actions, despite intra-class appearance variations and occlusion challenges.
Soufiane Lamghari, Guillaume-Alexandre Bilodeau, Nicolas Saunier
ICPR2
2020 MFST: Multi-Features Siamese Tracker
abstract
Siamese trackers have recently achieved interesting results due to their balance between accuracy and speed. This success is mainly due to the fact that deep similarity networks were specifically designed to address the image similarity problem. Therefore, they are inherently more appropriate than classical CNNs for the tracking task. However, Siamese trackers rely on the last convolutional layers for similarity analysis and target search, which restricts their performance. In this paper, we argue that using a single convolutional layer as feature representation is not the optimal choice within the deep similarity framework, as multiple convolutional layers provide several abstraction levels in characterizing an object. Starting from this motivation, we present the Multi-Features Siamese Tracker (MFST), a novel tracking algorithm exploiting several hierarchical feature maps for robust deep similarity tracking. MFST proceeds by fusing hierarchical features to ensure a richer and more efficient representation. Moreover, we handle appearance variation by calibrating deep features extracted from two different CNN models. Based on this advanced feature representation, our algorithm achieves high tracking accuracy, while outperforming several state-of-the-art trackers, including standard Siamese trackers.
Zhenxi Li, Guillaume-Alexandre Bilodeau, Wassim Bouachir
ICPR2
2020 An Empirical Analysis of Visual Features for Multiple Object Tracking in Urban Scenes
abstract
This paper addresses the problem of selecting appearance features for multiple object tracking (MOT) in urban scenes. Over the years, a large number of features has been used for MOT. However, it is not clear whether some of them are better than others. Commonly used features are color histograms, histograms of oriented gradients, deep features from convolutional neural networks and re-identification (ReID) features. In this study, we assess how good these features are at discriminating objects enclosed by a bounding box in urban scene tracking scenarios. Several affinity measures, namely the L1, L2and the Bhattacharyya distances, Rank-1 counts and the cosine similarity, are also assessed for their impact on the discriminative power of the features. Results on several datasets show that features from ReID networks are the best for discriminating instances from one another regardless of the quality of the detector. If a ReID model is not available, color histograms may be selected if the detector has a good recall and there are few occlusions; otherwise, deep features are more robust to detectors with lower recall.
Mehdi Miah, Justine Pepin, Nicolas Saunier, Guillaume-Alexandre Bilodeau
ICPR4
2020 Deep 1D-Convnet for accurate Parkinson disease detection and severity prediction from gait
Imanne El Maachi, Guillaume-Alexandre Bilodeau, Wassim Bouachir
Expert Syst. Appl.2
2020 Robust face tracking using multiple appearance models and graph relational learning
Tanushri Chakravorty, Guillaume-Alexandre Bilodeau, Eric Granger
Mach. Vis. Appl.2
2019 Online Mutual Foreground Segmentation for Multispectral Stereo Videos
Pierre-Luc St-Charles, Guillaume-Alexandre Bilodeau, Robert Bergevin
Int. J. Comput. Vis.2
2019 From superpixel to human shape modelling for carried object detection
Farnoosh Ghadiri, Robert Bergevin, Guillaume-Alexandre Bilodeau
Pattern Recognit.3
2019 Domain-Specific Face Synthesis for Video Face Recognition From a Single Sample Per Person
abstract
In video surveillance, face recognition (FR) systems are employed to detect individuals of interest appearing over a distributed network of cameras. The performance of still-to-video FR systems can decline significantly because faces captured in unconstrained operational domain (OD) over multiple video cameras have a different underlying data distribution compared to faces captured under controlled conditions in the enrollment domain with a still camera. This is particularly true when individuals are enrolled to the system using a single reference still. To improve the robustness of these systems, it is possible to augment the reference set by generating synthetic faces based on the original still. However, without the knowledge of the OD, many synthetic images must be generated to account for all possible capture conditions. FR systems may, therefore, require complex implementations and yield lower accuracy when training on many less relevant images. This paper introduces an algorithm for domain-specific face synthesis (DSFS) that exploits the representative intra-class variation information available from the OD. Prior to operation (during camera calibration), a compact set of faces from unknown persons appearing in the OD is selected through affinity propagation clustering in the captured condition space (defined by pose and illumination estimation). The domain-specific variations of these face images are then projected onto the reference still of each individual by integrating an image-based face relighting technique inside the 3-D reconstruction framework. A compact set of synthetic faces is generated that resemble individuals of interest under the capture conditions relevant to the OD. In a particular implementation based on sparse representation classification, the synthetic faces generated with the DSFS are employed to form a cross-domain dictionary that accounts for structured sparsity, where the dictionary blocks combine the original and synthetic faces of each individual. Experimental results obtained with videos from the Chokepoint and COX-S2V data sets reveal that augmenting the reference gallery set of still-to-video FR systems using the proposed DSFS approach can provide a significantly higher level of accuracy compared with the state-of-the-art approaches, with only a moderate increase in its computational complexity.
Fania Mokhayeri, Eric Granger, Guillaume-Alexandre Bilodeau
IEEE Trans. Inf. Forensics Secur.3
2018 Tracking using Numerous Anchor Points
Tanushri Chakravorty, Guillaume-Alexandre Bilodeau, Eric Granger
Mach. Vis. Appl.2
2018 SPiKeS: Superpixel-Keypoints structure for robust visual tracking
François-Xavier Derue, Guillaume-Alexandre Bilodeau, Robert Bergevin
Mach. Vis. Appl.2
2017 Spatio-Temporal Consistency to Detect and Segment Carried Objects
Farnoosh Ghadiri, Robert Bergevin, Guillaume-Alexandre Bilodeau
BMVC3
2017 Dynamic Selection of Exemplar-SVMs for Watch-list Screening through Domain Adaptation
Saman Bashbaghi, Eric Granger, Robert Sabourin, Guillaume-Alexandre Bilodeau
ICPRAM4
2017 Robust watch-list screening using dynamic ensembles of SVMs based on multiple face representations
Saman Bashbaghi, Eric Granger, Robert Sabourin, Guillaume-Alexandre Bilodeau
Mach. Vis. Appl.4
2017 Dynamic ensembles of exemplar-SVMs for still-to-video face recognition
Saman Bashbaghi, Eric Granger, Robert Sabourin, Guillaume-Alexandre Bilodeau
Pattern Recognit.4
2016 Carried Object Detection Based on an Ensemble of Contour Exemplars
Farnoosh Ghadiri, Robert Bergevin, Guillaume-Alexandre Bilodeau
ECCV (7)3
2016 Online multi-object tracking by detection based on generative appearance models
Dorra Riahi, Guillaume-Alexandre Bilodeau
Comput. Vis. Image Underst.2
2016 Universal Background Subtraction Using Word Consensus Models
abstract
Background subtraction is often used as the first step in video analysis and smart surveillance applications. However, the issue of inconsistent performance across different scenarios due to a lack of flexibility remains a serious concern. To address this, we propose a novel non-parametric, pixel-level background modeling approach based on word dictionaries that draws from traditional codebooks and sample consensus approaches. In this new approach, the importance of each background sample (or word) is evaluated online based on their recurrence among all local observations. This helps build smaller pixel models that are better suited for long-term foreground detection. Combining these models with a frame-level dictionary and local feedback mechanisms leads us to our proposed background subtraction method, coined “PAWCS.” Experiments on the 2012 and 2014 versions of the ChangeDetection.net data set show that PAWCS outperforms 26 previously tested and published methods in terms of overall F-Measure as well as in most categories taken individually. Our results can be reproduced with a C++ implementation available online.
Pierre-Luc St-Charles, Guillaume-Alexandre Bilodeau, Robert Bergevin
IEEE Trans. Image Process.2
2016 Tracking All Road Users at Multimodal Urban Traffic Intersections
abstract
Because of the large variability of road user appearance in an urban setting, it is very challenging to track all of them with the purpose of obtaining precise and reliable trajectories. However, obtaining the trajectories of the various road users is very useful for many transportation applications. It is particularly essential for any task that requires higher level behavior interpretation, including new safety diagnosis methods that rely on the observation of road user interactions without a collision and therefore do not require waiting for collisions to happen. In this paper, we propose a tracking method that has been specifically designed to track the various road users that may be encountered in an urban environment. Since road users have very diverse shapes and appearances, our proposed method starts from background subtraction to extract the potential a priori unknown road users. Each of these road users is then tracked using a collection of keypoints inside the detected foreground regions, which allows the interpolation of object locations even during object merges or occlusions. A finite state machine handles fragmentation, splitting, and merging of the road users to correct and improve the resulting object trajectories. The proposed tracker was tested on several urban intersection videos and is shown to outperform an existing reference tracker used in transportation research.
Jean-Philippe Jodoin, Guillaume-Alexandre Bilodeau, Nicolas Saunier
IEEE Trans. Intell. Transp. Syst.2
2015 Ensembles of exemplar-SVMs for video face recognition from a single sample per person
abstract
Recognizing the face of target individuals in a watch-list is among the most challenging applications in video surveillance, especially when enrollment is based on one reference still facial image. Besides the limited representativeness of facial models used for matching, the appearance of faces captured in videos varies due to changes in illumination, pose, scales, etc., and to camera inter-operability. A multi-classifier system is proposed in this paper for robust still-to-video face recognition (FR) based on multiple diverse face representations. An individual-specific ensemble of exemplar-SVMs (e-SVMs) classifiers is assigned to each target person, where each classifier is trained using a high-quality reference face still versus many lower-quality faces of non-target individuals captured in videos. Diverse face representations are generated from different patches isolated in facial images and face descriptors that are robust to various nuisance factors (e.g., illumination and pose) commonly encountered in surveillance environments. Discriminant feature subsets, training samples, and ensemble fusion functions are selected using faces of non-target individuals captured in videos of the scene. Experiments on videos from the Chokepoint dataset reveal that the proposed ensemble of e-SVMs outperforms state-of-the-art FR systems specialized for the single sample per person problem.
Saman Bashbaghi, Eric Granger, Robert Sabourin, Guillaume-Alexandre Bilodeau
AVSS4
2015 Contextual object tracker with structure encoding
abstract
Motivated by the problem of object tracking in video sequences, this paper presents a new Contextual Object Tracker with Structural Encoding (CTSE). The novelty in our tracking approach lies in the application of contextual and structural information (that is specific to a target object) into a model-free tracker. This is first achieved by including features from a complementary region having correlated motion with the target object. Second, a local structure that represents a spatial constraint between features within the target object are included. SIFT keypoints are used as features to encode both these information. The tracking is done in three steps. Firstly, keypoints are detected and described to encode object structure. Secondly, they are matched in every frame. Finally, each matched keypoint votes for the target object location locally in a voting matrix by using the encoded object structure. The voting method gives more priority to the keypoints that have been matched more often and are closest to the target's center than the rest. The proposed tracker is competitive with state-of-the art trackers while being significantly faster. It ranks as first or second most accurate tracker in experiments with standard datasets.
Tanushri Chakravorty, Guillaume-Alexandre Bilodeau, Eric Granger
ICIP2
2015 Reproducible evaluation of Pan-Tilt-Zoom tracking
abstract
Tracking with a Pan-Tilt-Zoom (PTZ) camera has been a research topic in computer vision for many years. However, it is difficult to assess the progress that has been made because there is no standard evaluation methodology. The difficulty in evaluating PTZ tracking algorithms arises from their dynamic nature. In contrast to other forms of tracking, PTZ tracking involves both locating the target in the image and controlling the motors of the camera to aim it so that the target stays in its field of view. This type of tracking can only be performed online. In this paper, we propose a new evaluation framework based on a virtual PTZ camera. With this framework, tracking scenarios do not change for each experiment and we are able to replicate the main principles of online PTZ camera control and behavior including camera positioning delays, tracker processing delays, and numerical zoom. We tested our evaluation framework with the Camshift tracker to show its viability and to establish baseline results.
Gengjie Chen, Pierre-Luc St-Charles, Wassim Bouachir, Guillaume-Alexandre Bilodeau, Robert Bergevin
ICIP4
2015 Synthetic face generation under various operational conditions in video surveillance
abstract
In still-to-video face recognition (FR), the faces captured with surveillance cameras are matched against reference stills of target individuals enrolled to the system. FR is a challenging problem in video surveillance due to uncontrolled capture conditions (variations in pose, expression, illumination, blur, scale, etc.), and the limited number of reference stills to model target individuals. This paper introduces a new approach to generate multiple synthetic face images per reference still based on camera-specific capture conditions to deal with illumination variations. For each reference still, a diverse set of faces from non-target individuals appearing in the camera viewpoint are selected based on luminance and contrast distortion. These face images are then decomposed into detail layer and large scale layer using an edge-preserving image decomposition to obtain their illumination dependent component. Finally, the large scale layers of these images are morphed with each reference still image to generate multiple synthetic reference stills that incorporate illumination and contrast conditions. Experimental results obtained with the ChokePoint dataset reveal that these synthetic faces produce an enhanced face model. As the number of synthetic faces grows, the proposed approach provides a higher level of accuracy and robustness across a range of capture conditions.
Fania Mokhayeri, Eric Granger, Guillaume-Alexandre Bilodeau
ICIP3
2015 Multiple object tracking based on sparse generative appearance modeling
abstract
This paper addresses multiple object tracking which still remains a challenging problem because of factors like frequent occlusions, unknown number of targets and similarity in objects' appearance. We propose a novel approach for multiple object tracking using a multiple feature framework. The main focus of the proposed method is to build a robust appearance model. The appearance model of an object is built using a color model, a sparse appearance model, a motion model and spatial information. We validated the proposed algorithm on four publicly available videos with comparisons with state-of-the-art approaches. We demonstrate that our algorithm achieves competitive results.
Dorra Riahi, Guillaume-Alexandre Bilodeau
ICIP2
2015 Part-Based Tracking via Salient Collaborating Features
abstract
We present a novel part-based method for model-free tracking. In our model, key points are considered as elementary predictors, collaborating to localize the target. In order to differentiate reliable features from outliers and bad predictors, we define the notion of feature saliency including three factors: the persistence, the spatial consistency, and the predictive power of local features. Saliency information is learned during tracking to be used in several algorithmic steps: local predictions, global localization, feature removal, etc. By exploiting saliency information and key point structural properties, the proposed algorithm is able to track accurately generic objects, facing several difficulties such as occlusions, presence of distractors, and abrupt motion. The proposed tracker demonstrated a high robustness on challenging public datasets, outperforming significantly five recent state-of-the-art trackers.
Wassim Bouachir, Guillaume-Alexandre Bilodeau
WACV2
2015 A Self-Adjusting Approach to Change Detection Based on Background Word Consensus
abstract
Although there has long been interest in foreground background segmentation based on change detection for video surveillance applications, the issue of inconsistent performance across different scenarios remains a serious concern. To address this, we propose a new type of word based approach that regulates its own internal parameters using feedback mechanisms to withstand difficult conditions while keeping sensitivity intact in regular situations. Coined "PAWCS", this method's key advantages lie in its highly persistent and robust dictionary model based on color and local binary features as well as its ability to automatically adjust pixel-level segmentation behavior. Experiments using the 2012 Change Detection.net dataset show that it outranks numerous recently proposed solutions in terms of overall performance as well as in each category. A complete C++ implementation based on OpenCV is available online.
Pierre-Luc St-Charles, Guillaume-Alexandre Bilodeau, Robert Bergevin
WACV2
2015 Collaborative part-based tracking using salient local predictors
Wassim Bouachir, Guillaume-Alexandre Bilodeau
Comput. Vis. Image Underst.2
2015 Exploiting structural constraints for visual object tracking
Wassim Bouachir, Guillaume-Alexandre Bilodeau
Image Vis. Comput.2
2015 SuBSENSE: A Universal Change Detection Method With Local Adaptive Sensitivity
abstract
Foreground/background segmentation via change detection in video sequences is often used as a stepping stone in high-level analytics and applications. Despite the wide variety of methods that have been proposed for this problem, none has been able to fully address the complex nature of dynamic scenes in real surveillance tasks. In this paper, we present a universal pixel-level segmentation method that relies on spatiotemporal binary features as well as color information to detect changes. This allows camouflaged foreground objects to be detected more easily while most illumination variations are ignored. Besides, instead of using manually set, frame-wide constants to dictate model sensitivity and adaptation speed, we use pixel-level feedback loops to dynamically adjust our method's internal parameters without user intervention. These adjustments are based on the continuous monitoring of model fidelity and local segmentation noise levels. This new approach enables us to outperform all 32 previously tested state-of-the-art methods on the 2012 and 2014 versions of the ChangeDetection.net dataset in terms of overall F-Measure. The use of local binary image descriptors for pixel-level modeling also facilitates high-speed parallel implementations: our own version, which used no low-level or architecture-specific instruction, reached real-time processing speed on a midlevel desktop CPU. A complete C++ implementation based on OpenCV is available online.
Pierre-Luc St-Charles, Guillaume-Alexandre Bilodeau, Robert Bergevin
IEEE Trans. Image Process.2
2014 Watch-List Screening Using Ensembles Based on Multiple Face Representations
abstract
Still-to-video face recognition (FR) is an important function in watch list screening, where faces captured over a network of video surveillance cameras are matched against reference stills of target individuals. Recognizing faces in a watch list is a challenging problem in semi -- and unconstrained surveillance environments due to the lack of control over capture and operational conditions, and to the limited number of reference stills. This paper provides a performance baseline and guidelines for ensemble-based systems using a single high-quality reference still per individual, as found in many watch list screening applications. In particular, modular systems are considered, where an ensemble of template matchers based on multiple face representations is assigned to each individual of interest. During enrollment, multiple feature extraction (FE) techniques are applied to patches isolated in the reference still to generate diverse face-part representations that are robust to various nuisance factors (e.g., illumination and pose) encountered in video surveillance. The selection of relevant feature subsets, decision thresholds, and fusion functions of ensembles are achieved using faces of non-target individuals selected from reference videos (forming a universal background model). During operations, a face tracker gradually regroups faces captured from different people appearing in a scene, while each user-specific ensemble generates a decision per face capture. This leads to robust spatio-temporal FR when accumulated ensemble predictions surpass a detection threshold. Simulation results obtained with the Chokepoint video dataset show a significant improvement to accuracy, (1) when performing score-level fusion of matchers, where patches-based and FE techniques generate ensemble diversity, (2) when defining feature subsets and decision thresholds for each individual matcher of an ensemble using non-target videos, and (3) when accumulating positive detections over multiple frames.
Saman Bashbaghi, Eric Granger, Robert Sabourin, Guillaume-Alexandre Bilodeau
ICPR4
2014 Structure-aware keypoint tracking for partial occlusion handling
abstract
This paper introduces a novel keypoint-based method for visual object tracking. To represent the target, we use a new model combining color distribution with keypoints. The appearance model also incorporates the spatial layout of the keypoints, encoding the object structure learned during tracking. With this multi-feature appearance model, our Structure-Aware Tracker (SAT) estimates accurately the target location using three main steps. First, the search space is reduced to the most likely image regions with a probabilistic approach. Second, the target location is estimated in the reduced search space using deterministic keypoint matching. Finally, the location prediction is corrected by exploiting the keypoint structural model with a voting-based method. By applying our SAT on several tracking problems, we show that location correction based on structural constraints is a key technique to improve prediction in moderately crowded scenes, even if only a small part of the target is visible. We also conduct comparison with a number of state-of-the-art trackers and demonstrate the competitiveness of the proposed method.
Wassim Bouachir, Guillaume-Alexandre Bilodeau
WACV2
2014 Urban Tracker: Multiple object tracking in urban mixed traffic
abstract
In this paper, we study the problem of detecting and tracking multiple objects of various types in outdoor urban traffic scenes. This problem is especially challenging due to the large variation of road user appearances. To handle that variation, our system uses background subtraction to detect moving objects. In order to build the object tracks, an object model is built and updated through time inside a state machine using feature points and spatial information. When an occlusion occurs between multiple objects, the positions of feature points at previous observations are used to estimate the positions and sizes of the individual occluded objects. Our Urban Tracker algorithm is validated on four outdoor urban videos involving mixed traffic that includes pedestrians, cars, large vehicles, etc. Our method compares favorably to a current state of the art feature-based tracker for urban traffic scenes on pedestrians and mixed traffic.
Jean-Philippe Jodoin, Guillaume-Alexandre Bilodeau, Nicolas Saunier
WACV2
2014 Improving background subtraction using Local Binary Similarity Patterns
abstract
Most of the recently published background subtraction methods can still be classified as pixel-based, as most of their analysis is still only done using pixel-by-pixel comparisons. Few others might be regarded as spatial-based (or even spatiotemporal-based) methods, as they take into account the neighborhood of each analyzed pixel. Although the latter types can be viewed as improvements in many cases, most of the methods that have been proposed so far suffer in complexity, processing speed, and/or versatility when compared to their simpler pixel-based counterparts. In this paper, we present an adaptive background subtraction method, derived from the low-cost and highly efficient ViBe method, which uses a spatiotemporal binary similarity descriptor instead of simply relying on pixel intensities as its core component. We then test this method on multiple video sequences and show that by only replacing the core component of a pixel-based method it is possible to dramatically improve its overall performance while keeping memory usage, complexity and speed at acceptable levels for online applications.
Pierre-Luc St-Charles, Guillaume-Alexandre Bilodeau
WACV2
2014 A computationally efficient importance sampling tracking algorithm
Rana Farah, Qifeng Gan, J. M. Pierre Langlois, Guillaume-Alexandre Bilodeau, Yvon Savaria
Mach. Vis. Appl.4
2013 A LSS-based registration of stereo thermal-visible videos of multiple people using belief propagation
Atousa Torabi, Guillaume-Alexandre Bilodeau
Comput. Vis. Image Underst.2
2013 Local self-similarity-based registration of human ROIs in pairs of stereo thermal-visible videos
Atousa Torabi, Guillaume-Alexandre Bilodeau
Pattern Recognit.2
2013 Catching a Rat by Its Edglets
abstract
Computer vision is a noninvasive method for monitoring laboratory animals. In this article, we propose a robust tracking method that is capable of extracting a rodent from a frame under uncontrolled normal laboratory conditions. The method consists of two steps. First, a sliding window combines three features to coarsely track the animal. Then, it uses the edglets of the rodent to adjust the tracked region to the animal's boundary. The method achieves an average tracking error that is smaller than a representative state-of-the-art method.
Rana Farah, J. M. Pierre Langlois, Guillaume-Alexandre Bilodeau
IEEE Trans. Image Process.3
2012 An iterative integrated framework for thermal-visible image registration, sensor fusion, and people tracking for video surveillance applications
Atousa Torabi, Guillaume Massé, Guillaume-Alexandre Bilodeau
Comput. Vis. Image Underst.3
2012 Body temperature estimation of a moving subject from thermographic images
Guillaume-Alexandre Bilodeau, Atousa Torabi, Maxime Levesque, Charles Ouellet, J. M. Pierre Langlois, Pablo Lema, Lionel Carmant
Mach. Vis. Appl.1
2011 Comparative analysis of contrast enhancement algorithms in surveillance imaging
abstract
Image contrast enhancement methods play a key role in many image processing and vision applications. For surveillance applications, real-time contrast improvement over the whole image is required when videos are taken in poor lighting conditions. It is also necessary to highlight details in shadowed regions without introducing artifacts. In this paper, several state-of-the-art contrast enhancement methods are compared. Image quality is evaluated by means of objective metrics such as intensity contrast and brightness error, and by subjective assessment. Execution time is also measured. Experimental results show that a technique based on histogram modification presents a better trade-off considering both aspects.
Diana Carolina Gil, Rana Farah, J. M. Pierre Langlois, Guillaume-Alexandre Bilodeau, Yvon Savaria
ISCAS4
2011 Visible and infrared image registration using trajectories and composite foreground images
Guillaume-Alexandre Bilodeau, Atousa Torabi, François Morin
Image Vis. Comput.1
2011 People tracking using a network-based PTZ camera
Parisa Darvish Zadeh Varcheie, Guillaume-Alexandre Bilodeau
Mach. Vis. Appl.2
2009 Fuzzy Feature-Based Upper Body Tracking with IP PTZ Camera Control
Parisa Darvish Zadeh Varcheie, Guillaume-Alexandre Bilodeau
CIARP2
2007 Qualitative part-based models in content-based image retrieval
Guillaume-Alexandre Bilodeau, Robert Bergevin
Mach. Vis. Appl.1
2002 Part segmentation of objects in real images
Guillaume-Alexandre Bilodeau, Robert Bergevin
Pattern Recognit.1
2000 Generic Modeling of 3D Objects from Single 2D Images
abstract
Addresses the problem of building generic 3D models of structured objects on the basis of single 2D intensity images. In the context of the paper, generic modeling refers to the situation where analysis of the image information is performed on the sole basis of generic knowledge. That is, no a priori knowledge about the specific quantitative shape properties of the objects of interest is ever assumed. Moreover, images of interest are realistic. For instance, they may contain complex foreground 3D objects with textures and shadows, and a cluttered background. Objects are modeled by their constituent parts and connections. Therefore, a partly occluded object could be recognized from its model. Part models are based on geons, which are a set of qualitative generalized cylinders. An overview of the architecture of the modeling system is presented, along with the functionality of each subsystem and processing results.
Guillaume-Alexandre Bilodeau, Robert Bergevin
ICPR1