EDBT 2026 Demo / reviewers in the wild / expert
Pierre-Marc Jodoin
dblp:16/6189
· DBLP profile ↗
70ranked-venue papers
13as first author
17since 2021 · last 2026
0000-0002-6038-5753ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 12 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 16 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging rotational equivariance for reinforcement learning in tractographyabstractBrain tractography involves mapping diffusion-weighted images (DWI) onto streamlines representing neural fibre bundles. Recent research avenues have framed tractography into a reinforcement learning (RL) framework with actor-critic models. However, previous RL-based methods may compromise geometrical relations between the input (DWI) and output (tractogram). More specifically, 3D rotations applied to the input of RL-based tractography are not adequately reflected in the output, indicating a lack of SO(3) equivariance. This study aims to restore the equivariance present in previous non-learning-based methods (e.g., iFOD2 from MRtrix3) to RL-based tractography. To achieve this, we introduce SO(3) equivariant and invariant components for the actors (direction prediction model) and critics (Q-value prediction model), respectively. We employ an SE(3)-equivariant transformer as the next direction prediction function. The fact that both the input DWI and the output directional update can be represented as spherical signals that transform under representations of SO(3) makes this formulation a natural fit for the present problem. The contribution of this work is twofold. First, we discuss rotational equivariance in streamline tractography on a theoretical level. Second, we propose a method that combines RL-based tractography with a rotationally equivariant model. We evaluate the equivariance of the proposed method both locally and globally with phantom and in vivo data. The results show that the proposed method restores the equivariance of Track-to-Learn, which is the state-of-the-art for RL-based tractography. Our code is available at https://github.com/minnelab/SO3TrackToLearn. Fabian Leander Sinzinger, Antoine Théberge, Pierre-Marc Jodoin, Maxime Descoteaux, Rodrigo Moreno |
Medical Image Anal. | 3 |
| 2026 | BundleParc: Consistent white matter bundle parcellation without tractographyabstractTractometry, also known as tract profiling, is a powerful technique for probing microstructural properties along white matter (WM) tracts. A prerequisite for tractography-based tractometry is bundle parcellation—the subdivision of WM bundles into smaller segments where microstructural measures can be computed. However, existing parcellation methods lack consistency across bundles and timepoints, which reduces reproducibility and limits their utility for both longitudinal and cross-sectional studies. Moreover, these methods typically depend on tractography and bundle segmentation, two processes that are computationally expensive and often highly variable. In this work, we introduce BundleParc , a consistent and tractography-free bundle parcellation method. Instead of relying on streamline generation, BundleParc maps fiber orientation distribution function (fODF) volumes directly to label maps. Rigorous evaluation on research and clinical cohorts show that BundleParc is not only much simpler than state-of-the-art tract-based profiling methods, it is also consistently more accurate, robust and reproducible. With these results, BundleParc is a new solution for fast, easy-to-use, and off-the-shelf bundle segmentation and parcellation. • BundleParc produces parcellations for tractometry directly from FOD volumes. • BundleParc outperforms SOTA methods in accuracy, reproducibility in multiple cohorts. • BundleParc enables fast, reliable, robust and anatomically consistent parcellations. Antoine Théberge, Zineb El Yamani, Muhamed Barakovic, Stefano Magon, Joseph Yuan-Mou Yang, Maxime Descoteaux, François Rheault, Pierre-Marc Jodoin |
Medical Image Anal. | 8 |
| 2026 | Estimation of Segmental Longitudinal Strain in Transesophageal Echocardiography by Deep LearningabstractSegmental longitudinal strain (SLS) of the left ventricle (LV) is an important prognostic indicator for evaluating regional LV dysfunction, in particular for diagnosing and managing myocardial ischemia. Current techniques for strain estimation require significant manual intervention and expertise, limiting their efficiency and making them too resource-intensive for monitoring purposes. This study introduces the first automated pipeline, autoStrain, for SLS estimation in transesophageal echocardiography (TEE) using deep learning (DL) methods for motion estimation. We present a comparative analysis of two DL approaches: TeeFlow, based on the RAFT optical flow model for dense frame-to-frame predictions, and TeeTracker, based on the CoTracker point trajectory model for sparse long-sequence predictions. As ground truth motion data from real echocardiographic sequences are hardly accessible, we took advantage of a unique simulation pipeline (SIMUS) to generate a highly realistic synthetic TEE (synTEE) dataset of 80 patients with ground truth myocardial motion to train and evaluate both models. Our evaluation shows that TeeTracker outperforms TeeFlow in accuracy, achieving a mean distance error in motion estimation of 0.65 $\pm$ 0.20 mm on a synTEE test dataset. Clinical validation on 16 patients further demonstrated that SLS estimation with our autoStrain pipeline aligned with clinical references, achieving a mean difference (95% limits of agreement) of 1.09% (-8.90% to 11.09%). Incorporation of simulated ischemia in the synTEE data improved the accuracy of the models in quantifying abnormal deformation. Our findings indicate that integrating AI-driven motion estimation with TEE can significantly enhance the precision and efficiency of cardiac function assessment in clinical settings. Anders Austlid Taskén, Thierry Judge, Erik Andreas Rye Berg, Bjørnar Leangen Grenne, Frank Lindseth, Svend Aakhus, Pierre-Marc Jodoin, Nicolas Duchateau, Olivier Bernard 0001, Gabriel Kiss |
IEEE J. Biomed. Health Informatics | 8 |
| 2026 | Reinforcement Learning for Unsupervised Domain Adaptation in Spatio-Temporal Echocardiography SegmentationabstractDomain adaptation methods aim to bridge the gap between datasets by enabling knowledge transfer across domains, reducing the need for additional expert annotations. However, many approaches struggle with reliability in the target domain, an issue particularly critical in medical image segmentation, where accuracy and anatomical validity are essential. This challenge is further exacerbated in spatio-temporal data, where the lack of temporal consistency can significantly degrade segmentation quality, and particularly in echocardiography, where the presence of artifacts and noise can further hinder segmentation performance. To address these issues, we present RL4Seg3D, an unsupervised domain adaptation framework for 2D + time echocardiography segmentation. RL4Seg3D integrates novel reward functions and a fusion scheme to enhance key landmark precision in its segmentations while processing full-sized input videos. By leveraging reinforcement learning for image segmentation, our approach improves accuracy, anatomical validity, and temporal consistency while also providing, as a beneficial side effect, a robust uncertainty estimator, which can be used at test time to further enhance segmentation performance. We demonstrate the effectiveness of our framework on over 30,000 echocardiographic videos, showing that it outperforms standard domain adaptation techniques without the need for any labels on the target domain. Code is available at https://github.com/arnaudjudge/RL4Seg3D. Arnaud Judge, Nicolas Duchateau, Thierry Judge, Roman A. Sandler, Joseph Z. Sokol, Christian Desrosiers, Olivier Bernard 0001, Pierre-Marc Jodoin |
IEEE Trans. Medical Imaging | 8 |
| 2025 | Exploring the robustness of TractOracle methods in RL-based tractography
Jeremi Levesque, Antoine Théberge, Maxime Descoteaux, Pierre-Marc Jodoin |
Medical Image Anal. | 4 |
| 2024 | Domain Adaptation of Echocardiography Segmentation Via Reinforcement Learning
Arnaud Judge, Thierry Judge, Nicolas Duchateau, Roman A. Sandler, Joseph Z. Sokol, Olivier Bernard 0001, Pierre-Marc Jodoin |
MICCAI (9) | 7 |
| 2024 | TractOracle: Towards an Anatomically-Informed Reward Function for RL-Based Tractography
Antoine Théberge, Maxime Descoteaux, Pierre-Marc Jodoin |
MICCAI (2) | 3 |
| 2024 | What matters in reinforcement learning for tractography
Antoine Théberge, Christian Desrosiers, Arnaud Boré, Maxime Descoteaux, Pierre-Marc Jodoin |
Medical Image Anal. | 5 |
| 2023 | Asymmetric Contour Uncertainty Estimation for Medical Image Segmentation
Thierry Judge, Olivier Bernard 0001, Woo-Jin Cho Kim, Alberto Gómez 0002, Agisilaos Chartsias, Pierre-Marc Jodoin |
MICCAI (3) | 6 |
| 2023 | Generative Sampling in Bundle Tractography using Autoencoders (GESTA)
Jon Haitz Legarreta, Laurent Petit, Pierre-Marc Jodoin, Maxime Descoteaux |
Medical Image Anal. | 3 |
| 2022 | CRISP - Reliable Uncertainty Estimation for Medical Image Segmentation
Thierry Judge, Olivier Bernard 0001, Mihaela Porumb, Agisilaos Chartsias, Arian Beqiri, Pierre-Marc Jodoin |
MICCAI (8) | 6 |
| 2022 | ProstAttention-Net: A deep attention model for prostate cancer segmentation by aggressiveness in MRI scans
Audrey Duran, Gaspard Dussert, Olivier Rouvière, Tristan Jaouen, Pierre-Marc Jodoin, Carole Lartizien |
Medical Image Anal. | 5 |
| 2022 | Echocardiography Segmentation With Enforced Temporal ConsistencyabstractConvolutional neural networks (CNN) have demonstrated their ability to segment 2D cardiac ultrasound images. However, despite recent successes according to which the intra-observer variability on end-diastole and end-systole images has been reached, CNNs still struggle to leverage temporal information to provide accurate and temporally consistent segmentation maps across the whole cycle. Such consistency is required to accurately describe the cardiac function, a necessary step in diagnosing many cardiovascular diseases. In this paper, we propose a framework to learn the 2D+time apical long-axis cardiac shape such that the segmented sequences can benefit from temporal and anatomical consistency constraints. Our method is a post-processing that takes as input segmented echocardiographic sequences produced by any state-of-the-art method and processes it in two steps to (i) identify spatio-temporal inconsistencies according to the overall dynamics of the cardiac sequence and (ii) correct the inconsistencies. The identification and correction of cardiac inconsistencies relies on a constrained autoencoder trained to learn a physiologically interpretable embedding of cardiac shapes, where we can both detect and fix anomalies. We tested our framework on 98 full-cycle sequences from the CAMUS dataset, which are available alongside this paper. Our temporal regularization method not only improves the accuracy of the segmentation across the whole sequences, but also enforces temporal and anatomical consistency. Nathan Painchaud, Nicolas Duchateau, Olivier Bernard 0001, Pierre-Marc Jodoin |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Privacy Preserving for Medical Image Analysis via Non-Linear Deformation Proxy
Bach Ngoc Kim, Jose Dolz, Christian Desrosiers, Pierre-Marc Jodoin |
BMVC | 4 |
| 2021 | Filtering in tractography using autoencoders (FINTA)abstractCurrent brain white matter fiber tracking techniques show a number of problems, including: generating large proportions of streamlines that do not accurately describe the underlying anatomy; extracting streamlines that are not supported by the underlying diffusion signal; and under-representing some fiber populations, among others. In this paper, we describe a novel autoencoder-based learning method to filter streamlines from diffusion MRI tractography, and hence, to obtain more reliable tractograms. Our method, dubbed FINTA (Filtering in Tractography using Autoencoders) uses raw, unlabeled tractograms to train the autoencoder, and to learn a robust representation of brain streamlines. Such an embedding is then used to filter undesired streamline samples using a nearest neighbor algorithm. Our experiments on both synthetic and in vivo human brain diffusion MRI tractography data obtain accuracy scores exceeding the 90\% threshold on the test set. Results reveal that FINTA has a superior filtering performance compared to conventional, anatomy-based methods, and the RecoBundles state-of-the-art method. Additionally, we demonstrate that FINTA can be applied to partial tractograms without requiring changes to the framework. We also show that the proposed method generalizes well across different tracking methods and datasets, and shortens significantly the computation time for large (>1 M streamlines) tractograms. Together, this work brings forward a new deep learning framework in tractography based on autoencoders, which offers a flexible and powerful method for white matter filtering and bundling that could enhance tractometry and connectivity analyses. Jon Haitz Legarreta, Laurent Petit, François Rheault, Guillaume Theaud, Carl Lemaire, Maxime Descoteaux, Pierre-Marc Jodoin |
Medical Image Anal. | 7 |
| 2021 | Track-to-Learn: A general framework for tractography with deep reinforcement learning
Antoine Théberge, Christian Desrosiers, Maxime Descoteaux, Pierre-Marc Jodoin |
Medical Image Anal. | 4 |
| 2021 | Privacy-Net: An Adversarial Approach for Identity-Obfuscated Segmentation of Medical ImagesabstractThis paper presents a client/server privacy-preserving network in the context of multicentric medical image analysis. Our approach is based on adversarial learning which encodes images to obfuscate the patient identity while preserving enough information for a target task. Our novel architecture is composed of three components: 1) an encoder network which removes identity-specific features from input medical images, 2) a discriminator network that attempts to identify the subject from the encoded images, 3) a medical image analysis network which analyzes the content of the encoded images (segmentation in our case). By simultaneously fooling the discriminator and optimizing the medical analysis network, the encoder learns to remove privacy-specific features while keeping those essentials for the target task. Our approach is illustrated on the problem of segmenting brain MRI from the large-scale Parkinson Progression Marker Initiative (PPMI) dataset. Using longitudinal data from PPMI, we show that the discriminator learns to heavily distort input images while allowing for highly accurate segmentation results. Our results also demonstrate that an encoder trained on the PPMI dataset can be used for segmenting other datasets, without the need for retraining. The code is made available at: https://github.com/bachkimn/Privacy-Net-An-Adversarial-Approach-forIdentity-Obfuscated-Segmentation-of-MedicalImages. Bach Ngoc Kim, Jose Dolz, Pierre-Marc Jodoin, Christian Desrosiers |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Cardiac Segmentation With Strong Anatomical GuaranteesabstractConvolutional neural networks (CNN) have had unprecedented success in medical imaging and, in particular, in medical image segmentation. However, despite the fact that segmentation results are closer than ever to the inter-expert variability, CNNs are not immune to producing anatomically inaccurate segmentations, even when built upon a shape prior. In this paper, we present a framework for producing cardiac image segmentation maps that are guaranteed to respect pre-defined anatomical criteria, while remaining within the inter-expert variability. The idea behind our method is to use a well-trained CNN, have it process cardiac images, identify the anatomically implausible results and warp these results toward the closest anatomically valid cardiac shape. This warping procedure is carried out with a constrained variational autoencoder (cVAE) trained to learn a representation of valid cardiac shapes through a smooth, yet constrained, latent space. With this cVAE, we can project any implausible shape into the cardiac latent space and steer it toward the closest correct shape. We tested our framework on short-axis MRI as well as apical two and four-chamber view ultrasound images, two modalities for which cardiac shapes are drastically different. With our method, CNNs can now produce results that are both within the inter-expert variability and always anatomically plausible without having to rely on a shape prior. Nathan Painchaud, Youssef Skandarani, Thierry Judge, Olivier Bernard 0001, Alain Lalande, Pierre-Marc Jodoin |
IEEE Trans. Medical Imaging | 6 |
| 2019 | Spectral Metric for Dataset Complexity AssessmentabstractIn this paper, we propose a new measure to gauge the complexity of image classification problems. Given an annotated image dataset, our method computes a complexity measure called the cumulative spectral gradient (CSG) which strongly correlates with the test accuracy of convolutional neural networks (CNN). The CSG measure is derived from the probabilistic divergence between classes in a spectral clustering framework. We show that this metric correlates with the overall separability of the dataset and thus its inherent complexity. As will be shown, our metric can be used for dataset reduction, to assess which classes are more difficult to disentangle, and approximate the accuracy one could expect to get with a CNN. Results obtained on 11 datasets and three CNN models reveal that our method is more accurate and faster than previous complexity measures. Frederic Branchaud-Charron, Andrew Achkar, Pierre-Marc Jodoin |
CVPR | 3 |
| 2019 | Structured Pruning of Neural Networks With Budget-Aware RegularizationabstractPruning methods have shown to be effective at reducing the size of deep neural networks while keeping accuracy almost intact. Among the most effective methods are those that prune a network while training it with a sparsity prior loss and learnable dropout parameters. A shortcoming of these approaches however is that neither the size nor the inference speed of the pruned network can be controlled directly; yet this is a key feature for targeting deployment of CNNs on low-power hardware. To overcome this, we introduce a budgeted regularized pruning framework for deep CNNs. Our approach naturally fits into traditional neural network training as it consists of a learnable masking layer, a novel budget-aware objective function, and the use of knowledge distillation. We also provide insights on how to prune a residual network and how this can lead to new architectures. Experimental results reveal that CNNs pruned with our method are more accurate and less compute-hungry than state-of-the-art methods. Also, our approach is more effective at preventing accuracy collapse in case of severe pruning; this allows pruning factors of up to 16× without significant accuracy drop. Carl Lemaire, Andrew Achkar, Pierre-Marc Jodoin |
CVPR | 3 |
| 2019 | Cardiac MRI Segmentation with Strong Anatomical Guarantees
Nathan Painchaud, Youssef Skandarani, Thierry Judge, Olivier Bernard 0001, Alain Lalande, Pierre-Marc Jodoin |
MICCAI (2) | 6 |
| 2019 | Convolutional Neural Network With Shape Prior Applied to Cardiac MRI SegmentationabstractIn this paper, we present a novel convolutional neural network architecture to segment images from a series of short-axis cardiac magnetic resonance slices (CMRI). The proposed model is an extension of the U-net that embeds a cardiac shape prior and involves a loss function tailored to the cardiac anatomy. Since the shape prior is computed offline only once, the execution of our model is not limited by its calculation. Our system takes as input raw magnetic resonance images, requires no manual preprocessing or image cropping and is trained to segment the endocardium and epicardium of the left ventricle, the endocardium of the right ventricle, as well as the center of the left ventricle. With its multiresolution grid architecture, the network learns both high and low-level features useful to register the shape prior as well as accurately localize the borders of the cardiac regions. Experimental results obtained on the Automatic Cardiac Diagnostic Challenge - Medical Image Computing and Computer Assisted Intervention (ACDC-MICCAI) 2017 dataset show that our model segments multislices CMRI (left and right ventricle contours) in 0.18 s with an average Dice coefficient of [Formula: see text] and an average 3-D Hausdorff distance of [Formula: see text] mm. Clément Zotti, Zhiming Luo, Alain Lalande, Pierre-Marc Jodoin |
IEEE J. Biomed. Health Informatics | 4 |
| 2019 | Deep Learning for Segmentation Using an Open Large-Scale Dataset in 2D EchocardiographyabstractDelineation of the cardiac structures from 2D echocardiographic images is a common clinical task to establish a diagnosis. Over the past decades, the automation of this task has been the subject of intense research. In this paper, we evaluate how far the state-of-the-art encoder-decoder deep convolutional neural network methods can go at assessing 2D echocardiographic images, i.e., segmenting cardiac structures and estimating clinical indices, on a dataset, especially, designed to answer this objective. We, therefore, introduce the cardiac acquisitions for multi-structure ultrasound segmentation dataset, the largest publicly-available and fully-annotated dataset for the purpose of echocardiographic assessment. The dataset contains two and four-chamber acquisitions from 500 patients with reference measurements from one cardiologist on the full dataset and from three cardiologists on a fold of 50 patients. Results show that encoder-decoder-based architectures outperform state-of-the-art non-deep learning methods and faithfully reproduce the expert analysis for the end-diastolic and end-systolic left ventricular volumes, with a mean correlation of 0.95 and an absolute mean error of 9.5 ml. Concerning the ejection fraction of the left ventricle, results are more contrasted with a mean correlation coefficient of 0.80 and an absolute mean error of 5.6%. Although these results are below the inter-observer scores, they remain slightly worse than the intra-observer's ones. Based on this observation, areas for improvement are defined, which open the door for accurate and fully-automatic analysis of 2D echocardiographic images. Sarah Leclerc, Erik Smistad, João Pedrosa, Andreas Østvik, Frederic Cervenansky, Florian Espinosa, Torvald Espeland, Erik Andreas Rye Berg, Pierre-Marc Jodoin, Thomas Grenier, Carole Lartizien, Jan D'hooge, Lasse Løvstakken, Olivier Bernard 0001 |
IEEE Trans. Medical Imaging | 9 |
| 2018 | Traffic Analytics With Low-Frame-Rate VideosabstractIn this paper, we investigate the possibility of monitoring highway traffic based on videos whose frame rate is too low to accurately estimate motion features. The goal of the proposed method is to recognize traffic conditions instead of measuring them, as is usually the case. The main advantage of our approach comes from its ability to process low-frame-rate videos for which motion features cannot be estimated. Our method takes advantage of the highly redundant nature of traffic scenes that are pictured from a top-down perspective showing vehicles on a predominant asphalted road surrounded by background objects. Due to the limited variety of objects pictured in traffic scenes, our method gets to learn features that are specific to such images. With these features, our method is able to segment traffic images, classify traffic scenes, and estimate traffic density without requiring motion features. Different convolutional neural network models are proposed to segment traffic images in three different classes (Road, Car, and Background), classify traffic images into different categories (Empty, Fluid, Heavy, and Jam), and predict traffic density. We also propose a procedure to perform transfer learning of any of these models to new traffic scenes. Zhiming Luo, Pierre-Marc Jodoin, Songzhi Su, Shaozi Li, Hugo Larochelle |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | MIO-TCD: A New Benchmark Dataset for Vehicle Classification and LocalizationabstractThe ability to train on a large dataset of labeled samples is critical to the success of deep learning in many domains. In this paper, we focus on motor vehicle classification and localization from a single video frame and introduce the "MIOvision Traffic Camera Dataset" (MIO-TCD) in this context. MIO-TCD is the largest dataset for motorized traffic analysis to date. It includes 11 traffic object classes such as cars, trucks, buses, motorcycles, bicycles, pedestrians. It contains 786,702 annotated images acquired at different times of the day and different periods of the year by hundreds of traffic surveillance cameras deployed across Canada and the United States. The dataset consists of two parts: a "localization dataset", containing 137,743 full video frames with bounding boxes around traffic objects, and a "classification dataset", containing 648,959 crops of traffic objects from the 11 classes. We also report results from the 2017 CVPR MIO-TCD Challenge, that leveraged this dataset, and compare them with results for state-of-the-art deep learning architectures. These results demonstrate the viability of deep learning methods for vehicle localization and classification from a single video frame in real-life traffic scenarios. The topperforming methods achieve both accuracy and Kappa score above 96% on the classification dataset and mean-average precision of 77% on the localization dataset. We also identify scenarios in which state-of-the-art methods still fail and we suggest avenues to address these challenges. Both the dataset and detailed results are publicly available on-line [1]. Zhiming Luo, Frederic Branchaud-Charron, Carl Lemaire, Janusz Konrad, Shaozi Li, Akshaya Mishra, Andrew Achkar, Justin A. Eichel, Pierre-Marc Jodoin |
IEEE Trans. Image Process. | 9 |
| 2018 | Deep Learning Techniques for Automatic MRI Cardiac Multi-Structures Segmentation and Diagnosis: Is the Problem Solved?abstractDelineation of the left ventricular cavity, myocardium, and right ventricle from cardiac magnetic resonance images (multi-slice 2-D cine MRI) is a common clinical task to establish diagnosis. The automation of the corresponding tasks has thus been the subject of intense research over the past decades. In this paper, we introduce the "Automatic Cardiac Diagnosis Challenge" dataset (ACDC), the largest publicly available and fully annotated dataset for the purpose of cardiac MRI (CMR) assessment. The dataset contains data from 150 multi-equipments CMRI recordings with reference measurements and classification from two medical experts. The overarching objective of this paper is to measure how far state-of-the-art deep learning methods can go at assessing CMRI, i.e., segmenting the myocardium and the two ventricles as well as classifying pathologies. In the wake of the 2017 MICCAI-ACDC challenge, we report results from deep learning methods provided by nine research groups for the segmentation task and four groups for the classification task. Results show that the best methods faithfully reproduce the expert analysis, leading to a mean value of 0.97 correlation score for the automatic extraction of clinical indices and an accuracy of 0.96 for automatic diagnosis. These results clearly open the door to highly accurate and fully automatic analysis of cardiac CMRI. We also identify scenarios for which deep learning methods are still failing. Both the dataset and detailed results are publicly available online, while the platform will remain open for new submissions. Olivier Bernard 0001, Alain Lalande, Clément Zotti, Frederic Cervenansky, Xin Yang 0009, Pheng-Ann Heng, Irem Cetin, Karim Lekadir, Oscar Camara 0001, Miguel Ángel González Ballester, Gerard Sanroma, Sandy Napel, Steffen E. Petersen, Georgios Tziritas, Ilias Grinias, Mahendra Khened, Alex Varghese, Ganapathy Krishnamurthi, Marc-Michel Rohé, Xavier Pennec, Maxime Sermesant, Fabian Isensee, Paul F. Jaeger, Klaus H. Maier-Hein, Peter M. Full, Ivo Wolf, Sandy Engelhardt, Christian F. Baumgartner, Lisa M. Koch, Jelmer M. Wolterink, Ivana Isgum, Yeonggul Jang, Yoonmi Hong, Jay Patravali, Shubham Jain 0006, Olivier Humbert, Pierre-Marc Jodoin |
IEEE Trans. Medical Imaging | 37 |
| 2017 | Non-local Deep Features for Salient Object DetectionabstractSaliency detection aims to highlight the most relevant objects in an image. Methods using conventional models struggle whenever salient objects are pictured on top of a cluttered background while deep neural nets suffer from excess complexity and slow evaluation speeds. In this paper, we propose a simplified convolutional neural network which combines local and global information through a multi-resolution 4×5 grid structure. Instead of enforcing spacial coherence with a CRF or superpixels as is usually the case, we implemented a loss function inspired by the Mumford-Shah functional which penalizes errors on the boundary. We trained our model on the MSRA-B dataset, and tested it on six different saliency benchmark datasets. Results show that our method is on par with the state-of-the-art while reducing computation time by a factor of 18 to 100 times, enabling near real-time, high performance saliency detection. Zhiming Luo, Akshaya Kumar Mishra, Andrew Achkar, Justin A. Eichel, Shaozi Li, Pierre-Marc Jodoin |
CVPR | 6 |
| 2017 | Brain tumor segmentation with Deep Neural Networks
Mohammad Havaei, Axel Davy, David Warde-Farley, Antoine Biard, Aaron C. Courville, Yoshua Bengio, Christopher Joseph Pal, Pierre-Marc Jodoin, Hugo Larochelle |
Medical Image Anal. | 8 |
| 2017 | ISLES 2015 - A public evaluation benchmark for ischemic stroke lesion segmentation from multispectral MRI
Oskar Maier, Bjoern Menze, Janina von der Gablentz, Levin Häni, Mattias P. Heinrich, Matthias Liebrand, Stefan Winzeck, Abdul Basit 0007, Paul Bentley, Liang Chen 0018, Daan Christiaens, Francis Dutil, Karl Egger, Chaolu Feng, Ben Glocker, Michael Götz, Tom Haeck, Hanna-Leena Halme, Mohammad Havaei, Khan M. Iftekharuddin, Pierre-Marc Jodoin |
Medical Image Anal. | 21 |
| 2017 | Interactive deep learning method for segmenting moving objects
Yi Wang 0025, Zhiming Luo, Pierre-Marc Jodoin |
Pattern Recognit. Lett. | 3 |
| 2017 | Improving pedestrian detection using motion-guided filtering
Yi Wang 0025, Sébastien Piérard, Songzhi Su, Pierre-Marc Jodoin |
Pattern Recognit. Lett. | 4 |
| 2017 | Extensive Benchmark and Survey of Modeling Methods for Scene Background InitializationabstractScene background initialization is the process by which a method tries to recover the background image of a video without foreground objects in it. Having a clear understanding about which approach is more robust and/or more suited to a given scenario is of great interest to many end users or practitioners. The aim of this paper is to provide an extensive survey of scene background initialization methods as well as a novel benchmarking framework. The proposed framework involves several evaluation metrics and state-of-the-art methods, as well as the largest video data set ever made for this purpose. The data set consists of several camera-captured videos that: 1) span categories focused on various background initialization challenges; 2) are obtained with different cameras of different lengths, frame rates, spatial resolutions, lighting conditions, and levels of compression; and 3) contain indoor and outdoor scenes. The wide variety of our data set prevents our analysis from favoring a certain family of background initialization methods over others. Our evaluation framework allows us to quantitatively identify solved and unsolved issues related to scene background initialization. We also identify scenarios for which state-of-the-art methods systematically fail. Pierre-Marc Jodoin, Lucia Maddalena 0001, Alfredo Petrosino, Yi Wang 0025 |
IEEE Trans. Image Process. | 1 |
| 2016 | CBDF: Compressed Binary Discriminative Feature
Li-Chuan Geng, Pierre-Marc Jodoin, Songzhi Su, Shaozi Li |
Neurocomputing | 2 |
| 2016 | Retrieval in Long-Surveillance Videos Using User-Described Motion and Object AttributesabstractWe present a content-based retrieval method for long-surveillance videos in wide-area (airborne) and near-field [closed-circuit television (CCTV)] imagery. Our goal is to retrieve video segments, with a focus on detecting objects moving on routes, that match user-defined events of interest. The sheer size and remote locations where surveillance videos are acquired necessitates highly compressed representations that are also meaningful for supporting user-defined queries. To address these challenges, we archive long-surveillance video through lightweight processing based on low-level local spatiotemporal extraction of motion and object 2. These are then hashed into an inverted index using locality-sensitive hashing. This local approach allows for query flexibility and leads to significant gains in compression. Our second task is to extract partial matches to user-created queries and assemble them into full matches using dynamic programming (DP). DP assembles the indexed low-level features into a video segment that matches the query route by exploiting causality. We examine CCTV and airborne footage, whose low contrast makes motion extraction more difficult. We generate robust motion estimates for airborne data using a tracklets generation algorithm, while we use the Horn and Schunck approach to generate motion estimates for CCTV. Our approach handles long routes, low contrasts, and occlusion. We derive bounds on the rate of false positives and demonstrate the effectiveness of the approach for counting, motion pattern recognition, and abandoned object applications. Greg Castañón, Mohamed A. Elgharib, Venkatesh Saligrama, Pierre-Marc Jodoin |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Standardized Evaluation System for Left Ventricular Segmentation Algorithms in 3D EchocardiographyabstractReal-time 3D Echocardiography (RT3DE) has been proven to be an accurate tool for left ventricular (LV) volume assessment. However, identification of the LV endocardium remains a challenging task, mainly because of the low tissue/blood contrast of the images combined with typical artifacts. Several semi and fully automatic algorithms have been proposed for segmenting the endocardium in RT3DE data in order to extract relevant clinical indices, but a systematic and fair comparison between such methods has so far been impossible due to the lack of a publicly available common database. Here, we introduce a standardized evaluation framework to reliably evaluate and compare the performance of the algorithms developed to segment the LV border in RT3DE. A database consisting of 45 multivendor cardiac ultrasound recordings acquired at different centers with corresponding reference measurements from three experts are made available. The algorithms from nine research groups were quantitatively evaluated and compared using the proposed online platform. The results showed that the best methods produce promising results with respect to the experts' measurements for the extraction of clinical indices, and that they offer good segmentation precision in terms of mean distance error in the context of the experts' variability range. The platform remains open for new submissions. Olivier Bernard 0001, Johan G. Bosch, Brecht Heyde, Martino Alessandrini, Daniel Barbosa 0001, Sorina Camarasu-Pop, Frederic Cervenansky, Sébastien Valette, Oana Mirea, Michaël Bernier, Pierre-Marc Jodoin, Jaime Santo Domingos, Richard V. Stebbing, Kevin Keraudren, Ozan Oktay, Jose Caballero, Daniel Rueckert, Fausto Milletari, Seyed-Ahmad Ahmadi, Erik Smistad, Frank Lindseth, Maartje van Stralen, Örjan Smedby, Erwan Donal, Mark Monaghan, Alex Papachristidis, Marcel L. Geleijnse, Elena Galli, Jan D'hooge |
IEEE Trans. Medical Imaging | 11 |
| 2015 | Traffic analysis without motion featuresabstractIn this paper, we investigate the possibility of monitoring traffic without using any motion features. The goal of our system is to process videos with ultra-low frame rate, i.e. videos for which reliable motion features cannot be computed. In this work, we investigate how 2D spatial features combined with a machine learning method can assess traffic conditions such as fluid traffic, dense traffic, and traffic jam. The underlying hypothesis that we ought to validate is that traffic images are heavily characterized by their 2D spatial textures. In that perspective, we tested different 2D texture features and machine learning methods to see how accurate such an approach can be. We also performed a regression on the image descriptor in order to estimate traffic density. Experimental results obtained on the UCSD traffic dataset reveal that our approach generalizes well to various weather and lighting conditions. It even outperforms state-of-the-art traffic analysis methods relying on spatio-temporal features. Zhiming Luo, Pierre-Marc Jodoin, Shaozi Li, Songzhi Su |
ICIP | 2 |
| 2015 | High-speed transition patterns for video projection, 3D reconstruction, and copyright protection
Jonathan Boisvert, Marc-Antoine Drouin, Pierre-Marc Jodoin |
Pattern Recognit. | 3 |
| 2015 | Machine learning and pattern recognition models in change detection
Djamel Bouchaffra, Mohamed Cheriet, Pierre-Marc Jodoin, Diane M. Beck |
Pattern Recognit. | 3 |
| 2015 | Signal processing and learning methods for 3D semantic analysis
Yue Gao 0002, Rongrong Ji, Xinbo Gao 0001, Pierre-Marc Jodoin |
Signal Process. | 4 |
| 2015 | Novel Graph Cuts Method for Multi-Frame Super-ResolutionabstractIn this letter, we propose a new graph cuts multi-frame super resolution method. The method is carried out in 3 steps. First, we project each high-resolution pixel p onto the low-resolution images and select low-resolution pixels which fall within the zone of influence of p. Second, we weigh the contribution of the low-resolution pixels via a soft switching function and add them to construct a virtual low resolution pixel. The high resolution image is then recovered after minimizing a Maximum a posteriori Markov Random Field (MAP-MRF) energy function. This is done by approximating our energy function to make it graph representable and minimize it with a graph cuts α-expansion algorithm. Experimental results show that our approach outperforms state-of-the-art methods. Dongxiao Zhang, Pierre-Marc Jodoin, Cuihua Li, Yun-Dong Wu, Guo-Rong Cai |
IEEE Signal Process. Lett. | 2 |
| 2014 | Lying-pose detection with training dataset expansionabstractWe propose a rotation and scale invariant method to locate people lying on the ground. Unlike conventional human-shape detection methods which assume that all human shapes are in upright position, a person lying on the ground can have arbitrary orientation and pose. Accounting for every possible body configuration would thus require a huge training dataset that would be challenging to gather. In this paper, we propose a method which increases the size of a small training dataset and allows to detect multiple body poses. To do so, our method increases the size of the dataset with a geometric distortion method followed by a rejection sampling method. Then, it automatically identifies K body configurations in the training set, realign it in upright position and trains K SVM classifiers, one for each body configuration. Lying pose detection is then performed by considering a max pooling strategy across all K SVM classifiers. Daoxun Xia, Songzhi Su, Shaozi Li, Pierre-Marc Jodoin |
ICIP | 4 |
| 2014 | Efficient Interactive Brain Tumor Segmentation as Within-Brain kNN ClassificationabstractWe consider the problem of brain tumor segmentation from magnetic resonance (MR) images. This task is most frequently tackled using machine learning methods that generalize across brains, by learning from training brain images in order to generalize to novel test brains. However this approach faces many obstacles that threaten its performance, such as the ability to properly perform multi-brain registration or brain-atlas alignment, or to extract appropriate high-dimensional features that support good generalization. These operations are both nontrivial and time-consuming, limiting the practicality of these approaches in a clinical context. In this paper, we propose to side step these issues by approaching the problem as one of within brain generalization. Specifically, we propose a semi-automatic method that segments a given brain by training and generalizing within that brain only, based on some minimum user interaction. We investigate how k nearest neighbors (kNN), arguably the simplest machine learning method available, combined with the simplest feature vector possible (raw MR signal + (x,y,z) position) can be combined into a method that is both simple, accurate and fast. Results obtained on the online BRATS dataset reveal that our method is fast and second best in terms of the complete and core test set tumor segmentation. Mohammad Havaei, Pierre-Marc Jodoin, Hugo Larochelle |
ICPR | 2 |
| 2014 | A Novel Video Dataset for Change Detection BenchmarkingabstractChange detection is one of the most commonly encountered low-level tasks in computer vision and video processing. A plethora of algorithms have been developed to date, yet no widely accepted, realistic, large-scale video data set exists for benchmarking different methods. Presented here is a unique change detection video data set consisting of nearly 90 000 frames in 31 video sequences representing six categories selected to cover a wide range of challenges in two modalities (color and thermal infrared). A distinguishing characteristic of this benchmark video data set is that each frame is meticulously annotated by hand for ground-truth foreground, background, and shadow area boundaries-an effort that goes much beyond a simple binary label denoting the presence of change. This enables objective and precise quantitative comparison and ranking of video-based change detection algorithms. This paper discusses various aspects of the new data set, quantitative performance metrics used, and comparative results for over two dozen change detection algorithms. It draws important conclusions on solved and remaining issues in change detection, and describes future challenges for the scientific community. The data set, evaluation tools, and algorithm rankings are available to the public on a website and will be updated with feedback from academia and industry in the future. Nil Goyette, Pierre-Marc Jodoin, Fatih Porikli, Janusz Konrad, Prakash Ishwar |
IEEE Trans. Image Process. | 2 |
| 2013 | Meta-tracking for video scene understandingabstractThis paper presents a novel method to extract dominant motion patterns (MPs) and the main entry/exit areas from a surveillance video. The method first computes motion histograms for each pixel and then converts it into orientation distribution functions (ODFs). Given these ODFs, a novel particle meta-tracking procedure is launched which produces meta-tracks, i.e. particle trajectories. As opposed to conventional tracking which focuses on individual moving objects, meta-tracking uses particles to follow the dominant flow of the traffic. In a last step, a novel method is used to simultaneously identify the main entry/exit areas and recover the predominant MPs. The meta-tracking procedure is a unique way to connect low-level motion features to long-range MPs. This kind of tracking is inspired by brain fiber tractography which has long been used to find dominant connections in the brain. Our method is fast, simple to implement, and works both on sparse and extremely crowded scenes. It also works on highly structured scenes (highways, traffic-light corners, etc.) as well as on chaotic scenes. Pierre-Marc Jodoin, Yannick Benezeth, Yi Wang 0025 |
AVSS | 1 |
| 2013 | Perspective-SIFT: An efficient tool for low-altitude remote sensing image registration
Guo-Rong Cai, Pierre-Marc Jodoin, Shaozi Li, Yun-Dong Wu, Songzhi Su, Zhenkun Huang |
Signal Process. | 2 |
| 2012 | Real-Time Activity Search of Surveillance VideoabstractWe present a fast and flexible content-based retrieval method for surveillance video. Designing a video search robust to uncertain activity duration, high variability in object shapes and scene content is challenging. We propose a two-step approach to video search. First, local motion features are inserted into an inverted index using locality-sensitive hashing (LSH). Second, we utilize a novel optimization approach based on edit distance to minimize temporal distortion, limited obscuration and imperfect queries. This approach assembles the local features stored in the index into a video segment which matches the query video. Pre-processing of archival video is performed in real-time, and retrieval speed scales as a function of the number of matches rather than video length. We demonstrate the effectiveness of the approach for counting, motion pattern recognition and abandoned object applications using a pair of challenging video datasets. Greg Castañón, Venkatesh Saligrama, André-Louis Caron, Pierre-Marc Jodoin |
AVSS | 4 |
| 2012 | Exploratory search of long surveillance videosabstractWe present a fast and flexible content-based retrieval method for surveillance video. Designing a video search robust to uncertain activity duration, high variability in object shapes and scene content is challenging. We propose a two-step approach to video search. First, local features are inserted into an inverted index using locality-sensitive hashing (LSH). Second, we utilize a novel dynamic programming (DP) approach to robustify against temporal distortion, limited obscuration and imperfect queries. DP exploits causality to assemble the local features stored in the index into a video segment which matches the query video. Pre-processing of archival video is performed in real-time, and retrieval speed scales as a function of the number of matches rather than video length. We derive bounds on the rate of false positives, demonstrate the effectiveness of the approach for counting, motion pattern recognition and abandoned object applications using seven challenging video datasets and compare with existing work. Greg Castañón, André-Louis Caron, Venkatesh Saligrama, Pierre-Marc Jodoin |
ACM Multimedia | 4 |
| 2012 | Camera-projector matching using unstructured video
Marc-Antoine Drouin, Pierre-Marc Jodoin, Julien Prémont |
Mach. Vis. Appl. | 2 |
| 2012 | Behavior SubtractionabstractBackground subtraction has been a driving engine for many computer vision and video analytics tasks. Although its many variants exist, they all share the underlying assumption that photometric scene properties are either static or exhibit temporal stationarity. While this works in many applications, the model fails when one is interested in discovering changes in scene dynamics instead of changes in scene's photometric properties; the detection of unusual pedestrian or motor traffic patterns are but two examples. We propose a new model and computational framework that assume the dynamics of a scene, not its photometry, to be stationary, i.e., a dynamic background serves as the reference for the dynamics of an observed scene. Central to our approach is the concept of an event, which we define as short-term scene dynamics captured over a time window at a specific spatial location in the camera field of view. Unlike in our earlier work, we compute events by time-aggregating vector object descriptors that can combine multiple features, such as object size, direction of movement, speed, etc. We characterize events probabilistically, but use low-memory, low-complexity surrogates in a practical implementation. Using these surrogates amounts to behavior subtraction, a new algorithm for effective and efficient temporal anomaly detection and localization. Behavior subtraction is resilient to spurious background motion, such as due to camera jitter, and is content-blind, i.e., it works equally well on humans, cars, animals, and other objects in both uncluttered and highly cluttered scenes. Clearly, treating video as a collection of events rather than colored pixels opens new possibilities for video analytics. Pierre-Marc Jodoin, Venkatesh Saligrama, Janusz Konrad |
IEEE Trans. Image Process. | 1 |
| 2011 | Abnormality detection using low-level co-occurring events
Yannick Benezeth, Pierre-Marc Jodoin, Venkatesh Saligrama |
Pattern Recognit. Lett. | 2 |
| 2011 | Image Multidistortion EstimationabstractWe present a method for estimating the amount of noise and blur in a distorted image. Our method is based on the multiscale structural similarity (MS-SSIM) framework that, although designed to measure image quality, is used to estimate the amount of blur and noise in a degraded image given a reference image. We show that there exists a bijective mapping between the 2-D noise/blur space and the 3-D MS-SSIM space, which allows to recover distortion parameters. That mapping allows to formulate the multidistortion-estimation problem as a classical optimization problem. Various search strategies such as Newton, simplex, NewUOA, and brute-force search are presented and compared. We also show that a bicubic patch can be used to approximate the bijective mapping between the noise/blur space and the 3-D MS-SSIM space. Interestingly, the use of such a patch reduces the processing time by a factor of 40 without significantly reducing precision. Based on quantitative results, we show that the amount of different types of blur and noise in a distorted image can be recovered with accuracy of roughly 2% and 8%, respectively. Our methods are compared with four state-of-the-art noise- and blur-estimation techniques. André-Louis Caron, Pierre-Marc Jodoin |
IEEE Trans. Image Process. | 2 |
| 2010 | Human Detection with a Multi-sensors Stereovision System
Yannick Benezeth, Pierre-Marc Jodoin, Bruno Emile, Hélène Laurent, Christophe Rosenberger |
ICISP | 2 |
| 2010 | Search Strategies for Image Multi-distortion EstimationabstractIn this paper, we present a method for estimating the amount of Gaussian noise and Gaussian blur in a distorted image. Our method is based on the MS-SSIM framework which, although designed to measure image quality, is used to estimate the amount of blur and noise in a degraded image given a reference image. Various search strategies such as Newton, Simplex, and brute force search are presented and rigorously compared. Based on quantitative results, we show that the amount of blur and noise in a distorted image can be recovered with an accuracy up to 0.95% and 5.40%, respectively. To our knowledge, such precision has never been achieved before. André-Louis Caron, Pierre-Marc Jodoin, Christophe Charrier |
ICPR | 2 |
| 2010 | Activity Based Matching in Distributed Camera NetworksabstractIn this paper, we consider the problem of finding correspondences between distributed cameras that have partially overlapping field of views. When multiple cameras with adaptable orientations and zooms are deployed, as in many wide area surveillance applications, identifying correspondence between different activities becomes a fundamental issue. We propose a correspondence method based upon activity features that, unlike photometric features, have certain geometry independence properties. The proposed method is robust to pose, illumination and geometric effects, unsupervised (does not require any calibration objects). In addition, these features are amenable to low communication bandwidth and distributed network applications. We present quantitative and qualitative results with synthetic and real life examples, and compare the proposed method with scale invariant feature transform (SIFT) based method. We show that our method significantly outperforms the SIFT method when cameras have significantly different orientations. We then describe extensions of our method in a number of directions including topology reconstruction, camera calibration, and distributed anomaly detection. Erhan Baki Ermis, Pierre Clarot, Pierre-Marc Jodoin, Venkatesh Saligrama |
IEEE Trans. Image Process. | 3 |
| 2009 | Abnormal events detection based on spatio-temporal co-occurencesabstractWe explore a location based approach for behavior modeling and abnormality detection. In contrast to the conventional object based approach where an object may first be tagged, identified, classified, and tracked, we proceed directly with event characterization and behavior modeling at the pixel(s) level based on motion labels obtained from background subtraction. Since events are temporally and spatially dependent, this calls for techniques that account for statistics of spatiotemporal events. Based on motion labels, we learn co-occurrence statistics for normal events across space-time. For one (or many) key pixel(s), we estimate a co-occurrence matrix that accounts for any two active labels which co-occur simultaneously within the same spatiotemporal volume. This co-occurrence matrix is then used as a potential function in a Markov random field (MRF) model to describe the probability of observations within the same spatiotemporal volume. The MRF distribution implicitly accounts for speed, direction, as well as the average size of the objects passing in front of each key pixel. Furthermore, when the spatiotemporal volume is large enough, the co-occurrence distribution contains the average normal path followed by moving objects. The learned normal co-occurrence distribution can be used for abnormal detection. Our method has been tested on various outdoor videos representing various challenges. Yannick Benezeth, Pierre-Marc Jodoin, Venkatesh Saligrama, Christophe Rosenberger |
CVPR | 2 |
| 2009 | Optical-flow based on an edge-avoidance procedure
Pierre-Marc Jodoin, Max Mignotte |
Comput. Vis. Image Underst. | 1 |
| 2009 | Foreground-Adaptive Background SubtractionabstractBackground subtraction is a powerful mechanism for detecting change in a sequence of images that finds many applications. The most successful background subtraction methods apply probabilistic models to background intensities evolving in time; nonparametric and mixture-of-Gaussians models are but two examples. The main difficulty in designing a robust background subtraction algorithm is the selection of a detection threshold. In this paper, we adapt this threshold to varying video statistics by means of two statistical models. In addition to a nonparametric background model, we introduce a foreground model based on small spatial neighborhood to improve discrimination sensitivity. We also apply a Markov model to change labels to improve spatial coherence of the detections. The proposed methodology is applicable to other background models as well. J. Mike McHugh, Janusz Konrad, Venkatesh Saligrama, Pierre-Marc Jodoin |
IEEE Signal Process. Lett. | 4 |
| 2008 | Motion segmentation and abnormal behavior detection via behavior clusteringabstractWe consider a change detection problem in video surveillance applications and propose busy-idle rates, meaningful and easy to compute features, to characterize the behavior profile of a given pixel. We describe the geometry independence property of these features, and use them to model the typical behavior that is observed in training sequences. Using a small number of samples for each pixel we generate behavior clusters, wherein pixels with similar behavior profiles fall into the same cluster. We then generate probabilistic models corresponding to behavior clusters, and use these models to perform abnormal behavior detection. Erhan Baki Ermis, Venkatesh Saligrama, Pierre-Marc Jodoin, Janusz Konrad |
ICIP | 3 |
| 2008 | Motion detection with an unstable cameraabstractFast and accurate motion detection in the presence of camera jitter is known to be a difficult problem. Existing statistical methods often produce abundant false positives since jitter-induced motion is difficult to differentiate from scene-induced motion. Although frame alignment by means of camera motion compensation can help resolve such ambiguities, the additional steps of motion estimation and compensation increase the complexity of the overall algorithm. In this paper, we address camera jitter by applying background subtraction to scene dynamics instead of scene photometry. In our method, an object is assumed moving if its dynamical behavior is different from the average dynamics observed in a reference sequence. Our method is conceptually simple, fast, requires little memory, and is easy to train, even on videos containing moving objects. It has been tested and performs well on indoor and outdoor sequences with strong camera jitter. Pierre-Marc Jodoin, Janusz Konrad, Venkatesh Saligrama, Vincent Veilleux-Gaboury |
ICIP | 1 |
| 2008 | Markovian method for 2D, 3D and 4D segmentation of MRIabstractMagnetic resonance imaging (MRI) is well adapted for early detection of diseases such as aortic aneuryms or dissections. In this paper, we present a new Markovian method which evolves an active contour for 2D, 3D and 4D (3D + time) segmentation. As opposed to other Markovian contour-based methods, our approach considers an implicit contour as the boundary of a 2D region. The regions are modeled via a Markov random field (MRF) and their computation is based on the maximum a posteriori probability criterion solved using an ICM algorithm. Our method depends on only one parameter that controls region boundary smoothness, is fast, easy to implement and can accommodate different likelihood functions to handle images with very different characteristics. Results on real and synthetic MRI are presented. Pierre-Marc Jodoin, Alain Lalande, Yvon Voisin, Olivier Bouchot, Eric Steinmetz |
ICIP | 1 |
| 2008 | Motion detection with false discovery rate controlabstractVisual surveillance applications such as object identification, object tracking, and anomaly detection require reliable motion detection as an initial processing step. Such a detection is often accomplished by means of background subtraction which can be as simple as thresholding of intensity difference between movement-free background and current frame. However, more effective background subtraction methods employ probabilistic modeling of the background followed by probability thresholding. In this case, the balance between false positives and false negatives (misses) is controlled by a threshold that needs to be adjusted heuristically depending on object sparsity. In this paper, we propose a different detection method that is based on false discovery rate control, a multiple-comparison procedure that applies thresholding in significance-score rather than probability space. The proposed approach allows explicit control of false positives and automatically adapts to object sparsity. The new method offers a qualitative improvement in real scenarios as well as a measurable performance gain over non-adaptive techniques when tested on synthetic sequences. J. Mike McHugh, Janusz Konrad, Venkatesh Saligrama, Pierre-Marc Jodoin, David A. Castañón |
ICIP | 4 |
| 2008 | Review and evaluation of commonly-implemented background subtraction algorithmsabstractLocating moving objects in a video sequence is the first step of many computer vision applications. Among the various motion-detection techniques, background subtraction methods are commonly implemented, especially for applications relying on a fixed camera. Since the basic inter-frame difference with global threshold is often a too simplistic method, more elaborate (and often probabilistic) methods have been proposed. These methods often aim at making the detection process more robust to noise, background motion and camera jitter. In this paper, we present commonly-implemented background subtraction algorithms and we evaluate them quantitatively. In order to gauge performances of each method, tests are performed on a wide range of real, synthetic and semi-synthetic video sequences representing different challenges. Yannick Benezeth, Pierre-Marc Jodoin, Bruno Emile, Hélène Laurent, Christophe Rosenberger |
ICPR | 2 |
| 2007 | Statistical Background Subtraction Using Spatial CuesabstractMost statistical background subtraction techniques are based on the analysis of temporal color/intensity distribution. However, learning statistics on a series of time frames can be problematic, especially when no frame absent of moving objects is available or when the available memory is not sufficient to store the series of frames needed for learning. In this letter, we propose a spatial variation to the traditional temporal framework. The proposed framework allows statistical motion detection with methods trained on one background frame instead of a series of frames as is usually the case. Our framework includes two spatial background subtraction approaches suitable for different applications. The first approach is meant for scenes having a nonstatic background due to noise, camera jitter or animation in the scene (e.g.,waving trees, fluttering leaves). This approach models each pixel with two PDFs: one unimodal PDF and one multimodal PDF, both trained on one background frame. In this way, the method can handle backgrounds with static and nonstatic areas. The second spatial approach is designed to use as little processing time and memory as possible. Based on the assumption that neighboring pixels often share similar temporal distribution, this second approach models the background with one global mixture of Gaussians. Pierre-Marc Jodoin, Max Mignotte, Janusz Konrad |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Segmentation Framework Based on Label Field FusionabstractIn this paper, we put forward a novel fusion framework that mixes together label fields instead of observation data as is usually the case. Our framework takes as input two label fields: a quickly estimated and to-be-refined segmentation map and a spatial region map that exhibits the shape of the main objects of the scene. These two label fields are fused together with a global energy function that is minimized with a deterministic iterative conditional mode algorithm. As explained in the paper, the energy function may implement a pure fusion strategy or a fusion-reaction function. In the latter case, a data-related term is used to make the optimization problem well posed. We believe that the conceptual simplicity, the small number of parameters, the use of a simple and fast deterministic optimizer that admits a natural implementation on a parallel architecture are among the main advantages of our approach. Our fusion framework is adapted to various computer vision applications among which are motion segmentation, motion estimation and occlusion detection. Pierre-Marc Jodoin, Max Mignotte, Christophe Rosenberger |
IEEE Trans. Image Process. | 1 |
| 2006 | Detecting Half-Occlusion with a Fast Region-Based Fusion ProcedureabstractThis paper presents a novel region-based approach for detecting occlusion between two consecutive frames. Based on a generalization of Marr and Poggio’s uniqueness assumption, the explicit goal of our method is to reduce the number of false positives while optimizing the hit rate. To do so, our method relies on a fusion procedure that blends together two segmentation maps: one pre-estimated occlusion binary map and one color segmentation map. While the occlusion map is obtained after a simple thresholding procedure, the color segmentation map is obtained with an unsupervised Markovian approach. Assuming that the color segmentation regions exhibit more precise edges, the occlusion areas are iteratively modified to fit the colorregion shapes. Since our method has been entirely implemented on a parallel architecture (a Graphics Processor Unit), its processing times are remarkably low. Our method is compared with other occlusion approaches both quantitatively and qualitatively on scenes that represent different challenges. 1 Pierre-Marc Jodoin, Christophe Rosenberger, Max Mignotte |
BMVC | 1 |
| 2006 | Optical-Flow Based on an Edge-Avoidance ProcedureabstractThis paper presents a differential optical flow method which accounts for two typical motion-estimation problems : (1) flow regularization within regions of uniform motion while (2) preserving sharp edges near motion discontinuities i.e., where motion is mul-timodal by nature. The method proposed is a modified version of the well known Lucas Kanade (LK) algorithm. Based on documented assumptions, our method computes motion with a classical least-square fit on a local neighborhood shifted away from where motion is likely to be multimodal. This edge-avoidance procedure is based on the non-parametric mean-shift algorithm which shifts the LK integration window away from local sharp edges. Our method also locally regularizes motion by performing a fusion of local motion estimates. Our method is compared with other edge-preserving methods on image sequences representing different challenges. Pierre-Marc Jodoin, Max Mignotte |
ICIP | 1 |
| 2006 | Light and Fast Statistical Motion Detection Method Based on Ergodic ModelabstractIn this paper, we propose a light and fast pixel-based statistical motion detection method based on a background subtraction procedure. The statistical representation of the background relies on its spatial color distributions herein modeled by a mixture of Gaussians. The Gaussian parameters are obtained after segmenting one reference frame with an unsupervised Bayesian approach whose parameter estimation step is ensured by the K-means and the iterated conditional estimation (ICE) algorithms. Since the motion detection function only depends on a global mixture of M Gaussians, only a few bits per pixel need to be stored in memory. Our method achieves real-time performances, especially when look up tables are used to store pre-calculated data. Results have been obtained on synthetic and real video sequences and compared with other statistical methods. Pierre-Marc Jodoin, Max Mignotte, Janusz Konrad |
ICIP | 1 |
| 2004 | Unsupervised motion detection using a markovian temporal model with global spatial constraintsabstractIn this work, we propose an unsupervised Bayesian model for the detection of moving objects from dynamic scenes. This unsupervised solution is a three-step approach that uses a statistical model of an interframe gradient norm field (as likelihood model) with a local regularization term (as prior model) combined with strong intraframe spatial constraints. In the first step, the spatial constraints are estimated by making an unsupervised Markovian spatial over-segmentation of two input frames. In the second step, the interframe gradient (derived from the input frames) is restored to minimize undesired noise. In the last step, an unsupervised Markovian temporal segmentation (with global spatial constraints) is performed to generate the desired motion label field. The maximum a posteriori (MAP) estimation of the label field associated with the spatial segmentations (in the first step) and the motion label field (in the third step) is performed by a classical Iterative Conditional Mode (ICM) algorithm. An Iterative Conditional Estimation (ICE) procedure is exploited for estimating the parameters of the spatial model and the region-constrained temporal model. This new statistical method of motion detection has been successfully applied to real dynamic scenes and seems to be well suited for the temporal detection of noisy image sequences. Pierre-Marc Jodoin, Max Mignotte |
ICIP | 1 |
| 2004 | An energy-basfd framework using global spatial constraints for the stereo correspondence problemabstractThis paper investigates the use of a region-based approach for the stereo matching problem. We have stated this problem in a commonly adopted global energy-based framework. Our energy-based model mixes a local and robust regularization term with global spatial constraints. These constraints are related to a (precomputed) partition into homogeneous regions with identical disparity. In practice, our approach assigns a single disparity to regions instead of individual pixels. These regions, used to globally constrain the ill-posed nature of our minimization problem, are estimated by combining an unsupervised Markovian segmentation and a roughly estimated disparity map. This disparity map is computed with a basic winner-take-all (WTA) procedure. The proposed global energy function seems to be well suited to find good disparity discontinuities at object boundaries, especially when the number of disparities is large. An iterated conditional modes (ICM) algorithm is used to optimize this global energy function. We provide experimental results on real stereo image pairs. A quality measure, based on ground truth data, is used to evaluate the performance of our algorithm. Results indicate that our approach is fast and performs well compared to other existing methods. Pierre-Marc Jodoin, Max Mignotte |
ICIP | 1 |
| 2004 | Fast hierarchical importance sampling with blue noise propertiesabstractThis paper presents a novel method for efficiently generating a good sampling pattern given an importance density over a 2D domain. A Penrose tiling is hierarchically subdivided creating a sufficiently large number of sample points. These points are numbered using the Fibonacci number system, and these numbers are used to threshold the samples against the local value of the importance density. Pre-computed correction vectors, obtained using relaxation, are used to improve the spectral characteristics of the sampling pattern. The technique is deterministic and very fast; the sampling time grows linearly with the required number of samples. We illustrate our technique with importance-based environment mapping, but the technique is versatile enough to be used in a large variety of computer graphics applications, such as light transport calculations, digital halftoning, geometry processing, and various rendering techniques. Victor Ostromoukhov, Charles Donohue, Pierre-Marc Jodoin |
ACM Trans. Graph. | 3 |