Thierry Bouwmans

dblp:48/5367 · DBLP profile ↗
← Back
48ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0003-4018-8856ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A novel method for anomaly detection using graph signal embedding: Application to anomalous sound detection
abstract
Anomalous Sound Detection (ASD) is an emerging field that combines audio signal processing and anomaly detection to provide accurate detection of anomalous sound events. Conventional ASD encounters two main problems: a) the inherently low amount of anomalous samples, eg. traffic accidents compared to any other event on the road, which creates an unbalanced dataset, and b) the limited ability of standard audio features, such as MFCC, to characterise anomalies, especially in highly noisy environments, such as roads, factories, cocktail party scenes, etc. In this context, we propose a Graph Signal Processing (GSP) approach based on graph embedding and filtering before undertaking anomaly detection methods, such as One-Class SVM, Isolation Forests and Autoencoders. First, acoustic features such as MFCC or log-Energy are extracted and embedded into a graph. Second, the constructed graphs are -optionally- filtered using a graph filterbank. Third, the magnitude of the joint Fourier transform is computed for each filtered graph. Finally, the features thus extracted are used to train one of the standard semi-supervised anomaly detection method using only normal data. To evaluate the advantages of the proposed approach, we tested using different standard anomaly detection methods, such as OC-SVM, isolation forests and autoencoders. Besides, experiments were conducted on different kinds of sounds and anomalies. Each anomaly detection method is carried out with and without graph embedding. In case of graph embedding, different types of graph topologies are tested, such as Path, Ring, Fully-Connected, Sensor-Network, … , etc. Experiments demonstrate the effectiveness of the proposed Graph Embedding and Filtering technique in improving the results of anomaly detection for all the methods.
Zied Mnasri, Thierry Bouwmans
Expert Syst. Appl.2
2026 Graph refactored domain adversarial learning for underwater image enhancement
Meghna Kapoor, Badri N. Subudhi, Thierry Bouwmans, Ankur Bansal
Pattern Recognit. Lett.3
2025 Graph-based Moving Object Segmentation for underwater videos using semi-supervised learning
abstract
International audience
Meghna Kapoor, Wieke Prummel, Jhony-Heriberto Giraldo-Zuluaga, Badri N. Subudhi, Anastasia Zakharova, Thierry Bouwmans, Ankur Bansal
Comput. Vis. Image Underst.6
2025 Long-term spatio-temporal graph attention network for traffic forecasting
abstract
• Comprehensive overview of numerous recent works. • Holistic understanding of spatial and temporal dependencies for traffic forecasting. • Analysis of long-term traffic patterns to reveal consistent behavioral trends. • Combination of local and global spatial data for detailed traffic pattern insights. • An attention-based decoder that focuses on the most relevant extracted features. Accurate traffic flow prediction is a critical component of intelligent transportation systems and smart cities, playing an essential role in traffic control, transportation planning, and infrastructure development. Numerous recent research studies highlight the need to enhance prediction accuracy by addressing complex temporal and spatial dependencies. However, due to the complexity of these spatio-temporal patterns, achieving accurate traffic predictions is still a main challenge in long-term scenarios. In this context, we first provide a comprehensive overview of the traffic forecasting to locate where research is going on. Then, we develop a Long-Term Spatio-Temporal Graph Attention Network (LSTGAN) architecture designed to analyze long-term historical data to address the above issue. This architecture encodes several previous time steps and extracts temporal patterns using convolutional layers. These features are then combined with the spatial features captured by a spatial attention module and a graph convolution layer to be processed by a temporal attention decoder responsible for making predictions. Experiments on METR-LA and PEMS-BAY datasets show that our proposed architecture outperforms most existing state-of-the-art baselines.
Brahim Remmouche, Doulkifli Boukraâ, Anastasia Zakharova, Thierry Bouwmans, Mokhtar Taffar
Expert Syst. Appl.4
2025 AI4RDD: Artificial Intelligence and Rare Disease Diagnosis: A proposal to improve the anamnesis process
Serena Lembo, Paola Barra, Luigi Di Biasi, Thierry Bouwmans, Genny Tortora
Image Vis. Comput.4
2024 Robust and efficient FISTA-based method for moving object detection under background movements
Maryam Amoozegar, Masoumeh Akbarizadeh, Thierry Bouwmans
Knowl. Based Syst.3
2024 A Multi-Scale Contrast Preserving Encoder-Decoder Architecture for Local Change Detection From Thermal Video Scenes
abstract
This article presents a new deep-learning architecture based on an encoder-decoder framework that retains contrast while performing background subtraction (BS) on thermal videos. The proposed scheme consists of three consecutive blocks: the encoder, the Multi-Scale Contrast Preservation (MSCP) block, and the decoder. The encoder network employs a hybrid of convolution and atrous convolution blocks to preserve both sparse and dense features, with a skip connection. The encoder, combined with the MSCP block, maintains multi-scale contrast features with reduced training loss. Furthermore, the decoder network accurately projects the extracted features at different layers into pixel-level detail. The proposed end-to-end model efficiently provides a binary map for the corresponding thermal video scene. The efficiency of the proposed algorithm is validated on two large-scale datasets, namely CDnet 2014 and the Tripura University Video Dataset at Night Time (TU-VDN). Both qualitative and quantitative results demonstrate that MSCP outperforms thirty-eight existing BS schemes.
Badri N. Subudhi, Veerakumar Thangaraj, Vinit Jakhetiya, Thierry Bouwmans
IEEE Trans. Inf. Forensics Secur.5
2023 On the Trade-off between Over-smoothing and Over-squashing in Deep Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have succeeded in various computer science applications, yet deep GNNs underperform their shallow counterparts despite deep learning's success in other domains. Over-smoothing and over-squashing are key challenges when stacking graph convolutional layers, hindering deep representation learning and information propagation from distant nodes. Our work reveals that over-smoothing and over-squashing are intrinsically related to the spectral gap of the graph Laplacian, resulting in an inevitable trade-off between these two issues, as they cannot be alleviated simultaneously. To achieve a suitable compromise, we propose adding and removing edges as a viable approach. We introduce the Stochastic Jost and Liu Curvature Rewiring (SJLR) algorithm, which is computationally efficient and preserves fundamental properties compared to previous curvature-based methods. Unlike existing approaches, SJLR performs edge addition and removal during GNN training while maintaining the graph unchanged during testing. Comprehensive comparisons demonstrate SJLR's competitive performance in addressing over-smoothing and over-squashing.
Jhony-Heriberto Giraldo-Zuluaga, Konstantinos Skianis, Thierry Bouwmans, Fragkiskos D. Malliaros
CIKM3
2023 Time-Varying Signals Recovery Via Graph Neural Networks
abstract
The recovery of time-varying graph signals is a fundamental problem with numerous applications in sensor networks and forecasting in time series. Effectively capturing the spatiotemporal information in these signals is essential for the downstream tasks. Previous studies have used the smoothness of the temporal differences of such graph signals as an initial assumption. Nevertheless, this smoothness assumption could result in a degradation of performance in the corresponding application when the prior does not hold. In this work, we relax the requirement of this hypothesis by including a learning module. We propose a Time Graph Neural Network (TimeGNN) for the recovery of time-varying graph signals. Our algorithm uses an encoder-decoder architecture with a specialized loss composed of a mean squared error function and a Sobolev smoothness operator. TimeGNN shows competitive performance against previous methods in real datasets.
Jhon A. Castro-Correa, Jhony-Heriberto Giraldo-Zuluaga, Anindya Mondal, Mohsen Badiey, Thierry Bouwmans, Fragkiskos D. Malliaros
ICASSP5
2023 Higher-Order Sparse Convolutions in Graph Neural Networks
abstract
Graph Neural Networks (GNNs) have been applied to many problems in computer sciences. Capturing higher-order relationships between nodes is crucial to increase the expressive power of GNNs. However, existing methods to capture these relationships could be infeasible for large-scale graphs. In this work, we introduce a new higher-order sparse convolution based on the Sobolev norm of graph signals. Our Sparse Sobolev GNN (S-SobGNN) computes a cascade of filters on each layer with increasing Hadamard powers to get a more diverse set of functions, and then a linear combination layer weights the embeddings of each filter. We evaluate S-SobGNN in several applications of semi-supervised learning. S-SobGNN shows competitive performance in all applications as compared to several state-of-the-art methods.
Jhony-Heriberto Giraldo-Zuluaga, Sajid Javed, Arif Mahmood, Fragkiskos D. Malliaros, Thierry Bouwmans
ICASSP5
2023 Inductive Graph Neural Networks for Moving Object Segmentation
abstract
Moving Object Segmentation (MOS) is a challenging problem in computer vision, particularly in scenarios with dynamic backgrounds, abrupt lighting changes, shadows, camouflage, and moving cameras. While graph-based methods have shown promising results in MOS, they have mainly relied on transductive learning which assumes access to the entire training and testing data for evaluation. However, this assumption is not realistic in real-world applications where the system needs to handle new data during deployment. In this paper, we propose a novel Graph Inductive Moving Object Segmentation (GraphIMOS) algorithm based on a Graph Neural Network (GNN) architecture. Our approach builds a generic model capable of performing prediction on newly added data frames using the already trained model. GraphI-MOS outperforms previous inductive learning methods and is more generic than previous transductive techniques. Our proposed algorithm enables the deployment of graph-based MOS models in real-world applications.
Wieke Prummel, Jhony-Heriberto Giraldo-Zuluaga, Anastasia Zakharova, Thierry Bouwmans
ICIP4
2023 Uncertainty clustering internal validity assessment using Fréchet distance for unsupervised learning
abstract
Knowing the number of clusters a priori is one of the most challenging aspects of unsupervised learning. Clustering Internal Validity Indices (CIVIs) evaluate partitions in unsupervised algorithms based on metrics like compactness, separation, and density. However, specialized CIVIs for specific applications have been designed, and there is no general CIVI that works in all scenarios. The absence of CIVIs based on crisp uncertainty metrics is especially critical in decision-making processes that involve ambiguity, non-convex distributions, outliers, and overlapping data. To address this problem, we propose a novel Uncertainty Fréchet (UF) CIVI that assesses the certainty of a well-defined partition. UF leverages uncertainty fingerprints based on Type-2 fuzzy Gaussian Mixture Models (T2FGMM) and the Fréchet distance between clusters to introduce a metric that evaluates partition quality. We integrate UF into a merging methodology that combines similar clusters within a partition, allowing us to determine the number of clusters without the need to run the clustering algorithms iteratively as other CIVIs require. We undertake a comprehensive evaluation of our proposal on 5,250 convex, 36 non-convex synthetic datasets, and five benchmark real datasets. In addition, we apply UF in a real-world scenario that involves high uncertainty: Passive Acoustic Monitoring (PAM) of ecosystems, which aims to study ecological transformations through acoustic recordings. The results show that UF exhibits notable performance in synthetic and real-world scenarios, obtaining an Adjusted Mutual Information (AMI) score higher than 0.88 for normal, uniform, gamma, and triangular distribution datasets. In the PAM application, UF identifies the transformation of ecosystems through sound using clustering algorithms and UF, achieving an F1 score of 0.84. Therefore, results show that the UF index is a suitable tool for researchers and practitioners working with highly uncertain data.
Nestor Rendon, Jhony-Heriberto Giraldo-Zuluaga, Thierry Bouwmans, Susana Rodríguez-Buritica, Edison Ramirez, Claudia Isaza
Eng. Appl. Artif. Intell.3
2022 An End to End Encoder-Decoder Network with Multi-scale Feature Pulling for Detecting Local Changes From Video Scene
abstract
Local change detection for moving object detection is an essential step in any computer vision task. The most well-known technique is background subtraction BGS. However, the performance of BGS is strongly dependent on the background construction. The background construction to be robust in the presence of various challenges: dynamic backgrounds, illumination changes, camera jitter, etc. In this paper, we propose a novel encoder-decoder-based end-to-end deep learning framework for BGS. Thus, we explore a VGG-19 deep network with a transfer learning strategy as an encoder that deeply learned and extracted the features at different levels. We herewith propose a Multi-scale Feature Pulling MFP block which can retain the features at the various scales of the challenging video scenes. We also design a decoder network which is a stack of several transposed convolutional layers which precisely predict that each pixel of the target frame belongs to the background or foreground. The efficiency of the proposed algorithm is validated on the CDNet-2014 dataset by comparing its results against seventeen state-of-the-art techniques.
Badri N. Subudhi, Thierry Bouwmans, Vinit Jakheytiya, Veerakumar Thangaraj
AVSS3
2022 Hypergraph Convolutional Networks for Weakly-Supervised Semantic Segmentation
abstract
Semantic segmentation is a fundamental topic in computer vision. Several deep learning methods have been proposed for semantic segmentation with outstanding results. However, these models require a lot of densely annotated images. To address this problem, we propose a new algorithm that uses Hy-perGraph Convolutional Networks for Weakly-supervised Semantic Segmentation (HyperGCN-WSS). Our algorithm constructs spatial and k-Nearest Neighbor (k-NN) graphs from the images in the dataset to generate the hypergraphs. Then, we train a specialized HyperGraph Convolutional Network (HyperGCN) architecture using some weak signals. The outputs of the HyperGCN are denominated pseudo-labels, which are later used to train a DeepLab model for semantic segmentation. HyperGCN-WSS is evaluated on the PASCAL VOC 2012 dataset for semantic segmentation, using scribbles or clicks as weak signals. Our algorithm shows competitive performance against previous methods.
Jhony-Heriberto Giraldo-Zuluaga, Vincenzo Mariano Scarrica, Antonino Staiano, Francesco Camastra, Thierry Bouwmans
ICIP5
2022 Moving objects segmentation using generative adversarial modeling
Maryam Sultana, Arif Mahmood, Thierry Bouwmans, Muhammad Haris Khan, Soon Ki Jung
Neurocomputing3
2022 SemiSegSAR: A Semi-Supervised Segmentation Algorithm for Ship SAR Images
abstract
Automatic ship segmentation from high-resolution Synthetic Aperture Radar (SAR) remote sensing images has been a topic of interest that has gradually gained attention over the years due to the abundance of earth observation sensors. Recently, deep learning methods have provided a breakthrough increasing the performance greatly by using large amount of labeled data. Yet, the high cost related to the samples labeling and their scarcity result in significant limitation of their wide use. Therefore, it is crucial to overcome the unlabeled inputs challenge and develop semi-supervised learning approaches to enhance the machine learning models capacity. Our letter proposes a semi-supervised segmentation algorithm for SAR images named SemiSegSAR based on the use of Graph Signal Processing. This method includes instance segmentation; texture and statistical SAR features to represent the nodes of the graph; K-nearest neighbors to construct the graph; and Sobolev minimization algorithm to tackle the problem of semi-supervised semantic segmentation. The proposed algorithm is trained and tested using the publicly available SSDD and HRSID ship detection datasets. Experiments show that SemiSegSAR outperforms the current state-of-the-art semi-supervised and supervised methods while requiring only few labeled data.
Marwa Chendeb, Jhony-Heriberto Giraldo-Zuluaga, Mina Al-Saad, Muna Darweesh, Thierry Bouwmans
IEEE Geosci. Remote. Sens. Lett.5
2022 Graph Moving Object Segmentation
abstract
Moving Object Segmentation (MOS) is a fundamental task in computer vision. Due to undesirable variations in the background scene, MOS becomes very challenging for static and moving camera sequences. Several deep learning methods have been proposed for MOS with impressive performance. However, these methods show performance degradation in the presence of unseen videos; and usually, deep learning models require large amounts of data to avoid overfitting. Recently, graph learning has attracted significant attention in many computer vision applications since they provide tools to exploit the geometrical structure of data. In this work, concepts of graph signal processing are introduced for MOS. First, we propose a new algorithm that is composed of segmentation, background initialization, graph construction, unseen sampling, and a semi-supervised learning method inspired by the theory of recovery of graph signals. Second, theoretical developments are introduced, showing one bound for the sample complexity in semi-supervised learning, and two bounds for the condition number of the Sobolev norm. Our algorithm has the advantage of requiring less labeled data than deep learning methods while having competitive results on both static and moving camera videos. Our algorithm is also adapted for Video Object Segmentation (VOS) tasks and is evaluated on six publicly available datasets outperforming several state-of-the-art methods in challenging conditions.
Jhony-Heriberto Giraldo-Zuluaga, Sajid Javed, Thierry Bouwmans
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 Semi-Supervised Background Subtraction Of Unseen Videos: Minimization Of The Total Variation Of Graph Signals
abstract
Recently, several successful methods based on deep neural networks have been proposed for background subtraction. These deep neural algorithms have almost perfect performance, relying in the availability of ground-truth frames of the tested videos during the training step. However, the performance of some of these algorithms drops significantly when tested on unseen videos. In this paper, concepts of semi-supervised learning are introduced in the problem of background subtraction for unseen videos. We propose a new algorithm named Graph-BGS-TV, this method uses: Mask R-CNN for instances segmentation; temporal median filter for background initialization; motion, texture, and intensity features for representing the nodes of a graph; k-nearest neighbors for the construction of the graph; and finally a total variation minimization algorithm to solve the problem of background subtraction. GraphBGS-TV is tested in the change detection dataset, outperforming unsupervised and supervised methods in the challenges “PTZ” and “shadows”.
Jhony-Heriberto Giraldo-Zuluaga, Thierry Bouwmans
ICIP2
2020 Dual Information-Based Background Model For Moving Object Detection
abstract
In this article, a novel pixel based object detection framework is proposed that leverages dual type pixel-level information to construct the background model. The first type of information is initially used intensity histograms over a training set of a few initial video frames. Finally, it is formed by gathering all the minimum and maximum values of contiguous non-zero frequencies of the temporal intensity histogram. The second type of information constitutes a set having only the discrete pixel values. Subsequently, a pixel-level periodic updating scheme is used to make the model robust and flexible enough to recognize and detect foregrounds in various critical background environments. This dual format model produces effective results over many state-of-the-art methods in a large variety of challenging real-life video sequences.
Sujoy Madhab Roy, Thierry Bouwmans
ICIP2
2020 Dynamic Background Subtraction Using Least Square Adversarial Learning
abstract
Dynamic Background Subtraction (BS) is a fundamental problem in many vision-based applications. BS in real complex environments has several challenging conditions like illumination variations, shadows, camera jitters, and bad weather. In this study, we aim to address the challenges of BS in complex scenes by exploiting conditional least squares adversarial networks. During training, a scene-specific conditional least squares adversarial network with two additional regularizations including L1-Loss and Perceptual-Loss is employed to learn the dynamic background variations. The given input to the model is video frames conditioned on corresponding ground truth to learn the dynamic changes in complex scenes. Afterwards, testing is performed on unseen test video frames so that the generator would conduct dynamic background subtraction. The proposed method consisting of three loss-terms including least squares adversarial loss, L1-Loss and Perceptual-Loss is evaluated on two benchmark datasets CDnet2014 and BMC. The results of our proposed method show improved performance on both datasets compared with 10 existing state-of-the-art methods.
Maryam Sultana, Arif Mahmood, Thierry Bouwmans, Soon Ki Jung
ICIP3
2020 GraphBGS: Background Subtraction via Recovery of Graph Signals
abstract
Background subtraction is a fundamental preprocessing task in computer vision. This task becomes challenging in real scenarios due to variations in the background for both static and moving camera sequences. Several deep learning methods for background subtraction have been proposed in the literature with competitive performances. However, these models show performance degradation when tested on unseen videos; and they require huge amount of data to avoid overfitting. Recently, graph-based algorithms have been successful approaching unsupervised and semi-supervised learning problems. Furthermore, the theory of graph signal processing and semi-supervised learning have been combined leading to new insights in the field of machine learning. In this paper, concepts of recovery of graph signals are introduced in the problem of background subtraction. We propose a new algorithm called Graph BackGround Subtraction (GraphBGS), which is composed of: instance segmentation, background initialization, graph construction, graph sampling, and a semi-supervised algorithm inspired from the theory of recovery of graph signals. Our algorithm has the advantage of requiring less labeled data than deep learning methods while having competitive results on both: static and moving camera videos. GraphBGS outperforms unsupervised and supervised methods in several challenging conditions on the publicly available Change Detection (CDNet2014), and UCSD background subtraction databases.
Jhony-Heriberto Giraldo-Zuluaga, Thierry Bouwmans
ICPR2
2020 Deep detector classifier (DeepDC) for moving objects segmentation and classification in video surveillance
abstract
In this study, the authors present a new approach to segment and classify moving objects in video sequences by combining an unsupervised anomaly discovery framework called DeepSphere and generative adversarial networks. The proposed deep detector classifier employs and validates DeepSphere, which aims mainly to identify the anomalous cases in the spatial and temporal context in order to perform foreground objects segmentation. For post‐processing, some morphological operations are considered to better segment and extract the desired objects. Finally, they take advantage of the power of generative models, which recognise the problem of semi‐supervised learning as a specific missing data imputation task in order to classify the segmented objects. They evaluate the method with multiple datasets and the results confirm the effectiveness of the proposed approach, which achieves superior performance over the state‐of‐the‐art methods having the capabilities of segmenting and classifying moving objects from videos surveillance.
Sirine Ammar, Thierry Bouwmans, Nizar Zaghden, Mahmoud Neji
IET Image Process.2
2020 Introduction to the Special Issue on Multimodal Machine Learning for Human Behavior Analysis
abstract
No abstract available.
Shengping Zhang, Huiyu Zhou 0001, Dong Xu 0001, M. Emre Celebi 0001, Thierry Bouwmans
ACM Trans. Multim. Comput. Commun. Appl.5
2019 Deep neural network concepts for background subtraction: A systematic review and comparative evaluation
Thierry Bouwmans, Sajid Javed, Maryam Sultana, Soon Ki Jung
Neural Networks1
2019 Moving Object Detection in Complex Scene Using Spatiotemporal Structured-Sparse RPCA
abstract
Moving object detection is a fundamental step in various computer vision applications. Robust Principal Component Analysis (RPCA) based methods have often been employed for this task. However, the performance of these methods deteriorates in the presence of dynamic background scenes, camera jitter, camouflaged moving objects, and/or variations in illumination. It is because of an underlying assumption that the elements in the sparse component are mutually independent, and thus the spatiotemporal structure of the moving objects is lost. To address this issue, we propose a spatiotemporal structured sparse RPCA algorithm for moving objects detection, where we impose spatial and temporal regularization on the sparse component in the form of graph Laplacians. Each Laplacian corresponds to a multi-feature graph constructed over superpixels in the input matrix. We enforce the sparse component to act as eigenvectors of the spatial and temporal graph Laplacians while minimizing the RPCA objective function. These constraints incorporate a spatiotemporal subspace structure within the sparse component. Thus, we obtain a novel objective function for separating moving objects in the presence of complex backgrounds. The proposed objective function is solved using a linearized alternating direction method of multipliers based batch optimization. Moreover, we also propose an online optimization algorithm for real-time applications. We evaluated both the batch and online solutions using six publicly available datasets that included most of the aforementioned challenges. Our experiments demonstrated the superior performance of the proposed algorithms compared with the current state-of-the-art methods.
Sajid Javed, Arif Mahmood, Somaya Al-Máadeed, Thierry Bouwmans, Soon Ki Jung
IEEE Trans. Image Process.4
2018 On the Applications of Robust PCA in Image and Video Processing
abstract
Robust principal component analysis (RPCA) via decomposition into low-rank plus sparse matrices offers a powerful framework for a large variety of applications such as image processing, video processing, and 3-D computer vision. Indeed, most of the time these applications require to detect sparse outliers from the observed imagery data that can be approximated by a low-rank matrix. Moreover, most of the time experiments show that RPCA with additional spatial and/or temporal constraints often outperforms the state-of-the-art algorithms in these applications. Thus, the aim of this paper is to survey the applications of RPCA in computer vision. In the first part of this paper, we review representative image processing applications as follows: 1) low-level imaging such as image recovery and denoising, image composition, image colorization, image alignment and rectification, multifocus image, and face recognition; 2) medical imaging such as dynamic magnetic resonance imaging (MRI) for acceleration of data acquisition, background suppression, and learning of interframe motion fields; and 3) imaging for 3-D computer vision with additional depth information such as in structure from motion (SfM) and 3-D motion recovery. In the second part, we present the applications of RPCA in video processing which utilize additional spatial and temporal information compared to image processing. Specifically, we investigate video denoising and restoration, hyperspectral video, and background/foreground separation. Finally, we provide perspectives on possible future research directions and algorithmic frameworks that are suitable for these applications.
Thierry Bouwmans, Sajid Javed, Hongyang Zhang 0001, Zhouchen Lin, Ricardo Otazo
Proc. IEEE1
2018 Rethinking PCA for Modern Data Sets: Theory, Algorithms, and Applications
abstract
The papers in this special issue introduce the reader to the theory, algorithms, and applications of principal component analysis (PCA) and its many extensions. The aim of PCA is to reduce the dimensionality of multivariate data while preserving as much of the relevant information as possible. It is often the first step in various types of exploratory data analysis, predictive modeling, and classification and clustering tasks, and finds applications in biomedical imaging, computer vision, process fault detection, recommendation systems’ design, and many more domains.
Namrata Vaswani, Yuejie Chi, Thierry Bouwmans
Proc. IEEE3
2018 Spatiotemporal Low-Rank Modeling for Complex Scene Background Initialization
abstract
Background modeling constitutes the building block of many computer-vision tasks. Traditional schemes model the background as a low rank matrix with corrupted entries. These schemes operate in batch mode and do not scale well with the data size. Moreover, without enforcing spatiotemporal information in the low-rank component, and because of occlusions by foreground objects and redundancy in video data, the design of a background initialization method robust against outliers is very challenging. To overcome these limitations, this paper presents a spatiotemporal low-rank modeling method on dynamic video clips for estimating the robust background model. The proposed method encodes spatiotemporal constraints by regularizing spectral graphs. Initially, a motion-compensated binary matrix is generated using optical flow information to remove redundant data and to create a set of dynamic frames from the input video sequence. Then two graphs are constructed, one between frames for temporal consistency and the other between features for spatial consistency, to encode the local structure for continuously promoting the intrinsic behavior of the low-rank model against outliers. These two terms are then incorporated in the iterative Matrix Completion framework for improved segmentation of background. Rigorous evaluation on severely occluded and dynamic background sequences demonstrates the superior performance of the proposed method over state-of-the-art approaches.
Sajid Javed, Arif Mahmood, Thierry Bouwmans, Soon Ki Jung
IEEE Trans. Circuits Syst. Video Technol.3
2017 Multi-feature fusion based background subtraction for video sequences with strong background changes
abstract
Current background subtraction algorithms are sensitive to sudden changes. In this paper, we propose a multi-feature fusion scheme to background subtraction for video sequences with strong background changes. We reconstruct the whole videos frame by frame by fusing several video features. In this fusing step, we design an energy function based on enforcing every features with an equal weight. By comparing reconstruction videos with the original videos, pixels with small differences are classified as background pixels. Thus, we can identify background areas in advance and then we construct a contour-based mask combining mechanism. Experimental results conducted on the OTCBVS, BMC 2012 and PETS 2001 datasets show that our method improves the performance of the Zivkovic's GMM and SubSENSE for video sequences with strong background changes.
Zhenkun Huang, Ruimin Hu, Thierry Bouwmans
ICIP3
2017 Scene background initialization: A taxonomy
Thierry Bouwmans, Lucia Maddalena 0001, Alfredo Petrosino
Pattern Recognit. Lett.1
2017 Editorial-Scene background modeling and initialization
abstract
International audience
Alfredo Petrosino, Lucia Maddalena 0001, Thierry Bouwmans
Pattern Recognit. Lett.3
2017 Superpixel-based online wagging one-class ensemble for feature selection in foreground/background separation
Caroline Pacheco do Espírito Silva, Thierry Bouwmans, Carl Frélicot
Pattern Recognit. Lett.2
2017 Guest Editorial Introduction to the Special Issue on Group and Crowd Behavior Analysis for Intelligent Multicamera Video Surveillance
abstract
Despite significant progress in human behavior analysis over the past few years, most of today’s state-of-the-art algorithms focus on analyzing individual behavior in a simple environment monitored by a single camera. Recently, the widespread availability of cameras and a growing need for public safety have shifted the attention of researchers in video surveillance from individual behavior analysis to group and crowd behavior analysis in multicamera networks. Group behavior analysis provides a novel level for describing events, which are semantically more meaningful, highlighting barely visible relational connections among people. Crowd behavior analysis can also be used for anomaly detection such as panic scenarios, dangerous situations, and illegal behaviors in public spaces.
Hongxun Yao, Andrea Cavallaro, Thierry Bouwmans, Zhengyou Zhang
IEEE Trans. Circuits Syst. Video Technol.3
2017 Background-Foreground Modeling Based on Spatiotemporal Sparse Subspace Clustering
abstract
Background estimation and foreground segmentation are important steps in many high-level vision tasks. Many existing methods estimate background as a low-rank component and foreground as a sparse matrix without incorporating the structural information. Therefore, these algorithms exhibit degraded performance in the presence of dynamic backgrounds, photometric variations, jitter, shadows, and large occlusions. We observe that these backgrounds often span multiple manifolds. Therefore, constraints that ensure continuity on those manifolds will result in better background estimation. Hence, we propose to incorporate the spatial and temporal sparse subspace clustering into the robust principal component analysis (RPCA) framework. To that end, we compute a spatial and temporal graph for a given sequence using motion-aware correlation coefficient. The information captured by both graphs is utilized by estimating the proximity matrices using both the normalized Euclidean and geodesic distances. The low-rank component must be able to efficiently partition the spatiotemporal graphs using these Laplacian matrices. Embedded with the RPCA objective function, these Laplacian matrices constrain the background model to be spatially and temporally consistent, both on linear and nonlinear manifolds. The solution of the proposed objective function is computed by using the linearized alternating direction method with adaptive penalty optimization scheme. Experiments are performed on challenging sequences from five publicly available datasets and are compared with the 23 existing state-of-the-art methods. The results demonstrate excellent performance of the proposed algorithm for both the background estimation and foreground segmentation.
Sajid Javed, Arif Mahmood, Thierry Bouwmans, Soon Ki Jung
IEEE Trans. Image Process.3
2016 GMM Background Modeling Using Divergence-Based Weight Updating
Juan D. Pulgarin-Giraldo, Andrés Marino Álvarez-Meza, David Insuasti-Ceballos, Thierry Bouwmans, Germán Castellanos-Domínguez
CIARP4
2016 Motion-Aware Graph Regularized RPCA for background modeling of complex scenes
abstract
Computing a background model from a given sequence of video frames is a prerequisite for many computer vision applications. Recently, this problem has been posed as learning a low-dimensional subspace from high dimensional data. Many contemporary subspace segmentation methods have been proposed to overcome the limitations of the methods developed for simple background scenes. Unfortunately, because of the absence of motion information and without preserving intrinsic geometric structure of video data, most existing algorithms do not provide promising nature of the low-rank component for complex scenes. Such as largely occluded background by foreground objects, superfluity in video frames in order to cope with intermittent motion of foreground objects, sudden lighting condition variation, and camera jitter sequences. To overcome these difficulties, we propose a motion-aware regularization of graphs on low-rank component for video background modeling. We compute optical flow and use this information to make a motion-aware matrix. In order to learn the locality and similarity information within a video we compute inter-frame and intra-frame graphs which we use to preserve geometric information in the low-rank component. Finally, we use linearized alternating direction method with parallel splitting and adaptive penalty to incorporate the preceding steps to recover the model of the background. Experimental evaluations on challenging sequences demonstrate promising results over state-of-the-art methods.
Sajid Javed, Soon Ki Jung, Arif Mahmood, Thierry Bouwmans
ICPR4
2016 Online Weighted One-Class Ensemble for feature selection in background/foreground separation
abstract
Background subtraction (BS) is one of the key steps for detecting moving objects in video surveillance applications. In the last few years, many BS methods have been developed to handle the different challenges met in video surveillance but the role and the relevance of the visual features used has been less investigated. In this paper, we present an Online Weighted Ensemble of One-Class SVMs (Support Vector Machines) able to select suitable features for each pixel to distinguish the foreground objects from the background. In addition, our proposal uses a mechanism to update the relative importance of each feature over time. Moreover, a heuristic approach is used to reduce the complexity of the background model maintenance while maintaining the robustness of the background model. Results on two datasets show the pertinence of the approach.
Caroline Pacheco do Espírito Silva, Thierry Bouwmans, Carl Frélicot
ICPR2
2015 Double-constrained RPCA based on saliency maps for foreground detection in automated maritime surveillance
abstract
The development of automated video-surveillance applications for maritime environment is a very difficult task due to the complexity of the scenes: moving water, waves, etc. The motion of the objects of interest (i.e. ships or boats) can be mixed with the dynamic behavior of the background (non-regular patterns). In this paper, a double-constrained Robust Principal Component Analysis (RPCA), named SCM-RPCA (Shape and Confidence Map-based RPCA), is proposed to improve the object foreground detection in maritime scenes. The sparse component is constrained by shape and confidence maps both extracted from spatial saliency maps. The experimental results in the UCSD and MarDT data sets indicate a better enhancement of the object foreground mask when compared with some related RPCA methods.
Andrews Sobral, Thierry Bouwmans, El-hadi Zahzah
AVSS2
2014 OR-PCA with MRF for Robust Foreground Detection in Highly Dynamic Backgrounds
Sajid Javed, Seon Ho Oh, Andrews Sobral, Thierry Bouwmans, Soon Ki Jung
ACCV (3)4
2014 Robust PCA via Principal Component Pursuit: A review for a comparative evaluation in video surveillance
Thierry Bouwmans, El-hadi Zahzah
Comput. Vis. Image Underst.1
2014 Special issue on background modeling for foreground detection in real-world dynamic scenes
Thierry Bouwmans, Jordi Gonzàlez 0001, Caifeng Shan, Massimo Piccardi, Larry Davis 0001
Mach. Vis. Appl.1
2012 Foreground detection based on low-rank and block-sparse matrix decomposition
abstract
Foreground detection is the first step in video surveillance system to detect moving objects. Principal Components Analysis (PCA) shows a nice framework to separate moving objects from the background but without a mechanism of robust analysis, the moving objects may be absorbed into the background model. This drawback can be solved by recent researches on Robust Principal Component Analysis (RPCA). The background sequence is then modeled by a low rank subspace that can gradually change over time, while the moving foreground objects constitute the correlated sparse outliers. In this paper, we propose to use a RPCA method based on low-rank and block-sparse matrix decomposition to achieve foreground detection. This decomposition enforces the low-rankness of the background and the block-sparsity aspect of the foreground. Experimental results on different datasets show the pertinence of the proposed approach.
Charles Guyon, Thierry Bouwmans, El-hadi Zahzah
ICIP2
2012 Foreground detection via robust low rank matrix factorization including spatial constraint with Iterative reweighted regression
Charles Guyon, Thierry Bouwmans, El-hadi Zahzah
ICPR2
2012 Background subtraction via incremental maximum margin criterion: a discriminative subspace approach
Diana Farcas, Cristina Marghes, Thierry Bouwmans
Mach. Vis. Appl.3
2008 Fuzzy integral for moving object detection
abstract
Detection of moving objects is the first step in many applications using video sequences like video-surveillance, optical motion capture and multimedia application. The process mainly used is the background subtraction which one key step is the foreground detection. The goal is to classify pixels of the current image as foreground or background. Some critical situations as shadows, illumination variations can occur in the scene and generate a false classification of image pixels. To deal with the uncertainty in the classification issue, we propose to use the Choquet integral as aggregation operator. Experiments on different data sets in video surveillance have shown a robustness of the proposed method against some critical situations when fusing color and texture features. Different color spaces have been tested to improve the insensitivity of the detection to the illumination changes. Then, the algorithm has been compared with another fuzzy approach based on the Sugeno integral and has proved its robustness.
Fida El Baf, Thierry Bouwmans, Bertrand Vachon
FUZZ-IEEE2
2008 A fuzzy approach for background subtraction
abstract
Background Subtraction is a widely used approach to detect moving objects from static cameras. Many different methods have been proposed over the recent years and can be classified following different mathematical model: determinist model, statistical model or filter model. The presence of critical situations i.e. noise, illumination changes and structural background changes introduce two main problems: The first one is the uncertainty in the classification of the pixel in foreground and background. The second one is the imprecision in the localization of the moving object. In this context, we propose a fuzzy approach for background subtraction. For this, we use the Choquet integral in the foreground detection and propose fuzzy adaptive background maintenance. Results show the pertinence of our approach.
Fida El Baf, Thierry Bouwmans, Bertrand Vachon
ICIP2
2001 A new stereomatching algorithm based on linear features and the fuzzy integral
André Bigand, Thierry Bouwmans, Jean-Paul Dubus
Pattern Recognit. Lett.2
2001 Extraction of line segments from fuzzy images
André Bigand, Thierry Bouwmans, Jean-Paul Dubus
Pattern Recognit. Lett.2