EDBT 2026 Demo / reviewers in the wild / expert
François Brémond
dblp:90/6418
· DBLP profile ↗
182ranked-venue papers
4as first author
60since 2021 · last 2026
0000-0003-2988-2142ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 149 · 3 first-author · 46 since 2021Artificial intelligence and machine learning · 73 · 37 since 2021Systems, architecture and hardware · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSecurity and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Open Your Eyes to See More: Dual Perspective Contrastive Learning for Skeleton-Based Action Understanding
Di Yang 0002, Quan Kong, Gianpiero Francesca, François Brémond |
FG | 5 |
| 2026 | Dual the Reasoning, Double the Insight with TambI: A Self-Supervised Framework for Skeleton Action Representation
Snehashis Majhi, Di Yang 0002, Quan Kong, Gianpiero Francesca, François Brémond |
ICPR (3) | 6 |
| 2026 | Denoise, Divide, Distill, and Predict D3P: Towards Forecasting Long-horizon Real-world Anomaly from NormalcyabstractForecasting abnormal human behavior (AHB) in unconstrained real-world environments is critical for enabling proactive safety interventions. Unlike short-term anomaly detection, long-horizon forecasting offers a vital reaction window but remains underexplored due to three core challenges: (i) noisy, complex human–agent interactions; (ii) weak temporal coupling between normal observations and distant anomalies; and (iii) data scarcity limiting the scalability of autoregressive models. To address these, we propose ${{\mathcal{D}}^3}{\mathcal{P}}$ (Denoise, Divide, Distill, and Predict), a novel encoder–decoder framework that bridges denoised pasts with distilled autoregressive futures. Our Differential Past Encoder (DiPE) disentangles scene-level and object-level dynamics via differential attention, suppressing irrelevant interactions and enhancing discriminative cues. The Distilled Future Auto-Regressive Decoder (D-FAD) adopts a divide-and-conquer strategy, segmenting future queries into temporal chunks for sequential prediction, while leveraging distillation to balance robustness and latency. We validate our approach on the AHB-F benchmark, the only dataset dedicated to abnormal behavior forecasting, and further integrate D-FAD with several state-of-the-art methods. In all cases, our framework consistently outperforms prior work in both forecasting accuracy and computational efficiency. Quentin Mérilleau, Snehashis Majhi, Antitza Dantcheva, Quan Kong, Lorenzo Garattoni, Gianpiero Francesca, François Brémond |
WACV | 7 |
| 2026 | MuSACo: Multimodal Subject-Specific Selection and Adaptation for Expression Recognition with Co-TrainingabstractPersonalized expression recognition (ER) involves adapting a machine learning model to subject-specific data for improved recognition of expressions with considerable inter-personal variability. Subject-specific ER can benefit significantly from multi-source domain adaptation (MSDA) methods – where each domain corresponds to a specific subject – to improve model accuracy and robustness. Despite promising results, state-of-the-art MSDA approaches often overlook multimodal information or blend sources into a single domain, limiting subject diversity and failing to explicitly capture unique subject-specific characteristics. To address these limitations, we introduce MuSACo, a multi-modal subject-specific selection and adaptation method for ER based on co-training. It leverages complementary information across multiple modalities and multiple source domains for subject-specific adaptation. This makes MuSACo particularly relevant for affective computing applications in digital health, such as patient-specific assessment for stress or pain, where subject-level nuances are crucial. MuSACo selects source subjects relevant to the target and generates pseudo-labels using the dominant modality for class-aware learning, in conjunction with a class-agnostic loss to learn from less confident target samples. Finally, source features from each modality are aligned, while only confident target features are combined. Experimental results on challenging multimodal ER datasets – BioVid, StressID, and BAH – show that MuSACo outperforms UDA (blending) and state-of-the-art MSDA methods. Our code is available: https://github.com/osamazeeshan/MuSACo Muhammad Osama Zeeshan, Natacha Gillet, Alessandro L. Koerich, Marco Pedersoli, François Brémond, Eric Granger |
WACV | 5 |
| 2025 | SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily LivingabstractThe introduction of vision-language models like CLIP has enabled the development of foundational video models capable of generalizing to unseen videos and human actions. However, these models are typically trained on web videos, which often fail to capture the challenges present in Activities of Daily Living (ADL) videos. Existing works address ADL-specific challenges, such as similar appearances, subtle motion patterns, and multiple viewpoints, by combining 3D skeletons and RGB videos. However, these approaches are not integrated with language, limiting their ability to generalize to unseen action classes. In this paper, we introduce SKI models, which integrate 3D skeletons into the vision-language embedding space. SKI models leverage a skeleton-language model, SkeletonCLIP, to infuse skeleton information into Vision Language Models (VLMs) and Large Vision Language Models (LVLMs) through collaborative training. Notably, SKI models do not require skeleton data during inference, enhancing their robustness for real-world applications. The effectiveness of SKI models is validated on three popular ADL datasets for zero-shot action recognition and video caption generation tasks. Arkaprava Sinha, Dominick Reilly, François Brémond, Pu Wang 0001, Srijan Das |
AAAI | 3 |
| 2025 | Are Attention Maps Richer than we Imagined for Action Recognition?abstractDeep learning models are becoming more general and robust by the day. Specifically, image foundation models have recently shown exponential growth. In this work, we introduce a way to exploit this growth in the field of video classification. The basic idea here is that if we have a good understanding of space, we should not require complicated spatio-temporal processing. We introduce Attention Map (AM) flow, a way to identify the location of local changes between two frames in a video, without adding additional parameters specifically for it. We utilise adapters, which have been growing in popularity in the field of parameterefficient transfer learning. These help us incorporate AM flow in a pretrained image model without the need of finetuning it. With just these changes and minimal temporal processing, an image model is able to achieve state-of-the- art results on popular action recognition datasets with low training time and requiring minimal pretraining. This work explores the theory behind this idea and the intricacies involved. Through relevant experiments, we show the efficacy of this method and discuss various ideas to take this work forward. We use kinetics-400, something-something v2 and Toyota smarthome datasets and achieve state-of-the-art or comparable results. We also show that video models suffer from extensive pretraining on multiple datasets and a large training time, but our work answers these problems. actionrecognition transformers image-to-video-models Tanay Agrawal, Abid Ali 0002, Antitza Dantcheva, François Brémond |
AVSS | 4 |
| 2025 | Just Dance with pi! A Poly-modal Inductor for Weakly-supervised Video Anomaly DetectionabstractWeakly-supervised methods for video anomaly detection (VAD) are conventionally based merely on RGB spatio-temporal features, which continues to limit their reliability in real-world scenarios. This is due to the fact that RGB-features are not sufficiently distinctive in setting apart categories such as shoplifting from visually similar events. Therefore, towards robust complex real-world VAD, it is essential to augment RGB spatio-temporal features by additional modalities. Motivated by this, we introduce the Poly-modal Induced framework for VAD: "PI-VAD" (or π-VAD), a novel approach that augments RGB representations by five additional modalities. Specifically, the modalities include sensitivity to fine-grained motion (Pose), three dimensional scene and entity representation (Depth), surrounding objects (Panoptic masks), global motion (optical flow), as well as language cues (VLM). Each modality represents an axis of a polygon, streamlined to add salient cues to RGB. π-VAD includes two plug-in modules, namely Pseudo-modality Generation module and Cross Modal Induction module, which generate modality-specific prototypical representation and, thereby, induce multi-modal information into RGB cues. These modules operate by performing anomaly-aware auxiliary tasks and necessitate five modality backbones – only during training. Notably, π-VAD achieves state-of-the-art accuracy on three prominent VAD datasets encompassing real-world scenarios, without requiring the computational overhead of five modality backbones at inference. Snehashis Majhi, Giacomo D'Amicantonio, Antitza Dantcheva, Quan Kong, Lorenzo Garattoni, Gianpiero Francesca, Egor Bondarev, François Brémond |
CVPR | 8 |
| 2025 | LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of LivingabstractCurrent Large Language Vision Models (LLVMs) trained on web videos perform well in general video understanding but struggle with fine-grained details, complex human-object interactions (HOI), and view-invariant representation learning essential for Activities of Daily Living (ADL). This limitation stems from a lack of specialized ADL video instruction-tuning datasets and insufficient modality integration to capture discriminative action representations. To address this, we propose a semi-automated framework for curating ADL datasets, creating ADL-X, a multiview, multimodal RGBS instruction-tuning dataset. Additionally, we introduce LLAVIDAL, an LLVM integrating videos, 3D skeletons, and HOIs to model ADL’s complex spatiotemporal relationships. For training LLAVIDAL a simple joint alignment of all modalities yields suboptimal results; thus, we propose a Multimodal Progressive (MMPro) training strategy, incorporating modalities in stages following a curriculum. We also establish ADL MCQ and video description benchmarks to assess LLVM performance in ADL tasks. Trained on ADL-X, LLAVIDAL achieves state-of-the-art performance across ADL benchmarks. Code and data will be made publicly available at https://adl-x.github.io/. Dominick Reilly, Rajatsubhra Chakraborty, Arkaprava Sinha, Manish Kumar Govind, Pu Wang 0001, François Brémond, Le Xue, Srijan Das |
CVPR | 6 |
| 2025 | Scaling Action Detection: AdaTAD++ with Transformer-Enhanced Temporal-Spatial Adaptation
Tanay Agrawal, Abid Ali 0002, Antitza Dantcheva, François Brémond |
ICCV | 4 |
| 2025 | Mixture of Experts Guided by Gaussian Splatters Matters: A New Approach to Weakly-Supervised Video Anomaly DetectionabstractVideo Anomaly Detection (VAD) is a challenging task due to the variability of anomalous events and the limited availability of labeled data. Under the Weakly-Supervised VAD (WSVAD) paradigm, only video-level labels are provided during training, while predictions are made at the frame level. Although state-of-the-art models perform well on simple anomalies (e.g., explosions), they struggle with complex real-world events (e.g., shoplifting). This difficulty stems from two key issues: (1) the inability of current models to address the diversity of anomaly types, as they process all categories with a shared model, overlooking category-specific features; and (2) the weak supervision signal, which lacks precise temporal information, limiting the ability to capture nuanced anomalous patterns blended with normal events. To address these challenges, we propose Gaussian Splatting-guided Mixture of Experts (GS-MoE), a novel framework that employs a set of expert models, each specialized in capturing specific anomaly types. These experts are guided by a temporal Gaussian splatting loss, enabling the model to leverage temporal consistency and enhance weak supervision. The Gaussian splatting approach encourages a more precise and comprehensive representation of anomalies by focusing on temporal segments most likely to contain abnormal events. The predictions from these specialized experts are integrated through a mixture-of-experts mechanism to model complex relationships across diverse anomaly patterns. Our approach achieves state-of-the-art performance, with a 91.58% AUC on the UCF-Crime dataset, and demonstrates superior results on XD-Violence and MSAD datasets. By leveraging category-specific expertise and temporal guidance, GS-MoE sets a new benchmark for VAD under weak supervision. Giacomo D'Amicantonio, Snehashis Majhi, Quan Kong, Lorenzo Garattoni, Gianpiero Francesca, François Brémond, Egor Bondarev |
ICCV | 6 |
| 2025 | Identifying Surgical Instruments in Pedagogical Cataract Surgery Videos through an Optimized Aggregation NetworkabstractInstructional cataract surgery videos are crucial for ophthalmologists and trainees to observe surgical details repeatedly. This paper presents a deep learning model for real-time identification of surgical instruments in these videos, using a custom dataset scraped from open-access sources. Inspired by the architecture of YOLOV9, the model employs a Programmable Gradient Information (PGI) mechanism and a novel Generally-Optimized Efficient Layer Aggregation Network (Go-ELAN) to address the information bottleneck problem, enhancing Minimum Average Precision (mAP) at higher Non-Maximum Suppression Intersection over Union (NMS IoU) scores. The Go-ELAN YOLOV9 model, evaluated against YOLO v5, v7, v8, v9 vanilla, Laptool and DETR, achieves a superior mAP of 73.74 at IoU 0.5 on a dataset of 615 images with 10 instrument classes, demonstrating the effectiveness of the proposed model. Sanya Sinha, Michal Balazia, François Brémond |
IPAS | 3 |
| 2025 | MultiMediate '25: Cross-cultural Multi-domain Engagement EstimationabstractEstimating momentary conversational engagement is central to assistive, socially aware AI systems, yet models are typically trained and evaluated within a single domain, limiting real-world robustness. The MultiMediate '25 challenge advances engagement estimation to more challenging, cross-cultural, and multi-domain settings. Building on prior challenge editions, we expand beyond NOXI as the sole training source by introducing NOXI-J, a new multilingual corpus covering Japanese and Chinese interactions, enabling both training and evaluation in diverse linguistic contexts. Although NOXI-J conceptually extends NOXI, we treat it as a distinct domain because linguistic, cultural, capture, and annotation differences induce measurable distribution shifts. In this paper, we present new annotations, precomputed multi-modal features (visual, vocal, and verbal), baseline evaluations, and an analysis of the best performing challenge solutions. Beyond accuracy, we quantify fairness using Conditional Demographic Disparity for gender and language. Our baselines confirm strong in-domain performance (e.g., paralinguistic eGeMAPS and video-transformer features) and reveal notable cross-domain drops, underscoring the challenge of cultural, linguistic, and interactional shifts. Fairness analyses indicate generally small discrepancies for our baselines. We observe the largest disparities for the proposed challenge solutions on the Chinese language test set. All annotations, features, code, and leaderboards are made publicly available to foster sustained progress on robust and fair engagement estimation. Daksitha Withanage, Marius Funk, Michal Balazia, Huajian Qiu, Shogo Okada, François Brémond, Jan Alexandersson, Andreas Bulling, Elisabeth André, Philipp Müller 0001 |
ACM Multimedia | 6 |
| 2025 | Loose Social-Interaction Recognition in Real-World Therapy ScenariosabstractThe computer vision community has explored dyadic interactions for atomic actions such as pushing, carrying-object, etc. However, with the advancement in deep learning models, there is a need to explore more complex dyadic situations such as loose interactions. These are interactions where two people perform certain atomic activities to complete a global action irrespective of temporal synchronisation and physical engagement, like cooking-together for example. Analysing these types of dyadic-interactions has several useful applications in the medical domain for social-skills development and mental health diagnosis. To achieve this, we propose a novel dual-path architecture to capture the loose interaction between two individuals. Our model learns global abstract features from each stream via a CNNs backbone and fuses them using a new Global-Layer-Attention module based on a cross-attention strategy. We evaluate our model on real-world autism diagnoses such as our Loose-Interaction dataset, and the publicly available Autism dataset for loose interactions. Our network achieves baseline results on the Loose-Interaction and SOTA results on the Autism datasets. Moreover, we study different social interactions by experimenting on a publicly available dataset i.e. NTU-RGB+D (interactive classes from both NTU-60 and NTU-120). We have found that different interactions require different network designs. We also compare a slightly different version of our method (details in Section 3.6) by incorporating time information to address tight interactions achieving SOTA results. Abid Ali 0002, Rui Dai 0001, Ashish Marisetty, Guillaume Astruc, Monique Thonnat, Jean-Marc Odobez, Susanne Thümmler, François Brémond |
WACV | 8 |
| 2025 | CM3T: Framework for Efficient Multimodal Learning for Inhomogeneous Interaction DatasetsabstractChallenges in cross-learning involve inhomogeneous or even inadequate amount of training data and lack of resources for retraining large pretrained models. Inspired by transfer learning techniques in NLP, adapters and prefix tuning, this paper presents a new model-agnostic plugin architecture for cross-learning, called CM3T, that adapts transformer-based models to new or missing information. We introduce two adapter blocks: multi-head vision adapters for transfer learning and cross-attention adapters for multimodal learning. Training becomes substantially efficient as the backbone and other plugins do not need to be finetuned along with these additions. Comparative and ablation studies on three datasets Epic-Kitchens-100, MPIIGroupInteraction and UDIVA v0.5 show efficacy of this framework on different recording settings and tasks. With only 12.8% trainable parameters compared to the backbone to process video input and only 22.3% trainable parameters for two additional modalities, we achieve comparable and even better results than the state-of-the-art. CM3T has no specific requirements for training or pretraining and is a step towards bridging the gap between a general model and specific practical applications of video classification. Tanay Agrawal, Mohammed Guermal, Michal Balazia, François Brémond |
WACV | 4 |
| 2025 | Guess Future Anomalies from Normalcy: Forecasting Abnormal Behavior in Real-World VideosabstractForecasting Abnormal Human Behavior (AHB) aims to predict unusual behavior in advance by analyzing early patterns of normal human interactions. Unlike typical action prediction methods, this task focuses on observing only normal interactions to predict both, short and long term future abnormal behavior. Despite its affirmative impact on society, AHB prediction remains under-explored in current research. This is primarily due to the challenges involved in anticipating complex human behaviors and interactions with surrounding agents in real-world situations. Further, there exists an underlying uncertainty between the early normal patterns and the future abnormal behavior, thereby making the prediction harder. To address these challenges, we introduce a novel transformer model that improves early interaction modeling by accounting for uncertainties in both, observations and future outcomes. To the best of our knowledge, we are the first to explore the task. Therefore, we present a new comprehensive dataset referred to as “AHB-F”††Code, Models, Dataset: https://github.com/snehashismajhi/AHB-F, which features real-world scenarios with complex human interactions. The AHB-F has a deterministic evaluation protocol that ensures only normal frames to be observed for long and short term future prediction. We extensively evaluate and compare competitive action anticipation methods on our benchmark. Our results show that our method consistently outperforms existing action anticipation approaches, both in quantitative and qualitative evaluations. Snehashis Majhi, Mohammed Guermal, Antitza Dantcheva, Quan Kong, Lorenzo Garattoni, Gianpiero Francesca, François Brémond |
WACV | 7 |
| 2025 | Anti-Forgetting Adaptation for Unsupervised Person Re-IdentificationabstractRegular unsupervised domain adaptive person re-identification (ReID) focuses on adapting a model from a source domain to a fixed target domain. However, an adapted ReID model can hardly retain previously-acquired knowledge and generalize to unseen data. In this paper, we propose a Dual-level Joint Adaptation and Anti-forgetting (DJAA) framework, which incrementally adapts a model to new domains without forgetting source domain and each adapted target domain. We explore the possibility of using prototype and instance-level consistency to mitigate the forgetting during the adaptation. Specifically, we store a small number of representative image samples and corresponding cluster prototypes in a memory buffer, which is updated at each adaptation step. With the buffered images and prototypes, we regularize the image-to-image similarity and image-to-prototype similarity to rehearse old knowledge. After the multi-step adaptation, the model is tested on all seen domains and several unseen domains to validate the generalization ability of our method. Extensive experiments demonstrate that our proposed method significantly improves the anti-forgetting, generalization and backward-compatible ability of an unsupervised person ReID model. Hao Chen 0061, François Brémond, Nicu Sebe, Shiliang Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | EEG classification with limited data: A deep clustering approach
Mohsen Tabejamaat, Hoda Mohammadzade, Farhood Negin, François Brémond |
Pattern Recognit. | 4 |
| 2024 | Weakly-Supervised Autism Severity Assessment in Long VideosabstractAutism Spectrum Disorder (ASD) is a diverse collection of neurobiological conditions marked by challenges in social communication and reciprocal interactions, as well as repetitive and stereotypical behaviors. Atypical behavior patterns in a long, untrimmed video can serve as biomarkers for children with ASD. In this paper, we propose a video-based weakly-supervised method that takes spatio-temporal features of long videos to learn typical and atypical behaviors for autism detection. On top of that, we propose a shallow TCN-MLP network, which is designed to further categorize the severity score. We evaluate our method on actual evaluation videos of children with autism collected and annotated (for severity score) by clinical professionals. Experimental results demonstrate the effectiveness of behaviors biomarkers that could help clinicians in autism spectrum analysis. Abid Ali 0002, Camilla Barbini, Séverine Dubuisson, Jean-Marc Odobez, François Brémond, Susanne Thümmler |
CBMI | 6 |
| 2024 | MultiMediate'24: Multi-Domain Engagement EstimationabstractEstimating the momentary level of participant's engagement is an important prerequisite for assistive systems that support human interactions. Previous work has addressed this task in within-domain evaluation scenarios, i.e. training and testing on the same dataset. This is in contrast to real-life scenarios where domain shifts between training and testing data frequently occur. With MultiMediate'24, we present the first challenge addressing multi-domain engagement estimation. As training data, we utilise the NOXI database of dyadic novice-expert interactions. In addition to within-domain test data, we add two new test domains. First, we introduce recordings following the NOXI protocol but covering languages that are not present in the NOXI training data. Second, we collected novel engagement annotations on the MPIIGroupInteraction dataset which consists of group discussions between three to four people. In this way, MultiMediate'24 evaluates the ability of approaches to generalise across factors such as language and cultural background, group size, task, and screen-mediated vs. face-to-face interaction. This paper describes the MultiMediate'24 challenge and presents baseline results. In addition, we discuss selected challenge solutions. Philipp Müller 0001, Michal Balazia, Tobias Baur 0001, Michael Dietz, Alexander Heimerl, Anna Penzkofer, Dominik Schiller, François Brémond, Jan Alexandersson, Elisabeth André, Andreas Bulling |
ACM Multimedia | 8 |
| 2024 | P-Age: Pexels Dataset for Robust Spatio-Temporal Apparent Age ClassificationabstractAge estimation is a challenging task that has numerous applications. In this paper, we propose a new direction for age classification that utilizes a video-based model to address challenges such as occlusions, low-resolution, and lighting conditions. To address these challenges, we propose AgeFormer which utilizes spatio-temporal information on the dynamics of the entire body dominating facebased methods for age classification. Our novel two-stream architecture uses TimeSformer and EfficientNet as backbones, to effectively capture both facial and body dynamics information for efficient and accurate age estimation in videos. Furthermore, to fill the gap in predicting age in real-world situations from videos, we construct a video dataset called Pexels Age (P-Age) for age classification. The proposed method achieves superior results compared to existing face-based age estimation methods and is evaluated in situations where the face is highly occluded, blurred, or masked. The method is also cross-tested on a variety of challenging video datasets such as Charades, Smarthome, and Thumos-14. The code and dataset is available at https://github.com/Ashish013/AgeFormer. Abid Ali 0002, Ashish Marisetty, François Brémond |
WACV | 3 |
| 2024 | JOADAA: joint online action detection and action anticipationabstractAction anticipation involves forecasting future actions by connecting past events to future ones. However, this reasoning ignores the real-life hierarchy of events which is considered to be composed of three main parts: past, present, and future. We argue that considering these three main parts and their dependencies could improve performance. On the other hand, online action detection is the task of predicting actions in a streaming manner. In this case, one has access only to the past and present information. Therefore, in online action detection (OAD) the existing approaches miss semantics or future information which limits their performance. To sum up, for both of these tasks, the complete set of knowledge (past-present-future) is missing, which makes it challenging to infer action dependencies, therefore having low performances. To address this limitation, we propose to fuse both tasks into a single uniform architecture. By combining action anticipation and online action detection, our approach can cover the missing dependencies of future information in online action detection. This method referred to as JOADAA, presents a uniform model that jointly performs action anticipation and online action detection. We validate our proposed model on three challenging datasets: THUMOS’14, which is a sparsely annotated dataset with one action per time step, CHARADES, and Multi-THUMOS, two densely annotated datasets with more complex scenarios. JOADAA achieves SOTA results on these benchmarks for both tasks. Mohammed Guermal, Abid Ali 0002, Rui Dai 0001, François Brémond |
WACV | 4 |
| 2024 | OE-CTST: Outlier-Embedded Cross Temporal Scale Transformer for Weakly-supervised Video Anomaly DetectionabstractVideo anomaly detection in real-world scenarios is challenging due to the complex temporal blending of long and short-length anomalies with normal ones. Further, it is more difficult to detect those due to : (i) Distinctive features characterizing the short and long anomalies with sharp and progressive temporal cues respectively; (ii) Lack of precise temporal information (i.e. weak-supervision) limits the temporal dynamics modeling of anomalies from normal events. In this paper, we propose a novel ‘temporal transformer’ framework for weakly-supervised anomaly detection: OE-CTST†. The proposed framework has two major components: (i) Outlier Embedder (OE) and (ii) Cross Temporal Scale Transformer (CTST). First, OE generates anomaly-aware temporal position encoding to allow the transformer to effectively model the temporal dynamics among the anomalies and normal events. Second, CTST encodes the cross-correlation between multi-temporal scale features to benefit short and long length anomalies by modeling the global temporal relations. The proposed OE-CTST is validated on three publicly available datasets i.e. UCF-Crime, XD-Violence, and IITB-Corridor, outperforming recently reported state-of-the-art approaches. Snehashis Majhi, Rui Dai 0001, Quan Kong, Lorenzo Garattoni, Gianpiero Francesca, François Brémond |
WACV | 6 |
| 2024 | Human-Scene Network: A novel baseline with self-rectifying loss for weakly supervised video anomaly detection
Snehashis Majhi, Rui Dai 0001, Quan Kong, Lorenzo Garattoni, Gianpiero Francesca, François Brémond |
Comput. Vis. Image Underst. | 6 |
| 2024 | UnWarpME: Unsupervised warping map estimation in real-world scenarios
Mohsen Tabejamaat, Farhood Negin, François Brémond |
Expert Syst. Appl. | 3 |
| 2024 | View-Invariant Skeleton Action Representation Learning via Motion Retargeting
Di Yang 0002, Yaohui Wang 0001, Antitza Dantcheva, Lorenzo Garattoni, Gianpiero Francesca, François Brémond |
Int. J. Comput. Vis. | 6 |
| 2024 | Improving texture integrity through second-order constraints on warping maps
Mohsen Tabejamaat, Farhood Negin, François Brémond |
Neurocomputing | 3 |
| 2024 | Synthetic Data in Human Analysis: A SurveyabstractDeep neural networks have become prevalent in human analysis, boosting the performance of applications, such as biometric recognition, action recognition, as well as person re-identification. However, the performance of such networks scales with the available training data. In human analysis, the demand for large-scale datasets poses a severe challenge, as data collection is tedious, time-expensive, costly and must comply with data protection laws. Current research investigates the generation of synthetic data as an efficient and privacy-ensuring alternative to collecting real data in the field. This survey introduces the basic definitions and methodologies, essential when generating and employing synthetic data for human analysis. We summarise current state-of-the-art methods and the main benefits of using synthetic data. We also provide an overview of publicly available synthetic datasets and generation models. Finally, we discuss limitations, as well as open research problems in this field. This survey is intended for researchers and practitioners in the field of human analysis. Indu Joshi, Marcel Grimmer, Christian Rathgeb, Christoph Busch 0001, François Brémond, Antitza Dantcheva |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | LIA: Latent Image AnimatorabstractPrevious animation techniques mainly focus on leveraging explicit structure representations (e.g., meshes or keypoints) for transferring motion from driving videos to source images. However, such methods are challenged with large appearance variations between source and driving data, as well as require complex additional modules to respectively model appearance and motion. Towards addressing these issues, we introduce the Latent Image Animator (LIA), streamlined to animate high-resolution images. LIA is designed as a simple autoencoder that does not rely on explicit representations. Motion transfer in the pixel space is modeled as linear navigation of motion codes in the latent space. Specifically such navigation is represented as an orthogonal motion dictionary learned in a self-supervised manner based on proposed Linear Motion Decomposition (LMD). Extensive experimental results demonstrate that LIA outperforms state-of-the-art on VoxCeleb, TaichiHD, and TED-talk datasets with respect to video quality and spatio-temporal consistency. In addition LIA is well equipped for zero-shot high-resolution image animation. Code, models, and demo video are available at https://github.com/wyhsirius/LIA. Yaohui Wang 0001, Di Yang 0002, François Brémond, Antitza Dantcheva |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Self-Supervised Video Representation Learning via Latent Time NavigationabstractSelf-supervised video representation learning aimed at maximizing similarity between different temporal segments of one video, in order to enforce feature persistence over time. This leads to loss of pertinent information related to temporal relationships, rendering actions such as `enter' and `leave' to be indistinguishable. To mitigate this limitation, we propose Latent Time Navigation (LTN), a time parameterized contrastive learning strategy that is streamlined to capture fine-grained motions. Specifically, we maximize the representation similarity between different video segments from one video, while maintaining their representations time-aware along a subspace of the latent representation code including an orthogonal basis to represent temporal changes. Our extensive experimental analysis suggests that learning video representations by LTN consistently improves performance of action classification in fine-grained and human-oriented tasks (e.g., on Toyota Smarthome dataset). In addition, we demonstrate that our proposed model, when pre-trained on Kinetics-400, generalizes well onto the unseen real world video benchmark datasets UCF101 and HMDB51, achieving state-of-the-art performance in action recognition. Di Yang 0002, Yaohui Wang 0001, Quan Kong, Antitza Dantcheva, Lorenzo Garattoni, Gianpiero Francesca, François Brémond |
AAAI | 7 |
| 2023 | Attributes-Aware Network for Temporal Action Detection
Rui Dai 0001, Srijan Das, Michael S. Ryoo, François Brémond |
BMVC | 4 |
| 2023 | LAC - Latent Action Composition for Skeleton-based Action SegmentationabstractSkeleton-based action segmentation requires recognizing composable actions in untrimmed videos. Current approaches decouple this problem by first extracting local visual features from skeleton sequences and then processing them by a temporal model to classify frame-wise actions. However, their performances remain limited as the visual features cannot sufficiently express composable actions. In this context, we propose Latent Action Composition (LAC)1, a novel self-supervised framework aiming at learning from synthesized composable motions for skeleton-based action segmentation. LAC is composed of a novel generation module towards synthesizing new sequences. Specifically, we design a linear latent space in the generator to represent primitive motion. New composed motions can be synthe-sized by simply performing arithmetic operations on latent representations of multiple input skeleton sequences. LAC leverages such synthesized sequences, which have large diversity and complexity, for learning visual representations of skeletons in both sequence and frame spaces via contrastive learning. The resulting visual encoder has a high expressive power and can be effectively transferred onto action segmentation tasks by end-to-end fine-tuning without the need for additional temporal models. We conduct a study focusing on transfer-learning and we show that representations learned from pre-trained LAC outperform the state-of-the-art by a large margin on TSU, Charades, PKU-MMD datasets. Di Yang 0002, Yaohui Wang 0001, Antitza Dantcheva, Quan Kong, Lorenzo Garattoni, Gianpiero Francesca, François Brémond |
ICCV | 7 |
| 2023 | MultiMediate '23: Engagement Estimation and Bodily Behaviour Recognition in Social InteractionsabstractAutomatic analysis of human behaviour is a fundamental prerequisite for the creation of machines that can effectively interact with- and support humans in social interactions. In MultiMediate'23, we address two key human social behaviour analysis tasks for the first time in a controlled challenge: engagement estimation and bodily behaviour recognition in social interactions. This paper describes the MultiMediate'23 challenge and presents novel sets of annotations for both tasks. For engagement estimation we collected novel annotations on the NOvice eXpert Interaction (NOXI) database. For bodily behaviour recognition, we annotated test recordings of the MPIIGroupInteraction corpus with the BBSI annotation scheme. In addition, we present baseline results for both challenge tasks. Philipp Müller 0001, Michal Balazia, Tobias Baur 0001, Michael Dietz, Alexander Heimerl, Dominik Schiller, Mohammed Guermal, Dominike Thomas, François Brémond, Jan Alexandersson, Elisabeth André, Andreas Bulling |
ACM Multimedia | 9 |
| 2023 | StressID: a Multimodal Dataset for Stress IdentificationabstractStressID is a new dataset specifically designed for stress identification fromunimodal and multimodal data. It contains videos of facial expressions, audiorecordings, and physiological signals. The video and audio recordings are acquiredusing an RGB camera with an integrated microphone. The physiological datais composed of electrocardiography (ECG), electrodermal activity (EDA), andrespiration signals that are recorded and monitored using a wearable device. Thisexperimental setup ensures a synchronized and high-quality multimodal data col-lection. Different stress-inducing stimuli, such as emotional video clips, cognitivetasks including mathematical or comprehension exercises, and public speakingscenarios, are designed to trigger a diverse range of emotional responses. Thefinal dataset consists of recordings from 65 participants who performed 11 tasks,as well as their ratings of perceived relaxation, stress, arousal, and valence levels.StressID is one of the largest datasets for stress identification that features threedifferent sources of data and varied classes of stimuli, representing more than39 hours of annotated data in total. StressID offers baseline models for stressclassification including a cleaning, feature extraction, and classification phase foreach modality. Additionally, we provide multimodal predictive models combiningvideo, audio, and physiological inputs. The data and the code for the baselines areavailable at https://project.inria.fr/stressid/. Hava Chaptoukaev, Valeriya Strizhkova, Michele Panariello, Bianca Dalpaos, Aglind Reka, Valeria Manera, Susanne Thümmler, Esma Ismailova, Nicholas W. D. Evans, François Brémond, Massimiliano Todisco, Maria A. Zuluaga, Laura M. Ferrari |
NeurIPS | 10 |
| 2023 | Multimodal Vision Transformers with Forced Attention for Behavior AnalysisabstractHuman behavior understanding requires looking at minute details in the large context of a scene containing multiple input modalities. It is necessary as it allows the design of more human-like machines. While transformer approaches have shown great improvements, they face multiple challenges such as lack of data or background noise. To tackle these, we introduce the Forced Attention (FAt) Transformer which utilize forced attention with a modified backbone for input encoding and a use of additional inputs. In addition to improving the performance on different tasks and inputs, the modification requires less time and memory resources. We provide a model for a generalised feature extraction for tasks concerning social signals and behavior analysis. Our focus is on understanding behavior in videos where people are interacting with each other or talking into the camera which simulates the first person point of view in social interaction. FAt Transformers are applied to two downstream tasks: personality recognition and body language recognition. We achieve state-of-the-art results for Udiva v0.5, First Impressions v2 and MPII Group Interaction datasets. We further provide an extensive ablation study of the proposed architecture. Tanay Agrawal, Michal Balazia, Philipp Müller 0001, François Brémond |
WACV | 4 |
| 2023 | Face attribute analysis from structured light: an end-to-end approach
Vikas Thamizharasan, Abhijit Das 0001, Daniele Battaglino, François Brémond, Antitza Dantcheva |
Multim. Tools Appl. | 4 |
| 2023 | Learning Invariance From Generated Variance for Unsupervised Person Re-IdentificationabstractThis work focuses on unsupervised representation learning in person re-identification (ReID). Recent self-supervised contrastive learning methods learn invariance by maximizing the representation similarity between two augmented views of a same image. However, traditional data augmentation may bring to the fore undesirable distortions on identity features, which is not always favorable in id-sensitive ReID tasks. In this article, we propose to replace traditional data augmentation with a generative adversarial network (GAN) that is targeted to generate augmented views for contrastive learning. A 3D mesh guided person image generator is proposed to disentangle a person image into id-related and id-unrelated features. Deviating from previous GAN-based ReID methods that only work in id-unrelated space (pose and camera style), we conduct GAN-based augmentation on both id-unrelated and id-related features. We further propose specific contrastive losses to help our network learn invariance from id-unrelated and id-related augmentations. By jointly training the generative and the contrastive modules, our method achieves new state-of-the-art unsupervised person ReID performance on mainstream large-scale benchmarks. Hao Chen 0061, Yaohui Wang 0001, Benoit Lagadec, Antitza Dantcheva, François Brémond |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Toyota Smarthome Untrimmed: Real-World Untrimmed Videos for Activity DetectionabstractDesigning activity detection systems that can be successfully deployed in daily-living environments requires datasets that pose the challenges typical of real-world scenarios. In this paper, we introduce a new untrimmed daily-living dataset that features several real-world challenges: Toyota Smarthome Untrimmed (TSU). TSU contains a wide variety of activities performed in a spontaneous manner. The dataset contains dense annotations including elementary, composite activities and activities involving interactions with objects. We provide an analysis of the real-world challenges featured by our dataset, highlighting the open issues for detection algorithms. We show that current state-of-the-art methods fail to achieve satisfactory performance on the TSU dataset. Therefore, we propose a new baseline method for activity detection to tackle the novel challenges provided by our dataset. This method leverages one modality (i.e. optic flow) to generate the attention weights to guide another modality (i.e RGB) to better detect the activity boundaries. This is particularly beneficial to detect activities characterized by high temporal variance. We show that the method we propose outperforms state-of-the-art methods on TSU and on another popular challenging dataset, Charades. Rui Dai 0001, Srijan Das, Saurav Sharma, Luca Minciullo, Lorenzo Garattoni, François Brémond, Gianpiero Francesca |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | MS-TCT: Multi-Scale Temporal ConvTransformer for Action DetectionabstractAction detection is a significant and challenging task, especially in densely-labelled datasets of untrimmed videos. Such data consist of complex temporal relations including composite or co-occurring actions. To detect actions in these complex settings, it is critical to capture both shortterm and long-term temporal information efficiently. To this end, we propose a novel ‘ConvTransformer’ network for action detection: MS-TCT11Code/Models: https://github.com/dairui01/MS-TCT. This network comprises of three main components: (1) a Temporal Encoder module which explores global and local temporal relations at multiple temporal resolutions, (2) a Temporal Scale Mixer module which effectively fuses multi-scale features, creating a unified feature representation, and (3) a Classification module which learns a center-relative position of each action instance in time, and predicts frame-level classification scores. Our experimental results on multiple challenging datasets such as Charades, TSU and MultiTHUMOS, validate the effectiveness of the proposed method, which outperforms the state-of-the-art methods on all three datasets. Rui Dai 0001, Srijan Das, Kumara Kahatapitiya, Michael S. Ryoo, François Brémond |
CVPR | 5 |
| 2022 | Latent Image Animator: Learning to Animate Images via Latent Space Navigation
Yaohui Wang 0001, Di Yang 0002, François Brémond, Antitza Dantcheva |
ICLR | 3 |
| 2022 | THORN: Temporal Human-Object Relation Network for Action RecognitionabstractMost action recognition models treat human activities as unitary events. However, human activities often follow a certain hierarchy. In fact, many human activities are compositional. Also, these actions are mostly human-object interactions. In this paper we propose to recognize human action by leveraging the set of interactions that define an action. In this work, we present an end-to-end network: THORN, that can leverage important human-object and object-object interactions to predict actions. This model is built on top of a 3D backbone network. The key components of our model are: 1) An object representation filter for modeling object. 2) An object relation reasoning module to capture object relations. 3) A classification layer to predict the action labels. To show the robustness of THORN, we evaluate it on EPIC-Kitchen55 and EGTEA Gaze+, two of the largest and most challenging first-person and human-object interaction datasets. THORN achieves state-of-the-art performance on both datasets. Mohammed Guermal, Rui Dai 0001, François Brémond |
ICPR | 3 |
| 2022 | Bodily Behaviors in Social Interaction: Novel Annotations and State-of-the-Art EvaluationabstractBody language is an eye-catching social signal and its automatic analysis can significantly advance artificial intelligence systems to understand and actively participate in social interactions. While computer vision has made impressive progress in low-level tasks like head and body pose estimation, the detection of more subtle behaviors such as gesturing, grooming, or fumbling is not well explored. In this paper we present BBSI, the first set of annotations of complex Bodily Behaviors embedded in continuous Social Interactions in a group setting. Based on previous work in psychology, we manually annotated 26 hours of spontaneous human behavior in the MPIIGroupInteraction dataset with 15 distinct body language classes. We present comprehensive descriptive statistics on the resulting dataset as well as results of annotation quality evaluations. For automatic detection of these behaviors, we adapt the Pyramid Dilated Attention Network (PDAN), a state-of-the-art approach for human action detection. We perform experiments using four variants of spatial-temporal features as input to PDAN: Two-Stream Inflated 3D CNN, Temporal Segment Networks, Temporal Shift Module and Swin Transformer. Results are promising and indicate a great room for improvement in this difficult task. Representing a key piece in the puzzle towards automatic understanding of social behavior, BBSI is fully available to the research community. Michal Balazia, Philipp Müller 0001, Ákos Levente Tánczos, August von Liechtenstein, François Brémond |
ACM Multimedia | 5 |
| 2022 | VPN++: Rethinking Video-Pose Embeddings for Understanding Activities of Daily LivingabstractMany attempts have been made towards combining RGB and 3D poses for the recognition of Activities of Daily Living (ADL). ADL may look very similar and often necessitate to model fine-grained details to distinguish them. Because the recent 3D ConvNets are too rigid to capture the subtle visual patterns across an action, this research direction is dominated by methods combining RGB and 3D Poses. But the cost of computing 3D poses from RGB stream is high in the absence of appropriate sensors. This limits the usage of aforementioned approaches in real-world applications requiring low latency. Then, how to best take advantage of 3D Poses for recognizing ADL? To this end, we propose an extension of a pose driven attention mechanism: Video-Pose Network (VPN), exploring two distinct directions. One is to transfer the Pose knowledge into RGB through a feature-level distillation and the other towards mimicking pose driven attention through an attention-level distillation. Finally, these two approaches are integrated into a single model, we call VPN++. It is worth noting that VPN++ exploits the pose embeddings at training via distillation but not at inference. We show that VPN++ is not only effective but also provides a high speed up and high resilience to noisy Poses. VPN++, with or without 3D Poses, outperforms the representative baselines on 4 public datasets. Code is available at https://github.com/srijandas07/vpnplusplus. Srijan Das, Rui Dai 0001, Di Yang 0002, François Brémond |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | A Spatio-Temporal Approach for Apathy ClassificationabstractApathy is characterized by symptoms such as reduced emotional response, lack of motivation, and limited social interaction. Current methods for apathy diagnosis require the patient’s presence in a clinic and time consuming clinical interviews, which are costly and inconvenient for both, patients and clinical staff, hindering among other large-scale diagnostics. In this work, we propose a novel spatio-temporal framework for apathy classification, which is streamlined to analyze facial dynamics and emotion in videos. Specifically, we divide the videos into smaller clips, and proceed to extract associated facial dynamics and emotion-based features. Statistical representations/descriptors based on each feature and clip serve as input of the proposed Gated Recurrent Unit (GRU)-architecture. Temporal representations of individual features at the lower level of the proposed architecture are combined at deeper layers of the proposed GRU architecture, in order to obtain the final feature-set for apathy classification. Based on extensive experiments, we show that fusion of characteristics such as emotion and facial dynamics in proposed deep-bi-directional GRU obtains an accuracy of 95.34% in apathy classification. Abhijit Das 0001, Xuesong Niu, Antitza Dantcheva, S. L. Happy, Hu Han 0001, Radia Zeghari, Philippe Robert, Shiguang Shan, François Brémond, Xilin Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2021 | From Multimodal to Unimodal Attention in Transformers using Knowledge DistillationabstractMultimodal Deep Learning has garnered much interest, and transformers have triggered novel approaches, thanks to the cross-attention mechanism. Here we propose an approach to deal with two key existing challenges: the high computational resource demanded and the issue of missing modalities. We introduce for the first time the concept of knowledge distillation in transformers to use only one modality at inference time. We report a full study analyzing multiple student-teacher configurations, levels at which distillation is applied, and different methodologies. With the best configuration, we improved the state-of-the-art accuracy by 3%, we reduced the number of parameters by 2.5 times and the inference time by 22%. Such performance-computation tradeoff can be exploited in many applications and we aim at opening a new research area where the deployment of complex models with limited resources is demanded Dhruv Agarwal 0002, Tanay Agrawal, Laura M. Ferrari, François Brémond |
AVSS | 4 |
| 2021 | DAM: Dissimilarity Attention Module for Weakly-supervised Video Anomaly DetectionabstractVideo anomaly detection under weak supervision is complicated due to the difficulties in identifying the anomaly and normal instances during training, hence, resulting in non-optimal margin of separation. In this paper, we propose a framework consisting of Dissimilarity Attention Module (DAM) to discriminate the anomaly instances from normal ones both at feature level and score level. In order to decide instances to be normal or anomaly, DAM takes local spatio-temporal (i.e. clips within a video) dissimilarities into account rather than the global temporal context of a video. This allows the framework to detect anomalies in real-time (i.e. online) scenarios without the need of extra window buffer time. Further more, we adopt two-variants of DAM for learning the dissimilarities between successive video clips. The proposed framework along with DAM is validated on two large scale anomaly detection datasets i.e. UCF-Crime and ShanghaiTech, outperforming the online state-of-the-art approaches by 1.5% and 3.4% respectively. The source code and models will be available at https://github.com/snehashismajhi/DAM-Anomaly-Detection Snehashis Majhi, Srijan Das, François Brémond |
AVSS | 3 |
| 2021 | TrichTrack: Multi-Object Tracking of Small-Scale Trichogramma WaspsabstractTrichogramma wasps behaviors are studied extensively due to their effectiveness as biological control agents across the globe. However, to our knowledge, the field of intra/inter-species Trichogramma behavior is yet to be explored thoroughly. To study these behaviors it is crucial to identify and track Trichogramma individuals over a long period in a lab setup. For this, we propose a robust tracking pipeline named TrichTrack. Due to the unavailability of labeled data, we train our detector using an iterative weakly supervised method. We also use a weakly supervised method to train a Re-Identification (ReID) network by leveraging noisy tracklet sampling. This enables us to distinguish Trichogramma individuals that are indistinguishable from human eyes. We also develop a two-staged tracking module that filters out the easy association to improve its efficiency. Our method outperforms existing insect trackers on most of the MOTMetrics, specifically on ID switches and fragmentations. Vishal Pani, Martin Bernet, Vincent Calcagno, Louise Van Oudenhove, François Brémond |
AVSS | 5 |
| 2021 | FLAME: Facial Landmark Heatmap Activated Multimodal Gaze Estimationabstract3D gaze estimation is about predicting the line of sight of a person in 3D space. Person-independent models for the same lack precision due to anatomical differences of subjects, whereas person-specific calibrated techniques add strict constraints on scalability. To overcome these issues, we propose a novel technique, Facial Landmark Heatmap Activated Multimodal Gaze Estimation (FLAME), as a way of combining eye anatomical information using eye land-mark heatmaps to obtain precise gaze estimation without any person-specific calibration. Our evaluation demonstrates a competitive performance of about 10% improvement on benchmark datasets ColumbiaGaze and EYEDIAP. We also conduct an ablation study to validate our method. Neelabh Sinha, Michal Balazia, François Brémond |
AVSS | 3 |
| 2021 | CTRN: Class-Temporal Relational Network for Action Detection
Rui Dai 0001, Srijan Das, François Brémond |
BMVC | 3 |
| 2021 | Guided Flow Field Estimation by Generating Independent Patches
Mohsen Tabejamaat, Farhood Negin, François Brémond |
BMVC | 3 |
| 2021 | UNIK: A Unified Framework for Real-world Skeleton-based Action Recognition
Di Yang 0002, Yaohui Wang 0001, Antitza Dantcheva, Lorenzo Garattoni, Gianpiero Francesca, François Brémond |
BMVC | 6 |
| 2021 | Joint Generative and Contrastive Learning for Unsupervised Person Re-IdentificationabstractRecent self-supervised contrastive learning provides an effective approach for unsupervised person re-identification (ReID) by learning invariance from different views (transformed versions) of an input. In this paper, we incorporate a Generative Adversarial Network (GAN) and a contrastive learning module into one joint training framework. While the GAN provides online data augmentation for contrastive learning, the contrastive module learns view-invariant features for generation. In this context, we propose a mesh-based view generator. Specifically, mesh projections serve as references towards generating novel views of a person. In addition, we propose a view-invariant loss to facilitate contrastive learning between original and generated views. Deviating from previous GAN-based unsupervised ReID methods involving domain adaptation, we do not rely on a labeled source dataset, which makes our method more flexible. Extensive experimental results show that our method significantly outperforms state-of-the-art methods under both, fully unsupervised and unsupervised domain adaptive settings on several large scale ReID dat-sets. Source code and models are available under https://github.com/chenhao2345/GCL. Hao Chen 0061, Yaohui Wang 0001, Benoit Lagadec, Antitza Dantcheva, François Brémond |
CVPR | 5 |
| 2021 | Weakly-supervised Joint Anomaly Detection and ClassificationabstractAnomaly activities such as robbery, explosion, accidents, etc. need immediate actions for preventing loss of human life and property in real world surveillance systems. Although the recent automation in surveillance systems are capable of detecting the anomalies, but they still need human efforts for categorizing the anomalies and taking necessary preventive actions. This is due to the lack of methodology performing both anomaly detection and classification for real world scenarios. Thinking of a fully automatized surveillance system, which is capable of both detecting and classifying the anomalies that need immediate actions, a joint anomaly detection and classification method is a pressing need. The task of joint detection and classification of anomalies becomes challenging due to the unavailability of dense annotated videos pertaining to anomalous classes, which is a crucial factor for training modern deep architecture. Furthermore, doing it through manual human effort seems impossible. Thus, we propose a method that jointly handles the anomaly detection and classification in a single framework by adopting a weakly-supervised learning paradigm. In weakly-supervised learning instead of dense temporal annotations, only video-level labels are sufficient for learning. The proposed model is validated on a large-scale publicly available UCF-Crime dataset, achieving state-of-the-art results. The source code and models will be available at https://github.com/snehashismajhi/JointDetectClassify. Snehashis Majhi, Srijan Das, François Brémond, Ratnakar Dash, Pankaj Kumar Sa |
FG | 3 |
| 2021 | Emotion Editing in Head Reenactment Videos using Latent Space ManipulationabstractVideo generation greatly benefits from integrating facial expressions, as they are highly pertinent in social interaction and hence increase realism in generated talking head videos. Motivated by this, we propose a method for editing emotions in head reenactment videos that is streamlined to modify the latent space of a pre-trained neural head reenactment system. Specifically, our method seeks to disentangle emotions from the latent pose and identity representation. The proposed learning process is based on cycle consistency and image reconstruction losses. Our results suggest that despite its simplicity, such learning successfully decomposes emotion from pose and identity. Our method reproduces facial mimics of a person from a driving video, as well as allows for emotion editing in the reenactment video. We compare our method to the state-of-art for altering emotions in reenactment videos, producing more realistic results that the state-of-art. Valeriya Strizhkova, Yaohui Wang 0001, David Anghelone, Di Yang 0002, Antitza Dantcheva, François Brémond |
FG | 6 |
| 2021 | Self-Supervised Video Pose Representation Learning for Occlusion- Robust Action RecognitionabstractAction recognition based on human pose has witnessed increasing attention due to its robustness to changes in appearances, environments, and view-points. Despite associated progress, one remaining challenge has to do with occlusion in real-world videos that hinders the visibility of all joints. Such occlusion impedes representation of such scenes by models that have been trained on full-body pose data, obtained in laboratory conditions with specific sensors. To address this, as a first contribution, we introduce OR- VPE, a novel video pose embedding network that is streamlined to learn an occlusion-robust representation for pose sequences in videos. In order to enable our embedding network to handle partially visible joints, we propose to incorporate a sub-graph data augmentation mechanism during training, which simulates occlusions, into a video pose encoder based on Graph Convolutional Networks (GCNs). As a second contribution, we apply a contrastive learning module to train the video pose representation in a self-supervised manner without the necessity of action annotations. This is achieved by maximizing the mutual information of the same pose sequence pruned into different spatio-temporal subgraphs. Experimental analyses show that compared to training the same encoder from scratch, our proposed OR-VPE, with pre-training on a large-scale dataset, NTU-RGB+D 120, improves the performance of the downstream action classification on Toyota Smarthome, N-UCLA and Penn Action datasets. Di Yang 0002, Yaohui Wang 0001, Antitza Dantcheva, Lorenzo Garattoni, Gianpiero Francesca, François Brémond |
FG | 6 |
| 2021 | ICE: Inter-instance Contrastive Encoding for Unsupervised Person Re-identificationabstractUnsupervised person re-identification (ReID) aims at learning discriminative identity features without annotations. Recently, self-supervised contrastive learning has gained increasing attention for its effectiveness in unsupervised representation learning. The main idea of instance contrastive learning is to match a same instance in different augmented views. However, the relationship between different instances has not been fully explored in previous contrastive methods, especially for instance-level contrastive loss. To address this issue, we propose Interinstance Contrastive Encoding (ICE) that leverages interinstance pairwise similarity scores to boost previous classlevel contrastive ReID methods. We first use pairwise similarity ranking as one-hot hard pseudo labels for hard instance contrast, which aims at reducing intra-class variance. Then, we use similarity scores as soft pseudo labels to enhance the consistency between augmented and original views, which makes our model more robust to augmentation perturbations. Experiments on several large-scale person ReID datasets validate the effectiveness of our proposed unsupervised method ICE, which is competitive with even supervised methods. Code is made available at https://github.com/chenhao2345/ICE. Hao Chen 0061, Benoit Lagadec, François Brémond |
ICCV | 3 |
| 2021 | Learning an Augmented RGB Representation with Cross-Modal Knowledge Distillation for Action DetectionabstractIn video understanding, most cross-modal knowledge distillation (KD) methods are tailored for classification tasks, focusing on the discriminative representation of the trimmed videos. However, action detection requires not only categorizing actions, but also localizing them in untrimmed videos. Therefore, transferring knowledge pertaining to temporal relations is critical for this task which is missing in the previous cross-modal KD frameworks. To this end, we aim at learning an augmented RGB representation for action detection, taking advantage of additional modalities at training time through KD. We propose a KD frame-work consisting of two levels of distillation. On one hand, atomic-level distillation encourages the RGB student to learn the sub-representation of the actions from the teacher in a contrastive manner. On the other hand, sequence-level distillation encourages the student to learn the temporal knowledge from the teacher, which consists of transferring the Global Contextual Relations and the Action Boundary Saliency. The result is an Augmented-RGB stream that can achieve competitive performance as the two-stream network while using only RGB at inference time. Extensive experimental analysis shows that our proposed distillation frame-work is generic and outperforms other popular cross-modal distillation methods in action detection task. Rui Dai 0001, Srijan Das, François Brémond |
ICCV | 3 |
| 2021 | Enhancing Diversity in Teacher-Student Networks via Asymmetric branches for Unsupervised Person Re-identificationabstractThe objective of unsupervised person re-identification (Re-ID) is to learn discriminative features without laborintensive identity annotations. State-of-the-art unsupervised Re-ID methods assign pseudo labels to unlabeled images in the target domain and learn from these noisy pseudo labels. Recently introduced Mean Teacher Model is a promising way to mitigate the label noise. However, during the training, self-ensembled teacher-student networks quickly converge to a consensus which leads to a local minimum. We explore the possibility of using an asymmetric structure inside neural network to address this problem. First, asymmetric branches are proposed to extract features in different manners, which enhances the feature diversity in appearance signatures. Then, our proposed cross-branch supervision allows one branch to get supervision from the other branch, which transfers distinct knowledge and enhances the weight diversity between teacher and student networks. Extensive experiments show that our proposed method can significantly surpass the performance of previous work on both unsupervised domain adaptation and fully unsupervised Re-ID tasks. Hao Chen 0061, Benoit Lagadec, François Brémond |
WACV | 3 |
| 2021 | PDAN: Pyramid Dilated Attention Network for Action DetectionabstractHandling long and complex temporal information is an important challenge for action detection tasks. This challenge is further aggravated by densely distributed actions in untrimmed videos. Previous action detection methods fail in selecting the key temporal information in long videos. To this end, we introduce the Dilated Attention Layer (DAL). Compared to previous temporal convolution layer, DAL allocates attentional weights to local frames in the kernel, which enables it to learn better local representation across time. Furthermore, we introduce Pyramid Dilated Attention Network (PDAN) which is built upon DAL. With the help of multiple DALs with different dilation rates, PDAN can model short-term and long-term temporal relations simultaneously by focusing on local segments at the level of low and high temporal receptive fields. This property enables PDAN to handle complex temporal relations between different action instances in long untrimmed videos. To corroborate the effectiveness and robustness of our method, we evaluate it on three densely annotated, multi-label datasets: Mul-tiTHUMOS, Charades and Toyota Smarthome Untrimmed (TSU) dataset. PDAN is able to outperform previous state-of-the-art methods on all these datasets."Time abides long enough for those who make use of it. Rui Dai 0001, Srijan Das, Luca Minciullo, Lorenzo Garattoni, Gianpiero Francesca, François Brémond |
WACV | 6 |
| 2021 | Selective Spatio-Temporal Aggregation Based Pose Refinement System: Towards Understanding Human Activities in Real-World VideosabstractTaking advantage of human pose data for understanding human activities has attracted much attention these days. However, state-of-the-art pose estimators struggle in obtaining high-quality 2D or 3D pose data due to occlusion, truncation and low-resolution in real-world un-annotated videos. Hence, in this work, we propose 1) a Selective Spatio-Temporal Aggregation mechanism, named SST-A, that refines and smooths the keypoint locations extracted by multiple expert pose estimators, 2) an effective weakly-supervised self-training framework which leverages the aggregated poses as pseudo ground-truth in-stead of handcrafted annotations for real-world pose estimation. Extensive experiments are conducted for evaluating not only the upstream pose refinement but also the downstream action recognition performance on four datasets, Toyota Smarthome, NTU-RGB+D, Charades, and Kinetics-50. We demonstrate that the skeleton data refined by our Pose-Refinement system (SSTA-PRS) is effective at boosting various existing action recognition models, which achieves competitive or state-of-the-art performance. Di Yang 0002, Rui Dai 0001, Yaohui Wang 0001, Rupayan Mallick, Luca Minciullo, Gianpiero Francesca, François Brémond |
WACV | 7 |
| 2021 | Expression recognition with deep features extracted from holistic and part-based models
S. L. Happy, Antitza Dantcheva, François Brémond |
Image Vis. Comput. | 3 |
| 2020 | G3AN: Disentangling Appearance and Motion for Video GenerationabstractCreating realistic human videos entails the challenge of being able to simultaneously generate both appearance, as well as motion. To tackle this challenge, we introduce G3AN, a novel spatio-temporal generative model, which seeks to capture the distribution of high dimensional video data and to model appearance and motion in disentangled manner. The latter is achieved by decomposing appearance and motion in a three-stream Generator, where the main stream aims to model spatio-temporal consistency, whereas the two auxiliary streams augment the main stream with multi-scale appearance and motion features, respectively. An extensive quantitative and qualitative analysis shows that our model systematically and significantly outperforms state-of-the-art methods on the facial expression datasets MUG and UvA-NEMO, as well as the Weizmann and UCF101 datasets on human action. Additional analysis on the learned latent representations confirms the successful decomposition of appearance and motion. Yaohui Wang 0001, Piotr Bilinski, François Brémond, Antitza Dantcheva |
CVPR | 3 |
| 2020 | VPN: Learning Video-Pose Embedding for Activities of Daily Living
Srijan Das, Saurav Sharma, Rui Dai 0001, François Brémond, Monique Thonnat |
ECCV (9) | 4 |
| 2020 | Semi-supervised Emotion Recognition using Inconsistently Annotated DataabstractExpression recognition remains challenging, predominantly due to (a) lack of sufficient data, (b) subtle emotion intensity, (c) subjective and inconsistent annotation, as well as due to (d) in-the-wild data containing variations in pose, intensity, and occlusion. To address such challenges in a unified framework, we propose a self-training based semi-supervised convolutional neural network (CNN) framework, which directly addresses the problem of (a) limited data by leveraging information from unannotated samples. Our method uses `successive label smoothing' to adapt to the subtle expressions and improve the model performance for (b) low-intensity expression samples. Further, we address (c) inconsistent annotations by assigning sample weights during loss computation, thereby ignoring the effect of incorrect ground-truth. We observe significant performance improvement in in-the-wild datasets by leveraging the information from the in-the-lab datasets, related to challenge (d). Associated to that, experiments on four publicly available datasets demonstrate large performance gains in cross-database performance, as well as show that the proposed method achieves to learn different expression intensities, even when trained with categorical samples. S. L. Happy, Antitza Dantcheva, François Brémond |
FG | 3 |
| 2020 | Apathy Classification by Exploiting Task RelatednessabstractApathy is characterized by symptoms such as reduced emotional response, lack of motivation, and limited social interaction. Current methods for apathy diagnosis require the patient's presence in a clinic and time consuming clinical interviews, which are costly and inconvenient for both patients and clinical staff, hindering among others large-scale diagnostics. In this work we propose a multi-task learning (MTL) framework for apathy classification based on facial analysis, entailing both emotion and facial movements. In addition, it leverages information from other auxiliary tasks (i.e., clinical scores), which might be closely or distantly related to the main task of apathy classification. Our proposed MTL approach (termed MTL+) improves apathy classification by jointly learning model weights and the relatedness of the auxiliary tasks to the main task in an iterative manner. Our results on 90 video sequences acquired from 45 subjects obtained an apathy classification accuracy of up to 80%, using the concatenated emotion and motion features. Our results further demonstrate the improved performance of MTL+ over MTL. S. L. Happy, Antitza Dantcheva, Abhijit Das 0001, François Brémond, Radia Zeghari, Philippe Robert |
FG | 4 |
| 2020 | How Unique Is a Face: An Investigative StudyabstractFace recognition has been widely accepted as a means of identification in applications ranging from border control to security in the banking sector. Surprisingly, while widely accepted, we still lack the understanding of uniqueness or distinctiveness of faces as biometric modality. In this work, we study the impact of factors such as image resolution, feature representation, database size, age and gender on uniqueness denoted by the Kullback-Leibler divergence between genuine and impostor distributions. Towards understanding the impact, we present experimental results on the datasets AT&T, LFW, IMDb-Face, as well as ND-TWINS, with the feature extraction algorithms VGGFace, VGG16, ResNet50, InceptionV3, MobileNet and DenseNet121, that reveal the quantitative impact of the named factors. While these are early results, our findings indicate the need for a better understanding of the concept of biometric uniqueness and its implication on face recognition. Michal Balazia, S. L. Happy, François Brémond, Antitza Dantcheva |
ICPR | 3 |
| 2020 | Learning Discriminative and Generalizable Representations by Spatial-Channel Partition for Person Re-IdentificationabstractIn Person Re-Identification (Re-ID) task, combining local and global features is a common strategy to overcome missing key parts and misalignment on models based only on global features. Using this combination, neural networks yield impressive performance in Re-ID task. Previous part-based models mainly focus on spatial partition strategies. Recently, operations on channel information, such as Group Normalization and Channel Attention, have brought significant progress to various visual tasks. However, channel partition has not drawn much attention in Person Re-ID. In this paper, we conduct a study to exploit the potential of channel partition in Re-ID task. Based on this study, we propose an end-to-end Spatial and Channel partition Representation network (SCR) in order to better exploit both spatial and channel information. Experiments conducted on three mainstream image-based evaluation protocols including Market-1501, DukeMTMC-ReID and CUHK03 and one video-based evaluation protocol MARS validate the performance of our model, which outperforms previous state-of- the-art in both single and cross domain Re-ID tasks. Hao Chen 0061, Benoit Lagadec, François Brémond |
WACV | 3 |
| 2020 | Looking deeper into Time for Activities of Daily Living RecognitionabstractIn this paper, we introduce a new approach for Activities of Daily Living (ADL) recognition. In order to discriminate between activities with similar appearance and motion, we focus on their temporal structure. Actions with subtle and similar motion are hard to disambiguate since long-range temporal information is hard to encode. So, we propose an end-to-end Temporal Model to incorporate long-range temporal information without losing subtle details. The temporal structure is represented globally by different temporal granularities and locally by temporal segments. We also propose a two-level pose driven attention mechanism to take into account the relative importance of the segments and granularities. We validate our approach on 2 public datasets: a 3D human activity dataset (NTU-RGB+D) and a human-object interaction dataset (Northwestern-UCLA Multiview Action 3D). Our Temporal Model can also be incorporated with any existing 3D CNN (including attention based) as a backbone which reveals its robustness. Srijan Das, Monique Thonnat, François Brémond |
WACV | 3 |
| 2020 | A One-and-Half Stage Pedestrian DetectorabstractPedestrian detection is a specific instance of the more general problem of object detection in computer vision. A balance between detection accuracy and speed is a desirable trait for pedestrian detection systems in many applications such as self-driving cars. In this paper, we follow the wisdom of " and less is often more" to achieve this balance. We propose a lightweight mechanism based on semantic segmentation to reduce the number of anchors to be processed. We furthermore unify this selection with the intra-anchor feature pooling strategy adopted in high performance two-stage detectors such as Faster-RCNN. Such a strategy is avoided in one-stage detectors like SSD in favour of faster inference but at the cost of reducing the accuracy vis-à-vis two-stage detectors. However our anchor selection renders it practical to use feature pooling without giving up the inference speed.Our proposed approach succeeds in detecting pedestrians with state-of-art performance on caltech-reasonable and ciypersons datasets with inference speeds of ~ 32 fps. Ujjwal, Aziz Dziri, Bertrand Leroy, François Brémond |
WACV | 4 |
| 2020 | ImaGINator: Conditional Spatio-Temporal GAN for Video GenerationabstractGenerating human videos based on single images entails the challenging simultaneous generation of realistic and visual appealing appearance and motion. In this context, we propose a novel conditional GAN architecture, namely ImaGINator, which given a single image, a condition (label of a facial expression or action) and noise, decomposes appearance and motion in both latent and high level feature spaces, generating realistic videos. This is achieved by (i) a novel spatio-temporal fusion scheme, which generates dynamic motion, while retaining appearance throughout the full video sequence by transmitting appearance (originating from the single image) through all layers of the network. In addition, we propose (ii) a novel transposed (1+2)D convolution, factorizing the transposed 3D convolutional filters into separate transposed temporal and spatial components, which yields significantly gains in video quality and speed. We extensively evaluate our approach on the facial expression datasets MUG and UvA-NEMO, as well as on the action datasets NATOPS and Weizmann. We show that our approach achieves significantly better quantitative and qualitative results than the state-of-the-art. The source code and models are available under https://github.com/wyhsirius/ImaGINator. Yaohui Wang 0001, Piotr Bilinski, François Brémond, Antitza Dantcheva |
WACV | 3 |
| 2019 | SkeleMotion: A New Representation of Skeleton Joint Sequences based on Motion Information for 3D Action RecognitionabstractDue to the availability of large-scale skeleton datasets, 3D human action recognition has recently called the attention of computer vision community. Many works have focused on encoding skeleton data as skeleton image representations based on spatial structure of the skeleton joints, in which the temporal dynamics of the sequence is encoded as variations in columns and the spatial structure of each frame is represented as rows of a matrix. To further improve such representations, we introduce a novel skeleton image representation to be used as input of Convolutional Neural Networks (CNNs), named SkeleMotion. The proposed approach encodes the temporal dynamics by explicitly computing the magnitude and orientation values of the skeleton joints. Different temporal scales are employed to compute motion values to aggregate more temporal dynamics to the representation making it able to capture long-range joint interactions involved in actions as well as filtering noisy motion values. Experimental results demonstrate the effectiveness of the proposed representation on 3D action recognition outperforming the state-of-the-art on NTU RGB+D 120 dataset. Carlos Antônio Caetano Jr., Jessica Sena, François Brémond, Jefersson A. dos Santos, William Robson Schwartz |
AVSS | 3 |
| 2019 | Self-Attention Temporal Convolutional Network for Long-Term Daily Living Activity DetectionabstractIn this paper, we address the detection of daily living activities in long-term untrimmed videos. The detection of daily living activities is challenging due to their long temporal components, low inter-class variation and high intra-class variation. To tackle these challenges, recent approaches based on Temporal Convolutional Networks (TCNs) have been proposed. Such methods can capture long-term temporal patterns using a hierarchy of temporal convolutional filters, pooling and up sampling steps. However, as one of the important features of convolutional networks, TCNs process a local neighborhood across time which leads to inefficiency in modeling the long-range dependencies between these temporal patterns of the video. In this paper, we propose Self-Attention - Temporal Convolutional Network (SA-TCN), which is able to capture both complex activity patterns and their dependencies within long-term untrimmed videos. We evaluate our proposed model on DAily Home LIfe Activity Dataset (DAHLIA) and Breakfast datasets. Our proposed method achieves state-of-the-art performance on both DAHLIA and Breakfast dataset. Rui Dai 0001, Luca Minciullo, Lorenzo Garattoni, Gianpiero Francesca, François Brémond |
AVSS | 5 |
| 2019 | Spatial Attention for Pedestrian DetectionabstractAchieving high detection accuracy and high inference speed is important for a pedestrian detection system in self-driving applications. There exists a trade-off between detection accuracy and inference speed in modern convolutional object detectors. In this paper, we propose a novel pedestrian detection system, which leverages spatial attention and a two-level cascade of classification and bounding box regression to balance the trade-off. Our proposed spatial attention module reduces the search space for pedestrians by selecting a small set of anchor boxes for further processing. Furthermore, we present a two-level cascade of bounding box classification and regression and demonstrate its effectiveness for improved accuracy. We demonstrate the performance of our system on 2 public datasets-caltech-reasonable and citypersons; with state-of-art performance. Our ablation studies confirm the usefulness of our spatial attention and cascade modules. Ujjwal, Aziz Dziri, Bertrand Leroy, François Brémond |
AVSS | 4 |
| 2019 | Characterizing the State of Apathy with Facial Expression and Motion AnalysisabstractReduced emotional response, lack of motivation, and limited social interaction comprise the major symptoms of apathy. Current methods for apathy diagnosis require the patient's presence in a clinic, and time consuming clinical interviews and questionnaires involving medical personnel, which are costly and logistically inconvenient for patients and clinical staff, hindering among other large scale diagnostics. In this paper we introduce a novel machine learning framework to classify apathetic and non-apathetic patients based on analysis of facial dynamics, entailing both emotion and facial movement. Our approach caters to the challenging setting of current apathy assessment interviews, which include short video clips with wide face pose variations, very low-intensity expressions, and insignificant inter-class variations. We test our algorithm on a dataset consisting of 90 video sequences acquired from 45 subjects and obtained an accuracy of 84% in apathy classification. Based on extensive experiments, we show that the fusion of emotion and facial local motion produces the best feature set for apathy classification. In addition, we train regression models to predict the clinical scores related to the mental state examination (MMSE) and the neuropsychiatric apathy inventory (NPI) using the motion and emotion features. Our results suggest that the performance can be further improved by appending the predicted clinical scores to the video-based feature representation. S. L. Happy, Antitza Dantcheva, Abhijit Das 0001, Radia Zeghari, Philippe Robert, François Brémond |
FG | 6 |
| 2019 | Toyota Smarthome: Real-World Activities of Daily LivingabstractThe performance of deep neural networks is strongly influenced by the quantity and quality of annotated data. Most of the large activity recognition datasets consist of data sourced from the web, which does not reflect challenges that exist in activities of daily living. In this paper, we introduce a large real-world video dataset for activities of daily living: Toyota Smarthome. The dataset consists of 16K RGB+D clips of 31 activity classes, performed by seniors in a smarthome. Unlike previous datasets, videos were fully unscripted. As a result, the dataset poses several challenges: high intra-class variation, high class imbalance, simple and composite activities, and activities with similar motion and variable duration. Activities were annotated with both coarse and fine-grained labels. These characteristics differentiate Toyota Smarthome from other datasets for activity recognition. As recent activity recognition approaches fail to address the challenges posed by Toyota Smarthome, we present a novel activity recognition method with attention mechanism. We propose a pose driven spatio-temporal attention mechanism through 3D ConvNets. We show that our novel method outperforms state-of-the-art methods on benchmark datasets, as well as on the Toyota Smarthome dataset. We release the dataset for research use. Srijan Das, Rui Dai 0001, Michal Koperski, Luca Minciullo, Lorenzo Garattoni, François Brémond, Gianpiero Francesca |
ICCV | 6 |
| 2019 | A New Hybrid Architecture for Human Activity Recognition from RGB-D Videos
Srijan Das, Monique Thonnat, Kaustubh Sakhalkar, Michal Koperski, François Brémond, Gianpiero Francesca |
MMM (2) | 5 |
| 2019 | Where to Focus on for Human Action Recognition?abstractIn this paper, we present a new attention model for the recognition of human action from RGB-D videos. We propose an attention mechanism based on 3D articulated pose. The objective is to focus on the most relevant body parts involved in the action. For action classification, we propose a classification network compounded of spatio-temporal subnetworks modeling the appearance of human body parts and RNN attention subnetwork implementing our attention mechanism. Furthermore, we train our proposed network end-to-end using a regularized cross-entropy loss, leading to a joint training of the RNN delivering attention globally to the whole set of spatio-temporal features, extracted from 3D ConvNets. Our method outperforms the State-of-the-art methods on the largest human activity recognition dataset available to-date (NTU RGB+D Dataset) which is also multi-views and on a human action recognition dataset with object interaction (Northwestern-UCLA Multiview Action 3D Dataset). Srijan Das, Arpit Chaudhary, François Brémond, Monique Thonnat |
WACV | 3 |
| 2019 | Cross Domain Residual Transfer Learning for Person Re-IdentificationabstractThis paper presents a novel way to transfer model weights from one domain to another using residual learning framework instead of direct fine-tuning. It also argues for hybrid models that use learned (deep) features and statistical metric learning for multi-shot person re-identification when training sets are small. This is in contrast to popular end-to-end neural network based models or models that use hand-crafted features with adaptive matching models (neural nets or statistical metrics). Our experiments demonstrate that a hybrid model with residual transfer learning can yield significantly better re-identification performance than an end-to-end model when training set is small. On iLIDS-VID and PRID datasets, we achieve rank-1 recognition rates of 89.8% and 95%, respectively, which is a significant improvement over state-of-the-art. Furqan Khan, François Brémond |
WACV | 2 |
| 2019 | Magnitude-Orientation Stream network and depth information applied to activity recognition
Carlos Antônio Caetano Jr., Victor C. de Melo, François Brémond, Jefersson A. dos Santos, William Robson Schwartz |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | A Weakly Supervised learning technique for classifying facial expressions
S. L. Happy, Antitza Dantcheva, François Brémond |
Pattern Recognit. Lett. | 3 |
| 2018 | Deep-Temporal LSTM for Daily Living Action RecognitionabstractIn this paper, we propose to improve the traditional use of RNNs by employing a many to many model for video classification. We analyze the importance of modeling spatial layout and temporal encoding for daily living action recognition. Many RGB methods focus only on short term temporal information obtained from optical flow. Skeleton based methods on the other hand show that modeling long term skeleton evolution improves action recognition accuracy. In this work, we propose a deep-temporal LSTM architecture which extends standard LSTM and allows better encoding of temporal information. In addition, we propose to fuse 3D skeleton geometry with deep static appearance. We validate our approach on public available CAD60, MSRDailyActivity3D and NTU-RGB+D, achieving competitive performance as compared to the state-of-the art. Srijan Das, Michal Koperski, François Brémond, Gianpiero Francesca |
AVSS | 3 |
| 2018 | Online Detection of Long-Term Daily Living Activities by Weakly Supervised Recognition of Sub-ActivitiesabstractIn this paper, we address detection of activities in long-term untrimmed videos. Detecting temporal delineation of activities is important to analyze large-scale videos. However, there are still challenges yet to be overcome in order to have an accurate temporal segmentation of activities. Detection of daily-living activities is even more challenging due to their high intra-class and low inter-class variations, complex temporal relationships of sub-activities performed in realistic settings. To tackle these problems, we propose an online activity detection framework based on the discovery of sub-activities. We consider a long-term activity as a sequence of short-term sub-activities. Then we utilize a weakly supervised classifier trained on discovered sub-activities which allows us to predict an ongoing activity before being completely observed. To achieve a more precise segmentation a greedy post-processing technique based on Markov models is employed. We evaluate our framework on DAHLIA and GAADRD daily living activity datasets where we achieve state-of-the-art results on detection of activities. Farhood Negin, Abhishek Goel, Abdelrahman G. Abubakr, François Brémond, Gianpiero Francesca |
AVSS | 4 |
| 2018 | Cascade-Dispatched Classifier Ensemble and Regressor for Pedestrian DetectionabstractThis paper focuses on ensemble classifiers for pedestrian detection. Ensemble learning is widely used in this field for context disambiguation or via a cascade-of-rejectors. However, applying the typical, parallel, instance of it remains disappointing in most cases. Our work studies the mechanisms that hinder the efficiency of ensemble classifiers for pedestrian detection, and, based on our findings, we introduce a structured classifier ensemble that improves performance without loss of speed. We also harness this principle for context disambiguation via the application of a regressor to pedestrian detection. Experiments on the INRIA and Caltech-USA datasets validate the approach. Rémi Trichet, François Brémond |
AVSS | 2 |
| 2018 | Late Fusion of Multiple Convolutional Layers for Pedestrian DetectionabstractWe propose a system design for pedestrian detection by leveraging the power of multiple convolutional layers explicitly. We quantify the effect of different convolutional layers on the detection of pedestrians of varying scales and occlusion level. We show that earlier convolutional layers are better at handling small-scale and partially occluded pedestrians. We take cue from these conclusions and propose a pedestrian detection system design based on Faster-RCNN which leverages multiple convolutional layers by late fusion. In our design, we introduce height-awareness in the loss function to make the network emphasize on pedestrian heights which are misclassified during the training process. The proposed system design achieves a log-average miss-rate of 9.25% on the caltech-reasonable dataset. This is within 1.5% of the current state-of-art approach, while being a more compact system. Ujjwal, Aziz Dziri, Bertrand Leroy, François Brémond |
AVSS | 4 |
| 2018 | Residual Transfer Learning for Multiple Object TrackingabstractTo address the Multiple Object Tracking (MOT) challenge, we propose to enhance the tracklet appearance features, given by a Convolutional Neural Network (CNN), based on the Residual Transfer Learning (RTL) method. Considering that object classification and tracking are significantly different tasks at high level. And that traditional fine-tuning limits the possible variations in all the layers of the network since it changes the last convolutional layers. Beyond that, our proposed method provides more flexibility in terms of modelling the difference between these two tasks with a four-stage training. This transfer approach increases the feature performance compared to traditional CNN fine-tuning. Experiments on the MOT17 challenge show competitive results with the current state-of-the-art methods. Juan Diego Gonzales Zuniga, Thi Lan Anh Nguyen, François Brémond |
AVSS | 3 |
| 2018 | Show me your face and I will tell you your height, weight and body mass indexabstractBody height, weight, as well as the associated and composite body mass index (BMI) are human attributes of pertinence due to their use in a number of applications including surveillance, re-identification, image retrieval systems, as well as healthcare. Previous work on automated estimation of height, weight and BMI has predominantly focused on 2D and 3D full-body images and videos. Little attention has been given to the use of face for estimating such traits. Motivated by the above, we here explore the possibility of estimating height, weight and BMI from single-shot facial images by proposing a regression method based on the 50-layers ResNet-architecture. In addition, we present a novel dataset consisting of 1026 subjects and show results, which suggest that facial images contain discriminatory information pertaining to height, weight and BMI, comparable to that of body-images and videos. Finally, we perform a gender-based analysis of the prediction of height, weight and BMI. Antitza Dantcheva, François Brémond, Piotr Bilinski |
ICPR | 2 |
| 2018 | Online temporal detection of daily-living human activities in long untrimmed video streamsabstractMany approaches were proposed to solve the problem of activity recognition in short clipped videos, which achieved impressive results with hand-crafted and deep features. However, it is not practical to have clipped videos in real life, where cameras provide continuous video streams in applications such as robotics, video surveillance, and smart-homes. Here comes the importance of activity detection to help recognizing and localizing each activity happening in long videos. Activity detection can be defined as the ability to localize starting and ending of each human activity happening in the video, in addition to recognizing each activity label. A more challenging category of human activities is the daily-living activities, such as eating, reading, cooking, etc, which have low inter-class variation and environment where actions are performed are similar. In this work we focus on solving the problem of detection of daily-living activities in untrimmed video streams. We introduce new online activity detection pipeline that utilizes single sliding window approach in a novel way; the classifier is trained with sub-parts of training activities, and an online frame-level early detection is done for sub-parts of long activities during detection. Finally, a greedy Markov model based post processing algorithm is applied to remove false detection and achieve better results. We test our approaches on two daily-living datasets, DAHLIA and GAADRD, outperforming state of the art results by more than 10%. Abhishek Goel, Abdelrahman G. Abubakr, Michal Koperski, François Brémond, Gianpiero Francesca |
IPAS | 4 |
| 2018 | Learning to Represent Spatio-Temporal Features for Fine Grained Action RecognitionabstractConvolutional neural networks have pushed the boundaries of action recognition in videos, especially with the introduction of 3D convolutions. But it is an open ended question on how efficiently a 3D CNN can model temporal information? which we try to investigate and introduce a new optical flow representation to improve the motion stream. We use the baseline inflated 3D CNN networks and separate the convolutional filters into spatial and temporal, which reduces the number of parameters with minimal loss of accuracy. We evaluate our approach on NTU RGBD dataset which is the largest human action dataset and outperform the state-of-the-art by a large margin. Kaustubh Sakhalkar, François Brémond |
IPAS | 2 |
| 2018 | Recognition of Daily Activities by embedding hand-crafted features within a semantic analysisabstractThe recognition of complex actions is still a challenging task in Computer Vision especially in daily living scenarios, where problems like occlusion and limited field of view are very common. Recognition of Activity Daily Living (ADL) could improve the quality of life and supporting independent and healthy living of older or/and impaired people by using information and communication technologies at home, at the workplace and in public spaces. This paper proposes to embed spatio-temporal information into ontology models to improve action recognition using visual words. Actions detected by visual words are implemented as Primitive States in the scenario and then used as Components of Composite States to merge them with spatio-temporal patterns that the people display while performing ADLs. In a challenging dataset, such as SmartHome, where a high variance intra-class and low variance inter-class is present, recognition results for some actions improve in precision and recall thanks to spatial information. Francesco Verrini, Carlos Fernando Crispim, Manuela Chessa, Fabio Solari, François Brémond |
IPAS | 5 |
| 2018 | LBP Channels for Pedestrian DetectionabstractThis paper introduces a new channel descriptor for pedestrian detection. This type of descriptor usually selects a set of one-valued filters within the enormous set of all possible filters for improved efficiency. The main claim underpinning this paper is that the recent works on channel-based features restrict the filter space search, therefore bringing along the obsolescence of one-valued filter representation. To prove our claim, we introduce a 12-valued filter representation based on local binary patterns. Indeed, various improvements now allow for this texture feature to provide a very discriminative, yet compact descriptor. Filter selection boasting new combination restrictions as well as a reverse selection process are also presented to choose the best filters. experiments on the INRIA and Caltech-USA datasets validate the approach. Rémi Trichet, François Brémond |
WACV | 2 |
| 2018 | PRAXIS: Towards automatic cognitive assessment using gesture recognition
Farhood Negin, Pau Rodríguez, Michal Koperski, Adlen Kerboua, Jordi Gonzàlez 0001, Jeremy Bourgeois, Emmanuelle Chapoulie, Philippe Robert, François Brémond |
Expert Syst. Appl. | 9 |
| 2017 | Multi-Object tracking using multi-channel part appearance representationabstractAppearance based multi-object tracking (MOT) is a challenging task, specially in complex scenes where objects have similar appearance or are occluded by background or other objects. Such factors motivate researchers to propose effective trackers which should satisfy real-time processing and object trajectory recovery criteria. In order to handle both mentioned requirements, we propose a robust online multi-object tracking method that extends the features and methods proposed for re-identification to MOT. The proposed tracker combines a local and a global tracker in a comprehensive two-step framework. In the local tracking step, we use the frame-to-frame association to generate online object trajectories. Each object trajectory is called tracklet and is represented by a set of multi-modal feature distributions modeled by GMMs. In the global tracking step, occlusions and mis-detections are recovered by tracklet bipartite association method based on learning Mahalanobis metric between GMM components using KISSME metric learning algorithm. Experiments on two public datasets show that our tracker performs well when compared to state-of-the-art tracking algorithms. Thi Lan Anh Nguyen, Furqan Muhammad Khan, Farhood Negin, François Brémond |
AVSS | 4 |
| 2017 | Action recognition based on a mixture of RGB and depth based skeletonabstractIn this paper, we study how different skeleton extraction methods affect the performance of action recognition. As shown in previous work skeleton information can be exploited for action recognition. Nevertheless, skeleton detection problem is already hard and very often it is difficult to obtain reliable skeleton information from videos. In this paper, we compare two skeleton detection methods: the depth-map based method used with Kinect camera and RGB based method that uses Deep Convolutional Neural Networks. In order to balance the pros and cons of mentioned skeleton detection methods w.r.t. action recognition task, we propose a fusion of classifiers trained based on each skeleton detection method. Such fusion lead to performance improvement. We validate our approach on CAD-60 and MSRDailyActivity3D, achieving state-of-the-art results. Srijan Das, Michal Koperski, François Brémond, Gianpiero Francesca |
AVSS | 3 |
| 2017 | Efficient Video Summarization Using Principal Person Appearance for Video-Based Person Re-Identification
Seongro Yoon, Furqan Khan, François Brémond |
BMVC | 3 |
| 2017 | Multi-shot Person Re-Identification Using Part Appearance MixtureabstractAppearance based person re-identification in real-world video surveillance systems is a challenging problem for many reasons, including ineptness of existing low level features under significant viewpoint, illumination, or camera characteristic changes to robustly describe a person's appearance. One approach to handle appearance variability is to learn similarity metrics or ranking functions to implicitly model appearance transformation between cameras for each camera pair, or group, in the system. The alternative, that this paper follows, is the more fundamental approach of improving appearance descriptors, called signatures, to cater for high appearance variance and occlusions. The novel signature representation for multi-shot person reidentification presented in this paper uses multiple appearance models, each describing appearance as a probability distribution of a low-level feature for a certain portion of individual's body. Combined with metric learning, rank-1 recognition rates of 92:5% and 79:5% are achieved on PRID2011 [12] and iLIDS-VID [34] datasets, respectively. Furqan Muhammad Khan, François Brémond |
WACV | 2 |
| 2017 | Toward Abnormal Trajectory and Event Detection in Video SurveillanceabstractIn this paper, we present a unified approach for abnormal behavior detection and group behavior analysis in video scenes. Existing approaches for abnormal behavior detection do either use trajectory-based or pixel-based methods. Unlike these approaches, we propose an integrated pipeline that incorporates the output of object trajectory analysis and pixel-based analysis for abnormal behavior inference. This enables to detect abnormal behaviors related to speed and direction of object trajectories, as well as complex behaviors related to finer motion of each object. By applying our approach on three different data sets, we show that our approach is able to detect several types of abnormal group behaviors with less number of false alarms compared with existing approaches. Serhan Cosar, Giuseppe Donatiello, Vania Bogorny, Carolina Gárate, Luis Otávio Alvares, François Brémond |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2017 | Globality-Locality-Based Consistent Discriminant Feature Ensemble for Multicamera TrackingabstractSpatiotemporal data association and fusion is a well-known NP-hard problem even in a small number of cameras and frames. Although it is difficult to be tractable, solving them is pivotal for tracking in a multicamera network. Most approaches model association maladaptively toward properties and contents of video, and hence they produce suboptimal associations and association errors propagate over time to adversely affect fusion. In this paper, we present an online multicamera multitarget tracking framework that performs adaptive tracklet correspondence by analyzing and understanding contents and properties of video. Unlike other methods that work only on synchronous videos, our approach uses dynamic time warping to establish correspondence even if videos have linear or nonlinear time asynchronous relationship. Association is a two-stage process based on geometric and appearance descriptor space ranked by their inter- and intra-camera consistency and discriminancy. Fusion is reinforced by weighting the associated tracklets with a confidence score calculated using reliability of individual camera tracklets. Our robust ranking and election learning algorithm dynamically selects appropriate features for any given video. Our method establishes that, given the right ensemble of features, even computationally efficient optimization yields better accuracy in tracking over time and provides faster convergence that is suitable for real-time application. For evaluation on RGB, we benchmark on multiple sequences in PETS 2009 and we achieve performance that is on par with the state of the art. For evaluating on RGB-D, we built a new data set. Kanishka Nithin, François Brémond |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Gender Estimation Based on Smile-DynamicsabstractAutomated gender estimation has numerous applications, including video surveillance, human-computer interaction, anonymous customized advertisement, and image retrieval. Most commonly, the underlying algorithms analyze the facial appearance for clues of gender. In this paper, we propose a novel method for gender estimation, which exploits dynamic features gleaned from smiles and we proceed to show that: a) facial dynamics incorporate clues for gender dimorphism and b) while for adult individuals appearance features are more accurate than dynamic features, for subjects under 18 years facial dynamics can outperform appearance features. In addition, we fuse proposed dynamics-based approach with state-of-the-art appearance-based algorithms, predominantly improving performance of the latter. Results show that smile-dynamics include pertinent and complementary to appearance gender information. Antitza Dantcheva, François Brémond |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2017 | Exploiting Feature Correlations by Brownian Statistics for People Detection and RecognitionabstractCharacterizing an image region by its feature intercorrelations is a modern trend in computer vision. In this paper, we introduce a new image descriptor that can be seen as a natural extension of a standard covariance descriptor with the advantage of capturing nonlinear and nonmonotone dependencies. Inspired from the recent advances in mathematical statistics of Brownian motion, we can express highly complex structural information in a compact and computationally efficient manner. We show that our Brownian covariance descriptor can capture richer image characteristics than the covariance descriptor. Additionally, a detailed analysis of the Brownian manifold reveals that opposite to the classical covariance descriptor, the proposed descriptor lies in a relatively flat manifold, which can be treated as a Euclidean. This brings significant boost in the efficiency of the descriptor. The effectiveness and the generality of our approach is validated on two challenging vision tasks, pedestrian classification, and person reidentification. The experiments are carried out on multiple datasets achieving promising results. Slawomir Bak, Marco San-Biagio, Ratnesh Kumar 0003, Vittorio Murino, François Brémond |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2016 | Human violence recognition and detection in surveillance videosabstractIn this paper, we focus on the important topic of violence recognition and detection in surveillance videos. Our goal is to determine if a violence occurs in a video (recognition) and when it happens (detection). Firstly, we propose an extension of the Improved Fisher Vectors (IFV) for videos, which allows to represent a video using both local features and their spatio-temporal positions. Then, we study the popular sliding window approach for violence detection, and we re-formulate the Improved Fisher Vectors and use the summed area table data structure to speed up the approach. We present an extensive evaluation, comparison and analysis of the proposed improvements on 4 state-of-the-art datasets. We show that the proposed improvements make the violence recognition more accurate (as compared to the standard IFV, IFV with spatio-temporal grid, and other state-of-the-art methods) and make the violence detection significantly faster. Piotr Bilinski, François Brémond |
AVSS | 2 |
| 2016 | Exploring depth information for head detection with depth imagesabstractHead detection may be more demanding than face recognition and pedestrian detection in the scenarios where a face turns away or body parts are occluded in the view of a sensor, but locating people is needed. In this paper, we introduce an efficient head detection approach for single depth images at low computational expense. First, a novel head descriptor is developed and used to classify pixels as head or non-head. We use depth values to guide each window size, to eliminate false positives of head centers, and to cluster head pixels, which significantly reduce the computation costs of searching for appropriate parameters. High head detection performance was achieved in experiments - 90% accuracy for our dataset containing heads with different body postures, head poses, and distances to a Kinect2 sensor, and above 70% precision on a public dataset composed of a few daily activities, which is higher than using a head-shoulder detector with HOG feature for depth images. François Brémond, Hugues Thomas |
AVSS | 2 |
| 2016 | Semi-supervised understanding of complex activities from temporal conceptsabstractMethods for action recognition have evolved considerably over the past years and can now automatically learn and recognize short term actions with satisfactory accuracy. Nonetheless, the recognition of complex activities - compositions of actions and scene objects - is still an open problem due to the complex temporal and composite structure of this category of events. Existing methods focus either on simple activities or oversimplify the modeling of complex activities by targeting only whole-part relations between its sub-parts (e.g., actions). In this paper, we propose a semi-supervised approach that learns complex activities from the temporal patterns of concept compositions (e.g., “slicing-tomato” before “pouring into-pan”). We demonstrate that our method outperforms prior work in the task of automatic modeling and recognition of complex activities learned out of the interaction of 218 distinct concepts. Carlos Fernando Crispim, Michal Koperski, Serhan Cosar, François Brémond |
AVSS | 4 |
| 2016 | Unsupervised data association for metric learning in the context of multi-shot person re-identificationabstractAppearance based person re-identification is a challenging task, specially due to difficulty in capturing high intra-person appearance variance across cameras when inter-person similarity is also high. Metric learning is often used to address deficiency of low-level features by learning view specific re-identification models. The models are often acquired using a supervised algorithm. This is not practical for real-world surveillance systems because annotation effort is view dependent. In this paper, we propose a strategy to automatically generate labels for person tracks to learn similarity metric for multi-shot person re-identification task. We demonstrate on multiple challenging datasets that the proposed labeling strategy significantly improves performance of two baseline methods and the extent of improvement is comparable to that of manual annotations in the context of KISSME algorithm [14]. Furqan Muhammad Khan, François Brémond |
AVSS | 2 |
| 2016 | Modeling spatial layout of features for real world scenario RGB-D action recognitionabstractDepth information improves skeleton detection, thus skeleton based methods are the most popular methods in RGB-D action recognition. But skeleton detection working range is limited in terms of distance and view-point. Most of the skeleton based action recognition methods ignore fact that skeleton may be missing. Local points-of-interest (POIs) do not require skeleton detection. But they fail if they cannot detect enough POIs e.g. amount of motion in action is low. Most of them ignore spatial-location of features. We cope with the above problems by employing people detector instead of skeleton detector. We propose method to encode spatial-layout of features inside bounding box. We also introduce descriptor which encodes static information for actions with low amount of motion. We validate our approach on: 3 public data-sets. The results show that our method is competitive to skeleton based methods, while requiring much simpler people detection instead of skeleton detection. Michal Koperski, François Brémond |
AVSS | 2 |
| 2016 | A hybrid framework for online recognition of activities of daily living in real-world settingsabstractMany supervised approaches report state-of-the-art results for recognizing short-term actions in manually clipped videos by utilizing fine body motion information. The main downside of these approaches is that they are not applicable in real world settings. The challenge is different when it comes to unstructured scenes and long-term videos. Unsupervised approaches have been used to model the long-term activities but the main pitfall is their limitation to handle subtle differences between similar activities since they mostly use global motion information. In this paper, we present a hybrid approach for long-term human activity recognition with more precise recognition of activities compared to unsupervised approaches. It enables processing of long-term videos by automatically clipping and performing online recognition. The performance of our approach has been tested on two Activities of Daily Living (ADL) datasets. Experimental results are promising compared to existing approaches. Farhood Negin, Michal Koperski, Carlos Fernando Crispim, François Brémond, Serhan Cosar, Konstantinos Avgerinakis |
AVSS | 4 |
| 2016 | Multi-object tracking of pedestrian driven by contextabstractThe characteristics like density of objects, their contrast with respect to surrounding background, their occlusion level and many more describe the context of the scene. The variation of the context represents ambiguous task to be solved by tracker. In this paper we present a new long term tracking framework boosted by context around each tracklet. The framework works by first learning the database of optimal tracker parameters for various context offline. During the testing, the context surrounding each tracklet is extracted and match against database to select best tracker parameters. The tracker parameters are tuned for each tracklet in the scene to highlight its discrimination with respect to surrounding context rather than tuning the parameters for whole scene. The proposed framework is trained on 9 public video sequences and tested on 3 unseen sets. It outperforms the state-of-art pedestrian trackers in scenarios of motion changes, appearance changes and occlusion of objects. Thi Lan Anh Nguyen, François Brémond, Jana Trojanová |
AVSS | 2 |
| 2016 | Image-based gender estimation from body and face across distancesabstractGender estimation has received increased attention due to its use in a number of pertinent security and commercial applications. Automated gender estimation algorithms are mainly based on extracting representative features from face images. In this work we study gender estimation based on information deduced jointly from face and body, extracted from single-shot images. The approach addresses challenging settings such as low-resolution-images, as well as settings when faces are occluded. Specifically the face-based features include local binary patterns (LBP) and scale-invariant feature transform (SIFT) features, projected into a PCA space. The features of the novel body-based algorithm proposed in this work include continuous shape information extracted from body silhouettes and texture information retained by HOG descriptors. Support Vector Machines (SVMs) are used for classification for body and face features. We conduct experiments on images extracted from video-sequences of the Multi-Biometric Tunnel database, emphasizing on three distance-settings: close, medium and far, ranging from full body exposure (far setting) to head and shoulders exposure (close setting). The experiments suggest that while face-based gender estimation performs best in the close-distance-setting, body-based gender estimation performs best when a large part of the body is visible. Finally we present two score-level-fusion schemes of face and body-based features, outperforming the two individual modalities in most cases. Ester Gonzalez-Sosa, Antitza Dantcheva, Rubén Vera-Rodríguez, Jean-Luc Dugelay, François Brémond, Julian Fierrez |
ICPR | 5 |
| 2016 | Semantic Event Fusion of Different Visual Modality Concepts for Activity RecognitionabstractCombining multimodal concept streams from heterogeneous sensors is a problem superficially explored for activity recognition. Most studies explore simple sensors in nearly perfect conditions, where temporal synchronization is guaranteed. Sophisticated fusion schemes adopt problem-specific graphical representations of events that are generally deeply linked with their training data and focused on a single sensor. This paper proposes a hybrid framework between knowledge-driven and probabilistic-driven methods for event representation and recognition. It separates semantic modeling from raw sensor data by using an intermediate semantic representation, namely concepts. It introduces an algorithm for sensor alignment that uses concept similarity as a surrogate for the inaccurate temporal information of real life scenarios. Finally, it proposes the combined use of an ontology language, to overcome the rigidity of previous approaches at model definition, and a probabilistic interpretation for ontological models, which equips the framework with a mechanism to handle noisy and ambiguous concept observations, an ability that most knowledge-driven methods lack. We evaluate our contributions in multimodal recordings of elderly people carrying out IADLs. Results demonstrated that the proposed framework outperforms baseline methods both in event recognition performance and in delimiting the temporal boundaries of event instances. Carlos Fernando Crispim, Vincent Buso, Konstantinos Avgerinakis, Georgios Meditskos, Alexia Briassouli, Jenny Benois-Pineau, Ioannis Kompatsiaris, François Brémond |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2015 | Robust global tracker based on an online estimation of tracklet descriptor reliabilityabstractThe complex scene conditions such as light change, high density of mobile objects or object occlusion can cause object mis-detections. When a tracker can not recover these mis-detections, the trajectory of an object is fragmented into some short trajectories called tracklets. As a result, tracking quality is reduced remarkably. In this paper, we propose a new approach to improve the tracking quality by a global tracker which merges all tracklets belonging to an object in the whole video. Particularly, we compute descriptor reliability over time based on their discrimination. On the other hand, a motion model is also combined with appearance descriptors in a flexible way to improve the tracking quality. The proposed approach is evaluated on four benchmark datasets. The obtained results show the robustness and effectiveness of our approach compared to tracking as well as tracklet linking approaches from state of the art. Thi Lan Anh Nguyen, Duc Phu Chau, François Brémond |
AVSS | 3 |
| 2015 | Minimizing hallucination in histogram of Oriented GradientsabstractHistogram of Oriented Gradients is one of the most extensively used image descriptors in computer vision. It has successfully been applied to various vision tasks such as localization, classification and recognition. As it mainly captures gradient strengths in an image, it is sensitive to local variations in illumination and contrast. In the result, a normalization of this descriptor turns out to be essential for good performance [3, 4]. Although different normalization schemes have been investigated, all of them usually employ L1 or L2-norm. In this paper we show that an incautious application of L-like norms to the HOG descriptor might produce a hallucination effect. To overcome this issue, we propose a new normalization scheme that effectively minimizes hallucinations. This scheme is built upon a detailed analysis of the gradient distribution resulting in adding an extra bin with a specific value that increases HOG distinctiveness. We validated our approach on person re-identification and action recognition, demonstrating significant boost in the performance. Javier Ortiz 0005, Slawomir Bak, Michal Koperski, François Brémond |
AVSS | 4 |
| 2015 | Adaptive Neuro-Fuzzy Controller for Multi-object Tracker
Duc Phu Chau, K. Subramanian 0001, François Brémond |
ICVS | 3 |
| 2015 | Video Covariance Matrix Logarithm for Human Action Recognition in Videos
Piotr Bilinski, François Brémond |
IJCAI | 2 |
| 2015 | Unsupervised discovery of human activities from long-time videosabstractIn this study, the authors propose a complete framework based on a hierarchical activity model to understand and recognise activities of daily living in unstructured scenes. At each particular time of a long‐time video, the framework extracts a set of space‐time trajectory features describing the global position of an observed person and the motion of his/her body parts. Human motion information is gathered in a new feature that the authors call perceptual feature chunks (PFCs). The set of PFCs is used to learn, in an unsupervised way, particular regions of the scene (topology) where the important activities occur. Using topologies and PFCs, the video is broken into a set of small events (‘primitive events’) that have a semantic meaning. The sequences of ‘primitive events’ and topologies are used to construct hierarchical models for activities. The proposed approach has been tested with the medical field application to monitor patients suffering from Alzheimer's and dementia. The authors have compared their approach to their previous study and a rule‐based approach. Experimental results show that the framework achieves better performance than existing works and has the potential to be used as a monitoring tool in medical field applications. Salma Elloumi, Serhan Cosar, Guido Pusiol, François Brémond, Monique Thonnat |
IET Comput. Vis. | 4 |
| 2014 | Global tracker: An online evaluation framework to improve tracking qualityabstractEvaluating the quality of tracking outputs is an important task in video analysis. This paper presents a new framework for estimating both detection and tracking quality during runtime. If anomalies are detected in the tracking output results, they are categorized as natural phenomena or real errors using contextual information. As this framework should be generic and work on any kind of system (single camera, camera network), a reacquisition step using a constrained clustering algorithm is also performed in order to keep track of the object even if it leaves the scene and comes back or appears on another camera. The framework is evaluated on two datasets using different kinds of tracking algorithms. Julien Badie, François Brémond |
AVSS | 2 |
| 2014 | Improving person re-identification by viewpoint cuesabstractRe-identifying people in a network of cameras requires an invariant human representation. State of the art algorithms are likely to fail in real-world scenarios due to serious perspective changes. Most of existing approaches focus on invariant and discriminative features, while ignoring the body alignment issue. In this paper we propose 3 methods for improving the performance of person re-identification. We focus on eliminating perspective distortions by using 3D scene information. Perspective changes are minimized by affine transformations of cropped images containing the target (1). Further we estimate the human pose for (2) clustering data from a video stream and (3) weighting image features. The pose is estimated using 3D scene information and motion of the target. We validated our approach on a publicly available dataset with a network of 8 cameras. The results demonstrated significant increase in the re-identification performance over the state of the art. Slawomir Bak, Sofia Zaidenberg, Bernard Boulay, François Brémond |
AVSS | 4 |
| 2014 | Representing visual appearance by video Brownian covariance descriptor for human action recognitionabstractThis paper addresses a problem of recognizing human actions in video sequences. Recent studies have shown that methods which use bag-of-features and space-time features achieve high recognition accuracy. Such methods extract both appearance-based and motion-based features. This paper focuses only on appearance features. We propose to model relationships between different pixel-level appearance features such as intensity and gradient using Brownian covariance, which is a natural extension of classical covariance measure. While classical covariance can model only linear relationships, Brownian covariance models all kinds of possible relationships. We propose a method to compute Brownian covariance on space-time volume of a video sequence. We show that proposed Video Brownian Covariance (VBC) descriptor carries complementary information to the Histogram of Oriented Gradients (HOG) descriptor. The fusion of these two descriptors gives a significant improvement in performance on three challenging action recognition datasets. Piotr Bilinski, Michal Koperski, Slawomir Bak, François Brémond |
AVSS | 4 |
| 2014 | Background subtraction in people detection framework for RGB-D camerasabstractIn this paper, we propose a background subtraction algorithm specific for depth videos from RGB-D cameras. Embedded in a people detection framework, it does not classify foreground / background at pixel level but provides useful information for the framework to remove noise. Noise is only removed when the framework has all the information from background subtraction, classification and object tracking. In our experiment, our background subtraction algorithm outperforms GMM, a popular background subtraction algorithm, in detecting people and removing noise. Anh-Tuan Nghiem, François Brémond |
AVSS | 2 |
| 2014 | 3D trajectories for action recognitionabstractRecent development in affordable depth sensors opens new possibilities in action recognition problem. Depth information improves skeleton detection, therefore many authors focused on analyzing pose for action recognition. But still skeleton detection is not robust and fail in more challenging scenarios, where sensor is placed outside of optimal working range and serious occlusions occur. In this paper we investigate state-of-the-art methods designed for RGB videos, which have proved their performance. Then we extend current state-of-the-art algorithms to benefit from depth information without need of skeleton detection. In this paper we propose two novel video descriptors. First combines motion and 3D information. Second improves performance on actions with low movement rate. We validate our approach on challenging MSR Daily Activty 3D dataset. Michal Koperski, Piotr Bilinski, François Brémond |
ICIP | 3 |
| 2014 | Gait Recognition Based on Modified Phase Only Correlation
Imad Rida, Ahmed Bouridane, Samer Al Kork, François Brémond |
ICISP | 4 |
| 2014 | Brownian descriptor: A rich meta-feature for appearance matchingabstractThis paper introduces an image region descriptor and applies it to the problem of appearance matching. The proposed descriptor can be seen as a natural extension of covariance. Driven by recent studies in mathematical statistics related to Brownian motion, we design the Brownian descriptor. In contrast to the classical covariance descriptor, which measures the degree of linear relationship between features, our novel descriptor measures the degree of all kinds of possible relationships between features. We argue that the proposed covariance is a richer descriptor than the classical covariance, especially when fusing non-linearly dependent features. We evaluate our approach on tracking related applications, demonstrating that the Brownian descriptor outperforms the classical covariance in terms of matching accuracy and efficiency. Slawomir Bak, Ratnesh Kumar 0003, François Brémond |
WACV | 3 |
| 2014 | Automatic tracker selection w.r.t object detection performanceabstractThe tracking algorithm performance depends on video content. This paper presents a new multi-object tracking approach which is able to cope with video content variations. First the object detection is improved using Kanade-Lucas-Tomasi (KLT) feature tracking. Second, for each mobile object, an appropriate tracker is selected among a KLT-based tracker and a discriminative appearance-based tracker. This selection is supported by an online tracking evaluation. The approach has been experimented on three public video datasets. The experimental results show a better performance of the proposed approach compared to recent state of the art trackers. Duc Phu Chau, François Brémond, Monique Thonnat |
WACV | 2 |
| 2014 | Online parameter tuning for object tracking algorithms
Duc Phu Chau, Monique Thonnat, François Brémond, Etienne Corvée |
Image Vis. Comput. | 3 |
| 2013 | Online tracking parameter adaptation based on evaluationabstractParameter tuning is a common issue for many tracking algorithms. In order to solve this problem, this paper proposes an online parameter tuning to adapt a tracking algorithm to various scene contexts. In an offline training phase, this approach learns how to tune the tracker parameters to cope with different contexts. In the online control phase, once the tracking quality is evaluated as not good enough, the proposed approach computes the current context and tunes the tracking parameters using the learned values. The experimental results show that the proposed approach improves the performance of the tracking algorithm and outperforms recent state of the art trackers. This paper brings two contributions: (1) an online tracking evaluation, and (2) a method to adapt online tracking parameters to scene contexts. Duc Phu Chau, Julien Badie, François Brémond |
AVSS | 3 |
| 2013 | Evaluation of a monitoring system for event recognition of older peopleabstractPopulation aging has been motivating academic research and industry to develop technologies for the improvement of older people's quality of life, medical diagnosis, and support on frailty cases. Most of available research prototypes for older people monitoring focus on fall detection or gait analysis and rely on wearable, environmental, or video sensors. We present an evaluation of a research prototype of a video monitoring system for event recognition of older people. The prototype accuracy is evaluated for the recognition of physical tasks (e.g., Up and Go test) and instrumental activities of daily living (e.g., watching TV, writing a check) of participants of a clinical protocol for Alzheimer's disease study (29 participants). The prototype uses as input a 2D RGB camera, and its performance is compared to the use of a RGB-D camera. The experimentation results show the proposed approach has a competitive performance to the use of a RGB-D camera, even outperforming it on event recognition precision. The use of a 2D-camera is advantageous, as the camera field of view can be much larger and cover an entire room where at least a couple of RGB-D cameras would be necessary. Carlos Fernando Crispim, Vasanth Bathrinarayanan, Baptiste Fosty, Alexandra König, Rim Romdhane, Monique Thonnat, François Brémond |
AVSS | 7 |
| 2013 | Activity recognition and uncertain knowledge in video scenesabstractActivity recognition has been a growing research topic in the last years and its application varies from automatic recognition of social interaction such as shaking hands, parking lot surveillance, traffic monitoring and the detection of abandoned luggage. This paper describes a probabilistic framework for uncertainty handling in a description-based event recognition approach. The proposed approach allows the flexible modeling of composite events with complex temporal constraints. It uses probability theory to provide a consistent framework for dealing with uncertain knowledge for the recognition of complex events. We validate the event recognition accuracy of the proposed algorithm on real-world videos. The experimental results show that our system can successfully recognize activities with a high recognition rate. We conclude by comparing our algorithm with the state of the art and showing how the definition of event models and the probabilistic reasoning can influence the results of real-time event recognition. Rim Romdhane, Carlos Fernando Crispim, François Brémond, Monique Thonnat |
AVSS | 3 |
| 2013 | Automatic Parameter Adaptation for Multi-object Tracking
Duc Phu Chau, Monique Thonnat, François Brémond |
ICVS | 3 |
| 2013 | Hierarchical and incremental event learning approach based on concept formation models
Marcos Zúñiga, François Brémond, Monique Thonnat |
Neurocomputing | 2 |
| 2012 | Recovering People Tracking Errors Using Enhanced Covariance-Based SignaturesabstractThis paper presents a new approach for tracking multiple persons in a single camera. This approach focuses on recovering tracked individuals that have been lost and are detected again, after being miss-detected (e.g. occluded) or after leaving the scene and coming back. In order to correct tracking errors, a multi-cameras re-identification method is adapted, with a real-time constraint. The proposed approach uses a highly discriminative human signature based on covariance matrix, improved using background subtraction, and a people detection confidence. The problem of linking several tracklets belonging to the same individual is also handled as a ranking problem using a learned parameter. The objective is to create clusters of tracklets describing the same individual. The evaluation is performed on PETS2009 dataset showing promising results. Julien Badie, Slawomir Bak, Silviu-Tudor Serban, François Brémond |
AVSS | 4 |
| 2012 | Contextual Statistics of Space-Time Ordered Features for Human Action RecognitionabstractThe bag-of-words approach with local spatio-temporal features have become a popular video representation for action recognition. Recent methods have typically focused on capturing global and local statistics of features. However, existing approaches ignore relations between the features, particularly space-time arrangement of features, and thus may not be discriminative enough. Therefore, we propose a novel figure-centric representation which captures both local density of features and statistics of space-time ordered features. Using two benchmark datasets for human action recognition, we demonstrate that our representation enhances the discriminative power of features and improves action recognition performance, achieving 96.16% recognition rate on popular KTH action dataset and 93.33% on challenging ADL dataset. Piotr Bilinski, François Brémond |
AVSS | 2 |
| 2012 | Online Learning of Activities from VideoabstractThe present work introduces a new method for activity extraction from video. To achieve this, we focus on the modelling of context by developing an algorithm that automatically learns the main activity zones of the observed scene by taking as input the trajectories of detected mobiles. Automatically learning the context of the scene (activity zones) allows first to extract a knowledge on the occupancy of the different areas of the scene. In a second step, learned zones are employed to extract people activities by relating mobile trajectories to the learned zones, in this way, the activity of a person can be summarised as the series of zones that the person has visited. For the analysis of the trajectory, a multiresolution analysis is set such that a trajectory is segmented into a series of tracklets based on changing speed points thus allowing differentiating when people stop to interact with elements of the scene or other persons. Tracklets allow thus to extract behavioural information. Starting and ending tracklet points are fed to a simple yet advantageous incremental clustering algorithm to create an initial partition of the scene. Similarity relations between resulting clusters are modeled employing fuzzy relations. These can then be aggregated with typical soft-computing algebra. A clustering algorithm based on the transitive closure calculation of the fuzzy relations allows building the final structure of the scene. To allow for incremental learning and update of activity zones (and thus people activities), fuzzy relations are defined with online learning terms. We present results obtained on real videos from different activity domains. Jose Luis Patino, François Brémond, Monique Thonnat |
AVSS | 2 |
| 2012 | Qualitative Evaluation of Detection and Tracking PerformanceabstractA new evaluation approach for detection and tracking systems is presented in this work. Given an algorithm that detects people and simultaneously tracks them, we evaluate its output by considering the complexity of the input scene. Some videos used for the evaluation are recorded using the Kinect sensor which provides for an automated ground truth acquisition system. To analyze the algorithm performance, a number of reasons due to which an algorithm might fail is investigated and quantified over the entire video sequence. A set of features called Scene Complexity measures are obtained for each input frame. The variability in the algorithm performance is modeled by these complexity measures using a polynomial regression model. From the regression statistics, we show that we can compare the performance of two different algorithms and also quantify the relative influence of the scene complexity measures on a given algorithm. Swaminathan Sankaranarayanan, François Brémond, David M. J. Tax |
AVSS | 2 |
| 2012 | A Generic Framework for Video Understanding Applied to Group Behavior RecognitionabstractThis paper presents an approach to detect and track groups of people in video-surveillance applications, and to automatically recognize their behavior. This method keeps track of individuals moving together by maintaining a spacial and temporal group coherence. First, people are individually detected and tracked. Second, their trajectories are analyzed over a temporal window and clustered using the Mean-Shift algorithm. A coherence value describes how well a set of people can be described as a group. Furthermore, we propose a formal event description language. The group events recognition approach is successfully validated on 4 camera views from 3 datasets: an airport, a subway, a shopping center corridor and an entrance hall. Sofia Zaidenberg, Bernard Boulay, François Brémond |
AVSS | 3 |
| 2012 | Learning to Match Appearances by Correlations in a Covariance Metric Space
Slawomir Bak, Guillaume Charpiat, Etienne Corvée, François Brémond, Monique Thonnat |
ECCV (3) | 4 |
| 2012 | Multi-target tracking by discriminative analysis on Riemannian manifoldabstractThis paper addresses the problem of multi-target tracking in crowded scenes from a single camera. We propose an algorithm for learning discriminative appearance models for different targets. These appearance models are based on covariance descriptor extracted from tracklets given by a short-term tracking algorithm. Short-term tracking relies on object descriptors tuned by a controller which copes with context variation over time. We link tracklets by using discriminative analysis on a Riemannian manifold. Our evaluation shows that by applying this discriminative analysis, we can reduce false alarms and identity switches, not only for tracking in a single camera but also for matching object appearances between non-overlapping cameras. Slawomir Bak, Duc Phu Chau, Julien Badie, Etienne Corvée, François Brémond, Monique Thonnat |
ICIP | 5 |
| 2012 | Boosted human re-identification using Riemannian manifolds
Slawomir Bak, Etienne Corvée, François Brémond, Monique Thonnat |
Image Vis. Comput. | 3 |
| 2012 | Recognizing Gestures by Learning Local Motion Signatures of HOG DescriptorsabstractWe introduce a new gesture recognition framework based on learning local motion signatures (LMSs) of HOG descriptors introduced by [1]. Our main contribution is to propose a new probabilistic learning-classification scheme based on a reliable tracking of local features. After the generation of these LMSs computed on one individual by tracking Histograms of Oriented Gradient (HOG) [2] descriptor, we learn a codebook of video-words (i.e., clusters of LMSs) using k-means algorithm on a learning gesture video database. Then, the video-words are compacted to a code-book of codewords by the Maximization of Mutual Information (MMI) algorithm. At the final step, we compare the LMSs generated for a new gesture w.r.t. the learned code-book via the k-nearest neighbors (k-NN) algorithm and a novel voting strategy. Our main contribution is the handling of the N to N mapping between codewords and gesture labels within the proposed voting strategy. Experiments have been carried out on two public gesture databases: KTH [3] and IXMAS [4]. Results show that the proposed method outperforms recent state-of-the-art methods. Mohamed Bécha Kaâniche, François Brémond |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Multiple-shot human re-identification by Mean Riemannian Covariance GridabstractHuman re-identification is defined as a requirement to determine whether a given individual has already appeared over a network of cameras. This problem is particularly hard by significant appearance changes across different camera views. In order to re-identify people a human signature should handle difference in illumination, pose and camera parameters. We propose a new appearance model combining information from multiple images to obtain highly discriminative human signature, called Mean Riemannian Covariance Grid (MRCG). The method is evaluated and compared with the state of the art using benchmark video sequences from the ETHZ and the i-LIDS datasets. We demonstrate that the proposed approach outperforms state of the art methods. Finally, the results of our approach are shown on two other more pertinent datasets. Slawomir Bak, Etienne Corvée, François Brémond, Monique Thonnat |
AVSS | 3 |
| 2011 | Evaluation of Local Descriptors for Action Recognition in Videos
Piotr Bilinski, François Brémond |
ICVS | 2 |
| 2011 | A Cognitive Vision System for Nuclear Fusion Device Monitoring
Vincent Martin 0001, Victor Moncada, Jean-Marcel Travere, Thierry Loarer, François Brémond, Guillaume Charpiat, Monique Thonnat |
ICVS | 5 |
| 2011 | Unsupervised Activity Extraction on Long-Term Video Recordings Employing Soft Computing Relations
Jose Luis Patino, Murray Evans, James M. Ferryman, François Brémond, Monique Thonnat |
ICVS | 4 |
| 2011 | Unsupervised Discovery, Modeling, and Analysis of Long Term Activities
Guido Pusiol, François Brémond, Monique Thonnat |
ICVS | 2 |
| 2011 | Probabilistic Recognition of Complex Event
Rim Romdhane, Bernard Boulay, François Brémond, Monique Thonnat |
ICVS | 3 |
| 2011 | Online learning neural tracker
Suresh Sundaram 0002, François Brémond, Monique Thonnat, Hyoung Joong Kim |
Neurocomputing | 2 |
| 2010 | Person Re-identification Using Haar-based and DCD-based SignatureabstractIn many surveillance systems there is a requirement to determine whether a given person of interest has already been observed over a network of cameras. This paper presents two approaches for this person re-identification problem. In general the human appearance obtained in one camera is usually different from the ones obtained in another camera. In order to re-identify people the human signature should handle difference in illumination, pose and camera parameters. Our appearance models are based on hoar-like features and dominant color descriptors. The AdaBoost scheme is applied to both descriptors to achieve the most invariant and discriminative signature. The methods are evaluated using benchmark video sequences with different camera views where people are automatically detected using Histograms of Oriented Gradients (HOG). The reidentification performance is presented using the cumulative matching characteristic (CMC) curve. Slawomir Bak, Etienne Corvée, François Brémond, Monique Thonnat |
AVSS | 3 |
| 2010 | Person Re-identification Using Spatial Covariance Regions of Human Body PartsabstractIn many surveillance systems there is a requirement to determine whether a given person of interest has already been observed over a network of cameras. This is the person re-identification problem. The human appearance obtained in one camera is usually different from the ones obtained in another camera. In order to re-identify people the human signature should handle difference in illumination, pose and camera parameters. We propose a new appearance model based on spatial covariance regions extracted from human body parts. The new spatial pyramid scheme is applied to capture the correlation between human body parts in order to obtain a discriminative human signature. The human body parts are automatically detected using Histograms of Oriented Gradients (HOG). The method is evaluated using benchmark video sequences from i-LIDS Multiple-Camera Tracking Scenario data set. The re-identification performance is presented using the cumulative matching characteristic (CMC) curve. Finally, we show that the proposed approach outperforms state of the art methods. Slawomir Bak, Etienne Corvée, François Brémond, Monique Thonnat |
AVSS | 3 |
| 2010 | Body Parts Detection for People Tracking Using Trees of Histogram of Oriented Gradient DescriptorsabstractVision algorithms face many challenging issues when it comes to analyze human activities in video surveillance applications.For instance, occlusions makes the detection and tracking of people a hard task to perform. Hence advanced and adapted solutions are required to analyze the content of video sequences. We here present a people detection algorithm based on a hierarchical tree of Histogram of Oriented Gradients referred to as HOG. The detection is coupled with independently trained body part detectors to enhance the detection performance and to reach state of the art performances. We adopt a person tracking scheme which calculates HOG dissimilarities between detected persons throughout a sequence. The algorithms are tested in videos with challenging situations such as occlusions. False alarms are further reduced by using 2D and 3D information of moving objects segmented from a background reference frame. Etienne Corvée, François Brémond |
AVSS | 2 |
| 2010 | Intelligent Video Systems: A Review of Performance Evaluation Metrics that Use Mapping ProceduresabstractIn Intelligent Video Systems, most of the recent advanced performance evaluation metrics perform a stage of mapping data between the system results and ground truth. This paper aims to review these metrics using a proposed framework. It will focus on metrics for events detection, objects detection and objects tracking systems. Xavier Desurmont, Cyril Carincotte, François Brémond |
AVSS | 3 |
| 2010 | Video Activity Extraction and Reporting with Incremental Unsupervised LearningabstractThe present work presents a new method for activity extraction and reporting from video based on the aggregation of fuzzy relations. Trajectory clustering is first employed mainly to discover the points of entry and exit of mobiles appearing in the scene. In a second step, proximity relations between resulting clusters of detected mobiles and contextual elements from the scene are modeled employing fuzzy relations. These can then be aggregated employing typical soft-computing algebra. A clustering algorithm based on the transitive closure calculation of the fuzzy relations allows building the structure of the scene and characterises the ongoing different activities of the scene. Discovered activity zones can be reported as activity maps with different granularities thanks to the analysis of the transitive closure matrix. Taking advantage of the soft relation properties, activity zones and related activities can be labeled in a more human-like language. We present results obtained on real videos corresponding to apron monitoring in the Toulouse airport in France. Jose Luis Patino, François Brémond, Murray Evans, Ali Shahrokni, James M. Ferryman |
AVSS | 2 |
| 2010 | Trajectory Based Activity DiscoveryabstractHAL is a multi-disciplinary open access archive for the deposit and dissemination of sci-entific research documents, whether they are pub-lished or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L’archive ouverte pluridisciplinaire HAL, est destinée au dépôt et a ̀ la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d’enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. Guido Pusiol, François Brémond, Monique Thonnat |
AVSS | 2 |
| 2010 | A Framework Dealing with Uncertainty for Complex Event RecognitionabstractThis paper presents a constraint-based approach for video event recognition with probabilistic reasoning for handling uncertainty. The main advantage of constraint-based approaches is the possibility for human expert to model composite events with complex temporal constraints. But the approaches are usually deterministic and do not enable the convenient mechanism of probability reasoning to handle the uncertainty. The first advantage of the proposed approach is the ability to model and recognize composite events with complex temporal constraints. The second advantage is that probability theory provides a consistent framework for dealing with uncertain knowledge for a robust and reliable recognition of complex event. This approach is evaluated with 4 real healthcare videos and a public video ETISEO'06. The results are compared with state of the art method. The comparison shows that the proposed approach improves significantly the process of recognition and characterizes the likelihood of the recognized events. Rim Romdhane, François Brémond, Monique Thonnat |
AVSS | 2 |
| 2010 | An Activity Monitoring System for Real Elderly at Home: Validation StudyabstractSince the population of the elderly grows highly, the improvement of the quality of life of elderly at home is of a great importance. This can be achieved through the development of technologies for monitoring their activities at home. In this context, we propose an activity monitoring system which aims to achieve behavior analysis of elderly people. The proposed system consists of an approach combining heterogeneous sensor data to recognize activities at home. This approach combines data provided by video cameras with data provided by environmental sensors attached to house furnishings. In this paper, we validate the proposed activity monitoring system for the recognition of a set of daily activities (e.g. using kitchen equipment, preparing meal) for 9 real elderly volunteers living in an experimental apartment. We compare the behavioral profile between the 9 elderly volunteers. This study shows that the proposed system is thoroughly accepted by the elderly and it is also well appreciated by the medical staff. Nadia Zouba, François Brémond, Monique Thonnat |
AVSS | 2 |
| 2010 | Gesture recognition by learning local motion signaturesabstractThis paper overviews a new gesture recognition framework based on learning local motion signatures (LMSs) introduced by [5]. After the generation of these LMSs computed on one individual by tracking Histograms of Oriented Gradient (HOG) [2] descriptor, we learn a codebook of video-words (i.e. clusters of LMSs) using k-means algorithm on a learning gesture video database. Then the video-words are compacted to a codebook of code-words by the Maximization of Mutual Information (MMI) algorithm. At the final step, we compare the LMSs generated for a new gesture w.r.t. the learned codebook via the k-nearest neighbors (k-NN) algorithm and a novel voting strategy. Our main contribution is the handling of the N to N mapping between code-words and gesture labels with the proposed voting strategy. Experiments have been carried out on two public gesture databases: KTH [16] and IXMAS [19]. Results show that the proposed method outperforms recent state-of-the-art methods. Mohamed Bécha Kaâniche, François Brémond |
CVPR | 2 |
| 2010 | On-Line Video Recognition and Counting of Harmful InsectsabstractThis article is concerned with on-line counting of harmful insects of certain species in videos in the framework of in situ video-surveillance that aims at the early detection of prominent pest attacks in greenhouse crops. The video-processing challenges that need to be coped with concern mainly the low spatial resolution and color contrast of the objects of interest in the videos, the outdoor issues and the video-processing which needs to be done in quasi-real time. Thus, we propose an approach which makes use of a pattern recognition algorithm to extract the locations of the harmful insects of interest in a video, which we combine with some video-processing algorithms in order to achieve an on-line video-surveillance solution. The system has been validated off-line on the whiteflie species (one potential harmful insect) and has shown acceptable performance in terms of accuracy versus computational time. Ikhlef Bechar, Sabine Moisan, Monique Thonnat, François Brémond |
ICPR | 4 |
| 2010 | Activity discovery from video employing soft computing relationsabstractThe present work presents a novel approach for activity extraction and knowledge discovery from video. Spatial and temporal properties from detected mobile objects are modeled employing fuzzy relations. These can then be aggregated employing typical soft-computing algebra. A clustering algorithm based on the transitive closure calculation of the fuzzy relations allows finding spatio-temporal patterns of activity. We employ trajectory-based analysis of mobiles in the video to discover the points of entry and exit of mobiles appearing in the scene and ultimately deduce the different areas of activity in the scene. These areas can be reported as activity maps with different granularities thanks to the analysis of the transitive closure matrix of the mobile fuzzy spatial relations. Discovered activity zones and spatio-temporal patterns of activity can be labeled in a human-like language. We present results obtained on real videos corresponding to apron monitoring in the Toulouse airport in France. Jose Luis Patino, François Brémond, Monique Thonnat |
IJCNN | 2 |
| 2009 | Tracking HoG Descriptors for Gesture RecognitionabstractWe introduce a new HoG (Histogram of Oriented Gradients) tracker for Gesture Recognition. Our main contribution is to build HoG trajectory descriptors (representing local motion) which are used for gesture recognition. First,we select for each individual in the scene a set of corner points to determine textured regions where to compute 2D HoG descriptors. Second, we track these 2D HoG descriptors in order to build temporal HoG descriptors. Lost descriptors are replaced by newly detected ones. Finally, we extract the local motion descriptors to learn offline a set of given gestures.Then, a new video can be classified according to the gesture occurring in the video. Results shows that the tracker performs well compared to KLT tracker. The generated local motion descriptors are validated through gesture learning-classification using the KTH action database. Mohamed Bécha Kaâniche, François Brémond |
AVSS | 2 |
| 2009 | Multisensor Fusion for Monitoring Elderly Activities at HomeabstractIn this paper we propose a new multisensor based activity recognition approach which uses video cameras and environmental sensors in order to recognize interesting elderly activities at home. This approach aims to provide accuracy and robustness to the activity recognition system. In the proposed approach, we choose to perform fusion at the high-level (event level) by combining video events with environmental events. To measure the accuracy of the proposed approach, we have tested a set of human activities in an experimental laboratory. The experiment consists of a scenario of daily activities performed by fourteen volunteers (aged from 60 to 85 years). Each volunteer has been observed during 4 hours and 14 video scenes have been acquired by 4 video cameras (about ten frames per second). The fourteen volunteers were asked to perform a set of household activities, such as preparing a meal, taking a meal, washing dishes, cleaning the kitchen, and watching TV. Each volunteer was alone in the laboratory during the experiment. Nadia Zouba, François Brémond, Monique Thonnat |
AVSS | 2 |
| 2009 | Incremental Video Event Learning
Marcos Zúñiga, François Brémond, Monique Thonnat |
ICVS | 2 |
| 2009 | Surveillance Video Indexing and Retrieval Using Object Features and Semantic EventsabstractIn this paper, we propose an approach for surveillance video indexing and retrieval. The objective of this approach is to answer five main challenges we have met in this domain: (1) the lack of means for finding data from the indexed databases, (2) the lack of approaches working at different abstraction levels, (3) imprecise indexing, (4) incomplete indexing, (5) the lack of user-centered search. We propose a new data model containing two main types of extracted video contents: physical objects and events. Based on this data model, we present a new rich and flexible query language. This language works at different abstraction levels, provides both exact and approximate matching and takes into account users' interest. In order to work with the imprecise indexing, two new methods respectively for object representation and object matching are proposed. Videos from two projects which have been partially indexed are used to validate the proposed approach. We have analyzed both query language usage and retrieval results. The obtained retrieval results analyzed by the average normalized ranks are promising. The retrieval results at the object level are compared with another state of the art approach. Thi-Lan Le, Monique Thonnat, Alain Boucher, François Brémond |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2008 | Crowd Behavior Recognition for Video Surveillance
Shobhit Saxena, François Brémond, Monique Thonnat, Ruihua Ma |
ACIVS | 2 |
| 2008 | Commentary Paper 2 on Action Signature: A Novel Holistic Representation for Action RecognitionabstractThis paper describes a method for action recognition using a classification algorithm based on mixtures of von Mises distributions processing action signatures. An action signature is a 1D sequence of angles, forming a trajectory, which are extracted from a 2D map of adjusted orientations (subtracting the average orientation) of the gradient of the motion-history image. To obtain the action signature, the authors scan the image along the direction given by the average gradient orientation, selecting only the points for which the motion energy is equal to 1. François Brémond, Mohamed Bécha Kaâniche |
AVSS | 1 |
| 2008 | Commentary Paper on "Learning and Classification of Trajectories in Dynamic Scenes: A General Framework for Live Video Analysis"abstractThe paper describes a general platform for live video analysis. The first stage of the platform is to build a topological scene description by learning the location of nodes (i.e. zones), which are called points of interest. There are two kinds of points of interest, the entry-exit zones (areas where moving object appear and disappear in the scene) and the stopping zones (areas where the moving objects have slow speed or remain in a circle of radius R for more than t seconds). The zones are modelled by 2D Gaussian methods. The routes between nodes are learned considering only the spatial location of trajectories in the image scene and using fuzzy C means (FCM) clusterization of the trajectories that begin in an entry zone, end in an exit zone and do not remain in a stop zone. The main trajectory cluster points are aligned using dynamic time warping and merged if the Euclidean distance is lower than a threshold.The second stage consists in the modelling of the paths by introducing not only the spatial location of the trajectories but the dynamics as well to analyze behaviour. The spatio-temporal path properties are encoded using Hidden Markov Models. The platform makes one model for each cluster computed in the previous stage. The training of each HMM model is done with the paths associated to each FCM cluster. The platform adds new models by using a batch update procedure. Trajectories that do not fit in any of the models are collected and re-clustered periodically. The HMMs are updated using maximum likehood linear regression (MLLR). Each time a new trajectory is classified into a path, a transformation is learned and applied to the mean of each of the HMM states updating its corresponding path model.The last stage comprises the behaviour analysis. Each novel trajectory detected is classified into a path by comparison with all the HMMs using forward-backward procedure finding the HMM with the maximum likelihood. Anomalous trajectories are recognized deciding that its likelihood is low, by comparing the likelihood with a decision threshold. The decision threshold is learned during training. The platform also provides an online tracking analysis method. In this case a small window of the last trajectory points is analysed, this window is constantly updated with incoming points. The live tracking classification is done considering only the most recent points of the window by comparing the points with the HMMs to estimate the likelihood at each time the window is updated. The platform can detect abnormalities during live tracking by a similar method used with complete trajectories. Path prediction is described using the HMMs by calculating the top 3 best fit paths determined with the HMM likelihoods and then estimating the probability of the incomplete trajectory to remain in one of those paths. François Brémond, Guido Pusiol |
AVSS | 1 |
| 2008 | Shadow Removal in Indoor ScenesabstractIn this paper, we propose a shadow removal algorithm for indoor scenes. This algorithm uses three types of constraints: chromaticity consistency, texture consistency and range of shadow intensity. The chromaticity consistency is verified in both HSV and RGB color spaces. The texture verification is based on the local coherency (over a pixel neighbourhood) of intensity reduction ratio between shadows and background. Finally, for the range of shadow intensity, we define a localized lower bound of the intensity reduction ratio so that dark mobile objects are not classified as shadows. Because the chromaticity constraint is only correct if the chromaticity of ambient light is the same as that of diffuse light, our algorithms only works in the indoor scenes. Anh-Tuan Nghiem, François Brémond, Monique Thonnat |
AVSS | 2 |
| 2008 | A Query Language Combining Object Features and Semantic Events for Surveillance Video Retrieval
Thi-Lan Le, Monique Thonnat, Alain Boucher, François Brémond |
MMM | 4 |
| 2007 | ETISEO, performance evaluation for video surveillance systemsabstractThis paper presents the results of ETISEO, a performance evaluation project for video surveillance systems. Many other projects have already evaluated the performance of video surveillance systems, but more on an end-user point of view. ETISEO aims at studying the dependency between algorithms and the video characteristics. Firstly we describe ETISEO methodology which consists in addressing each video processing problem separately. Secondly, we present the main evaluation metrics of ETISEO as well as their benefits, limitations and conditions of use. Finally, we discuss about the contributions of ETISEO to the evaluation community. Anh-Tuan Nghiem, François Brémond, Monique Thonnat, Valéry Valentin |
AVSS | 2 |
| 2007 | Video understanding for complex activity recognition
Florent Fusier, Valéry Valentin, François Brémond, Monique Thonnat, Mark Borg, David Thirde, James M. Ferryman |
Mach. Vis. Appl. | 3 |
| 2007 | Real-time control of video surveillance systems with program supervision techniques
Benoît Georis, François Brémond, Monique Thonnat |
Mach. Vis. Appl. | 2 |
| 2006 | Evaluation and Knowledge Representation Formalisms to Improve Video UnderstandingabstractThis article presents a methodology to build efficient real-time semantic video understanding systems addressing real world problems. In our case, semantic video under- standing consists in the recognition of predefined scenario models in a given application domain starting from a pixel analysis up to a symbolic description of what is happening in the scene viewed by cameras. This methodology proposes to use evaluation to acquire knowledge of programs and to represent this knowledge with appropriate formalisms. First, to obtain efficiency, a formalism enables to model video processing programs and their associated parameter adaptation rules. These rules are written by experts after performing a technical evaluation. Second, a scenario for- malism enables experts to model their needs and to easily refine their scenario models to adapt them to real-life situa- tions. This refinement is performed with an end-user evalu- ation. This second part ensures that systems match end-user expectations. Results are reported for scenario recognition performances on real video sequences taken from a bank agency monitoring application. Benoît Georis, Magale Maziere, François Brémond |
ICVS | 3 |
| 2006 | A Real-Time Scene Understanding System for Airport Apron MonitoringabstractThis paper presents a distributed multi-camera visual surveillance system for automatic scene interpretation of airport aprons. The system comprises two main modules Scene Tracking and Scene Understanding. The Scene Tracking module is responsible for detecting, tracking and classifying the objects on the apron. The Scene Understanding module performs high level interpretation of the apron activities by applying cognitive spatio-temporal reasoning. The performance of the complete system is demonstrated for a range of representative test scenarios. David Thirde, Mark Borg, James M. Ferryman, Florent Fusier, Valéry Valentin, François Brémond, Monique Thonnat |
ICVS | 6 |
| 2006 | An APRIORI-based Method for Frequent Composite Event Discovery in VideosabstractWe propose a method for discovery of composite events in videos. The algorithm processes a set of primitive events such as simple spatial relations between objects obtained from a tracking system and outputs frequent event patterns which can be interpreted as frequent composite events. We use the APRIORI algorithm from the field of data mining for efficient detection of frequent patterns. We adapt this algorithm to handle temporal uncertainty in the data without losing its computational effectiveness. It is formulated as a generic framework in which the context knowledge is clearly separated from the method in form of a similarity measure for comparison between two video activities and a library of primitive events serving as a basis for the composite events. Alexander Toshev, François Brémond, Monique Thonnat |
ICVS | 2 |
| 2006 | Applying 3D human model in a posture recognition system
Bernard Boulay, François Brémond, Monique Thonnat |
Pattern Recognit. Lett. | 2 |
| 2005 | Video surveillance for aircraft activity monitoringabstractThis paper presents a complete visual surveillance system for the automatic scene interpretation of airport aprons. The system comprises two modules scene tracking and scene understanding. The scene tracking module, comprising a bottom-up methodology, and the scene understanding module, comprising a video event representation and recognition scheme, have been demonstrated to be a valid approach for apron monitoring. Mark Borg, David Thirde, James M. Ferryman, Florent Fusier, Valéry Valentin, François Brémond, Monique Thonnat |
AVSS | 6 |
| 2005 | Shape recognition based on a video and multi-sensor systemabstractWe present in this paper a real-time system for shape recognition. The proposed system is a video and multi-sensor platform that is able to classify the mobile objects evolving in the scene into several expected categories. The key of the recognition method is to compute mobile object properties thanks to the camera and sensors and then to use Bayesian classifiers. A learning phase based on ground truth data is used to train the Bayesian classifiers. Our recognition method has been integrated into an existing access control device used in public transportation (subway) at RATP (Regie Autonome des Transports Parisiens) to improve safety and comfort, to prevent fraud and to count people for statistical matters. The expected categories in this case are mainly "adult", "child", "suitcase" and "two adults close to each other". Huy-Binh Bui Ngoc, François Brémond, Monique Thonnat, Jean-Claude Faure |
AVSS | 2 |
| 2004 | Video-based event recognition: activity representation and probabilistic recognition methods
Somboon Hongeng, Ramakant Nevatia, François Brémond |
Comput. Vis. Image Underst. | 3 |
| 2003 | Recurrent Bayesian Network for the Recognition of Human Behaviors from Video
Nicolas Moënne-Loccoz, François Brémond, Monique Thonnat |
ICVS | 2 |
| 2003 | Automatic Video Interpretation: A Recognition Algorithm for Temporal Scenarios Based on Pre-compiled Scenario Models
Van-Thinh Vu, François Brémond, Monique Thonnat |
ICVS | 2 |
| 2003 | Automatic Video Interpretation: A Novel Algorithm for Temporal Scenario Recognition
Van-Thinh Vu, François Brémond, Monique Thonnat |
IJCAI | 2 |
| 2002 | Group Behavior Recognition With Multiple CamerasabstractWe propose in this paper an approach for recognizing group of people behaviors using multiple cameras with overlapping FOVs (Field Of View). In this context, Behavior recognition first relies on low level motion detection and frame to frame tracking which generate a graph of mobile objects for each camera. Second, to take advantage of all cameras observing the same scene, a combination mechanism is performed to combine the graphs computed for each camera into a global one. This global graph is then used for long term tracking of groups of people evolving in the scene. Finally, the result of the group tracking is used by a higher level module which recognizes predefined scenarios corresponding to specific group behaviors. This article focuses on the graphs combination mechanism and on the recognition of group behaviors. At the end, results on these two algorithms are described. Frédéric Cupillard, François Brémond, Monique Thonnat |
WACV | 2 |
| 2001 | Tracking multiple individuals for video communicationabstractWe propose a new interpretation platform dedicated to video communication. Basically, a video communication system takes an image flow from a camera and broadcasts it in a computer network. The goal of interpretation is to enable the video communication system to adapt automatically the broadcasted image flow, using image filtering, blurring and zooming. For that we need to detect and track individuals in office scenes, then to understand their behaviour. We focus on a new tracking method based on a 3D model of the scene, on explicit models of individuals and on the computation of several possible paths for each individual. The main issues of the tracking algorithm are presented. Finally we show the results of our algorithm for several sequences illustrating office activities in everyday situations. Alberto Avanzi, François Brémond, Monique Thonnat |
ICIP (2) | 2 |
| 2001 | Event Detection and Analysis from Video StreamsabstractWe present a system which takes as input a video stream obtained from an airborne moving platform and produces an analysis of the behavior of the moving objects in the scene. To achieve this functionality, our system relies on two modular blocks. The first one detects and tracks moving regions in the sequence. It uses a set of features at multiple scales to stabilize the image sequence, that is, to compensate for the motion of the observer, then extracts regions with residual motion and uses an attribute graph representation to infer their trajectories. The second module takes as input these trajectories, together with user-provided information in the form of geospatial context and goal context to instantiate likely scenarios. We present details of the system, together with results on a number of real video sequences and also provide a quantitative analysis of the results. Gérard G. Medioni, Isaac Cohen, François Brémond, Somboon Hongeng, Ramakant Nevatia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2000 | Representation and Optimal Recognition of Human ActivitiesabstractTowards the goal of realizing a generic automatic human activity recognition system, a new formalism is proposed. Activities are described by a chained hierarchical representation using three type of entities: image features, mobile object properties and scenarios. Taking image features of tracked moving regions from an image sequence as input, mobile object properties are first computed by specific methods while noise is suppressed by statistical methods. Scenarios are recognized from mobile object properties based on Bayesian analysis. Several scenarios are recognized by an algorithm using a probabilistic finite-state automaton (a variant of structured HMM). A demonstration of the optimality of this recognition method is discussed. Finally, the validity and the effectiveness of our approach is demonstrated on both real-world and perturbed data. Somboon Hongeng, François Brémond, Ramakant Nevatia |
CVPR | 2 |
| 2000 | Bayesian Framework for Video Surveillance ApplicationabstractThe goal of this paper is to describe and demonstrate the application of Bayesian networks in a generic automatic video surveillance system. Taking image features of tracked moving regions from an image sequence as input, mobile object properties are first computed and noise is suppressed by statistical methods. The probability that a scenario occurs is then computed from these mobile object properties through several layers of naive Bayesian classifiers (or a Bayesian network). Several issues and solutions regarding the efficiency of the Bayesian network are discussed. For example, the parameters of the networks, which represent rare activities (typical of video surveillance applications), can be learned from image sequences of similar scenarios which are more common. We demonstrate the effectiveness of our approach by training the networks with 600 image frames belonging to one domain of interest and applying them to image sequences in a different domain. Somboon Hongeng, François Brémond, Ramakant Nevatia |
ICPR | 2 |
| 1998 | Issues of representing context illustrated by video-surveillance applications
François Brémond, Monique Thonnat |
Int. J. Hum. Comput. Stud. | 1 |
| 1998 | Tracking multiple nonrigid objects in video sequencesabstractThis paper presents a method to track multiple nonrigid objects in video sequences. First, we present related works on tracking methods. Second, we describe our proposed approach. We use the notion of target to represent the perception of object motion. To handle the particularities of nonrigid objects we define a target as an individually tracked moving region or as a group of moving regions globally tracked. Then we explain how to compute the trajectory of a target and how to compute the correspondences between known targets and moving regions newly detected. In the case of an ambiguous correspondence we define a compound target to freeze the associations between targets and moving regions until a more accurate information is available. Finally we provide an example to illustrate the way we have implemented the proposed tracking method for video-surveillance applications. François Brémond, Monique Thonnat |
IEEE Trans. Circuits Syst. Video Technol. | 1 |