Antitza Dantcheva

dblp:13/2986 · DBLP profile ↗
← Back
54ranked-venue papers
8as first author
34since 2021 · last 2026
0000-0003-0107-7029ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 6 first-author · 23 since 2021Artificial intelligence and machine learning · 36 · 1 first-author · 25 since 2021Security and privacy · 4 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Generalizable Deepfake Detection via Simplicity-Bias-Aware CLIP Adaptation
Charbel Yahchouchi, Noemi Roggero, Laurent Saroul, Antitza Dantcheva
ICPR (10)4
2026 Denoise, Divide, Distill, and Predict D3P: Towards Forecasting Long-horizon Real-world Anomaly from Normalcy
abstract
Forecasting abnormal human behavior (AHB) in unconstrained real-world environments is critical for enabling proactive safety interventions. Unlike short-term anomaly detection, long-horizon forecasting offers a vital reaction window but remains underexplored due to three core challenges: (i) noisy, complex human–agent interactions; (ii) weak temporal coupling between normal observations and distant anomalies; and (iii) data scarcity limiting the scalability of autoregressive models. To address these, we propose ${{\mathcal{D}}^3}{\mathcal{P}}$ (Denoise, Divide, Distill, and Predict), a novel encoder–decoder framework that bridges denoised pasts with distilled autoregressive futures. Our Differential Past Encoder (DiPE) disentangles scene-level and object-level dynamics via differential attention, suppressing irrelevant interactions and enhancing discriminative cues. The Distilled Future Auto-Regressive Decoder (D-FAD) adopts a divide-and-conquer strategy, segmenting future queries into temporal chunks for sequential prediction, while leveraging distillation to balance robustness and latency. We validate our approach on the AHB-F benchmark, the only dataset dedicated to abnormal behavior forecasting, and further integrate D-FAD with several state-of-the-art methods. In all cases, our framework consistently outperforms prior work in both forecasting accuracy and computational efficiency.
Quentin Mérilleau, Snehashis Majhi, Antitza Dantcheva, Quan Kong, Lorenzo Garattoni, Gianpiero Francesca, François Brémond
WACV3
2026 AI killed the Video Star. Audio-Driven Diffusion Model for Expressive Talking Head Generation
Baptiste Chopin, Tashvik Dhamija, Pranav Balaji, Yaohui Wang 0001, Antitza Dantcheva
Int. J. Comput. Vis.5
2025 Are Attention Maps Richer than we Imagined for Action Recognition?
abstract
Deep learning models are becoming more general and robust by the day. Specifically, image foundation models have recently shown exponential growth. In this work, we introduce a way to exploit this growth in the field of video classification. The basic idea here is that if we have a good understanding of space, we should not require complicated spatio-temporal processing. We introduce Attention Map (AM) flow, a way to identify the location of local changes between two frames in a video, without adding additional parameters specifically for it. We utilise adapters, which have been growing in popularity in the field of parameterefficient transfer learning. These help us incorporate AM flow in a pretrained image model without the need of finetuning it. With just these changes and minimal temporal processing, an image model is able to achieve state-of-the- art results on popular action recognition datasets with low training time and requiring minimal pretraining. This work explores the theory behind this idea and the intricacies involved. Through relevant experiments, we show the efficacy of this method and discuss various ideas to take this work forward. We use kinetics-400, something-something v2 and Toyota smarthome datasets and achieve state-of-the-art or comparable results. We also show that video models suffer from extensive pretraining on multiple datasets and a large training time, but our work answers these problems. actionrecognition transformers image-to-video-models
Tanay Agrawal, Abid Ali 0002, Antitza Dantcheva, François Brémond
AVSS3
2025 Just Dance with pi! A Poly-modal Inductor for Weakly-supervised Video Anomaly Detection
abstract
Weakly-supervised methods for video anomaly detection (VAD) are conventionally based merely on RGB spatio-temporal features, which continues to limit their reliability in real-world scenarios. This is due to the fact that RGB-features are not sufficiently distinctive in setting apart categories such as shoplifting from visually similar events. Therefore, towards robust complex real-world VAD, it is essential to augment RGB spatio-temporal features by additional modalities. Motivated by this, we introduce the Poly-modal Induced framework for VAD: "PI-VAD" (or π-VAD), a novel approach that augments RGB representations by five additional modalities. Specifically, the modalities include sensitivity to fine-grained motion (Pose), three dimensional scene and entity representation (Depth), surrounding objects (Panoptic masks), global motion (optical flow), as well as language cues (VLM). Each modality represents an axis of a polygon, streamlined to add salient cues to RGB. π-VAD includes two plug-in modules, namely Pseudo-modality Generation module and Cross Modal Induction module, which generate modality-specific prototypical representation and, thereby, induce multi-modal information into RGB cues. These modules operate by performing anomaly-aware auxiliary tasks and necessitate five modality backbones – only during training. Notably, π-VAD achieves state-of-the-art accuracy on three prominent VAD datasets encompassing real-world scenarios, without requiring the computational overhead of five modality backbones at inference.
Snehashis Majhi, Giacomo D'Amicantonio, Antitza Dantcheva, Quan Kong, Lorenzo Garattoni, Gianpiero Francesca, Egor Bondarev, François Brémond
CVPR3
2025 Scaling Action Detection: AdaTAD++ with Transformer-Enhanced Temporal-Spatial Adaptation
Tanay Agrawal, Abid Ali 0002, Antitza Dantcheva, François Brémond
ICCV3
2025 Guess Future Anomalies from Normalcy: Forecasting Abnormal Behavior in Real-World Videos
abstract
Forecasting Abnormal Human Behavior (AHB) aims to predict unusual behavior in advance by analyzing early patterns of normal human interactions. Unlike typical action prediction methods, this task focuses on observing only normal interactions to predict both, short and long term future abnormal behavior. Despite its affirmative impact on society, AHB prediction remains under-explored in current research. This is primarily due to the challenges involved in anticipating complex human behaviors and interactions with surrounding agents in real-world situations. Further, there exists an underlying uncertainty between the early normal patterns and the future abnormal behavior, thereby making the prediction harder. To address these challenges, we introduce a novel transformer model that improves early interaction modeling by accounting for uncertainties in both, observations and future outcomes. To the best of our knowledge, we are the first to explore the task. Therefore, we present a new comprehensive dataset referred to as “AHB-F”††Code, Models, Dataset: https://github.com/snehashismajhi/AHB-F, which features real-world scenarios with complex human interactions. The AHB-F has a deterministic evaluation protocol that ensures only normal frames to be observed for long and short term future prediction. We extensively evaluate and compare competitive action anticipation methods on our benchmark. Our results show that our method consistently outperforms existing action anticipation approaches, both in quantitative and qualitative evaluations.
Snehashis Majhi, Mohammed Guermal, Antitza Dantcheva, Quan Kong, Lorenzo Garattoni, Gianpiero Francesca, François Brémond
WACV3
2025 LEO: Generative Latent Image Animator for Human Video Synthesis
Yaohui Wang 0001, Xin Ma 0031, Cunjian Chen, Antitza Dantcheva, Bo Dai 0002, Yu Qiao 0001
Int. J. Comput. Vis.5
2025 Beyond the visible: A survey on cross-spectral face recognition
David Anghelone, Cunjian Chen, Arun Ross, Antitza Dantcheva
Neurocomputing4
2024 DiffTV: Identity-Preserved Thermal-to-Visible Face Translation via Feature Alignment and Dual-Stage Conditions
abstract
The thermal-to-visible (T2V) face translation task is essential for enabling face verification in low-light or dark conditions by converting thermal infrared faces into their visible counterparts. However, this task faces two primary challenges. First, the inherent differences between the modalities hinder the effective use of thermal information to guide RGB face reconstruction. Second, translated RGB faces often lack the identity details of the corresponding visible faces, such as skin color. To tackle these challenges, we introduce DiffTV, the first Latent Diffusion Model (LDM) specifically designed for T2V facial image translation with a focus on preserving identity. Our approach proposes a novel heterogeneous feature alignment strategy that bridges the modal gap and extracts both coarse-and fine-grained identity features consistent with visible images. Furthermore, a dual-stage condition injection strategy introduces control information to guide identity-preserved translation. Experimental results demonstrate the superior performance of DiffTV, particularly in scenarios where maintaining identity integrity is critical.
Guiqin Zhao, Guoli Wang 0004, Zejin Wang, Antitza Dantcheva, Lan Du 0002, Cunjian Chen
ACM Multimedia6
2024 View-Invariant Skeleton Action Representation Learning via Motion Retargeting
Di Yang 0002, Yaohui Wang 0001, Antitza Dantcheva, Lorenzo Garattoni, Gianpiero Francesca, François Brémond
Int. J. Comput. Vis.3
2024 Unsupervised domain alignment of fingerprint denoising models using pseudo annotations
Indu Joshi, Tushar Prakash, Antitza Dantcheva, Sumantra Dutta Roy, Prem Kumar Kalra
Multim. Tools Appl.4
2024 Synthetic Data in Human Analysis: A Survey
abstract
Deep neural networks have become prevalent in human analysis, boosting the performance of applications, such as biometric recognition, action recognition, as well as person re-identification. However, the performance of such networks scales with the available training data. In human analysis, the demand for large-scale datasets poses a severe challenge, as data collection is tedious, time-expensive, costly and must comply with data protection laws. Current research investigates the generation of synthetic data as an efficient and privacy-ensuring alternative to collecting real data in the field. This survey introduces the basic definitions and methodologies, essential when generating and employing synthetic data for human analysis. We summarise current state-of-the-art methods and the main benefits of using synthetic data. We also provide an overview of publicly available synthetic datasets and generation models. Finally, we discuss limitations, as well as open research problems in this field. This survey is intended for researchers and practitioners in the field of human analysis.
Indu Joshi, Marcel Grimmer, Christian Rathgeb, Christoph Busch 0001, François Brémond, Antitza Dantcheva
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 LIA: Latent Image Animator
abstract
Previous animation techniques mainly focus on leveraging explicit structure representations (e.g., meshes or keypoints) for transferring motion from driving videos to source images. However, such methods are challenged with large appearance variations between source and driving data, as well as require complex additional modules to respectively model appearance and motion. Towards addressing these issues, we introduce the Latent Image Animator (LIA), streamlined to animate high-resolution images. LIA is designed as a simple autoencoder that does not rely on explicit representations. Motion transfer in the pixel space is modeled as linear navigation of motion codes in the latent space. Specifically such navigation is represented as an orthogonal motion dictionary learned in a self-supervised manner based on proposed Linear Motion Decomposition (LMD). Extensive experimental results demonstrate that LIA outperforms state-of-the-art on VoxCeleb, TaichiHD, and TED-talk datasets with respect to video quality and spatio-temporal consistency. In addition LIA is well equipped for zero-shot high-resolution image animation. Code, models, and demo video are available at https://github.com/wyhsirius/LIA.
Yaohui Wang 0001, Di Yang 0002, François Brémond, Antitza Dantcheva
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Self-Supervised Video Representation Learning via Latent Time Navigation
abstract
Self-supervised video representation learning aimed at maximizing similarity between different temporal segments of one video, in order to enforce feature persistence over time. This leads to loss of pertinent information related to temporal relationships, rendering actions such as `enter' and `leave' to be indistinguishable. To mitigate this limitation, we propose Latent Time Navigation (LTN), a time parameterized contrastive learning strategy that is streamlined to capture fine-grained motions. Specifically, we maximize the representation similarity between different video segments from one video, while maintaining their representations time-aware along a subspace of the latent representation code including an orthogonal basis to represent temporal changes. Our extensive experimental analysis suggests that learning video representations by LTN consistently improves performance of action classification in fine-grained and human-oriented tasks (e.g., on Toyota Smarthome dataset). In addition, we demonstrate that our proposed model, when pre-trained on Kinetics-400, generalizes well onto the unseen real world video benchmark datasets UCF101 and HMDB51, achieving state-of-the-art performance in action recognition.
Di Yang 0002, Yaohui Wang 0001, Quan Kong, Antitza Dantcheva, Lorenzo Garattoni, Gianpiero Francesca, François Brémond
AAAI4
2023 LAC - Latent Action Composition for Skeleton-based Action Segmentation
abstract
Skeleton-based action segmentation requires recognizing composable actions in untrimmed videos. Current approaches decouple this problem by first extracting local visual features from skeleton sequences and then processing them by a temporal model to classify frame-wise actions. However, their performances remain limited as the visual features cannot sufficiently express composable actions. In this context, we propose Latent Action Composition (LAC)1, a novel self-supervised framework aiming at learning from synthesized composable motions for skeleton-based action segmentation. LAC is composed of a novel generation module towards synthesizing new sequences. Specifically, we design a linear latent space in the generator to represent primitive motion. New composed motions can be synthe-sized by simply performing arithmetic operations on latent representations of multiple input skeleton sequences. LAC leverages such synthesized sequences, which have large diversity and complexity, for learning visual representations of skeletons in both sequence and frame spaces via contrastive learning. The resulting visual encoder has a high expressive power and can be effectively transferred onto action segmentation tasks by end-to-end fine-tuning without the need for additional temporal models. We conduct a study focusing on transfer-learning and we show that representations learned from pre-trained LAC outperform the state-of-the-art by a large margin on TSU, Charades, PKU-MMD datasets.
Di Yang 0002, Yaohui Wang 0001, Antitza Dantcheva, Quan Kong, Lorenzo Garattoni, Gianpiero Francesca, François Brémond
ICCV3
2023 ANYRES: Generating High-Resolution visible-face images from Low-Resolution thermal-face images
abstract
Cross-spectral Face Recognition (CFR) aims to compare facial images across different modalities, i.e., the visible and thermal spectra. CFR is more challenging than traditional face recognition (FR) due to the profound modality gap in-between spectra. As related applications range from night-vision FR to robust presentation attacks detection, acquisition involves capturing images at various distances, represented by different image resolutions. Prior approaches have addressed CFR by considering a fixed resolution, necessitating that a subject stands at a precise distance from a given sensor during acquisition, which constitutes an impractical scenario in real-life. Towards loosening this constraint, we propose ANYRES, a unified model endowed with the ability to handle a wide range of input resolutions. ANYRES generates high resolution visible images from low resolution thermal images, placing emphasis on maintaining the cross-spectral identity. We demonstrate the effectiveness of the method and present extensive FR experiments on multi-spectral paired face datasets.
David Anghelone, Sarah Lannes, Antitza Dantcheva
ICME3
2023 Face attribute analysis from structured light: an end-to-end approach
Vikas Thamizharasan, Abhijit Das 0001, Daniele Battaglino, François Brémond, Antitza Dantcheva
Multim. Tools Appl.5
2023 Learning Invariance From Generated Variance for Unsupervised Person Re-Identification
abstract
This work focuses on unsupervised representation learning in person re-identification (ReID). Recent self-supervised contrastive learning methods learn invariance by maximizing the representation similarity between two augmented views of a same image. However, traditional data augmentation may bring to the fore undesirable distortions on identity features, which is not always favorable in id-sensitive ReID tasks. In this article, we propose to replace traditional data augmentation with a generative adversarial network (GAN) that is targeted to generate augmented views for contrastive learning. A 3D mesh guided person image generator is proposed to disentangle a person image into id-related and id-unrelated features. Deviating from previous GAN-based ReID methods that only work in id-unrelated space (pose and camera style), we conduct GAN-based augmentation on both id-unrelated and id-related features. We further propose specific contrastive losses to help our network learn invariance from id-unrelated and id-related augmentations. By jointly training the generative and the contrastive modules, our method achieves new state-of-the-art unsupervised person ReID performance on mainstream large-scale benchmarks.
Hao Chen 0061, Yaohui Wang 0001, Benoit Lagadec, Antitza Dantcheva, François Brémond
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 TFLD: Thermal Face and Landmark Detection for Unconstrained Cross-spectral Face Recognition
abstract
Automated thermal-to-visible face recognition has received increased attention due to benefits related to low-light applications. Towards improvement of related matching accuracy, we hereby present TFLD, a detector of face and landmarks operating in the thermal spectrum. Our proposed TFLD is based on the architecture of YOLOv5, integrating sequential modules for face and landmark detection. We introduce a thermal face restoration scheme, in order to enhance thermal image quality and hence detection accuracy. We address data scarcity by transferring landmarks in paired visible and thermal images. Our experimental results showcase that our proposed detector accurately detects faces, as well as landmarks in a wide range of adversarial conditions. Further, TFLD achieves promising results on three benchmark multi-spectral face and landmark datasets, namely ARL-VTF, SF-TL54 and RWTH-Aachen; thereby improving the matching accuracy in cross-spectral face recognition by providing robust face alignment based on estimated facial landmarks.
David Anghelone, Sarah Lannes, Valeriya Strizhkova, Philippe Faure, Cunjian Chen, Antitza Dantcheva
IJCB6
2022 Attention-Guided Generative Adversarial Network for Explainable Thermal to Visible Face Recognition
abstract
Thermal to visible face image translation aims at synthesizing high-fidelity visible face images from thermal counterparts, placing emphasis on preserving the identity of the faces. While remarkable progress has been achieved related to the quality of synthetic images, as well as related to associated face matching accuracy, interpreting the generation process from thermal to visible face images remains an open challenge. Towards tackling this challenge, we present a novel generic attention-guided generative adversarial network (AG-GAN) for thermal to visible image translation. The AG-GAN framework is based on an encoder network that directly generates attention feature maps from an input thermal image in either, supervised or unsupervised fashion. A decoder network takes the attention maps and applies adaptive layer-instance normalization, in order to reconstruct the corresponding visible image. We show that solving thermal to visible image translation tasks through AG-GAN significantly improves the cross-spectral face matching accuracy, as well as inherently supports model explanation.
Cunjian Chen, David Anghelone, Philippe Faure, Antitza Dantcheva
IJCB4
2022 Latent Image Animator: Learning to Animate Images via Latent Space Navigation
Yaohui Wang 0001, Di Yang 0002, François Brémond, Antitza Dantcheva
ICLR4
2022 On restoration of degraded fingerprints
Indu Joshi, Ayush Utkarsh, Pravendra Singh, Antitza Dantcheva, Sumantra Dutta Roy, Prem Kumar Kalra
Multim. Tools Appl.4
2022 A Spatio-Temporal Approach for Apathy Classification
abstract
Apathy is characterized by symptoms such as reduced emotional response, lack of motivation, and limited social interaction. Current methods for apathy diagnosis require the patient’s presence in a clinic and time consuming clinical interviews, which are costly and inconvenient for both, patients and clinical staff, hindering among other large-scale diagnostics. In this work, we propose a novel spatio-temporal framework for apathy classification, which is streamlined to analyze facial dynamics and emotion in videos. Specifically, we divide the videos into smaller clips, and proceed to extract associated facial dynamics and emotion-based features. Statistical representations/descriptors based on each feature and clip serve as input of the proposed Gated Recurrent Unit (GRU)-architecture. Temporal representations of individual features at the lower level of the proposed architecture are combined at deeper layers of the proposed GRU architecture, in order to obtain the final feature-set for apathy classification. Based on extensive experiments, we show that fusion of characteristics such as emotion and facial dynamics in proposed deep-bi-directional GRU obtains an accuracy of 95.34% in apathy classification.
Abhijit Das 0001, Xuesong Niu, Antitza Dantcheva, S. L. Happy, Hu Han 0001, Radia Zeghari, Philippe Robert, Shiguang Shan, François Brémond, Xilin Chen 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 UNIK: A Unified Framework for Real-world Skeleton-based Action Recognition
Di Yang 0002, Yaohui Wang 0001, Antitza Dantcheva, Lorenzo Garattoni, Gianpiero Francesca, François Brémond
BMVC3
2021 Joint Generative and Contrastive Learning for Unsupervised Person Re-Identification
abstract
Recent self-supervised contrastive learning provides an effective approach for unsupervised person re-identification (ReID) by learning invariance from different views (transformed versions) of an input. In this paper, we incorporate a Generative Adversarial Network (GAN) and a contrastive learning module into one joint training framework. While the GAN provides online data augmentation for contrastive learning, the contrastive module learns view-invariant features for generation. In this context, we propose a mesh-based view generator. Specifically, mesh projections serve as references towards generating novel views of a person. In addition, we propose a view-invariant loss to facilitate contrastive learning between original and generated views. Deviating from previous GAN-based unsupervised ReID methods involving domain adaptation, we do not rely on a labeled source dataset, which makes our method more flexible. Extensive experimental results show that our method significantly outperforms state-of-the-art methods under both, fully unsupervised and unsupervised domain adaptive settings on several large scale ReID dat-sets. Source code and models are available under https://github.com/chenhao2345/GCL.
Hao Chen 0061, Yaohui Wang 0001, Benoit Lagadec, Antitza Dantcheva, François Brémond
CVPR4
2021 Explainable Thermal to Visible Face Recognition Using Latent-Guided Generative Adversarial Network
abstract
One of the main challenges in performing thermal-to-visible face image translation is preserving the identity across different spectral bands. Existing work does not effectively disentangle the identity from other confounding factors. In this paper, we propose a Latent-Guided Generative Adversarial Network (LG-GAN) to explicitly decompose an input image into identity code that is spectral-invariant and style code that is spectral-dependent. By using such a disentanglement, we are able to analyze the identity preservation by interpreting and visualizing the identity code. We present extensive face recognition experiments on two challenging Visible-Thermal face datasets. We show that the learned identity code is effective in preserving the identity, thus offering useful insights on interpreting and explaining thermal-to-visible face image translation.
David Anghelone, Cunjian Chen, Philippe Faure, Arun Ross, Antitza Dantcheva
FG5
2021 Demystifying Attention Mechanisms for Deepfake Detection
abstract
Manipulated images and videos, i.e., deepfakes have become increasingly realistic due to the tremendous progress of deep learning methods. However, such manipulation has triggered social concerns, necessitating the introduction of robust and reliable methods for deepfake detection. In this work, we explore a set of attention mechanisms and adapt them for the task of deepfake detection. Generally, attention mechanisms in videos modulate the representation learned by a convolutional neural network (CNN) by focusing on the salient regions across space-time. In our scenario, we aim at learning discriminative features to take into account the temporal evolution of faces to spot manipulations. To this end, we address the two research questions ‘How to use attention mechanisms?’, and ‘What type of attention is effective for the task of deepfake detection?’ Towards answering these questions, we provide a detailed study and experiments on videos tampered by four manipulation techniques, as included in the FaceForensics++ dataset. We investigate three scenarios, where the networks are trained to detect (a) all manipulated videos, (b) each manipulation technique individually, as well as (c) the veracity of videos pertaining to manipulation techniques not included in the train set.
Abhijit Das 0001, Srijan Das, Antitza Dantcheva
FG3
2021 BVPNet: Video-to-BVP Signal Prediction for Remote Heart Rate Estimation
abstract
In this paper, we propose a new method for remote photoplethysmography (rPPG) based heart rate (HR) estimation. In particular, our proposed method BVPNet is streamlined to predict the blood volume pulse (BVP) signals from face videos. Towards this, we firstly define ROIs based on facial landmarks and then extract the raw temporal signal from each ROI. Then the extracted signals are pre-processed via first-order difference and Butterworth filter and combined to form a Spatial-Temporal map (STMap). We then propose to revise U-Net, in order to predict BVP signals from the STMap. BVPNet takes into account both temporal and frequency domain losses in order to learn better than conventional models. Our experimental results suggest that our BVPNet outperforms the state-of-the-art methods on two publicly available datasets (MMSE-HR and VIPL-HR).
Abhijit Das 0001, Hao Lu 0009, Hu Han 0001, Antitza Dantcheva, Shiguang Shan, Xilin Chen 0001
FG4
2021 Emotion Editing in Head Reenactment Videos using Latent Space Manipulation
abstract
Video generation greatly benefits from integrating facial expressions, as they are highly pertinent in social interaction and hence increase realism in generated talking head videos. Motivated by this, we propose a method for editing emotions in head reenactment videos that is streamlined to modify the latent space of a pre-trained neural head reenactment system. Specifically, our method seeks to disentangle emotions from the latent pose and identity representation. The proposed learning process is based on cycle consistency and image reconstruction losses. Our results suggest that despite its simplicity, such learning successfully decomposes emotion from pose and identity. Our method reproduces facial mimics of a person from a driving video, as well as allows for emotion editing in the reenactment video. We compare our method to the state-of-art for altering emotions in reenactment videos, producing more realistic results that the state-of-art.
Valeriya Strizhkova, Yaohui Wang 0001, David Anghelone, Di Yang 0002, Antitza Dantcheva, François Brémond
FG5
2021 Self-Supervised Video Pose Representation Learning for Occlusion- Robust Action Recognition
abstract
Action recognition based on human pose has witnessed increasing attention due to its robustness to changes in appearances, environments, and view-points. Despite associated progress, one remaining challenge has to do with occlusion in real-world videos that hinders the visibility of all joints. Such occlusion impedes representation of such scenes by models that have been trained on full-body pose data, obtained in laboratory conditions with specific sensors. To address this, as a first contribution, we introduce OR- VPE, a novel video pose embedding network that is streamlined to learn an occlusion-robust representation for pose sequences in videos. In order to enable our embedding network to handle partially visible joints, we propose to incorporate a sub-graph data augmentation mechanism during training, which simulates occlusions, into a video pose encoder based on Graph Convolutional Networks (GCNs). As a second contribution, we apply a contrastive learning module to train the video pose representation in a self-supervised manner without the necessity of action annotations. This is achieved by maximizing the mutual information of the same pose sequence pruned into different spatio-temporal subgraphs. Experimental analyses show that compared to training the same encoder from scratch, our proposed OR-VPE, with pre-training on a large-scale dataset, NTU-RGB+D 120, improves the performance of the downstream action classification on Toyota Smarthome, N-UCLA and Penn Action datasets.
Di Yang 0002, Yaohui Wang 0001, Antitza Dantcheva, Lorenzo Garattoni, Gianpiero Francesca, François Brémond
FG3
2021 Data Uncertainty Guided Noise-aware Preprocessing Of Fingerprints
abstract
The effectiveness of fingerprint-based authentication systems on good quality fingerprints is established long back. However, the performance of standard fingerprint matching systems on noisy and poor quality fingerprints is far from satisfactory. Towards this, we propose a data uncertainty-based framework which enables the state-of-the-art fingerprint pre-processing models to quantify noise present in the input image and identify fingerprint regions with background noise and poor ridge clarity. Quantification of noise helps the model two folds: firstly, it makes the objective function adaptive to the noise in a particular input fingerprint and consequently, helps to achieve robust performance on noisy and distorted fingerprint regions. Secondly, it provides a noise variance map which indicates noisy pixels in the input fingerprint image. The predicted noise variance map enables the end-users to understand erroneous predictions due to noise present in the input image. Extensive experimental evaluation on 13 publicly available fingerprint databases, across different architectural choices and two fingerprint processing tasks demonstrate effectiveness of the proposed framework.
Indu Joshi, Ayush Utkarsh, Riya Kothari, Vinod K. Kurmi, Antitza Dantcheva, Sumantra Dutta Roy, Prem Kumar Kalra
IJCNN5
2021 Sensor-invariant Fingerprint ROI Segmentation Using Recurrent Adversarial Learning
abstract
A fingerprint region of interest (roi) segmentation algorithm is designed to separate the foreground fingerprint from the background noise. All the learning based state-of-the-art fingerprint roi segmentation algorithms proposed in the literature are benchmarked on scenarios when both training and testing databases consist of fingerprint images acquired from the same sensors. However, when testing is conducted on a different sensor, the segmentation performance obtained is often unsatisfactory. As a result, every time a new fingerprint sensor is used for testing, the fingerprint roi segmentation model needs to be re-trained with the fingerprint image acquired from the new sensor and its corresponding manually marked ROI. Manually marking fingerprint ROI is expensive because firstly, it is time consuming and more importantly, requires domain expertise. In order to save the human effort in generating annotations required by state-of-the-art, we propose a fingerprint roi segmentation model which aligns the features of fingerprint images derived from the unseen sensor such that they are similar to the ones obtained from the fingerprints whose ground truth roi masks are available for training. Specifically, we propose a recurrent adversarial learning based feature alignment network that helps the fingerprint roi segmentation model to learn sensor-invariant features. Consequently, sensor-invariant features learnt by the proposed roi segmentation model help it to achieve improved segmentation performance on fingerprints acquired from the new sensor. Experiments on publicly available FVC databases demonstrate the efficacy of the proposed work.
Indu Joshi, Ayush Utkarsh, Riya Kothari, Vinod K. Kurmi, Antitza Dantcheva, Sumantra Dutta Roy, Prem Kumar Kalra
IJCNN5
2021 Expression recognition with deep features extracted from holistic and part-based models
S. L. Happy, Antitza Dantcheva, François Brémond
Image Vis. Comput.2
2020 G3AN: Disentangling Appearance and Motion for Video Generation
abstract
Creating realistic human videos entails the challenge of being able to simultaneously generate both appearance, as well as motion. To tackle this challenge, we introduce G3AN, a novel spatio-temporal generative model, which seeks to capture the distribution of high dimensional video data and to model appearance and motion in disentangled manner. The latter is achieved by decomposing appearance and motion in a three-stream Generator, where the main stream aims to model spatio-temporal consistency, whereas the two auxiliary streams augment the main stream with multi-scale appearance and motion features, respectively. An extensive quantitative and qualitative analysis shows that our model systematically and significantly outperforms state-of-the-art methods on the facial expression datasets MUG and UvA-NEMO, as well as the Weizmann and UCF101 datasets on human action. Additional analysis on the learned latent representations confirms the successful decomposition of appearance and motion.
Yaohui Wang 0001, Piotr Bilinski, François Brémond, Antitza Dantcheva
CVPR4
2020 Semi-supervised Emotion Recognition using Inconsistently Annotated Data
abstract
Expression recognition remains challenging, predominantly due to (a) lack of sufficient data, (b) subtle emotion intensity, (c) subjective and inconsistent annotation, as well as due to (d) in-the-wild data containing variations in pose, intensity, and occlusion. To address such challenges in a unified framework, we propose a self-training based semi-supervised convolutional neural network (CNN) framework, which directly addresses the problem of (a) limited data by leveraging information from unannotated samples. Our method uses `successive label smoothing' to adapt to the subtle expressions and improve the model performance for (b) low-intensity expression samples. Further, we address (c) inconsistent annotations by assigning sample weights during loss computation, thereby ignoring the effect of incorrect ground-truth. We observe significant performance improvement in in-the-wild datasets by leveraging the information from the in-the-lab datasets, related to challenge (d). Associated to that, experiments on four publicly available datasets demonstrate large performance gains in cross-database performance, as well as show that the proposed method achieves to learn different expression intensities, even when trained with categorical samples.
S. L. Happy, Antitza Dantcheva, François Brémond
FG2
2020 Apathy Classification by Exploiting Task Relatedness
abstract
Apathy is characterized by symptoms such as reduced emotional response, lack of motivation, and limited social interaction. Current methods for apathy diagnosis require the patient's presence in a clinic and time consuming clinical interviews, which are costly and inconvenient for both patients and clinical staff, hindering among others large-scale diagnostics. In this work we propose a multi-task learning (MTL) framework for apathy classification based on facial analysis, entailing both emotion and facial movements. In addition, it leverages information from other auxiliary tasks (i.e., clinical scores), which might be closely or distantly related to the main task of apathy classification. Our proposed MTL approach (termed MTL+) improves apathy classification by jointly learning model weights and the relatedness of the auxiliary tasks to the main task in an iterative manner. Our results on 90 video sequences acquired from 45 subjects obtained an apathy classification accuracy of up to 80%, using the concatenated emotion and motion features. Our results further demonstrate the improved performance of MTL+ over MTL.
S. L. Happy, Antitza Dantcheva, Abhijit Das 0001, François Brémond, Radia Zeghari, Philippe Robert
FG2
2020 A video is worth more than 1000 lies. Comparing 3DCNN approaches for detecting deepfakes
abstract
Manipulated images and videos have become increasingly realistic due to the tremendous progress of deep convolutional neural networks (CNNs). While technically intriguing, such progress raises a number of social concerns related to the advent and spread of fake information and fake news. Such concerns necessitate the introduction of robust and reliable methods for fake image and video detection. Towards this in this work, we study the ability of state of the art video CNNs including 3D ResNet, 3D ResNeXt, and I3D in detecting manipulated videos. We present related experimental results on videos tampered by four manipulation techniques, as included in the FaceForensics++ dataset. We investigate three scenarios, where the networks are trained to detect (a) all manipulated videos, as well as (b) separately each manipulation technique individually. Finally and deviating from previous works, we conduct cross-manipulation results, where we (c) detect the veracity of videos pertaining to manipulation-techniques not included in the train set. Our findings clearly indicate the need for a better understanding of manipulation methods and the importance of designing algorithms that can successfully generalize onto unknown manipulations.
Yaohui Wang 0001, Antitza Dantcheva
FG2
2020 How Unique Is a Face: An Investigative Study
abstract
Face recognition has been widely accepted as a means of identification in applications ranging from border control to security in the banking sector. Surprisingly, while widely accepted, we still lack the understanding of uniqueness or distinctiveness of faces as biometric modality. In this work, we study the impact of factors such as image resolution, feature representation, database size, age and gender on uniqueness denoted by the Kullback-Leibler divergence between genuine and impostor distributions. Towards understanding the impact, we present experimental results on the datasets AT&T, LFW, IMDb-Face, as well as ND-TWINS, with the feature extraction algorithms VGGFace, VGG16, ResNet50, InceptionV3, MobileNet and DenseNet121, that reveal the quantitative impact of the named factors. While these are early results, our findings indicate the need for a better understanding of the concept of biometric uniqueness and its implication on face recognition.
Michal Balazia, S. L. Happy, François Brémond, Antitza Dantcheva
ICPR4
2020 ImaGINator: Conditional Spatio-Temporal GAN for Video Generation
abstract
Generating human videos based on single images entails the challenging simultaneous generation of realistic and visual appealing appearance and motion. In this context, we propose a novel conditional GAN architecture, namely ImaGINator, which given a single image, a condition (label of a facial expression or action) and noise, decomposes appearance and motion in both latent and high level feature spaces, generating realistic videos. This is achieved by (i) a novel spatio-temporal fusion scheme, which generates dynamic motion, while retaining appearance throughout the full video sequence by transmitting appearance (originating from the single image) through all layers of the network. In addition, we propose (ii) a novel transposed (1+2)D convolution, factorizing the transposed 3D convolutional filters into separate transposed temporal and spatial components, which yields significantly gains in video quality and speed. We extensively evaluate our approach on the facial expression datasets MUG and UvA-NEMO, as well as on the action datasets NATOPS and Weizmann. We show that our approach achieves significantly better quantitative and qualitative results than the state-of-the-art. The source code and models are available under https://github.com/wyhsirius/ImaGINator.
Yaohui Wang 0001, Piotr Bilinski, François Brémond, Antitza Dantcheva
WACV4
2019 Characterizing the State of Apathy with Facial Expression and Motion Analysis
abstract
Reduced emotional response, lack of motivation, and limited social interaction comprise the major symptoms of apathy. Current methods for apathy diagnosis require the patient's presence in a clinic, and time consuming clinical interviews and questionnaires involving medical personnel, which are costly and logistically inconvenient for patients and clinical staff, hindering among other large scale diagnostics. In this paper we introduce a novel machine learning framework to classify apathetic and non-apathetic patients based on analysis of facial dynamics, entailing both emotion and facial movement. Our approach caters to the challenging setting of current apathy assessment interviews, which include short video clips with wide face pose variations, very low-intensity expressions, and insignificant inter-class variations. We test our algorithm on a dataset consisting of 90 video sequences acquired from 45 subjects and obtained an accuracy of 84% in apathy classification. Based on extensive experiments, we show that the fusion of emotion and facial local motion produces the best feature set for apathy classification. In addition, we train regression models to predict the clinical scores related to the mental state examination (MMSE) and the neuropsychiatric apathy inventory (NPI) using the motion and emotion features. Our results suggest that the performance can be further improved by appending the predicted clinical scores to the video-based feature representation.
S. L. Happy, Antitza Dantcheva, Abhijit Das 0001, Radia Zeghari, Philippe Robert, François Brémond
FG2
2019 Robust Remote Heart Rate Estimation from Face Utilizing Spatial-temporal Attention
abstract
In this work, we propose an end-to-end approach for robust remote heart rate (HR) measurement gleaned from facial videos. Specifically the approach is based on remote photoplethysmography (rPPG), which constitutes a pulse triggered perceivable chromatic variation, sensed in RGB-face videos. Consequently, rPPGs can be affected in less-constrained settings. To unpin the shortcoming, the proposed algorithm utilizes a spatio-temporal attention mechanism, which places focus on the salient features included in rPPG-signals. In addition, we propose an effective rPPG augmentation approach, generating multiple rPPG signals with varying HRs from a single face video. Experimental results on the public datasets VIPL-HR and MMSE-HR show that the proposed method outperforms state-of-the-art algorithms in remote HR estimation.
Xuesong Niu, Xingyuan Zhao, Hu Han 0001, Abhijit Das 0001, Antitza Dantcheva, Shiguang Shan, Xilin Chen 0001
FG5
2019 Improving Face Sketch Recognition via Adversarial Sketch-Photo Transformation
abstract
Face sketch-photo transformation has broad applications in forensics, law enforcement, and digital entertainment, particular for face recognition systems that are designed for photo-to-photo matching. While there are a number of methods for face photo-to-sketch transformation, studies on sketch-to-photo transformation remain limited. In this paper, we propose a novel conditional CycleGAN for face sketch-to-photo transformation. Specifically, we leverage the advantages of CycleGAN and conditional GANs and design a feature-level loss to assure the high quality of the generated face photos from sketches. The generated face photos are used, as a replacement of face sketches, and particularly for face identification against a gallery set of mugshot photos. Experimental results on the public-domain database CUFSF show that the proposed approach is able to generate realistic photos from sketches, and the generated photos are instrumental in improving the sketch identification accuracy against a large gallery set.
Shikang Yu, Hu Han 0001, Shiguang Shan, Antitza Dantcheva, Xilin Chen 0001
FG4
2019 A Weakly Supervised learning technique for classifying facial expressions
S. L. Happy, Antitza Dantcheva, François Brémond
Pattern Recognit. Lett.2
2018 Show me your face and I will tell you your height, weight and body mass index
abstract
Body height, weight, as well as the associated and composite body mass index (BMI) are human attributes of pertinence due to their use in a number of applications including surveillance, re-identification, image retrieval systems, as well as healthcare. Previous work on automated estimation of height, weight and BMI has predominantly focused on 2D and 3D full-body images and videos. Little attention has been given to the use of face for estimating such traits. Motivated by the above, we here explore the possibility of estimating height, weight and BMI from single-shot facial images by proposing a regression method based on the 50-layers ResNet-architecture. In addition, we present a novel dataset consisting of 1026 subjects and show results, which suggest that facial images contain discriminatory information pertaining to height, weight and BMI, comparable to that of body-images and videos. Finally, we perform a gender-based analysis of the prediction of height, weight and BMI.
Antitza Dantcheva, François Brémond, Piotr Bilinski
ICPR1
2017 Gender Estimation Based on Smile-Dynamics
abstract
Automated gender estimation has numerous applications, including video surveillance, human-computer interaction, anonymous customized advertisement, and image retrieval. Most commonly, the underlying algorithms analyze the facial appearance for clues of gender. In this paper, we propose a novel method for gender estimation, which exploits dynamic features gleaned from smiles and we proceed to show that: a) facial dynamics incorporate clues for gender dimorphism and b) while for adult individuals appearance features are more accurate than dynamic features, for subjects under 18 years facial dynamics can outperform appearance features. In addition, we fuse proposed dynamics-based approach with state-of-the-art appearance-based algorithms, predominantly improving performance of the latter. Results show that smile-dynamics include pertinent and complementary to appearance gender information.
Antitza Dantcheva, François Brémond
IEEE Trans. Inf. Forensics Secur.1
2016 Image-based gender estimation from body and face across distances
abstract
Gender estimation has received increased attention due to its use in a number of pertinent security and commercial applications. Automated gender estimation algorithms are mainly based on extracting representative features from face images. In this work we study gender estimation based on information deduced jointly from face and body, extracted from single-shot images. The approach addresses challenging settings such as low-resolution-images, as well as settings when faces are occluded. Specifically the face-based features include local binary patterns (LBP) and scale-invariant feature transform (SIFT) features, projected into a PCA space. The features of the novel body-based algorithm proposed in this work include continuous shape information extracted from body silhouettes and texture information retained by HOG descriptors. Support Vector Machines (SVMs) are used for classification for body and face features. We conduct experiments on images extracted from video-sequences of the Multi-Biometric Tunnel database, emphasizing on three distance-settings: close, medium and far, ranging from full body exposure (far setting) to head and shoulders exposure (close setting). The experiments suggest that while face-based gender estimation performs best in the close-distance-setting, body-based gender estimation performs best when a large part of the body is visible. Finally we present two score-level-fusion schemes of face and body-based features, outperforming the two individual modalities in most cases.
Ester Gonzalez-Sosa, Antitza Dantcheva, Rubén Vera-Rodríguez, Jean-Luc Dugelay, François Brémond, Julian Fierrez
ICPR2
2016 What Else Does Your Biometric Data Reveal? A Survey on Soft Biometrics
abstract
Recent research has explored the possibility of extracting ancillary information from primary biometric traits viz., face, fingerprints, hand geometry, and iris. This ancillary information includes personal attributes, such as gender, age, ethnicity, hair color, height, weight, and so on. Such attributes are known as soft biometrics and have applications in surveillance and indexing biometric databases. These attributes can be used in a fusion framework to improve the matching accuracy of a primary biometric system (e.g., fusing face with gender information), or can be used to generate qualitative descriptions of an individual (e.g., young Asian female with dark eyes and brown hair). The latter is particularly useful in bridging the semantic gap between human and machine descriptions of the biometric data. In this paper, we provide an overview of soft biometrics and discuss some of the techniques that have been proposed to extract them from the image and the video data. We also introduce a taxonomy for organizing and classifying soft biometric attributes, and enumerate the strengths and limitations of these attributes in the context of an operational biometric system. Finally, we discuss open research problems in this field. This survey is intended for researchers and practitioners in the field of biometrics.
Antitza Dantcheva, Petros Elia, Arun Ross
IEEE Trans. Inf. Forensics Secur.1
2015 Assessment of female facial beauty based on anthropometric, non-permanent and acquisition characteristics
Antitza Dantcheva, Jean-Luc Dugelay
Multim. Tools Appl.1
2011 Frontal-to-side face re-identification based on hair, skin and clothes patches
abstract
Despite recent advances, face-recognition algorithms are still challenged when applied in the setting of video surveillance systems which inherently introduce variations in the pose of subjects. The present work addresses this problem, and seeks to provide a recognition algorithm that is specifically suited for a frontal-to-side re-identification setting. Deviating from classical biometric approaches, the proposed method considers color- and texture- based soft biometric traits, specifically those taken from patches of hair, skin and clothes. The proposed method and the suitability of these patch-based traits are then validated both analytically and empirically.
Antitza Dantcheva, Jean-Luc Dugelay
AVSS1
2011 On the reliability of eye color as a soft biometric trait
abstract
This work studies eye color as a soft biometric trait and provides a novel insight about the influence of pertinent factors in this context, like color spaces, illumination and presence of glasses. A motivation for the paper is the fact that the human iris color is an essential facial trait for Caucasians, which can be employed in iris pattern recognition systems for pruning the search or in soft biometrics systems for person re-identification. Towards studying iris color as a soft biometric trait, we consider a system for automatic detection of eye color, based on standard facial images. The system entails automatic iris localization, followed by classification based on Gaussian Mixture Models with Expectation Maximization. We finally provide related detection results on the UBIRIS2 database employable in a real time eye color detection system.
Antitza Dantcheva, Nesli Erdogmus, Jean-Luc Dugelay
WACV1
2011 Bag of soft biometrics for person identification - New trends and challenges
Antitza Dantcheva, Carmelo Velardo, Angela D'Angelo, Jean-Luc Dugelay
Multim. Tools Appl.1
2010 BIOFACE: a biometric face demonstrator
abstract
In this paper, a demonstrator called BIOFACE incorporating several facial biometric techniques is described. It includes the well established Eigenfaces and the recently published Tomofaces techniques, which perform face recognition based on facial appearance and dynamics, respectively. Both techniques are based on the space dimensionality reduction and the enrollment requires the projection of several positive face samples to the reduced space. Alternatively, BIOFACE also performs face recognition based on the matching of Scale Invariant Feature Transform (SIFT) features.
Mourad Ouaret, Antitza Dantcheva, Rui Min 0002, Lionel Daniel, Jean-Luc Dugelay
ACM Multimedia2
2010 Person recognition using a bag of facial soft biometrics (BoFSB)
abstract
This work introduces the novel idea of using a bag of facial soft biometrics for person verification and identification. The novel tool inherits the non-intrusiveness and computational efficiency of soft biometrics, which allow for fast and enrolment-free biometric analysis, even in the absence of consent and cooperation of the surveillance subject. In conjunction with the proposed system design and detection algorithms, we also proceed to shed some light on the statistical properties of different parameters that are pertinent to the proposed system, as well as provide insight on general design aspects in soft-biometric systems, and different aspects regarding efficient resource allocation.
Antitza Dantcheva, Jean-Luc Dugelay, Petros Elia
MMSP1