VLDB 2026 Research / reviewers in the wild / expert
Stefanos Zafeiriou
dblp:25/1885 · also Stefanos P. Zafeiriou
· DBLP profile ↗
255ranked-venue papers
25as first author
62since 2021 · last 2026
0000-0002-5222-1740ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 199 · 14 first-author · 52 since 2021Graphics, computer vision, multimedia, augmented reality and games · 162 · 12 first-author · 40 since 2021Security and privacy · 6 · 2 first-authorDatabases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PRISM: A Unified Framework for Photorealistic Reconstruction and Intrinsic Scene Modeling
Alara Dirik, Tuanfeng Y. Wang, Duygu Ceylan, Stefanos Zafeiriou, Anna Frühstück |
ICPR (4) | 4 |
| 2025 | Arc2Avatar: Generating Expressive 3D Avatars from a Single Image via ID GuidanceabstractInspired by the effectiveness of 3D Gaussian Splatting (3DGS) in reconstructing detailed 3D scenes within multiview setups and the emergence of large 2D human foundation models, we introduce Arc2Avatar, the first Score Distillation Sampling (SDS) based method utilizing a human face foundation model as guidance with just a single image as input. To achieve that, we extend such a model for diverse-view human head generation by fine-tuning on synthetic data and modifying its conditioning. Our avatars maintain a dense correspondence with a human face mesh template, allowing blendshape-based expression generation. This is achieved through a modified 3DGS approach, connectivity regularizers, and a strategic initialization tailored for our task. Additionally, we propose an optional efficient SDSbased correction step to refine the blendshape expressions. Experiments demonstrate that Arc2Avatar achieves stateof-the-art realism and identity preservation, effectively addressing color issues by allowing the use of very low guidance, enabled by our strong identity prior and initialization strategy, without compromising detail. Code and models are available on our project page. Dimitrios Gerogiannis, Foivos Paraperas Papantoniou, Rolandos Alexandros Potamias, Alexander Lattas, Stefanos Zafeiriou |
CVPR | 5 |
| 2025 | WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wildabstractIn recent years, 3D hand pose estimation methods have garnered significant attention due to their extensive applications in human-computer interaction, virtual reality, and robotics. In contrast, there has been a notable gap in hand detection pipelines, posing significant challenges in constructing effective real-world multi-hand reconstruction systems. In this work, we present a data-driven pipeline for efficient multi-hand reconstruction in the wild. The proposed pipeline is composed of two components: a real-time fully convolutional hand localization and a high-fidelity transformer-based 3D hand reconstruction model. To tackle the limitations of previous methods and build a robust and stable detection network, we introduce a large-scale dataset with over than 2M in-the-wild hand images with diverse lighting, illumination, and occlusion conditions. Our approach outperforms previous methods in both efficiency and accuracy on popular 2D and 3D benchmarks. Finally, we showcase the effectiveness of our pipeline to achieve smooth 3D hand tracking from monocular videos, without utilizing any temporal components. Code, models, and dataset are available on our project page. Rolandos Alexandros Potamias, Jinglei Zhang 0003, Jiankang Deng, Stefanos Zafeiriou |
CVPR | 4 |
| 2025 | S^3-Face: SSS-Compliant Facial Reflectance Estimation via Diffusion PriorsabstractRecent 3D face reconstruction methods have made remarkable advancements, yet achieving high-quality facial reflectance from monocular input remains challenging. Existing methods rely on the light-stage captured data to learn facial reflectance models. However, limited subject diversity in these datasets poses challenges in achieving good generalization and broad applicability. This motivates us to explore whether the extensive priors captured in recent generative diffusion models (e.g., Stable Diffusion) can enable more generalizable facial reflectance estimation as these models have been pre-trained on large-scale internet image collections containing rich visual patterns. In this paper, we introduce the use of Stable Diffusion as a prior for facial reflectance estimation, achieving robust results with minimal captured data for fine-tuning. We present S3-Face, a comprehensive framework capable of producing SSS-compliant skin reflectance from in-the-wild images. Our method adopts a two-stage training approach: in the first stage, DSN-Net is trained to predict diffuse albedo, specular albedo, and normal maps from in-the-wild images using a novel joint reflectance attention module. In the second stage, HM-Net is trained to generate hemoglobin and melanin maps based on the diffuse albedo predicted in the first stage, yielding SSS-compliant and detailed reflectance maps. Extensive experiments demonstrate that our method achieves strong generalization and produces high-fidelity, SSS-compliant facial reflectance estimation. Xingyu Ren, Jiankang Deng, Yuhao Cheng, Wenhan Zhu, Yichao Yan, Xiaokang Yang 0001, Stefanos Zafeiriou, Chao Ma 0004 |
CVPR | 7 |
| 2025 | Dyn-HaMR: Recovering 4D Interacting Hand Motion from a Dynamic CameraabstractWe propose Dyn-HaMR, to the best of our knowledge, the first approach to reconstruct 4D global hand motion from monocular videos recorded by dynamic cameras in the wild. Reconstructing accurate 3D hand meshes from monocular videos is a crucial task for understanding human behaviour, with significant applications in augmented and virtual reality (AR/VR). However, existing methods for monocular hand reconstruction typically rely on a weak perspective camera model, which simulates hand motion within a limited camera frustum. As a result, these approaches struggle to recover the full 3D global trajectory and often produce noisy or incorrect depth estimations, particularly when the video is captured by dynamic or moving cameras, which is common in egocentric scenarios. Our DynHaMR consists of a multi-stage, multi-objective optimization pipeline, that factors in (i) simultaneous localization and mapping (SLAM) to robustly estimate relative camera motion, (ii) an interacting-hand prior for generative infilling and to refine the interaction dynamics, ensuring plausible recovery under (self-)occlusions, and (iii) hierarchical initialization through a combination of state-of-the-art hand tracking methods. Through extensive evaluations on both in-the-wild and indoor datasets, we show that our approach significantly outperforms state-of-the-art methods in terms of 4D global mesh recovery. This establishes a new benchmark for hand motion reconstruction from monocular video with moving cameras. Our project page is at https://dyn-hamr.github.io/. Zhengdi Yu, Stefanos Zafeiriou, Tolga Birdal |
CVPR | 2 |
| 2025 | Large Learning Rates Simultaneously Achieve Robustness to Spurious Correlations and Compressibility
Melih Barsbey, Lucas Prieto, Stefanos Zafeiriou, Tolga Birdal |
ICCV | 3 |
| 2025 | SpinMeRound: Consistent Multi-View Identity Generation Using Diffusion Models
Stathis Galanakis, Alexander Lattas, Stylianos Moschoglou, Bernhard Kainz, Stefanos Zafeiriou |
ICCV | 5 |
| 2025 | ImHead: A Large-Scale Implicit Morphable Model for Localized Head ModelingabstractOver the last years, 3D morphable models (3DMMs) have emerged as a state-of-the-art methodology for modeling and generating expressive 3D avatars. However, given their reliance on a strict topology, along with their linear nature, they struggle to represent complex full-head shapes. Following the advent of deep implicit functions, we propose imHead, a novel implicit 3DMM that not only models expressive 3D head avatars but also facilitates localized editing of the facial features. Previous methods directly divided the latent space into local components accompanied by an identity encoding to capture the global shape variations, leading to expensive latent sizes. In contrast, we retain a single compact identity space and introduce an intermediate region-specific latent representation to enable local edits. To train imHead, we curate a large-scale dataset of 4K distinct identities, making a step-towards large scale 3D head modeling. Under a series of experiments we demonstrate the expressive power of the proposed model to represent diverse identities and expressions outperforming previous approaches. Additionally, the proposed approach provides an interpretable solution for 3D face manipulation, allowing the user to make localized edits. Rolandos Alexandros Potamias, Stathis Galanakis, Jiankang Deng, Athanasios Papaioannou, Stefanos Zafeiriou |
ICCV | 5 |
| 2025 | Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language Generator
Ronglai Zuo, Rolandos Alexandros Potamias, Evangelos Ververas, Jiankang Deng, Stefanos Zafeiriou |
ICCV | 5 |
| 2025 | Are Large Brainwave Foundation Models Capable Yet ? Insights from Fine-TuningabstractFoundation Models have demonstrated significant success across various domains in Artificial Intelligence (AI), yet their capabilities for brainwave modeling remain unclear. In this paper, we comprehensively evaluate current Large Brainwave Foundation Models (LBMs) through systematic fine-tuning experiments across multiple Brain-Computer Interface (BCI) benchmark tasks, including memory tasks and sleep stage classification. Our extensive analysis shows that state-of-the-art LBMs achieve only marginal improvements (0.5\%) over traditional deep architectures while requiring significantly more parameters (millions vs thousands), raising important questions about their efficiency and applicability in BCI contexts. Moreover, through detailed ablation studies and Low-Rank Adaptation (LoRA), we significantly reduce trainable parameters without performance degradation, while demonstrating that architectural and training inefficiencies limit LBMs' current capabilities. Our experiments span both full model fine-tuning and parameter-efficient adaptation techniques, providing insights into optimal training strategies for BCI applications. We pioneer the application of LoRA to LBMs, revealing that performance benefits generally emerge when adapting multiple neural network components simultaneously. These findings highlight the critical need for domain-specific development strategies to advance LBMs, suggesting that current architectures may require redesign to fully leverage the potential of foundation models in brainwave analysis. Na Lee, Konstantinos Barmpas, Yannis Panagakis, Dimitrios A. Adamos, Nikolaos A. Laskaris, Stefanos Zafeiriou |
ICML | 6 |
| 2025 | FitDiff: Robust Monocular 3D Facial Shape and Reflectance Estimation using Diffusion ModelsabstractThe remarkable progress in 3D face reconstruction has resulted in high-detail and photorealistic facial representations. Recently, Diffusion Models have revolutionized the capabilities of generative methods by surpassing the performance of GANs. In this work, we present FitDiff, a diffusion-based 3D facial avatar generative model. Lever-aging diffusion principles, our model accurately generates relightable facial avatars, utilizing an identity embedding extracted from an “in-the-wild” 2D facial image. The introduced multimodal diffusion model is the first to concurrently output facial reflectance maps (diffuse and specular albedo and normals) and shapes, showcasing great generalization capabilities. It is solely trained on an annotated subset of a public facial dataset, paired with 3D reconstructions. We revisit the typical 3D facial fitting approach by guiding a reverse diffusion process using perceptual and face recognition losses. Being the first 3D LDM conditioned on face recognition embeddings, FitDiff reconstructs relightable human avatars, that can be used as-is in common rendering engines, starting only from an unconstrained facial image, and achieving state-of-the-art performance. Stathis Galanakis, Alexander Lattas, Stylianos Moschoglou, Stefanos Zafeiriou |
WACV | 4 |
| 2024 | Distribution Matching for Multi-Task Learning of Classification Tasks: A Large-Scale Study on Faces & BeyondabstractMulti-Task Learning (MTL) is a framework, where multiple related tasks are learned jointly and benefit from a shared representation space, or parameter transfer. To provide sufficient learning support, modern MTL uses annotated data with full, or sufficiently large overlap across tasks, i.e., each input sample is annotated for all, or most of the tasks. However, collecting such annotations is prohibitive in many real applications, and cannot benefit from datasets available for individual tasks. In this work, we challenge this setup and show that MTL can be successful with classification tasks with little, or non-overlapping annotations, or when there is big discrepancy in the size of labeled data per task. We explore task-relatedness for co-annotation and co-training, and propose a novel approach, where knowledge exchange is enabled between the tasks via distribution matching. To demonstrate the general applicability of our method, we conducted diverse case studies in the domains of affective computing, face recognition, species recognition, and shopping item classification using nine datasets. Our large-scale study of affective tasks for basic expression recognition and facial action unit detection illustrates that our approach is network agnostic and brings large performance improvements compared to the state-of-the-art in both tasks and across all studied databases. In all case studies, we show that co-training via task-relatedness is advantageous and prevents negative transfer (which occurs when MT model's performance is worse than that of at least one single-task model). Dimitris Kollias, Viktoriia Sharmanska, Stefanos Zafeiriou |
AAAI | 3 |
| 2024 | Neural Sign Actors: A diffusion model for 3D sign language production from textabstractSign Languages (SL) serve as the primary mode of communication for the Deaf and Hard of Hearing communities. Deep learning methods for SL recognition and translation have achieved promising results. However, Sign Language Production (SLP) poses a challenge as the generated motions must be realistic and have precise semantic meaning. Most SLP methods rely on 2D data, which hinders their re-alism. In this work, a diffusion-based SLP model is trained on a curated large-scale dataset of 4D signing avatars and their corresponding text transcripts. The proposed method can generate dynamic sequences of 3D avatars from an un-constrained domain of discourse using a diffusion process formed on a novel and anatomically informed graph neu-ral network defined on the SMPL-X body skeleton. Through quantitative and qualitative experiments, we show that the proposed method considerably outperforms previous meth-ods of SLP. This work makes an important step towards re-alistic neural sign avatars, bridging the communication gap between Deaf and hearing communities.11Project page: https://baltatzisv.github.io/neural-sign-actors/ Vasileios Baltatzis, Rolandos Alexandros Potamias, Evangelos Ververas, Guanxiong Sun, Jiankang Deng, Stefanos Zafeiriou |
CVPR | 6 |
| 2024 | Locally Adaptive Neural 3D Morphable ModelsabstractWe present the Locally Adaptive Morphable Model (LAMM), a highly flexible Auto-Encoder (AE) framework for learning to generate and manipulate 3D meshes. We train our architecture following a simple self-supervised training scheme in which input displacements over a set of sparse control vertices are used to overwrite the encoded geometry in order to transform one training sample into another. During inference, our model produces a dense output that adheres locally to the specified sparse geom-etry while maintaining the overall appearance of the en-coded object. This approach results in state-of-the-art per-formance in both disentangling manipulated geometry and 3D mesh reconstruction. To the best of our knowledge LAMM is the first end-to-end framework that enables direct local control of 3D vertex geometry in a single forward pass. A very efficient computational graph allows our net-work to train with only afraction of the memory required by previous methods and run faster during inference, generating 12k vertex meshes at >60fps on a single CPU thread. We further leverage local geometry control as a primitive for higher level editing operations and present a set of derivative capabilities such as swapping and sampling object parts. Code and pretrained models can be found at https://github.com/michaeltrs/LAMM. Michail Tarasiou, Rolandos Alexandros Potamias, Eimear O' Sullivan, Stylianos Ploumpis, Stefanos Zafeiriou |
CVPR | 5 |
| 2024 | Design2Cloth: 3D Cloth Generation from 2D MasksabstractIn recent years, there has been a significant shift in the field of digital avatar research, towards modeling, animating and reconstructing clothed human representations, as a key step towards creating realistic avatars. However, current 3D cloth generation methods are garment specific or trained completely on synthetic data, hence lacking fine details and realism. In this work, we make a step to-wards automatic realistic garment design and propose De-sign2Cloth, a high fidelity 3D generative model trained on a real world dataset from more than 2000 subject scans. To provide vital contribution to the fashion industry, we developed a user-friendly adversarial model capable of gener-ating diverse and detailed clothes simply by drawing a 2D cloth mask. Under a series of both qualitative and quantitative experiments, we showcase that Design2Cloth outper-forms current state-of-the-art cloth generative models by a large margin. In addition to the generative properties of our network, we showcase that the proposed method can be used to achieve high quality reconstructions from single in-the-wild images and 3D scans. Dataset, code and pre-trained model will become publicly available.11Project Page: https://jiali-zheng.github.io/Design2Cloth/ Rolandos Alexandros Potamias, Stefanos Zafeiriou |
CVPR | 3 |
| 2024 | AnimateMe: 4D Facial Expressions via Diffusion Models
Dimitrios Gerogiannis, Foivos Paraperas Papantoniou, Rolandos Alexandros Potamias, Alexander Lattas, Stylianos Moschoglou, Stylianos Ploumpis, Stefanos Zafeiriou |
ECCV (77) | 7 |
| 2024 | Arc2Face: A Foundation Model for ID-Consistent Human Faces
Foivos Paraperas Papantoniou, Alexander Lattas, Stylianos Moschoglou, Jiankang Deng, Bernhard Kainz, Stefanos Zafeiriou |
ECCV (37) | 6 |
| 2024 | ShapeFusion: A 3D Diffusion Model for Localized Shape Editing
Rolandos Alexandros Potamias, Michail Tarasiou, Stylianos Ploumpis, Stefanos Zafeiriou |
ECCV (14) | 4 |
| 2024 | 3DGazeNet: Generalizing 3D Gaze Estimation with Weak-Supervision from Synthetic Views
Evangelos Ververas, Polydefkis Gkagkos, Jiankang Deng, Michail C. Doukas, Jia Guo 0003, Stefanos Zafeiriou |
ECCV (21) | 6 |
| 2024 | SAGS: Structure-Aware 3D Gaussian Splatting
Evangelos Ververas, Rolandos Alexandros Potamias, Jifei Song, Jiankang Deng, Stefanos Zafeiriou |
ECCV (19) | 5 |
| 2024 | ID-to-3D: Expressive ID-guided 3D Heads via Score Distillation SamplingabstractWe propose ID-to-3D, a method to generate identity- and text-guided 3D human heads with disentangled expressions, starting from even a single casually captured ‘in-the-wild’ image of a subject. The foundation of our approach is anchored in compositionality, alongside the use of task-specific 2D diffusion models as priors for optimization. First, we extend a foundational model with a lightweight expression-aware and ID-aware architecture, and create 2D priors for geometric and texture generation, via fine-tuning only 0.2% of its available training parameters. Then, we jointly leverage a neural parametric representation for the expression of each subject and a multi-stage generation of highly detailed geometry and albedo texture. This combination of strong face identity embeddings and our neural representation enables accurate reconstruction of not only facial features but also accessories and hair, and can be meshed to provide render-ready assets for gaming and telepresence. Our results achieve an unprecedented level of id-consistent and high-quality texture and geometry generation, generalizing to a ‘world’ of unseen 3D identities, without relying on large 3D captured datasets of human assets. Francesca Babiloni, Alexander Lattas, Jiankang Deng, Stefanos Zafeiriou |
NeurIPS | 4 |
| 2024 | UV-free Texture Generation with Denoising and Geodesic Heat DiffusionabstractSeams, distortions, wasted UV space, vertex-duplication, and varying resolution over the surface are the most prominent issues of the standard UV-based texturing of meshes. These issues are particularly acute when automatic UV-unwrapping techniques are used. For this reason, instead of generating textures in automatically generated UV-planes like most state-of-the-art methods, we propose to represent textures as coloured point-clouds whose colours are generated by a denoising diffusion probabilistic model constrained to operate on the surface of 3D objects. Our sampling and resolution agnostic generative model heavily relies on heat diffusion over the surface of the meshes for spatial communication between points. To enable processing of arbitrarily sampled point-cloud textures and ensure long-distance texture consistency we introduce a fast re-sampling of the mesh spectral properties used during the heat diffusion and introduce a novel heat-diffusion-based self-attention mechanism. Our code and pre-trained models are available at github.com/simofoti/UV3-TeD. Simone Foti, Stefanos Zafeiriou, Tolga Birdal |
NeurIPS | 2 |
| 2023 | FitMe: Deep Photorealistic 3D Morphable Model AvatarsabstractIn this paper, we introduce FitMe, a facial reflectance model and a differentiable rendering optimization pipeline, that can be used to acquire high-fidelity renderable human avatars from single or multiple images. The model consists of a multi-modal style-based generator, that captures facial appearance in terms of diffuse and specular reflectance, and a PCA-based shape model. We employ a fast differentiable rendering process that can be used in an optimization pipeline, while also achieving photorealistic facial shading. Our optimization process accurately captures both the facial reflectance and shape in high-detail, by exploiting the expressivity of the style-based latent representation and of our shape model. FitMe achieves state-of-the-art reflectance acquisition and identity preservation on single “in-the-wild” facial images, while it produces impressive scan-like results, when given multiple unconstrained facial images pertaining to the same identity. In contrast with recent implicit avatar reconstructions, FitMe requires only one minute and produces relightable mesh and texture-based avatars, that can be used by end-user applications. Alexander Lattas, Stylianos Moschoglou, Stylianos Ploumpis, Baris Gecer, Jiankang Deng, Stefanos Zafeiriou |
CVPR | 6 |
| 2023 | Handy: Towards a High Fidelity 3D Hand Shape and Appearance ModelabstractOver the last few years, with the advent of virtual and augmented reality, an enormous amount of research has been focused on modeling, tracking and reconstructing human hands. Given their power to express human behavior, hands have been a very important, but challenging component of the human body. Currently, most of the state-of-the-art reconstruction and pose estimation methods rely on the low polygon MANO model. Apart from its low polygon count, MANO model was trained with only 31 adult subjects, which not only limits its expressive power but also imposes unnecessary shape reconstruction constraints on pose estimation methods. Moreover, hand appearance remains almost unexplored and neglected from the majority of hand reconstruction methods. In this work, we propose “Handy”, a large-scale model of the human hand, modeling both shape and appearance composed of over 1200 subjects which we make publicly available for the benefit of the research community. In contrast to current models, our proposed hand model was trained on a dataset with large diversity in age, gender, and ethnicity, which tackles the limitations of MANO and accurately reconstructs out-of-distribution samples. In order to create a high quality texture model, we trained a powerful GAN, which preserves high frequency details and is able to generate high resolution hand textures. To showcase the capabilities of the proposed model, we built a synthetic dataset of textured hands and trained a hand pose estimation network to reconstruct both the shape and appearance from single images. As it is demonstrated in an extensive series of quantitative as well as qualitative experiments, our model proves to be robust against the state-of-the-art and realistically captures the 3D hand shape and pose along with a high frequency detailed texture even in adverse “in-the-wild” conditions. Rolandos Alexandros Potamias, Stylianos Ploumpis, Stylianos Moschoglou, Vasileios Triantafyllou, Stefanos Zafeiriou |
CVPR | 5 |
| 2023 | ViTs for SITS: Vision Transformers for Satellite Image Time SeriesabstractIn this paper we introduce the Temporo-Spatial Vision Transformer (TSViT), a fully-attentional model for general Satellite Image Time Series (SITS) processing based on the Vision Transformer (ViT). TSViT splits a SITS record into non-overlapping patches in space and time which are tokenized and subsequently processed by a factorized temporo-spatial encoder. We argue, that in contrast to natural images, a temporal-then-spatial factorization is more intuitive for SITS processing and present experimental evidence for this claim. Additionally, we enhance the model's discriminative power by introducing two novel mechanisms for acquisition-time-specific temporal positional encodings and multiple learnable class tokens. The effect of all novel design choices is evaluated through an extensive ablation study. Our proposed architecture achieves state-of-the-art performance, surpassing previous approaches by a significant margin in three publicly available SITS semantic segmentation and classification datasets. All model, training and evaluation codes can be found at https://github.com/michaeltrs/DeepSatModels. Michail Tarasiou, Erik Chavez, Stefanos Zafeiriou |
CVPR | 3 |
| 2023 | Adaptive Spiral Layers for Efficient 3D Representation Learning on MeshesabstractThe success of deep learning models on structured data has generated significant interest in extending their application to non-Euclidean domains. In this work, we introduce a novel intrinsic operator suitable for representation learning on 3D meshes. Our operator is specifically tailored to adapt its behavior to the irregular structure of the underlying graph and effectively utilize its long-range dependencies, while at the same time ensuring computational efficiency and ease of optimization. In particular, inspired by the framework of Spiral Convolution, which extracts and transforms the vertices in the 3D mesh following a local spiral ordering, we propose a general operator that dynamically adjusts the length of the spiral trajectory and the parameters of the transformation for each processed vertex and mesh. Then, we use polyadic decomposition to factorize its dense weight tensor into a sequence of lighter linear layers that separately process features and vertices information, hence significantly reducing the computational complexity without introducing any stringent inductive biases. Notably, we leverage dynamic gating to achieve spatial adaptivity and induce global reasoning with constant time complexity benefitting from an efficient dynamic pooling mechanism based on Summed-Area-tables. Used as a drop-in replacement on existing architectures for shape correspondence our operator significantly improves the performance-efficiency trade-off, and in 3D shape generation with morphable models achieves state-of-the-art performance with a three-fold reduction in the number of parameters required. Project page: https://github.com/Fb2221/DFC Francesca Babiloni, Matteo Maggioni, Thomas Tanay, Jiankang Deng, Ales Leonardis, Stefanos Zafeiriou |
ICCV | 6 |
| 2023 | Relightify: Relightable 3D Faces from a Single Image via Diffusion ModelsabstractFollowing the remarkable success of diffusion models on image generation, recent works have also demonstrated their impressive ability to address a number of inverse problems in an unsupervised way, by properly constraining the sampling process based on a conditioning input. Motivated by this, in this paper, we present the first approach to use diffusion models as a prior for highly accurate 3D facial BRDF reconstruction from a single image. We start by leveraging a high-quality UV dataset of facial reflectance (diffuse and specular albedo and normals), which we render under varying illumination settings to simulate natural RGB textures and, then, train an unconditional diffusion model on concatenated pairs of rendered textures and reflectance components. At test time, we fit a 3D morphable model to the given image and unwrap the face in a partial UV texture. By sampling from the diffusion model, while retaining the observed texture part intact, the model inpaints not only the self-occluded areas but also the unknown reflectance components, in a single sequence of denoising steps. In contrast to existing methods, we directly acquire the observed texture from the input image, thus, resulting in more faithful and consistent reflectance estimation. Through a series of qualitative and quantitative comparisons, we demonstrate superior performance in both texture completion as well as reflectance reconstruction tasks. Foivos Paraperas Papantoniou, Alexander Lattas, Stylianos Moschoglou, Stefanos Zafeiriou |
ICCV | 4 |
| 2023 | Spatio-temporal Prompting Network for Robust Video Feature ExtractionabstractFrame quality deterioration is one of the main challenges in the field of video understanding. To compensate for the information loss caused by deteriorated frames, recent approaches exploit transformer-based integration modules to obtain spatio-temporal information. However, these integration modules are heavy and complex. Furthermore, each integration module is specifically tailored for its target task, making it difficult to generalise to multiple tasks. In this paper, we present a neat and unified framework, called Spatio-Temporal Prompting Network (STPN). It can efficiently extract robust and accurate video features by dynamically adjusting the input features in the backbone network. Specifically, STPN predicts several video prompts containing spatio-temporal information of neighbour frames. Then, these video prompts are prepended to the patch embeddings of the current frame as the updated input for video feature extraction. Moreover, STPN is easy to generalise to various video tasks because it does not contain task-specific modules. Without bells and whistles, STPN achieves state-of-the-art performance on three widely-used datasets for different video understanding tasks, i.e., ImageNetVID for video object detection, YouTubeVIS for video instance segmentation, and GOT-10k for visual object tracking. Codes are available at https://github.com/guanxiongsun/STPN Guanxiong Sun, Zhaoyu Zhang 0001, Jiankang Deng, Stefanos Zafeiriou, Yang Hua 0001 |
ICCV | 5 |
| 2023 | Dynamic Neural PortraitsabstractWe present Dynamic Neural Portraits, a novel approach to the problem of full-head reenactment. Our method generates photo-realistic video portraits by explicitly controlling head pose, facial expressions and eye gaze. Our proposed architecture is different from existing methods that rely on GAN-based image-to-image translation networks for transforming renderings of 3D faces into photo-realistic images. Instead, we build our system upon a 2D coordinate-based MLP with controllable dynamics. Our intuition to adopt a 2D-based representation, as opposed to recent 3D NeRF-like systems, stems from the fact that video portraits are captured by monocular stationary cameras, therefore, only a single viewpoint of the scene is available. Primarily, we condition our generative model on expression blendshapes, nonetheless, we show that our system can be successfully driven by audio features as well. Our experiments demonstrate that the proposed method is 270 times faster than recent NeRF-based reenactment methods, with our networks achieving speeds of 24 fps for resolutions up to 1024×1024, while outperforming prior works in terms of visual quality. Michail C. Doukas, Stylianos Ploumpis, Stefanos Zafeiriou |
WACV | 3 |
| 2023 | 3DMM-RF: Convolutional Radiance Fields for 3D Face ModelingabstractFacial 3D Morphable Models are a main computer vision subject with countless applications and have been highly optimized in the last two decades. The tremendous improvements of deep generative networks have created various possibilities for improving such models and have attracted wide interest. Moreover, the recent advances in neural radiance fields, are revolutionising novel-view synthesis of known scenes. In this work, we present a facial 3D Morphable Model, which exploits both of the above, and can accurately model a subject’s identity, pose and expression and render it in arbitrary illumination. This is achieved by utilizing a powerful deep style-based generator to overcome two main weaknesses of neural radiance fields, their rigidity and rendering speed. We introduce a style-based generative network that synthesizes in one pass all and only the required rendering samples of a neural radiance field. We create a vast labelled synthetic dataset of facial renders, and train the network, so that it can accurately model and generalize on facial identity, pose and appearance. Finally, we show that this model can accurately be fit to "in-the-wild" facial images of arbitrary pose and illumination, extract the facial characteristics, and be used to re-render the face in controllable conditions. Stathis Galanakis, Baris Gecer, Alexander Lattas, Stefanos Zafeiriou |
WACV | 4 |
| 2023 | Linear Complexity Self-Attention With $3{\mathrm{rd}}$3 rd Order PolynomialsabstractSelf-attention mechanisms and non-local blocks have become crucial building blocks for state-of-the-art neural architectures thanks to their unparalleled ability in capturing long-range dependencies in the input. However their cost is quadratic with the number of spatial positions hence making their use impractical in many real case applications. In this work, we analyze these methods through a polynomial lens, and we show that self-attention can be seen as a special case of a 3 rd order polynomial. Within this polynomial framework, we are able to design polynomial operators capable of accessing the same data pattern of non-local and self-attention blocks while reducing the complexity from quadratic to linear. As a result, we propose two modules (Poly-NL and Poly-SA) that can be used as "drop-in" replacements for more-complex non-local and self-attention layers in state-of-the-art CNNs and ViT architectures. Our modules can achieve comparable, if not better, performance across a wide range of computer vision tasks while keeping a complexity equivalent to a standard linear layer. Francesca Babiloni, Ioannis Marras, Jiankang Deng, Filippos Kokkinos, Matteo Maggioni, Grigorios Chrysos 0002, Philip Torr 0001, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Improving Graph Neural Network Expressivity via Subgraph Isomorphism CountingabstractWhile Graph Neural Networks (GNNs) have achieved remarkable results in a variety of applications, recent studies exposed important shortcomings in their ability to capture the structure of the underlying graph. It has been shown that the expressive power of standard GNNs is bounded by the Weisfeiler-Leman (WL) graph isomorphism test, from which they inherit proven limitations such as the inability to detect and count graph substructures. On the other hand, there is significant empirical evidence, e.g. in network science and bioinformatics, that substructures are often intimately related to downstream tasks. To this end, we propose "Graph Substructure Networks" (GSN), a topologically-aware message passing scheme based on substructure encoding. We theoretically analyse the expressive power of our architecture, showing that it is strictly more expressive than the WL test, and provide sufficient conditions for universality. Importantly, we do not attempt to adhere to the WL hierarchy; this allows us to retain multiple attractive properties of standard GNNs such as locality and linear network complexity, while being able to disambiguate even hard instances of graph isomorphism. We perform an extensive experimental evaluation on graph classification and regression tasks and obtain state-of-the-art results in diverse real-world settings including molecular graphs and social networks. Giorgos Bouritsas, Fabrizio Frasca, Stefanos Zafeiriou, Michael M. Bronstein |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Free-HeadGAN: Neural Talking Head Synthesis With Explicit Gaze ControlabstractWe present Free-HeadGAN, a person-generic neural talking head synthesis system. We show that modeling faces with sparse 3D facial landmarks is sufficient for achieving state-of-the-art generative performance, without relying on strong statistical priors of the face, such as 3D Morphable Models. Apart from 3D pose and facial expressions, our method is capable of fully transferring the eye gaze, from a driving actor to a source identity. Our complete pipeline consists of three components: a canonical 3D key-point estimator that regresses 3D pose and expression-related deformations, a gaze estimation network and a generator that is built upon the architecture of HeadGAN. We further experiment with an extension of our generator to accommodate few-shot learning using an attention mechanism, in case multiple source images are available. Compared to recent methods for reenactment and motion transfer, our system achieves higher photo-realism combined with superior identity preservation, while offering explicit gaze control. Michail C. Doukas, Evangelos Ververas, Viktoriia Sharmanska, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | EDFace-Celeb-1 M: Benchmarking Face Hallucination With a Million-Scale DatasetabstractRecent deep face hallucination methods show stunning performance in super-resolving severely degraded facial images, even surpassing human ability. However, these algorithms are mainly evaluated on non-public synthetic datasets. It is thus unclear how these algorithms perform on public face hallucination datasets. Meanwhile, most of the existing datasets do not well consider the distribution of races, which makes face hallucination methods trained on these datasets biased toward some specific races. To address the above two problems, in this paper, we build a public Ethnically Diverse Face dataset, EDFace-Celeb-1 M, and design a benchmark task for face hallucination. Our dataset includes 1.7 million photos that cover different countries, with relatively balanced race composition. To the best of our knowledge, it is the largest-scale and publicly available face hallucination dataset in the wild. Associated with this dataset, this paper also contributes various evaluation protocols and provides comprehensive analysis to benchmark the existing state-of-the-art methods. The benchmark evaluations demonstrate the performance and limitations of state-of-the-art algorithms. https://github.com/HDCVLab/EDFace-Celeb-1M. Kaihao Zhang, Dongxu Li 0003, Wenhan Luo, Jingyu Liu 0004, Jiankang Deng, Wei Liu 0005, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | 4DME: A Spontaneous 4D Micro-Expression Dataset With MultimodalitiesabstractMicro-expressions (ME) are a special form of facial expressions which may occur when people try to hide their true feelings for some reasons. MEs are important clues to reveal people’s true feelings, but are difficult or impossible to be captured by ordinary persons with naked-eyes as they are very short and subtle. It is expected that robust computer vision methods can be developed to automatically analyze MEs which requires lots of ME data. The current ME datasets are insufficient, and mostly contain only one single form of 2D color videos. Researches on 4D data of ordinary facial expressions have prospered, but so far no 4D data is available in ME study. In the current study, we introduce the 4DME dataset: a new spontaneous ME dataset which includes 4D data along with three other video modalities. Both micro- and macro-expression clips are labeled out in 4DME, and 22 AU labels and five categories of emotion labels are annotated. Experiments are carried out using three 2D-based methods and one 4D-based method to provide baseline results. The results indicate that the 4D data can potentially benefit ME recognition. The 4DME dataset could be used for developing 4D-based approaches, or exploring fusion of multiple video sources (e.g., texture and depth) for the task of ME analysis in future. Besides, we also emphasize the importance of forming a clear and unified criteria of ME annotation for future ME data collection studies. Several key questions related with ME annotation are listed and discussed in depth, especially about the relationship between AUs and ME emotion categories. A preliminary AU-Emo mapping table is proposed with justified explanations and supportive experimental results. Several unsolved issues are also summarized for future work. Shiyang Cheng 0001, Yante Li, Muzammil Behzad, Jie Shen 0008, Stefanos Zafeiriou, Maja Pantic, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 6 |
| 2023 | Inverse Image Frequency for Long-Tailed Image RecognitionabstractThe long-tailed distribution is a common phenomenon in the real world. Extracted large scale image datasets inevitably demonstrate the long-tailed property and models trained with imbalanced data can obtain high performance for the over-represented categories, but struggle for the under-represented categories, leading to biased predictions and performance degradation. To address this challenge, we propose a novel de-biasing method named Inverse Image Frequency (IIF). IIF is a multiplicative margin adjustment transformation of the logits in the classification layer of a convolutional neural network. Our method achieves stronger performance than similar works and it is especially useful for downstream tasks such as long-tailed instance segmentation as it produces fewer false positive detections. Our extensive experiments show that IIF surpasses the state of the art on many long-tailed benchmarks such as ImageNet-LT, CIFAR-LT, Places-LT and LVIS, reaching 55.8% top-1 accuracy with ResNet50 on ImageNet-LT and 26.3% segmentation AP with MaskRCNN ResNet50 on LVIS. Code available at https://github.com/kostas1515/iif. Konstantinos Panagiotis Alexandridis, Shan Luo 0001, Anh Nguyen 0003, Jiankang Deng, Stefanos Zafeiriou |
IEEE Trans. Image Process. | 5 |
| 2023 | Redesigning Multi-Scale Neural Network for Crowd CountingabstractPerspective distortions and crowd variations make crowd counting a challenging task in computer vision. To tackle it, many previous works have used multi-scale architecture in deep neural networks (DNNs). Multi-scale branches can be either directly merged (e.g. by concatenation) or merged through the guidance of proxies (e.g. attentions) in the DNNs. Despite their prevalence, these combination methods are not sophisticated enough to deal with the per-pixel performance discrepancy over multi-scale density maps. In this work, we redesign the multi-scale neural network by introducing a hierarchical mixture of density experts, which hierarchically merges multi-scale density maps for crowd counting. Within the hierarchical structure, an expert competition and collaboration scheme is presented to encourage contributions from all scales; pixel-wise soft gating nets are introduced to provide pixel-wise soft weights for scale combinations in different hierarchies. The network is optimized using both the crowd density map and the local counting map, where the latter is obtained by local integration on the former. Optimizing both can be problematic because of their potential conflicts. We introduce a new relative local counting loss based on relative count differences among hard-predicted local regions in an image, which proves to be complementary to the conventional absolute error loss on the density map. Experiments show that our method achieves the state-of-the-art performance on five public datasets, i.e. ShanghaiTech, UCF_CC_50, JHU-CROWD++, NWPU-Crowd and Trancos. Our codes will be available at https://github.com/ZPDu/Redesigning-Multi-Scale-Neural-Network-for-Crowd-Counting. Zhipeng Du, Miaojing Shi, Jiankang Deng, Stefanos Zafeiriou |
IEEE Trans. Image Process. | 4 |
| 2022 | 3D human tongue reconstruction from single "in-the-wild" imagesabstract3D face reconstruction from a single image is a task that has garnered increased interest in the Computer Vision community, especially due to its broad use in a number of applications such as realistic 3D avatar creation, pose invariant face recognition and face hallucination. Since the introduction of the 3D Morphable Model in the late 90's, we witnessed an explosion of research aiming at particularly tackling this task. Nevertheless, despite the increasing level of detail in the 3D face reconstructions from single images mainly attributed to deep learning advances, finer and highly deformable components of the face such as the tongue are still absent from all 3D face models in the literature, although being very important for the realness of the 3D avatar representations. In this work we present the first, to the best of our knowledge, end-to-end trainable pipeline that accurately reconstructs the 3D face together with the tongue. Moreover, we make this pipeline robust in “in-the-wild” images by introducing a novel GAN method tailored for 3D tongue surface generation. Finally, we make publicly available to the community the first diverse tongue dataset, consisting of 1,800 raw scans of 700 individuals varying in gender, age, and ethnicity backgrounds**Project url: www.github.com/steliosploumpis/tongue. As we demonstrate in an extensive series of quantitative as well as qualitative experiments, our model proves to be robust and realistically captures the 3D tongue structure, even in adverse “in-the- wild” conditions. Stylianos Ploumpis, Stylianos Moschoglou, Vasileios Triantafyllou, Stefanos Zafeiriou |
CVPR | 4 |
| 2022 | Neural Mesh SimplificationabstractDespite the advent in rendering, editing and preprocessing methods of 3D meshes, their real-time execution remains still infeasible for large-scale meshes. To ease and accelerate such processes, mesh simplification methods have been introduced with the aim to reduce the mesh resolution while preserving its appearance. In this work we attempt to tackle the novel task of learnable and differentiable mesh simplification. Compared to traditional simplification approaches that collapse edges in a greedy iterative manner, we propose a fast and scalable method that simplifies a given mesh in one-pass. The proposed method unfolds in three steps. Initially, a subset of the input vertices is sampled using a sophisticated extension of random sampling. Then, we train a sparse attention network to propose candidate triangles based on the edge connectivity of the sampled vertices. Finally, a classification network estimates the probability that a candidate triangle will be included in the final mesh. The fast, lightweight and differentiable properties of the proposed method makes it possible to be plugged in every learnable pipeline without introducing a significant overhead. We evaluate both the sampled vertices and the generated triangles under several appearance error measures and compare its performance against several state-of-the-art baselines. Furthermore, we showcase that the running performance can be up to 10× faster than traditional methods. Rolandos Alexandros Potamias, Stylianos Ploumpis, Stefanos Zafeiriou |
CVPR | 3 |
| 2022 | Decoupled Multi-task Learning with Cyclical Self-Regulation for Face ParsingabstractThis paper probes intrinsic factors behind typical failure cases (e.g. spatial inconsistency and boundary confusion) produced by the existing state-of-the-art method in face parsing. To tackle these problems, we propose a novel Decoupled Multi-task Learning with Cyclical Self-Regulation (DML-CSR) for face parsing. Specifically, DML-CSR designs a multi-task model which comprises face parsing, binary edge, and category edge detection. These tasks only share low-level encoder weights without high-level interactions between each other, enabling to decouple auxiliary modules from the whole network at the inference stage. To address spatial inconsistency, we develop a dynamic dual graph convolutional network to capture global contextual information without using any extra pooling operation. To handle boundary confusion in both single and multiple face scenarios, we exploit binary and category edge detection to jointly obtain generic geometric structure and fine-grained semantic clues of human faces. Besides, to prevent noisy labels from degrading model generalization during training, cyclical self-regulation is proposed to self-ensemble several model instances to get a new model and the resulting model then is used to self-distill subsequent models, through alternating iterations. Experiments show that our method achieves the new state-of-the-art performance on the Helen, CelebAMask-HQ, and Lapa datasets. The source code is available at https://github.com/deepinsight/insightface/tree/master/parsing/dml_csr. Qingping Zheng, Jiankang Deng, Ying Li 0017, Stefanos Zafeiriou |
CVPR | 5 |
| 2022 | MimicME: A Large Scale Diverse 4D Database for Facial Expression Analysis
Athanasios Papaioannou, Baris Gecer, Shiyang Cheng 0001, Grigorios Chrysos 0002, Jiankang Deng, Eftychia Fotiadou, Christos Kampouris, Dimitris Kollias, Stylianos Moschoglou, Kritaphat Songsri-in, Stylianos Ploumpis, George Trigeorgis, Panagiotis Tzirakis, Evangelos Ververas, Allan Ponniah, Anastasios Roussos, Stefanos Zafeiriou |
ECCV (8) | 18 |
| 2022 | Revisiting Point Cloud Simplification: A Learnable Feature Preserving Approach
Rolandos Alexandros Potamias, Giorgos Bouritsas, Stefanos Zafeiriou |
ECCV (2) | 3 |
| 2022 | Sample and Computation Redistribution for Efficient Face Detection
Jia Guo 0003, Jiankang Deng, Alexander Lattas, Stefanos Zafeiriou |
ICLR | 4 |
| 2022 | Deepsatdata: Building Large Scale Datasets of Satellite Images for Training Machine Learning ModelsabstractThis paper presents DeepSatData a free and open source pipeline for automatically generating satellite imagery datasets for training machine learning models. The implementation presented can be used to query, download and process freely available Sentinel-2 data for the generation of large scale datasets required for training deep neural networks (DNN). We discuss design considerations faced from the point of view of DNN training and evaluation such as checking the quality of ground truth data and assessing the scalability of the approach. Accompanying code is made publicly available in https://github.com/michaeltrs/DeepSatData. Michail Tarasiou, Stefanos Zafeiriou |
IGARSS | 2 |
| 2022 | Physically-Based Face Rendering for NIR-VIS Face RecognitionabstractNear infrared (NIR) to Visible (VIS) face matching is challenging due to the significant domain gaps as well as a lack of sufficient data for cross-modality model training. To overcome this problem, we propose a novel method for paired NIR-VIS facial image generation. Specifically, we reconstruct 3D face shape and reflectance from a large 2D facial dataset and introduce a novel method of transforming the VIS reflectance to NIR reflectance. We then use a physically-based renderer to generate a vast, high-resolution and photorealistic dataset consisting of various poses and identities in the NIR and VIS spectra. Moreover, to facilitate the identity feature learning, we propose an IDentity-based Maximum Mean Discrepancy (ID-MMD) loss, which not only reduces the modality gap between NIR and VIS images at the domain level but encourages the network to focus on the identity features instead of facial details, such as poses and accessories. Extensive experiments conducted on four challenging NIR-VIS face recognition benchmarks demonstrate that the proposed method can achieve comparable performance with the state-of-the-art (SOTA) methods without requiring any existing NIR-VIS face recognition datasets. With slightly fine-tuning on the target NIR-VIS face recognition datasets, our method can significantly surpass the SOTA performance. Code and pretrained models are released under the insightface GitHub. Yunqi Miao, Alexander Lattas, Jiankang Deng, Jungong Han, Stefanos Zafeiriou |
NeurIPS | 5 |
| 2022 | Deep Polynomial Neural NetworksabstractDeep convolutional neural networks (DCNNs) are currently the method of choice both for generative, as well as for discriminative learning in computer vision and machine learning. The success of DCNNs can be attributed to the careful selection of their building blocks (e.g., residual blocks, rectifiers, sophisticated normalization schemes, to mention but a few). In this paper, we propose Π-Nets, a new class of function approximators based on polynomial expansions. Π-Nets are polynomial neural networks, i.e., the output is a high-order polynomial of the input. The unknown parameters, which are naturally represented by high-order tensors, are estimated through a collective tensor factorization with factors sharing. We introduce three tensor decompositions that significantly reduce the number of parameters and show how they can be efficiently implemented by hierarchical neural networks. We empirically demonstrate that Π-Nets are very expressive and they even produce good results without the use of non-linear activation functions in a large battery of tasks and signals, i.e., images, graphs, and audio. When used in conjunction with activation functions, Π-Nets produce state-of-the-art results in three challenging tasks, i.e., image generation, face verification and 3D mesh representation learning. The source code is available at https://github.com/grigorisg9gr/polynomial_nets. Grigorios Chrysos 0002, Stylianos Moschoglou, Giorgos Bouritsas, Jiankang Deng, Yannis Panagakis, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | ArcFace: Additive Angular Margin Loss for Deep Face Recognition
Jiankang Deng, Jia Guo 0003, Jing Yang 0038, Niannan Xue, Irene Kotsia, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Fast-GANFIT: Generative Adversarial Network for High Fidelity 3D Face ReconstructionabstractA lot of work has been done towards reconstructing the 3D facial structure from single images by capitalizing on the power of deep convolutional neural networks (DCNNs). In the recent works, the texture features either correspond to components of a linear texture space or are learned by auto-encoders directly from in-the-wild images. In all cases, the quality of the facial texture reconstruction is still not capable of modeling facial texture with high-frequency details. In this paper, we take a radically different approach and harness the power of generative adversarial networks (GANs) and DCNNs in order to reconstruct the facial texture and shape from single images. That is, we utilize GANs to train a very powerful facial texture prior from a large-scale 3D texture dataset. Then, we revisit the original 3D Morphable Models (3DMMs) fitting making use of non-linear optimization to find the optimal latent parameters that best reconstruct the test image but under a new perspective. In order to be robust towards initialisation and expedite the fitting process, we propose a novel self-supervised regression based approach. We demonstrate excellent results in photorealistic and identity preserving 3D face reconstructions and achieve for the first time, to the best of our knowledge, facial texture reconstruction with high-frequency details. Baris Gecer, Stylianos Ploumpis, Irene Kotsia, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | AvatarMe++: Facial Shape and BRDF Inference With Photorealistic Rendering-Aware GANsabstractOver the last years, with the advent of Generative Adversarial Networks (GANs), many face analysis tasks have accomplished astounding performance, with applications including, but not limited to, face generation and 3D face reconstruction from a single "in-the-wild" image. Nevertheless, to the best of our knowledge, there is no method which can produce render-ready high-resolution 3D faces from "in-the-wild" images and this can be attributed to the: (a) scarcity of available data for training, and (b) lack of robust methodologies that can successfully be applied on very high-resolution data. In this paper, we introduce the first method that is able to reconstruct photorealistic render-ready 3D facial geometry and BRDF from a single "in-the-wild" image. To achieve this, we capture a large dataset of facial shape and reflectance, which we have made public. Moreover, we define a fast and photorealistic differentiable rendering methodology with accurate facial skin diffuse and specular reflection, self-occlusion and subsurface scattering approximation. With this, we train a network that disentangles the facial diffuse and specular reflectance components from a mesh and texture with baked illumination, scanned or reconstructed with a 3DMM fitting method. As we demonstrate in a series of qualitative and quantitative experiments, our method outperforms the existing arts by a significant margin and reconstructs authentic, 4K by 6K-resolution 3D faces from a single low-resolution image, that are ready to be rendered in various applications and bridge the uncanny valley. Alexander Lattas, Stylianos Moschoglou, Stylianos Ploumpis, Baris Gecer, Abhijeet Ghosh, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Guest Editorial: Non-Euclidean Machine LearningabstractOver the past decade, deep learning has had a revolutionary impact on a broad range of fields such as computer vision and image processing, computational photography, medical imaging and speech and language analysis and synthesis etc. Deep learning technologies are estimated to have added billions in business value, created new markets, and transformed entire industrial segments. Most of today’s successful deep learning methods such as Convolutional Neural Networks (CNNs) rely on classical signal processing models that limit their applicability to data with underlying Euclidean grid-like structure, e.g., images or acoustic signals. Yet, many applications deal with non-Euclidean (graph- or manifold-structured) data. For example, in social network analysis the users and their attributes are generally modeled as signals on the vertices of graphs. In biology protein-to-protein interactions are modeled as graphs. In computer vision & graphics 3D objects are modeled as meshes or point clouds. Furthermore, a graph representation is a very natural way to describe interactions between objects or signals. The classical deep learning paradigm on Euclidean domains falls short in providing appropriate tools for such kind of data. Until recently, the lack of deep learning models capable of correctly dealing with non-Euclidean data has been a major obstacle in these fields. This special section addresses the need to bring together leading efforts in non-Euclidean deep learning across all communities. From the papers that the special received twelve were selected for publication. The selected papers can naturally fall in three distinct categories: (a) methodologies that advance machine learning on data that are represented as graphs, (b) methodologies that advance machine learning on manifold-valued data, and (c) applications of machine learning methodologies on non-Euclidean spaces in computer vision and medical imaging. We briefly review the accepted papers in each of the groups. Stefanos Zafeiriou, Michael M. Bronstein, Taco Cohen, Oriol Vinyals, Jure Leskovec, Pietro Liò, Joan Bruna, Marco Gori |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Context-Self Contrastive Pretraining for Crop Type Semantic SegmentationabstractIn this paper, we propose a fully supervised pre-training scheme based on contrastive learning particularly tailored to dense classification tasks. The proposedContext-Self Contrastive Loss(CSCL) learns an embedding space that makes semantic boundaries pop-up by use of a similarity metric between every location in a training sample and its local context. For crop type semantic segmentation fromSatellite Image Time Series(SITS) we find performance at parcel boundaries to be a critical bottleneck and explain how CSCL tackles the underlying cause of that problem, improving the state-of-the-art performance in this task. Additionally, using images from theSentinel-2(S2) satellite missions we compile the largest, to our knowledge, SITS dataset densely annotated by crop type and parcel identities, which we make publicly available together with the data generation pipeline. Using that data we find CSCL, even with minimal pre-training, to improve all respective baselines and present a process for semantic segmentation at greater resolution than that of the input images for obtaining crop classes at a more granular level. The code and instructions to download the data can be found in https://github.com/michaeltrs/DeepSatModels. Michail Tarasiou, Riza Alp Güler, Stefanos Zafeiriou |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | AutoHash: Learning Higher-Order Feature Interactions for Deep CTR PredictionabstractFeature combinations are essential for the success of many web applications, such as personalised recommendation and online advertising. State-of-the-art methods usually model explicit feature interactions to help neural networks reduce the number of parameters and achieve better performance. However, their explicit feature interactions are often restricted to the second-order due to computational complexity. In this work, we propose efficient ways to represent explicit high-order feature combinations as well as prune redundant features in the mean time. To begin with, we make novel use of the Count Sketch algorithm within a DNN classifier such that high-order feature combinations can be compactly represented. After that, to combat the problem of redundant features which degrade the prediction performance, we introduce an adaptive hashing algorithm, AutoHash, which can automatically select meaningful features to interact at high orders according to the specific dataset in question. This is an AutoML approach. Experiments on three well-known public datasets demonstrate that AutoHash is significantly superior to state-of-the-art methods. Meanwhile, due to its efficient scheme of automatically selecting useful high-order feature interactions, AutoHash has less model complexity and can be trained in an end-to-end manner with less training time than state-of-the-art methods. Niannan Xue, Bin Liu 0072, Huifeng Guo, Ruiming Tang, Fengwei Zhou, Stefanos Zafeiriou, Jun Wang 0012, Zhenguo Li |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2021 | Binary Graph Neural NetworksabstractGraph Neural Networks (GNNs) have emerged as a powerful and flexible framework for representation learning on irregular data. As they generalize the operations of classical CNNs on grids to arbitrary topologies, GNNs also bring much of the implementation challenges of their Euclidean counterparts. Model size, memory footprint, and energy consumption are common concerns for many real-world applications. Network binarization allocates a single bit to parameters and activations, thus dramatically reducing the memory requirements (up to 32x compared to single-precision floating-point numbers) and maximizing the benefits of fast SIMD instructions on modern hardware for measurable speedups. However, in spite of the large body of work on binarization for classical CNNs, this area remains largely unexplored in geometric deep learning. In this paper, we present and evaluate different strategies for the binarization of graph neural networks. We show that through careful design of the models, and control of the training process, binary graph neural networks can be trained at only a moderate cost in accuracy on challenging benchmarks. In particular, we present the first dynamic graph neural network in Hamming space, able to leverage efficient k-NN search on binary vectors to speed-up the construction of the dynamic graph. We further verify that the binary models offer significant savings on embedded devices. Our code is publicly available on Github1. Mehdi Bahri, Gaétan Bahl, Stefanos Zafeiriou |
CVPR | 3 |
| 2021 | Variational Prototype Learning for Deep Face RecognitionabstractDeep face recognition has achieved remarkable improvements due to the introduction of margin-based softmax loss, in which the prototype stored in the last linear layer represents the center of each class. In these methods, training samples are enforced to be close to positive prototypes and far apart from negative prototypes by a clear margin. However, we argue that prototype learning only employs sample-to-prototype comparisons without considering sample-to-sample comparisons during training and the low loss value gives us an illusion of perfect feature embedding, impeding the further exploration of SGD. To this end, we propose Variational Prototype Learning (VPL), which represents every class as a distribution instead of a point in the latent space. By identifying the slow feature drift phenomenon, we directly inject memorized features into prototypes to approximate variational prototype sampling. The proposed VPL can simulate sample-to-sample comparisons within the classification framework, encouraging the SGD solver to be more exploratory, while boosting performance. Moreover, VPL is conceptually simple, easy to implement, computationally efficient and memory saving. We present extensive experimental results on popular benchmarks, which demonstrate the superiority of the proposed VPL method over the state-of-the-art competitors. Jiankang Deng, Jia Guo 0003, Jing Yang 0038, Alexander Lattas, Stefanos Zafeiriou |
CVPR | 5 |
| 2021 | OSTeC: One-Shot Texture CompletionabstractThe last few years have witnessed the great success of non-linear generative models in synthesizing high-quality photorealistic face images. Many recent 3D facial texture reconstruction and pose manipulation from a single image approaches still rely on large and clean face datasets to train image-to-image Generative Adversarial Networks (GANs). Yet the collection of such a large scale high-resolution 3D texture dataset is still very costly and difficult to maintain age/ethnicity balance. Moreover, regression-based approaches suffer from generalization to the in-the-wild conditions and are unable to fine-tune to a target-image. In this work, we propose an unsupervised approach for one-shot 3D facial texture completion that does not re-quire large-scale texture datasets, but rather harnesses the knowledge stored in 2D face generators. The proposed approach rotates an input image in 3D and fill-in the unseen regions by reconstructing the rotated image in a 2D face generator, based on the visible parts. Finally, we stitch the most visible textures at different angles in the UV image-plane. Further, we frontalize the target image by projecting the completed texture into the generator. The qualitative and quantitative experiments demonstrate that the completed UV textures and frontalized images are of high quality, resembles the original identity, can be used to train a texture GAN model for 3DMM fitting and improve pose-invariant face recognition.1 Baris Gecer, Jiankang Deng, Stefanos Zafeiriou |
CVPR | 3 |
| 2021 | Speech Emotion Recognition Using Semantic InformationabstractSpeech emotion recognition is a crucial problem manifesting in a multitude of applications such as human computer interaction and education. Although several advancements have been made in the recent years, especially with the advent of Deep Neural Networks (DNN), most of the studies in the literature fail to consider the semantic information in the speech signal. In this paper, we propose a novel framework that can capture both the semantic and the paralinguistic information in the signal. In particular, our framework is comprised of a semantic feature extractor, that captures the semantic information, and a paralinguistic feature extractor, that captures the paralinguistic information. Both semantic and paraliguistic features are then combined to a unified representation using a novel attention mechanism. The unified feature vector is passed through a LSTM to capture the temporal dynamics in the signal, before the final prediction. To validate the effectiveness of our framework, we use the popular SEWA dataset of the AVEC challenge series and compare with the three winning papers. Our model provides state-of-the-art results in the valence and liking dimensions.1 Panagiotis Tzirakis, Anh Nguyen 0003, Stefanos Zafeiriou, Björn W. Schuller |
ICASSP | 3 |
| 2021 | Poly-NL: Linear Complexity Non-local Layers With 3rd Order PolynomialsabstractSpatial self-attention layers, in the form of Non-Local blocks, introduce long-range dependencies in Convolutional Neural Networks by computing pairwise similarities among all possible positions. Such pairwise functions underpin the effectiveness of non-local layers, but also determine a complexity that scales quadratically with respect to the input size both in space and time. This is a severely limiting factor that practically hinders the applicability of non-local blocks to even moderately sized inputs. Previous works focused on reducing the complexity by modifying the underlying matrix operations, however in this work we aim to retain full expressiveness of non-local layers while keeping complexity linear. We overcome the efficiency limitation of non-local blocks by framing them as special cases of 3rd order polynomial functions. This fact enables us to formulate novel fast Non-Local blocks, capable of reducing the complexity from quadratic to linear with no loss in performance, by replacing any direct computation of pairwise similarities with element-wise multiplications. The proposed method, which we dub as "Poly-NL", is competitive with state-of-the-art performance across image recognition, instance segmentation, and face detection tasks, while having considerably less computational overhead. Francesca Babiloni, Ioannis Marras, Filippos Kokkinos, Jiankang Deng, Grigorios Chrysos 0002, Stefanos Zafeiriou |
ICCV | 6 |
| 2021 | HeadGAN: One-shot Neural Head Synthesis and EditingabstractRecent attempts to solve the problem of head reenactment using a single reference image have shown promising results. However, most of them either perform poorly in terms of photo-realism, or fail to meet the identity preservation problem, or do not fully transfer the driving pose and expression. We propose HeadGAN, a novel system that conditions synthesis on 3D face representations, which can be extracted from any driving video and adapted to the facial geometry of any reference image, disentangling identity from expression. We further improve mouth movements, by utilising audio features as a complementary input. The 3D face representation enables HeadGAN to be further used as an efficient method for compression and reconstruction and a tool for expression and pose editing. Michail C. Doukas, Stefanos Zafeiriou, Viktoriia Sharmanska |
ICCV | 2 |
| 2021 | Shape My Face: Registering 3D Face Scans by Surface-to-Surface TranslationabstractAbstract Standard registration algorithms need to be independently applied to each surface to register, following careful pre-processing and hand-tuning. Recently, learning-based approaches have emerged that reduce the registration of new scans to running inference with a previously-trained model. The potential benefits are multifold: inference is typically orders of magnitude faster than solving a new instance of a difficult optimization problem, deep learning models can be made robust to noise and corruption, and the trained model may be re-used for other tasks, e.g. through transfer learning. In this paper, we cast the registration task as a surface-to-surface translation problem, and design a model to reliably capture the latent geometric information directly from raw 3D face scans. We introduce Shape-My-Face (SMF), a powerful encoder-decoder architecture based on an improved point cloud encoder, a novel visual attention mechanism, graph convolutional decoders with skip connections, and a specialized mouth model that we smoothly integrate with the mesh convolutions. Compared to the previous state-of-the-art learning algorithms for non-rigid registration of face scans, SMF only requires the raw data to be rigidly aligned (with scaling) with a pre-defined face template. Additionally, our model provides topologically-sound meshes with minimal supervision, offers faster training time, has orders of magnitude fewer trainable parameters, is more robust to noise, and can generalize to previously unseen datasets. We extensively evaluate the quality of our registrations on diverse data. We demonstrate the robustness and generalizability of our model with in-the-wild face scans across different modalities, sensor types, and resolutions. Finally, we show that, by learning to register scans, SMF produces a hybrid linear and non-linear morphable model. Manipulation of the latent space of SMF allows for shape generation, and morphing applications such as expression transfer in-the-wild. We train SMF on a dataset of human faces comprising 9 large-scale databases on commodity hardware. Mehdi Bahri, Eimear O' Sullivan, Shunwang Gong, Feng Liu 0037, Xiaoming Liu 0002, Michael M. Bronstein, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 7 |
| 2021 | Towards a Complete 3D Morphable Model of the Human HeadabstractThree-dimensional morphable models (3DMMs) are powerful statistical tools for representing the 3D shapes and textures of an object class. Here we present the most complete 3DMM of the human head to date that includes face, cranium, ears, eyes, teeth and tongue. To achieve this, we propose two methods for combining existing 3DMMs of different overlapping head parts: (i). use a regressor to complete missing parts of one model using the other, and (ii). use the Gaussian Process framework to blend covariance matrices from multiple models. Thus, we build a new combined face-and-head shape model that blends the variability and facial detail of an existing face model (the LSFM) with the full head modelling capability of an existing head model (the LYHM). Then we construct and fuse a highly-detailed ear model to extend the variation of the ear shape. Eye and eye region models are incorporated into the head model, along with basic models of the teeth, tongue and inner mouth cavity. The new model achieves state-of-the-art performance. We use our model to reconstruct full head representations from single, unconstrained images allowing us to parameterize craniofacial shape and texture, along with the ear shape, eye gaze and eye color. Stylianos Ploumpis, Evangelos Ververas, Eimear O' Sullivan, Stylianos Moschoglou, Haoyang Wang 0002, Nick E. Pears, William A. P. Smith, Baris Gecer, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2021 | Tensor Methods in Computer Vision and Deep LearningabstractTensors, or multidimensional arrays, are data structures that can naturally represent visual data of multiple dimensions. Inherently able to efficiently capture structured, latent semantic spaces and high-order interactions, tensors have a long history of applications in a wide span of computer vision problems. With the advent of the deep learning paradigm shift in computer vision, tensors have become even more fundamental. Indeed, essential ingredients in modern deep learning architectures, such as convolutions and attention mechanisms, can readily be considered as tensor mappings. In effect, tensor methods are increasingly finding significant applications in deep learning, including the design of memory and compute efficient network architectures, improving robustness to random noise and adversarial attacks, and aiding the theoretical understanding of deep networks. This article provides an in-depth and practical review of tensors and tensor methods in the context of representation learning and deep learning, with a particular focus on visual data analysis and computer vision applications. Concretely, besides fundamental work in tensor-based visual data analysis methods, we focus on recent developments that have brought on a gradual increase in tensor methods, especially in deep learning architectures and their implications in computer vision applications. To further enable the newcomer to grasp such concepts quickly, we provide companion Python notebooks, covering key aspects of this article and implementing them, step-by-step with TensorLy. Yannis Panagakis, Jean Kossaifi, Grigorios Chrysos 0002, James Oldfield 0001, Mihalis A. Nicolaou, Anima Anandkumar, Stefanos Zafeiriou |
Proc. IEEE | 7 |
| 2021 | Exploiting Multi-CNN Features in CNN-RNN Based Dimensional Emotion Recognition on the OMG in-the-Wild DatasetabstractThis article presents a novel CNN-RNN based approach, which exploits multiple CNN features for dimensional emotion recognition in-the-wild, utilizing the One-Minute Gradual-Emotion (OMG-Emotion) dataset. Our approach includes first pre-training with the relevant and large in size, Aff-Wild and Aff-Wild2 emotion databases. Low-, mid- and high-level features are extracted from the trained CNN component and are exploited by RNN subnets in a multi-task framework. Their outputs constitute an intermediate level prediction; final estimates are obtained as the mean or median values of these predictions. Fusion of the networks is also examined for boosting the obtained performance, at Decision-, or at Model-level; in the latter case a RNN was used for the fusion. Our approach, although using only the visual modality, outperformed state-of-the-art methods that utilized audio and visual modalities. Some of our developments have been submitted to the OMG-Emotion Challenge, ranking second among the technologies which used only visual information for valence estimation; ranking third overall. Through extensive experimentation, we further show that arousal estimation is greatly improved when low-level features are combined with high-level ones. Dimitris Kollias, Stefanos Zafeiriou |
IEEE Trans. Affect. Comput. | 2 |
| 2020 | VA-StarGAN: Continuous Affect Generation
Dimitris Kollias, Stefanos Zafeiriou |
ACIVS | 2 |
| 2020 | TESA: Tensor Element Self-Attention via MatricizationabstractRepresentation learning is a fundamental part of modern computer vision, where abstract representations of data are encoded as tensors optimized to solve problems like image segmentation and inpainting. Recently, self-attention in the form of Non-Local Block has emerged as a powerful technique to enrich features, by capturing complex interdependencies in feature tensors. However, standard self-attention approaches leverage only spatial relationships, drawing similarities between vectors and overlooking correlations between channels. In this paper, we introduce a new method, called Tensor Element Self-Attention (TESA) that generalizes such work to capture interdependencies along all dimensions of the tensor using matricization. An order R tensor produces R results, one for each dimension. The results are then fused to produce an enriched output which encapsulates similarity among tensor elements. Additionally, we analyze self-attention mathematically, providing new perspectives on how it adjusts the singular values of the input feature tensor. With these new insights, we present experimental results demonstrating how TESA can benefit diverse problems including classification and instance segmentation. By simply adding a TESA module to existing networks, we substantially improve competitive baselines and set new state-of-the-art results for image inpainting on Celeb and low light raw-to-rgb image translation on SID. Francesca Babiloni, Ioannis Marras, Gregory Slabaugh, Stefanos Zafeiriou |
CVPR | 4 |
| 2020 | P-nets: Deep Polynomial Neural NetworksabstractDeep Convolutional Neural Networks (DCNNs) is currently the method of choice both for generative, as well as for discriminative learning in computer vision and machine learning. The success of DCNNs can be attributed to the careful selection of their building blocks (e.g., residual blocks, rectifiers, sophisticated normalization schemes, to mention but a few). In this paper, we propose Π-Nets, a new class of DCNNs. Π-Nets are polynomial neural networks, i.e., the output is a high-order polynomial of the input. Π-Nets can be implemented using special kind of skip connections and their parameters can be represented via high-order tensors. We empirically demonstrate that Π-Nets have better representation power than standard DCNNs and they even produce good results without the use of non-linear activation functions in a large battery of tasks and signals, i.e., images, graphs, and audio. When used in conjunction with activation functions, Π-Nets produce state-of-the-art results in challenging tasks, such as image generation. Lastly, our framework elucidates why recent generative models, such as StyleGAN, improve upon their predecessors, e.g., ProGAN. Grigorios Chrysos 0002, Stylianos Moschoglou, Giorgos Bouritsas, Yannis Panagakis, Jiankang Deng, Stefanos Zafeiriou |
CVPR | 6 |
| 2020 | RetinaFace: Single-Shot Multi-Level Face Localisation in the WildabstractThough tremendous strides have been made in uncontrolled face detection, accurate and efficient 2D face alignment and 3D face reconstruction in-the-wild remain an open challenge. In this paper, we present a novel single-shot, multi-level face localisation method, named RetinaFace, which unifies face box prediction, 2D facial landmark localisation and 3D vertices regression under one common target: point regression on the image plane. To fill the data gap, we manually annotated five facial landmarks on the WIDER FACE dataset and employed a semi-automatic annotation pipeline to generate 3D vertices for face images from the WIDER FACE, AFLW and FDDB datasets. Based on extra annotations, we propose a mutually beneficial regression target for 3D face reconstruction, that is predicting 3D vertices projected on the image plane constrained by a common 3D topology. The proposed 3D face reconstruction branch can be easily incorporated, without any optimisation difficulty, in parallel with the existing box and 2D landmark regression branches during joint training. Extensive experimental results show that RetinaFace can simultaneously achieve stable face detection, accurate 2D face alignment and robust 3D face reconstruction while being efficient through single-shot inference. Jiankang Deng, Jia Guo 0003, Evangelos Ververas, Irene Kotsia, Stefanos Zafeiriou |
CVPR | 5 |
| 2020 | Geometrically Principled Connections in Graph Neural NetworksabstractGraph convolution operators bring the advantages of deep learning to a variety of graph and mesh processing tasks previously deemed out of reach. With their continued success comes the desire to design more powerful architectures, often by adapting existing deep learning techniques to non-Euclidean data. In this paper, we argue geometry should remain the primary driving force behind innovation in the emerging field of geometric deep learning. We relate graph neural networks to widely successful computer graphics and data approximation models: radial basis functions (RBFs). We conjecture that, like RBFs, graph convolution layers would benefit from the addition of simple functions to the powerful convolution kernels. We introduce affine skip connections, a novel building block formed by combining a fully connected layer with any graph convolution operator. We experimentally demonstrate the effectiveness of our technique, and show the improved performance is the consequence of more than the increased number of parameters. Operators equipped with the affine skip connection markedly outperform their base performance on every task we evaluated, i.e., shape reconstruction, dense shape correspondence, and graph classification. We hope our simple and effective approach will serve as a solid baseline and help ease future research in graph neural networks. Shunwang Gong, Mehdi Bahri, Michael M. Bronstein, Stefanos Zafeiriou |
CVPR | 4 |
| 2020 | DeepFaceFlow: In-the-Wild Dense 3D Facial Motion EstimationabstractDense 3D facial motion capture from only monocular in-the-wild pairs of RGB images is a highly challenging problem with numerous applications, ranging from facial expression recognition to facial reenactment. In this work, we propose DeepFaceFlow, a robust, fast, and highly-accurate framework for the dense estimation of 3D non-rigid facial flow between pairs of monocular images. Our DeepFaceFlow framework was trained and tested on two very large-scale facial video datasets, one of them of our own collection and annotation, with the aid of occlusion-aware and 3D-based loss function. We conduct comprehensive experiments probing different aspects of our approach and demonstrating its improved performance against state-of-the-art flow and 3D reconstruction methods. Furthermore, we incorporate our framework in a full-head state-of-the-art facial video synthesis method and demonstrate the ability of our method in better representing and capturing the facial dynamics, resulting in a highly-realistic facial video synthesis. Given registered pairs of images, our framework generates 3D flow maps at 60 fps. Mohammad Rami Koujan, Anastasios Roussos, Stefanos Zafeiriou |
CVPR | 3 |
| 2020 | Weakly-Supervised Mesh-Convolutional Hand Reconstruction in the WildabstractWe introduce a simple and effective network architecture for monocular 3D hand pose estimation consisting of an image encoder followed by a mesh convolutional decoder that is trained through a direct 3D hand mesh reconstruction loss. We train our network by gathering a large-scale dataset of hand action in YouTube videos and use it as a source of weak supervision. Our weakly-supervised mesh convolutions-based system largely outperforms state-of-the-art methods, even halving the errors on the in the wild benchmark. The dataset and additional resources are available at https://arielai.com/mesh_hands. Dominik Kulon, Riza Alp Güler, Iasonas Kokkinos, Michael M. Bronstein, Stefanos Zafeiriou |
CVPR | 5 |
| 2020 | AvatarMe: Realistically Renderable 3D Facial Reconstruction "In-the-Wild"abstractOver the last years, with the advent of Generative Adversarial Networks (GANs), many face analysis tasks have accomplished astounding performance, with applications including, but not limited to, face generation and 3D face reconstruction from a single "in-the-wild" image. Nevertheless, to the best of our knowledge, there is no method which can produce high-resolution photorealistic 3D faces from "in-the-wild" images and this can be attributed to the: (a) scarcity of available data for training, and (b) lack of robust methodologies that can successfully be applied on very high-resolution data. In this paper, we introduce AvatarMe, the first method that is able to reconstruct photorealistic 3D faces from a single "in-the-wild" image with an increasing level of detail. To achieve this, we capture a large dataset of facial shape and reflectance and build on a state-of-the-art 3D texture and shape reconstruction method and successively refine its results, while generating the per-pixel diffuse and specular components that are required for realistic rendering. As we demonstrate in a series of qualitative and quantitative experiments, AvatarMe outperforms the existing arts by a significant margin and reconstructs authentic, 4K by 6K-resolution 3D faces from a single low-resolution image that, for the first time, bridges the uncanny valley. Alexander Lattas, Stylianos Moschoglou, Baris Gecer, Stylianos Ploumpis, Vasileios Triantafyllou, Abhijeet Ghosh, Stefanos Zafeiriou |
CVPR | 7 |
| 2020 | BLSM: A Bone-Level Skinned Model of the Human Mesh
Haoyang Wang 0002, Riza Alp Güler, Iasonas Kokkinos, George Papandreou, Stefanos Zafeiriou |
ECCV (5) | 5 |
| 2020 | Sub-center ArcFace: Boosting Face Recognition by Large-Scale Noisy Web Faces
Jiankang Deng, Jia Guo 0003, Tongliang Liu, Mingming Gong, Stefanos Zafeiriou |
ECCV (11) | 5 |
| 2020 | Synthesizing Coupled 3D Face Modalities by Trunk-Branch Generative Adversarial Networks
Baris Gecer, Alexander Lattas, Stylianos Ploumpis, Jiankang Deng, Athanasios Papaioannou, Stylianos Moschoglou, Stefanos Zafeiriou |
ECCV (29) | 7 |
| 2020 | Reconstructing the Noise Variance Manifold for Image Denoising
Ioannis Marras, Grigorios Chrysos 0002, Ioannis Alexiou, Gregory Slabaugh, Stefanos Zafeiriou |
ECCV (9) | 5 |
| 2020 | Learning to Generate Customized Dynamic 3D Facial Expressions
Rolandos Alexandros Potamias, Stylianos Ploumpis, Giorgos Bouritsas, Evangelos Ververas, Stefanos Zafeiriou |
ECCV (29) | 6 |
| 2020 | Analysing Affective Behavior in the First ABAW 2020 CompetitionabstractThe Affective Behavior Analysis in-the-wild (ABAW) 2020 Competition is the first Competition aiming at automatic analysis of the three main behavior tasks of valence-arousal estimation, basic expression recognition and action unit detection. It is split into three Challenges, each one addressing a respective behavior task. For the Challenges, we provide a common benchmark database, Aff-Wild2, which is a large scale in-the-wild database and the first one annotated for all these three tasks. In this paper, we describe this Competition, to be held in conjunction with the IEEE Conference on Face and Gesture Recognition, May 2020, in Buenos Aires, Argentina. We present the three Challenges, with the utilized Competition corpora. We outline the evaluation metrics, present both the baseline system and the top-3 performing teams' methodologies per Challenge and finally present their obtained results. More information regarding the Competition, the leaderboard of each Challenge and details for accessing the utilized database, are provided in the Competition site: http://ibug.doc.ic.ac.uk/ resources/fg-2020-competition-affective-behavior-analysis. Dimitris Kollias, Attila Schulc, Elnar Hajiyev, Stefanos Zafeiriou |
FG | 4 |
| 2020 | Head2Head: Video-based Neural Head SynthesisabstractIn this paper, we propose a novel machine learning architecture for facial reenactment. In particular, contrary to the model-based approaches or recent frame-based methods that use Deep Convolutional Neural Networks (DCNNs) to generate individual frames, we propose a novel method that (a) exploits the special structure of facial motion (paying particular attention to mouth motion) and (b) enforces temporal consistency. We demonstrate that the proposed method can transfer facial expressions, pose and gaze of a source actor to a target video in a photo-realistic fashion more accurately than state-of-the-art methods. Mohammad Rami Koujan, Michail C. Doukas, Anastasios Roussos, Stefanos Zafeiriou |
FG | 4 |
| 2020 | ReenactNet: Real-time Full Head ReenactmentabstractVideo-to-video synthesis is a challenging problem aiming at learning a translation function between a sequence of semantic maps and a photo-realistic video depicting the characteristics of a driving video. We propose a head-to-head system of our own implementation capable of fully transferring the human head 3D pose, facial expressions and eye gaze from a source to a target actor, while preserving the identity of the target actor. Our system produces high-fidelity, temporally-smooth and photo-realistic synthetic videos faithfully transferring the human time-varying head attributes from the source to the target actor. Our proposed implementation: 1) works in real time (~20 fps), 2) runs on a commodity laptop with a webcam as the only input, 3) is interactive, allowing the participant to drive a target person, e.g. a celebrity, politician, etc, instantly by varying their expressions, head pose, and eye gaze, and visualising the synthesised video concurrently. Mohammad Rami Koujan, Michail C. Doukas, Anastasios Roussos, Stefanos Zafeiriou |
FG | 4 |
| 2020 | Face Video Generation from a Single Image and LandmarksabstractIn this paper, we are concerned with the challenging problem of producing a full image sequence of a deformable face given only an image and generic facial motions encoded by a set of sparse landmarks. To this end, we build upon recent breakthroughs in image-to-image translation such as pix2pix, CycleGAN and StarGAN which learn Deep Convolutional Neural Networks (DCNNs) that learn to map aligned pairs of images between different domains (i.e., having different labels) and propose a new architecture which is not driven any more by labels but by spatial maps, facial landmarks. In particular, we propose the MotionGAN which transforms an input face image into a new one according to a heatmap of target landmarks. We show that it is possible to create very realistic face videos using a single image and a set of target landmarks. Furthermore, our method can be used to edit a facial image with arbitrary motions according to landmarks (e.g., expression, speech, etc.). This provides much more flexibility to face editing, expression transfer, facial video creation, etc. than models based on discrete expressions, audio or action units. Kritaphat Songsri-in, Stefanos Zafeiriou |
FG | 2 |
| 2020 | 3D Landmark Localization in Point Clouds for the Human Earabstract3D landmark localization plays an important role in many aspects of 3D data processing, from morphometric analysis to the initialization of mesh registration algorithms. In this work we address the problem of landmark localization in 3D point clouds by extending leading 2D landmark localization algorithms to the 3D domain. By leveraging the PointNet++ architecture, we can construct an architecture that is invariant to the ordering of the data. Input point clouds are segmented into background and landmark regions, and offset vectors are calculated within the landmark regions to refine predicted landmark locations. We demonstrate a high landmark localization accuracy, even as the number of points in the input point cloud decreases. By making use of a 3D morphable model as a novel means of data augmentation, improved landmark localization accuracy and consistency can be obtained. We present our results for landmark localization on the human ear. Eimear O' Sullivan, Stefanos Zafeiriou |
FG | 2 |
| 2020 | Synthesising 3D Facial Motion from "In-the-Wild" SpeechabstractSynthesising 3D facial motion from speech is a crucial problem manifesting in a multitude of applications such as computer games and movies. Recently proposed methods tackle this problem in controlled conditions of speech. In this paper, we introduce the first methodology for 3D facial motion synthesis from speech captured in arbitrary recording conditions (“in-the-wild”) and independent of the speaker. For our purposes, we captured 4D sequences of people uttering 500 words, contained in the Lip Reading in the Wild (LRW) words, a publicly available large-scale in-the-wild dataset, and built a set of 3D blendshapes appropriate for speech. We correlate the 3D shape parameters of the speech blendshapes to the LRW audio samples by means of a novel time-warping technique, named Deep Canonical Attentional Warping (DCAW), that can simultaneously learn hierarchical non-linear representations and a warping path in an end-to-end manner. We thoroughly evaluate our proposed methods, and show the ability of a deep learning model to synthesise 3D facial motion in handling different speakers and continuous speech signals in uncontrolled conditions1. Panagiotis Tzirakis, Athanasios Papaioannou, Alexander Lattas, Michail Tarasiou, Björn W. Schuller, Stefanos Zafeiriou |
FG | 6 |
| 2020 | Extracting Deep Local Features to Detect Manipulated Images of Human FacesabstractRecent developments in computer vision and machine learning have made it possible to create realistic manipulated videos of human faces, raising the issue of ensuring adequate protection against the malevolent effects unlocked by such capabilities. In this paper we propose local image features which are shared across manipulated regions as a key element for the automatic detection of manipulated face images. We also design a lightweight architecture with the correct structural biases for extracting such features and derive a multitask training scheme that consistently outperforms image class supervision alone. The trained networks achieve state-of-the-art results in the FaceForensics++ dataset using significantly reduced number of parameters and are shown to work well in detecting fully generated face images. Michail Tarasiou, Stefanos Zafeiriou |
ICIP | 2 |
| 2020 | Ear Cartilage Inference for Reconstructive Surgery with Convolutional Mesh Autoencoders
Eimear O' Sullivan, Lara van de Lande, Antonia Osolos, Silvia Schievano, David J. Dunaway, Neil Bulstrode, Stefanos Zafeiriou |
MICCAI (3) | 7 |
| 2020 | AutoGroup: Automatic Feature Grouping for Modelling Explicit High-Order Feature Interactions in CTR PredictionabstractModelling feature interactions is key in Click-Through Rate (CTR) predictions. State-of-the-art models usually include explicit feature interactions to better model non-linearity in a deep network, but enumerating all feature combinations of high orders is not efficient and brings challenges to network optimization. In this work, we use AutoML to seek useful high-order feature interactions to train on without manual feature selection. For this purpose, an end-to-end model, AutoGroup, is proposed, which casts the selection of feature interactions as a structural optimization problem. In a nutshell, AutoGroup first automatically groups useful features into a number of feature sets. Then, it generates interactions of any order from these feature sets using a novel interaction function. The main contribution of AutoGroup is that it performs both dimensionality reduction and feature selection which are not seen in previous models. Offline experiments on three public large-scale benchmark datasets demonstrate the superior performance and efficiency of AutoGroup over state-of-the-art models. Furthermore, a ten-day online A/B test verifies that AutoGroup can be reliably deployed in production and outperform the commercial baseline by 10% on average in terms of CTR and CVR. Bin Liu 0072, Niannan Xue, Huifeng Guo, Ruiming Tang, Stefanos Zafeiriou, Xiuqiang He 0001, Zhenguo Li |
SIGIR | 5 |
| 2020 | Facial Expression Synthesis using a Global-Local Multilinear FrameworkabstractAbstract We present a practical method to synthesize plausible 3D facial expressions for a particular target subject. The ability to synthesize an entire facial rig from a single neutral expression has a large range of applications both in computer graphics and computer vision, ranging from the efficient and cost‐effective creation of CG characters to scalable data generation for machine learning purposes. Unlike previous methods based on multilinear models, the proposed approach is capable to extrapolate well outside the sample pool, which allows it to plausibly predict the identity of the target subject and create artifact free expression shapes while requiring only a small input dataset. We introduce global‐local multilinear models that leverage the strengths of expression‐specific and identity‐specific local models combined with coarse motion estimations from a global model. Experimental results show that we achieve high‐quality, plausible facial expression synthesis results for an individual that outperform existing methods both quantitatively and qualitatively. Mengjiao Wang 0002, Derek Bradley, Stefanos Zafeiriou, Thabo Beeler |
Comput. Graph. Forum | 3 |
| 2020 | RoCGAN: Robust Conditional GANabstractAbstract Conditional image generation lies at the heart of computer vision and conditional generative adversarial networks (cGAN) have recently become the method of choice for this task, owing to their superior performance. The focus so far has largely been on performance improvement, with little effort in making cGANs more robust to noise. However, the regression (of the generator) might lead to arbitrarily large errors in the output, which makes cGANs unreliable for real-world applications. In this work, we introduce a novel conditional GAN model, calledRoCGAN, which leverages structure in the target space of the model to address the issue. Specifically, we augment the generator with an unsupervised pathway, which promotes the outputs of the generator to span the target manifold, even in the presence of intense noise. We prove that RoCGAN share similar theoretical properties as GAN and establish with both synthetic and real data the merits of our model. We perform a thorough experimental validation on large scale datasets for natural scenes and faces and observe that our model outperforms existing cGAN architectures by a large margin. We also empirically demonstrate the performance of our approach in the face of two types of noise (adversarial and Bernoulli). Grigorios Chrysos 0002, Jean Kossaifi, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 3 |
| 2020 | Deep Neural Network Augmentation: Generating Faces for Affect AnalysisabstractAbstract This paper presents a novel approach for synthesizing facial affect; either in terms of the six basic expressions (i.e., anger, disgust, fear, joy, sadness and surprise), or in terms of valence (i.e., how positive or negative is an emotion) and arousal (i.e., power of the emotion activation). The proposed approach accepts the following inputs:(i) a neutral 2D image of a person; (ii) a basic facial expression or a pair of valence-arousal (VA) emotional state descriptors to be generated, or a path of affect in the 2D VA space to be generated as an image sequence. In order to synthesize affect in terms of VA, for this person, 600,000 frames from the 4DFAB database were annotated. The affect synthesis is implemented by fitting a 3D Morphable Model on the neutral image, then deforming the reconstructed face and adding the inputted affect, and blending the new face with the given affect into the original image. Qualitative experiments illustrate the generation of realistic images, when the neutral image is sampled from fifteen well known lab-controlled or in-the-wild databases, including Aff-Wild, AffectNet, RAF-DB; comparisons with generative adversarial networks (GANs) show the higher quality achieved by the proposed approach. Then, quantitative experiments are conducted, in which the synthesized images are used for data augmentation in training deep neural networks to perform affect recognition over all databases; greatly improved performances are achieved when compared with state-of-the-art methods, as well as with GAN-based data augmentation, in all cases. Dimitris Kollias, Shiyang Cheng 0001, Evangelos Ververas, Irene Kotsia, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 5 |
| 2020 | 3DFaceGAN: Adversarial Nets for 3D Face Representation, Generation, and TranslationabstractAbstract Over the past few years, Generative Adversarial Networks (GANs) have garnered increased interest among researchers in Computer Vision, with applications including, but not limited to, image generation, translation, imputation, and super-resolution. Nevertheless, no GAN-based method has been proposed in the literature that can successfully represent, generate or translate 3D facial shapes (meshes). This can be primarily attributed to two facts, namely that (a) publicly available 3D face databases are scarce as well as limited in terms of sample size and variability (e.g., few subjects, little diversity in race and gender), and (b) mesh convolutions for deep networks present several challenges that are not entirely tackled in the literature, leading to operator approximations and model instability, often failing to preserve high-frequency components of the distribution. As a result, linear methods such as Principal Component Analysis (PCA) have been mainly utilized towards 3D shape analysis, despite being unable to capture non-linearities and high frequency details of the 3D face—such as eyelid and lip variations. In this work, we present 3DFaceGAN, the first GAN tailored towards modeling the distribution of 3D facial surfaces, while retaining the high frequency details of 3D face shapes. We conduct an extensive series of both qualitative and quantitative experiments, where the merits of 3DFaceGAN are clearly demonstrated against other, state-of-the-art methods in tasks such as 3D shape representation, generation, and translation. Stylianos Moschoglou, Stylianos Ploumpis, Mihalis A. Nicolaou, Athanasios Papaioannou, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 5 |
| 2020 | SliderGAN: Synthesizing Expressive Face Images by Sliding 3D Blendshape ParametersabstractAbstract Image-to-image (i2i) translation is the dense regression problem of learning how to transform an input image into an output using aligned image pairs. Remarkable progress has been made in i2i translation with the advent of deep convolutional neural networks and particular using the learning paradigm of generative adversarial networks (GANs). In the absence of paired images, i2i translation is tackled with one or multiple domain transformations (i.e., CycleGAN, StarGAN etc.). In this paper, we study the problem of image-to-image translation, under a set of continuous parameters that correspond to a model describing a physical process. In particular, we propose the SliderGAN which transforms an input face image into a new one according to the continuous values of a statistical blendshape model of facial motion. We show that it is possible to edit a facial image according to expression and speech blendshapes, using sliders that control the continuous values of the blendshape model. This provides much more flexibility in various tasks, including but not limited to face editing, expression transfer and face neutralisation, comparing to models based on discrete expressions or action units. Evangelos Ververas, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 2 |
| 2019 | Expression, Affect, Action Unit Recognition: Aff-Wild2, Multi-Task Learning and ArcFace
Dimitris Kollias, Stefanos Zafeiriou |
BMVC | 2 |
| 2019 | Single Image 3D Hand Reconstruction with Mesh Convolutions
Dominik Kulon, Haoyang Wang 0002, Riza Alp Güler, Michael M. Bronstein, Stefanos Zafeiriou |
BMVC | 5 |
| 2019 | ArcFace: Additive Angular Margin Loss for Deep Face RecognitionabstractRecently, a popular line of research in face recognition is adopting margins in the well-established softmax loss function to maximize class separability. In this paper, we first introduce an Additive Angular Margin Loss (ArcFace), which not only has a clear geometric interpretation but also significantly enhances the discriminative power. Since ArcFace is susceptible to the massive label noise, we further propose sub-center ArcFace, in which each class contains K sub-centers and training samples only need to be close to any of the K positive sub-centers. Sub-center ArcFace encourages one dominant sub-class that contains the majority of clean faces and non-dominant sub-classes that include hard or noisy faces. Based on this self-propelled isolation, we boost the performance through automatically purifying raw web faces under massive real-world noise. Besides discriminative feature embedding, we also explore the inverse problem, mapping feature vectors to face images. Without training any additional generator or discriminator, the pre-trained ArcFace model can generate identity-preserved face images for both subjects inside and outside the training data only by using the network gradient and Batch Normalization (BN) priors. Extensive experiments demonstrate that ArcFace can enhance the discriminative feature embedding as well as strengthen the generative face synthesis. Jiankang Deng, Jia Guo 0003, Niannan Xue, Stefanos Zafeiriou |
CVPR | 4 |
| 2019 | GANFIT: Generative Adversarial Network Fitting for High Fidelity 3D Face ReconstructionabstractIn the past few years, a lot of work has been done towards reconstructing the 3D facial structure from single images by capitalizing on the power of Deep Convolutional Neural Networks (DCNNs). In the most recent works, differentiable renderers were employed in order to learn the relationship between the facial identity features and the parameters of a 3D morphable model for shape and texture. The texture features either correspond to components of a linear texture space or are learned by auto-encoders directly from in-the-wild images. In all cases, the quality of the facial texture reconstruction of the state-of-the-art methods is still not capable of modeling textures in high fidelity. In this paper, we take a radically different approach and harness the power of Generative Adversarial Networks (GANs) and DCNNs in order to reconstruct the facial texture and shape from single images. That is, we utilize GANs to train a very powerful generator of facial texture in UV space. Then, we revisit the original 3D Morphable Models (3DMMs) fitting approaches making use of non-linear optimization to find the optimal latent parameters that best reconstruct the test image but under a new perspective. We optimize the parameters with the supervision of pretrained deep identity features through our end-to-end differentiable framework. We demonstrate excellent results in photorealistic and identity preserving 3D face reconstructions and achieve for the first time, to the best of our knowledge, facial texture reconstruction with high-frequency details. Baris Gecer, Stylianos Ploumpis, Irene Kotsia, Stefanos Zafeiriou |
CVPR | 4 |
| 2019 | Combining 3D Morphable Models: A Large Scale Face-And-Head ModelabstractThree-dimensional Morphable Models (3DMMs) are powerful statistical tools for representing the 3D surfaces of an object class. In this context, we identify an interesting question that has previously not received research attention: is it possible to combine two or more 3DMMs that (a) are built using different templates that perhaps only partly overlap, (b) have different representation capabilities and (c) are built from different datasets that may not be publicly-available? In answering this question, we make two contributions. First, we propose two methods for solving this problem: i. use a regressor to complete missing parts of one model using the other, ii. use the Gaussian Process framework to blend covariance matrices from multiple models. Second, as an example application of our approach, we build a new head and face model that combines the variability and facial detail of the LSFM with the full head modelling of the LYHM. The resulting combined model achieves state-of-the-art performance and outperforms existing head models by a large margin. Finally, as an application experiment, we reconstruct full head representations from single, unconstrained images by utilizing our proposed large-scale model in conjunction with the Face-Warehouse blendshapes for handling expressions. Stylianos Ploumpis, Haoyang Wang 0002, Nick E. Pears, William A. P. Smith, Stefanos Zafeiriou |
CVPR | 5 |
| 2019 | Dense 3D Face Decoding Over 2500FPS: Joint Texture & Shape Convolutional Mesh Decodersabstract3D Morphable Models (3DMMs) are statistical models that represent facial texture and shape variations using a set of linear bases and more particular Principal Component Analysis (PCA). 3DMMs were used as statistical priors for reconstructing 3D faces from images by solving non-linear least square optimization problems. Recently, 3DMMs were used as generative models for training non-linear mappings (\ie, regressors) from image to the parameters of the models via Deep Convolutional Neural Networks (DCNNs). Nevertheless, all of the above methods use either fully connected layers or 2D convolutions on parametric unwrapped UV spaces leading to large networks with many parameters. In this paper, we present the first, to the best of our knowledge, non-linear 3DMMs by learning joint texture and shape auto-encoders using direct mesh convolutions. We demonstrate how these auto-encoders can be used to train very light-weight models that perform Coloured Mesh Decoding (CMD) in-the-wild at a speed of over 2500 FPS. Jiankang Deng, Irene Kotsia, Stefanos Zafeiriou |
CVPR | 4 |
| 2019 | Reparameterising 3D Statistical Shape Modelsabstract3D statistical shape models are widely used in modelling 3D shapes such as human faces and bodies. The limitation of such model is that, once built, the model can only represent 3D shape instances of a fixed mesh topology. While some applications may require a shape model of a different mesh topology, the model building pipeline has to be repeated with the new template, which could be time and computational resource consuming. In other cases only the statistical model is available and access to the original data is not possible. In this paper, we present a method to reparameterise a given 3D statistical shape model to any topology without using any training data. We also show that the reparameterised model achieves comparable performance as the original model. Haoyang Wang 0002, Stefanos Zafeiriou |
FG | 2 |
| 2019 | Time-series Clustering with Jointly Learning Deep Representations, Clusters and Temporal BoundariesabstractClustering and segmentation of temporal data is an important task across several fields, with prominent applications in computer vision and machine learning such as face and gesture segmentation. Several related methods have been proposed in literature, focusing on learning temporal boundaries and clusters, with recent works focusing on learning deep representations for clustering. However, none of the proposed methods is suitable for jointly learning segments, clusters, as well as representations. In this paper, we propose the first methodology that simultaneously discovers suitable deep representations, as well as clusters and temporal boundaries, with the clustering process providing supervisory cues for updating temporal boundaries and training the proposed deep learning architecture. We demonstrate the power of the proposed approach on a human motion segmentation task using the CMU-MMAC database. Our method provides the best results with respect to normalized mutual information compared to other clustering algorithms. Panagiotis Tzirakis, Mihalis A. Nicolaou, Björn W. Schuller, Stefanos Zafeiriou |
FG | 4 |
| 2019 | Neural 3D Morphable Models: Spiral Convolutional Networks for 3D Shape Representation Learning and GenerationabstractGenerative models for 3D geometric data arise in many important applications in 3D computer vision and graphics. In this paper, we focus on 3D deformable shapes that share a common topological structure, such as human faces and bodies. Morphable Models and their variants, despite their linear formulation, have been widely used for shape representation, while most of the recently proposed nonlinear approaches resort to intermediate representations, such as 3D voxel grids or 2D views. In this work, we introduce a novel graph convolutional operator, acting directly on the 3D mesh, that explicitly models the inductive bias of the fixed underlying graph. This is achieved by enforcing consistent local orderings of the vertices of the graph, through the spiral operator, thus breaking the permutation invariance property that is adopted by all the prior work on Graph Neural Networks. Our operator comes by construction with desirable properties (anisotropic, topology-aware, lightweight, easy-to-optimise), and by using it as a building block for traditional deep generative architectures, we demonstrate state-of-the-art results on a variety of 3D shape datasets compared to the linear Morphable Model and other graph convolutional operators. Giorgos Bouritsas, Sergiy Bokhnyak, Stylianos Ploumpis, Stefanos Zafeiriou, Michael M. Bronstein |
ICCV | 4 |
| 2019 | Robust Conditional Generative Adversarial Networks
Grigorios Chrysos 0002, Jean Kossaifi, Stefanos Zafeiriou |
ICLR (Poster) | 3 |
| 2019 | Motion Deblurring of FacesabstractFace analysis lies at the heart of computer vision with remarkable progress in the past decades. Face recognition and tracking are tackled by building invariance to fundamental modes of variation such as illumination, 3D pose. A much less standing mode of variation is motion deblurring, which however presents substantial challenges in face analysis. Recent approaches either make oversimplifying assumptions, e.g. in cases of joint optimization with other tasks, or fail to preserve the highly structured shape/identity information. We introduce a two-step architecture tailored to the challenges of motion deblurring: the first step restores the low frequencies; the second restores the high frequencies, while ensuring that the outputs span the natural images manifold. Both steps are implemented with a supervised data-driven method; to train those we devise a method for creating realistic motion blur by averaging a variable number of frames. The averaged images originate from the $$2MF^2$$ dataset with $$19$$ million facial frames, which we introduce for the task. Considering deblurring as an intermediate step, we conduct a thorough experimentation on high-level face analysis tasks, i.e. landmark localization and face verification, on blurred images. The experimental evaluation demonstrates the superiority of our method. Grigorios Chrysos 0002, Paolo Favaro, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 3 |
| 2019 | The Menpo Benchmark for Multi-pose 2D and 3D Facial Landmark Localisation and TrackingabstractIn this article, we present the Menpo 2D and Menpo 3D benchmarks, two new datasets for multi-pose 2D and 3D facial landmark localisation and tracking. In contrast to the previous benchmarks such as 300W and 300VW, the proposed benchmarks contain facial images in both semi-frontal and profile pose. We introduce an elaborate semi-automatic methodology for providing high-quality annotations for both the Menpo 2D and Menpo 3D benchmarks. In Menpo 2D benchmark, different visible landmark configurations are designed for semi-frontal and profile faces, thus making the 2D face alignment full-pose. In Menpo 3D benchmark, a united landmark configuration is designed for both semi-frontal and profile faces based on the correspondence with a 3D face model, thus making face alignment not only full-pose but also corresponding to the real-world 3D space. Based on the considerable number of annotated images, we organised Menpo 2D Challenge and Menpo 3D Challenge for face alignment under large pose variations in conjunction with CVPR 2017 and ICCV 2017, respectively. The results of these challenges demonstrate that recent deep learning architectures, when trained with the abundant data, lead to excellent results. We also provide a very simple, yet effective solution, named Cascade Multi-view Hourglass Model, to 2D and 3D face alignment. In our method, we take advantage of all 2D and 3D facial landmark annotations in a joint way. We not only capitalise on the correspondences between the semi-frontal and profile 2D facial landmarks but also employ joint supervision from both 2D and 3D facial landmarks. Finally, we discuss future directions on the topic of face alignment. Jiankang Deng, Anastasios Roussos, Grigorios Chrysos 0002, Evangelos Ververas, Irene Kotsia, Jie Shen 0008, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 7 |
| 2019 | Special Issue on Machine Vision
Tae-Kyun Kim 0001, Stefanos Zafeiriou, Ben Glocker, Stefan Leutenegger |
Int. J. Comput. Vis. | 2 |
| 2019 | Deep Affect Prediction in-the-Wild: Aff-Wild Database and Challenge, Deep Architectures, and BeyondabstractAutomatic understanding of human affect using visual signals is of great importance in everyday human–machine interactions. Appraising human emotional states, behaviors and reactions displayed in real-world settings, can be accomplished using latent continuous dimensions (e.g., the circumplex model of affect). Valence (i.e., how positive or negative is an emotion) and arousal (i.e., power of the activation of the emotion) constitute popular and effective representations for affect. Nevertheless, the majority of collected datasets this far, although containing naturalistic emotional states, have been captured in highly controlled recording conditions. In this paper, we introduce the Aff-Wild benchmark for training and evaluating affect recognition algorithms. We also report on the results of the First Affect-in-the-wild Challenge (Aff-Wild Challenge) that was recently organized in conjunction with CVPR 2017 on the Aff-Wild database, and was the first ever challenge on the estimation of valence and arousal in-the-wild. Furthermore, we design and extensively train an end-to-end deep neural architecture which performs prediction of continuous emotion dimensions based on visual cues. The proposed deep learning architecture, AffWildNet, includes convolutional and recurrent neural network layers, exploiting the invariant properties of convolutional features, while also modeling temporal dynamics that arise in human behavior via the recurrent layers. The AffWildNet produced state-of-the-art results on the Aff-Wild Challenge. We then exploit the AffWild database for learning features, which can be used as priors for achieving best performances both for dimensional, as well as categorical emotion recognition, using the RECOLA, AFEW-VA and EmotiW 2017 datasets, compared to all other methods designed for the same goal. The database and emotion recognition models are available at http://ibug.doc.ic.ac.uk/resources/first-affect-wild-challenge . Dimitris Kollias, Panagiotis Tzirakis, Mihalis A. Nicolaou, Athanasios Papaioannou, Guoying Zhao 0001, Björn W. Schuller, Irene Kotsia, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 8 |
| 2019 | An Adversarial Neuro-Tensorial Approach for Learning Disentangled RepresentationsabstractSeveral factors contribute to the appearance of an object in a visual scene, including pose, illumination, and deformation, among others. Each factor accounts for a source of variability in the data, while the multiplicative interactions of these factors emulate the entangled variability, giving rise to the rich structure of visual object appearance. Disentangling such unobserved factors from visual data is a challenging task, especially when the data have been captured in uncontrolled recording conditions (also referred to as “in-the-wild”) and label information is not available. In this paper, we propose a pseudo-supervised deep learning method for disentangling multiple latent factors of variation in face images captured in-the-wild. To this end, we propose a deep latent variable model, where the multiplicative interactions of multiple latent factors of variation are explicitly modelled by means of multilinear (tensor) structure. We demonstrate that the proposed approach indeed learns disentangled representations of facial expressions and pose, which can be used in various applications, including face editing, as well as 3D face reconstruction and classification of facial expression, identity and pose. Mengjiao Wang 0002, Zhixin Shu, Shiyang Cheng 0001, Yannis Panagakis, Dimitris Samaras, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 6 |
| 2019 | Robust Kronecker Component AnalysisabstractDictionary learning and component analysis models are fundamental for learning compact representations that are relevant to a given task (feature extraction, dimensionality reduction, denoising, etc.). The model complexity is encoded by means of specific structure, such as sparsity, low-rankness, or nonnegativity. Unfortunately, approaches like K-SVD - that learn dictionaries for sparse coding via Singular Value Decomposition (SVD) - are hard to scale to high-volume and high-dimensional visual data, and fragile in the presence of outliers. Conversely, robust component analysis methods such as the Robust Principal Component Analysis (RPCA) are able to recover low-complexity (e.g., low-rank) representations from data corrupted with noise of unknown magnitude and support, but do not provide a dictionary that respects the structure of the data (e.g., images), and also involve expensive computations. In this paper, we propose a novel Kronecker-decomposable component analysis model, coined as Robust Kronecker Component Analysis (RKCA), that combines ideas from sparse dictionary learning and robust component analysis. RKCA has several appealing properties, including robustness to gross corruption; it can be used for low-rank modeling, and leverages separability to solve significantly smaller problems. We design an efficient learning algorithm by drawing links with a restricted form of tensor factorization, and analyze its optimality and low-rankness properties. The effectiveness of the proposed approach is demonstrated on real-world applications, namely background subtraction and image denoising and completion, by performing a thorough comparison with the current state of the art. Mehdi Bahri, Yannis Panagakis, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Non-Negative Matrix Factorizations for Multiplex Network AnalysisabstractNetworks have been a general tool for representing, analyzing, and modeling relational data arising in several domains. One of the most important aspect of network analysis is community detection or network clustering. Until recently, the major focus have been on discovering community structure in single (i.e., monoplex) networks. However, with the advent of relational data with multiple modalities, multiplex networks, i.e., networks composed of multiple layers representing different aspects of relations, have emerged. Consequently, community detection in multiplex network, i.e., detecting clusters of nodes shared by all layers, has become a new challenge. In this paper, we propose Network Fusion for Composite Community Extraction (NF-CCE), a new class of algorithms, based on four different non-negative matrix factorization models, capable of extracting composite communities in multiplex networks. Each algorithm works in two steps: first, it finds a non-negative, low-dimensional feature representation of each network layer; then, it fuses the feature representation of layers into a common non-negative, low-dimensional feature representation via collective factorization. The composite clusters are extracted from the common feature representation. We demonstrate the superior performance of our algorithms over the state-of-the-art methods on various types of multiplex networks, including biological, social, economic, citation, phone communication, and brain multiplex networks. Vladimir Gligorijevic, Yannis Panagakis, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Side Information for Face Completion: A Robust PCA ApproachabstractRobust principal component analysis (RPCA) is a powerful method for learning low-rank feature representation of various visual data. However, for certain types as well as significant amount of error corruption, it fails to yield satisfactory results; a drawback that can be alleviated by exploiting domain-dependent prior knowledge or information. In this paper, we propose two models for the RPCA that take into account such side information, even in the presence of missing values. We apply this framework to the task of UV completion which is widely used in pose-invariant face recognition. Moreover, we construct a generative adversarial network (GAN) to extract side information as well as subspaces. These subspaces not only assist in the recovery but also speed up the process in case of large-scale data. We quantitatively and qualitatively evaluate the proposed approaches through both synthetic data and eight real-world datasets to verify their effectiveness. Niannan Xue, Jiankang Deng, Shiyang Cheng 0001, Yannis Panagakis, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2019 | Fast Multilevel Algorithms for Compressive Principal Component PursuitabstractRecovering a low-rank matrix from highly corrupted measurements arises in compressed sensing of structured high-dimensional signals (e.g., videos and hyperspectral images among others). Robust principal component analysis, solved via principal component pursuit (PCP), recovers a low-rank matrix from sparse corruptions that are of unknown value and support by decomposing the observation matrix into two terms, a low-rank matrix and a sparse one, accounting for sparse noise and outliers. In the more general setting, where only a fraction of the data matrix has been observed, low-rank matrix recovery is achieved by solving the compressive principal component pursuit (CPCP). Both PCP and CPCP are well-studied convex programs, and numerous iterative algorithms have been proposed for their optimization. Nevertheless, these algorithms involve singular value decomposition (SVD) at each iteration, which renders their applicability challenging in the case of massive data. In this paper, we propose a multilevel approach for the solution of PCP and CPCP problems. The core principle behind our algorithm is to apply SVD in models of lower dimensionality than the original one and then lift its solution to the original problem dimension. Hence, our methods rely on the assumption that the low-rank component can be represented in a lower dimensional space. We show that the proposed algorithms are easy to implement and converge at the same rate but with much lower iteration cost. Numerical experiments on numerous synthetic and real problems indicate that the proposed multilevel algorithms are several times faster than their original counterparts, namely, PCP and CPCP. Vahan Hovhannisyan, Yannis Panagakis, Panos Parpas, Stefanos Zafeiriou |
SIAM J. Imaging Sci. | 4 |
| 2019 | Editorial of Special Issue on Human Behaviour Analysis "In-the-Wild"abstractThe papers in this special section focus on human face and body image analysis, one of the most researched objects. One of the main reasons behind this popularity lies in the numerous applications of automatic face and body gesture analysis algorithms, that span several fields such as Human-Computer and Human-Robot Interaction (facial expression/body gesture recognition for automatic analysis of affect), medicine and healthcare (detection of emotional and cognitive disorders), as well as biometrics (face recognition, gait recognition). The papers in this section focus on recent efforts towards catalysing progress in automatic analysis of human behaviour in uncontrolled, “in-the-wild” conditions. We summarize research efforts towards the development of research methodologies, database collections and benchmarks, as well as algorithms and systems for machine analysis of human behaviour, focusing on facial expressions, body gestures, speech, as well as various other sensors. We are delighted that the special issue includes authors both from academia as well as the industry. Mihalis A. Nicolaou, Stefanos Zafeiriou, Irene Kotsia, Guoying Zhao 0001, Jeffrey F. Cohn |
IEEE Trans. Affect. Comput. | 2 |
| 2019 | Joint Multi-View Face Alignment in the WildabstractThe de facto algorithm for facial landmark estimation involves running a face detector with a subsequent deformable model fitting on the bounding box. This encompasses two basic problems: i) the detection and deformable fitting steps are performed independently, while the detector might not provide best-suited initialization for the fitting step, ii) the face appearance varies hugely across different poses, which makes the deformable face fitting very challenging and thus distinct models have to be used (e.g., one for profile and one for frontal faces). In this work, we propose the first, to the best of our knowledge, joint multi-view convolutional network to handle large pose variations across faces in-the-wild, and elegantly bridge face detection and facial landmark localization tasks. Existing joint face detection and landmark localization methods focus only on a very small set of landmarks. By contrast, our method can detect and align a large number of landmarks for semi-frontal (68 landmarks) and profile (39 landmarks) faces. We evaluate our model on a plethora of datasets including standard static image datasets such as IBUG, 300W, COFW, and the latest Menpo Benchmark for both semi-frontal and profile faces. Significant improvement over state-of-the-art methods on deformable face tracking is witnessed on 300VW benchmark. We also demonstrate state-ofthe- art results for face detection on FDDB and MALF datasets. Jiankang Deng, George Trigeorgis, Stefanos Zafeiriou |
IEEE Trans. Image Process. | 4 |
| 2018 | Informed Non-Convex Robust Principal Component Analysis With FeaturesabstractWe revisit the problem of robust principal component analysis with features acting as prior side information. To this aim, a novel, elegant, non-convex optimization approach is proposed to decompose a given observation matrix into a low-rank core and the corresponding sparse residual. Rigorous theoretical analysis of the proposed algorithm results in exact recovery guarantees with low computational complexity. Aptly designed synthetic experiments demonstrate that our method is the first to wholly harness the power of non-convexity over convexity in terms of both recoverability and speed. That is, the proposed non-convex approach is more accurate and faster compared to the best available algorithms for the problem under study. Two real-world applications, namely image classification and face denoising further exemplify the practical superiority of the proposed method. Niannan Xue, Jiankang Deng, Yannis Panagakis, Stefanos Zafeiriou |
AAAI | 4 |
| 2018 | Multi-Attribute Probabilistic Linear Discriminant Analysis for 3D Facial Shapes
Stylianos Moschoglou, Stylianos Ploumpis, Mihalis A. Nicolaou, Stefanos Zafeiriou |
ACCV (3) | 4 |
| 2018 | Stacked Dense U-Nets with Dual Transformers for Robust Face Alignment
Jia Guo 0003, Jiankang Deng, Niannan Xue, Stefanos Zafeiriou |
BMVC | 4 |
| 2018 | 4DFAB: A Large Scale 4D Database for Facial Expression Analysis and Biometric ApplicationsabstractThe progress we are currently witnessing in many computer vision applications, including automatic face analysis, would not be made possible without tremendous efforts in collecting and annotating large scale visual databases. To this end, we propose 4DFAB, a new large scale database of dynamic high-resolution 3D faces (over 1,800,000 3D meshes). 4DFAB contains recordings of 180 subjects captured in four different sessions spanning over a five-year period. It contains 4D videos of subjects displaying both spontaneous and posed facial behaviours. The database can be used for both face and facial expression recognition, as well as behavioural biometrics. It can also be used to learn very powerful blendshapes for parametrising facial behaviour. In this paper, we conduct several experiments and demonstrate the usefulness of the database for various applications. The database will be made publicly available for research purposes. Shiyang Cheng 0001, Irene Kotsia, Maja Pantic, Stefanos Zafeiriou |
CVPR | 4 |
| 2018 | UV-GAN: Adversarial Facial UV Map Completion for Pose-Invariant Face RecognitionabstractRecently proposed robust 3D face alignment methods establish either dense or sparse correspondence between a 3D face model and a 2D facial image. The use of these methods presents new challenges as well as opportunities for facial texture analysis. In particular, by sampling the image using the fitted model, a facial UV can be created. Unfortunately, due to self-occlusion, such a UV map is always incomplete. In this paper, we propose a framework for training Deep Convolutional Neural Network (DCNN) to complete the facial UV map extracted from in-the-wild images. To this end, we first gather complete UV maps by fitting a 3D Morphable Model (3DMM) to various multiview image and video datasets, as well as leveraging on a new 3D dataset with over 3,000 identities. Second, we devise a meticulously designed architecture that combines local and global adversarial DCNNs to learn an identity-preserving facial UV completion model. We demonstrate that by attaching the completed UV to the fitted mesh and generating instances of arbitrary poses, we can increase pose variations for training deep face recognition/verification models, and minimise pose discrepancy during testing, which lead to better performance. Experiments on both controlled and in-the-wild UV datasets prove the effectiveness of our adversarial UV completion model. We achieve state-of-the-art verification accuracy, 94.05%, under the CFP frontal-profile protocol only by combining pose augmentation during training and pose discrepancy reduction during testing. We will release the first in-the-wild UV dataset (we refer as WildUV) that comprises of complete facial UV maps from 1,892 identities for research purposes. Jiankang Deng, Shiyang Cheng 0001, Niannan Xue, Stefanos Zafeiriou |
CVPR | 5 |
| 2018 | Cascade Multi-View Hourglass Model for Robust 3D Face AlignmentabstractEstimating the 3D facial landmarks from a 2D image remains a challenging problem. Even though state-of-the-art 2D alignment methods are able to predict accurate landmarks for semi-frontal faces, the majority of them fail to provide semantically consistent landmarks for profile faces. A de facto solution to this problem is through 3D face alignment that preserves correspondence across different poses. In this paper, we proposed a Cascade Multi-view Hourglass Model for 3D face alignment, where the first Hourglass model is explored to jointly predict semi-frontal and profile 2D facial landmarks, after removing spatial transformations, another Hourglass model is employed to estimate the 3D facial shapes. To improve the capacity without sacrificing the computational complexity, the original residual bottleneck block in the Hourglass model is replaced by a parallel, multi-scale inception-resnet block. Extensive experiments on two challenging 3D face alignment datasets, AFLW2000-3D and Menpo-3D, show the robustness of the proposed method under continuous pose changes. Jiankang Deng, Shiyang Cheng 0001, Stefanos Zafeiriou |
FG | 4 |
| 2018 | Deep and Deformable: Convolutional Mixtures of Deformable Part-Based ModelsabstractDeep Convolutional Neural Networks (DCNNs) are currently the method of choice for tasks such that objects and parts detections. Before the advent of DCNNs the method of choice for part detection in a supervised setting (i.e., when part annotations are available) were strongly supervised Deformable Part-based Models (DPMs) on Histograms of Gradients (HoGs) features. Recently, efforts were made to combine the powerful DCNNs features with DPMs which provide an explicit way to model relation between parts. Nevertheless, none of the proposed methodologies provides a unification of DCNNs with strongly supervised DPMs. In this paper, we propose, to the best of our knowledge, the first methodology that jointly trains a strongly supervised DPM and in the same time learns the optimal DCNN features. The proposed methodology not only exploits the relationship between parts but also contains an inherent mechanism for mining of hard-negatives. We demonstrate the power of the proposed approach in facial landmark detection "in-the-wild" where we provide state-of-the-art results for the problem of facial landmark localisation in standard benchmarks such as 300W and 300VW. Kritaphat Songsri-in, George Trigeorgis, Stefanos Zafeiriou |
FG | 3 |
| 2018 | Improve Accurate Pose Alignment and Action Localization by Dense Pose EstimationabstractIn this work we explore the use of shape-based representations as an auxiliary source of supervision for pose estimation and action recognition. We show that shape-based representations can act as a source of `privileged information' that complements and extends the pure landmark-level annotations. We explore 2D shape-based supervision signals, such as Support Vector Shape. Our experiments show that shape-based supervision signals substantially improve pose alignment accuracy in the form of a cascade architecture. We outperform state-of-the-art methods on the MPII and LSP datasets, while using substantially shallower networks. For action localization in untrimmed videos, our method introduces additional classification signals based on the structured segment networks (SSN) and further improved the performance. To be specific, dense human pose and landmarks localization signals are involved in detection progress. We applied out network to all frames of videos alongside with output from SSN to further improve detection accuracy, especially for pose related and sparsely annotated videos. The method in general achieves state-of-the-art performance on Activity Detection Task for ActivityNet Challenge2017 test set and witnesses remarkable improvement on pose related and sparsely annotated categories e.g. sports. Jiankang Deng, Stefanos Zafeiriou |
FG | 3 |
| 2018 | Training Deep Neural Networks with Different Datasets In-the-wild: The Emotion Recognition ParadigmabstractA novel procedure is presented in this paper, for training a deep convolutional and recurrent neural network, taking into account both the available training data set and some information extracted from similar networks trained with other relevant data sets. This information is included in an extended loss function used for the network training, so that the network can have an improved performance when applied to the other data sets, without forgetting the learned knowledge from the original data set. Facial expression and emotion recognition in-the-wild is the test bed application that is used to demonstrate the improved performance achieved using the proposed approach. In this framework, we provide an experimental study on categorical emotion recognition using datasets from a very recent related emotion recognition challenge. Dimitris Kollias, Stefanos Zafeiriou |
IJCNN | 2 |
| 2018 | The INTERSPEECH 2018 Computational Paralinguistics Challenge: Atypical & Self-Assessed Affect, Crying & Heart BeatsabstractThe INTERSPEECH 2018 Computational Paralinguistics Challenge addresses four different problems for the first time in a research competition under well-defined conditions: In the Atypical Affect Sub-Challenge, four basic emotions annotated in the speech of handicapped subjects have to be classified; in the Self-Assessed Affect Sub-Challenge, valence scores given by the speakers themselves are used for a three-class classification problem; in the Crying Sub-Challenge, three types of infant vocalisations have to be told apart; and in the Heart Beats Sub-Challenge, three different types of heart beats have to be determined.We describe the Sub-Challenges, their conditions, and baseline feature extraction and classifiers, which include data-learnt (supervised) feature representations by end-to-end learning, the 'usual' ComParE and BoAW features, and deep unsupervised representation learning using the AUDEEP toolkit for the first time in the challenge series. Björn W. Schuller, Stefan Steidl, Anton Batliner, Peter B. Marschik, Harald Baumeister, Fengquan Dong, Simone Hantke, Florian B. Pokorny, Eva-Maria Rathner, Katrin D. Bartl-Pokorny, Christa Einspieler, Dajie Zhang, Alice Baird, Shahin Amiriparian, Kun Qian 0003, Zhao Ren, Maximilian Schmitt, Panagiotis Tzirakis, Stefanos Zafeiriou |
INTERSPEECH | 19 |
| 2018 | Large Scale 3D Morphable ModelsabstractWe present large scale facial model (LSFM)-a 3D Morphable Model (3DMM) automatically constructed from 9663 distinct facial identities. To the best of our knowledge LSFM is the largest-scale Morphable Model ever constructed, containing statistical information from a huge variety of the human population. To build such a large model we introduce a novel fully automated and robust Morphable Model construction pipeline, informed by an evaluation of state-of-the-art dense correspondence techniques. The dataset that LSFM is trained on includes rich demographic information about each subject, allowing for the construction of not only a global 3DMM model but also models tailored for specific age, gender or ethnicity groups. We utilize the proposed model to perform age classification from 3D shape alone and to reconstruct noisy out-of-sample data in the low-dimensional model space. Furthermore, we perform a systematic analysis of the constructed 3DMM models that showcases their quality and descriptive power. The presented extensive qualitative and quantitative evaluations reveal that the proposed 3DMM achieves state-of-the-art results, outperforming existing models by a large margin. Finally, for the benefit of the research community, we make publicly available the source code of the proposed automatic 3DMM construction pipeline, as well as the constructed global 3DMM and a variety of bespoke models tailored by age, gender and ethnicity. James Booth 0001, Anastasios Roussos, Allan Ponniah, David J. Dunaway, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 5 |
| 2018 | A Comprehensive Performance Evaluation of Deformable Face Tracking "In-the-Wild"abstractRecently, technologies such as face detection, facial landmark localisation and face recognition and verification have matured enough to provide effective and efficient solutions for imagery captured under arbitrary conditions (referred to as "in-the-wild"). This is partially attributed to the fact that comprehensive "in-the-wild" benchmarks have been developed for face detection, landmark localisation and recognition/verification. A very important technology that has not been thoroughly evaluated yet is deformable face tracking "in-the-wild". Until now, the performance has mainly been assessed qualitatively by visually assessing the result of a deformable face tracking technology on short videos. In this paper, we perform the first, to the best of our knowledge, thorough evaluation of state-of-the-art deformable face tracking pipelines using the recently introduced 300 VW benchmark. We evaluate many different architectures focusing mainly on the task of on-line deformable face tracking. In particular, we compare the following general strategies: (a) generic face detection plus generic facial landmark localisation, (b) generic model free tracking plus generic facial landmark localisation, as well as (c) hybrid approaches using state-of-the-art face detection, model free tracking and facial landmark localisation technologies. Our evaluation reveals future avenues for further research on the topic. Grigorios Chrysos 0002, Epameinondas Antonakos, Patrick Snape, Akshay Asthana, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 5 |
| 2018 | 3D Reconstruction of "In-the-Wild" Faces in Images and Videosabstract3D Morphable Models (3DMMs) are powerful statistical models of 3D facial shape and texture, and are among the state-of-the-art methods for reconstructing facial shape from single images. With the advent of new 3D sensors, many 3D facial datasets have been collected containing both neutral as well as expressive faces. However, all datasets are captured under controlled conditions. Thus, even though powerful 3D facial shape models can be learnt from such data, it is difficult to build statistical texture models that are sufficient to reconstruct faces captured in unconstrained conditions ("in-the-wild"). In this paper, we propose the first "in-the-wild" 3DMM by combining a statistical model of facial identity and expression shape with an "in-the-wild" texture model. We show that such an approach allows for the development of a greatly simplified fitting procedure for images and videos, as there is no need to optimise with regards to the illumination parameters. We have collected three new benchmarks that combine "in-the-wild" images and video with ground truth 3D facial geometry, the first of their kind, and report extensive quantitative evaluations using them that demonstrate our method is state-of-the-art. James Booth 0001, Anastasios Roussos, Evangelos Ververas, Epameinondas Antonakos, Stylianos Ploumpis, Yannis Panagakis, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2018 | PD2T: Person-Specific Detection, Deformable TrackingabstractFace detection/alignment methods have reached a satisfactory state in static images captured under arbitrary conditions. Such methods typically perform (joint) fitting for each frame and are used in commercial applications; however in the majority of the real-world scenarios the dynamic scenes are of interest. We argue that generic fitting per frame is suboptimal (it discards the informative correlation of sequential frames) and propose to learn person-specific statistics from the video to improve the generic results. To that end, we introduce a meticulously studied pipeline, which we name PD2T, that performs person-specific detection and landmark localisation. We carry out extensive experimentation with a diverse set of i) generic fitting results, ii) different objects (human faces, animal faces) that illustrate the powerful properties of our proposed pipeline and experimentally verify that PD2T outperforms all the compared methods. Grigorios Chrysos 0002, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Recovering Joint and Individual Components in Facial DataabstractA set of images depicting faces with different expressions or in various ages consists of components that are shared across all images (i.e., joint components) imparting to the depicted object the properties of human faces as well as individual components that are related to different expressions or age groups. Discovering the common (joint) and individual components in facial images is crucial for applications such as facial expression transfer and age progression. The problem is rather challenging when dealing with images captured in unconstrained conditions in the presence of sparse non-Gaussian errors of large magnitude (i.e., sparse gross errors or outliers) and contain missing data. In this paper, we investigate the use of a method recently introduced in statistics, the so-called Joint and Individual Variance Explained (JIVE) method, for the robust recovery of joint and individual components in visual facial data consisting of an arbitrary number of views. Since the JIVE is not robust to sparse gross errors, we propose alternatives, which are (1) robust to sparse gross, non-Gaussian noise, (2) able to automatically find the individual components rank, and (3) can handle missing data. We demonstrate the effectiveness of the proposed methods to several computer vision applications, namely facial expression synthesis and 2D and 3D face age progression 'in-the-wild'. Christos Sagonas, Evangelos Ververas, Yannis Panagakis, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Deep Canonical Time Warping for Simultaneous Alignment and Representation Learning of SequencesabstractMachine learning algorithms for the analysis of time-series often depend on the assumption that utilised data are temporally aligned. Any temporal discrepancies arising in the data is certain to lead to ill-generalisable models, which in turn fail to correctly capture properties of the task at hand. The temporal alignment of time-series is thus a crucial challenge manifesting in a multitude of applications. Nevertheless, the vast majority of algorithms oriented towards temporal alignment are either applied directly on the observation space or simply utilise linear projections-thus failing to capture complex, hierarchical non-linear representations that may prove beneficial, especially when dealing with multi-modal data (e.g., visual and acoustic information). To this end, we present Deep Canonical Time Warping (DCTW), a method that automatically learns non-linear representations of multiple time-series that are (i) maximally correlated in a shared subspace, and (ii) temporally aligned. Furthermore, we extend DCTW to a supervised setting, where during training, available labels can be utilised towards enhancing the alignment process. By means of experiments on four datasets, we show that the representations learnt significantly outperform state-of-the-art methods in temporal alignment, elegantly handling scenarios with heterogeneous feature sets, such as the temporal alignment of acoustic and visual information. George Trigeorgis, Mihalis A. Nicolaou, Björn W. Schuller, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Disentangling the Modes of Variation in Unlabelled DataabstractStatistical methods are of paramount importance in discovering the modes of variation in visual data. The Principal Component Analysis (PCA) is probably the most prominent method for extracting a single mode of variation in the data. However, in practice, several factors contribute to the appearance of visual objects including pose, illumination, and deformation, to mention a few. To extract these modes of variations from visual data, several supervised methods, such as the TensorFaces relying on multilinear (tensor) decomposition have been developed. The main drawbacks of such methods is that they require both labels regarding the modes of variations and the same number of samples under all modes of variations (e.g., the same face under different expressions, poses etc.). Therefore, their applicability is limited to well-organised data, usually captured in well-controlled conditions. In this paper, we propose a novel general multilinear matrix decomposition method that discovers the multilinear structure of possibly incomplete sets of visual data in unsupervised setting (i.e., without the presence of labels). We also propose extensions of the method with sparsity and low-rank constraints in order to handle noisy data, captured in unconstrained conditions. Besides that, a graph-regularised variant of the method is also developed in order to exploit available geometric or label information for some modes of variations. We demonstrate the applicability of the proposed method in several computer vision tasks, including Shape from Shading (SfS) (in the wild and with occlusion removal), expression transfer, and estimation of surface normals from images captured in the wild. Mengjiao Wang 0002, Yannis Panagakis, Patrick Snape, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | IPST: Incremental Pictorial Structures for Model-Free Tracking of Deformable ObjectsabstractModel-free tracking is a well-studied task in computer vision. Typically, a rectangular bounding box containing a single object is provided in the first (few) frame(s) and then the method tracks the object in the rest frames. However, for deformable objects (e.g. faces, bodies) the single bounding box scenario is sub-optimal; a part-based approach would be more effective. The current state-of-the-art part-based approach is incrementally trained discriminative Deformable Part Models (DPM). Nevertheless, training discriminative DPMs with one or a few examples poses a huge challenge. We argue that a generative model is a better fit for the task. We utilise the powerful pictorial structures, which we augment with incremental updates to account for object adaptations. Our proposed incremental pictorial structures, which we call IPST, are experimentally validated in different scenarios. In a thorough experimentation we demonstrate that IPST outperforms the existing model-free methods in facial landmark tracking, body tracking, animal tracking (newly introduced to verify the strength in ad hoc cases). Grigorios Chrysos 0002, Epameinondas Antonakos, Stefanos Zafeiriou |
IEEE Trans. Image Process. | 3 |
| 2017 | 3D Face Morphable Models "In-the-Wild"abstract3D Morphable Models (3DMMs) are powerful statistical models of 3D facial shape and texture, and among the state-of-the-art methods for reconstructing facial shape from single images. With the advent of new 3D sensors, many 3D facial datasets have been collected containing both neutral as well as expressive faces. However, all datasets are captured under controlled conditions. Thus, even though powerful 3D facial shape models can be learnt from such data, it is difficult to build statistical texture models that are sufficient to reconstruct faces captured in unconstrained conditions (in-the-wild). In this paper, we propose the first, to the best of our knowledge, in-the-wild 3DMM by combining a powerful statistical model of facial shape, which describes both identity and expression, with an in-the-wild texture model. We show that the employment of such an in-the-wild texture model greatly simplifies the fitting procedure, because there is no need to optimise with regards to the illumination parameters. Furthermore, we propose a new fast algorithm for fitting the 3DMM in arbitrary images. Finally, we have captured the first 3D facial database with relatively unconstrained conditions and report quantitative evaluations with state-of-the-art performance. Complementary qualitative reconstruction results are demonstrated on standard in-the-wild facial databases. James Booth 0001, Epameinondas Antonakos, Stylianos Ploumpis, George Trigeorgis, Yannis Panagakis, Stefanos Zafeiriou |
CVPR | 6 |
| 2017 | DenseReg: Fully Convolutional Dense Shape Regression In-the-WildabstractIn this paper we propose to learn a mapping from image pixels into a dense template grid through a fully convolutional network. We formulate this task as a regression problem and train our network by leveraging upon manually annotated facial landmarks "in-the-wild". We use such landmarks to establish a dense correspondence field between a three-dimensional object template and the input image, which then serves as the ground-truth for training our regression system. We show that we can combine ideas from semantic segmentation with regression networks, yielding a highly-accurate quantized regression architecture. Our system, called DenseReg, allows us to estimate dense image-to-template correspondences in a fully convolutional manner. As such our network can provide useful correspondence information as a stand-alone system, while when used as an initialization for Statistical Deformable Models we obtain landmark localization results that largely outperform the current state-of-the-art on the challenging 300W benchmark. We thoroughly evaluate our method on a host of facial analysis tasks, and demonstrate its use for other correspondence estimation tasks, such as the human body and the human ear. DenseReg code is made available at http://alpguler.com/DenseReg.html along with supplementary materials. Riza Alp Güler, George Trigeorgis, Epameinondas Antonakos, Patrick Snape, Stefanos Zafeiriou, Iasonas Kokkinos |
CVPR | 5 |
| 2017 | Robust Joint and Individual Variance ExplainedabstractDiscovering the common (joint) and individual subspaces is crucial for analysis of multiple data sets, including multi-view and multi-modal data. Several statistical machine learning methods have been developed for discovering the common features across multiple data sets. The most well studied family of the methods is that of Canonical Correlation Analysis (CCA) and its variants. Even though the CCA is a powerful tool, it has several drawbacks that render its application challenging for computer vision applications. That is, it discovers only common features and not individual ones, and it is sensitive to gross errors present in visual data. Recently, efforts have been made in order to develop methods that discover individual and common components. Nevertheless, these methods are mainly applicable in two sets of data. In this paper, we investigate the use of a recently proposed statistical method, the so-called Joint and Individual Variance Explained (JIVE) method, for the recovery of joint and individual components in an arbitrary number of data sets. Since, the JIVE is not robust to gross errors, we propose alternatives, which are both robust to non-Gaussian noise of large magnitude, as well as able to automatically find the rank of the individual components. We demonstrate the effectiveness of the proposed approach to two computer vision applications, namely facial expression synthesis and face age progression in-the-wild. Christos Sagonas, Yannis Panagakis, Alina Leidinger, Stefanos Zafeiriou |
CVPR | 4 |
| 2017 | Face Normals "In-the-Wild" Using Fully Convolutional NetworksabstractIn this work we pursue a data-driven approach to the problem of estimating surface normals from a single intensity image, focusing in particular on human faces. We introduce new methods to exploit the currently available facial databases for dataset construction and tailor a deep convolutional neural network to the task of estimating facial surface normals in-the-wild. We train a fully convolutional network that can accurately recover facial normals from images including a challenging variety of expressions and facial poses. We compare against state-of-the-art face Shape-from-Shading and 3D reconstruction techniques and show that the proposed network can recover substantially more accurate and realistic normals. Furthermore, in contrast to other existing face-specific surface recovery methods, we do not require the solving of an explicit alignment step due to the fully convolutional nature of our network. George Trigeorgis, Patrick Snape, Iasonas Kokkinos, Stefanos Zafeiriou |
CVPR | 4 |
| 2017 | Learning the Multilinear Structure of Visual DataabstractStatistical decomposition methods are of paramount importance in discovering the modes of variations of visual data. Probably the most prominent linear decomposition method is the Principal Component Analysis (PCA), which discovers a single mode of variation in the data. However, in practice, visual data exhibit several modes of variations. For instance, the appearance of faces varies in identity, expression, pose etc. To extract these modes of variations from visual data, several supervised methods, such as the TensorFaces, that rely on multilinear (tensor) decomposition (e.g., Higher Order SVD) have been developed. The main drawbacks of such methods is that they require both labels regarding the modes of variations and the same number of samples under all modes of variations (e.g., the same face under different expressions, poses etc.). Therefore, their applicability is limited to well-organised data, usually captured in well-controlled conditions. In this paper, we propose the first general multilinear method, to the best of our knowledge, that discovers the multilinear structure of visual data in unsupervised setting. That is, without the presence of labels. We demonstrate the applicability of the proposed method in two applications, namely Shape from Shading (SfS) and expression transfer. Mengjiao Wang 0002, Yannis Panagakis, Patrick Snape, Stefanos Zafeiriou |
CVPR | 4 |
| 2017 | A Joint Discriminative Generative Model for Deformable Model Construction and ClassificationabstractDiscriminative classification models have been successfully applied for various computer vision tasks such as object and face detection and recognition. However, deformations can change objects coordinate space and perturb robust similarity measurement, which is the essence of all classification algorithms. The common approach to deal with deformations is either to seek for deformation invariant features or to develop models that describe objects deformations. However, the former approach requires a huge amount of data and a good amount of engineering to be properly trained, while the latter require considerable human effort in the form of carefully annotated data. In this paper, we propose a method that jointly learns with minimal human intervention a generative deformable model using only a simple shape model of the object and images automatically downloaded from the Internet, and also extracts features appropriate for classification. The proposed algorithm is applied on various classification problems such as “in-thewild” face recognition, gender classification and eye glasses detection on data retrieved by querying into a web image search engine. We demonstrate that not only it outperforms other automatic methods by large margins, but also performs comparably with supervised methods trained on thousands of manually annotated data. Ioannis Marras, Symeon Nikitidis, Stefanos Zafeiriou, Maja Pantic |
FG | 3 |
| 2017 | Deformable Models of Ears in-the-Wild for Alignment and RecognitionabstractEars have been discovered to have biometric importance for identifying people and/or verifying their identity. This is largely because of their complex inner shape structure, which is not only unique but also long-lasting regardless of ageing. In this paper, we make two important contributions relevant to analysis of ear in imagery captured in unconstrained conditions. That is, we present (a) the first, to the best of our knowledge, annotated database with ear landmarks and use it in order to build statistical deformable ear models in-the-wild and (b) a database of 2058 labelled unconstrained ear images with 231 subjects and use it for ear recognition/verification. We perform extensive comparisons for ear alignment using many state-of-the-art techniques and extensive experiments. Finally, we conducted extensive experiments for ear recognition using both handcrafted, as well as learned features (i.e., using deep learning). All annotated data and code will be publicly available. Stefanos Zafeiriou |
FG | 2 |
| 2017 | Dynamic Probabilistic Linear Discriminant Analysis for video classificationabstractComponent Analysis (CA) comprises of statistical techniques that decompose signals into appropriate latent components, relevant to a task-at-hand (e.g., clustering, segmentation, classification). Recently, an explosion of research in CA has been witnessed, with several novel probabilistic models proposed (e.g., Probabilistic Principal CA, Probabilistic Linear Discriminant Analysis (PLDA), Probabilistic Canonical Correlation Analysis). PLDA is a popular generative probabilistic CA method, that incorporates knowledge regarding class-labels and furthermore introduces class-specific and sample-specific latent spaces. While PLDA has been shown to outperform several state-of-the-art methods, it is nevertheless a static model; any feature-level temporal dependencies that arise in the data are ignored. As has been repeatedly shown, appropriate modelling of temporal dynamics is crucial for the analysis of temporal data (e.g., videos). In this light, we propose the first, to the best of our knowledge, probabilistic LDA formulation that models dynamics, the so-called Dynamic-PLDA (DPLDA). DPLDA is a generative model suitable for video classification and is able to jointly model the label information (e.g., face identity, consistent over videos of the same subject), as well as dynamic variations of each individual video. Experiments on video classification tasks such as face and facial expression recognition show the efficacy of the proposed method. Alessandro Fabris, Mihalis A. Nicolaou, Irene Kotsia, Stefanos Zafeiriou |
ICASSP | 4 |
| 2017 | The unconstrained ear recognition challengeabstractIn this paper we present the results of the Unconstrained Ear Recognition Challenge (UERC), a group benchmarking effort centered around the problem of person recognition from ear images captured in uncontrolled conditions. The goal of the challenge was to assess the performance of existing ear recognition techniques on a challenging large-scale dataset and identify open problems that need to be addressed in the future. Five groups from three continents participated in the challenge and contributed six ear recognition techniques for the evaluation, while multiple baselines were made available for the challenge by the UERC organizers. A comprehensive analysis was conducted with all participating approaches addressing essential research questions pertaining to the sensitivity of the technology to head rotation, flipping, gallery size, large-scale recognition and others. The top performer of the UERC was found to ensure robust performance on a smaller part of the dataset (with 180 subjects) regardless of image characteristics, but still exhibited a significant performance drop when the entire dataset comprising 3,704 subjects was used for testing. Ziga Emersic, Dejan Stepec, Vitomir Struc, Peter Peer, Anjith George, Adil M. Ahmad, Elshibani Omar, Terrance E. Boult, Reza Safdari, Stefanos Zafeiriou, Doggucan Yaman, Fevziye Irem Eyiokur, Hazim Kemal Ekenel |
IJCB | 11 |
| 2017 | Robust Kronecker-Decomposable Component Analysis for Low-Rank Modeling
Mehdi Bahri, Yannis Panagakis, Stefanos Zafeiriou |
ICCV | 3 |
| 2017 | Side Information in Robust Principal Component Analysis: Algorithms and ApplicationsabstractRobust Principal Component Analysis (RPCA) aims at recovering a low-rank subspace from grossly corrupted high-dimensional (often visual) data and is a cornerstone in many machine learning and computer vision applications. Even though RPCA has been shown to be very successful in solving many rank minimisation problems, there are still cases where degenerate or suboptimal solutions are obtained. This is likely to be remedied by taking into account of domain-dependent prior knowledge. In this paper, we propose two models for the RPCA problem with the aid of side information on the low-rank structure of the data. The versatility of the proposed methods is demonstrated by applying them to four applications, namely background subtraction, facial image denoising, face and facial expression recognition. Experimental results on synthetic and five real world datasets indicate the robustness and effectiveness of the proposed methods on these application domains, largely outperforming six previous approaches. Niannan Xue, Yannis Panagakis, Stefanos Zafeiriou |
ICCV | 3 |
| 2017 | The INTERSPEECH 2017 Computational Paralinguistics Challenge: Addressee, Cold & SnoringabstractThe INTERSPEECH 2017 Computational Paralinguistics Challenge addresses three different problems for the first time in research competition under well-defined conditions: In the Addressee sub-challenge, it has to be determined whether speech produced by an adult is directed towards another adult or towards a child; in the Cold sub-challenge, speech under cold has to be told apart from ‘healthy’ speech; and in the Snoring subchallenge, four different types of snoring have to be classified. In this paper, we describe these sub-challenges, their conditions, and the baseline feature extraction and classifiers, which include data-learnt feature representations by end-to-end learning with convolutional and recurrent neural networks, and bag-of-audiowords for the first time in the challenge series Björn W. Schuller, Stefan Steidl, Anton Batliner, Elika Bergelson, Jarek Krajewski, Christoph Janott, Andrei Amatuni, Marisa Casillas, Amanda Seidl, Melanie Soderstrom, Anne S. Warlaumont, Guillermo Hidalgo, Sebastian Schnieder, Clemens Heiser, Winfried Hohenhorst, Michael Herzog, Maximilian Schmitt, Kun Qian 0003, Yue Zhang 0014, George Trigeorgis, Panagiotis Tzirakis, Stefanos Zafeiriou |
INTERSPEECH | 22 |
| 2017 | A Unified Framework for Compositional Fitting of Active Appearance ModelsabstractActive appearance models (AAMs) are one of the most popular and well-established techniques for modeling deformable objects in computer vision. In this paper, we study the problem of fitting AAMs using compositional gradient descent (CGD) algorithms. We present a unified and complete view of these algorithms and classify them with respect to three main characteristics: (i) cost function; (ii) type of composition; and (iii) optimization method. Furthermore, we extend the previous view by: (a) proposing a novel Bayesian cost function that can be interpreted as a general probabilistic formulation of the well-known project-out loss; (b) introducing two new types of composition, asymmetric and bidirectional, that combine the gradients of both image and appearance model to derive better convergent and more robust CGD algorithms; and (c) providing new valuable insights into existent CGD algorithms by reinterpreting them as direct applications of the Schur complement and the Wiberg method. Finally, in order to encourage open research and facilitate future comparisons with our work, we make the implementation of the algorithms studied in this paper publicly available as part of the Menpo Project ( http://www.menpo.org ). Joan Alabort-i-Medina, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 2 |
| 2017 | Robust Statistical Frontalization of Human and Animal FacesabstractThe unconstrained acquisition of facial data in real-world conditions may result in face images with significant pose variations, illumination changes, and occlusions, affecting the performance of facial landmark localization and recognition methods. In this paper, a novel method, robust to pose, illumination variations, and occlusions is proposed for joint face frontalization and landmark localization. Unlike the state-of-the-art methods for landmark localization and pose correction, where large amount of manually annotated images or 3D facial models are required, the proposed method relies on a small set of frontal images only. By observing that the frontal facial image of both humans and animals, is the one having the minimum rank of all different poses, a model which is able to jointly recover the frontalized version of the face as well as the facial landmarks is devised. To this end, a suitable optimization problem is solved, concerning minimization of the nuclear norm (convex surrogate of the rank function) and the matrix $$\ell _1$$ norm accounting for occlusions. The proposed method is assessed in frontal view reconstruction of human and animal faces, landmark localization, pose-invariant face recognition, face verification in unconstrained conditions, and video inpainting by conducting experiment on 9 databases. The experimental results demonstrate the effectiveness of the proposed method in comparison to the state-of-the-art methods for the target problems. Christos Sagonas, Yannis Panagakis, Stefanos Zafeiriou, Maja Pantic |
Int. J. Comput. Vis. | 3 |
| 2017 | Statistical non-rigid ICP algorithm and its application to 3D face alignment
Shiyang Cheng 0001, Ioannis Marras, Stefanos Zafeiriou, Maja Pantic |
Image Vis. Comput. | 3 |
| 2017 | The Conflict Escalation Resolution (CONFER) Database
Christos Georgakis 0001, Yannis Panagakis, Stefanos Zafeiriou, Maja Pantic |
Image Vis. Comput. | 3 |
| 2017 | A Deep Matrix Factorization Method for Learning Attribute RepresentationsabstractSemi-Non-negative Matrix Factorization is a technique that learns a low-dimensional representation of a dataset that lends itself to a clustering interpretation. It is possible that the mapping between this new representation and our original data matrix contains rather complex hierarchical information with implicit lower-level hidden attributes, that classical one level clustering methodologies cannot interpret. In this work we propose a novel model, Deep Semi-NMF, that is able to learn such hidden representations that allow themselves to an interpretation of clustering according to different, unknown attributes of a given dataset. We also present a semi-supervised version of the algorithm, named Deep WSF, that allows the use of (partial) prior information for each of the known attributes of a dataset, that allows the model to be used on datasets with mixed attribute knowledge. Finally, we show that our models are able to learn low-dimensional representations that are better suited for clustering, but also classification, outperforming Semi-Non-negative Matrix Factorization, but also other state-of-the-art methodologies variants. George Trigeorgis, Konstantinos Bousmalis, Stefanos Zafeiriou, Björn W. Schuller |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Nonnegative Decompositions for Dynamic Visual Data AnalysisabstractThe analysis of high-dimensional, possibly temporally misaligned, and time-varying visual data is a fundamental task in disciplines, such as image, vision, and behavior computing. In this paper, we focus on dynamic facial behavior analysis and in particular on the analysis of facial expressions. Distinct from the previous approaches, where sets of facial landmarks are used for face representation, raw pixel intensities are exploited for: 1) unsupervised analysis of the temporal phases of facial expressions and facial action units (AUs) and 2) temporal alignment of a certain facial behavior displayed by two different persons. To this end, the slow features nonnegative matrix factorization (SFNMF) is proposed in order to learn slow varying parts-based representations of time varying sequences capturing the underlying dynamics of temporal phenomena, such as facial expressions. Moreover, the SFNMF is extended in order to handle two temporally misaligned data sequences depicting the same visual phenomena. To do so, the dynamic time warping is incorporated into the SFNMF, allowing the temporal alignment of the data sets onto the subspace spanned by the estimated nonnegative shared latent features amongst the two visual sequences. Extensive experimental results in two video databases demonstrate the effectiveness of the proposed methods in: 1) unsupervised detection of the temporal phases of posed and spontaneous facial events and 2) temporal alignment of facial expressions, outperforming by a large margin the state-of-the-art methods that they are compared to. Lazaros Zafeiriou, Yannis Panagakis, Maja Pantic, Stefanos Zafeiriou |
IEEE Trans. Image Process. | 4 |
| 2016 | A 3D Morphable Model Learnt from 10, 000 FacesabstractWe present Large Scale Facial Model (LSFM) - a 3D Morphable Model (3DMM) automatically constructed from 9,663 distinct facial identities. To the best of our knowledge LSFM is the largest-scale Morphable Model ever constructed, containing statistical information from a huge variety of the human population. To build such a large model we introduce a novel fully automated and robust Morphable Model construction pipeline. The dataset that LSFM is trained on includes rich demographic information about each subject, allowing for the construction of not only a global 3DMM but also models tailored for specific age, gender or ethnicity groups. As an application example, we utilise the proposed model to perform age classification from 3D shape alone. Furthermore, we perform a systematic analysis of the constructed 3DMMs that showcases their quality and descriptive power. The presented extensive qualitative and quantitative evaluations reveal that the proposed 3DMM achieves state-of-the-art results, outperforming existing models by a large margin. Finally, for the benefit of the research community, we make publicly available the source code of the proposed automatic 3DMM construction pipeline. In addition, the constructed global 3DMM and a variety of bespoke models tailored by age, gender and ethnicity are available on application to researchers involved in medically oriented research. James Booth 0001, Anastasios Roussos, Stefanos Zafeiriou, Allan Ponniah, David J. Dunaway |
CVPR | 3 |
| 2016 | Deep Canonical Time WarpingabstractMachine learning algorithms for the analysis of timeseries often depend on the assumption that the utilised data are temporally aligned. Any temporal discrepancies arising in the data is certain to lead to ill-generalisable models, which in turn fail to correctly capture the properties of the task at hand. The temporal alignment of time-series is thus a crucial challenge manifesting in a multitude of applications. Nevertheless, the vast majority of algorithms oriented towards the temporal alignment of time-series are applied directly on the observation space, or utilise simple linear projections. Thus, they fail to capture complex, hierarchical non-linear representations which may prove to be beneficial towards the task of temporal alignment, particularly when dealing with multi-modal data (e.g., aligning visual and acoustic information). To this end, we present the Deep Canonical Time Warping (DCTW), a method which automatically learns complex non-linear representations of multiple time-series, generated such that (i) they are highly correlated, and (ii) temporally in alignment. By means of experiments on four real datasets, we show that the representations learnt via the proposed DCTW significantly outperform state-of-the-art methods in temporal alignment, elegantly handling scenarios with highly heterogeneous features, such as the temporal alignment of acoustic and visual features. George Trigeorgis, Mihalis A. Nicolaou, Stefanos Zafeiriou, Björn W. Schuller |
CVPR | 3 |
| 2016 | Mnemonic Descent Method: A Recurrent Process Applied for End-to-End Face AlignmentabstractCascaded regression has recently become the method of choice for solving non-linear least squares problems such as deformable image alignment. Given a sizeable training set, cascaded regression learns a set of generic rules that are sequentially applied to minimise the least squares problem. Despite the success of cascaded regression for problems such as face alignment and head pose estimation, there are several shortcomings arising in the strategies proposed thus far. Specifically, (a) the regressors are learnt independently, (b) the descent directions may cancel one another out and (c) handcrafted features (e.g., HoGs, SIFT etc.) are mainly used to drive the cascade, which may be sub-optimal for the task at hand. In this paper, we propose a combined and jointly trained convolutional recurrent neural network architecture that allows the training of an end-to-end to system that attempts to alleviate the aforementioned drawbacks. The recurrent module facilitates the joint optimisation of the regressors by assuming the cascades form a nonlinear dynamical system, in effect fully utilising the information between all cascade levels by introducing a memory unit that shares information across all levels. The convolutional module allows the network to extract features that are specialised for the task at hand and are experimentally shown to outperform hand-crafted features. We show that the application of the proposed architecture for the problem of face alignment results in a strong improvement over the current state-of-the-art. George Trigeorgis, Patrick Snape, Mihalis A. Nicolaou, Epameinondas Antonakos, Stefanos Zafeiriou |
CVPR | 5 |
| 2016 | Joint Unsupervised Deformable Spatio-Temporal Alignment of SequencesabstractTypically, the problems of spatial and temporal alignment of sequences are considered disjoint. That is, in order to align two sequences, a methodology that (non)-rigidly aligns the images is first applied, followed by temporal alignment of the obtained aligned images. In this paper, we propose the first, to the best of our knowledge, methodology that can jointly spatio-temporally align two sequences, which display highly deformable texture-varying objects. We show that by treating the problems of deformable spatial and temporal alignment jointly, we achieve better results than considering the problems independent. Furthermore, we show that deformable spatio-temporal alignment of faces can be performed in an unsupervised manner (i.e., without employing face trackers or building person-specific deformable models). Lazaros Zafeiriou, Epameinondas Antonakos, Stefanos Zafeiriou, Maja Pantic |
CVPR | 3 |
| 2016 | Estimating Correspondences of Deformable Objects "In-the-Wild"abstractDuring the past few years we have witnessed the development of many methodologies for building and fitting Statistical Deformable Models (SDMs). The construction of accurate SDMs requires careful annotation of images with regards to a consistent set of landmarks. However, the manual annotation of a large amount of images is a tedious, laborious and expensive procedure. Furthermore, for several deformable objects, e.g. human body, it is difficult to define a consistent set of landmarks, and, thus, it becomes impossible to train humans in order to accurately annotate a collection of images. Nevertheless, for the majority of objects, it is possible to extract the shape by object segmentation or even by shape drawing. In this paper, we show for the first time, to the best of our knowledge, that it is possible to construct SDMs by putting object shapes in dense correspondence. Such SDMs can be built with much less effort for a large battery of objects. Additionally, we show that, by sampling the dense model, a part-based SDM can be learned with its parts being in correspondence. We employ our framework to develop SDMs of human arms and legs, which can be used for the segmentation of the outline of the human body, as well as to provide better and more consistent annotations for body joints. Epameinondas Antonakos, Joan Alabort-i-Medina, Anastasios Roussos, Stefanos Zafeiriou |
CVPR | 5 |
| 2016 | Fine-Grained Material Classification Using Micro-geometry and Reflectance
Christos Kampouris, Stefanos Zafeiriou, Abhijeet Ghosh, Sotiris Malassiotis |
ECCV (5) | 2 |
| 2016 | Adieu features? End-to-end speech emotion recognition using a deep convolutional recurrent networkabstractThe automatic recognition of spontaneous emotions from speech is a challenging task. On the one hand, acoustic features need to be robust enough to capture the emotional content for various styles of speaking, and while on the other, machine learning algorithms need to be insensitive to outliers while being able to model the context. Whereas the latter has been tackled by the use of Long Short-Term Memory (LSTM) networks, the former is still under very active investigations, even though more than a decade of research has provided a large set of acoustic descriptors. In this paper, we propose a solution to the problem of ‘context-aware’ emotional relevant feature extraction, by combining Convolutional Neural Networks (CNNs) with LSTM networks, in order to automatically learn the best representation of the speech signal directly from the raw time representation. In this novel work on the so-called end-to-end speech emotion recognition, we show that the use of the proposed topology significantly outperforms the traditional approaches based on signal processing techniques for the prediction of spontaneous and natural emotions on the RECOLA database. George Trigeorgis, Fabien Ringeval, Raymond Brueckner, Erik Marchi, Mihalis A. Nicolaou, Björn W. Schuller, Stefanos Zafeiriou |
ICASSP | 7 |
| 2016 | Adaptive cascaded regressionabstractThe two predominant families of deformable models for the task of face alignment are: (i) discriminative cascaded regression models, and (ii) generative models optimised with Gauss-Newton. Although these approaches have been found to work well in practise, they each suffer from convergence issues. Cascaded regression has no theoretical guarantee of convergence to a local minimum and thus may fail to recover the fine details of the object. Gauss-Newton optimisation is not robust to initialisations that are far from the optimal solution. In this paper, we propose the first, to the best of our knowledge, attempt to combine the best of these two worlds under a unified model and report state-of-the-art performance on the most recent facial benchmark challenge. Epameinondas Antonakos, Patrick Snape, George Trigeorgis, Stefanos Zafeiriou |
ICIP | 4 |
| 2016 | Fusion and community detection in multi-layer graphsabstractRelational data arising in many domains can be represented by networks (or graphs) with nodes capturing entities and edges representing relationships between these entities. Community detection in networks has become one of the most important problems having a broad range of applications. Until recently, the vast majority of papers have focused on discovering community structures in a single network. However, with the emergence of multi-view network data in many real-world applications and consequently with the advent of multilayer graph representation, community detection in multi-layer graphs has become a new challenge. Multi-layer graphs provide complementary views of connectivity patterns of the same set of vertices. Fusion of the network layers is expected to achieve better clustering performance. In this paper, we propose two novel methods, coined as WSSNMTF (Weighted Simultaneous Symmetric Non-Negative Matrix Tri-Factorization) and NG-WSSNMTF (Natural Gradient WSSNMTF), for fusion and clustering of multi-layer graphs. Both methods are robust with respect to missing edges and noise. We compare the performance of the proposed methods with two baseline methods, as well as with three state-of-the-art methods on synthetic and three real-world datasets. The experimental results indicate superior performance of the proposed methods. Vladimir Gligorijevic, Yannis Panagakis, Stefanos Zafeiriou |
ICPR | 3 |
| 2016 | Back to the future: A fully automatic method for robust age progressionabstractIt has been shown that significant age difference between a probe and gallery face image can decrease the matching accuracy. If the face images can be normalized in age, there can be a huge impact on the face verification accuracy and thus many novel applications such as matching driver's license, passport and visa images with the real person's images can be effectively implemented. Face progression can address this issue by generating a face image for a specific age. Many researchers have attempted to address this problem focusing on predicting older faces from a younger face. In this paper, we propose a novel method for robust and automatic face progression in totally unconstrained conditions. Our method takes into account that faces belonging to the same age-groups share age patterns such as wrinkles while faces across different age-groups share some common patterns such as expressions and skin colors. Given training images of K different age-groups the proposed method learns to recover K low-rank age and one low-rank common components. These extracted components from the learning phase are used to progress an input face to younger as well as older ages in bidirectional fashion. Using standard datasets, we demonstrate that the proposed progression method outperforms state-of-the-art age progression methods and also improves matching accuracy in a face verification protocol that includes age progression. Christos Sagonas, Yannis Panagakis, Saritha Arunkumar, Nalini K. Ratha, Stefanos Zafeiriou |
ICPR | 5 |
| 2016 | Editorial of special issue on spontaneous facial behaviour analysis
Stefanos Zafeiriou, Guoying Zhao 0001, Matti Pietikäinen, Rama Chellappa, Irene Kotsia, Jeffrey F. Cohn |
Comput. Vis. Image Underst. | 1 |
| 2016 | Capturing correlations of local features for image representation
Xiaopeng Hong, Guoying Zhao 0001, Stefanos Zafeiriou, Maja Pantic, Matti Pietikäinen |
Neurocomputing | 3 |
| 2016 | 300 Faces In-The-Wild Challenge: database and results
Christos Sagonas, Epameinondas Antonakos, Georgios Tzimiropoulos, Stefanos Zafeiriou, Maja Pantic |
Image Vis. Comput. | 4 |
| 2016 | A robust similarity measure for volumetric image registration with outliersabstractImage registration under challenging realistic conditions is a very important area of research. In this paper, we focus on algorithms that seek to densely align two volumetric images according to a global similarity measure. Despite intensive research in this area, there is still a need for similarity measures that are robust to outliers common to many different types of images. For example, medical image data is often corrupted by intensity inhomogeneities and may contain outliers in the form of pathologies. In this paper we propose a global similarity measure that is robust to both intensity inhomogeneities and outliers without requiring prior knowledge of the type of outliers. We combine the normalised gradients of images with the cosine function and show that it is theoretically robust against a very general class of outliers. Experimentally, we verify the robustness of our measures within two distinct algorithms. Firstly, we embed our similarity measures within a proof-of-concept extension of the Lucas–Kanade algorithm for volumetric data. Finally, we embed our measures within a popular non-rigid alignment framework based on free-form deformations and show it to be robust against both simulated tumours and intensity inhomogeneities. Patrick Snape, Stefan Pszczólkowski, Stefanos Zafeiriou, Georgios Tzimiropoulos, Christian Ledig, Daniel Rueckert |
Image Vis. Comput. | 3 |
| 2016 | 300 W: Special issue on facial landmark localisation "in-the-wild"
Stefanos Zafeiriou, Georgios Tzimiropoulos, Maja Pantic |
Image Vis. Comput. | 1 |
| 2016 | Robust Correlated and Individual Component AnalysisabstractRecovering correlated and individual components of two, possibly temporally misaligned, sets of data is a fundamental task in disciplines such as image, vision, and behavior computing, with application to problems such as multi-modal fusion (via correlated components), predictive analysis, and clustering (via the individual ones). Here, we study the extraction of correlated and individual components under real-world conditions, namely i) the presence of gross non-Gaussian noise and ii) temporally misaligned data. In this light, we propose a method for the Robust Correlated and Individual Component Analysis (RCICA) of two sets of data in the presence of gross, sparse errors. We furthermore extend RCICA in order to handle temporal incongruities arising in the data. To this end, two suitable optimization problems are solved. The generality of the proposed methods is demonstrated by applying them onto 4 applications, namely i) heterogeneous face recognition, ii) multi-modal feature fusion for human behavior analysis (i.e., audio-visual prediction of interest and conflict), iii) face clustering, and iv) thetemporal alignment of facial expressions. Experimental results on 2 synthetic and 7 real world datasets indicate the robustness and effectiveness of the proposed methodson these application domains, outperforming other state-of-the-art methods in the field. Yannis Panagakis, Mihalis A. Nicolaou, Stefanos Zafeiriou, Maja Pantic |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | MAGMA: Multilevel Accelerated Gradient Mirror Descent Algorithm for Large-Scale Convex Composite MinimizationabstractComposite convex optimization models arise in several applications and are especially prevalent in inverse problems with a sparsity inducing norm and in general convex optimization with simple constraints. The most widely used algorithms for convex composite models are accelerated first order methods; however, they can take a large number of iterations to compute an acceptable solution for large-scale problems. In this paper we propose speeding up first order methods by taking advantage of the structure present in many applications and in image processing in particular. Our method is based on multilevel optimization methods and exploits the fact that many applications that give rise to large-scale models can be modeled using varying degrees of fidelity. We use Nesterov's acceleration techniques together with the multilevel approach to achieve an $\mathcal{O}(1/\sqrt{\epsilon})$ convergence rate, where $\epsilon$ denotes the desired accuracy. The proposed method has a better convergence rate than any other existing multilevel method for convex problems and in addition has the same rate as accelerated methods, which is known to be optimal for first order methods. Moreover, as our numerical experiments show, on large-scale face recognition problems our algorithm is several times faster than the state of the art. Vahan Hovhannisyan, Panos Parpas, Stefanos Zafeiriou |
SIAM J. Imaging Sci. | 3 |
| 2016 | Probabilistic Slow Features for Behavior AnalysisabstractA recently introduced latent feature learning technique for time-varying dynamic phenomena analysis is the so-called slow feature analysis (SFA). SFA is a deterministic component analysis technique for multidimensional sequences that, by minimizing the variance of the first-order time derivative approximation of the latent variables, finds uncorrelated projections that extract slowly varying features ordered by their temporal consistency and constancy. In this paper, we propose a number of extensions in both the deterministic and the probabilistic SFA optimization frameworks. In particular, we derive a novel deterministic SFA algorithm that is able to identify linear projections that extract the common slowest varying features of two or more sequences. In addition, we propose an expectation maximization (EM) algorithm to perform inference in a probabilistic formulation of SFA and similarly extend it in order to handle two and more time-varying data sequences. Moreover, we demonstrate that the probabilistic SFA (EM-SFA) algorithm that discovers the common slowest varying latent space of multiple sequences can be combined with dynamic time warping techniques for robust sequence time-alignment. The proposed SFA algorithms were applied for facial behavior analysis, demonstrating their usefulness and appropriateness for this task. Lazaros Zafeiriou, Mihalis A. Nicolaou, Stefanos Zafeiriou, Symeon Nikitidis, Maja Pantic |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Unifying holistic and Parts-Based Deformable Model fittingabstractThe construction and fitting of deformable models that capture the degrees of freedom of articulated objects is one of the most popular areas of research in computer vision. Two of the most popular approaches are: Holistic Deformable Models (HDMs), which try to represent the object as a whole, and Parts-Based Deformable Models (PBDMs), which model object parts independently. Both models have been shown to have their own advantages. In this paper we try to marry the previous two approaches into a unified one that potentially combines the advantages of both. We do so by merging the well-established frameworks of Active Appearance Models (holistic) and Constrained Local Models (part-based) using a novel probabilistic formulation of the fitting problem. We show that our unified holistic and part-based formulation achieves state-of-the-art results in the problem of face alignment in-the-wild. Finally, in order to encourage open research and facilitate future comparisons with the proposed method, our code will be made publicly available to the research community. Joan Alabort-i-Medina, Stefanos Zafeiriou |
CVPR | 2 |
| 2015 | Active Pictorial StructuresabstractIn this paper we present a novel generative deformable model motivated by Pictorial Structures (PS) and Active Appearance Models (AAMs) for object alignment in-the-wild. Inspired by the tree structure used in PS, the proposed Active Pictorial Structures (APS)1model the appearance of the object using multiple graph-based pairwise normal distributions (Gaussian Markov Random Field) between the patches extracted from the regions around adjacent landmarks. We show that this formulation is more accurate than using a single multivariate distribution (Principal Component Analysis) as commonly done in the literature. APS employ a weighted inverse compositional Gauss-Newton optimization with fixed Jacobian and Hessian that achieves close to real-time performance and state-of-the-art results. Finally, APS have a spring-like graph-based deformation prior term that makes them robust to bad initializations. We present extensive experiments on the task of face alignment, showing that APS outperform current state-of-the-art methods. To the best of our knowledge, the proposed method is the first weighted inverse compositional technique that proves to be so accurate and efficient at the same time. Epameinondas Antonakos, Joan Alabort-i-Medina, Stefanos Zafeiriou |
CVPR | 3 |
| 2015 | Automatic construction Of robust spherical harmonic subspacesabstractIn this paper we propose a method to automatically recover a class specific low dimensional spherical harmonic basis from a set of in-the-wild facial images. We combine existing techniques for uncalibrated photometric stereo and low rank matrix decompositions in order to robustly recover a combined model of shape and identity. We build this basis without aid from a 3D model and show how it can be combined with recent efficient sparse facial feature localisation techniques to recover dense 3D facial shape. Unlike previous works in the area, our method is very efficient and is an order of magnitude faster to train, taking only a few minutes to build a model with over 2000 images. Furthermore, it can be used for real-time recovery of facial shape. Patrick Snape, Yannis Panagakis, Stefanos Zafeiriou |
CVPR | 3 |
| 2015 | Robust Statistical Face FrontalizationabstractRecently, it has been shown that excellent results can be achieved in both facial landmark localization and pose-invariant face recognition. These breakthroughs are attributed to the efforts of the community to manually annotate facial images in many different poses and to collect 3D facial data. In this paper, we propose a novel method for joint frontal view reconstruction and landmark localization using a small set of frontal images only. By observing that the frontal facial image is the one having the minimum rank of all different poses, an appropriate model which is able to jointly recover the frontalized version of the face as well as the facial landmarks is devised. To this end, a suitable optimization problem, involving the minimization of the nuclear norm and the matrix l1 norm is solved. The proposed method is assessed in frontal face reconstruction, face landmark localization, pose-invariant face recognition, and face verification in unconstrained conditions. The relevant experiments have been conducted on 8 databases. The experimental results demonstrate the effectiveness of the proposed method in comparison to the state-of-the-art methods for the target problems. Christos Sagonas, Yannis Panagakis, Stefanos Zafeiriou, Maja Pantic |
ICCV | 3 |
| 2015 | Face FlowabstractIn this paper, we propose a method for the robust and efficient computation of multi-frame optical flow in an expressive sequence of facial images. We formulate a novel energy minimisation problem for establishing dense correspondences between a neutral template and every frame of a sequence. We exploit the highly correlated nature of human expressions by representing dense facial motion using a deformation basis. Furthermore, we exploit the even higher correlation between deformations in a given input sequence by imposing a low-rank prior on the coefficients of the deformation basis, yielding temporally consistent optical flow. Our proposed model-based formulation, in conjunction with the inverse compositional strategy and low-rank matrix optimisation that we adopt, leads to a highly efficient algorithm for calculating facial flow. As experimental evaluation, we show quantitative experiments on a challenging novel benchmark of face sequences, with dense ground truth optical flow provided by motion capture data. We also provide qualitative results on a real sequence displaying fast motion and occlusions. Extensive quantitative and qualitative comparisons demonstrate that the proposed method outperforms state-of-the-art optical flow and dense non-rigid registration techniques, whilst running an order of magnitude faster. Patrick Snape, Anastasios Roussos, Yannis Panagakis, Stefanos Zafeiriou |
ICCV | 4 |
| 2015 | A survey on face detection in the wild: Past, present and future
Stefanos Zafeiriou, Cha Zhang, Zhengyou Zhang |
Comput. Vis. Image Underst. | 1 |
| 2015 | From Pixels to Response Maps: Discriminative Image Filtering for Face Alignment in the WildabstractWe propose a face alignment framework that relies on the texture model generated by the responses of discriminatively trained part-based filters. Unlike standard texture models built from pixel intensities or responses generated by generic filters (e.g. Gabor), our framework has two important advantages. First, by virtue of discriminative training, invariance to external variations (like identity, pose, illumination and expression) is achieved. Second, we show that the responses generated by discriminatively trained filters (or patch-experts) are sparse and can be modeled using a very small number of parameters. As a result, the optimization methods based on the proposed texture model can better cope with unseen variations. We illustrate this point by formulating both part-based and holistic approaches for generic face alignment and show that our framework outperforms the state-of-the-art on multiple "wild" databases. The code and dataset annotations are available for research purposes from http://ibug.doc.ic.ac.uk/resources. Akshay Asthana, Stefanos Zafeiriou, Georgios Tzimiropoulos, Shiyang Cheng 0001, Maja Pantic |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Variational Infinite Hidden Conditional Random FieldsabstractHidden conditional random fields (HCRFs) are discriminative latent variable models which have been shown to successfully learn the hidden structure of a given classification problem. An Infinite hidden conditional random field is a hidden conditional random field with a countably infinite number of hidden states, which rids us not only of the necessity to specify a priori a fixed number of hidden states available but also of the problem of overfitting. Markov chain Monte Carlo (MCMC) sampling algorithms are often employed for inference in such models. However, convergence of such algorithms is rather difficult to verify, and as the complexity of the task at hand increases the computational cost of such algorithms often becomes prohibitive. These limitations can be overcome by variational techniques. In this paper, we present a generalized framework for infinite HCRF models, and a novel variational inference approach on a model based on coupled Dirichlet Process Mixtures, the HCRF-DPM. We show that the variational HCRF-DPM is able to converge to a correct number of represented hidden states, and performs as well as the best parametric HCRFs-chosen via cross-validation-for the difficult tasks of recognizing instances of agreement, disagreement, and pain in audiovisual sequences. Konstantinos Bousmalis, Stefanos Zafeiriou, Louis-Philippe Morency, Maja Pantic, Zoubin Ghahramani |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Feature-Based Lucas-Kanade and Active Appearance ModelsabstractLucas-Kanade and active appearance models are among the most commonly used methods for image alignment and facial fitting, respectively. They both utilize nonlinear gradient descent, which is usually applied on intensity values. In this paper, we propose the employment of highly descriptive, densely sampled image features for both problems. We show that the strategy of warping the multichannel dense feature image at each iteration is more beneficial than extracting features after warping the intensity image at each iteration. Motivated by this observation, we demonstrate robust and accurate alignment and fitting performance using a variety of powerful feature descriptors. Especially with the employment of histograms of oriented gradient and scale-invariant feature transform features, our method significantly outperforms the current state-of-the-art results on in-the-wild databases. Epameinondas Antonakos, Joan Alabort-i-Medina, Georgios Tzimiropoulos, Stefanos Zafeiriou |
IEEE Trans. Image Process. | 4 |
| 2015 | Online Kernel Slow Feature Analysis for Temporal Video Segmentation and TrackingabstractSlow feature analysis (SFA) is a dimensionality reduction technique which has been linked to how visual brain cells work. In recent years, the SFA was adopted for computer vision tasks. In this paper, we propose an exact kernel SFA (KSFA) framework for positive definite and indefinite kernels in Krein space. We then formulate an online KSFA which employs a reduced set expansion. Finally, by utilizing a special kind of kernel family, we formulate exact online KSFA for which no reduced set is required. We apply the proposed system to develop a SFA-based change detection algorithm for stream data. This framework is employed for temporal video segmentation and tracking. We test our setup on synthetic and real data streams. When combined with an online learning tracking system, the proposed change detection approach improves upon tracking setups that do not utilize change detection. Stephan Liwicki, Stefanos Zafeiriou, Maja Pantic |
IEEE Trans. Image Process. | 2 |
| 2014 | Generalised Scalable Robust Principal Component Analysis
Georgios Papamakarios, Yannis Panagakis, Stefanos Zafeiriou |
BMVC | 3 |
| 2014 | Bayesian Active Appearance ModelsabstractIn this paper we provide the first, to the best of our knowledge, Bayesian formulation of one of the most successful and well-studied statistical models of shape and texture, i.e. Active Appearance Models (AAMs). To this end, we use a simple probabilistic model for texture generation assuming both Gaussian noise and a Gaussian prior over a latent texture space. We retrieve the shape parameters by formulating a novel cost function obtained by marginalizing out the latent texture space. This results in a fast implementation when compared to other simultaneous algorithms for fitting AAMs, mainly due to the removal of the calculation of texture parameters. We demonstrate that, contrary to what is believed regarding the performance of AAMs in generic fitting scenarios, optimization of the proposed cost function produces results that outperform discriminatively trained state-of-the-art methods in the problem of facial alignment "in the wild". Joan Alabort-i-Medina, Stefanos Zafeiriou |
CVPR | 2 |
| 2014 | Automatic Construction of Deformable Models In-the-WildabstractDeformable objects are everywhere. Faces, cars, bicycles, chairs etc. Recently, there has been a wealth of research on training deformable models for object detection, part localization and recognition using annotated data. In order to train deformable models with good generalization ability, a large amount of carefully annotated data is required, which is a highly time consuming and costly task. We propose the first - to the best of our knowledge - method for automatic construction of deformable models using images captured in totally unconstrained conditions, recently referred to as "in-the-wild". The only requirements of the method are a crude bounding box object detector and a-priori knowledge of the object's shape (e.g. a point distribution model). The object detector can be as simple as the Viola-Jones algorithm (e.g. even the cheapest digital camera features a robust face detector). The 2D shape model can be created by using only a few shape examples with deformations. In our experiments on facial deformable models, we show that the proposed automatically built model not only performs well, but also outperforms discriminative models trained on carefully annotated data. To the best of our knowledge, this is the first time it is shown that an automatically constructed model can perform as well as methods trained directly on annotated data. Epameinondas Antonakos, Stefanos Zafeiriou |
CVPR | 2 |
| 2014 | Incremental Face Alignment in the WildabstractThe development of facial databases with an abundance of annotated facial data captured under unconstrained 'in-the-wild' conditions have made discriminative facial deformable models the de facto choice for generic facial landmark localization. Even though very good performance for the facial landmark localization has been shown by many recently proposed discriminative techniques, when it comes to the applications that require excellent accuracy, such as facial behaviour analysis and facial motion capture, the semi-automatic person-specific or even tedious manual tracking is still the preferred choice. One way to construct a person-specific model automatically is through incremental updating of the generic model. This paper deals with the problem of updating a discriminative facial deformable model, a problem that has not been thoroughly studied in the literature. In particular, we study for the first time, to the best of our knowledge, the strategies to update a discriminative model that is trained by a cascade of regressors. We propose very efficient strategies to update the model and we show that is possible to automatically construct robust discriminative person and imaging condition specific models 'in-the-wild' that outperform state-of-the-art generic face alignment strategies. Akshay Asthana, Stefanos Zafeiriou, Shiyang Cheng 0001, Maja Pantic |
CVPR | 2 |
| 2014 | Full-Angle Quaternions for Robustly Matching Vectors of 3D RotationsabstractIn this paper we introduce a new distance for robustly matching vectors of 3D rotations. A special representation of 3D rotations, which we coin full-angle quaternion (FAQ), allows us to express this distance as Euclidean. We apply the distance to the problems of 3D shape recognition from point clouds and 2D object tracking in color video. For the former, we introduce a hashing scheme for scale and translation which outperforms the previous state-of-the-art approach on a public dataset. For the latter, we incorporate online subspace learning with the proposed FAQ representation to highlight the benefits of the new representation. Stephan Liwicki, Minh-Tri Pham, Stefanos Zafeiriou, Maja Pantic, Björn Stenger |
CVPR | 3 |
| 2014 | Merging SVMs with Linear Discriminant Analysis: A Combined ModelabstractA key problem often encountered by many learning algorithms in computer vision dealing with high dimensional data is the so called "curse of dimensionality" which arises when the available training samples are less than the input feature space dimensionality. To remedy this problem, we propose a joint dimensionality reduction and classification framework by formulating an optimization problem within the maximum margin class separation task. The proposed optimization problem is solved using alternative optimization where we jointly compute the low dimensional maximum margin projections and the separating hyperplanes in the projection subspace. Moreover, in order to reduce the computational cost of the developed optimization algorithm we incorporate orthogonality constraints on the derived projection bases and show that the resulting combined model is an alternation between identifying the optimal separating hyperplanes and performing a linear discriminant analysis on the support vectors. Experiments on face, facial expression and object recognition validate the effectiveness of the proposed method against state-of-the-art dimensionality reduction algorithms. Symeon Nikitidis, Stefanos Zafeiriou, Maja Pantic |
CVPR | 2 |
| 2014 | RAPS: Robust and Efficient Automatic Construction of Person-Specific Deformable ModelsabstractThe construction of Facial Deformable Models (FDMs) is a very challenging computer vision problem, since the face is a highly deformable object and its appearance drastically changes under different poses, expressions, and illuminations. Although several methods for generic FDMs construction, have been proposed for facial landmark localization in still images, they are insufficient for tasks such as facial behaviour analysis and facial motion capture where perfect landmark localization is required. In this case, person-specific FDMs (PSMs) are mainly employed, requiring manual facial landmark annotation for each person and person-specific training. In this paper, a novel method for the automatic construction of PSMs is proposed. To this end, an orthonormal subspace which is suitable for facial image reconstruction is learnt. Next, to correct the fittings of a generic model, image congealing (i.e., batch image aliment) is performed by employing only the learnt orthonormal subspace. Finally, the corrected fittings are used to construct the PSM. The image congealing problem is solved by formulating a suitable sparsity regularized rank minimization problem. The proposed method outperforms the state-of-the art methods that is compared to, in terms of both landmark localization accuracy and computational time. Christos Sagonas, Yannis Panagakis, Stefanos Zafeiriou, Maja Pantic |
CVPR | 3 |
| 2014 | Kernel-PCA Analysis of Surface Normals for Shape-from-ShadingabstractWe propose a kernel-based framework for computing components from a set of surface normals. This framework allows us to easily demonstrate that component analysis can be performed directly upon normals. We link previously proposed mapping functions, the azimuthal equidistant projection (AEP) and principal geodesic analysis (PGA), to our kernel-based framework. We also propose a new mapping function based upon the cosine distance between normals. We demonstrate the robustness of our proposed kernel when trained with noisy training sets. We also compare our kernels within an existing shape-from-shading (SFS) algorithm. Our spherical representation of normals, when combined with the robust properties of cosine kernel, produces a very robust subspace analysis technique. In particular, our results within SFS show a substantial qualitative and quantitative improvement over existing techniques. Patrick Snape, Stefanos Zafeiriou |
CVPR | 2 |
| 2014 | Joint Unsupervised Face Alignment and Behaviour Analysis
Lazaros Zafeiriou, Epameinondas Antonakos, Stefanos Zafeiriou, Maja Pantic |
ECCV (4) | 3 |
| 2014 | Robust Canonical Correlation Analysis: Audio-visual fusion for learning continuous interestabstractThe problem of automatically estimating the interest level of a subject has been gaining attention by researchers, mostly due to the vast applicability of interest detection. In this work, we obtain a set of continuous interest annotations for the SE-MAINE database, which we analyse also in terms of emotion dimensions such as valence and arousal. Most importantly, we propose a robust variant of Canonical Correlation Analysis (RCCA) for performing audio-visual fusion, which we apply to the prediction of interest. RCCA recovers a low-rank subspace which captures the correlations of fused modalities, while isolating gross errors in the data without making any assumptions regarding Gaussianity. We experimentally show that RCCA is more appropriate than other standard fusion techniques (such as l2-CCA and feature-level fusion), since it both captures interactions between modalities while also decontaminating the obtained subspace from errors which are dominant in real-world problems. Mihalis A. Nicolaou, Yannis Panagakis, Stefanos Zafeiriou, Maja Pantic |
ICASSP | 3 |
| 2014 | HOG active appearance modelsabstractWe propose the combination of dense Histogram of Oriented Gradients (HOG) features with Active Appearance Models (AAMs). We employ the efficient Inverse Compositional optimization technique and show results for the task of face fitting. By taking advantage of the descriptive characteristics of HOG features, we build robust and accurate AAMs that generalize well to unseen faces with illumination, identity, pose and occlusion variations. Our experiments on challenging in-the-wild databases show that HOG AAMs significantly outperfrom current state-of-the-art results of discriminative methods trained on larger databases. Epameinondas Antonakos, Joan Alabort-i-Medina, Georgios Tzimiropoulos, Stefanos Zafeiriou |
ICIP | 4 |
| 2014 | Optimal UV spaces for facial morphable model constructionabstractEstablishing inter-mesh dense correspondence is a key step in the process of constructing Morphable Models. The most successful approaches to date reduce this 3D correspondence problem to a 2D image morphing one by applying an interpolant in UV space - a space in which the manifold of the face is flattened into a contiguous 2D atlas. Contiguous UV spaces are natural products of the laser scanning devices popular for use in Morphable Model construction, but a wide gamut of devices can now be used to capture 3D data that do not yield a UV space representation conducive to warping. In this paper we explore how to optimally construct contiguous UV spaces from both annotated and non-annotated facial meshes with efficient cylindrical and spherical projection operations. James Booth 0001, Stefanos Zafeiriou |
ICIP | 2 |
| 2014 | 3D facial geometric features for constrained local modelabstractWe propose a 3D Constrained Local Model framework for deformable face alignment in depth image. Our framework exploits the intrinsic 3D geometric information in depth data by utilizing robust histogram-based 3D geometric features that are based on normal vectors. In addition, we demonstrate the fusion of intensity data and 3D features that further improves the facial landmark localization accuracy. The experiments are conducted on publicly available FRGC database. The results show that our 3D features based CLM completely outperforms the raw depth features based CLM in term of fitting accuracy and robustness, and the fusion of intensity and 3D depth feature further improves the performance. Another benefit is that the proposed 3D features in our framework do not require any pre-processing procedure on the data. Shiyang Cheng 0001, Stefanos Zafeiriou, Akshay Asthana, Maja Pantic |
ICIP | 2 |
| 2014 | Slow features nonnegative matrix factorization for temporal data decompositionabstractIn this paper, we combine the principles of temporal slowness and nonnegative parts-based learning into a single framework that aims to learn slow varying parts-based representations of time varying sequences. We demonstrate that the proposed algorithm arises naturally by embedding the Slow Features Analysis trace optimization problem in the nonnegative subspace learning framework and derive novel multiplicative update rules for its optimization. The usefulness of the developed algorithm is demonstrated for unsupervised facial behaviour dynamics analysis on MMI database. Lazaros Zafeiriou, Symeon Nikitidis, Stefanos Zafeiriou, Maja Pantic |
ICIP | 3 |
| 2014 | A Deep Semi-NMF Model for Learning Hidden RepresentationsabstractSemi-NMF is a matrix factorization technique that learns a low-dimensional representation of a dataset that lends itself to a clustering interpretation. It is possible that the mapping between this new representation and our original features contains rather complex hierarchical information with implicit lower-level hidden attributes, that classical one level clustering methodologies can not interpret. In this work we propose a novel model, Deep Semi-NMF, that is able to learn such hidden representations that allow themselves to an interpretation of clustering according to different, unknown attributes of a given dataset. We show that by doing so, our model is able to learn low-dimensional representations that are better suited for clustering, outperforming Semi-NMF, but also other NMF variants. George Trigeorgis, Konstantinos Bousmalis, Stefanos Zafeiriou, Björn W. Schuller |
ICML | 3 |
| 2014 | Menpo: A Comprehensive Platform for Parametric Image Alignment and Visual Deformable ModelsabstractThe Menpo Project, hosted at http://www.menpo.io, is a BSD-licensed software platform providing a complete and comprehensive solution for annotating, building, fitting and evaluating deformable visual models from image data. Menpo is a powerful and flexible cross-platform framework written in Python that works on Linux, OS X and Windows. Menpo has been designed to allow for easy adaptation of Lucas-Kanade (LK) parametric image alignment techniques, and goes a step further in providing all the necessary tools for building and fitting state-of-the-art deformable models such as Active Appearance Models (AAMs), Constrained Local Models (CLMs) and regression-based methods (such as the Supervised Descent Method (SDM)). These methods are extensively used for facial point localisation although they can be applied to many other deformable objects. Menpo makes it easy to understand and evaluate these complex algorithms, providing tools for visualisation, analysis, and performance assessment. A key challenge in building deformable models is data annotation; Menpo expedites this process by providing a simple web-based annotation tool hosted at http://www.landmarker.io. The Menpo Project is thoroughly documented and provides extensive examples for all of its features. We believe the project is ideal for researchers, practitioners and students alike. Joan Alabort-i-Medina, Epameinondas Antonakos, James Booth 0001, Patrick Snape, Stefanos Zafeiriou |
ACM Multimedia | 5 |
| 2014 | Real-time generic face tracking in the wild with CUDAabstractWe present a robust real-time face tracking system based on the Constrained Local Models framework by adopting the novel regression-based Discriminative Response Map Fitting (DRMF) method. By exploiting the algorithm's potential parallelism, we present a hybrid CPU-GPU implementation capable of achieving real-time performance at 30 to 45 FPS, on ordinary consumer-grade computers. We have made the software publicly available for research purposes Shiyang Cheng 0001, Akshay Asthana, Stefanos Zafeiriou, Jie Shen 0008, Maja Pantic |
MMSys | 3 |
| 2014 | A Unified Framework for Probabilistic Component Analysis
Mihalis A. Nicolaou, Stefanos Zafeiriou, Maja Pantic |
ECML/PKDD (2) | 2 |
| 2014 | Optimal illumination directions for faces and rough surfaces for single and multiple light imaging using class-specific prior knowledge
Vasileios Argyriou, Stefanos Zafeiriou, Maria Petrou |
Comput. Vis. Image Underst. | 2 |
| 2014 | Online learning and fusion of orientation appearance models for robust rigid object tracking
Ioannis Marras, Georgios Tzimiropoulos, Stefanos Zafeiriou, Maja Pantic |
Image Vis. Comput. | 3 |
| 2014 | Special issue on "Multi-biometrics and Mobile-biometrics: Recent Advances and Future Research"
Lei Zhang 0006, Tieniu Tan, Arun Ross, Stefanos Zafeiriou |
Image Vis. Comput. | 4 |
| 2014 | Active Orientation Models for Face Alignment In-the-WildabstractWe present Active Orientation Models (AOMs), generative models of facial shape and appearance, which extend the well-known paradigm of Active Appearance Models (AAMs) for the case of generic face alignment under unconstrained conditions. Robustness stems from the fact that the proposed AOMs employ a statistically robust appearance model based on the principal components of image gradient orientations. We show that when incorporated within standard optimization frameworks for AAM learning and fitting, this kernel Principal Component Analysis results in robust algorithms for model fitting. At the same time, the resulting optimization problems maintain the same computational cost. As a result, the main similarity of AOMs with AAMs is the computational complexity. In particular, the project-out version of AOMs is as computationally efficient as the standard project-out inverse compositional algorithm, which is admittedly one of the fastest algorithms for fitting AAMs. We verify experimentally that: 1) AOMs generalize well to unseen variations and 2) outperform all other state-of-the-art AAM methods considered by a large margin. This performance improvement brings AOMs at least in par with other contemporary methods for face alignment. Finally, we provide MATLAB code at http://ibug.doc.ic.ac.uk/resources. Georgios Tzimiropoulos, Joan Alabort-i-Medina, Stefanos Zafeiriou, Maja Pantic |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2014 | Principal Component Analysis With Complex Kernel: The Widely Linear ModelabstractNonlinear complex representations, via the use of complex kernels, can be applied to model and capture the nonlinearities of complex data. Even though the theoretical tools of complex reproducing kernel Hilbert spaces (CRKHS) have been recently successfully applied to the design of digital filters and regression and classification frameworks, there is a limited research on component analysis and dimensionality reduction in CRKHS. The aim of this brief is to properly formulate the most popular component analysis methodology, i.e., Principal Component Analysis (PCA), in CRKHS. In particular, we define a general widely linear complex kernel PCA framework. Furthermore, we show how to efficiently perform widely linear PCA in small sample sized problems. Finally, we show the usefulness of the proposed framework in robust reconstruction using Euler data representation. Athanasios Papaioannou, Stefanos Zafeiriou |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | Robust Discriminative Response Map Fitting with Constrained Local ModelsabstractWe present a novel discriminative regression based approach for the Constrained Local Models (CLMs) framework, referred to as the Discriminative Response Map Fitting (DRMF) method, which shows impressive performance in the generic face fitting scenario. The motivation behind this approach is that, unlike the holistic texture based features used in the discriminative AAM approaches, the response map can be represented by a small set of parameters and these parameters can be very efficiently used for reconstructing unseen response maps. Furthermore, we show that by adopting very simple off-the-shelf regression techniques, it is possible to learn robust functions from response maps to the shape parameters updates. The experiments, conducted on Multi-PIE, XM2VTS and LFPW database, show that the proposed DRMF method outperforms state-of-the-art algorithms for the task of generic face fitting. Moreover, the DRMF method is computationally very efficient and is real-time capable. The current MATLAB implementation takes 1 second per image. To facilitate future comparisons, we release the MATLAB code and the pre-trained models for research purposes. Akshay Asthana, Stefanos Zafeiriou, Shiyang Cheng 0001, Maja Pantic |
CVPR | 2 |
| 2013 | Robust Canonical Time Warping for the Alignment of Grossly Corrupted SequencesabstractTemporal alignment of human behaviour from visual data is a very challenging problem due to a numerous reasons, including possible large temporal scale differences, inter/intra subject variability and, more importantly, due to the presence of gross errors and outliers. Gross errors are often in abundance due to incorrect localization and tracking, presence of partial occlusion etc. Furthermore, such errors rarely follow a Gaussian distribution, which is the de-facto assumption in machine learning methods. In this paper, building on recent advances on rank minimization and compressive sensing, a novel, robust to gross errors temporal alignment method is proposed. While previous approaches combine the dynamic time warping (DTW) with low-dimensional projections that maximally correlate two sequences, we aim to learn two underlying projection matrices (one for each sequence), which not only maximally correlate the sequences but, at the same time, efficiently remove the possible corruptions in any datum in the sequences. The projections are obtained by minimizing the weighted sum of nuclear and ℓ1norms, by solving a sequence of convex optimization problems, while the temporal alignment is found by applying the DTW in an alternating fashion. The superiority of the proposed method against the state-of-the-art time alignment methods, namely the canonical time warping and the generalized time warping, is indicated by the experimental results on both synthetic and real datasets. Yannis Panagakis, Mihalis A. Nicolaou, Stefanos Zafeiriou, Maja Pantic |
CVPR | 3 |
| 2013 | On One-Shot Similarity Kernels: Explicit Feature Maps and PropertiesabstractKernels have been a common tool of machine learning and computer vision applications for modeling non-linearities and/or the design of robust Robustness may refer to either the presence of outliers and noise or to the robustness to a class of transformations (e.g., translation). similarity measures between objects. Arguably, the class of positive semi-definite (psd) kernels, widely known as Mercer's Kernels, constitutes one of the most well-studied cases. For every psd kernel there exists an associated feature map to an arbitrary dimensional Hilbert space mathcal H, the so-called feature space. The main reason behind psd kernels' popularity is the fact that classification/regression techniques (such as Support Vector Machines (SVMs)) and component analysis algorithms (such as Kernel Principal Component Analysis (KPCA)) can be devised in mathcal H, without an explicit definition of the feature map, only by using the kernel (the so-called kernel trick). Recently, due to the development of very efficient solutions for large scale linear SVMs and for incremental linear component analysis, the research towards finding feature map approximations for classes of kernels has attracted significant interest. In this paper, we attempt the derivation of explicit feature maps of a recently proposed class of kernels, the so-called one-shot similarity kernels. We show that for this class of kernels either there exists an explicit representation in feature space or the kernel can be expressed in such a form that allows for exact incremental learning. We theoretically explore the properties of these kernels and show how these kernels can be used for the development of robust visual tracking, recognition and deformable fitting algorithms. Stefanos Zafeiriou, Irene Kotsia |
ICCV | 1 |
| 2013 | Learning Slow Features for Behaviour AnalysisabstractA recently introduced latent feature learning technique for time varying dynamic phenomena analysis is the so called Slow Feature Analysis (SFA). SFA is a deterministic component analysis technique for multi-dimensional sequences that by minimizing the variance of the first order time derivative approximation of the input signal finds uncorrelated projections that extract slowly-varying features ordered by their temporal consistency and constancy. In this paper, we propose a number of extensions in both the deterministic and the probabilistic SFA optimization frameworks. In particular, we derive a novel deterministic SFA algorithm that is able to identify linear projections that extract the common slowest varying features of two or more sequences. In addition, we propose an Expectation Maximization (EM) algorithm to perform inference in a probabilistic formulation of SFA and similarly extend it in order to handle two and more time varying data sequences. Moreover, we demonstrate that the probabilistic SFA (EMSFA) algorithm that discovers the common slowest varying latent space of multiple sequences can be combined with dynamic time warping techniques for robust sequence time alignment. The proposed SFA algorithms were applied for facial behavior analysis demonstrating their usefulness and appropriateness for this task. Lazaros Zafeiriou, Mihalis A. Nicolaou, Stefanos Zafeiriou, Symeon Nikitidis, Maja Pantic |
ICCV | 3 |
| 2013 | Correlated-spaces regression for learning continuous emotion dimensionsabstractAdopting continuous dimensional annotations for affective analysis has been gaining rising attention by researchers over the past years. Due to the idiosyncratic nature of this problem, many subproblems have been identified, spanning from the fusion of multiple continuous annotations to exploiting output-correlations amongst emotion dimensions. In this paper, we firstly empirically answer several important questions which have found partial or no answer at all so far in related literature. In more detail, we study the correlation of each emotion dimension (i) with respect to other emotion dimensions, (ii) to basic emotions (e.g., happiness, anger). As a measure for comparison, we use video and audio features. Interestingly enough, we find that (i) each emotion dimension is more correlated with other emotion dimensions rather than with face and audio features, and similarly (ii) that each basic emotion is more correlated with emotion dimensions than with audio and video features. A similar conclusion holds for discrete emotions which are found to be highly correlated to emotion dimensions as compared to audio and/or video features. Motivated by these findings, we present a novel regression algorithm (Correlated-Spaces Regression, CSR), inspired by Canonical Correlation Analysis (CCA) which learns output-correlations and performs supervised dimensionality reduction and multimodal fusion by (i) projecting features extracted from all modalities and labels onto a common space where their inter-correlation is maximised and (ii) learning mappings from the projected feature space onto the projected, uncorrelated label space. Mihalis A. Nicolaou, Stefanos Zafeiriou, Maja Pantic |
ACM Multimedia | 2 |
| 2013 | Variational Hidden Conditional Random Fields with Coupled Dirichlet Process Mixtures
Konstantinos Bousmalis, Stefanos Zafeiriou, Louis-Philippe Morency, Maja Pantic, Zoubin Ghahramani |
ECML/PKDD (2) | 2 |
| 2013 | Euler Principal Component AnalysisabstractPrincipal Component Analysis (PCA) is perhaps the most prominent learning tool for dimensionality reduction in pattern recognition and computer vision. However, the ℓ 2-norm employed by standard PCA is not robust to outliers. In this paper, we propose a kernel PCA method for fast and robust PCA, which we call Euler-PCA (e-PCA). In particular, our algorithm utilizes a robust dissimilarity measure based on the Euler representation of complex numbers. We show that Euler-PCA retains PCA’s desirable properties while suppressing outliers. Moreover, we formulate Euler-PCA in an incremental learning framework which allows for efficient computation. In our experiments we apply Euler-PCA to three different computer vision applications for which our method performs comparably with other state-of-the-art approaches. Stephan Liwicki, Georgios Tzimiropoulos, Stefanos Zafeiriou, Maja Pantic |
Int. J. Comput. Vis. | 3 |
| 2013 | A sparse representation method for determining the optimal illumination directions in Photometric Stereo
Vasileios Argyriou, Stefanos Zafeiriou, Barbara Villarini, Maria Petrou |
Signal Process. | 2 |
| 2013 | High order pLSA for indexing tagged images
Spiros Nikolopoulos, Stefanos Zafeiriou, Ioannis Patras, Ioannis Kompatsiaris |
Signal Process. | 2 |
| 2013 | Guest Editorial Introduction to the Special Issue on Modern Control for Computer GamesabstractA typical gaming scenario, as developed in the past 20 years, involves a player interacting with a game using a specialized input device, such as a joystic, a mouse, a keyboard, etc. Recent technological advances and new sensors (for example, low cost commodity depth cameras) have enabled the introduction of more elaborated approaches in which the player is now able to interact with the game using his body pose, facial expressions, actions, and even his physiological signals. A new era of games has already started, employing computer vision techniques, brain-computer interfaces systems, haptic and wearable devices. The future lies in games that will be intelligent enough not only to extract the player's commands provided by his speech and gestures but also his behavioral cues, as well as his/her emotional states, and adjust their game plot accordingly in order to ensure more realistic and satisfactory gameplay experience. This special issue on modern control for computer games discusses several interdisciplinary factors that influence a user's input to a game, something directly linked to the gaming experience. These include, but are not limited to, the following: behavioral affective gaming, user satisfaction and perception, motion capture and scene modeling, and complete software frameworks that address several challenges risen in such scenarios. Vasileios Argyriou, Irene Kotsia, Stefanos Zafeiriou, Maria Petrou |
IEEE Trans. Cybern. | 3 |
| 2013 | Face Recognition and Verification Using Photometric Stereo: The Photoface Database and a Comprehensive EvaluationabstractThis paper presents a new database suitable for both 2-D and 3-D face recognition based on photometric stereo (PS): the Photoface database. The database was collected using a custom-made four-source PS device designed to enable data capture with minimal interaction necessary from the subjects. The device, which automatically detects the presence of a subject using ultrasound, was placed at the entrance to a busy workplace and captured 1839 sessions of face images with natural pose and expression. This meant that the acquired data is more realistic for everyday use than existing databases and is, therefore, an invaluable test bed for state-of-the-art recognition algorithms. The paper also presents experiments of various face recognition and verification algorithms using the albedo, surface normals, and recovered depth maps. Finally, we have conducted experiments in order to demonstrate how different methods in the pipeline of PS (i.e., normal field computation and depth map reconstruction) affect recognition and verification performance. These experiments help to 1) demonstrate the usefulness of PS, and our device in particular, for minimal-interaction face recognition, and 2) highlight the optimal reconstruction and recognition algorithms for use with natural-expression PS data. The database can be downloaded from http://www.uwe.ac.uk/research/Photoface. Stefanos Zafeiriou, Gary A. Atkinson, Mark F. Hansen, William A. P. Smith, Vasileios Argyriou, Maria Petrou, Melvyn L. Smith, Lyndon N. Smith |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2013 | Infinite Hidden Conditional Random Fields for Human Behavior AnalysisabstractHidden conditional random fields (HCRFs) are discriminative latent variable models that have been shown to successfully learn the hidden structure of a given classification problem (provided an appropriate validation of the number of hidden states). In this brief, we present the infinite HCRF (iHCRF), which is a nonparametric model based on hierarchical Dirichlet processes and is capable of automatically learning the optimal number of hidden states for a classification task. We show how we learn the model hyperparameters with an effective Markov-chain Monte Carlo sampling technique, and we explain the process that underlines our iHCRF model with the Restaurant Franchise Rating Agencies analogy. We show that the iHCRF is able to converge to a correct number of represented hidden states, and outperforms the best finite HCRFs--chosen via cross-validation--for the difficult tasks of recognizing instances of agreement, disagreement, and pain. Moreover, the iHCRF manages to achieve this performance in significantly less total training, validation, and testing time. Konstantinos Bousmalis, Stefanos Zafeiriou, Louis-Philippe Morency, Maja Pantic |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2012 | Incremental Slow Feature Analysis with Indefinite Kernel for Online Temporal Video Segmentation
Stephan Liwicki, Stefanos Zafeiriou, Maja Pantic |
ACCV (2) | 2 |
| 2012 | Generic Active Appearance Models Revisited
Georgios Tzimiropoulos, Joan Alabort-i-Medina, Stefanos Zafeiriou, Maja Pantic |
ACCV (3) | 3 |
| 2012 | Binary Pattern Analysis for 3D Facial Action Unit DetectionabstractIn this paper we propose new binary pattern features for use in the problem of 3D facial action unit (AU) detection. Two representations of 3D facial geometries are employed, the depth map and the Azimuthal Projection Distance Image (APDI). To these the traditional Local Binary Pattern is applied, along with Local Phase Quantisation, Gabor filters and Monogenic filters, followed by the binary pattern feature extraction method. Feature vectors are formed for each feature type through concatenation of histograms formed from the resulting binary numbers. Feature selection is then performed using a two-stage GentleBoost approach. Finally, we apply Support Vector Machines as classifiers for detection of each AU. This system is tested in two ways. First we perform 10-fold cross-validation on the Bosphorus database, and then we perform cross-database testing by training on this database and then testing on apex frames from the D3DFACS database, achieving promising results in both. Georgia Sandbach, Stefanos Zafeiriou, Maja Pantic |
BMVC | 2 |
| 2012 | Subspace Learning in Krein Spaces: Complete Kernel Fisher Discriminant Analysis with Indefinite Kernels
Stefanos Zafeiriou |
ECCV (4) | 1 |
| 2012 | Local normal binary patterns for 3D facial action unit detectionabstractThis paper proposes a new feature descriptor, local normal binary patterns (LNBPs), which is exploited for detection of facial action units (AUs). After LNBPs have been employed to form descriptor vectors, which capture the detailed shape of the action, feature selection is performed via a Gentle-Boost (GB) algorithm, and support vector machines (SVMs) are trained to detect each AU. This process was tested on the Bosphorus database, alongside the same test using 3D local binary pattern (3DLBP) descriptors which apply the LBP operator to the depth map of the face. LNBP descriptors were demonstrated to outperform 3DLBPs in detection of many individual AUs. Finally, feature fusion was used to combine the benefits of the 3DLBPs and each of the LNBP descriptors, with the best result achieving a mean ROC AuC of 96.35. Georgia Sandbach, Stefanos Zafeiriou, Maja Pantic |
ICIP | 2 |
| 2012 | Recognition of 3D facial expression dynamics
Georgia Sandbach, Stefanos Zafeiriou, Maja Pantic, Daniel Rueckert |
Image Vis. Comput. | 2 |
| 2012 | Static and dynamic 3D facial expression recognition: A comprehensive survey
Georgia Sandbach, Stefanos Zafeiriou, Maja Pantic, Lijun Yin 0001 |
Image Vis. Comput. | 2 |
| 2012 | 3D facial behaviour analysis and understanding
Stefanos Zafeiriou, Lijun Yin 0001 |
Image Vis. Comput. | 1 |
| 2012 | Subspace Learning from Image Gradient OrientationsabstractWe introduce the notion of subspace learning from image gradient orientations for appearance-based object recognition. As image data are typically noisy and noise is substantially different from Gaussian, traditional subspace learning from pixel intensities very often fails to estimate reliably the low-dimensional subspace of a given data population. We show that replacing pixel intensities with gradient orientations and the ℓ₂ norm with a cosine-based distance measure offers, to some extend, a remedy to this problem. Within this framework, which we coin Image Gradient Orientations (IGO) subspace learning, we first formulate and study the properties of Principal Component Analysis of image gradient orientations (IGO-PCA). We then show its connection to previously proposed robust PCA techniques both theoretically and experimentally. Finally, we derive a number of other popular subspace learning techniques, namely, Linear Discriminant Analysis (LDA), Locally Linear Embedding (LLE), and Laplacian Eigenmaps (LE). Experimental results show that our algorithms significantly outperform popular methods such as Gabor features and Local Binary Patterns and achieve state-of-the-art performance for difficult problems such as illumination and occlusion-robust face recognition. In addition to this, the proposed IGO-methods require the eigendecomposition of simple covariance matrices and are as computationally efficient as their corresponding ℓ₂ norm intensity-based counterparts. Matlab code for the methods presented in this paper can be found at http://ibug.doc.ic.ac.uk/resources. Georgios Tzimiropoulos, Stefanos Zafeiriou, Maja Pantic |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Efficient Online Subspace Learning With an Indefinite Kernel for Visual Tracking and RecognitionabstractWe propose an exact framework for online learning with a family of indefinite (not positive) kernels. As we study the case of nonpositive kernels, we first show how to extend kernel principal component analysis (KPCA) from a reproducing kernel Hilbert space to Krein space. We then formulate an incremental KPCA in Krein space that does not require the calculation of preimages and therefore is both efficient and exact. Our approach has been motivated by the application of visual tracking for which we wish to employ a robust gradient-based kernel. We use the proposed nonlinear appearance model learned online via KPCA in Krein space for visual tracking in many popular and difficult tracking scenarios. We also show applications of our kernel framework for the problem of face recognition. Stephan Liwicki, Stefanos Zafeiriou, Georgios Tzimiropoulos, Maja Pantic |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2012 | Regularized Kernel Discriminant Analysis With a Robust Kernel for Face Recognition and VerificationabstractWe propose a robust approach to discriminant kernel-based feature extraction for face recognition and verification. We show, for the first time, how to perform the eigen analysis of the within-class scatter matrix directly in the feature space. This eigen analysis provides the eigenspectrum of its range space and the corresponding eigenvectors as well as the eigenvectors spanning its null space. Based on our analysis, we propose a kernel discriminant analysis (KDA) which combines eigenspectrum regularization with a feature-level scheme (ER-KDA). Finally, we combine the proposed ER-KDA with a nonlinear robust kernel particularly suitable for face recognition/verification applications which require robustness against outliers caused by occlusions and illumination changes. We applied the proposed framework to several popular databases (Yale, AR, XM2VTS) and achieved state-of-the-art performance for most of our experiments. Stefanos Zafeiriou, Georgios Tzimiropoulos, Maria Petrou, Tania Stathaki |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2011 | Fast and robust appearance-based trackingabstractWe introduce a fast and robust subspace-based approach to appearance-based object tracking. The core of our approach is based on Fast Robust Correlation (FRC), a recently proposed technique for the robust estimation of large translational displacements. We show how the basic principles of FRC can be naturally extended to formulate a robust version of Principal Component Analysis (PCA) which can be efficiently implemented incrementally and therefore is particularly suitable for robust real-time appearance-based object tracking. Our experimental results demonstrate that the proposed approach outperforms other state-of-the-art holistic appearance-based trackers on several popular video sequences. Stephan Liwicki, Stefanos Zafeiriou, Georgios Tzimiropoulos, Maja Pantic |
FG | 2 |
| 2011 | A dynamic approach to the recognition of 3D facial expressions and their temporal modelsabstractIn this paper we propose a method that exploits 3D motion-based features between frames of 3D facial geometry sequences for dynamic facial expression recognition. An expressive sequence is modeled to contain an onset followed by an apex and an offset. Feature selection methods are applied in order to extract features for each of the onset and offset segments of the expression. These features are then used to train a Hidden Markov Model in order to model the full temporal dynamics of the expression. The proposed fully automatic system was tested in a subset of the BU-4DFE database for the recognition of happiness, anger and surprise. Comparisons with a similar system based on the motion extracted from facial intensity images was also performed. The attained results suggest that the use of the 3D information does indeed improve the recognition accuracy when compared to the 2D data. Georgia Sandbach, Stefanos Zafeiriou, Maja Pantic, Daniel Rueckert |
FG | 2 |
| 2011 | Principal component analysis of image gradient orientations for face recognitionabstractWe introduce the notion of Principal Component Analysis (PCA) of image gradient orientations. As image data is typically noisy, but noise is substantially different from Gaussian, traditional PCA of pixel intensities very often fails to estimate reliably the low-dimensional subspace of a given data population. We show that replacing intensities with gradient orientations and the ℓ2norm with a cosine-based distance measure offers, to some extend, a remedy to this problem. Our scheme requires the eigen-decomposition of a covariance matrix and is as computationally efficient as standard ℓ2intensity-based PCA. We demonstrate some of its favorable properties for the application of face recognition. Georgios Tzimiropoulos, Stefanos Zafeiriou, Maja Pantic |
FG | 2 |
| 2011 | Audiovisual classification of vocal outbursts in human conversation using Long-Short-Term Memory networksabstractWe investigate classification of non-linguistic vocalisations with a novel audiovisual approach and Long Short-Term Memory (LSTM) Recurrent Neural Networks as highly successful dynamic sequence classifiers. As database of evaluation serves this year's Paralinguistic Challenge's Audiovisual Interest Corpus of human-to-human natural conversation. For video-based analysis we compare shape and appearance based features. These are fused in an early manner with typical audio descriptors. The results show significant improvements of LSTM networks over a static approach based on Support Vector Machines. More important, we can show a significant gain in performance when fusing audio and visual shape features. Florian Eyben, Stavros Petridis, Björn W. Schuller, Georgios Tzimiropoulos, Stefanos Zafeiriou, Maja Pantic |
ICASSP | 5 |
| 2011 | Robust and efficient parametric face alignmentabstractWe propose a correlation-based approach to parametric object alignment particularly suitable for face analysis applications which require efficiency and robustness against occlusions and illumination changes. Our algorithm registers two images by iteratively maximizing their correlation coefficient using gradient ascent. We compute this correlation coefficient from complex gradients which capture the orientation of image structures rather than pixel intensities. The maximization of this gradient correlation coefficient results in an algorithm which is as computationally efficient as ℓ2norm-based algorithms, can be extended within the inverse compositional framework (without the need for Hessian recomputation) and is robust to outliers. To the best of our knowledge, no other algorithm has been proposed so far having all three features. We show the robustness of our algorithm for the problem of face alignment in the presence of occlusions and non-uniform illumination changes. The code that reproduces the results of our paper can be found at http://ibug.doc.ic.ac.uk/resources. Georgios Tzimiropoulos, Stefanos Zafeiriou, Maja Pantic |
ICCV | 2 |
| 2011 | 2.5D Elastic graph matching
Stefanos Zafeiriou, Maria Petrou |
Comput. Vis. Image Underst. | 1 |
| 2011 | Nonnegative tensor factorization as an alternative Csiszar-Tusnady procedure: algorithms, convergence, probabilistic interpretations and novel probabilistic tensor latent variable analysis algorithms
Stefanos Zafeiriou, Maria Petrou |
Data Min. Knowl. Discov. | 1 |
| 2010 | Bidirectional relighting for 3D-aided 2D face recognitionabstractIn this paper, we present a new method for bidirectional relighting for 3D-aided 2D face recognition under large pose and illumination changes. During subject enrollment, we build subject-specific 3D annotated models by using the subjects' raw 3D data and 2D texture. During authentication, the probe 2D images are projected onto a normalized image space using the subject-specific 3D model in the gallery. Then, a bidirectional relighting algorithm and two similarity metrics (a view-dependent complex wavelet structural similarity and a global similarity) are employed to compare the gallery and probe. We tested our algorithms on the UHDB11 and UHDB12 databases that contain 3D data with probe images under large lighting and pose variations. The experimental results show the robustness of our approach in recognizing faces in difficult situations. George Toderici, Georgios Passalis, Stefanos Zafeiriou, Georgios Tzimiropoulos, Maria Petrou, Theoharis Theoharis, Ioannis A. Kakadiaris |
CVPR | 3 |
| 2010 | Nonnegative Embeddings and Projections for Dimensionality Reduction and Information VisualizationabstractIn this paper, we propose novel algorithms for low dimensionality nonnegative embedding of vectorial and/or relational data, as well as nonnegative projections for dimensionality reduction. We start by introducing a novel algorithm for Metric Multidimensional Scaling (MMS). We propose algorithms for Nonnegative Locally Linear Embedding (NLLE) and Nonnegative Laplacian Eigenmaps (NLE). By reformulating the problem of MMS, NLLE and NLE for finding projections we propose algorithms for Nonnegative Principal Component Analysis (NPCA), for Nonnegative Orthogonal Neighbourhood Preserving Projections (NONPP) and Nonnegative Orthogonal Locality Preserving Projections (NOLPP). We demonstrate some first preliminary results of the proposed methods in data visualization. Stefanos Zafeiriou, Nikolaos A. Laskaris |
ICPR | 1 |
| 2010 | Robust FFT-Based Scale-Invariant Image Registration with Image GradientsabstractWe present a robust FFT-based approach to scale-invariant image registration. Our method relies on FFT-based correlation twice: once in the log-polar Fourier domain to estimate the scaling and rotation and once in the spatial domain to recover the residual translation. Previous methods based on the same principles are not robust. To equip our scheme with robustness and accuracy, we introduce modifications which tailor the method to the nature of images. First, we derive efficient log-polar Fourier representations by replacing image functions with complex gray-level edge maps. We show that this representation both captures the structure of salient image features and circumvents problems related to the low-pass nature of images, interpolation errors, border effects, and aliasing. Second, to recover the unknown parameters, we introduce the normalized gradient correlation. We show that, using image gradients to perform correlation, the errors induced by outliers are mapped to a uniform distribution for which our normalized gradient correlation features robust performance. Exhaustive experimentation with real images showed that, unlike any other Fourier-based correlation techniques, the proposed method was able to estimate translations, arbitrary rotations, and scale factors up to 6. Georgios Tzimiropoulos, Vasileios Argyriou, Stefanos Zafeiriou, Tania Stathaki |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2010 | Image replica detection system utilizing R-trees and linear discriminant analysis
Spiros Nikolopoulos, Stefanos Zafeiriou, Nikos Nikolaidis 0001, Ioannis Pitas |
Pattern Recognit. | 2 |
| 2010 | Nonlinear Non-Negative Component Analysis AlgorithmsabstractIn this paper, general solutions for nonlinear non-negative component analysis for data representation and recognition are proposed. Motivated by a combination of the non-negative matrix factorization (NMF) algorithm and kernel theory, which has lead to a recently proposed NMF algorithm in a polynomial feature space, we propose a general framework where one can build a nonlinear non-negative component analysis method using kernels, the so-called projected gradient kernel non-negative matrix factorization (PGKNMF). In the proposed approach, arbitrary positive definite kernels can be adopted while at the same time it is ensured that the limit point of the procedure is a stationary point of the optimization problem. Moreover, we propose fixed point algorithms for the special case of Gaussian radial basis function (RBF) kernels. We demonstrate the power of the proposed methods in face and facial expression recognition applications. Stefanos Zafeiriou, Maria Petrou |
IEEE Trans. Image Process. | 1 |
| 2009 | Nonlinear Nonnegative Component AnalysisabstractIn this paper general solutions for nonlinear nonnegative component analysis for data representation and recognition are proposed. That is, motivated by a combination of the Nonnegative Matrix Factorization (NMF) algorithm and kernel theory, which has lead to an NMF algorithm in a polynomial feature space, we propose a general framework where one can build a nonlinear nonnegative component analysis using kernels, the so-called projected gradient Kernel Nonnegative Matrix Factorization (PGKNMF). In the proposed approach, arbitrary positive kernels can be adopted while at the same time it is ensured that the limit point of the procedure is a stationary point of the optimization problem. Moreover, we propose fixed point algorithms for the special case of Radial Basis Function (RBF) kernels. We demonstrate the power of the proposed methods in face and facial expression recognition applications. Stefanos Zafeiriou, Maria Petrou |
CVPR | 1 |
| 2009 | Use of random time-intervals (RTIs) generation for biometric verification
Nikolaos A. Laskaris, Stefanos Zafeiriou, Lambrini Garefa |
Pattern Recognit. | 2 |
| 2009 | Novel Multiclass Classifiers Based on the Minimization of the Within-Class VarianceabstractIn this paper, a novel class of multiclass classifiers inspired by the optimization of Fisher discriminant ratio and the support vector machine (SVM) formulation is introduced. The optimization problem of the so-called minimum within-class variance multiclass classifiers (MWCVMC) is formulated and solved in arbitrary Hilbert spaces, defined by Mercer's kernels, in order to find multiclass decision hyperplanes/surfaces. Afterwards, MWCVMCs are solved using indefinite kernels and dissimilarity measures via pseudo-Euclidean embedding. The power of the proposed approach is first demonstrated in the facial expression recognition of the seven basic facial expressions (i.e., anger, disgust, fear, happiness, sadness, and surprise plus the neutral state) problem in the presence of partial facial occlusion by using a pseudo-Euclidean embedding of Hausdorff distances and the MWCVMC. The experiments indicated a recognition accuracy rate achieved up to 99%. The MWCVMC classifiers are also applied to face recognition and other classification problems using Mercer's kernels. Irene Kotsia, Ioannis Pitas, Stefanos Zafeiriou |
IEEE Trans. Neural Networks | 3 |
| 2009 | Discriminant Nonnegative Tensor Factorization AlgorithmsabstractNonnegative matrix factorization (NMF) has proven to be very successful for image analysis, especially for object representation and recognition. NMF requires the object tensor (with valence more than one) to be vectorized. This procedure may result in information loss since the local object structure is lost due to vectorization. Recently, in order to remedy this disadvantage of NMF methods, nonnegative tensor factorizations (NTF) algorithms that can be applied directly to the tensor representation of object collections have been introduced. In this paper, we propose a series of unsupervised and supervised NTF methods. That is, we extend several NMF methods using arbitrary valence tensors. Moreover, by incorporating discriminant constraints inside the NTF decompositions, we present a series of discriminant NTF methods. The proposed approaches are tested for face verification and facial expression recognition, where it is shown that they outperform other popular subspace approaches. Stefanos Zafeiriou |
IEEE Trans. Neural Networks | 1 |
| 2008 | Texture and Shape Information Fusion for Facial Action Unit RecognitionabstractA novel method that fuses texture and shape information to achieve Facial Action Unit (FAU) recognition from video sequences is proposed. In order to extract the texture information, a subspace method based on Discriminant Non- negative Matrix Factorization (DNMF) is applied on the difference images of the video sequence, calculated taking under consideration the neutral and the most expressive frame, to extract the desired classification label. The shape information consists of the deformed Candide facial grid (more specifically the grid node displacements between the neutral and the most expressive facial expression frame) that corresponds to the facial expression depicted in the video sequence. The shape information is afterwards classified using a two-class Support Vector Machine (SVM) system. The fusion of texture and shape information is performed using Median Radial Basis Functions (MRBFs) Neural Networks (NNs) in order to detect the set of present FAUs. The accuracy achieved in the Cohn-Kanade database is equal to 92.1% when recognizing the 17 FAUs that are responsible for facial expression development. Irene Kotsia, Stefanos Zafeiriou, Nikos Nikolaidis 0001, Ioannis Pitas |
ACHI | 2 |
| 2008 | Motivating class-specific nonlinear projections for single and multiple view face verificationabstractIn this paper we motivate the use of class-specific nonlinear subspace methods for face verification. The problem of face verification is considered as a two-class problem (genuine versus impostor class). The typical Fisher's linear discriminant analysis (FLDA) gives only one or two projections in a two-class problem. This is a very strict limitation to the search of discriminant dimensions. As for the FLDA for N class problems (N > 2) the transformation is not person specific. In order to remedy these limitations of FLDA, exploit the individuality of human faces and take into consideration the fact that the distribution of facial images, under different viewpoints, illumination variations and facial expression is highly complex and non-linear, novel kernel discriminant algorithms are used. The new method was tested in the face verification problem using single and multiple view datasets and found to outperform other commonly used kernel approaches. Georgios Goudelis, Stefanos Zafeiriou, Anastasios Tefas, Nikos Nikolaidis 0001, Ioannis Pitas |
ICIP | 2 |
| 2008 | Texture and shape information fusion for facial expression and facial action unit recognition
Irene Kotsia, Stefanos Zafeiriou, Ioannis Pitas |
Pattern Recognit. | 2 |
| 2008 | Beyond FCM: Graph-theoretic post-processing algorithms for learning and representing the data structure
Nikolaos A. Laskaris, Stefanos Zafeiriou |
Pattern Recognit. | 2 |
| 2008 | On the Improvement of Support Vector Techniques for Clustering by Means of Whitening TransformabstractIn this letter, we suggest a novel method for clustering, based on finding the smallest enclosing hyperellipse in arbitrary Hilbert spaces. In particular, we show that the one class support vector method that finds the minimum bounding hypersphere, under the whitening transform, becomes a method for finding the minimum bounding hyperellipse. Afterwards, we generalize the method in order to find the minimum bounding hyperellipse in arbitrary Hilbert spaces. We illustrate the power of the proposed methods in clustering applications. Stefanos Zafeiriou, Nikolaos A. Laskaris |
IEEE Signal Process. Lett. | 1 |
| 2008 | Camera Motion Estimation Using a Novel Online Vector Field Model in Particle FiltersabstractIn this paper, a novel algorithm for parametric camera motion estimation is introduced. More particularly, a novel stochastic vector field model is proposed, which can handle smooth motion patterns derived from long periods of stable camera motion and can also cope with rapid camera motion changes and periods when the camera remains still. The stochastic vector field model is established from a set of noisy measurements, such as motion vectors derived, e.g., from block matching techniques, in order to provide an estimation of the subsequent camera motion in the form of a motion vector field. A set of rules for a robust and online update of the camera motion model parameters is also proposed, based on the expectation maximization algorithm. The proposed model is embedded in a particle filters framework in order to predict the future camera motion based on current and prior observations. We estimate the subsequent camera motion by finding the optimum affine transform parameters so that, when applied to the current video frame, the resulting motion vector field to approximate the one estimated by the stochastic model. Extensive experimental results verify the usefulness of the proposed scheme in camera motion pattern classification and in the accurate estimation of the 2D affine camera transform motion parameters. Moreover, the camera motion estimation has been incorporated into an object tracker in order to investigate if the new schema improves its tracking efficiency, when camera motion and tracked object motion are combined. Symeon Nikitidis, Stefanos Zafeiriou, Ioannis Pitas |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Discriminant Graph Structures for Facial Expression RecognitionabstractIn this paper, a series of advances in elastic graph matching for facial expression recognition are proposed. More specifically, a new technique for the selection of the most discriminant facial landmarks for every facial expression (discriminant expression-specific graphs) is applied. Furthermore, a novel kernel-based technique for discriminant feature extraction from graphs is presented. This feature extraction technique remedies some of the limitations of the typical kernel Fisher discriminant analysis (KFDA) which provides a subspace of very limited dimensionality (i.e., one or two dimensions) in two-class problems. The proposed methods have been applied to the Cohn-Kanade database in which very good performance has been achieved in a fully automatic manner. Stefanos Zafeiriou, Ioannis Pitas |
IEEE Trans. Multim. | 1 |
| 2007 | Discriminant Graph Structures for Face VerificationabstractElastic graph matching is one of the most well known techniques for frontal face recognition/verification and one of the few techniques that can be combined successfully with fully automatic face localization and alignment methods. In this paper, we propose an algorithm for finding the most discriminant features upon a person's face and a person-specific graph is placed in the spatial coordinates that correspond to these discriminant features. We illustrate the improvements in performance by applying the proposed method in frontal face verification using the XM2VTS database. Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas |
ICASSP (1) | 1 |
| 2007 | A Novel Kernel Discriminant Analysis for Face VerificationabstractIn this paper a novel non-linear subspace method for face verification is proposed. The problem of face verification is considered as a two-class problem (genuine versus impostor class). The typical Fisher's linear discriminant analysis (FLDA) gives only one or two projections in a two-class problem. This is a very strict limitation to the search of discriminant dimensions. As for the FLDA for N class problems (N is greater than two) the transformation is not person specific. In order to remedy these limitations of FLDA, exploit the individuality of human faces and take into consideration the fact that the distribution of facial images, under different viewpoints, illumination variations and facial expression is highly complex and non-linear, novel kernel discriminant algorithms are proposed. The new methods are tested in the face verification problem using the XM2VTS database where it is verified that they outperform other commonly used kernel approaches. Georgios Goudelis, Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas |
ICIP (4) | 2 |
| 2007 | The discriminant elastic graph matching algorithm applied to frontal face verification
Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas |
Pattern Recognit. | 1 |
| 2007 | Class-Specific Kernel-Discriminant Analysis for Face VerificationabstractIn this paper, novel nonlinear subspace methods for face verification are proposed. The problem of face verification is considered as a two-class problem (genuine versus impostor class). The typical Fisher's linear discriminant analysis (FLDA) gives only one or two projections in a two-class problem. This is a very strict limitation to the search of discriminant dimensions. As for the FLDA forNclass problems (Nis greater than two), the transformation is not person specific. In order to remedy these limitations of FLDA, exploit the individuality of human faces and take into consideration the fact that the distribution of facial images, under different viewpoints, illumination variations, and facial expression is highly complex and nonlinear, novel kernel-discriminant algorithms are proposed. The new methods are tested in the face verification problem using the XM2VTS, AR, ORL, Yale, and UMIST databases where it is verified that they outperform other commonly used kernel approaches such as kernel-PCA (KPCA), kernel direct discriminant analysis (KDDA), complete kernel Fisher's discriminant analysis (CKFDA), the two-class KDDA, CKFDA, and other two-class and multiclass variants of kernel-discriminant analysis based on Fisher's criterion. Georgios Goudelis, Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2007 | A Novel Discriminant Non-Negative Matrix Factorization Algorithm With Applications to Facial Image Characterization ProblemsabstractThe methods introduced so far regarding discriminant non-negative matrix factorization (DNMF) do not guarantee convergence to a stationary limit point. In order to remedy this limitation, a novel DNMF method is presented that uses projected gradients. The proposed algorithm employs some extra modifications that make the method more suitable for classification tasks. The usefulness of the proposed technique to frontal face verification and facial expression recognition problems is demonstrated. Irene Kotsia, Stefanos Zafeiriou, Ioannis Pitas |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2007 | Learning Discriminant Person-Specific Facial Models Using Expandable GraphsabstractIn this paper, a novel algorithm for finding discriminant person-specific facial models is proposed and tested for frontal face verification. The most discriminant features of a person's face are found and a deformable model is placed in the spatial coordinates that correspond to these discriminant features. The discriminant deformable models, for verifying the person's identity, that are learned through this procedure are elastic graphs that are dense in the facial areas considered discriminant for a specific person and sparse in other less significant facial areas. The discriminant graphs are enhanced by a discriminant feature selection method for the graph nodes in order to find the most discriminant jet features. The proposed approach significantly enhances the performance of elastic graph matching in frontal face verification Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2007 | Minimum Class Variance Support Vector MachinesabstractIn this paper, a modified class of support vector machines (SVMs) inspired from the optimization of Fisher's discriminant ratio is presented, the so-called minimum class variance SVMs (MCVSVMs). The MCVSVMs optimization problem is solved in cases in which the training set contains less samples that the dimensionality of the training vectors using dimensionality reduction through principal component analysis (PCA). Afterward, the MCVSVMs are extended in order to find nonlinear decision surfaces by solving the optimization problem in arbitrary Hilbert spaces defined by Mercer's kernels. In that case, it is shown that, under kernel PCA, the nonlinear optimization problem is transformed into an equivalent linear MCVSVMs problem. The effectiveness of the proposed approach is demonstrated by comparing it with the standard SVMs and other classifiers, like kernel Fisher discriminant analysis in facial image characterization problems like gender determination, eyeglass, and neutral facial expression detection. Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas |
IEEE Trans. Image Process. | 1 |
| 2006 | Image Replica Detection using R-Trees and Linear Discriminant AnalysisabstractIn this paper a novel system for image replica detection is presented. The system uses color-based descriptors in order to extract robust features for image representation. These features are used for indexing the images in a database using an R-tree. When a query about whether a test image is a replica of an image in the database is submitted, the R-tree is traversed and a set of candidate images is retrieved. Then, in order to obtain a single result and at the same time reduce the number of decision errors the system is enhanced with linear discriminant analysis (LDA). The conducted experiments show that the proposed approach is very promising Spiros Nikolopoulos, Stefanos Zafeiriou, Panagiotis Sidiropoulos, Nikos Nikolaidis 0001, Ioannis Pitas |
ICME | 2 |
| 2006 | Exploiting discriminant information in nonnegative matrix factorization with application to frontal face verificationabstractIn this paper, two supervised methods for enhancing the classification accuracy of the Nonnegative Matrix Factorization (NMF) algorithm are presented. The idea is to extend the NMF algorithm in order to extract features that enforce not only the spatial locality, but also the separability between classes in a discriminant manner. The first method employs discriminant analysis in the features derived from NMF. In this way, a two-phase discriminant feature extraction procedure is implemented, namely NMF plus Linear Discriminant Analysis (LDA). The second method incorporates the discriminant constraints inside the NMF decomposition. Thus, a decomposition of a face to its discriminant parts is obtained and new update rules for both the weights and the basis images are derived. The introduced methods have been applied to the problem of frontal face verification using the well-known XM2VTS database. Both methods greatly enhance the performance of NMF for frontal face verification. Stefanos Zafeiriou, Anastasios Tefas, Ioan Buciu, Ioannis Pitas |
IEEE Trans. Neural Networks | 1 |
| 2005 | Exploiting discriminant information in elastic graph matchingabstractIn this paper, we investigate the use of discriminant techniques in the elastic graph matching (EGM) algorithm. First we use discriminant analysis in the feature vectors of the nodes in order to find the most discriminant features. The similarity measure for discriminant feature vectors and the node deformation are combined in a discriminant manner in order to form a local similarity measure between nodes. Moreover, the local similarity values at the nodes of the elastic graph, are weighted by coefficients that are also derived by some discriminant analysis in order to form a total similarity measure between faces. We illustrate the improvements in performance in frontal face verification using a modified multiscale morphological analysis. Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas |
ICIP (3) | 1 |
| 2005 | Blind Robust Watermarking Schemes for Copyright Protection of 3D Mesh ObjectsabstractIn this paper, two novel methods suitable for blind 3D mesh object watermarking applications are proposed. The first method is robust against 3D rotation, translation, and uniform scaling. The second one is robust against both geometric and mesh simplification attacks. A pseudorandom watermarking signal is cast in the 3D mesh object by deforming its vertices geometrically, without altering the vertex topology. Prior to watermark embedding and detection, the object is rotated and translated so that its center of mass and its principal component coincide with the origin and the z-axis of the Cartesian coordinate system. This geometrical transformation ensures watermark robustness to translation and rotation. Robustness to uniform scaling is achieved by restricting the vertex deformations to occur only along the r coordinate of the corresponding (r, theta, phi) spherical coordinate system. In the first method, a set of vertices that correspond to specific angles theta is used for watermark embedding. In the second method, the samples of the watermark sequence are embedded in a set of vertices that correspond to a range of angles in the theta domain in order to achieve robustness against mesh simplifications. Experimental results indicate the ability of the proposed method to deal with the aforementioned attacks. Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2004 | A blind robust watermarking scheme for copyright protection of 3d mesh models
Stefanos Zafeiriou, Anastasios Tefas, Ioannis Pitas |
ICIP | 1 |