EDBT 2026 Demo / reviewers in the wild / expert
Yannis Panagakis
dblp:99/8307 · also Ioannis Panagakis, Yiannis Panagakis
· DBLP profile ↗
83ranked-venue papers
9as first author
27since 2021 · last 2026
0000-0003-0153-5210ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 64 · 5 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Disentangling Local and Global Semantics in Diffusion Models for Image EditingabstractAbstract Diffusion models have achieved state-of-the-art image synthesis, yet unlike GANs, they lack a well-structured latent space for intuitive image editing. Existing diffusion-based editing methods often rely on supervised fine-tuning or text-based guidance, while recent unsupervised techniques leveraging the model’s bottleneck layer suffer from one or more key limitations: (i) they focus only on global attributes, (ii) fail to disentangle local and global semantics, or (iii) require extensive human intervention. To fill this gap, we first propose an unsupervised method for localized image editing in pre-trained unconditional diffusion models that disentangles local and global semantics in the model’s latent space. Given an input image and a user-specified region of interest, our approach uses the denoising network’s Jacobian to map that region to a corresponding latent subspace. We then separate this subspace into shared (global) and region-specific components to uncover latent directions that control local attributes. These directions generalize across images, enabling semantically consistent edits without retraining. We go one step further by extending our method to minimize manual supervision by automatically inferring edit directions from a single reference image and generating region masks without human input. Experiments on multiple datasets show that our method yields more localized, high-fidelity edits than state-of-the-art approaches. Manos Plitsis, Theodoros Kouzelis, Panagiotis Koromilas, Vassilis Katsouros, Mihalis A. Nicolaou, Yannis Panagakis |
Int. J. Comput. Vis. | 6 |
| 2025 | Are Large Brainwave Foundation Models Capable Yet ? Insights from Fine-TuningabstractFoundation Models have demonstrated significant success across various domains in Artificial Intelligence (AI), yet their capabilities for brainwave modeling remain unclear. In this paper, we comprehensively evaluate current Large Brainwave Foundation Models (LBMs) through systematic fine-tuning experiments across multiple Brain-Computer Interface (BCI) benchmark tasks, including memory tasks and sleep stage classification. Our extensive analysis shows that state-of-the-art LBMs achieve only marginal improvements (0.5\%) over traditional deep architectures while requiring significantly more parameters (millions vs thousands), raising important questions about their efficiency and applicability in BCI contexts. Moreover, through detailed ablation studies and Low-Rank Adaptation (LoRA), we significantly reduce trainable parameters without performance degradation, while demonstrating that architectural and training inefficiencies limit LBMs' current capabilities. Our experiments span both full model fine-tuning and parameter-efficient adaptation techniques, providing insights into optimal training strategies for BCI applications. We pioneer the application of LoRA to LBMs, revealing that performance benefits generally emerge when adapting multiple neural network components simultaneously. These findings highlight the critical need for domain-specific development strategies to advance LBMs, suggesting that current architectures may require redesign to fully leverage the potential of foundation models in brainwave analysis. Na Lee, Konstantinos Barmpas, Yannis Panagakis, Dimitrios A. Adamos, Nikolaos A. Laskaris, Stefanos Zafeiriou |
ICML | 3 |
| 2025 | Synthetic Speech Source Tracing using Metric Learning
Dimitrios Koutsianos, Stavros Zacharopoulos, Yannis Panagakis, Themos Stafylakis |
INTERSPEECH | 3 |
| 2024 | Enabling Local Editing in Diffusion Models by Joint and Individual Component Analysis
Theodoros Kouzelis, Emmanouil Plitsis, Mihalis A. Nicolaou, Yannis Panagakis |
BMVC | 4 |
| 2024 | Investigating Personalization Methods in Text to Music GenerationabstractIn this work, we investigate the personalization of text-to-music diffusion models in a few-shot setting. Motivated by recent advances in the computer vision domain, we are the first to explore the combination of pre-trained text-to-audio diffusers with two established personalization methods. We experiment with the effect of audio-specific data augmentation on the overall system performance and assess different training strategies. For evaluation, we construct a novel dataset with prompts and music clips. We consider both embedding-based and music-specific metrics for quantitative evaluation, as well as a user study for qualitative evaluation. Our analysis shows that similarity metrics are in accordance with user preferences and that current personalization approaches tend to learn rhythmic music concepts more easily than melody. The code, dataset, and example material of this study are open to the research community1. Manos Plitsis, Theodoros Kouzelis, Georgios Paraskevopoulos, Vassilis Katsouros, Yannis Panagakis |
ICASSP | 5 |
| 2024 | Bridging Mini-Batch and Asymptotic Analysis in Contrastive Learning: From InfoNCE to Kernel-Based LossesabstractWhat do different contrastive learning (CL) losses actually optimize for? Although multiple CL methods have demonstrated remarkable representation learning capabilities, the differences in their inner workings remain largely opaque. In this work, we analyse several CL families and prove that, under certain conditions, they admit the same minimisers when optimizing either their batch-level objectives or their expectations asymptotically. In both cases, an intimate connection with the hyperspherical energy minimisation (HEM) problem resurfaces. Drawing inspiration from this, we introduce a novel CL objective, coined Decoupled Hyperspherical Energy Loss (DHEL). DHEL simplifies the problem by decoupling the target hyperspherical energy from the alignment of positive examples while preserving the same theoretical guarantees. Going one step further, we show the same results hold for another relevant CL family, namely kernel contrastive learning (KCL), with the additional advantage of the expected loss being independent of batch size, thus identifying the minimisers in the non-asymptotic regime. Empirical results demonstrate improved downstream performance and robustness across combinations of different batch sizes and hyperparameters and reduced dimensionality collapse, on several computer vision datasets. Panagiotis Koromilas, Giorgos Bouritsas, Theodoros Giannakopoulos, Mihalis A. Nicolaou, Yannis Panagakis |
ICML | 5 |
| 2024 | Multilinear Mixture of Experts: Scalable Expert Specialization through FactorizationabstractThe Mixture of Experts (MoE) paradigm provides a powerful way to decompose dense layers into smaller, modular computations often more amenable to human interpretation, debugging, and editability. However, a major challenge lies in the computational cost of scaling the number of experts high enough to achieve fine-grained specialization. In this paper, we propose the Multilinear Mixture of Experts (μMoE) layer to address this, focusing on vision models. μMoE layers enable scalable expert specialization by performing an implicit computation on prohibitively large weight tensors entirely in factorized form. Consequently, μMoEs (1) avoid the restrictively high inference-time costs of dense MoEs, yet (2) do not inherit the training issues of the popular sparse MoEs' discrete (non-differentiable) expert routing. We present both qualitative and quantitative evidence that scaling μMoE layers when fine-tuning foundation models for vision tasks leads to more specialized experts at the class-level, further enabling manual bias correction in CelebA attribute classification. Finally, we show qualitative results demonstrating the expert specialism achieved when pre-training large GPT2 and MLP-Mixer models with parameter-matched μMoE blocks at every layer, maintaining comparable accuracy. Our code is available at: https://github.com/james-oldfield/muMoE. James Oldfield 0001, Markos Georgopoulos, Grigorios Chrysos 0002, Christos Tzelepis, Yannis Panagakis, Mihalis A. Nicolaou, Jiankang Deng, Ioannis Patras |
NeurIPS | 5 |
| 2024 | Scale Equivariant Graph MetanetworksabstractThis paper pertains to an emerging machine learning paradigm: learning higher- order functions, i.e. functions whose inputs are functions themselves, particularly when these inputs are Neural Networks (NNs). With the growing interest in architectures that process NNs, a recurring design principle has permeated the field: adhering to the permutation symmetries arising from the connectionist structure of
NNs. However, are these the sole symmetries present in NN parameterizations? Zooming into most practical activation functions (e.g. sine, ReLU, tanh) answers this question negatively and gives rise to intriguing new symmetries, which we collectively refer to as scaling symmetries, that is, non-zero scalar multiplications and divisions of weights and biases. In this work, we propose Scale Equivariant Graph MetaNetworks - ScaleGMNs, a framework that adapts the Graph Metanetwork (message-passing) paradigm by incorporating scaling symmetries and thus rendering neuron and edge representations equivariant to valid scalings. We introduce novel building blocks, of independent technical interest, that allow for equivariance or invariance with respect to individual scalar multipliers or their product and use them in all components of ScaleGMN. Furthermore, we prove that, under certain expressivity conditions, ScaleGMN can simulate the forward and backward pass of any input feedforward neural network. Experimental results demonstrate that our method advances the state-of-the-art performance for several datasets and activation functions, highlighting the power of scaling symmetries as an inductive bias for NN processing. The source code is publicly available at https://github.com/jkalogero/scalegmn. Ioannis Kalogeropoulos, Giorgos Bouritsas, Yannis Panagakis |
NeurIPS | 3 |
| 2024 | Benchmarking Counterfactual Image GenerationabstractGenerative AI has revolutionised visual content editing, empowering users to effortlessly modify images and videos. However, not all edits are equal. To perform realistic edits in domains such as natural image or medical imaging, modifications must respect causal relationships inherent to the data generation process. Such image editing falls into the counterfactual image generation regime. Evaluating counterfactual image generation is substantially complex: not only it lacks observable ground truths, but also requires adherence to causal constraints. Although several counterfactual image generation methods and evaluation metrics exist a comprehensive comparison within a unified setting is lacking. We present a comparison framework to thoroughly benchmark counterfactual image generation methods. We evaluate the performance of three conditional image generation model families developed within the Structural Causal Model (SCM) framework. We incorporate several metrics that assess diverse aspects of counterfactuals, such as composition, effectiveness, minimality of interventions, and image realism. We integrate all models that have been used for the task at hand and expand them to novel datasets and causal graphs, demonstrating the superiority of Hierarchical VAEs across most datasets and metrics. Our framework is implemented in a user-friendly Python package that can be extended to incorporate additional SCMs, causal methods, generative models, and datasets for the community to build on. Code: https://github.com/gulnazaki/counterfactual-benchmark. Thomas Melistas, Nikos Spyrou, Nefeli Gkouti, Pedro Sanchez, Athanasios Vlontzos, Yannis Panagakis, Giorgos Papanastasiou, Sotirios A. Tsaftaris |
NeurIPS | 6 |
| 2024 | Locality-Preserving Directions for Interpreting the Latent Space of Satellite Image GANsabstractWe present a locality-aware method for interpreting the latent space of wavelet-based generative adversarial networks (GANs), that can well capture the large spatial and spectral variability that is characteristic of satellite imagery. By focusing on preserving locality, the proposed method is able to decompose the weight-space of pre-trained GANs and recover interpretable directions that correspond to high-level semantic concepts (such as urbanization, structure density, flora presence)-that can subsequently be used for guided synthesis of satellite imagery. In contrast to typically used approaches that focus on capturing the variability of the weight-space in a reduced dimensionality space (i.e., based on principal component analysis (PCA)), we show that preserving locality leads to vectors with different angles, that are more robust to artifacts and can better preserve class information. Via a set of quantitative and qualitative examples, we further show that the proposed approach can outperform both baseline geometric augmentations, as well as global, PCA-based approaches for data synthesis in the context of data augmentation for satellite scene classification. Georgia Kourmouli, Nikos Kostagiolas, Mihalis A. Nicolaou, Yannis Panagakis |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Bilinear Models of Parts and Appearances in Generative Adversarial NetworksabstractRecent advances in the understanding of Generative Adversarial Networks (GANs) have led to remarkable progress in visual editing and synthesis tasks, capitalizing on the rich semantics that are embedded in the latent spaces of pre-trained GANs. However, existing methods are often tailored to specific GAN architectures and are limited to either discovering global semantic directions that do not facilitate localized control, or require some form of supervision through manually provided regions or segmentation masks. In this light, we present an architecture-agnostic approach that jointly discovers factors representing spatial parts and their appearances in an entirely unsupervised fashion. These factors are obtained by applying a semi-nonnegative tensor factorization on the feature maps, which in turn enables context-aware local image editing with pixel-level control. In addition, we show that the discovered appearance factors correspond to saliency maps that localize concepts of interest, without using any labels. Experiments on a wide range of GAN architectures and datasets show that, in comparison to the state of the art, our method is far more efficient in terms of training time and, most importantly, provides much more accurate localized control. James Oldfield 0001, Christos Tzelepis, Yannis Panagakis, Mihalis A. Nicolaou, Ioannis Patras |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Large-Scale Unsupervised Audio Pre-Training for Video-to-Speech SynthesisabstractVideo-to-speech synthesis is the task of reconstructing the speech signal from a silent video of a speaker. Previous approaches train on data from almost exclusively audio-visual datasets, i.e., every audio sample has a corresponding video sample. This precludes the use of abundant audio-only datasets which may not have a corresponding visual modality such as audiobooks, radio podcasts, and speech recognition datasets. In this paper we propose to train encoder-decoder models on more than 3,500 hours of audio data at 24kHz, and then use the pretrained decoders to initialize the audio decoders for the video-tospeech synthesis task. The pre-training step uses audio samples only and does not require labels or corresponding samples from other modalities (visual, text). We demonstrate that this improves the reconstructed speech and that it is an unexplored way to improve the quality of the generator in a cross-modal task while only requiring samples from one of the modalities. We conduct experiments using both raw audio and mel spectrograms as target outputs and benchmark our models with existing work. Triantafyllos Kefalas, Yannis Panagakis, Maja Pantic |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | MMATR: A Lightweight Approach for Multimodal Sentiment Analysis Based on Tensor MethodsabstractDespite the considerable research output on Multimodal Learning for Affect-related tasks, most of the current methods are very complex in terms of the number of trainable parameters, and thus do not constitute effective solutions for real-life applications. In this work we try to alleviate this gap in the literature by introducing the Multimodal Attention Tensor Regression (MMATR) network, a lightweight model that is based on: (i) a static input representation (2D matrix of dimensions time × features) for each modality, which helps to avoid high-parameterized sequential models by incorporating a CNN, (ii) the replacement of the usual pooling and flattening operations as well as the linear layers by tensor contraction and tensor regression layers that are able to reduce the number of parameters, while keeping the high-order structure of the multimodal data, and (iii) a bimodal attention layer that learns multimodal co-occurrences. By a set of experiments comparing with a variety of state-of-the-art techniques, we show that the proposed MMATR can achieve results competitive to the state-of-the-art in the task of Multimodal Sentiment Analysis, albeit having four orders of magnitude fewer parameters. Panagiotis Koromilas, Mihalis A. Nicolaou, Theodoros Giannakopoulos, Yannis Panagakis |
ICASSP | 4 |
| 2023 | PandA: Unsupervised Learning of Parts and Appearances in the Feature Maps of GANs
James Oldfield 0001, Christos Tzelepis, Yannis Panagakis, Mihalis A. Nicolaou, Ioannis Patras |
ICLR | 3 |
| 2023 | Parts of Speech-Grounded Subspaces in Vision-Language ModelsabstractLatent image representations arising from vision-language models have proved immensely useful for a variety of downstream tasks. However, their utility is limited by their entanglement with respect to different visual attributes. For instance, recent work has shown that CLIP image representations are often biased toward specific visual properties (such as objects or actions) in an unpredictable manner. In this paper, we propose to separate representations of the different visual modalities in CLIP’s joint vision-language space by leveraging the association between parts of speech and specific visual modes of variation (e.g. nouns relate to objects, adjectives describe appearance). This is achieved by formulating an appropriate component analysis model that learns subspaces capturing variability corresponding to a specific part of speech, while jointly minimising variability to the rest. Such a subspace yields disentangled representations of the different visual properties of an image or text in closed form while respecting the underlying geometry of the manifold on which the representations lie. What’s more, we show the proposed model additionally facilitates learning subspaces corresponding to specific visual appearances (e.g. artists’ painting styles), which enables the selective removal of entire visual themes from CLIP-based text-to-image synthesis. We validate the model both qualitatively, by visualising the subspace projections with a text-to-image model and by preventing the imitation of artists’ styles, and quantitatively, through class invariance metrics and improvements to baseline zero-shot classification. James Oldfield 0001, Christos Tzelepis, Yannis Panagakis, Mihalis A. Nicolaou, Ioannis Patras |
NeurIPS | 3 |
| 2023 | KAN-AV dataset for audio-visual face and speech analysis in the wildabstractHuman-computer interaction is becoming increasingly prevalent in daily life with the adoption of intelligent devices. These devices must be capable of interacting in diverse settings, such as environments with noise, music and differing illumination and occlusion conditions. They must also interact with a variety of end users across ages and backgrounds. Therefore, the machine learning community needs in-the-wild multi-modal datasets to develop models for face and speech analysis so that they can be applicable in most real world scenarios. However, most existing audio and audio-visual databases are captured in controlled conditions with few or no age and kinship labels. In this paper, we introduce the KAN-AV dataset which contains 98 h of audio-visual data from 970 identities across ages. Two thirds of the identities have kin relations in the dataset. The dataset is manually annotated with labels for kinship, age, and gender and is intended to drive future research in face and speech analysis. Triantafyllos Kefalas, Eftychia Fotiadou, Markos Georgopoulos, Yannis Panagakis, Pingchuan Ma 0001, Stavros Petridis, Themos Stafylakis, Maja Pantic |
Image Vis. Comput. | 4 |
| 2022 | Cluster-guided Image Synthesis with Unconditional ModelsabstractGenerative Adversarial Networks (GANs) are the driving force behind the state-of-the-art in image generation. Despite their ability to synthesize high-resolution photo-realistic images, generating content with on-demand conditioning of different granularity remains a challenge. This challenge is usually tackled by annotating massive datasets with the attributes of interest, a laborious task that is not always a viable option. Therefore, it is vital to introduce control into the generation process of unsupervised generative models. In this work, we focus on controllable image generation by leveraging GANs that are well-trained in an unsupervised fashion. To this end, we discover that the representation space of intermediate layers of the generator forms a number of clusters that separate the data according to semantically meaningful attributes (e.g., hair color and pose). By conditioning on the cluster assignments, the proposed method is able to control the semantic class of the generated image. Our approach enables sampling from each cluster by Implicit Maximum Likelihood Estimation (IMLE). We showcase the efficacy of our approach on faces (CelebA-HQ and FFHQ), animals (Imagenet) and objects (LSUN) using different pre-trained generative models. The results highlight the ability of our approach to condition image generation on attributes like gender, pose and hair style on faces, as well as a variety of features on different object classes. Markos Georgopoulos, James Oldfield 0001, Grigorios Chrysos 0002, Yannis Panagakis |
CVPR | 4 |
| 2022 | Augmenting Deep Classifiers with Polynomial Neural Networks
Grigorios Chrysos 0002, Markos Georgopoulos, Jiankang Deng, Jean Kossaifi, Yannis Panagakis, Anima Anandkumar |
ECCV (25) | 5 |
| 2022 | Deep Polynomial Neural NetworksabstractDeep convolutional neural networks (DCNNs) are currently the method of choice both for generative, as well as for discriminative learning in computer vision and machine learning. The success of DCNNs can be attributed to the careful selection of their building blocks (e.g., residual blocks, rectifiers, sophisticated normalization schemes, to mention but a few). In this paper, we propose Π-Nets, a new class of function approximators based on polynomial expansions. Π-Nets are polynomial neural networks, i.e., the output is a high-order polynomial of the input. The unknown parameters, which are naturally represented by high-order tensors, are estimated through a collective tensor factorization with factors sharing. We introduce three tensor decompositions that significantly reduce the number of parameters and show how they can be efficiently implemented by hierarchical neural networks. We empirically demonstrate that Π-Nets are very expressive and they even produce good results without the use of non-linear activation functions in a large battery of tasks and signals, i.e., images, graphs, and audio. When used in conjunction with activation functions, Π-Nets produce state-of-the-art results in three challenging tasks, i.e., image generation, face verification and 3D mesh representation learning. The source code is available at https://github.com/grigorisg9gr/polynomial_nets. Grigorios Chrysos 0002, Stylianos Moschoglou, Giorgos Bouritsas, Jiankang Deng, Yannis Panagakis, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Adversarial Learning of Disentangled and Generalizable Representations of Visual AttributesabstractRecently, a multitude of methods for image-to-image translation have demonstrated impressive results on problems, such as multidomain or multiattribute transfer. The vast majority of such works leverages the strengths of adversarial learning and deep convolutional autoencoders to achieve realistic results by well-capturing the target data distribution. Nevertheless, the most prominent representatives of this class of methods do not facilitate semantic structure in the latent space and usually rely on binary domain labels for test-time transfer. This leads to rigid models, unable to capture the variance of each domain label. In this light, we propose a novel adversarial learning method that: 1) facilitates the emergence of latent structure by semantically disentangling sources of variation and 2) encourages learning generalizable, continuous, and transferable latent codes that enable flexible attribute mixing. This is achieved by introducing a novel loss function that encourages representations to result in uniformly distributed class posteriors for disentangled attributes. In tandem with an algorithm for inducing generalizable properties, the resulting representations can be utilized for a variety of tasks such as intensity-preserving multiattribute image translation and synthesis, without requiring labeled test data. We demonstrate the merits of the proposed method by a set of qualitative and quantitative experiments on popular databases such as MultiPIE, RaFD, and BU-3DFE, where our method outperforms other state-of-the-art methods in tasks such as intensity-preserving multiattribute transfer and synthesis. James Oldfield 0001, Yannis Panagakis, Mihalis A. Nicolaou |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Tensor Component Analysis for Interpreting the Latent Space of GANs
James Oldfield 0001, Markos Georgopoulos, Yannis Panagakis, Mihalis A. Nicolaou, Ioannis Patras |
BMVC | 3 |
| 2021 | Defensive Tensorization
Adrian Bulat, Jean Kossaifi, Sourav Bhattacharya, Yannis Panagakis, Timothy M. Hospedales, Georgios Tzimiropoulos, Nicholas D. Lane, Maja Pantic |
BMVC | 4 |
| 2021 | Learning Separable Time-Frequency Filterbanks for Audio ClassificationabstractState-of-the-art audio classification systems often apply deep neural networks on hand-crafted features (e.g., spectrogram-based representations), instead of learning features directly from raw audio. Moreover, these audio networks have millions of unknown parameters need to be learned, which causes a great demand for computational resources and training data. In this paper, we aim to learn audio representations directly from raw audio, and at the same time mitigate its training burden by employing a light-weight architecture. In particular, we propose to learn separable filters, parametrized with only a few variables, namely center frequency and bandwidth, facilitating training and offering interpretability of learned representations. The generality of the proposed method is demonstrated by applying it onto two applications, namely 1) speaker identification and 2) acoustic event recognition. Experimental results indicate its effectiveness on these applications, especially when small amount of training data is available. Jie Pu, Yannis Panagakis, Maja Pantic |
ICASSP | 2 |
| 2021 | Conditional Generation Using Polynomial ExpansionsabstractGenerative modeling has evolved to a notable field of machine learning. Deep polynomial neural networks (PNNs) have demonstrated impressive results in unsupervised image generation, where the task is to map an input vector (i.e., noise) to a synthesized image. However, the success of PNNs has not been replicated in conditional generation tasks, such as super-resolution. Existing PNNs focus on single-variable polynomial expansions which do not fare well to two-variable inputs, i.e., the noise variable and the conditional variable. In this work, we introduce a general framework, called CoPE, that enables a polynomial expansion of two input variables and captures their auto- and cross-correlations. We exhibit how CoPE can be trivially augmented to accept an arbitrary number of input variables. CoPE is evaluated in five tasks (class-conditional generation, inverse problems, edges-to-image translation, image-to-image translation, attribute-guided generation) involving eight datasets. The thorough evaluation suggests that CoPE can be useful for tackling diverse conditional generation tasks. The source code of CoPE is available at https://github.com/grigorisg9gr/polynomialnetsforconditionalgeneration. Grigorios Chrysos 0002, Markos Georgopoulos, Yannis Panagakis |
NeurIPS | 3 |
| 2021 | Mitigating Demographic Bias in Facial Datasets with Style-Based Multi-attribute TransferabstractAbstract Deep learning has catalysed progress in tasks such as face recognition and analysis, leading to a quick integration of technological solutions in multiple layers of our society. While such systems have proven to be accurate by standard evaluation metrics and benchmarks, a surge of work has recently exposed the demographic bias that such algorithms exhibit–highlighting that accuracy does not entail fairness. Clearly, deploying biased systems under real-world settings can have grave consequences for affected populations. Indeed, learning methods are prone to inheriting, or even amplifying the bias present in a training set, manifested by uneven representation across demographic groups. In facial datasets, this particularly relates to attributes such as skin tone, gender, and age. In this work, we address the problem of mitigating bias in facial datasets by data augmentation. We propose a multi-attribute framework that can successfully transfer complex, multi-scale facial patterns even if these belong to underrepresented groups in the training set. This is achieved by relaxing the rigid dependence on a single attribute label, and further introducing a tensor-based mixing structure that captures multiplicative interactions between attributes in a multilinear fashion. We evaluate our method with an extensive set of qualitative and quantitative experiments on several datasets, with rigorous comparisons to state-of-the-art methods. We find that the proposed framework can successfully mitigate dataset bias, as evinced by extensive evaluations on established diversity metrics, while significantly improving fairness metrics such as equality of opportunity. Markos Georgopoulos, James Oldfield 0001, Mihalis A. Nicolaou, Yannis Panagakis, Maja Pantic |
Int. J. Comput. Vis. | 4 |
| 2021 | SEWA DB: A Rich Database for Audio-Visual Emotion and Sentiment Research in the WildabstractNatural human-computer interaction and audio-visual human behaviour sensing systems, which would achieve robust performance in-the-wild are more needed than ever as digital devices are increasingly becoming an indispensable part of our life. Accurately annotated real-world data are the crux in devising such systems. However, existing databases usually consider controlled settings, low demographic variability, and a single task. In this paper, we introduce the SEWA database of more than 2,000 minutes of audio-visual data of 398 people coming from six cultures, 50 percent female, and uniformly spanning the age range of 18 to 65 years old. Subjects were recorded in two different contexts: while watching adverts and while discussing adverts in a video chat. The database includes rich annotations of the recordings in terms of facial landmarks, facial action units (FAU), various vocalisations, mirroring, and continuously valued valence, arousal, liking, agreement, and prototypic examples of (dis)liking. This database aims to be an extremely valuable resource for researchers in affective computing and automatic human sensing and is expected to push forward the research in human behaviour analysis, including cultural studies. Along with the database, we provide extensive baseline experiments for automatic FAU detection and automatic valence, arousal, and (dis)liking intensity estimation. Jean Kossaifi, Robert Walecki, Yannis Panagakis, Jie Shen 0008, Maximilian Schmitt, Fabien Ringeval, Jing Han 0010, Vedhas Pandit, Antoine Toisoul, Björn W. Schuller, Kam Star, Elnar Hajiyev, Maja Pantic |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Tensor Methods in Computer Vision and Deep LearningabstractTensors, or multidimensional arrays, are data structures that can naturally represent visual data of multiple dimensions. Inherently able to efficiently capture structured, latent semantic spaces and high-order interactions, tensors have a long history of applications in a wide span of computer vision problems. With the advent of the deep learning paradigm shift in computer vision, tensors have become even more fundamental. Indeed, essential ingredients in modern deep learning architectures, such as convolutions and attention mechanisms, can readily be considered as tensor mappings. In effect, tensor methods are increasingly finding significant applications in deep learning, including the design of memory and compute efficient network architectures, improving robustness to random noise and adversarial attacks, and aiding the theoretical understanding of deep networks. This article provides an in-depth and practical review of tensors and tensor methods in the context of representation learning and deep learning, with a particular focus on visual data analysis and computer vision applications. Concretely, besides fundamental work in tensor-based visual data analysis methods, we focus on recent developments that have brought on a gradual increase in tensor methods, especially in deep learning architectures and their implications in computer vision applications. To further enable the newcomer to grasp such concepts quickly, we provide companion Python notebooks, covering key aspects of this article and implementing them, step-by-step with TensorLy. Yannis Panagakis, Jean Kossaifi, Grigorios Chrysos 0002, James Oldfield 0001, Mihalis A. Nicolaou, Anima Anandkumar, Stefanos Zafeiriou |
Proc. IEEE | 1 |
| 2020 | P-nets: Deep Polynomial Neural NetworksabstractDeep Convolutional Neural Networks (DCNNs) is currently the method of choice both for generative, as well as for discriminative learning in computer vision and machine learning. The success of DCNNs can be attributed to the careful selection of their building blocks (e.g., residual blocks, rectifiers, sophisticated normalization schemes, to mention but a few). In this paper, we propose Π-Nets, a new class of DCNNs. Π-Nets are polynomial neural networks, i.e., the output is a high-order polynomial of the input. Π-Nets can be implemented using special kind of skip connections and their parameters can be represented via high-order tensors. We empirically demonstrate that Π-Nets have better representation power than standard DCNNs and they even produce good results without the use of non-linear activation functions in a large battery of tasks and signals, i.e., images, graphs, and audio. When used in conjunction with activation functions, Π-Nets produce state-of-the-art results in challenging tasks, such as image generation. Lastly, our framework elucidates why recent generative models, such as StyleGAN, improve upon their predecessors, e.g., ProGAN. Grigorios Chrysos 0002, Stylianos Moschoglou, Giorgos Bouritsas, Yannis Panagakis, Jiankang Deng, Stefanos Zafeiriou |
CVPR | 4 |
| 2020 | Factorized Higher-Order CNNs With an Application to Spatio-Temporal Emotion EstimationabstractTraining deep neural networks with spatio-temporal (i.e., 3D) or multidimensional convolutions of higher-order is computationally challenging due to millions of unknown parameters across dozens of layers. To alleviate this, one approach is to apply low-rank tensor decompositions to convolution kernels in order to compress the network and reduce its number of parameters. Alternatively, new convolutional blocks, such as MobileNet, can be directly designed for efficiency. In this paper, we unify these two approaches by proposing a tensor factorization framework for efficient multidimensional (separable) convolutions of higher-order. Interestingly, the proposed framework enables a novel higher-order transduction, allowing to train a network on a given domain (e.g., 2D images or N-dimensional data in general) and using transduction to generalize to higher-order data such as videos (or (N+K)--dimensional data in general), capturing for instance temporal dynamics while preserving the learnt spatial information. We apply the proposed methodology, coined CP-Higher-Order Convolution (HO-CPConv), to spatio-temporal facial emotion analysis. Most existing facial affect models focus on static imagery and discard all temporal information. This is due to the above-mentioned burden of training 3D convolutional nets and the lack of large bodies of video data annotated by experts. We address both issues with our proposed framework. Initial training is first done on static imagery before using transduction to generalize to the temporal domain. We demonstrate superior performance on three challenging large scale affect estimation datasets, AffectNet, SEWA, and AFEW-VA. Jean Kossaifi, Antoine Toisoul, Adrian Bulat, Yannis Panagakis, Timothy M. Hospedales, Maja Pantic |
CVPR | 4 |
| 2020 | Speech-Driven Facial Animation Using Polynomial Fusion of FeaturesabstractSpeech-driven facial animation involves using a speech signal to generate realistic videos of talking faces. Recent deep learning approaches to facial synthesis rely on extracting low-dimensional representations and concatenating them, followed by a decoding step of the concatenated vector. This accounts for only first-order interactions of the features and ignores higher-order interactions. In this paper we propose a polynomial fusion layer that models the joint representation of the encodings by a higher-order polynomial, with the parameters modelled by a tensor decomposition. We demonstrate the suitability of this approach through experiments on generated videos evaluated on a range of metrics on video quality, audiovisual synchronisation and generation of blinks. Triantafyllos Kefalas, Konstantinos Vougioukas, Yannis Panagakis, Stavros Petridis, Jean Kossaifi, Maja Pantic |
ICASSP | 3 |
| 2020 | Learning Differentiable Sparse and Low Rank Networks for Audio-Visual Object LocalizationabstractParsimonious modelling, including sparsity and low rankness, has becomes a cornerstone in modern machine learning and signal processing. However, these modelling techniques have limited capabity to learn from large-scale data, and often require some pre-defined parameters to define their optimization procedure. In this paper, we propose a novel method to design specific deep neural networks for sparse and low-rank models, where the network can learn a data-adaptive model from training data. In particular, we design differentiable network units for sparse and low-rank matrices. Each layer of the network represents one iteration of the optimization process of sparse and low-rank models. The effectiveness of the proposed method is evaluated in the task of audio-visual object localization. Experimental results indicate the superior performance of the proposed method over traditional sparse and low-rank models. Jie Pu, Yannis Panagakis, Maja Pantic |
ICASSP | 2 |
| 2020 | Multilinear Latent Conditioning for Generating Unseen Attribute CombinationsabstractDeep generative models rely on their inductive bias to facilitate generalization, especially for problems with high dimensional data, like images. However, empirical studies have shown that variational autoencoders (VAE) and generative adversarial networks (GAN) lack the generalization ability that occurs naturally in human perception. For example, humans can visualize a woman smiling after only seeing a smiling man. On the contrary, the standard conditional VAE (cVAE) is unable to generate unseen attribute combinations. To this end, we extend cVAE by introducing a multilinear latent conditioning framework that captures the multiplicative interactions between the attributes. We implement two variants of our model and demonstrate their efficacy on MNIST, Fashion-MNIST and CelebA. Altogether, we design a novel conditioning framework that can be used with any architecture to synthesize unseen attribute combinations. Markos Georgopoulos, Grigorios Chrysos 0002, Maja Pantic, Yannis Panagakis |
ICML | 4 |
| 2020 | Investigating bias in deep face analysis: The KANFace dataset and empirical study
Markos Georgopoulos, Yannis Panagakis, Maja Pantic |
Image Vis. Comput. | 2 |
| 2020 | NAPS: Non-adversarial polynomial synthesis
Grigorios Chrysos 0002, Yannis Panagakis |
Pattern Recognit. Lett. | 2 |
| 2020 | Active Speaker Detection and Localization in Videos Using Low-Rank and Kernelized SparsityabstractA novel method for active speaker detection and localization in audio-visual recordings is proposed. The method relies on a specifically tailored matrix decomposition that exploits the intrinsic low-dimensional structure of audio-visual data, namely, the low-rank of the background visual/audio information and the sparsity of the correlated foreground components. Concretely, the data matrix of each modality is modeled as a superposition of two terms: 1) a low-rank matrix capturing the background information and 2) a kernelized sparse matrix capturing the non-linear correlated components among the audio and visual modalities and, hence, revealing the active speaker. To this end, we formulate an appropriate optimization problem that involves the minimization of nuclear- and matrix $\ell _1$-norms, and develop an efficient solver. Experimental results on active speaker detection and localization demonstrate the superior performance of the proposed method over other state-of-the-art approaches. Jie Pu, Yannis Panagakis, Maja Pantic |
IEEE Signal Process. Lett. | 2 |
| 2020 | Blind Audio-Visual Localization and Separation via Low-Rank and SparsityabstractThe ability to localize visual objects that are associated with an audio source and at the same time to separate the audio signal is a cornerstone in audio-visual signal-processing applications. However, available methods mainly focus on localizing only the visual objects, without audio separation abilities. Besides that, these methods often rely on either laborious preprocessing steps to segment video frames into semantic regions, or additional supervisions to guide their localization. In this paper, we aim to address the problem of visual source localization and audio separation in an unsupervised manner and avoid all preprocessing or post-processing steps. To this end, we devise a novel structured matrix decomposition method that decomposes the data matrix of each modality as a superposition of three terms: 1) a low-rank matrix capturing the background information; 2) a sparse matrix capturing the correlated components among the two modalities and, hence, uncovering the sound source in visual modality and the associated sound in audio modality; and 3) a third sparse matrix accounting for uncorrelated components, such as distracting objects in visual modality and irrelevant sound in audio modality. The generality of the proposed method is demonstrated by applying it onto three applications, namely: 1) visual localization of a sound source; 2) visually assisted audio separation; and 3) active speaker detection. Experimental results indicate the effectiveness of the proposed method on these application domains. Jie Pu, Yannis Panagakis, Stavros Petridis, Jie Shen 0008, Maja Pantic |
IEEE Trans. Cybern. | 2 |
| 2019 | Valence and Arousal Estimation In-The-Wild with Tensor MethodsabstractWhile it is relatively easy and natural for humans to detect and interpret non-verbal cues, it is a hard task for computer systems. Automatic recognition of emotions has been the subject of extensive studies in the past decade, but despite the various methods that have been implemented, the problem remains challenging. In particular, most existing works focus on predicting a set discrete stereotypical emotion categories. We are instead interested in predicting continuous values of valence and arousal, which are able to model accurately a broader range of spontaneous emotions. Moreover, as opposed to much of the prior work that focused on controlled (laboratory) conditions, we are interested in analysis in naturalist (in-the-wild) conditions. To do so, we propose to leverage the structure in the data using tensor methods. In addition to preserving the structure, these have the potential to also reduce the total number of parameters in the models, thus improving the computational performance. We first consider a model with analytic solution in the form of a Tucker Tensor Regression. We then investigate a deep, gradient based method, namely Tensor Regression Networks. We perform thorough experiments on two publicly available databases, AFEW-VA and SEWA, for facial affect estimation in-the-wild, in terms of valence and arousal levels. Experimental results demonstrate that tensor-based methods successfully leverage the structure in the data and on average outperform baseline methods. Anna Mitenkova, Jean Kossaifi, Yannis Panagakis, Maja Pantic |
FG | 3 |
| 2019 | Learning Low Rank and Sparse Models via Robust AutoencodersabstractRobust principal component analysis (RPCA), decomposes a data matrix into a superposition of a low-rank matrix and a sparse matrix under certain incoherent conditions. In this paper, we propose a nonlinear generalization of RPCA that uses two autoencoder networks to achieve such a decomposition, in which one autoencoder accounts for the low-rank component and the other for the sparse component. To this end, we provide a principled way of constructing these autoencoders for low-rank and sparse components. The generality of the proposed model is demonstrated by applying it onto three applications, namely 1) music/voice separation 2) image denoising and 3) video foreground separation. Experimental results indicate the effectiveness of the proposed model on these application domains. Jie Pu, Yannis Panagakis, Maja Pantic |
ICASSP | 2 |
| 2019 | Disentangling Geometry and Appearance with Regularised Geometry-Aware Generative Adversarial NetworksabstractDeep generative models have significantly advanced image generation, enabling generation of visually pleasing images with realistic texture. Apart from the texture, it is the shape geometry of objects that strongly dictates their appearance. However, currently available generative models do not incorporate geometric information into the image generation process. This often yields visual objects of degenerated quality. In this work, we propose a regularized Geometry-Aware Generative Adversarial Network (GAGAN) which disentangles appearance and shape in the latent space. This regularized GAGAN enables the generation of images with both realistic texture and shape. Specifically, we condition the generator on a statistical shape prior. The prior is enforced through mapping the generated images onto a canonical coordinate frame using a differentiable geometric transformation. In addition to incorporating geometric information, this constrains the search space and increases the model’s robustness. We show that our approach is versatile, able to generalise across domains (faces, sketches, hands and cats) and sample sizes (from as little as $$\sim \, 200{-}30{,}000$$ to more than 200, 000). We demonstrate superior performance through extensive quantitative and qualitative experiments in a variety of tasks and settings. Finally, we leverage our model to automatically and accurately detect errors or drifting in facial landmarks detection and tracking in-the-wild. Jean Kossaifi, Yannis Panagakis, Maja Pantic |
Int. J. Comput. Vis. | 3 |
| 2019 | An Adversarial Neuro-Tensorial Approach for Learning Disentangled RepresentationsabstractSeveral factors contribute to the appearance of an object in a visual scene, including pose, illumination, and deformation, among others. Each factor accounts for a source of variability in the data, while the multiplicative interactions of these factors emulate the entangled variability, giving rise to the rich structure of visual object appearance. Disentangling such unobserved factors from visual data is a challenging task, especially when the data have been captured in uncontrolled recording conditions (also referred to as “in-the-wild”) and label information is not available. In this paper, we propose a pseudo-supervised deep learning method for disentangling multiple latent factors of variation in face images captured in-the-wild. To this end, we propose a deep latent variable model, where the multiplicative interactions of multiple latent factors of variation are explicitly modelled by means of multilinear (tensor) structure. We demonstrate that the proposed approach indeed learns disentangled representations of facial expressions and pose, which can be used in various applications, including face editing, as well as 3D face reconstruction and classification of facial expression, identity and pose. Mengjiao Wang 0002, Zhixin Shu, Shiyang Cheng 0001, Yannis Panagakis, Dimitris Samaras, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 4 |
| 2019 | TensorLy: Tensor Learning in PythonabstractTensors are higher-order extensions of matrices. While matrix methods form the cornerstone of traditional machine learning and data analysis, tensor methods have been gaining increasing traction. However, software support for tensor operations is not on the same footing. In order to bridge this gap, we have developed TensorLy, a Python library that provides a high-level API for tensor methods and deep tensorized neural networks. TensorLy aims to follow the same standards adopted by the main projects of the Python scientific community, and to seamlessly integrate with them. Its BSD license makes it suitable for both academic and commercial applications. TensorLy's backend system allows users to perform computations with several libraries such as NumPy or PyTorch to name but a few. They can be scaled on multiple CPU or GPU machines. In addition, using the deep-learning frameworks as backend allows to easily design and train deep tensorized neural networks. TensorLy is available at https://github.com/tensorly/tensorly Jean Kossaifi, Yannis Panagakis, Anima Anandkumar, Maja Pantic |
J. Mach. Learn. Res. | 2 |
| 2019 | Robust Kronecker Component AnalysisabstractDictionary learning and component analysis models are fundamental for learning compact representations that are relevant to a given task (feature extraction, dimensionality reduction, denoising, etc.). The model complexity is encoded by means of specific structure, such as sparsity, low-rankness, or nonnegativity. Unfortunately, approaches like K-SVD - that learn dictionaries for sparse coding via Singular Value Decomposition (SVD) - are hard to scale to high-volume and high-dimensional visual data, and fragile in the presence of outliers. Conversely, robust component analysis methods such as the Robust Principal Component Analysis (RPCA) are able to recover low-complexity (e.g., low-rank) representations from data corrupted with noise of unknown magnitude and support, but do not provide a dictionary that respects the structure of the data (e.g., images), and also involve expensive computations. In this paper, we propose a novel Kronecker-decomposable component analysis model, coined as Robust Kronecker Component Analysis (RKCA), that combines ideas from sparse dictionary learning and robust component analysis. RKCA has several appealing properties, including robustness to gross corruption; it can be used for low-rank modeling, and leverages separability to solve significantly smaller problems. We design an efficient learning algorithm by drawing links with a restricted form of tensor factorization, and analyze its optimality and low-rankness properties. The effectiveness of the proposed approach is demonstrated on real-world applications, namely background subtraction and image denoising and completion, by performing a thorough comparison with the current state of the art. Mehdi Bahri, Yannis Panagakis, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Non-Negative Matrix Factorizations for Multiplex Network AnalysisabstractNetworks have been a general tool for representing, analyzing, and modeling relational data arising in several domains. One of the most important aspect of network analysis is community detection or network clustering. Until recently, the major focus have been on discovering community structure in single (i.e., monoplex) networks. However, with the advent of relational data with multiple modalities, multiplex networks, i.e., networks composed of multiple layers representing different aspects of relations, have emerged. Consequently, community detection in multiplex network, i.e., detecting clusters of nodes shared by all layers, has become a new challenge. In this paper, we propose Network Fusion for Composite Community Extraction (NF-CCE), a new class of algorithms, based on four different non-negative matrix factorization models, capable of extracting composite communities in multiplex networks. Each algorithm works in two steps: first, it finds a non-negative, low-dimensional feature representation of each network layer; then, it fuses the feature representation of layers into a common non-negative, low-dimensional feature representation via collective factorization. The composite clusters are extracted from the common feature representation. We demonstrate the superior performance of our algorithms over the state-of-the-art methods on various types of multiplex networks, including biological, social, economic, citation, phone communication, and brain multiplex networks. Vladimir Gligorijevic, Yannis Panagakis, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Side Information for Face Completion: A Robust PCA ApproachabstractRobust principal component analysis (RPCA) is a powerful method for learning low-rank feature representation of various visual data. However, for certain types as well as significant amount of error corruption, it fails to yield satisfactory results; a drawback that can be alleviated by exploiting domain-dependent prior knowledge or information. In this paper, we propose two models for the RPCA that take into account such side information, even in the presence of missing values. We apply this framework to the task of UV completion which is widely used in pose-invariant face recognition. Moreover, we construct a generative adversarial network (GAN) to extract side information as well as subspaces. These subspaces not only assist in the recovery but also speed up the process in case of large-scale data. We quantitatively and qualitatively evaluate the proposed approaches through both synthetic data and eight real-world datasets to verify their effectiveness. Niannan Xue, Jiankang Deng, Shiyang Cheng 0001, Yannis Panagakis, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2019 | Fast Multilevel Algorithms for Compressive Principal Component PursuitabstractRecovering a low-rank matrix from highly corrupted measurements arises in compressed sensing of structured high-dimensional signals (e.g., videos and hyperspectral images among others). Robust principal component analysis, solved via principal component pursuit (PCP), recovers a low-rank matrix from sparse corruptions that are of unknown value and support by decomposing the observation matrix into two terms, a low-rank matrix and a sparse one, accounting for sparse noise and outliers. In the more general setting, where only a fraction of the data matrix has been observed, low-rank matrix recovery is achieved by solving the compressive principal component pursuit (CPCP). Both PCP and CPCP are well-studied convex programs, and numerous iterative algorithms have been proposed for their optimization. Nevertheless, these algorithms involve singular value decomposition (SVD) at each iteration, which renders their applicability challenging in the case of massive data. In this paper, we propose a multilevel approach for the solution of PCP and CPCP problems. The core principle behind our algorithm is to apply SVD in models of lower dimensionality than the original one and then lift its solution to the original problem dimension. Hence, our methods rely on the assumption that the low-rank component can be represented in a lower dimensional space. We show that the proposed algorithms are easy to implement and converge at the same rate but with much lower iteration cost. Numerical experiments on numerous synthetic and real problems indicate that the proposed multilevel algorithms are several times faster than their original counterparts, namely, PCP and CPCP. Vahan Hovhannisyan, Yannis Panagakis, Panos Parpas, Stefanos Zafeiriou |
SIAM J. Imaging Sci. | 2 |
| 2018 | Informed Non-Convex Robust Principal Component Analysis With FeaturesabstractWe revisit the problem of robust principal component analysis with features acting as prior side information. To this aim, a novel, elegant, non-convex optimization approach is proposed to decompose a given observation matrix into a low-rank core and the corresponding sparse residual. Rigorous theoretical analysis of the proposed algorithm results in exact recovery guarantees with low computational complexity. Aptly designed synthetic experiments demonstrate that our method is the first to wholly harness the power of non-convexity over convexity in terms of both recoverability and speed. That is, the proposed non-convex approach is more accurate and faster compared to the best available algorithms for the problem under study. Two real-world applications, namely image classification and face denoising further exemplify the practical superiority of the proposed method. Niannan Xue, Jiankang Deng, Yannis Panagakis, Stefanos Zafeiriou |
AAAI | 3 |
| 2018 | GAGAN: Geometry-Aware Generative Adversarial NetworksabstractDeep generative models learned through adversarial training have become increasingly popular for their ability to generate naturalistic image textures. However, aside from their texture, the visual appearance of objects is significantly influenced by their shape geometry; information which is not taken into account by existing generative models. This paper introduces the Geometry-Aware Generative Adversarial Networks (GAGAN) for incorporating geometric information into the image generation process. Specifically, in GAGAN the generator samples latent variables from the probability space of a statistical shape model. By mapping the output of the generator to a canonical coordinate frame through a differentiable geometric transformation, we enforce the geometry of the objects and add an implicit connection from the prior to the generated object. Experimental results on face generation indicate that the GAGAN can generate realistic images of faces with arbitrary facial attributes such as facial expression, pose, and morphology, that are of better quality than current GAN-based methods. Our method can be used to augment any existing GAN architecture and improve the quality of the images generated. Jean Kossaifi, Yannis Panagakis, Maja Pantic |
CVPR | 3 |
| 2018 | Online Attention for Interpretable Conflict Estimation in Political DebatesabstractConflict arises naturally in dyadic interactions when involved individuals act on incompatible goals, interests, or actions. In this paper, the problem of conflict intensity estimation from audiovisual recordings is addressed. To this end, we propose an online attention-based neural network in order to learn a mapping from a sequence of audiovisual features to time-series describing conflict intensity. The proposed method is evaluated by conducting experiments in conflict intensity estimation by employing the CONFER dataset. Experimental results indicate the superiority of the proposed model compared to the state of the art. Furthermore, we demonstrate that by incorporating sparsity in the model, the origin of conflict can be traced back to specific key frames facilitating the interpretation of conflict escalation. Ruben Vereecken, Stavros Petridis, Yannis Panagakis, Maja Pantic |
FG | 3 |
| 2018 | Dynamic Behavior Analysis via Structured Rank MinimizationabstractHuman behavior and affect is inherently a dynamic phenomenon involving temporal evolution of patterns manifested through a multiplicity of non-verbal behavioral cues including facial expressions, body postures and gestures, and vocal outbursts. A natural assumption for human behavior modeling is that a continuous-time characterization of behavior is the output of a linear time-invariant system when behavioral cues act as the input (e.g., continuous rather than discrete annotations of dimensional affect). Here we study the learning of such dynamical system under real-world conditions, namely in the presence of noisy behavioral cues descriptors and possibly unreliable annotations by employing structured rank minimization. To this end, a novel structured rank minimization method and its scalable variant are proposed. The generalizability of the proposed framework is demonstrated by conducting experiments on 3 distinct dynamic behavior analysis tasks, namely (i) conflict intensity prediction, (ii) prediction of valence and arousal, and (iii) tracklet matching. The attained results outperform those achieved by other state-of-the-art methods for these tasks and, hence, evidence the robustness and effectiveness of the proposed approach. Christos Georgakis 0001, Yannis Panagakis, Maja Pantic |
Int. J. Comput. Vis. | 2 |
| 2018 | Modeling of facial aging and kinship: A survey
Markos Georgopoulos, Yannis Panagakis, Maja Pantic |
Image Vis. Comput. | 2 |
| 2018 | 3D Reconstruction of "In-the-Wild" Faces in Images and Videosabstract3D Morphable Models (3DMMs) are powerful statistical models of 3D facial shape and texture, and are among the state-of-the-art methods for reconstructing facial shape from single images. With the advent of new 3D sensors, many 3D facial datasets have been collected containing both neutral as well as expressive faces. However, all datasets are captured under controlled conditions. Thus, even though powerful 3D facial shape models can be learnt from such data, it is difficult to build statistical texture models that are sufficient to reconstruct faces captured in unconstrained conditions ("in-the-wild"). In this paper, we propose the first "in-the-wild" 3DMM by combining a statistical model of facial identity and expression shape with an "in-the-wild" texture model. We show that such an approach allows for the development of a greatly simplified fitting procedure for images and videos, as there is no need to optimise with regards to the illumination parameters. We have collected three new benchmarks that combine "in-the-wild" images and video with ground truth 3D facial geometry, the first of their kind, and report extensive quantitative evaluations using them that demonstrate our method is state-of-the-art. James Booth 0001, Anastasios Roussos, Evangelos Ververas, Epameinondas Antonakos, Stylianos Ploumpis, Yannis Panagakis, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2018 | Recovering Joint and Individual Components in Facial DataabstractA set of images depicting faces with different expressions or in various ages consists of components that are shared across all images (i.e., joint components) imparting to the depicted object the properties of human faces as well as individual components that are related to different expressions or age groups. Discovering the common (joint) and individual components in facial images is crucial for applications such as facial expression transfer and age progression. The problem is rather challenging when dealing with images captured in unconstrained conditions in the presence of sparse non-Gaussian errors of large magnitude (i.e., sparse gross errors or outliers) and contain missing data. In this paper, we investigate the use of a method recently introduced in statistics, the so-called Joint and Individual Variance Explained (JIVE) method, for the robust recovery of joint and individual components in visual facial data consisting of an arbitrary number of views. Since the JIVE is not robust to sparse gross errors, we propose alternatives, which are (1) robust to sparse gross, non-Gaussian noise, (2) able to automatically find the individual components rank, and (3) can handle missing data. We demonstrate the effectiveness of the proposed methods to several computer vision applications, namely facial expression synthesis and 2D and 3D face age progression 'in-the-wild'. Christos Sagonas, Evangelos Ververas, Yannis Panagakis, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Disentangling the Modes of Variation in Unlabelled DataabstractStatistical methods are of paramount importance in discovering the modes of variation in visual data. The Principal Component Analysis (PCA) is probably the most prominent method for extracting a single mode of variation in the data. However, in practice, several factors contribute to the appearance of visual objects including pose, illumination, and deformation, to mention a few. To extract these modes of variations from visual data, several supervised methods, such as the TensorFaces relying on multilinear (tensor) decomposition have been developed. The main drawbacks of such methods is that they require both labels regarding the modes of variations and the same number of samples under all modes of variations (e.g., the same face under different expressions, poses etc.). Therefore, their applicability is limited to well-organised data, usually captured in well-controlled conditions. In this paper, we propose a novel general multilinear matrix decomposition method that discovers the multilinear structure of possibly incomplete sets of visual data in unsupervised setting (i.e., without the presence of labels). We also propose extensions of the method with sparsity and low-rank constraints in order to handle noisy data, captured in unconstrained conditions. Besides that, a graph-regularised variant of the method is also developed in order to exploit available geometric or label information for some modes of variations. We demonstrate the applicability of the proposed method in several computer vision tasks, including Shape from Shading (SfS) (in the wild and with occlusion removal), expression transfer, and estimation of surface normals from images captured in the wild. Mengjiao Wang 0002, Yannis Panagakis, Patrick Snape, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | 3D Face Morphable Models "In-the-Wild"abstract3D Morphable Models (3DMMs) are powerful statistical models of 3D facial shape and texture, and among the state-of-the-art methods for reconstructing facial shape from single images. With the advent of new 3D sensors, many 3D facial datasets have been collected containing both neutral as well as expressive faces. However, all datasets are captured under controlled conditions. Thus, even though powerful 3D facial shape models can be learnt from such data, it is difficult to build statistical texture models that are sufficient to reconstruct faces captured in unconstrained conditions (in-the-wild). In this paper, we propose the first, to the best of our knowledge, in-the-wild 3DMM by combining a powerful statistical model of facial shape, which describes both identity and expression, with an in-the-wild texture model. We show that the employment of such an in-the-wild texture model greatly simplifies the fitting procedure, because there is no need to optimise with regards to the illumination parameters. Furthermore, we propose a new fast algorithm for fitting the 3DMM in arbitrary images. Finally, we have captured the first 3D facial database with relatively unconstrained conditions and report quantitative evaluations with state-of-the-art performance. Complementary qualitative reconstruction results are demonstrated on standard in-the-wild facial databases. James Booth 0001, Epameinondas Antonakos, Stylianos Ploumpis, George Trigeorgis, Yannis Panagakis, Stefanos Zafeiriou |
CVPR | 5 |
| 2017 | Robust Joint and Individual Variance ExplainedabstractDiscovering the common (joint) and individual subspaces is crucial for analysis of multiple data sets, including multi-view and multi-modal data. Several statistical machine learning methods have been developed for discovering the common features across multiple data sets. The most well studied family of the methods is that of Canonical Correlation Analysis (CCA) and its variants. Even though the CCA is a powerful tool, it has several drawbacks that render its application challenging for computer vision applications. That is, it discovers only common features and not individual ones, and it is sensitive to gross errors present in visual data. Recently, efforts have been made in order to develop methods that discover individual and common components. Nevertheless, these methods are mainly applicable in two sets of data. In this paper, we investigate the use of a recently proposed statistical method, the so-called Joint and Individual Variance Explained (JIVE) method, for the recovery of joint and individual components in an arbitrary number of data sets. Since, the JIVE is not robust to gross errors, we propose alternatives, which are both robust to non-Gaussian noise of large magnitude, as well as able to automatically find the rank of the individual components. We demonstrate the effectiveness of the proposed approach to two computer vision applications, namely facial expression synthesis and face age progression in-the-wild. Christos Sagonas, Yannis Panagakis, Alina Leidinger, Stefanos Zafeiriou |
CVPR | 2 |
| 2017 | Learning the Multilinear Structure of Visual DataabstractStatistical decomposition methods are of paramount importance in discovering the modes of variations of visual data. Probably the most prominent linear decomposition method is the Principal Component Analysis (PCA), which discovers a single mode of variation in the data. However, in practice, visual data exhibit several modes of variations. For instance, the appearance of faces varies in identity, expression, pose etc. To extract these modes of variations from visual data, several supervised methods, such as the TensorFaces, that rely on multilinear (tensor) decomposition (e.g., Higher Order SVD) have been developed. The main drawbacks of such methods is that they require both labels regarding the modes of variations and the same number of samples under all modes of variations (e.g., the same face under different expressions, poses etc.). Therefore, their applicability is limited to well-organised data, usually captured in well-controlled conditions. In this paper, we propose the first general multilinear method, to the best of our knowledge, that discovers the multilinear structure of visual data in unsupervised setting. That is, without the presence of labels. We demonstrate the applicability of the proposed method in two applications, namely Shape from Shading (SfS) and expression transfer. Mengjiao Wang 0002, Yannis Panagakis, Patrick Snape, Stefanos Zafeiriou |
CVPR | 2 |
| 2017 | Temporal Archetypal Analysis for Action SegmentationabstractUnsupervised learning of invariant representations that efficiently describe high-dimensional time series has several applications in dynamic visual data analysis. Clearly, the problem becomes more challenging when dealing with multiple time series arising from different modalities. A prominent example of this multimodal setting is the human motion which can be represented by multimodal time series of pixel intensities, depth maps, and motion capture data. Here, we study, for the first time, the problem of unsupervised learning of temporally and modality invariant informative representations, referred to as archetypes, from multiple time series originating from different modalities. To this end a novel method, coined as temporal archetypal analysis, is proposed. The performance of the proposed method is assessed by conducting experiments in unsupervised action segmentation. Experimental results on three different real world datasets using single modal and multimodal visual representations indicate the robustness and effectiveness of the proposed methods, outperforming compared state-of-the-art methods by a large, in most of the cases, margin. Eftychia Fotiadou, Yannis Panagakis, Maja Pantic |
FG | 2 |
| 2017 | Audio-visual object localization and separation using low-rank and sparsityabstractThe ability to localize visual objects that are associated with an audio source and at the same time seperate the audio signal is a corner stone in several audio-visual signal processing applications. Past efforts usually focused on localizing only the visual objects, without audio separation abilities. Besides, they often rely computational expensive pre-processing steps to segment images pixels into object regions before applying localization approaches. We aim to address the problem of audio-visual source localization and separation in an unsupervised manner. The proposed approach employs low-rank in order to model the background visual and audio information and sparsity in order to extract the sparsely correlated components between the audio and visual modalities. In particular, this model decomposes each dataset into a sum of two terms: the low-rank matrices capturing the background uncorrelated information, while the sparse correlated components modelling the sound source in visual modality and the associated sound in audio modality. To this end a novel optimization problem, involving the minimization of nuclear norms and matrix ℓ1-norms is solved. We evaluated the proposed method in 1) visual localization and audio separation and 2) visual-assisted audio denoising. The experimental results demonstrate the effectiveness of the proposed method. Jie Pu, Yannis Panagakis, Stavros Petridis, Maja Pantic |
ICASSP | 2 |
| 2017 | Robust Kronecker-Decomposable Component Analysis for Low-Rank Modeling
Mehdi Bahri, Yannis Panagakis, Stefanos Zafeiriou |
ICCV | 2 |
| 2017 | Side Information in Robust Principal Component Analysis: Algorithms and ApplicationsabstractRobust Principal Component Analysis (RPCA) aims at recovering a low-rank subspace from grossly corrupted high-dimensional (often visual) data and is a cornerstone in many machine learning and computer vision applications. Even though RPCA has been shown to be very successful in solving many rank minimisation problems, there are still cases where degenerate or suboptimal solutions are obtained. This is likely to be remedied by taking into account of domain-dependent prior knowledge. In this paper, we propose two models for the RPCA problem with the aid of side information on the low-rank structure of the data. The versatility of the proposed methods is demonstrated by applying them to four applications, namely background subtraction, facial image denoising, face and facial expression recognition. Experimental results on synthetic and five real world datasets indicate the robustness and effectiveness of the proposed methods on these application domains, largely outperforming six previous approaches. Niannan Xue, Yannis Panagakis, Stefanos Zafeiriou |
ICCV | 2 |
| 2017 | Robust Statistical Frontalization of Human and Animal FacesabstractThe unconstrained acquisition of facial data in real-world conditions may result in face images with significant pose variations, illumination changes, and occlusions, affecting the performance of facial landmark localization and recognition methods. In this paper, a novel method, robust to pose, illumination variations, and occlusions is proposed for joint face frontalization and landmark localization. Unlike the state-of-the-art methods for landmark localization and pose correction, where large amount of manually annotated images or 3D facial models are required, the proposed method relies on a small set of frontal images only. By observing that the frontal facial image of both humans and animals, is the one having the minimum rank of all different poses, a model which is able to jointly recover the frontalized version of the face as well as the facial landmarks is devised. To this end, a suitable optimization problem is solved, concerning minimization of the nuclear norm (convex surrogate of the rank function) and the matrix $$\ell _1$$ norm accounting for occlusions. The proposed method is assessed in frontal view reconstruction of human and animal faces, landmark localization, pose-invariant face recognition, face verification in unconstrained conditions, and video inpainting by conducting experiment on 9 databases. The experimental results demonstrate the effectiveness of the proposed method in comparison to the state-of-the-art methods for the target problems. Christos Sagonas, Yannis Panagakis, Stefanos Zafeiriou, Maja Pantic |
Int. J. Comput. Vis. | 2 |
| 2017 | The Conflict Escalation Resolution (CONFER) Database
Christos Georgakis 0001, Yannis Panagakis, Stefanos Zafeiriou, Maja Pantic |
Image Vis. Comput. | 2 |
| 2017 | Nonnegative Decompositions for Dynamic Visual Data AnalysisabstractThe analysis of high-dimensional, possibly temporally misaligned, and time-varying visual data is a fundamental task in disciplines, such as image, vision, and behavior computing. In this paper, we focus on dynamic facial behavior analysis and in particular on the analysis of facial expressions. Distinct from the previous approaches, where sets of facial landmarks are used for face representation, raw pixel intensities are exploited for: 1) unsupervised analysis of the temporal phases of facial expressions and facial action units (AUs) and 2) temporal alignment of a certain facial behavior displayed by two different persons. To this end, the slow features nonnegative matrix factorization (SFNMF) is proposed in order to learn slow varying parts-based representations of time varying sequences capturing the underlying dynamics of temporal phenomena, such as facial expressions. Moreover, the SFNMF is extended in order to handle two temporally misaligned data sequences depicting the same visual phenomena. To do so, the dynamic time warping is incorporated into the SFNMF, allowing the temporal alignment of the data sets onto the subspace spanned by the estimated nonnegative shared latent features amongst the two visual sequences. Extensive experimental results in two video databases demonstrate the effectiveness of the proposed methods in: 1) unsupervised detection of the temporal phases of posed and spontaneous facial events and 2) temporal alignment of facial expressions, outperforming by a large margin the state-of-the-art methods that they are compared to. Lazaros Zafeiriou, Yannis Panagakis, Maja Pantic, Stefanos Zafeiriou |
IEEE Trans. Image Process. | 2 |
| 2016 | Fusion and community detection in multi-layer graphsabstractRelational data arising in many domains can be represented by networks (or graphs) with nodes capturing entities and edges representing relationships between these entities. Community detection in networks has become one of the most important problems having a broad range of applications. Until recently, the vast majority of papers have focused on discovering community structures in a single network. However, with the emergence of multi-view network data in many real-world applications and consequently with the advent of multilayer graph representation, community detection in multi-layer graphs has become a new challenge. Multi-layer graphs provide complementary views of connectivity patterns of the same set of vertices. Fusion of the network layers is expected to achieve better clustering performance. In this paper, we propose two novel methods, coined as WSSNMTF (Weighted Simultaneous Symmetric Non-Negative Matrix Tri-Factorization) and NG-WSSNMTF (Natural Gradient WSSNMTF), for fusion and clustering of multi-layer graphs. Both methods are robust with respect to missing edges and noise. We compare the performance of the proposed methods with two baseline methods, as well as with three state-of-the-art methods on synthetic and three real-world datasets. The experimental results indicate superior performance of the proposed methods. Vladimir Gligorijevic, Yannis Panagakis, Stefanos Zafeiriou |
ICPR | 2 |
| 2016 | Back to the future: A fully automatic method for robust age progressionabstractIt has been shown that significant age difference between a probe and gallery face image can decrease the matching accuracy. If the face images can be normalized in age, there can be a huge impact on the face verification accuracy and thus many novel applications such as matching driver's license, passport and visa images with the real person's images can be effectively implemented. Face progression can address this issue by generating a face image for a specific age. Many researchers have attempted to address this problem focusing on predicting older faces from a younger face. In this paper, we propose a novel method for robust and automatic face progression in totally unconstrained conditions. Our method takes into account that faces belonging to the same age-groups share age patterns such as wrinkles while faces across different age-groups share some common patterns such as expressions and skin colors. Given training images of K different age-groups the proposed method learns to recover K low-rank age and one low-rank common components. These extracted components from the learning phase are used to progress an input face to younger as well as older ages in bidirectional fashion. Using standard datasets, we demonstrate that the proposed progression method outperforms state-of-the-art age progression methods and also improves matching accuracy in a face verification protocol that includes age progression. Christos Sagonas, Yannis Panagakis, Saritha Arunkumar, Nalini K. Ratha, Stefanos Zafeiriou |
ICPR | 2 |
| 2016 | Robust Correlated and Individual Component AnalysisabstractRecovering correlated and individual components of two, possibly temporally misaligned, sets of data is a fundamental task in disciplines such as image, vision, and behavior computing, with application to problems such as multi-modal fusion (via correlated components), predictive analysis, and clustering (via the individual ones). Here, we study the extraction of correlated and individual components under real-world conditions, namely i) the presence of gross non-Gaussian noise and ii) temporally misaligned data. In this light, we propose a method for the Robust Correlated and Individual Component Analysis (RCICA) of two sets of data in the presence of gross, sparse errors. We furthermore extend RCICA in order to handle temporal incongruities arising in the data. To this end, two suitable optimization problems are solved. The generality of the proposed methods is demonstrated by applying them onto 4 applications, namely i) heterogeneous face recognition, ii) multi-modal feature fusion for human behavior analysis (i.e., audio-visual prediction of interest and conflict), iii) face clustering, and iv) thetemporal alignment of facial expressions. Experimental results on 2 synthetic and 7 real world datasets indicate the robustness and effectiveness of the proposed methodson these application domains, outperforming other state-of-the-art methods in the field. Yannis Panagakis, Mihalis A. Nicolaou, Stefanos Zafeiriou, Maja Pantic |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Discriminant Incoherent Component AnalysisabstractFace images convey rich information which can be perceived as a superposition of low-complexity components associated with attributes, such as facial identity, expressions, and activation of facial action units (AUs). For instance, low-rank components characterizing neutral facial images are associated with identity, while sparse components capturing non-rigid deformations occurring in certain face regions reveal expressions and AU activations. In this paper, the discriminant incoherent component analysis (DICA) is proposed in order to extract low-complexity components, corresponding to facial attributes, which are mutually incoherent among different classes (e.g., identity, expression, and AU activation) from training data, even in the presence of gross sparse errors. To this end, a suitable optimization problem, involving the minimization of nuclear-and l1 -norm, is solved. Having found an ensemble of class-specific incoherent components by the DICA, an unseen (test) image is expressed as a group-sparse linear combination of these components, where the non-zero coefficients reveal the class(es) of the respective facial attribute(s) that it belongs to. The performance of the DICA is experimentally assessed on both synthetic and real-world data. Emphasis is placed on face analysis tasks, namely, joint face and expression recognition, face recognition under varying percentages of training data corruption, subject-independent expression recognition, and AU detection by conducting experiments on four data sets. The proposed method outperforms all the methods that are compared with all the tasks and experimental settings. Christos Georgakis 0001, Yannis Panagakis, Maja Pantic |
IEEE Trans. Image Process. | 2 |
| 2015 | Automatic construction Of robust spherical harmonic subspacesabstractIn this paper we propose a method to automatically recover a class specific low dimensional spherical harmonic basis from a set of in-the-wild facial images. We combine existing techniques for uncalibrated photometric stereo and low rank matrix decompositions in order to robustly recover a combined model of shape and identity. We build this basis without aid from a 3D model and show how it can be combined with recent efficient sparse facial feature localisation techniques to recover dense 3D facial shape. Unlike previous works in the area, our method is very efficient and is an order of magnitude faster to train, taking only a few minutes to build a model with over 2000 images. Furthermore, it can be used for real-time recovery of facial shape. Patrick Snape, Yannis Panagakis, Stefanos Zafeiriou |
CVPR | 2 |
| 2015 | Robust Statistical Face FrontalizationabstractRecently, it has been shown that excellent results can be achieved in both facial landmark localization and pose-invariant face recognition. These breakthroughs are attributed to the efforts of the community to manually annotate facial images in many different poses and to collect 3D facial data. In this paper, we propose a novel method for joint frontal view reconstruction and landmark localization using a small set of frontal images only. By observing that the frontal facial image is the one having the minimum rank of all different poses, an appropriate model which is able to jointly recover the frontalized version of the face as well as the facial landmarks is devised. To this end, a suitable optimization problem, involving the minimization of the nuclear norm and the matrix l1 norm is solved. The proposed method is assessed in frontal face reconstruction, face landmark localization, pose-invariant face recognition, and face verification in unconstrained conditions. The relevant experiments have been conducted on 8 databases. The experimental results demonstrate the effectiveness of the proposed method in comparison to the state-of-the-art methods for the target problems. Christos Sagonas, Yannis Panagakis, Stefanos Zafeiriou, Maja Pantic |
ICCV | 2 |
| 2015 | Face FlowabstractIn this paper, we propose a method for the robust and efficient computation of multi-frame optical flow in an expressive sequence of facial images. We formulate a novel energy minimisation problem for establishing dense correspondences between a neutral template and every frame of a sequence. We exploit the highly correlated nature of human expressions by representing dense facial motion using a deformation basis. Furthermore, we exploit the even higher correlation between deformations in a given input sequence by imposing a low-rank prior on the coefficients of the deformation basis, yielding temporally consistent optical flow. Our proposed model-based formulation, in conjunction with the inverse compositional strategy and low-rank matrix optimisation that we adopt, leads to a highly efficient algorithm for calculating facial flow. As experimental evaluation, we show quantitative experiments on a challenging novel benchmark of face sequences, with dense ground truth optical flow provided by motion capture data. We also provide qualitative results on a real sequence displaying fast motion and occlusions. Extensive quantitative and qualitative comparisons demonstrate that the proposed method outperforms state-of-the-art optical flow and dense non-rigid registration techniques, whilst running an order of magnitude faster. Patrick Snape, Anastasios Roussos, Yannis Panagakis, Stefanos Zafeiriou |
ICCV | 3 |
| 2014 | Generalised Scalable Robust Principal Component Analysis
Georgios Papamakarios, Yannis Panagakis, Stefanos Zafeiriou |
BMVC | 2 |
| 2014 | RAPS: Robust and Efficient Automatic Construction of Person-Specific Deformable ModelsabstractThe construction of Facial Deformable Models (FDMs) is a very challenging computer vision problem, since the face is a highly deformable object and its appearance drastically changes under different poses, expressions, and illuminations. Although several methods for generic FDMs construction, have been proposed for facial landmark localization in still images, they are insufficient for tasks such as facial behaviour analysis and facial motion capture where perfect landmark localization is required. In this case, person-specific FDMs (PSMs) are mainly employed, requiring manual facial landmark annotation for each person and person-specific training. In this paper, a novel method for the automatic construction of PSMs is proposed. To this end, an orthonormal subspace which is suitable for facial image reconstruction is learnt. Next, to correct the fittings of a generic model, image congealing (i.e., batch image aliment) is performed by employing only the learnt orthonormal subspace. Finally, the corrected fittings are used to construct the PSM. The image congealing problem is solved by formulating a suitable sparsity regularized rank minimization problem. The proposed method outperforms the state-of-the art methods that is compared to, in terms of both landmark localization accuracy and computational time. Christos Sagonas, Yannis Panagakis, Stefanos Zafeiriou, Maja Pantic |
CVPR | 2 |
| 2014 | Robust Canonical Correlation Analysis: Audio-visual fusion for learning continuous interestabstractThe problem of automatically estimating the interest level of a subject has been gaining attention by researchers, mostly due to the vast applicability of interest detection. In this work, we obtain a set of continuous interest annotations for the SE-MAINE database, which we analyse also in terms of emotion dimensions such as valence and arousal. Most importantly, we propose a robust variant of Canonical Correlation Analysis (RCCA) for performing audio-visual fusion, which we apply to the prediction of interest. RCCA recovers a low-rank subspace which captures the correlations of fused modalities, while isolating gross errors in the data without making any assumptions regarding Gaussianity. We experimentally show that RCCA is more appropriate than other standard fusion techniques (such as l2-CCA and feature-level fusion), since it both captures interactions between modalities while also decontaminating the obtained subspace from errors which are dominant in real-world problems. Mihalis A. Nicolaou, Yannis Panagakis, Stefanos Zafeiriou, Maja Pantic |
ICASSP | 2 |
| 2014 | Elastic Net subspace clustering applied to pop/rock music structure analysis
Yannis Panagakis, Constantine Kotropoulos |
Pattern Recognit. Lett. | 1 |
| 2014 | Music genre classification via joint sparse low-rank representation of audio featuresabstractA novel framework for music genre classification, namely the joint sparse low-rank representation (JSLRR) is proposed in order to: 1) smooth the noise in the test samples, and 2) identify the subspaces that the test samples lie onto. An efficient algorithm is proposed for obtaining the JSLRR and a novel classifier is developed, which is referred to as the JSLRR-based classifier. Special cases of the JSLRR-based classifier are the joint sparse representation-based classifier and the low-rank representation-based one. The performance of the three aforementioned classifiers is compared against that of the sparse representation-based classifier, the nearest subspace classifier, the support vector machines, and the nearest neighbor classifier for music genre classification on six manually annotated benchmark datasets. The best classification results reported here are comparable with or slightly superior than those obtained by the state-of-the-art music genre classification methods. Yannis Panagakis, Constantine Kotropoulos, Gonzalo R. Arce |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2013 | Robust Canonical Time Warping for the Alignment of Grossly Corrupted SequencesabstractTemporal alignment of human behaviour from visual data is a very challenging problem due to a numerous reasons, including possible large temporal scale differences, inter/intra subject variability and, more importantly, due to the presence of gross errors and outliers. Gross errors are often in abundance due to incorrect localization and tracking, presence of partial occlusion etc. Furthermore, such errors rarely follow a Gaussian distribution, which is the de-facto assumption in machine learning methods. In this paper, building on recent advances on rank minimization and compressive sensing, a novel, robust to gross errors temporal alignment method is proposed. While previous approaches combine the dynamic time warping (DTW) with low-dimensional projections that maximally correlate two sequences, we aim to learn two underlying projection matrices (one for each sequence), which not only maximally correlate the sequences but, at the same time, efficiently remove the possible corruptions in any datum in the sequences. The projections are obtained by minimizing the weighted sum of nuclear and ℓ1norms, by solving a sequence of convex optimization problems, while the temporal alignment is found by applying the DTW in an alternating fashion. The superiority of the proposed method against the state-of-the-art time alignment methods, namely the canonical time warping and the generalized time warping, is indicated by the experimental results on both synthetic and real datasets. Yannis Panagakis, Mihalis A. Nicolaou, Stefanos Zafeiriou, Maja Pantic |
CVPR | 1 |
| 2013 | Music recommendation using hypergraphs and group sparsityabstractA challenging problem in multimedia recommendation is to model a variety of relations, such as social, friend, listening, or tagging ones in a unified framework and to exploit all these sources of information. In this paper, music recommendation problem is expressed as a hypergraph ranking problem, introducing group sparsity constraints. By doing so, one can control how the different data groups (i.e., sets of hypergraph vertices) affect the recommendation process. Experiments on a dataset collected from Last.fm demonstrate that the accuracy is significantly increased by exploiting the group structure of the data. Preliminary results are also presented for Greek folk music recommendation. Antonis Theodoridis, Constantine Kotropoulos, Yannis Panagakis |
ICASSP | 3 |
| 2012 | Automatic music tagging by low-rank representationabstractA novel multi-label annotation method is proposed and applied to music tagging. Each music recording is represented by its auditory temporal modulations (ATMs). Given a set of training music recordings represented by the tag-music recording matrix having zero-one (indicator) vectors of the tags associated with each recording in its columns along with the matrix of the ATM representations in its columns, a low-rank weight matrix is sought, such that the tag-music recording matrix is expressed as the product of the weight matrix and the matrix of the ATM representations plus an error matrix. Clearly, such a weight matrix captures the relationships between the labels (i.e., tags) and the audio features. It can be derived by solving a convex nuclear norm minimization problem, if the tag-music recording matrix and the matrix of the ATM representations are assumed to be jointly low-rank. Having found the weight matrix, the annotation vector for labeling any test music recording can be obtained by multiplying the weight matrix with its ATM representation. The just outlined method is referred to as low-rank representation based multi-label annotation (LRRMA). The LRRMA outperforms the state-of-the-art auto-tagging systems, when applied to the CAL500 dataset in a 5-fold cross-validation experimental protocol. Yannis Panagakis, Constantine Kotropoulos |
ICASSP | 1 |
| 2011 | Automatic music tagging via PARAFAC2abstractAutomatic music tagging is addressed by resorting to auditory temporal modulations and Parallel Factor Analysis 2 (PARAFAC2). The starting point is to represent each music recording by its auditory temporal modulations. Then, an irregular third order tensor is formed. The first slice contains the vectorized training temporal modulations, while the second slice contains the corresponding multi-label vectors. The PARAFAC2 is employed to effectively harness the multi-label information for dimensionality reduction. Any vectorized test auditory representation of temporal modulations is first projected onto the semantic space derived via the PARAFAC2 and the coefficient vector is obtained. Then, the annotation vector is obtained by multiplying this coefficient vector by the left singular vectors of the second slice (i.e., the slice associated to the label vector). The proposed framework, outperforms the state-of-the-art auto-tagging systems, when applied to the CAL500 dataset in a 10-fold cross-validation experimental protocol. Yannis Panagakis, Constantine Kotropoulos |
ICASSP | 1 |
| 2010 | Music genre classification via Topology Preserving Non-Negative Tensor Factorization and sparse representationsabstractMotivated by the rich, psycho-physiologically grounded properties of auditory cortical representations and the power of sparse representation-based classifiers, we propose a robust music genre classification framework. Its first pilar is a novel multilinear subspace analysis method that reduces the dimensionality of cortical representations of music signals, while preserving the topology of the cortical representations. Its second pilar is the sparse representation based classification, that models any test cortical representation as a sparse weighted sum of dictionary atoms, which stem from training cortical representations of known genre, by assuming that the representations of music recordings of the same genre are close enough in the tensor space they lie. Accordingly, the dimensionality reduction is made in a compatible manner to the working principle of the sparse-representation based classification. Music genre classification accuracy of 93.7% and 94.93% is reported on the GTZAN and the ISMIR2004 Genre datasets, respectively. Both accuracies outperform any accuracy ever reported for state of the art music genre classification algorithms applied to the aforementioned datasets. Yannis Panagakis, Constantine Kotropoulos |
ICASSP | 1 |
| 2010 | Ensemble Discriminant Sparse Projections Applied to Music Genre ClassificationabstractResorting to the rich, psycho-physiologically grounded, properties of the slow temporal modulations of music recordings, a novel classifier ensemble is built, which applies discriminant sparse projections. More specifically, over complete dictionaries are learned and sparse coefficient vectors are extracted to optimally approximate the slow temporal modulations of the training music recordings. The sparse coefficient vectors are then projected to the principal subspaces of their within-class and between-class covariance matrices. Decisions are taken with respect to the minimum Euclidean distance from the class mean sparse coefficient vectors, which undergo the aforementioned projections. The application of majority voting to the decisions taken by 10 individual classifiers, which are trained on the 10 training folds defined by stratified 10-fold cross-validation on the GTZAN dataset, yields a music genre classification accuracy of 84.96% on average. The latter exceeds by 2.46% the highest accuracy previously reported without employing any sparse representations. Constantine Kotropoulos, Gonzalo R. Arce, Yannis Panagakis |
ICPR | 3 |
| 2010 | Non-Negative Multilinear Principal Component Analysis of Auditory Temporal Modulations for Music Genre ClassificationabstractMotivated by psychophysiological investigations on the human auditory system, a bio-inspired two-dimensional auditory representation of music signals is exploited, that captures the slow temporal modulations. Although each recording is represented by a second-order tensor (i.e., a matrix), a third-order tensor is needed to represent a music corpus. Non-negative multilinear principal component analysis (NMPCA) is proposed for the unsupervised dimensionality reduction of the third-order tensors. The NMPCA maximizes the total tensor scatter while preserving the non-negativity of auditory representations. An algorithm for NMPCA is derived by exploiting the structure of the Grassmann manifold. The NMPCA is compared against three multilinear subspace analysis techniques, namely the non-negative tensor factorization, the high-order singular value decomposition, and the multilinear principal component analysis as well as their linear counterparts, i.e., the non-negative matrix factorization, the singular value decomposition, and the principal components analysis in extracting features that are subsequently classified by either support vector machine or nearest neighbor classifiers. Three different sets of experiments conducted on the GTZAN and the ISMIR2004 Genre datasets demonstrate the superiority of NMPCA against the aforementioned subspace analysis techniques in extracting more discriminating features, especially when the training set has small cardinality. The best classification accuracies reported in the paper exceed those obtained by the state-of-the-art music genre classification algorithms applied to both datasets. Yannis Panagakis, Constantine Kotropoulos, Gonzalo R. Arce |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | Audio-Assisted Movie Dialogue DetectionabstractAbstract—An audio-assisted system is investigated that detects if a movie scene is a dialogue or not. The system is based on actor indicator functions. That is, functions which define if an actor speaks at a certain time instant. In particular, the cross-correlation and the magnitude of the corresponding the cross-power spectral density of a pair of indicator functions are input to various classifiers, such as voted perceptrons, radial basis function networks, random trees, and support vector machines for dialogue/non-dialogue detection. To boost classifier efficiency AdaBoost is also exploited. The aforementioned classifiers are trained using ground truth indicator functions determined by human annotators for 41 dialogue and another 20 non-dialogue audio instances. For testing, actual indicator functions are derived by applying audio activity detection and actor clustering to audio recordings. 23 instances are randomly chosen among the aforementioned 41 dialogue instances, 17 of which correspond to dialogue scenes and 6 to non-dialogue ones. Accuracy ranging between 0.739 and 0.826 is reported. Index Terms—Audio activity detection, cross-correlation, crosspower spectral density, dialogue detection, indicator functions, speaker clustering. I. Margarita Kotti, Dimitrios Ververidis, Georgios Evangelopoulos, Yannis Panagakis, Constantine Kotropoulos, Petros Maragos, Ioannis Pitas |
IEEE Trans. Circuits Syst. Video Technol. | 4 |