VLDB 2026 Research / reviewers in the wild / expert
Mihalis A. Nicolaou
dblp:32/8615 · also Mihalis Nicolaou
· DBLP profile ↗
44ranked-venue papers
11as first author
17since 2021 · last 2026
0000-0001-9175-477XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 8 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Artificial intelligence-based heliostat characterisation through synthetic dataabstractThis work proposes a fast, non-intrusive methodology for heliostat mirror surface characterisation based on Deep Learning (DL) models trained exclusively on synthetic data. The approach uses solar heat flux distribution images generated through ray-tracing simulations at multiple sun positions to predict surface deformations. Among the evaluated architectures, a U-Net model incorporating injected heliostat and sun position data achieved the best performance, with a validation mean square error (MSE) of 0.0207, representing an 8% improvement over the next-best, which was the convolutional autoencoder. A refinement step using the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) is applied to further improve the prediction accuracy of local deformations, reducing the MSE by approximately 40%, albeit at an additional computational cost of around six hours per heliostat. The methodology is validated by applying a known real deformation to multiple heliostats at different positions within a representative solar tower field. For each heliostat, flux distribution images are obtained across different sun positions using ray-tracing software. These synthetic flux images then serve as the primary input for the DL model. Results show that the DL model can reconstruct the deformation pattern with high accuracy, (MSE loss of 0.015 for a real deformed heliostat) and near-instantaneous inference time. The proposed pipeline offers a scalable and efficient alternative to conventional optical calibration techniques. It is particularly suited for field-wide deployment in Concentrating Solar Power (CSP) systems, provided certain operational conditions are met, including clear-sky conditions, single-facet heliostat configurations, and minimal mirror soiling. Jorge Moreno García-Moreno, Kypros Milidonis, Georgia Kourmouli, Mihalis A. Nicolaou |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Disentangling Local and Global Semantics in Diffusion Models for Image EditingabstractAbstract Diffusion models have achieved state-of-the-art image synthesis, yet unlike GANs, they lack a well-structured latent space for intuitive image editing. Existing diffusion-based editing methods often rely on supervised fine-tuning or text-based guidance, while recent unsupervised techniques leveraging the model’s bottleneck layer suffer from one or more key limitations: (i) they focus only on global attributes, (ii) fail to disentangle local and global semantics, or (iii) require extensive human intervention. To fill this gap, we first propose an unsupervised method for localized image editing in pre-trained unconditional diffusion models that disentangles local and global semantics in the model’s latent space. Given an input image and a user-specified region of interest, our approach uses the denoising network’s Jacobian to map that region to a corresponding latent subspace. We then separate this subspace into shared (global) and region-specific components to uncover latent directions that control local attributes. These directions generalize across images, enabling semantically consistent edits without retraining. We go one step further by extending our method to minimize manual supervision by automatically inferring edit directions from a single reference image and generating region masks without human input. Experiments on multiple datasets show that our method yields more localized, high-fidelity edits than state-of-the-art approaches. Manos Plitsis, Theodoros Kouzelis, Panagiotis Koromilas, Vassilis Katsouros, Mihalis A. Nicolaou, Yannis Panagakis |
Int. J. Comput. Vis. | 5 |
| 2025 | Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of DecodersabstractMultilayer perceptrons (MLPs) are an integral part of large language models, yet their dense representations render them difficult to understand, edit, and steer. Recent methods learn interpretable approximations via neuron-level sparsity, yet fail to faithfully reconstruct the original mapping--significantly increasing model's next-token cross-entropy loss. In this paper, we advocate for moving to layer-level sparsity to overcome the accuracy trade-off in sparse layer approximation. Under this paradigm, we introduce Mixture of Decoders (MxDs). MxDs generalize MLPs and Gated Linear Units, expanding pre-trained dense layers into tens of thousands of specialized sublayers. Through a flexible form of tensor factorization, each sparsely activating MxD sublayer implements a linear transformation with full-rank weights--preserving the original decoders' expressive capacity even under heavy sparsity. Experimentally, we show that MxDs significantly outperform state-of-the-art methods (e.g., Transcoders) on the sparsity-accuracy frontier in language models with up to 3B parameters. Further evaluations on sparse probing and feature steering demonstrate that MxDs learn similarly specialized features of natural language--opening up a promising new avenue for designing interpretable yet faithful decompositions. Our code is included at: https://github.com/james-oldfield/MxD. James Oldfield 0001, Shawn Im, Yixuan Li 0001, Mihalis A. Nicolaou, Ioannis Patras, Grigorios Chrysos 0002 |
NeurIPS | 4 |
| 2024 | Enabling Local Editing in Diffusion Models by Joint and Individual Component Analysis
Theodoros Kouzelis, Emmanouil Plitsis, Mihalis A. Nicolaou, Yannis Panagakis |
BMVC | 3 |
| 2024 | Bridging Mini-Batch and Asymptotic Analysis in Contrastive Learning: From InfoNCE to Kernel-Based LossesabstractWhat do different contrastive learning (CL) losses actually optimize for? Although multiple CL methods have demonstrated remarkable representation learning capabilities, the differences in their inner workings remain largely opaque. In this work, we analyse several CL families and prove that, under certain conditions, they admit the same minimisers when optimizing either their batch-level objectives or their expectations asymptotically. In both cases, an intimate connection with the hyperspherical energy minimisation (HEM) problem resurfaces. Drawing inspiration from this, we introduce a novel CL objective, coined Decoupled Hyperspherical Energy Loss (DHEL). DHEL simplifies the problem by decoupling the target hyperspherical energy from the alignment of positive examples while preserving the same theoretical guarantees. Going one step further, we show the same results hold for another relevant CL family, namely kernel contrastive learning (KCL), with the additional advantage of the expected loss being independent of batch size, thus identifying the minimisers in the non-asymptotic regime. Empirical results demonstrate improved downstream performance and robustness across combinations of different batch sizes and hyperparameters and reduced dimensionality collapse, on several computer vision datasets. Panagiotis Koromilas, Giorgos Bouritsas, Theodoros Giannakopoulos, Mihalis A. Nicolaou, Yannis Panagakis |
ICML | 4 |
| 2024 | Multilinear Mixture of Experts: Scalable Expert Specialization through FactorizationabstractThe Mixture of Experts (MoE) paradigm provides a powerful way to decompose dense layers into smaller, modular computations often more amenable to human interpretation, debugging, and editability. However, a major challenge lies in the computational cost of scaling the number of experts high enough to achieve fine-grained specialization. In this paper, we propose the Multilinear Mixture of Experts (μMoE) layer to address this, focusing on vision models. μMoE layers enable scalable expert specialization by performing an implicit computation on prohibitively large weight tensors entirely in factorized form. Consequently, μMoEs (1) avoid the restrictively high inference-time costs of dense MoEs, yet (2) do not inherit the training issues of the popular sparse MoEs' discrete (non-differentiable) expert routing. We present both qualitative and quantitative evidence that scaling μMoE layers when fine-tuning foundation models for vision tasks leads to more specialized experts at the class-level, further enabling manual bias correction in CelebA attribute classification. Finally, we show qualitative results demonstrating the expert specialism achieved when pre-training large GPT2 and MLP-Mixer models with parameter-matched μMoE blocks at every layer, maintaining comparable accuracy. Our code is available at: https://github.com/james-oldfield/muMoE. James Oldfield 0001, Markos Georgopoulos, Grigorios Chrysos 0002, Christos Tzelepis, Yannis Panagakis, Mihalis A. Nicolaou, Jiankang Deng, Ioannis Patras |
NeurIPS | 6 |
| 2024 | Locality-Preserving Directions for Interpreting the Latent Space of Satellite Image GANsabstractWe present a locality-aware method for interpreting the latent space of wavelet-based generative adversarial networks (GANs), that can well capture the large spatial and spectral variability that is characteristic of satellite imagery. By focusing on preserving locality, the proposed method is able to decompose the weight-space of pre-trained GANs and recover interpretable directions that correspond to high-level semantic concepts (such as urbanization, structure density, flora presence)-that can subsequently be used for guided synthesis of satellite imagery. In contrast to typically used approaches that focus on capturing the variability of the weight-space in a reduced dimensionality space (i.e., based on principal component analysis (PCA)), we show that preserving locality leads to vectors with different angles, that are more robust to artifacts and can better preserve class information. Via a set of quantitative and qualitative examples, we further show that the proposed approach can outperform both baseline geometric augmentations, as well as global, PCA-based approaches for data synthesis in the context of data augmentation for satellite scene classification. Georgia Kourmouli, Nikos Kostagiolas, Mihalis A. Nicolaou, Yannis Panagakis |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Bilinear Models of Parts and Appearances in Generative Adversarial NetworksabstractRecent advances in the understanding of Generative Adversarial Networks (GANs) have led to remarkable progress in visual editing and synthesis tasks, capitalizing on the rich semantics that are embedded in the latent spaces of pre-trained GANs. However, existing methods are often tailored to specific GAN architectures and are limited to either discovering global semantic directions that do not facilitate localized control, or require some form of supervision through manually provided regions or segmentation masks. In this light, we present an architecture-agnostic approach that jointly discovers factors representing spatial parts and their appearances in an entirely unsupervised fashion. These factors are obtained by applying a semi-nonnegative tensor factorization on the feature maps, which in turn enables context-aware local image editing with pixel-level control. In addition, we show that the discovered appearance factors correspond to saliency maps that localize concepts of interest, without using any labels. Experiments on a wide range of GAN architectures and datasets show that, in comparison to the state of the art, our method is far more efficient in terms of training time and, most importantly, provides much more accurate localized control. James Oldfield 0001, Christos Tzelepis, Yannis Panagakis, Mihalis A. Nicolaou, Ioannis Patras |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | MMATR: A Lightweight Approach for Multimodal Sentiment Analysis Based on Tensor MethodsabstractDespite the considerable research output on Multimodal Learning for Affect-related tasks, most of the current methods are very complex in terms of the number of trainable parameters, and thus do not constitute effective solutions for real-life applications. In this work we try to alleviate this gap in the literature by introducing the Multimodal Attention Tensor Regression (MMATR) network, a lightweight model that is based on: (i) a static input representation (2D matrix of dimensions time × features) for each modality, which helps to avoid high-parameterized sequential models by incorporating a CNN, (ii) the replacement of the usual pooling and flattening operations as well as the linear layers by tensor contraction and tensor regression layers that are able to reduce the number of parameters, while keeping the high-order structure of the multimodal data, and (iii) a bimodal attention layer that learns multimodal co-occurrences. By a set of experiments comparing with a variety of state-of-the-art techniques, we show that the proposed MMATR can achieve results competitive to the state-of-the-art in the task of Multimodal Sentiment Analysis, albeit having four orders of magnitude fewer parameters. Panagiotis Koromilas, Mihalis A. Nicolaou, Theodoros Giannakopoulos, Yannis Panagakis |
ICASSP | 2 |
| 2023 | PandA: Unsupervised Learning of Parts and Appearances in the Feature Maps of GANs
James Oldfield 0001, Christos Tzelepis, Yannis Panagakis, Mihalis A. Nicolaou, Ioannis Patras |
ICLR | 4 |
| 2023 | Parts of Speech-Grounded Subspaces in Vision-Language ModelsabstractLatent image representations arising from vision-language models have proved immensely useful for a variety of downstream tasks. However, their utility is limited by their entanglement with respect to different visual attributes. For instance, recent work has shown that CLIP image representations are often biased toward specific visual properties (such as objects or actions) in an unpredictable manner. In this paper, we propose to separate representations of the different visual modalities in CLIP’s joint vision-language space by leveraging the association between parts of speech and specific visual modes of variation (e.g. nouns relate to objects, adjectives describe appearance). This is achieved by formulating an appropriate component analysis model that learns subspaces capturing variability corresponding to a specific part of speech, while jointly minimising variability to the rest. Such a subspace yields disentangled representations of the different visual properties of an image or text in closed form while respecting the underlying geometry of the manifold on which the representations lie. What’s more, we show the proposed model additionally facilitates learning subspaces corresponding to specific visual appearances (e.g. artists’ painting styles), which enables the selective removal of entire visual themes from CLIP-based text-to-image synthesis. We validate the model both qualitatively, by visualising the subspace projections with a text-to-image model and by preventing the imitation of artists’ styles, and quantitatively, through class invariance metrics and improvements to baseline zero-shot classification. James Oldfield 0001, Christos Tzelepis, Yannis Panagakis, Mihalis A. Nicolaou, Ioannis Patras |
NeurIPS | 4 |
| 2023 | Toward Explainable and Transferable Deep Downscaling of Atmospheric PollutantsabstractUnderstanding the intricate relationship between air quality and human health necessitates a comprehensive assessment of the spatial distribution of pollutants at high resolutions. Here, we employ deep learning methodologies for downscaling atmospheric concentrations of pollutants demonstrated to adversely affect human health (NO2and PM2.5). We train a model using data fusing multi-source information from the Copernicus Atmosphere Monitoring Service (CAMS) and in-situ observations from ground-based stations in North Italy. Our model demonstrates robust generalisation capabilities by effectively improving bias compared to ground-truth station observations when applied to the state of California, at an order of magnitude different downscaling resolution, despite not previously being trained on or exposed to this region. In addition to the demonstrated transferability, we developed and applied an occlusion-based method over the fused data sources for interpretability. Our results suggest that our model adeptly leverages auxiliary informed data encompassing past, present and future insights for inference. Thus, we are able to quantify the influence of the input variables on both predictability and on the uncertainty associated with the physical deterministic model. Giannis Ashiotis, Pantelis Georgiades, Theodoros Christoudias, Mihalis A. Nicolaou |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Deep Learning on the Sphere for Multi-model Ensembling of Significant Wave HeightabstractWhen working with geophysical variables on a global scale, a solution for processing data on the surface of a sphere is needed. At the same time, region-specific dynamics that deviate from the general behavior across the globe also need to be accounted for. Addressing these two necessities, we propose the first Deep Learning approach for multi-model ensembling that operates directly on the sphere. Our methodology allows to progressively allocate region-specific model complexity, guided by the clustering of the model forecasting errors. We evaluate our proposed method on a multi-model ensembling application of significant wave height, where the proposed method is shown to outperform 2D CNNs with less than half the parameters needed, while producing comparable results to models with more than 10 times the number of parameters. Andrea Littardi, Anders Hildeman, Mihalis A. Nicolaou |
ICASSP | 3 |
| 2022 | Adversarial Learning of Disentangled and Generalizable Representations of Visual AttributesabstractRecently, a multitude of methods for image-to-image translation have demonstrated impressive results on problems, such as multidomain or multiattribute transfer. The vast majority of such works leverages the strengths of adversarial learning and deep convolutional autoencoders to achieve realistic results by well-capturing the target data distribution. Nevertheless, the most prominent representatives of this class of methods do not facilitate semantic structure in the latent space and usually rely on binary domain labels for test-time transfer. This leads to rigid models, unable to capture the variance of each domain label. In this light, we propose a novel adversarial learning method that: 1) facilitates the emergence of latent structure by semantically disentangling sources of variation and 2) encourages learning generalizable, continuous, and transferable latent codes that enable flexible attribute mixing. This is achieved by introducing a novel loss function that encourages representations to result in uniformly distributed class posteriors for disentangled attributes. In tandem with an algorithm for inducing generalizable properties, the resulting representations can be utilized for a variety of tasks such as intensity-preserving multiattribute image translation and synthesis, without requiring labeled test data. We demonstrate the merits of the proposed method by a set of qualitative and quantitative experiments on popular databases such as MultiPIE, RaFD, and BU-3DFE, where our method outperforms other state-of-the-art methods in tasks such as intensity-preserving multiattribute transfer and synthesis. James Oldfield 0001, Yannis Panagakis, Mihalis A. Nicolaou |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Tensor Component Analysis for Interpreting the Latent Space of GANs
James Oldfield 0001, Markos Georgopoulos, Yannis Panagakis, Mihalis A. Nicolaou, Ioannis Patras |
BMVC | 4 |
| 2021 | Mitigating Demographic Bias in Facial Datasets with Style-Based Multi-attribute TransferabstractAbstract Deep learning has catalysed progress in tasks such as face recognition and analysis, leading to a quick integration of technological solutions in multiple layers of our society. While such systems have proven to be accurate by standard evaluation metrics and benchmarks, a surge of work has recently exposed the demographic bias that such algorithms exhibit–highlighting that accuracy does not entail fairness. Clearly, deploying biased systems under real-world settings can have grave consequences for affected populations. Indeed, learning methods are prone to inheriting, or even amplifying the bias present in a training set, manifested by uneven representation across demographic groups. In facial datasets, this particularly relates to attributes such as skin tone, gender, and age. In this work, we address the problem of mitigating bias in facial datasets by data augmentation. We propose a multi-attribute framework that can successfully transfer complex, multi-scale facial patterns even if these belong to underrepresented groups in the training set. This is achieved by relaxing the rigid dependence on a single attribute label, and further introducing a tensor-based mixing structure that captures multiplicative interactions between attributes in a multilinear fashion. We evaluate our method with an extensive set of qualitative and quantitative experiments on several datasets, with rigorous comparisons to state-of-the-art methods. We find that the proposed framework can successfully mitigate dataset bias, as evinced by extensive evaluations on established diversity metrics, while significantly improving fairness metrics such as equality of opportunity. Markos Georgopoulos, James Oldfield 0001, Mihalis A. Nicolaou, Yannis Panagakis, Maja Pantic |
Int. J. Comput. Vis. | 3 |
| 2021 | Tensor Methods in Computer Vision and Deep LearningabstractTensors, or multidimensional arrays, are data structures that can naturally represent visual data of multiple dimensions. Inherently able to efficiently capture structured, latent semantic spaces and high-order interactions, tensors have a long history of applications in a wide span of computer vision problems. With the advent of the deep learning paradigm shift in computer vision, tensors have become even more fundamental. Indeed, essential ingredients in modern deep learning architectures, such as convolutions and attention mechanisms, can readily be considered as tensor mappings. In effect, tensor methods are increasingly finding significant applications in deep learning, including the design of memory and compute efficient network architectures, improving robustness to random noise and adversarial attacks, and aiding the theoretical understanding of deep networks. This article provides an in-depth and practical review of tensors and tensor methods in the context of representation learning and deep learning, with a particular focus on visual data analysis and computer vision applications. Concretely, besides fundamental work in tensor-based visual data analysis methods, we focus on recent developments that have brought on a gradual increase in tensor methods, especially in deep learning architectures and their implications in computer vision applications. To further enable the newcomer to grasp such concepts quickly, we provide companion Python notebooks, covering key aspects of this article and implementing them, step-by-step with TensorLy. Yannis Panagakis, Jean Kossaifi, Grigorios Chrysos 0002, James Oldfield 0001, Mihalis A. Nicolaou, Anima Anandkumar, Stefanos Zafeiriou |
Proc. IEEE | 5 |
| 2020 | 3DFaceGAN: Adversarial Nets for 3D Face Representation, Generation, and TranslationabstractAbstract Over the past few years, Generative Adversarial Networks (GANs) have garnered increased interest among researchers in Computer Vision, with applications including, but not limited to, image generation, translation, imputation, and super-resolution. Nevertheless, no GAN-based method has been proposed in the literature that can successfully represent, generate or translate 3D facial shapes (meshes). This can be primarily attributed to two facts, namely that (a) publicly available 3D face databases are scarce as well as limited in terms of sample size and variability (e.g., few subjects, little diversity in race and gender), and (b) mesh convolutions for deep networks present several challenges that are not entirely tackled in the literature, leading to operator approximations and model instability, often failing to preserve high-frequency components of the distribution. As a result, linear methods such as Principal Component Analysis (PCA) have been mainly utilized towards 3D shape analysis, despite being unable to capture non-linearities and high frequency details of the 3D face—such as eyelid and lip variations. In this work, we present 3DFaceGAN, the first GAN tailored towards modeling the distribution of 3D facial surfaces, while retaining the high frequency details of 3D face shapes. We conduct an extensive series of both qualitative and quantitative experiments, where the merits of 3DFaceGAN are clearly demonstrated against other, state-of-the-art methods in tasks such as 3D shape representation, generation, and translation. Stylianos Moschoglou, Stylianos Ploumpis, Mihalis A. Nicolaou, Athanasios Papaioannou, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 3 |
| 2019 | Multimodal Joint Emotion and Game Context Recognition in League of Legends LivestreamsabstractVideo game streaming provides the viewer with a rich set of audio-visual data, conveying information both with regards to the game itself, through game footage and audio, as well as the streamer's emotional state and behaviour via webcam footage and audio. Analysing player behaviour and discovering correlations with game context is crucial for modelling and understanding important aspects of livestreams, but comes with a significant set of challenges - such as fusing multimodal data captured by different sensors in uncontrolled (`in-the-wild') conditions. Firstly, we present, to our knowledge, the first data set of League of Legends livestreams, annotated for both streamer affect and game context. Secondly, we propose a method that exploits tensor decompositions for high-order fusion of multimodal representations. The proposed method is evaluated on the problem of jointly predicting game context and player affect, compared with a set of baseline fusion approaches such as late and early fusion. Data and code are available at https://github.com/charlieringer/LoLEmoGameRecognition. Charles Ringer, James Alfred Walker, Mihalis A. Nicolaou |
CoG | 3 |
| 2019 | Time-series Clustering with Jointly Learning Deep Representations, Clusters and Temporal BoundariesabstractClustering and segmentation of temporal data is an important task across several fields, with prominent applications in computer vision and machine learning such as face and gesture segmentation. Several related methods have been proposed in literature, focusing on learning temporal boundaries and clusters, with recent works focusing on learning deep representations for clustering. However, none of the proposed methods is suitable for jointly learning segments, clusters, as well as representations. In this paper, we propose the first methodology that simultaneously discovers suitable deep representations, as well as clusters and temporal boundaries, with the clustering process providing supervisory cues for updating temporal boundaries and training the proposed deep learning architecture. We demonstrate the power of the proposed approach on a human motion segmentation task using the CMU-MMAC database. Our method provides the best results with respect to normalized mutual information compared to other clustering algorithms. Panagiotis Tzirakis, Mihalis A. Nicolaou, Björn W. Schuller, Stefanos Zafeiriou |
FG | 2 |
| 2019 | Deep Affect Prediction in-the-Wild: Aff-Wild Database and Challenge, Deep Architectures, and BeyondabstractAutomatic understanding of human affect using visual signals is of great importance in everyday human–machine interactions. Appraising human emotional states, behaviors and reactions displayed in real-world settings, can be accomplished using latent continuous dimensions (e.g., the circumplex model of affect). Valence (i.e., how positive or negative is an emotion) and arousal (i.e., power of the activation of the emotion) constitute popular and effective representations for affect. Nevertheless, the majority of collected datasets this far, although containing naturalistic emotional states, have been captured in highly controlled recording conditions. In this paper, we introduce the Aff-Wild benchmark for training and evaluating affect recognition algorithms. We also report on the results of the First Affect-in-the-wild Challenge (Aff-Wild Challenge) that was recently organized in conjunction with CVPR 2017 on the Aff-Wild database, and was the first ever challenge on the estimation of valence and arousal in-the-wild. Furthermore, we design and extensively train an end-to-end deep neural architecture which performs prediction of continuous emotion dimensions based on visual cues. The proposed deep learning architecture, AffWildNet, includes convolutional and recurrent neural network layers, exploiting the invariant properties of convolutional features, while also modeling temporal dynamics that arise in human behavior via the recurrent layers. The AffWildNet produced state-of-the-art results on the Aff-Wild Challenge. We then exploit the AffWild database for learning features, which can be used as priors for achieving best performances both for dimensional, as well as categorical emotion recognition, using the RECOLA, AFEW-VA and EmotiW 2017 datasets, compared to all other methods designed for the same goal. The database and emotion recognition models are available at http://ibug.doc.ic.ac.uk/resources/first-affect-wild-challenge . Dimitris Kollias, Panagiotis Tzirakis, Mihalis A. Nicolaou, Athanasios Papaioannou, Guoying Zhao 0001, Björn W. Schuller, Irene Kotsia, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 3 |
| 2019 | Editorial of Special Issue on Human Behaviour Analysis "In-the-Wild"abstractThe papers in this special section focus on human face and body image analysis, one of the most researched objects. One of the main reasons behind this popularity lies in the numerous applications of automatic face and body gesture analysis algorithms, that span several fields such as Human-Computer and Human-Robot Interaction (facial expression/body gesture recognition for automatic analysis of affect), medicine and healthcare (detection of emotional and cognitive disorders), as well as biometrics (face recognition, gait recognition). The papers in this section focus on recent efforts towards catalysing progress in automatic analysis of human behaviour in uncontrolled, “in-the-wild” conditions. We summarize research efforts towards the development of research methodologies, database collections and benchmarks, as well as algorithms and systems for machine analysis of human behaviour, focusing on facial expressions, body gestures, speech, as well as various other sensors. We are delighted that the special issue includes authors both from academia as well as the industry. Mihalis A. Nicolaou, Stefanos Zafeiriou, Irene Kotsia, Guoying Zhao 0001, Jeffrey F. Cohn |
IEEE Trans. Affect. Comput. | 1 |
| 2018 | Multi-Attribute Probabilistic Linear Discriminant Analysis for 3D Facial Shapes
Stylianos Moschoglou, Stylianos Ploumpis, Mihalis A. Nicolaou, Stefanos Zafeiriou |
ACCV (3) | 3 |
| 2018 | Less is More: Univariate Modelling to Detect Early Parkinson's Disease from Keystroke Dynamics
Antony Milne, Katayoun Farrahi, Mihalis A. Nicolaou |
DS | 3 |
| 2018 | Deep unsupervised multi-view detection of video game stream highlightsabstractWe consider the problem of automatic highlight-detection in video game streams. Currently, the vast majority of highlight-detection systems for games are triggered by the occurrence of hard-coded game events (e.g., score change, end-game), while most advanced tools and techniques are based on detection of highlights via visual analysis of game footage. We argue that in the context of game streaming, events that may constitute highlights are not only dependent on game footage, but also on social signals that are conveyed by the streamer during the play session (e.g., when interacting with viewers, or when commenting and reacting to the game). In this light, we present a multi-view unsupervised deep learning methodology for novelty-based highlight detection. The method jointly analyses both game footage and social signals such as the players facial expressions and speech, and shows promising results for generating highlights on streams of popular games such as Player Unknown's Battlegrounds. Charles Ringer, Mihalis A. Nicolaou |
FDG | 2 |
| 2018 | Deep Canonical Time Warping for Simultaneous Alignment and Representation Learning of SequencesabstractMachine learning algorithms for the analysis of time-series often depend on the assumption that utilised data are temporally aligned. Any temporal discrepancies arising in the data is certain to lead to ill-generalisable models, which in turn fail to correctly capture properties of the task at hand. The temporal alignment of time-series is thus a crucial challenge manifesting in a multitude of applications. Nevertheless, the vast majority of algorithms oriented towards temporal alignment are either applied directly on the observation space or simply utilise linear projections-thus failing to capture complex, hierarchical non-linear representations that may prove beneficial, especially when dealing with multi-modal data (e.g., visual and acoustic information). To this end, we present Deep Canonical Time Warping (DCTW), a method that automatically learns non-linear representations of multiple time-series that are (i) maximally correlated in a shared subspace, and (ii) temporally aligned. Furthermore, we extend DCTW to a supervised setting, where during training, available labels can be utilised towards enhancing the alignment process. By means of experiments on four datasets, we show that the representations learnt significantly outperform state-of-the-art methods in temporal alignment, elegantly handling scenarios with heterogeneous feature sets, such as the temporal alignment of acoustic and visual information. George Trigeorgis, Mihalis A. Nicolaou, Björn W. Schuller, Stefanos Zafeiriou |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Dynamic Probabilistic Linear Discriminant Analysis for video classificationabstractComponent Analysis (CA) comprises of statistical techniques that decompose signals into appropriate latent components, relevant to a task-at-hand (e.g., clustering, segmentation, classification). Recently, an explosion of research in CA has been witnessed, with several novel probabilistic models proposed (e.g., Probabilistic Principal CA, Probabilistic Linear Discriminant Analysis (PLDA), Probabilistic Canonical Correlation Analysis). PLDA is a popular generative probabilistic CA method, that incorporates knowledge regarding class-labels and furthermore introduces class-specific and sample-specific latent spaces. While PLDA has been shown to outperform several state-of-the-art methods, it is nevertheless a static model; any feature-level temporal dependencies that arise in the data are ignored. As has been repeatedly shown, appropriate modelling of temporal dynamics is crucial for the analysis of temporal data (e.g., videos). In this light, we propose the first, to the best of our knowledge, probabilistic LDA formulation that models dynamics, the so-called Dynamic-PLDA (DPLDA). DPLDA is a generative model suitable for video classification and is able to jointly model the label information (e.g., face identity, consistent over videos of the same subject), as well as dynamic variations of each individual video. Experiments on video classification tasks such as face and facial expression recognition show the efficacy of the proposed method. Alessandro Fabris, Mihalis A. Nicolaou, Irene Kotsia, Stefanos Zafeiriou |
ICASSP | 2 |
| 2016 | Deep Canonical Time WarpingabstractMachine learning algorithms for the analysis of timeseries often depend on the assumption that the utilised data are temporally aligned. Any temporal discrepancies arising in the data is certain to lead to ill-generalisable models, which in turn fail to correctly capture the properties of the task at hand. The temporal alignment of time-series is thus a crucial challenge manifesting in a multitude of applications. Nevertheless, the vast majority of algorithms oriented towards the temporal alignment of time-series are applied directly on the observation space, or utilise simple linear projections. Thus, they fail to capture complex, hierarchical non-linear representations which may prove to be beneficial towards the task of temporal alignment, particularly when dealing with multi-modal data (e.g., aligning visual and acoustic information). To this end, we present the Deep Canonical Time Warping (DCTW), a method which automatically learns complex non-linear representations of multiple time-series, generated such that (i) they are highly correlated, and (ii) temporally in alignment. By means of experiments on four real datasets, we show that the representations learnt via the proposed DCTW significantly outperform state-of-the-art methods in temporal alignment, elegantly handling scenarios with highly heterogeneous features, such as the temporal alignment of acoustic and visual features. George Trigeorgis, Mihalis A. Nicolaou, Stefanos Zafeiriou, Björn W. Schuller |
CVPR | 2 |
| 2016 | Mnemonic Descent Method: A Recurrent Process Applied for End-to-End Face AlignmentabstractCascaded regression has recently become the method of choice for solving non-linear least squares problems such as deformable image alignment. Given a sizeable training set, cascaded regression learns a set of generic rules that are sequentially applied to minimise the least squares problem. Despite the success of cascaded regression for problems such as face alignment and head pose estimation, there are several shortcomings arising in the strategies proposed thus far. Specifically, (a) the regressors are learnt independently, (b) the descent directions may cancel one another out and (c) handcrafted features (e.g., HoGs, SIFT etc.) are mainly used to drive the cascade, which may be sub-optimal for the task at hand. In this paper, we propose a combined and jointly trained convolutional recurrent neural network architecture that allows the training of an end-to-end to system that attempts to alleviate the aforementioned drawbacks. The recurrent module facilitates the joint optimisation of the regressors by assuming the cascades form a nonlinear dynamical system, in effect fully utilising the information between all cascade levels by introducing a memory unit that shares information across all levels. The convolutional module allows the network to extract features that are specialised for the task at hand and are experimentally shown to outperform hand-crafted features. We show that the application of the proposed architecture for the problem of face alignment results in a strong improvement over the current state-of-the-art. George Trigeorgis, Patrick Snape, Mihalis A. Nicolaou, Epameinondas Antonakos, Stefanos Zafeiriou |
CVPR | 3 |
| 2016 | Adieu features? End-to-end speech emotion recognition using a deep convolutional recurrent networkabstractThe automatic recognition of spontaneous emotions from speech is a challenging task. On the one hand, acoustic features need to be robust enough to capture the emotional content for various styles of speaking, and while on the other, machine learning algorithms need to be insensitive to outliers while being able to model the context. Whereas the latter has been tackled by the use of Long Short-Term Memory (LSTM) networks, the former is still under very active investigations, even though more than a decade of research has provided a large set of acoustic descriptors. In this paper, we propose a solution to the problem of ‘context-aware’ emotional relevant feature extraction, by combining Convolutional Neural Networks (CNNs) with LSTM networks, in order to automatically learn the best representation of the speech signal directly from the raw time representation. In this novel work on the so-called end-to-end speech emotion recognition, we show that the use of the proposed topology significantly outperforms the traditional approaches based on signal processing techniques for the prediction of spontaneous and natural emotions on the RECOLA database. George Trigeorgis, Fabien Ringeval, Raymond Brueckner, Erik Marchi, Mihalis A. Nicolaou, Björn W. Schuller, Stefanos Zafeiriou |
ICASSP | 5 |
| 2016 | Robust Correlated and Individual Component AnalysisabstractRecovering correlated and individual components of two, possibly temporally misaligned, sets of data is a fundamental task in disciplines such as image, vision, and behavior computing, with application to problems such as multi-modal fusion (via correlated components), predictive analysis, and clustering (via the individual ones). Here, we study the extraction of correlated and individual components under real-world conditions, namely i) the presence of gross non-Gaussian noise and ii) temporally misaligned data. In this light, we propose a method for the Robust Correlated and Individual Component Analysis (RCICA) of two sets of data in the presence of gross, sparse errors. We furthermore extend RCICA in order to handle temporal incongruities arising in the data. To this end, two suitable optimization problems are solved. The generality of the proposed methods is demonstrated by applying them onto 4 applications, namely i) heterogeneous face recognition, ii) multi-modal feature fusion for human behavior analysis (i.e., audio-visual prediction of interest and conflict), iii) face clustering, and iv) thetemporal alignment of facial expressions. Experimental results on 2 synthetic and 7 real world datasets indicate the robustness and effectiveness of the proposed methodson these application domains, outperforming other state-of-the-art methods in the field. Yannis Panagakis, Mihalis A. Nicolaou, Stefanos Zafeiriou, Maja Pantic |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Probabilistic Slow Features for Behavior AnalysisabstractA recently introduced latent feature learning technique for time-varying dynamic phenomena analysis is the so-called slow feature analysis (SFA). SFA is a deterministic component analysis technique for multidimensional sequences that, by minimizing the variance of the first-order time derivative approximation of the latent variables, finds uncorrelated projections that extract slowly varying features ordered by their temporal consistency and constancy. In this paper, we propose a number of extensions in both the deterministic and the probabilistic SFA optimization frameworks. In particular, we derive a novel deterministic SFA algorithm that is able to identify linear projections that extract the common slowest varying features of two or more sequences. In addition, we propose an expectation maximization (EM) algorithm to perform inference in a probabilistic formulation of SFA and similarly extend it in order to handle two and more time-varying data sequences. Moreover, we demonstrate that the probabilistic SFA (EM-SFA) algorithm that discovers the common slowest varying latent space of multiple sequences can be combined with dynamic time warping techniques for robust sequence time-alignment. The proposed SFA algorithms were applied for facial behavior analysis, demonstrating their usefulness and appropriateness for this task. Lazaros Zafeiriou, Mihalis A. Nicolaou, Stefanos Zafeiriou, Symeon Nikitidis, Maja Pantic |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2014 | Robust Canonical Correlation Analysis: Audio-visual fusion for learning continuous interestabstractThe problem of automatically estimating the interest level of a subject has been gaining attention by researchers, mostly due to the vast applicability of interest detection. In this work, we obtain a set of continuous interest annotations for the SE-MAINE database, which we analyse also in terms of emotion dimensions such as valence and arousal. Most importantly, we propose a robust variant of Canonical Correlation Analysis (RCCA) for performing audio-visual fusion, which we apply to the prediction of interest. RCCA recovers a low-rank subspace which captures the correlations of fused modalities, while isolating gross errors in the data without making any assumptions regarding Gaussianity. We experimentally show that RCCA is more appropriate than other standard fusion techniques (such as l2-CCA and feature-level fusion), since it both captures interactions between modalities while also decontaminating the obtained subspace from errors which are dominant in real-world problems. Mihalis A. Nicolaou, Yannis Panagakis, Stefanos Zafeiriou, Maja Pantic |
ICASSP | 1 |
| 2014 | A Unified Framework for Probabilistic Component Analysis
Mihalis A. Nicolaou, Stefanos Zafeiriou, Maja Pantic |
ECML/PKDD (2) | 1 |
| 2014 | Dynamic Probabilistic CCA for Analysis of Affective Behavior and Fusion of Continuous AnnotationsabstractFusing multiple continuous expert annotations is a crucial problem in machine learning and computer vision, particularly when dealing with uncertain and subjective tasks related to affective behavior. Inspired by the concept of inferring shared and individual latent spaces in Probabilistic Canonical Correlation Analysis (PCCA), we propose a novel, generative model that discovers temporal dependencies on the shared/individual spaces (Dynamic Probabilistic CCA, DPCCA). In order to accommodate for temporal lags, which are prominent amongst continuous annotations, we further introduce a latent warping process, leading to the DPCCA with Time Warpings (DPCTW) model. Finally, we propose two supervised variants of DPCCA/DPCTW which incorporate inputs (i.e., visual or audio features), both in a generative (SG-DPCCA) and discriminative manner (SD-DPCCA). We show that the resulting family of models (i) can be used as a unifying framework for solving the problems of temporal alignment and fusion of multiple annotations in time, (ii) can automatically rank and filter annotations based on latent posteriors or other model statistics, and (iii) that by incorporating dynamics, modeling annotation-specific biases, noise estimation, time warping and supervision, DPCTW outperforms state-of-the-art methods for both the aggregation of multiple, yet imperfect expert annotations as well as the alignment of affective behavior. Mihalis A. Nicolaou, Vladimir Pavlovic 0001, Maja Pantic |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Robust Canonical Time Warping for the Alignment of Grossly Corrupted SequencesabstractTemporal alignment of human behaviour from visual data is a very challenging problem due to a numerous reasons, including possible large temporal scale differences, inter/intra subject variability and, more importantly, due to the presence of gross errors and outliers. Gross errors are often in abundance due to incorrect localization and tracking, presence of partial occlusion etc. Furthermore, such errors rarely follow a Gaussian distribution, which is the de-facto assumption in machine learning methods. In this paper, building on recent advances on rank minimization and compressive sensing, a novel, robust to gross errors temporal alignment method is proposed. While previous approaches combine the dynamic time warping (DTW) with low-dimensional projections that maximally correlate two sequences, we aim to learn two underlying projection matrices (one for each sequence), which not only maximally correlate the sequences but, at the same time, efficiently remove the possible corruptions in any datum in the sequences. The projections are obtained by minimizing the weighted sum of nuclear and ℓ1norms, by solving a sequence of convex optimization problems, while the temporal alignment is found by applying the DTW in an alternating fashion. The superiority of the proposed method against the state-of-the-art time alignment methods, namely the canonical time warping and the generalized time warping, is indicated by the experimental results on both synthetic and real datasets. Yannis Panagakis, Mihalis A. Nicolaou, Stefanos Zafeiriou, Maja Pantic |
CVPR | 2 |
| 2013 | Learning Slow Features for Behaviour AnalysisabstractA recently introduced latent feature learning technique for time varying dynamic phenomena analysis is the so called Slow Feature Analysis (SFA). SFA is a deterministic component analysis technique for multi-dimensional sequences that by minimizing the variance of the first order time derivative approximation of the input signal finds uncorrelated projections that extract slowly-varying features ordered by their temporal consistency and constancy. In this paper, we propose a number of extensions in both the deterministic and the probabilistic SFA optimization frameworks. In particular, we derive a novel deterministic SFA algorithm that is able to identify linear projections that extract the common slowest varying features of two or more sequences. In addition, we propose an Expectation Maximization (EM) algorithm to perform inference in a probabilistic formulation of SFA and similarly extend it in order to handle two and more time varying data sequences. Moreover, we demonstrate that the probabilistic SFA (EMSFA) algorithm that discovers the common slowest varying latent space of multiple sequences can be combined with dynamic time warping techniques for robust sequence time alignment. The proposed SFA algorithms were applied for facial behavior analysis demonstrating their usefulness and appropriateness for this task. Lazaros Zafeiriou, Mihalis A. Nicolaou, Stefanos Zafeiriou, Symeon Nikitidis, Maja Pantic |
ICCV | 2 |
| 2013 | Correlated-spaces regression for learning continuous emotion dimensionsabstractAdopting continuous dimensional annotations for affective analysis has been gaining rising attention by researchers over the past years. Due to the idiosyncratic nature of this problem, many subproblems have been identified, spanning from the fusion of multiple continuous annotations to exploiting output-correlations amongst emotion dimensions. In this paper, we firstly empirically answer several important questions which have found partial or no answer at all so far in related literature. In more detail, we study the correlation of each emotion dimension (i) with respect to other emotion dimensions, (ii) to basic emotions (e.g., happiness, anger). As a measure for comparison, we use video and audio features. Interestingly enough, we find that (i) each emotion dimension is more correlated with other emotion dimensions rather than with face and audio features, and similarly (ii) that each basic emotion is more correlated with emotion dimensions than with audio and video features. A similar conclusion holds for discrete emotions which are found to be highly correlated to emotion dimensions as compared to audio and/or video features. Motivated by these findings, we present a novel regression algorithm (Correlated-Spaces Regression, CSR), inspired by Canonical Correlation Analysis (CCA) which learns output-correlations and performs supervised dimensionality reduction and multimodal fusion by (i) projecting features extracted from all modalities and labels onto a common space where their inter-correlation is maximised and (ii) learning mappings from the projected feature space onto the projected, uncorrelated label space. Mihalis A. Nicolaou, Stefanos Zafeiriou, Maja Pantic |
ACM Multimedia | 1 |
| 2012 | Dynamic Probabilistic CCA for Analysis of Affective Behaviour
Mihalis A. Nicolaou, Vladimir Pavlovic 0001, Maja Pantic |
ECCV (7) | 1 |
| 2012 | Output-associative RVM regression for dimensional and continuous emotion prediction
Mihalis A. Nicolaou, Hatice Gunes, Maja Pantic |
Image Vis. Comput. | 1 |
| 2011 | Output-associative RVM regression for dimensional and continuous emotion predictionabstractMany problems in machine learning and computer vision consist of predicting multi-dimensional output vectors given a specific set of input features. In many of these problems, there exist inherent temporal and spacial dependencies between the output vectors, as well as repeating output patterns and input-output associations, that can provide more robust and accurate predictors when modelled properly. With this intrinsic motivation, we propose a novel Output-Associative Relevance Vector Machine (OA-RVM) regression framework that augments the traditional RVM regression by being able to learn non-linear input and output dependencies. Instead of depending solely on the input patterns, OA-RVM models output structure and covariances within a predefined temporal window, thus capturing past, current and future context. As a result, output patterns manifested in the training data are captured within a formal probabilistic framework, and subsequently used during inference. As a proof of concept, we target the highly challenging problem of dimensional and continuous prediction of emotions from naturalistic facial expressions. We demonstrate the advantages of the proposed OA-RVM regression by performing both subject-dependent and subject-independent experiments using the SAL database. The experimental results show that OA-RVM regression outperforms the traditional RVM and SVM regression approaches in prediction accuracy, generating more robust and accurate models. Mihalis A. Nicolaou, Hatice Gunes, Maja Pantic |
FG | 1 |
| 2011 | A multi-layer hybrid framework for dimensional emotion classificationabstractThis paper investigates dimensional emotion prediction and classification from naturalistic facial expressions. Similarly to many pattern recognition problems, dimensional emotion classification requires generating multi-dimensional outputs. To date, classification for valence and arousal dimensions has been done separately, assuming that they are independent. However, various psychological findings suggest that these dimensions are correlated. We therefore propose a novel, multi-layer hybrid framework for emotion classification that is able to model inter-dimensional correlations. Firstly, we derive a novel geometric feature set based on the (a)symmetric spatio-temporal characteristics of facial expressions. Subsequently, we use the proposed feature set to train a multi-layer hybrid framework composed of a tem- poral regression layer for predicting emotion dimensions, a graphical model layer for modeling valence-arousal correlations, and a final classification and fusion layer exploiting informative statistics extracted from the lower layers. This framework (i) introduces the Auto-Regressive Coupled HMM (ACHMM), a graphical model specifically tailored to accommodate not only inter-dimensional correlations but also to exploit the internal dynamics of the actual observations, and (ii) replaces the commonly used Maximum Likelihood principle with a more robust final classification and fusion layer. Subject-independent experimental validation, performed on a naturalistic set of facial expressions, demonstrates the effectiveness of the derived feature set, and the robustness and flexibility of the proposed framework. Mihalis A. Nicolaou, Hatice Gunes, Maja Pantic |
ACM Multimedia | 1 |
| 2011 | Continuous Prediction of Spontaneous Affect from Multiple Cues and Modalities in Valence-Arousal SpaceabstractPast research in analysis of human affect has focused on recognition of prototypic expressions of six basic emotions based on posed data acquired in laboratory settings. Recently, there has been a shift toward subtle, continuous, and context-specific interpretations of affective displays recorded in naturalistic and real-world settings, and toward multimodal analysis and recognition of human affect. Converging with this shift, this paper presents, to the best of our knowledge, the first approach in the literature that: 1) fuses facial expression, shoulder gesture, and audio cues for dimensional and continuous prediction of emotions in valence and arousal space, 2) compares the performance of two state-of-the-art machine learning techniques applied to the target problem, the bidirectional Long Short-Term Memory neural networks (BLSTM-NNs), and Support Vector Machines for Regression (SVR), and 3) proposes an output-associative fusion framework that incorporates correlations and covariances between the emotion dimensions. Evaluation of the proposed approach has been done using the spontaneous SAL data from four subjects and subject-dependent leave-one-sequence-out cross validation. The experimental results obtained show that: 1) on average, BLSTM-NNs outperform SVR due to their ability to learn past and future context, 2) the proposed output-associative fusion framework outperforms feature-level and model-level fusion by modeling and learning correlations and patterns between the valence and arousal dimensions, and 3) the proposed system is well able to reproduce the valence and arousal ground truth obtained from human coders. Mihalis A. Nicolaou, Hatice Gunes, Maja Pantic |
IEEE Trans. Affect. Comput. | 1 |
| 2010 | Audio-Visual Classification and Fusion of Spontaneous Affective Data in Likelihood SpaceabstractThis paper focuses on audio-visual (using facial expression, shoulder and audio cues) classification of spontaneous affect, utilising generative models for classification (i) in terms of Maximum Likelihood Classification with the assumption that the generative model structure in the classifier is correct, and (ii) Likelihood Space Classification with the assumption that the generative model structure in the classifier may be incorrect, and therefore, the classification performance can be improved by projecting the results of generative classifiers onto likelihood space, and then using discriminative classifiers. Experiments are conducted by utilising Hidden Markov Models for single cue classification, and 2 and 3-chain coupled Hidden Markov Models for fusing multiple cues and modalities. For discriminative classification, we utilise Support Vector Machines. Results show that Likelihood Space Classification improves the performance (91.76%) of Maximum Likelihood Classification (79.1%). Thereafter, we introduce the concept of fusion in the likelihood space, which is shown to outperform the typically used model-level fusion, attaining a classification accuracy of 94.01% and further improving all previous results. Mihalis A. Nicolaou, Hatice Gunes, Maja Pantic |
ICPR | 1 |