Ioannis Patras

dblp:18/1556 · also Ioannis (Yiannis) Patras · DBLP profile ↗
← Back
175ranked-venue papers
12as first author
63since 2021 · last 2026
0000-0003-3913-4738ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 124 · 9 first-author · 39 since 2021Artificial intelligence and machine learning · 90 · 4 first-author · 40 since 2021Databases, data management, data science and information retrieval · 8 · 4 since 2021Human-computer interaction and ubiquitous computing · 8 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Computer networks · 2 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Confidence Should Be Calibrated More Than One Turn Deep
abstract
Zhaohan Zhang, Chengzhengxu Li, Xiaoming Liu, Chao Shen, Ziquan Liu, Ioannis Patras. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhaohan Zhang, Chengzhengxu Li, Chao Shen 0001, Ziquan Liu, Ioannis Patras
ACL (1)6
2026 GrACE: A Generative Approach to Better Confidence Elicitation and Efficient Test-Time Scaling in Large Language Models
abstract
Assessing the reliability of Large Language Models (LLMs) by confidence elicitation is a prominent approach to AI safety in highstakes applications, such as healthcare and finance.Existing methods either require expensive computational overhead or suffer from poor calibration, making them impractical and unreliable for real-world deployment.In this work, we propose GrACE, a Generative Approach to Confidence Elicitation that enables scalable and reliable confidence elicitation for LLMs.GrACE adopts a novel mechanism in which the model expresses confidence by the similarity between the last hidden state and the embedding of a special token appended to the vocabulary, in real-time.We fine-tune the model for calibrating the confidence with targets associated with accuracy.Extensive experiments show that the confidence produced by GrACE achieves the best discriminative capacity and calibration on open-ended generation tasks without resorting to additional sampling or an auxiliary model.Moreover, we propose two confidence-based strategies for test-time scaling with GrACE, which not only improve the accuracy of the final decision but also significantly reduce the number of required samples, highlighting its potential as a practical solution for deploying LLMs with reliable, on-the-fly confidence estimation.The code is available at: https://github.com/petezone/Grace.
Zhaohan Zhang, Ziquan Liu, Ioannis Patras
ACL (1)3
2026 HDD-Unet: A Unet-based architecture for low-light image enhancement
abstract
Low-light imaging has become a popular topic in image processing, with the quality enhancement of low light images being as a significant challenge, due to the difficulty in retaining colors, patterns, texture and style when generating a normal light image. Our objectives are mainly to firstly better preserve texture regions in image enhancement, while, secondly, preserving colors via color histogram blocks and, finally, to enhance the quality of image through dense denoising blocks. Our proposed novel framework, namely HDD-Unet, is a double Unet based on photorealistic style transfer for low-light image enhancement. The proposed low-light image enhancement method combines color histogram-based fusion, Haar wavelet pooling, dense-denoising blocks and U-net as a backbone architecture to enhance the contrast, reduce noise, and improve the visibility of low light images. Experimental results demonstrate that our proposed method outperforms existing methods in terms of PSNR and SSIM quantitative evaluation metrics, reaching or outperforming state-of-the-art accuracy, but with less resources. We also conduct an ablation study to investigate the impact of our approach on overexposed images, and systematic analysis on the late fusion weighting parameters. Multiple experiments were conducted with artificial noise inserted to accomplish more efficient comparison. The results show that the proposed framework enhances accurately images with various gamma corrections. The proposed method represents a significant advance in the field of low light image enhancement and has the potential to address several challenges associated with low light imaging. • Low-light image enhancement based on a combination of two U-Net branches with wavelet transformations, Adaptive Instance Normalization, color histogram blocks and denoising dense blocks. • Analytic framework creation for the HDD-Unet architecture. • Evaluation of the proposed method in a qualitative and a quantitative way.
Elissavet Batziou, Konstantinos Ioannidis, Ioannis Patras, Stefanos Vrochidis, Ioannis Kompatsiaris
Image Vis. Comput.3
2026 Unsupervised Object Localization driven by self-supervised foundation models: A comprehensive review
abstract
Object localization is a fundamental task in computer vision that traditionally requires labeled datasets for accurate results. Recent progress in self-supervised learning has enabled unsupervised object localization, reducing reliance on manual annotations. Unlike supervised encoders, which depend on annotated training data, self-supervised encoders learn semantic representations directly from large collections of unlabeled images. This makes them the natural foundation for unsupervised object localization, as they capture object-relevant features while eliminating the need for costly manual labels. These encoders produce semantically coherent patch embeddings. Grouping these embeddings reveals sets of patches that correspond to objects in an image. These patch sets can be converted into object masks or bounding boxes, enabling tasks such as single-object discovery, multi-object detection, and instance segmentation. By applying off-line mask clustering or using pre-trained vision-language models, unsupervised localization methods can assign semantic labels to discovered objects. This transforms initially class-agnostic objects (objects without class labels) into class-aware ones (objects with class labels), aligning these tasks with their supervised counterparts. This paper provides a structured review of unsupervised object localization methods in both class-agnostic and class-aware settings. In contrast, previous surveys have focused only on class-agnostic localization. We discuss state-of-the-art object discovery strategies based on self-supervised features and provide a detailed comparison of experimental results across a wide range of tasks, datasets, and evaluation metrics.
Sotirios Papadopoulos, Emmanouil Patsiouras, Konstantinos Ioannidis, Stefanos Vrochidis, Ioannis Kompatsiaris, Ioannis Patras
Image Vis. Comput.6
2025 Get Confused Cautiously: Textual Sequence Memorization Erasure with Selective Entropy Maximization
abstract
Large Language Models (LLMs) have been found to memorize and recite some of the textual sequences from their training set verbatim, raising broad concerns about privacy and copyright issues. This Textual Sequence Memorization (TSM) phenomenon leads to a high demand to regulate LLM output to prevent generating certain memorized text that a user wants to be forgotten. However, our empirical study reveals that existing methods for TSM erasure fail to unlearn large numbers of memorized samples without substantially jeopardizing the model utility. To achieve a better trade-off between the effectiveness of TSM erasure and model utility in LLMs, our paper proposes a new method, named Entropy Maximization with Selective Optimization (EMSO), where the model parameters are updated sparsely based on novel optimization and selection criteria, in a manner that does not require additional models or data other than that in the forget set. More specifically, we propose an entropy-based loss that is shown to lead to more stable optimization and better preserves model utility than existing methods. In addition, we propose a contrastive gradient metric that takes both the gradient magnitude and direction into consideration, so as to localize model parameters to update in a sparse model updating scehme. Extensive experiments across three model scales demonstrate that our method excels in handling large-scale forgetting requests while preserving model ability in language generation and understanding.
Zhaohan Zhang, Ziquan Liu, Ioannis Patras
COLING3
2025 Temporal Score Analysis for Understanding and Correcting Diffusion Artifacts
abstract
Visual artifacts remain a persistent challenge in diffusion models, even with training on massive datasets. Current solutions primarily rely on supervised detectors, yet lack understanding of why these artifacts occur in the first place. In our analysis, we identify three distinct phases in the diffusion generative process: Profiling, Mutation, and Refinement. Artifacts typically emerge during the Mutation phase, where certain regions exhibit anomalous score dynamics over time, causing abrupt disruptions in the normal evolution pattern. This temporal nature explains why existing methods focusing only on spatial uncertainty of the final output fail at effective artifact localization. Based on these insights, we propose ASCED (Abnormal Score Correction for Enhancing Diffusion), that detects artifacts by monitoring abnormal score dynamics during the diffusion process, with a trajectory-aware on-the-fly mitigation strategy that appropriate generation of noise in the detected areas. Unlike most existing methods that apply post hoc corrections, e.g., by applying a noising-denoising scheme after generation, our mitigation strategy operates seamlessly within the existing diffusion process. Extensive experiments demonstrate that our proposed approach effectively reduces artifacts across diverse domains, matching or surpassing existing supervised methods without additional training. Project page: YuCao16.github.io/ASCED.
Zengqun Zhao, Ioannis Patras, Shaogang Gong
CVPR3
2025 ReWind: Understanding Long Videos with Instructed Learnable Memory
abstract
Vision-Language Models (VLMs) are crucial for applications requiring integrated understanding textual and visual information. However, existing VLMs struggle with long videos due to computational inefficiency, memory limitations, and difficulties in maintaining coherent understanding across extended sequences. To address these challenges, we introduce ReWind, a novel memory-based VLM designed for efficient long video understanding while preserving temporal fidelity. ReWind operates in a two-stage framework. In the first stage, ReWind maintains a dynamic learnable memory module with a novel read-perceive-write cycle that stores and updates instruction-relevant visual information as the video unfolds. This module utilizes learnable queries and cross-attentions between memory contents and the input stream, ensuring low memory requirements by scaling linearly with the number of tokens. In the second stage, we propose an adaptive frame selection mechanism guided by the memory content to identify instruction-relevant key moments. It enriches the memory representations with detailed spatial information by selecting a few high-resolution frames, which are then combined with the memory contents and fed into a Large Language Model (LLM) to generate the final answer. We empirically demonstrate ReWind’s superior performance in visual question answering (VQA) and temporal grounding tasks, surpassing previous methods on long video benchmarks. Notably, ReWind achieves a +13% score gain and a +12% accuracy improvement on the MovieChat-1K VQA dataset and an +8% mIoU increase on Charades-STA for temporal grounding.
Anxhelo Diko, Tinghuai Wang, Wassim Swaileh, Shiyan Sun, Ioannis Patras
CVPR5
2025 AIM-Fair: Advancing Algorithmic Fairness via Selectively Fine-Tuning Biased Models with Contextual Synthetic Data
abstract
Recent advances in generative models have sparked research on improving model fairness with AI-generated data. However, existing methods often face limitations in the diversity and quality of synthetic data, leading to compromised fairness and overall model accuracy. Moreover, many approaches rely on the availability of demographic group labels, which are often costly to annotate. This paper proposes AIM-Fair, aiming to overcome these limitations and harness the potential of cutting-edge generative models in promoting algorithmic fairness. We investigate a fine-tuning paradigm starting from a biased model initially trained on real-world data without demographic annotations. This model is then fine-tuned using unbiased synthetic data generated by a state-of-the-art diffusion model to improve its fairness. Two key challenges are identified in this fine-tuning paradigm, 1) the low quality of synthetic data, which can still happen even with advanced generative models, and 2) the domain and bias gap between real and synthetic data. To address the limitation of synthetic data quality, we propose Contextual Synthetic Data Generation (CSDG) to generate data using a text-to-image diffusion model (T2I) with prompts generated by a context-aware LLM, ensuring both data diversity and control of bias in synthetic data. To resolve domain and bias shifts, we introduce a novel selective fine-tuning scheme in which only model parameters more sensitive to bias and less sensitive to domain shift are updated. Experiments on CelebA and UTKFace datasets show that our AIM-Fair improves model fairness while maintaining utility, outperforming both fully and partially fine-tuned approaches to model fairness. The code is available at https://github.com/zengqunzhao/AIM-Fair.
Zengqun Zhao, Ziquan Liu, Shaogang Gong, Ioannis Patras
CVPR5
2025 DiffusionAct: Controllable Diffusion Autoencoder for One-shot Face Reenactment
abstract
Video-driven neural face reenactment aims to synthesize realistic facial images that successfully preserve the identity and appearance of a source face, while transferring the target head pose and facial expressions. Existing GAN-based methods suffer from either distortions and visual artifacts or poor reconstruction quality, i.e., the background and several important appearance details, such as hair style/color, glasses and accessories, are not faithfully reconstructed. Recent advances in Diffusion Probabilistic Models (DPMs) enable the generation of high-quality realistic images. To this end, in this paper we present DiffusionAct, a novel method that leverages the photo-realistic image generation of diffusion models to perform neural face reenactment. Specifically, we propose to control the semantic space of a Diffusion Autoencoder (DiffAE), in order to edit the facial pose of the input images, defined as the head pose orientation and the facial expressions. Our method allows one-shot, self, and cross-subject reenactment, without requiring subject-specific fine-tuning. We compare against state-of-the-art GAN-, StyleGAN2-, and diffusion-based methods, showing better or on-par reenactment performance. Project page: https://stelabou.github.io/diffusionact/
Stella Bounareli, Christos Tzelepis, Vasileios Argyriou, Ioannis Patras, Georgios Tzimiropoulos
FG4
2025 Frequency-Guided Diffusion for Training-Free Text-Driven Image Translation
Zheng Gao 0003, Jifei Song, Zhensong Zhang, Jiankang Deng, Ioannis Patras
ICCV5
2025 VidCtx: Context-aware Video Question Answering with Image Models
abstract
To address computational and memory limitations of Large Multimodal Models in the Video Question-Answering task, several recent methods extract textual representations per frame (e.g., by captioning) and feed them to a Large Language Model (LLM) that processes them to produce the final response. However, in this way, the LLM does not have access to visual information and often has to process repetitive textual descriptions of nearby frames. To address those shortcomings, in this paper, we introduce VidCtx, a novel training-free VideoQA framework which integrates both modalities, i.e. both visual information from input frames and textual descriptions of others frames that give the appropriate context. More specifically, in the proposed framework a pre-trained Large Multimodal Model (LMM) is prompted to extract at regular intervals, question-aware textual descriptions (captions) of video frames. Those will be used as context when the same LMM will be prompted to answer the question at hand given as input a) a certain frame, b) the question and c) the context/caption of an appropriate frame. To avoid redundant information, we chose as context the descriptions of distant frames. Finally, a simple yet effective max pooling mechanism is used to aggregate the frame-level decisions. This methodology enables the model to focus on the relevant segments of the video and scale to a high number of frames. Experiments show that VidCtx achieves competitive performance among approaches that rely on open models on three public Video QA benchmarks, NExT-QA, IntentQA and STAR. Our code is available at https://github.com/IDT-ITI/VidCtx.
Andreas Goulas, Vasileios Mezaris, Ioannis Patras
ICME3
2025 VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning
abstract
Recognising emotions in context involves identifying an individual’s apparent emotions while considering contextual cues from the surrounding scene. Previous approaches to this task have typically designed explicit scene-encoding architectures or incorporated external scene-related information, such as captions. However, these methods often utilise limited contextual information or rely on intricate training pipelines to decouple noise from relevant information. In this work, we leverage the capabilities of Vision-and-Large-Language Models (VLLMs) to enhance in-context emotion classification in a more straightforward manner. Our proposed method follows a simple yet effective two-stage approach. First, we prompt VLLMs to generate natural language descriptions of the subject’s apparent emotion in relation to the visual context. Second, the descriptions, along with the visual input, are used to train a transformer-based architecture that fuses text and visual features before the final classification task. This method not only simplifies the training process but also significantly improves performance. Experimental results demonstrate that the textual descriptions effectively guide the model to constrain the noisy visual input, allowing our fused architecture to outperform individual modalities. Our approach achieves state-of-the-art performance across three datasets, BoLD, EMOTIC, and CAER-S, without bells and whistles. Our code will be made publicly available.
Alexandros Xenos, Niki Maria Foteinopoulou, Ioanna Ntinou, Ioannis Patras, Georgios Tzimiropoulos
IJCNN4
2025 Gen4Track: A Tuning-free Data Augmentation Framework via Self-correcting Diffusion Model for Vision-Language Tracking
abstract
The performance of current Vision-Language Tracking (VLT) models is constrained by the limited diversity and quantity of labeled data. Compared to constructing large-scale datasets, data augmentation offers a more cost-saving strategy for VLT by synthesizing new samples from existing data, rather than generating them from scratch. However, conventional techniques like rotation and flipping may disrupt scene composition, causing conflicts between visual layouts and textual annotations. Recent advances in generative models have inspired the use of synthetic videos for data augmentation. Yet, existing approaches fail to address the core concerns of data augmentation in VLT (shown in Fig. 1)-target location accuracy, text-video consistency, and video content coherency. To bridge the gap, we propose Gen4Track, a tuning-free data augmentation framework that leverages the self-correcting mechanism to dynamically generate high-quality video data with annotations. Our approach involves (1) optimizing the attention calculations in a frozen text-to-image diffusion model to synthesize coherent videos that satisfy specific conditions (e.g., spatial location, category, color, and style), and (2) implementing a self-correcting mechanism based on a Large Language Model (LLM) to improve text-video consistency. During video augmentation, we propose content-coherent self-attention and location-enhanced cross-attention mechanisms, ensuring that image-level editings are accurately and coherently propagated throughout the video. Then, with the goal of maximizing text-video consistency, we iteratively refine the augmentation instruction with our designed self-correcting mechanism for a more aligned video. Extensive experiments validate that Gen4Track significantly boosts the performance of SOTA VLT models (achieving improvements of up to 3.2% in SUC and 3.5% in PRE), opening a new chapter of training Vision-Language trackers with synthetic videos rather than manually annotated data.
Jiawei Ge 0002, Xin-Yu Zhang 0027, Jiuxin Cao, Xuelin Zhu, Qingqing Gao, Biwei Cao, Kun Wang 0057, Chang Liu 0113, Bo Liu 0004, Chen Feng 0028, Ioannis Patras
ACM Multimedia12
2025 Unveiling Open-set Noise: Theoretical Insights into Label Noise
abstract
Learning with Noisy Labels (LNL) reduces reliance on high-quality labeled data but often overlooks open-set noise, where noisy samples belong to unknown classes, unlike closed-set noise within known categories.This paper advances LNL by reformulating the problem to incorporate open-set noise through a complete noise transition matrix, enabling a theoretical comparison of its impact on classification error rates against closed-set noise. Our analysis reveals that open-set noise induces smaller error increases, with distinct effects from 'hard' (semantically similar to inliers) and 'easy' (dissimilar) variants. We evaluate entropy-based detection, finding it effective only for easy open-set noise, and propose solutions leveraging vision-language models and self-supervised learning to address hard noise challenges. For empirical validation, we introduce CIFAR100-O, ImageNet-O, and a WebVision open-set test set, enabling robust benchmarking of LNL methods under open-set noise conditions. Recognizing classification accuracy's limitations in capturing model robustness, we advocate out-of-distribution (OOD) detection as a complementary metric. Our theoretical and empirical results highlight the unique challenges of open-set noise, offering new tools and evaluation frameworks to enhance LNL robustness in real-world scenarios.
Chen Feng 0028, Nicu Sebe, Georgios Tzimiropoulos, Miguel R. D. Rodrigues, Ioannis Patras
ACM Multimedia5
2025 Vision-Language Pretraining for Variable-Shot Image Classification
Sotirios Papadopoulos, Konstantinos Ioannidis, Stefanos Vrochidis, Ioannis Kompatsiaris, Ioannis Patras
MMM (4)5
2025 Towards Interpretability Without Sacrifice: Faithful Dense Layer Decomposition with Mixture of Decoders
abstract
Multilayer perceptrons (MLPs) are an integral part of large language models, yet their dense representations render them difficult to understand, edit, and steer. Recent methods learn interpretable approximations via neuron-level sparsity, yet fail to faithfully reconstruct the original mapping--significantly increasing model's next-token cross-entropy loss. In this paper, we advocate for moving to layer-level sparsity to overcome the accuracy trade-off in sparse layer approximation. Under this paradigm, we introduce Mixture of Decoders (MxDs). MxDs generalize MLPs and Gated Linear Units, expanding pre-trained dense layers into tens of thousands of specialized sublayers. Through a flexible form of tensor factorization, each sparsely activating MxD sublayer implements a linear transformation with full-rank weights--preserving the original decoders' expressive capacity even under heavy sparsity. Experimentally, we show that MxDs significantly outperform state-of-the-art methods (e.g., Transcoders) on the sparsity-accuracy frontier in language models with up to 3B parameters. Further evaluations on sparse probing and feature steering demonstrate that MxDs learn similarly specialized features of natural language--opening up a promising new avenue for designing interpretable yet faithful decompositions. Our code is included at: https://github.com/james-oldfield/MxD.
James Oldfield 0001, Shawn Im, Yixuan Li 0001, Mihalis A. Nicolaou, Ioannis Patras, Grigorios Chrysos 0002
NeurIPS5
2025 λ-Orthogonality Regularization for Compatible Representation Learning
Simone Ricci, Niccolò Biondi, Federico Pernici, Ioannis Patras, Alberto Del Bimbo
NeurIPS4
2025 Enhancing Zero-Shot Facial Expression Recognition by LLM Knowledge Transfer
abstract
Current facial expression recognition (FER) models are often designed in a supervised learning manner and thus are constrained by the lack of large-scale facial expression images with high-quality annotations. Consequently, these models often fail to generalize well, performing poorly on unseen images in inference. Vision-language-based zero-shot models demonstrate a promising potential for addressing such challenges. However, these models lack task-specific knowledge and therefore are not optimized for the nuances of recognizing facial expressions. To bridge this gap, this work proposes a novel method, Exp-CLIP, to enhance zero-shot FER by transferring the task knowledge from large language models (LLMs). Specifically, based on the pre-trained vision-language encoders, we incorporate a projection head designed to map the initial joint vision-language space into a space that captures representations of facial actions. To train this projection head for subsequent zero-shot predictions, we propose to align the projected visual representations with task-specific semantic meanings derived from the LLM encoder, and the text instruction-based strategy is employed to customize the LLM knowledge. Given unlabelled facial data and efficient training of the projection head, Exp-CLIP achieves superior zero-shot results to the CLIP models and several other large vision-language models (LVLMs) on seven in-the-wild FER datasets. The code is available at https://github.com/zengqunzhao/Exp-CLIP.
Zengqun Zhao, Shaogang Gong, Ioannis Patras
WACV4
2025 Machine learning approaches for fine-grained symptom estimation in schizophrenia: A comprehensive review
abstract
Schizophrenia is a severe yet treatable mental disorder, and it is diagnosed using a multitude of primary and secondary symptoms. Diagnosis and treatment for each individual depends on the severity of the symptoms. Therefore, there is a need for accurate, personalised assessments. However, the process can be both time-consuming and subjective; hence, there is a motivation to explore automated methods that can offer consistent diagnosis and precise symptom assessments, thereby complementing the work of healthcare practitioners. Machine Learning has demonstrated impressive capabilities across numerous domains, including medicine; the use of Machine Learning in patient assessment holds great promise for healthcare professionals and patients alike, as it can lead to more consistent and accurate symptom estimation. This survey reviews methodologies utilising Machine Learning for diagnosing and assessing schizophrenia. Contrary to previous reviews that primarily focused on binary classification, this work recognises the complexity of the condition and, instead, offers an overview of Machine Learning methods designed for fine-grained symptom estimation. We cover multiple modalities, namely Medical Imaging, Electroencephalograms and Audio-Visual, as the illness symptoms can manifest in a patient's pathology and behaviour. Finally, we analyse the datasets and methodologies used in the studies and identify trends, gaps, as opportunities for future research.
Niki Maria Foteinopoulou, Ioannis Patras
Artif. Intell. Medicine2
2025 Introduction to the Special Issue on Realistic Synthetic Data: Generation, Learning, Evaluation
abstract
This is the foreword of our special issue volume on Realistic Synthetic Data: Generation, Learning, Evaluation organized with the ACM Transactions on Multimedia Computing, Communications, and Applications. It presents the target of the special issue that relates to synthetic data for various modalities, e.g., signals, images, volumes, audio, etc., controllable generation for learning from synthetic data, transfer learning and generalization of models, causality in data generation, addressing bias, limitations and trustworthiness in data generation, evaluation measures/protocols and benchmarks to assess quality of synthetic content, open synthetic datasets and software tools, and ethical aspects of synthetic data. The call for papers received a record number of 40 submissions out of which 15 were finally accepted for publication. This introduction provides an overview of the topics of each of the articles.
Bogdan Ionescu, Ioannis Patras, Henning Müller, Alberto Del Bimbo
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Self-Supervised Facial Representation Learning with Facial Region Awareness
abstract
Self-supervised pretraining has been proved to be effective in learning transferable representations that benefit various visual tasks. This paper asks this question: can self-supervised pretraining learn general facial representations for various facial analysis tasks? Recent efforts to-ward this goal are limited to treating each face image as a whole, i.e., learning consistent facial representations at the image-level, which overlooks the “consistency of local facial representations“(i.e., facial regions like eyes, nose, etc). In this work, we make a first attempt to propose a novel self-supervised facial representation learning framework to learn consistent global and local facial representations, Facial Region Awareness (FRA). Specifically, we explicitly enforce the consistency of facial regions by matching the local facial representations across views, which are extracted with learned heatmaps highlighting the facial regions. Inspired by the mask prediction in supervised semantic segmentation, we obtain the heatmaps via cosine simi-larity between the per-pixel projection of feature maps and “facial mask embeddings” computed from learnable positional embeddings, which leverage the attention mechanism to globally look up the facial image for facial regions. To learn such heatmaps, we formulate the learning of facial mask embeddings as a deep clustering problem by assigning the pixel features from the feature maps to them. The transfer learning results on facial classification and regression tasks show that our FRA outperforms previous pretrained models and more importantly, using ResNet as the unified backbone for various tasks, our FRA achieves comparable or even better performance compared with SOTA methods in facial analysis tasks.
Zheng Gao 0003, Ioannis Patras
CVPR2
2024 LAFS: Landmark-Based Facial Self-Supervised Learning for Face Recognition
abstract
In this work we focus on learning facial representations that can be adapted to train effective face recognition models, particularly in the absence of labels. Firstly, compared with existing labelled face datasets, a vastly larger magnitude of unlabeled faces exists in the real world. We explore the learning strategy of these unlabeled facial images through self-supervised pretraining to transfer gener-alized face recognition performance. Moreover, motivated by one recent finding, that is, the face saliency area is critical for face recognition, in contrast to utilizing random cropped blocks of images for constructing augmentations in pretraining, we utilize patches localized by extracted facial landmarks. This enables our method - namely LAndmark-based Facial Self-supervised learning (LAFS), to learn key representation that is more critical for face recognition. We also incorporate two landmark-specific augmen-tations which introduce more diversity of landmark information to further regularize the learning. With learned landmark-based facial representations, we further adapt the representation for face recognition with regularization mitigating variations in landmark positions. Our method achieves significant improvement over the state-of-the-art on multiple face recognition benchmarks, especially on more challenging few-shot scenarios. The code is available at https://github.com/szlbiubiubiulLAFS_CVPR2024.
Zhonglin Sun, Chen Feng 0028, Ioannis Patras, Georgios Tzimiropoulos
CVPR3
2024 Efficient Unsupervised Visual Representation Learning with Explicit Cluster Balancing
Ioannis Maniadis Metaxas, Georgios Tzimiropoulos, Ioannis Patras
ECCV (32)3
2024 EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Expression Recognition
abstract
Facial Expression Recognition (FER) is a crucial task in affective computing, but its conventional focus on the seven basic emotions limits its applicability to the complex and expanding emotional spectrum. To address the issue of new and unseen emotions present in dynamic in-the-wild FER, we propose a novel vision-language model that utilises sample-level text descriptions (i.e. captions of the context, expressions or emotional cues) as natural language supervision, aiming to enhance the learning of rich latent representations, for zero-shot classification. To test this, we evaluate using zero-shot classification of the model trained on sample-level descriptions on four popular dynamic FER datasets. Our findings show that this approach yields significant improvements when compared to baseline methods. Specifically, for zero-shot video FER, we outperform CLIP by over 10% in terms of Weighted Average Recall and 5% in terms of Unweighted Average Recall on several datasets. Furthermore, we evaluate the representations obtained from the network trained using sample-level descriptions on the downstream task of mental health symptom estimation, achieving performance comparable or superior to state-of-the-art methods and strong agreement with human experts. Namely, we achieved a Pearson's correlation coefficient of up to 0.85 for schizophrenia symptom severity estimation, which is comparable to human experts' agreement. The code is publicly available on https://github.com/NickyFot/EmoCLIP.git.
Niki Maria Foteinopoulou, Ioannis Patras
FG2
2024 CLIPCleaner: Cleaning Noisy Labels with CLIP
abstract
Learning with Noisy labels (LNL) poses a significant challenge for the Machine Learning community. Some of the most widely used approaches that select as clean samples for which the model itself (the in-training model) has high confidence, e.g., 'small loss', can suffer from the so called 'self-confirmation' bias. This bias arises because the in-training model, is at least partially trained on the noisy labels. Furthermore, in the classification case, an additional challenge arises because some of the label noise is between classes that are visually very similar ('hard noise'). This paper addresses these challenges by proposing a method (CLIPCleaner) that leverages CLIP, a powerful Vision-Language (VL) model for constructing a zero-shot classifier for efficient, offline, clean sample selection. This has the advantage that the sample selection is decoupled from the in-training model and that the sample selection is aware of the semantic and visual similarities between the classes due to the way that CLIP is trained. We provide theoretical justifications and empirical evidence to demonstrate the advantages of CLIP for LNL compared to conventional pre-trained models. Compared to current methods that combine iterative sample selection with various techniques, CLIPCleaner offers a simple, single-step approach that achieves competitive or superior performance on benchmark datasets. To the best of our knowledge, this is the first time a VL model has been used for sample selection to address the problem of Learning with Noisy Labels (LNL), highlighting their potential in the domain.
Chen Feng 0028, Georgios Tzimiropoulos, Ioannis Patras
ACM Multimedia3
2024 Multilinear Mixture of Experts: Scalable Expert Specialization through Factorization
abstract
The Mixture of Experts (MoE) paradigm provides a powerful way to decompose dense layers into smaller, modular computations often more amenable to human interpretation, debugging, and editability. However, a major challenge lies in the computational cost of scaling the number of experts high enough to achieve fine-grained specialization. In this paper, we propose the Multilinear Mixture of Experts (μMoE) layer to address this, focusing on vision models. μMoE layers enable scalable expert specialization by performing an implicit computation on prohibitively large weight tensors entirely in factorized form. Consequently, μMoEs (1) avoid the restrictively high inference-time costs of dense MoEs, yet (2) do not inherit the training issues of the popular sparse MoEs' discrete (non-differentiable) expert routing. We present both qualitative and quantitative evidence that scaling μMoE layers when fine-tuning foundation models for vision tasks leads to more specialized experts at the class-level, further enabling manual bias correction in CelebA attribute classification. Finally, we show qualitative results demonstrating the expert specialism achieved when pre-training large GPT2 and MLP-Mixer models with parameter-matched μMoE blocks at every layer, maintaining comparable accuracy. Our code is available at: https://github.com/james-oldfield/muMoE.
James Oldfield 0001, Markos Georgopoulos, Grigorios Chrysos 0002, Christos Tzelepis, Yannis Panagakis, Mihalis A. Nicolaou, Jiankang Deng, Ioannis Patras
NeurIPS8
2024 CemiFace: Center-based Semi-hard Synthetic Face Generation for Face Recognition
abstract
Privacy issue is a main concern in developing face recognition techniques. Although synthetic face images can partially mitigate potential legal risks while maintaining effective face recognition (FR) performance, FR models trained by face images synthesized by existing generative approaches frequently suffer from performance degradation problems due to the insufficient discriminative quality of these synthesized samples. In this paper, we systematically investigate what contributes to solid face recognition model training, and reveal that face images with certain degree of similarities to their identity centers show great effectiveness in the performance of trained FR models. Inspired by this, we propose a novel diffusion-based approach (namely **Ce**nter-based Se**mi**-hard Synthetic Face Generation (**CemiFace**) which produces facial samples with various levels of similarity to the subject center, thus allowing to generate face datasets containing effective discriminative samples for training face recognition. Experimental results show that with a modest degree of similarity, training on the generated dataset can produce competitive performance compared to previous generation methods. The code will be available at:https://github.com/szlbiubiubiu/CemiFace
Zhonglin Sun, Siyang Song, Ioannis Patras, Georgios Tzimiropoulos
NeurIPS3
2024 Improving Fairness using Vision-Language Driven Image Augmentation
abstract
Fairness is crucial when training a deep-learning discriminative model, especially in the facial domain. Models tend to correlate specific characteristics (such as age and skin color) with unrelated attributes (downstream tasks), resulting in biases which do not correspond to reality. It is common knowledge that these correlations are present in the data and are then transferred to the models during training (e.g., [35]). This paper proposes a method to mitigate these correlations to improve fairness. To do so, we learn interpretable and meaningful paths lying in the semantic space of a pre-trained diffusion model (DiffAE) [27] –such paths being supervised by contrastive text dipoles. That is, we learn to edit protected characteristics (age and skin color). These paths are then applied to augment images to improve the fairness of a given dataset. We test the proposed method on CelebA-HQ and UTKFace on several downstream tasks with age and skin color as protected characteristics. As a proxy for fairness, we compute the difference in accuracy with respect to the protected characteristics. Quantitative results show how the augmented images help the model improve the overall accuracy, the aforementioned metric, and the disparity of equal opportunity. Code is available at: https://github.com/Moreno98/Vision-Language-Bias-Control.
Moreno D'Incà, Christos Tzelepis, Ioannis Patras, Nicu Sebe
WACV3
2024 Self-Supervised Representation Learning with Cross-Context Learning between Global and Hypercolumn Features
abstract
Whilst contrastive learning yields powerful representations by matching different augmented views of the same instance, it lacks the ability to capture the similarities between different instances. One popular way to address this limitation is by learning global features (after the global pooling) to capture inter-instance relationships based on knowledge distillation, where the global features of the teacher are used to guide the learning of the global features of the student. Inspired by cross-modality learning, we extend this existing framework that only learns from global features by encouraging the global features and intermediate layer features to learn from each other. This leads to our novel self-supervised framework: cross-context learning between global and hypercolumn features (CGH), that enforces the consistency of instance relations between low-and high-level semantics. Specifically, we stack the intermediate feature maps to construct a "hypercolumn" representation so that we can measure instance relations using two contexts (hypercolumn and global feature) separately, and then use the relations of one context to guide the learning of the other. This cross-context learning allows the model to learn from the differences between the two contexts. The experimental results on linear classification and downstream tasks show that our method outperforms the state-of-the-art methods.
Zheng Gao 0003, Chen Feng 0028, Ioannis Patras
WACV3
2024 One-Shot Neural Face Reenactment via Finding Directions in GAN's Latent Space
abstract
Abstract In this paper, we present our framework for neural face/head reenactment whose goal is to transfer the 3D head orientation and expression of a target face to a source face. Previous methods focus on learning embedding networks for identity and head pose/expression disentanglement which proves to be a rather hard task, degrading the quality of the generated images. We take a different approach, bypassing the training of such networks, by using (fine-tuned) pre-trained GANs which have been shown capable of producing high-quality facial images. Because GANs are characterized by weak controllability, the core of our approach is a method to discover which directions in latent GAN space are responsible for controlling head pose and expression variations. We present a simple pipeline to learn such directions with the aid of a 3D shape model which, by construction, inherently captures disentangled directions for head pose, identity, and expression. Moreover, we show that by embedding real images in the GAN latent space, our method can be successfully used for the reenactment of real-world faces. Our method features several favorable properties including using a single source image (one-shot) and enabling cross-person reenactment. Extensive qualitative and quantitative results show that our approach typically produces reenacted faces of notably higher quality than those produced by state-of-the-art methods for the standard benchmarks of VoxCeleb1 & 2.
Stella Bounareli, Christos Tzelepis, Vasileios Argyriou, Ioannis Patras, Georgios Tzimiropoulos
Int. J. Comput. Vis.4
2024 Bilinear Models of Parts and Appearances in Generative Adversarial Networks
abstract
Recent advances in the understanding of Generative Adversarial Networks (GANs) have led to remarkable progress in visual editing and synthesis tasks, capitalizing on the rich semantics that are embedded in the latent spaces of pre-trained GANs. However, existing methods are often tailored to specific GAN architectures and are limited to either discovering global semantic directions that do not facilitate localized control, or require some form of supervision through manually provided regions or segmentation masks. In this light, we present an architecture-agnostic approach that jointly discovers factors representing spatial parts and their appearances in an entirely unsupervised fashion. These factors are obtained by applying a semi-nonnegative tensor factorization on the feature maps, which in turn enables context-aware local image editing with pixel-level control. In addition, we show that the discovered appearance factors correspond to saliency maps that localize concepts of interest, without using any labels. Experiments on a wide range of GAN architectures and datasets show that, in comparison to the state of the art, our method is far more efficient in terms of training time and, most importantly, provides much more accurate localized control.
James Oldfield 0001, Christos Tzelepis, Yannis Panagakis, Mihalis A. Nicolaou, Ioannis Patras
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 NoiseBox: Toward More Efficient and Effective Learning With Noisy Labels
abstract
Despite the large progress in supervised learning with neural networks, there are significant challenges in obtaining high-quality, large-scale and accurately labelled datasets. In such contexts, how to learn in the presence of noisy labels has received more and more attention. Addressing this relatively intricate problem to attain competitive results predominantly involves designing mechanisms that select samples that are expected to have reliable annotations. However, these methods typically involve multiple off-the-shelf techniques, resulting in intricate structures. Furthermore, they frequently make implicit or explicit assumptions about the noise modes/ratios within the dataset. Such assumptions can compromise model robustness and limit its performance under varying noise conditions. Unlike these methods, in this work, we propose an efficient and effective framework with minimal hyperparameters that achieves SOTA results in various benchmarks. Specifically, we design an efficient and concise training framework consisting of a subset expansion module responsible for exploring non-selected samples and a model training module to further reduce the impact of noise, called NoiseBox. Moreover, diverging from common sample selection methods based on the “small loss” mechanism, we introduce a novel sample selection method based on the neighbouring relationships and label consistency in the feature space. Without bells and whistles, such as model co-training, self-supervised pre-training and semi-supervised learning, and with robustness concerning the settings of its few hyper-parameters, our method significantly surpasses previous methods on both CIFAR10/CIFAR100 with synthetic noise and real-world noisy datasets such as Red Mini-ImageNet, WebVision, Clothing1M and ANIMAL-10N.
Chen Feng 0028, Georgios Tzimiropoulos, Ioannis Patras
IEEE Trans. Circuits Syst. Video Technol.3
2023 Prompting Visual-Language Models for Dynamic Facial Expression Recognition
Zengqun Zhao, Ioannis Patras
BMVC2
2023 Attribute-Preserving Face Dataset Anonymization via Latent Code Optimization
abstract
This work addresses the problem of anonymizing the identity of faces in a dataset of images, such that the privacy of those depicted is not violated, while at the same time the dataset is useful for downstream task such as for training machine learning models. To the best of our knowledge, we are the first to explicitly address this issue and deal with two major drawbacks of the existing state-of-the-art approaches, namely that they (i) require the costly training of additional, purpose-trained neural networks, and/or (ii) fail to retain the facial attributes of the original images in the anonymized counterparts, the preservation of which is of paramount importance for their use in downstream tasks. We accordingly present a task-agnostic anonymization procedure that directly optimizes the images' latent representation in the latent space of a pretrained GAN. By optimizing the latent codes directly, we ensure both that the identity is of a desired distance away from the original (with an identity obfuscation loss), whilst preserving the facial attributes (using a novel feature-matching loss in FaRL's [48] deep feature space). We demonstrate through a series of both qualitative and quantitative experiments that our method is capable of anonymizing the identity of the images whilst-crucially-better-preserving the facial attributes. We make the code and the pretrained models publicly available at: https://github.com/chi0tzp/FALCO.
Simone Barattin, Christos Tzelepis, Ioannis Patras, Nicu Sebe
CVPR3
2023 MaskCon: Masked Contrastive Learning for Coarse-Labelled Dataset
abstract
Deep learning has achieved great success in recent years with the aid of advanced neural network structures and large-scale human-annotated datasets. However, it is often costly and difficult to accurately and efficiently annotate large-scale datasets, especially for some specialized domains where fine-grained labels are required. In this setting, coarse labels are much easier to acquire as they do not require expert knowledge. In this work, we propose a contrastive learning method, called masked contrastive learning (MaskCon) to address the under-explored problem setting, where we learn with a coarse-labelled dataset in order to address a finer labelling problem. More specifically, within the contrastive learning framework, for each sample our method generates soft-labels with the aid of coarse labels against other samples and another augmented view of the sample in question. By contrast to self-supervised contrastive learning where only the sample's augmentations are considered hard positives, and in supervised contrastive learning where only samples with the same coarse labels are considered hard positives, we propose soft labels based on sample distances, that are masked by the coarse labels. This allows us to utilize both inter-sample relations and coarse labels. We demonstrate that our method can obtain as special cases many existing state-of-the-art works and that it provides tighter bounds on the generalization error. Experimentally, our method achieves significant improvement over the current state-of-the-art in various datasets, including CIFAR10, CIFAR100, ImageNet-1K, Standford Online Products and Stanford Cars196 datasets. Code and annotations are available at https://github.com/MrChenFeng/MaskCon_CVPR2023.
Chen Feng 0028, Ioannis Patras
CVPR2
2023 DivClust: Controlling Diversity in Deep Clustering
abstract
Clustering has been a major research topic in the field of machine learning, one to which Deep Learning has recently been applied with significant success. However, an aspect of clustering that is not addressed by existing deep clustering methods, is that of efficiently producing multiple, diverse partitionings for a given dataset. This is particularly important, as a diverse set of base clusterings are necessary for consensus clustering, which has been found to produce better and more robust results than relying on a single clustering. To address this gap, we propose Div-Clust, a diversity controlling loss that can be incorporated into existing deep clustering frameworks to produce multiple clusterings with the desired degree of diversity. We conduct experiments with multiple datasets and deep clustering frameworks and show that: a) our method effectively controls diversity across frameworks and datasets with very small additional computational cost, b) the sets of clusterings learned by DivClust include solutions that significantly outperform single-clustering baselines, and c) using an off-the-shelf consensus clustering algorithm, DivClust produces consensus clustering solutions that consistently outperform single-clustering baselines, effectively improving the performance of the base deep clustering framework. Code is available at https://github.com/ManiadisG/DivClust.
Ioannis Maniadis Metaxas, Georgios Tzimiropoulos, Ioannis Patras
CVPR3
2023 A Simple Baseline for Knowledge-Based Visual Question Answering
abstract
This paper is on the problem of Knowledge-Based Visual Question Answering (KB-VQA).Recent works have emphasized the significance of incorporating both explicit (through external databases) and implicit (through LLMs) knowledge to answer questions requiring external knowledge effectively.A common limitation of such approaches is that they consist of relatively complicated pipelines and often heavily rely on accessing GPT-3 API.Our main contribution in this paper is to propose a much simpler and readily reproducible pipeline which, in a nutshell, is based on efficient in-context learning by prompting LLaMA (1 and 2) using question-informative captions as contextual information.Contrary to recent approaches, our method is training-free, does not require access to external databases or APIs, and yet achieves state-of-the-art accuracy on the OK-VQA and A-OK-VQA datasets.Finally, we perform several ablation studies to understand important aspects of our method.
Alexandros Xenos, Themos Stafylakis, Ioannis Patras, Georgios Tzimiropoulos
EMNLP3
2023 StyleMask: Disentangling the Style Space of StyleGAN2 for Neural Face Reenactment
abstract
In this paper we address the problem of neural face reenactment, where, given a pair of a source and a target facial image, we need to transfer the target's pose (defined as the head pose and its facial expressions) to the source image, by preserving at the same time the source's identity characteristics (e.g., facial shape, hair style, etc), even in the challenging case where the source and the target faces belong to different identities. In doing so, we address some of the limitations of the state-of-the-art works, namely, a) that they depend on paired training data (i.e., source and target faces have the same identity), b) that they rely on labeled data during inference, and c) that they do not preserve identity in large head pose changes. More specifically, we propose a framework that, using unpaired randomly generated facial images, learns to disentangle the identity characteristics of the face from its pose by incorporating the recently introduced style space S [1] of StyleGAN2 [2], a latent representation space that exhibits remarkable disentanglement properties. By capitalizing on this, we learn to successfully mix a pair of source and target style codes using supervision from a 3D model. The resulting latent code, that is subsequently used for reenactment, consists of latent units corresponding to the facial pose of the target only and of units corresponding to the identity of the source only, leading to notable improvement in the reenactment performance compared to recent state-of-the-art methods. In comparison to state of the art, we quantitatively and qualitatively show that the proposed method produces higher quality results even on extreme pose variations. Finally, we report results on real images by first embedding them on the latent space of the pretrained generator. We make the code and the pretrained models publicly available at: https://github.com/StelaBou/StyleMask.
Stella Bounareli, Christos Tzelepis, Vasileios Argyriou, Ioannis Patras, Georgios Tzimiropoulos
FG4
2023 HyperReenact: One-Shot Reenactment via Jointly Learning to Refine and Retarget Faces
abstract
In this paper, we present our method for neural face reenactment, called HyperReenact, that aims to generate realistic talking head images of a source identity, driven by a target facial pose. Existing state-of-the-art face reenactment methods train controllable generative models that learn to synthesize realistic facial images, yet producing reenacted faces that are prone to significant visual artifacts, especially under the challenging condition of extreme head pose changes, or requiring expensive few-shot fine-tuning to better preserve the source identity characteristics. We propose to address these limitations by leveraging the photorealistic generation ability and the disentangled properties of a pretrained StyleGAN2 generator, by first inverting the real images into its latent space and then using a hypernetwork to perform: (i) refinement of the source identity characteristics and (ii) facial pose re-targeting, eliminating this way the dependence on external editing methods that typically produce artifacts. Our method operates under the one-shot setting (i.e., using a single source frame) and allows for cross-subject reenactment, without requiring any subject-specific fine-tuning. We compare our method both quantitatively and qualitatively against several state-of-the-art techniques on the standard benchmarks of VoxCeleb1 and VoxCeleb2, demonstrating the superiority of our approach in producing artifact-free images, exhibiting remarkable robustness even under extreme head pose changes. We make the code and the pretrained models publicly available at: https://github.com/StelaBou/HyperReenact.
Stella Bounareli, Christos Tzelepis, Vasileios Argyriou, Ioannis Patras, Georgios Tzimiropoulos
ICCV4
2023 Selecting A Diverse Set Of Aesthetically-Pleasing and Representative Video Thumbnails Using Reinforcement Learning
abstract
This paper presents a new reinforcement-based method for video thumbnail selection (called RL-DiVTS), that relies on estimates of the aesthetic quality, representativeness and visual diversity of a small set of selected frames, made with the help of tailored reward functions. The proposed method integrates a novel diversity-aware Frame Picking mechanism that performs a sequential frame selection and applies a reweighting process to demote frames that are visually-similar to the already selected ones. Experiments on two benchmark datasets (OVP and YouTube), using the top-3 matching evaluation protocol, show the competitiveness of RL-DiVTS against other SoA video thumbnail selection and summarization approaches from the literature.
Evlampios Apostolidis, Georgios Balaouras, Vasileios Mezaris, Ioannis Patras
ICIP4
2023 PandA: Unsupervised Learning of Parts and Appearances in the Feature Maps of GANs
James Oldfield 0001, Christos Tzelepis, Yannis Panagakis, Mihalis A. Nicolaou, Ioannis Patras
ICLR5
2023 NarSUM '23: The 2nd Workshop on User-Centric Narrative Summarization of Long Videos
abstract
With video capture devices becoming widely popular, the amount of video data generated per day has seen a rapid increase over the past few years. Browsing through hours of video data to retrieve useful information is a tedious and boring task. Video Summarization technology has played a crucial role in addressing this issue. It is a well-researched topic in the multimedia community. However, the focus so far has been limited to creating summary to videos which are short (only a few minutes). This workshop aims to call for researchers on relevant background to focus on novel solutions for user-centric narrative summarization of long videos. This workshop will also cover important aspects of video summarization research like what is "important" in a video, how to evaluate the goodness of a created summary, open challenges in video summarization, etc.
Mohan Kankanhalli, Ioannis Patras, Jianquan Liu, Yongkang Wong, Takahiro Komamizu, Satoshi Yamazaki, Karen Stephen, Kajal Kansal
ACM Multimedia2
2023 Low-Light Image Enhancement Based on U-Net and Haar Wavelet Pooling
Elissavet Batziou, Konstantinos Ioannidis, Ioannis Patras, Stefanos Vrochidis, Ioannis Kompatsiaris
MMM (2)3
2023 Parts of Speech-Grounded Subspaces in Vision-Language Models
abstract
Latent image representations arising from vision-language models have proved immensely useful for a variety of downstream tasks. However, their utility is limited by their entanglement with respect to different visual attributes. For instance, recent work has shown that CLIP image representations are often biased toward specific visual properties (such as objects or actions) in an unpredictable manner. In this paper, we propose to separate representations of the different visual modalities in CLIP’s joint vision-language space by leveraging the association between parts of speech and specific visual modes of variation (e.g. nouns relate to objects, adjectives describe appearance). This is achieved by formulating an appropriate component analysis model that learns subspaces capturing variability corresponding to a specific part of speech, while jointly minimising variability to the rest. Such a subspace yields disentangled representations of the different visual properties of an image or text in closed form while respecting the underlying geometry of the manifold on which the representations lie. What’s more, we show the proposed model additionally facilitates learning subspaces corresponding to specific visual appearances (e.g. artists’ painting styles), which enables the selective removal of entire visual themes from CLIP-based text-to-image synthesis. We validate the model both qualitatively, by visualising the subspace projections with a text-to-image model and by preventing the imitation of artists’ styles, and quantitatively, through class invariance metrics and improvements to baseline zero-shot classification.
James Oldfield 0001, Christos Tzelepis, Yannis Panagakis, Mihalis A. Nicolaou, Ioannis Patras
NeurIPS5
2023 "Just To See You Smile": SMILEY, a Voice-Guided GUY GAN
abstract
In this technical demonstration, we present SMILEY, a voice-guided virtual assistant. The system utilizes a deep neural architecture ContraCLIP to manipulate facial attributes using voice instructions, allowing for deeper speaker engagement and smoother customer experience when being used in the "virtual concierge" scenario. We validate the effectiveness of SMILEY and ContraCLIP via a successful real-world case study in Singapore and a large-scale quantitative evaluation.
Qi Yang 0005, Christos Tzelepis, Sergey I. Nikolenko, Ioannis Patras, Aleksandr Farseev
WSDM4
2023 Artistic neural style transfer using CycleGAN and FABEMD by adaptive information selection
Elissavet Batziou, Konstantinos Ioannidis, Ioannis Patras, Stefanos Vrochidis, Ioannis Kompatsiaris
Pattern Recognit. Lett.3
2022 SSR: An Efficient and Robust Framework for Learning with Unknown Label Noise
Chen Feng 0028, Georgios Tzimiropoulos, Ioannis Patras
BMVC3
2022 Adaptive Soft Contrastive Learning
abstract
Self-supervised learning has recently achieved great success in representation learning without human annotations. The dominant method – that is contrastive learning, is generally based on instance discrimination tasks, i.e., individual samples are treated as independent categories. However, presuming all the samples are different contradicts the natural grouping of similar samples in common visual datasets, e.g., multiple views of the same dog. To bridge the gap, this paper proposes an adaptive method that introduces soft inter-sample relations, namely Adaptive Soft Contrastive Learning (ASCL). More specifically, ASCL transforms the original instance discrimination task into a multi-instance soft discrimination task, and adaptively introduces inter-sample relations. As an effective and concise plug-in module for existing self-supervised learning frameworks, ASCL achieves the best performance on several benchmarks in terms of both performance and efficiency. Code is available at https://github.com/MrChenFeng/ASCL_ICPR2022.
Chen Feng 0028, Ioannis Patras
ICPR2
2022 Explaining video summarization based on the focus of attention
abstract
In this paper we propose a method for explaining video summarization. We start by formulating the problem as the creation of an explanation mask which indicates the parts of the video that influenced the most the estimates of a video summarization network, about the frames’ importance. Then, we explain how the typical analysis pipeline of attention-based networks for video summarization can be used to define explanation signals, and we examine various attention-based signals that have been studied as explanations in the NLP domain. We evaluate the performance of these signals by investigating the video summarization network’s input-output relationship according to different replacement functions, and utilizing measures that quantify the capability of explanations to spot the most and least influential parts of a video. We run experiments using an attention-based network (CA-SUM) and two datasets (SumMe and TVSum) for video summarization. Our evaluations indicate the advanced performance of explanations formed using the inherent attention weights, and demonstrate the ability of our method to explain the video summarization results using clues about the focus of the attention mechanism.
Evlampios Apostolidis, Georgios Balaouras, Vasileios Mezaris, Ioannis Patras
ISM4
2022 Summarizing Videos using Concentrated Attention and Considering the Uniqueness and Diversity of the Video Frames
abstract
In this work, we describe a new method for unsupervised video summarization. To overcome limitations of existing unsupervised video summarization approaches, that relate to the unstable training of Generator-Discriminator architectures, the use of RNNs for modeling long-range frames' dependencies and the ability to parallelize the training process of RNN-based network architectures, the developed method relies solely on the use of a self-attention mechanism to estimate the importance of video frames. Instead of simply modeling the frames' dependencies based on global attention, our method integrates a concentrated attention mechanism that is able to focus on non-overlapping blocks in the main diagonal of the attention matrix, and to enrich the existing information by extracting and exploiting knowledge about the uniqueness and diversity of the associated frames of the video. In this way, our method makes better estimates about the significance of different parts of the video, and drastically reduces the number of learnable parameters. Experimental evaluations using two benchmarking datasets (SumMe and TVSum) show the competitiveness of the proposed method against other state-of-the-art unsupervised summarization approaches, and demonstrate its ability to produce video summaries that are very close to the human preferences. An ablation study that focuses on the introduced components, namely the use of concentrated attention in combination with attention-based estimates about the frames' uniqueness and diversity, shows their relative contributions to the overall summarization performance.
Evlampios Apostolidis, Georgios Balaouras, Vasileios Mezaris, Ioannis Patras
ICMR4
2022 Learning from Label Relationships in Human Affect
abstract
Human affect and mental state estimation in an automated manner, face a number of difficulties, including learning from labels with poor or no temporal resolution, learning from few datasets with little data (often due to confidentiality constraints) and, (very) long, in-the-wild videos. For these reasons, deep learning methodologies tend to overfit, that is, arrive at latent representations with poor generalisation performance on the final regression task. To overcome this, in this work, we introduce two complementary contributions. First, we introduce a novel relational loss for multilabel regression and ordinal problems that regularises learning and leads to better generalisation. The proposed loss uses label vector inter-relational information to learn better latent representations by aligning batch label distances to the distances in the latent feature space. Second, we utilise a two-stage attention architecture that estimates a target for each clip by using features from the neighbouring clips as temporal context. We evaluate the proposed methodology on both continuous affect and schizophrenia severity estimation problems, as there are methodological and contextual parallels between the two. Experimental results demonstrate that the proposed methodology outperforms the baselines that are trained using the supervised regression loss, as well as pre-training the network architecture with an unsupervised contrastive loss. In the domain of schizophrenia, the proposed methodology outperforms previous state-of-the-art by a large margin, achieving a PCC of up to 78% performance close to that of human experts (85%) and much higher than previous works (uplift of up to 40%). In the case of affect recognition, we outperform previous vision-based methods in terms of CCC on both the OMG and the AMIGOS datasets. Specifically for AMIGOS, we outperform previous SoTA CCC for both arousal and valence by 9% and 13% respectively, and in the OMG dataset we outperform previous vision works by up to 5% for both arousal and valence.
Niki Maria Foteinopoulou, Ioannis Patras
ACM Multimedia2
2022 DnS: Distill-and-Select for Efficient and Accurate Video Indexing and Retrieval
abstract
Abstract In this paper, we address the problem of high performance and computationally efficient content-based video retrieval in large-scale datasets. Current methods typically propose either: (i) fine-grained approaches employing spatio-temporal representations and similarity calculations, achieving high performance at a high computational cost or (ii) coarse-grained approaches representing/indexing videos as global vectors, where the spatio-temporal structure is lost, providing low performance but also having low computational cost. In this work, we propose a Knowledge Distillation framework, called Distill-and-Select (DnS), that starting from a well-performing fine-grained Teacher Network learns: (a) Student Networks at different retrieval performance and computational efficiency trade-offs and (b) a Selector Network that at test time rapidly directs samples to the appropriate student to maintain both high retrieval performance and high computational efficiency. We train several students with different architectures and arrive at different trade-offs of performance and efficiency, i.e., speed and storage requirements, including fine-grained students that store/index videos using binary representations. Importantly, the proposed scheme allows Knowledge Distillation in large, unlabelled datasets—this leads to good students. We evaluate DnS on five public datasets on three different video retrieval tasks and demonstrate (a) that our students achieve state-of-the-art performance in several cases and (b) that the DnS framework provides an excellent trade-off between retrieval performance, computational speed, and storage space. In specific configurations, the proposed method achieves similar mAP with the teacher but is 20 times faster and requires 240 times less storage space. The collected dataset and implementation are publicly available: https://github.com/mever-team/distill-and-select .
Giorgos Kordopatis-Zilos, Christos Tzelepis, Symeon Papadopoulos, Ioannis Kompatsiaris, Ioannis Patras
Int. J. Comput. Vis.5
2021 Estimating continuous affect with label uncertainty
abstract
Continuous affect estimation is a problem where there is an inherent uncertainty and subjectivity in the labels that accompany data samples – typically, datasets use the average of multiple annotations or self-reporting to obtain ground truth labels. In this work, we propose a method for uncertainty-aware continuous affect estimation, that models explicitly the uncertainty of the ground truth label as a uni-variate Gaussian with mean equal to the ground truth label, and unknown variance. For each sample, the proposed neural network estimates not only the value of the target label (valence and arousal in our case), but also the variance. The network is trained with a loss that is defined as the KL-divergence between the estimation (valence/arousal) and the Gaussian around the ground truth. We show that, in two affect recognition problems with real data, the estimated variances are correlated with measures of uncertainty/error in the labels that are extracted either by considering multiple annotations of the data, or by manually cleaning the dataset.
Niki Maria Foteinopoulou, Christos Tzelepis, Ioannis Patras
ACII3
2021 Pairwise Ranking Network for Affect Recognition
abstract
In this work we study the problem of emotion recognition under the prism of preference learning. Affective datasets are typically annotated by assigning a single absolute label, i.e. a numerical value that describes the intensity of an emotional attribute, to each sample. Then, the majority of existing works on affect recognition employ sample-wise classification/regression methods to predict affective states, using those annotations. We take a different approach and use a deep network architecture that performs joint training on the tasks of classification/regression of samples and ordinal ranking between pairs of samples. By treating input samples in a pairwise manner, we leverage the auxiliary task of inferring the ordinal relation between their corresponding affective states. Incorporating the ranking objective allows capturing the inherently ordinal structure of emotions and learning the inter-sample relations, resulting in better generalization. Our method is incorporated into existing affect recognition architectures and evaluated on datasets of electroencephalograms (EEG) and images. We show that the approach proposed in this work leads to consistent performance gains when incorporated in classification/regression networks.
Georgios Zoumpourlis, Ioannis Patras
ACII2
2021 Tensor Component Analysis for Interpreting the Latent Space of GANs
James Oldfield 0001, Markos Georgopoulos, Yannis Panagakis, Mihalis A. Nicolaou, Ioannis Patras
BMVC5
2021 WarpedGANSpace: Finding non-linear RBF paths in GAN latent space
abstract
This work addresses the problem of discovering, in an unsupervised manner, interpretable paths in the latent space of pretrained GANs, so as to provide an intuitive and easy way of controlling the underlying generative factors. In doing so, it addresses some of the limitations of the state-of-the-art works, namely, a) that they discover directions that are independent of the latent code, i.e., paths that are linear, and b) that their evaluation relies either on visual inspection or on laborious human labeling. More specifically, we propose to learn non-linear warpings on the latent space, each one parametrized by a set of RBF-based latent space warping functions, and where each warping gives rise to a family of non-linear paths via the gradient of the function. Building on the work of [34], that discovers linear paths, we optimize the trainable parameters of the set of RBFs, so as that images that are generated by codes along different paths, are easily distinguishable by a discriminator network. This leads to easily distinguishable image transformations, such as pose and facial expressions in facial images. We show that linear paths can be derived as a special case of our method, and show experimentally that non-linear paths in the latent space lead to steeper, more disentangled and interpretable changes in the image space than in state-of-the art methods, both qualitatively and quantitatively. We make the code and the pretrained models publicly available at: https://github.com/chi0tzp/WarpedGANSpace.
Christos Tzelepis, Georgios Tzimiropoulos, Ioannis Patras
ICCV3
2021 Combining Global and Local Attention with Positional Encoding for Video Summarization
abstract
This paper presents a new method for supervised video summarization. To overcome drawbacks of existing RNN-based summarization architectures, that relate to the modeling of long-range frames’ dependencies and the ability to parallelize the training process, the developed model re-lies on the use of self-attention mechanisms to estimate the importance of video frames. Contrary to previous attention-based summarization approaches that model the frames’ dependencies by observing the entire frame sequence, our method combines global and local multi-head attention mechanisms to discover different modelings of the frames’ dependencies at different levels of granularity. Moreover, the utilized attention mechanisms integrate a component that encodes the temporal position of video frames - this is of major importance when producing a video summary. Experiments on two datasets (SumMe and TVSum) demonstrate the effectiveness of the proposed model compared to existing attention-based methods, and its competitiveness against other state-of-the-art supervised summarization approaches. An ablation study that focuses on our main proposed components, namely the use of global and local multi-head attention mechanisms in collaboration with an absolute positional encoding component, shows their relative contributions to the overall summarization performance.
Evlampios Apostolidis, Georgios Balaouras, Vasileios Mezaris, Ioannis Patras
ISM4
2021 Combining Adversarial and Reinforcement Learning for Video Thumbnail Selection
abstract
This paper presents a new method for unsupervised video thumbnail selection. The developed network architecture selects video thumbnails based on two criteria: the representativeness and the aesthetic quality of their visual content. Training relies on a combination of adversarial and reinforcement learning. The former is used to train a discriminator, whose goal is to distinguish the original from a reconstructed version of the video based on a small set of candidate thumbnails. The discriminator's feedback is a measure of the representativeness of the selected thumbnails. This measure is combined with estimates about the aesthetic quality of the thumbnails (made using a SoA Fully Convolutional Network) to form a reward and train the thumbnail selector via reinforcement learning. Experiments on two datasets (OVP and Youtube) show the competitiveness of the proposed method against other SoA approaches. An ablation study with respect to the adopted thumbnail selection criteria documents the importance of considering the aesthetics, and the contribution of this information when used in combination with measures about the representativeness of the visual content.
Evlampios Apostolidis, Eleni Adamantidou, Vasileios Mezaris, Ioannis Patras
ICMR4
2021 Few-Shot Action Localization without Knowing Boundaries
abstract
Learning to localize actions in long, cluttered, and untrimmed videos is a hard task, that in the literature has typically been addressed assuming the availability of large amounts of annotated training samples for each class -- either in a fully-supervised setting, where action boundaries are known, or in a weakly-supervised setting, where only class labels are known for each video. In this paper, we go a step further and show that it is possible to learn to localize actions in untrimmed videos when a) only one/few trimmed examples of the target action are available at test time, and b) when a large collection of videos with only class label annotation (some trimmed and some weakly annotated untrimmed ones) are available for training; with no overlap between the classes used during training and testing. To do so, we propose a network that learns to estimate Temporal Similarity Matrices (TSMs) that model a fine-grained similarity pattern between pairs of videos (trimmed or untrimmed), and uses them to generate Temporal Class Activation Maps (TCAMs) for seen or unseen classes. The TCAMs serve as temporal attention mechanisms to extract video-level representations of untrimmed videos, and to temporally localize actions at test time. To the best of our knowledge, we are the first to propose a weakly-supervised, one/few-shot action localization network that can be trained in an end-to-end fashion. Experimental results on THUMOS14 and ActivityNet1.2 datasets, show that our method achieves performance comparable or better to state-of-the-art fully-supervised, few-shot learning methods.
Ting-Ting Xie, Christos Tzelepis, Fan Fu, Ioannis Patras
ICMR4
2021 Video Summarization Using Deep Neural Networks: A Survey
abstract
Video summarization technologies aim to create a concise and complete synopsis by selecting the most informative parts of the video content. Several approaches have been developed over the last couple of decades, and the current state of the art is represented by methods that rely on modern deep neural network architectures. This work focuses on the recent advances in the area and provides a comprehensive survey of the existing deep-learning-based methods for generic video summarization. After presenting the motivation behind the development of technologies for video summarization, we formulate the video summarization task and discuss the main characteristics of a typical deep-learning-based analysis pipeline. Then, we suggest a taxonomy of the existing algorithms and provide a systematic review of the relevant literature that shows the evolution of the deep-learning-based video summarization technologies and leads to suggestions for future developments. We then report on protocols for the objective evaluation of video summarization algorithms, and we compare the performance of several deep-learning-based approaches. Based on the outcomes of these comparisons, as well as some documented considerations about the amount of annotated data and the suitability of evaluation protocols, we indicate potential future research directions.
Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris, Ioannis Patras
Proc. IEEE5
2021 SchiNet: Automatic Estimation of Symptoms of Schizophrenia from Facial Behaviour Analysis
abstract
Patients with schizophrenia often display impairments in the expression of emotion and speech and those are observed in their facial behaviour. Automatic analysis of patients’ facial expressions that is aimed at estimating symptoms of schizophrenia has received attention recently. However, the datasets that are typically used for training and evaluating the developed methods, contain only a small number of patients (4-34) and are recorded while the subjects were performing controlled tasks such as listening to life vignettes, or answering emotional questions. In this paper, we use videos of professional-patient interviews, in which symptoms were assessed in a standardised way as they should/may be assessed in practice, and which were recorded in realistic conditions (i.e., varying illumination levels and camera viewpoints) at the patients’ homes or at mental health services. We automatically analyse the facial behaviour of 91 out-patients – this is almost 3 times the number of patients in other studies – and propose SchiNet, a novel neural network architecture that estimates expression-related symptoms in two different assessment interviews. We evaluate the proposed SchiNet for patient-independent prediction of symptoms of schizophrenia. Experimental results show that some automatically detected facial expressions are significantly correlated to symptoms of schizophrenia, and that the proposed network for estimating symptom severity delivers promising results.
Mina Bishay, Petar Palasek, Stefan Priebe, Ioannis Patras
IEEE Trans. Affect. Comput.4
2021 AMIGOS: A Dataset for Affect, Personality and Mood Research on Individuals and Groups
abstract
We present AMIGOS- A dataset for Multimodal research of affect, personality traits and mood on Individuals and GrOupS. Different to other databases, we elicited affect using both short and long videos in two social contexts, one with individual viewers and one with groups of viewers. The database allows the multimodal study of the affective responses, by means of neuro-physiological signals of individuals in relation to their personality and mood, and with respect to the social context and videos' duration. The data is collected in two experimental settings. In the first one, 40 participants watched 16 short emotional videos. In the second one, the participants watched 4 long videos, some of them alone and the rest in groups. The participants' signals, namely, Electroencephalogram (EEG), Electrocardiogram (ECG) and Galvanic Skin Response (GSR), were recorded using wearable sensors. Participants' frontal HD video and both RGB and depth full body videos were also recorded. Participants emotions have been annotated with both self-assessment of affective levels (valence, arousal, control, familiarity, liking and basic emotions) felt during the videos as well as external-assessment of levels of valence and arousal. We present a detailed correlation analysis of the different dimensions as well as baseline methods and results for single-trial classification of valence and arousal, personality traits, mood and social context. The database is made publicly available.
Juan Abdon Miranda Correa, Mojtaba Khomami Abadi, Nicu Sebe, Ioannis Patras
IEEE Trans. Affect. Comput.4
2021 AC-SUM-GAN: Connecting Actor-Critic and Generative Adversarial Networks for Unsupervised Video Summarization
abstract
This paper presents a new method for unsupervised video summarization. The proposed architecture embeds an Actor-Critic model into a Generative Adversarial Network and formulates the selection of important video fragments (that will be used to form the summary) as a sequence generation task. The Actor and the Critic take part in a game that incrementally leads to the selection of the video key-fragments, and their choices at each step of the game result in a set of rewards from the Discriminator. The designed training workflow allows the Actor and Critic to discover a space of actions and automatically learn a policy for key-fragment selection. Moreover, the introduced criterion for choosing the best model after the training ends, enables the automatic selection of proper values for parameters of the training process that are not learned from the data (such as the regularization factor σ). Experimental evaluation on two benchmark datasets (SumMe and TVSum) demonstrates that the proposed AC-SUM-GAN model performs consistently well and gives SoA results in comparison to unsupervised methods, that are also competitive with respect to supervised methods.
Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris, Ioannis Patras
IEEE Trans. Circuits Syst. Video Technol.5
2020 Cycle-Consistent Adversarial Networks and Fast Adaptive Bi-dimensional Empirical Mode Decomposition for Style Transfer
abstract
Recently, research endeavors have shown the potentiality of Cycle-Consistent Adversarial Networks (CycleGAN) in style transfer. In Cycle-Consistent Adversarial Networks, the consistency loss is introduced to measure the difference between the original images and the reconstructed in both directions, forward and backward. In this work, the combination of Cycle-Consistent Adversarial Networks with Fast and Adaptive Bidimensional Empirical Mode Decomposition (FABEMD) is proposed to perform style transfer on images. In the proposed approach the cycle-consistency loss is modified to include the differences between the extracted Intrinsic Mode Functions (BIMFs) images. Instead of an estimation of pixel-to-pixel difference between the produced and input images, the FABEMD is applied and the extracted BIMFs are involved in the computation of the total cycle loss. This method enriches the computation of the total loss in a content-to-content and style-to-style comparison by connecting the spatial information to the frequency components. The experimental results reveal that the proposed method is efficient and produces qualitative results comparable to state-of-the-art methods.
Elissavet Batziou, Petros Alvanitopoulos, Konstantinos Ioannidis, Ioannis Patras, Stefanos Vrochidis, Ioannis Kompatsiaris
ICPR4
2020 Performance over Random: A Robust Evaluation Protocol for Video Summarization Methods
abstract
This paper proposes a new evaluation approach for video summarization algorithms. We start by studying the currently established evaluation protocol; this protocol, defined over the ground-truth annotations of the SumMe and TVSum datasets, quantifies the agreement between the user-defined and the automatically-created summaries with F-Score, and reports the average performance on a few different training/testing splits of the used dataset. We evaluate five publicly-available summarization algorithms under a large-scale experimental setting with 50 randomly-created data splits. We show that the results reported in the papers are not always congruent with their performance on the large-scale experiment, and that the F-Score cannot be used for comparing algorithms evaluated on different splits. We also show that the above shortcomings of the established evaluation protocol are due to the significantly varying levels of difficulty among the utilized splits, that affect the outcomes of the evaluations. Further analysis of these findings indicates a noticeable performance correlation among all algorithms and a random summarizer. To mitigate these shortcomings we propose an evaluation protocol that makes estimates about the difficulty of each used data split and utilizes this information during the evaluation process. Experiments involving different evaluation settings demonstrate the increased representativeness of performance results when using the proposed evaluation approach, and the increased reliability of comparisons when the examined methods have been evaluated on different data splits.
Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris, Ioannis Patras
ACM Multimedia5
2020 Unsupervised Video Summarization via Attention-Driven Adversarial Learning
Evlampios Apostolidis, Eleni Adamantidou, Alexandros I. Metsai, Vasileios Mezaris, Ioannis Patras
MMM (1)5
2019 TARN: Temporal Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition
Mina Bishay, Georgios Zoumpourlis, Ioannis Patras
BMVC3
2019 Your Fellows Matter: Affect Analysis across Subjects in Group Videos
abstract
Automatic affect analysis has become a well established research area in the last two decades. Recent works have started moving from individual to group scenarios. However, little attention has been paid to investigating how individuals in a group influence the affective states of each other. In this paper, we propose a novel framework for cross-subjects affect analysis in group videos. Specifically, we analyze the correlation of the affect among group members and investigate the automatic recognition of the affect of one subject using the behaviours expressed by another subject in the same group. A set of experiments are conducted using a recently collected database aimed at affect analysis in group settings. Our results show that (1) people in the same group do share more information in terms of behaviours and emotions than people in different groups; and (2) the affect of one subject in a group can be better predicted using the expressive behaviours of another subject within the same group than using that of a subject from a different group. This work is of great importance for affect recognition in group settings: when the information of one subject is unavailable due to occlusion, head/body poses etc., we can predict his/her affect by employing the expressive behaviours of the other subject(s).
Wenxuan Mou, Hatice Gunes, Ioannis Patras
FG3
2019 Can Automatic Facial Expression Analysis Be Used for Treatment Outcome Estimation in Schizophrenia?
abstract
Negative symptoms of schizophrenia include expressive deficits that are marked by a reduction in patients' behaviour. Analysing automatically non-verbal behaviour and exploiting the results for estimating symptom severity has drawn attention recently. However, those approaches are not accurate enough to be used for monitoring the changes in patient's symptom level during treatment interventions (i.e. the treatment outcome). In this paper, we propose a method that directly addresses the problem of Treatment Outcome Estimation (TOE) in schizophrenia - more specifically, is aimed at determining whether specific symptoms have improved or not by analysing jointly two videos of the same patient, one before and one after the treatment. The proposed architecture builds on Recurrent Neural Networks (RNNs) that learn differences in the patient behaviour before and after treatment. We validate the method in videotaped interviews for symptom assessment for 74 patients. Experimental results show that the proposed architecture achieves promising results for TOE in two different symptom assessment scales.
Mina Bishay, Stefan Priebe, Ioannis Patras
ICASSP3
2019 ViSiL: Fine-Grained Spatio-Temporal Video Similarity Learning
abstract
In this paper we introduce ViSiL, a Video Similarity Learning architecture that considers fine-grained Spatio-Temporal relations between pairs of videos -- such relations are typically lost in previous video retrieval approaches that embed the whole frame or even the whole video into a vector descriptor before the similarity estimation. By contrast, our Convolutional Neural Network (CNN)-based approach is trained to calculate video-to-video similarity from refined frame-to-frame similarity matrices, so as to consider both intra- and inter-frame relations. In the proposed method, pairwise frame similarity is estimated by applying Tensor Dot (TD) followed by Chamfer Similarity (CS) on regional CNN frame features - this avoids feature aggregation before the similarity calculation between frames. Subsequently, the similarity matrix between all video frames is fed to a four-layer CNN, and then summarized using Chamfer Similarity (CS) into a video-to-video similarity score - this avoids feature aggregation before the similarity calculation between videos and captures the temporal similarity patterns between matching frame sequences. We train the proposed network using a triplet loss scheme and evaluate it on five public benchmark datasets on four different video retrieval problems where we demonstrate large improvements in comparison to the state of the art. The implementation of ViSiL is publicly available.
Giorgos Kordopatis-Zilos, Symeon Papadopoulos, Ioannis Patras, Ioannis Kompatsiaris
ICCV3
2019 Exploring Feature Representation and Training Strategies in Temporal Action Localization
abstract
Temporal action localization has recently attracted significant interest in the Computer Vision community. However, despite the great progress, it is hard to identify which aspects of the proposed methods contribute most to the increase in localization performance. To address this issue, we conduct ablative experiments on feature extraction methods, fixed-size feature representation methods and training strategies, and report how each influences the overall performance. Based on our findings, we propose a two-stage detector that outperforms the state of the art in THUMOS14, achieving a mAP@tIoU=0.5 equal to 44.20%.
Tingting Xie, Xiaoshan Yang, Tianzhu Zhang 0001, Changsheng Xu, Ioannis Patras
ICIP5
2019 VERGE in VBS 2019
Stelios Andreadis, Anastasia Moumtzidou, Damianos Galanopoulos, Fotini Markatopoulou, Konstantinos Apostolidis, Thanassis Mavropoulos, Ilias Gialampoukidis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris, Ioannis Patras
MMM (2)11
2019 Multimodal Video Annotation for Retrieval and Discovery of Newsworthy Video in a News Verification Scenario
Lyndon J. B. Nixon, Evlampios Apostolidis, Fotini Markatopoulou, Ioannis Patras, Vasileios Mezaris
MMM (1)4
2019 Detecting Tampered Videos with Multimedia Forensics and Deep Learning
Markos Zampoglou, Fotini Markatopoulou, Grégoire Mercier, Despoina Touska, Evlampios Apostolidis, Symeon Papadopoulos, Roger Cozien, Ioannis Patras, Vasileios Mezaris, Ioannis Kompatsiaris
MMM (1)8
2019 Registration-free Face-SSD: Single shot analysis of smiles, facial attributes, and affect in the wild
Youngkyoon Jang, Hatice Gunes, Ioannis Patras
Comput. Vis. Image Underst.3
2019 A deep generic to specific recognition model for group membership analysis using non-verbal cues
abstract
Automatic understanding and analysis of groups has attracted increasing attention in the vision and multimedia communities in recent years. However, little attention has been paid to the automatic analysis of the non-verbal behaviors and how this can be utilized for analysis of group membership, i.e., recognizing which group each individual is part of. This paper presents a novel Support Vector Machine (SVM) based Deep Specific Recognition Model (DeepSRM) that is learned based on a generic recognition model . The generic recognition model refers to the model trained with data across different conditions, i.e., when people are watching movies of different types. Although the generic recognition model can provide a baseline for the recognition model trained for each specific condition, the different behaviors people exhibit in different conditions limit the recognition performance of the generic model. Therefore, the specific recognition model is proposed for each condition separately and built on top of the generic recognition model . A number of experiments are conducted using a database aiming to study group analysis while each group (i.e., four participants together) were watching a number of long movie segments. Our experimental results show that the proposed deep specific recognition model (44%) outperforms the generic recognition model (26%). The recognition of group membership also indicates that the non-verbal behaviors of individuals within a group share commonalities.
Wenxuan Mou, Christos Tzelepis, Vasileios Mezaris, Hatice Gunes, Ioannis Patras
Image Vis. Comput.5
2019 Implicit and Explicit Concept Relations in Deep Neural Networks for Multi-Label Video/Image Annotation
abstract
In this paper, we propose a deep convolutional neural network (DCNN) architecture that addresses the problem of video/image concept annotation by exploiting concept relations at two different levels. At the first level, we build on ideas from multi-task learning, and propose an approach to learn concept-specific representations that are sparse, linear combinations of representations of latent concepts. By enforcing the sharing of the latent concept representations, we exploit the implicit relations between the target concepts. At a second level, we build on ideas from structured output learning and propose the introduction, at training time, of a new cost term that explicitly models the correlations between the concepts. By doing so, we explicitly model the structure in the output space (i.e., the concept labels). Both of the above are implemented using standard convolutional layers and are incorporated in a single DCNN architecture that can then be trained end-to-end with standard back-propagation. Experiments on four large-scale video and image data sets show that the proposed DCNN improves concept annotation accuracy and outperforms the related state-of-the-art methods.
Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
IEEE Trans. Circuits Syst. Video Technol.3
2019 FIVR: Fine-Grained Incident Video Retrieval
abstract
This paper introduces the problem of Fine-grained Incident Video Retrieval (FIVR). Given a query video, the objective is to retrieve all associated videos, considering several types of associations that range from duplicate videos to videos from the same incident. FIVR offers a single framework that contains several retrieval tasks as special cases. To address the benchmarking needs of all such tasks, we construct and present a large-scale annotated video dataset, which we call FIVR-200K, and it comprises 225,960 videos. To create the dataset, we devise a process for the collection of YouTube videos based on major news events from recent years crawled from Wikipedia and deploy a retrieval pipeline for the automatic selection of query videos based on their estimated suitability as benchmarks. We also devise a protocol for the annotation of the dataset with respect to the four types of video associations defined by FIVR. Finally, we report the results of an experimental study on the dataset comparing five state-of-the-art methods developed based on a variety of visual descriptors, highlighting the challenges of the current problem.
Giorgos Kordopatis-Zilos, Symeon Papadopoulos, Ioannis Patras, Ioannis Kompatsiaris
IEEE Trans. Multim.3
2019 Alone versus In-a-group: A Multi-modal Framework for Automatic Affect Recognition
abstract
Recognition and analysis of human affect has been researched extensively within the field of computer science in the past two decades. However, most of the past research in automatic analysis of human affect has focused on the recognition of affect displayed by people in individual settings and little attention has been paid to the analysis of the affect expressed in group settings. In this article, we first analyze the affect expressed by each individual in terms of arousal and valence dimensions in both individual and group videos and then propose methods to recognize the contextual information, i.e., whether a person is alone or in-a-group by analyzing their face and body behavioral cues. For affect analysis, we first devise affect recognition models separately in individual and group videos and then introduce a cross-condition affect recognition model that is trained by combining the two different types of data. We conduct a set of experiments on two datasets that contain both individual and group videos. Our experiments show that (1) the proposed Volume Quantized Local Zernike Moments Fisher Vector outperforms other unimodal features in affect analysis; (2) the temporal learning model, Long-Short Term Memory Networks, works better than the static learning model, Support Vector Machine; (3) decision fusion helps to improve affect recognition, indicating that body behaviors carry emotional information that is complementary rather than redundant to the emotion content in facial behaviors; and (4) it is possible to predict the context, i.e., whether a person is alone or in-a-group, using their non-verbal behavioral cues.
Wenxuan Mou, Hatice Gunes, Ioannis Patras
ACM Trans. Multim. Comput. Commun. Appl.3
2018 Deep Mixture of MRFs for Human Pose Estimation
Ioannis Marras, Petar Palasek, Ioannis Patras
ACCV (3)3
2018 LikeNet: A Siamese Motion Estimation Network Trained in an Unsupervised Way
Aria Ahmadi, Ioannis Marras, Ioannis Patras
BMVC3
2018 A Multi-Task Cascaded Network for Prediction of Affect, Personality, Mood and Social Context Using EEG Signals
abstract
This paper presents a multi-task cascaded deep neural network which jointly predicts people's affective levels (valence and arousal) and personal factors using EEG signals recorded in response to presentation of affective multimedia content. Studied personal factors are the Big-five personality traits, mood (Positive Affect and Negative Affect Schedules) and social context (individual vs group). The cascaded network consists of two levels of prediction. The first level consists of a hybrid network designed to predict affective levels for individual video segments by combining the capabilities of CNN and RNN units. The second level consists of an RNN unit designed to predict personal factors based on the sequence of affective levels predictions of consecutive video segments. The first level reduces the dimensionality of the input while keeping affective information. The fusion of decisions of RNN and CNN networks produces an increase in performance of 3% and 4% f1-score respectively for valence and arousal recognition. And, our results for personal factors recognition, on average, outperform baseline studies by at least 2:7% mean f1-score.
Juan Abdon Miranda Correa, Ioannis Patras
FG2
2018 VERGE in VBS 2018
Anastasia Moumtzidou, Stelios Andreadis, Fotini Markatopoulou, Damianos Galanopoulos, Ilias Gialampoukidis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris, Ioannis Patras
MMM (2)9
2018 Linear Maximum Margin Classifier for Learning from Uncertain Data
abstract
In this paper, we propose a maximum margin classifier that deals with uncertainty in data input. More specifically, we reformulate the SVM framework such that each training example can be modeled by a multi-dimensional Gaussian distribution described by its mean vector and its covariance matrix-the latter modeling the uncertainty. We address the classification problem and define a cost function that is the expected value of the classical SVM cost when data samples are drawn from the multi-dimensional Gaussian distributions that form the set of the training examples. Our formulation approximates the classical SVM formulation when the training examples are isotropic Gaussians with variance tending to zero. We arrive at a convex optimization problem, which we solve efficiently in the primal form using a stochastic gradient descent approach. The resulting classifier, which we name SVM with Gaussian Sample Uncertainty (SVM-GSU), is tested on synthetic data and five publicly available and popular datasets; namely, the MNIST, WDBC, DEAP, TV News Channel Commercial Detection, and TRECVID MED datasets. Experimental results verify the effectiveness of the proposed method.
Christos Tzelepis, Vasileios Mezaris, Ioannis Patras
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 Background modelling based on generative unet
abstract
Background Modelling is a crucial step in background/foreground detection which could be used in video analysis, such as surveillance, people counting, face detection and pose estimation. Most methods need to choose the hyper parameters manually or use ground truth background masks (GT). In this work, we present an unsupervised deep background (BG) modelling method called BM-Unet which is based on a generative architecture that given a certain frame as input it generates as output the corresponding background image - to be more precise, a probabilistic heat map of the colour values. Our method learns parameters automatically and an augmented version of it that utilises colour, intensity differences and optical flow between a reference and a target frame is robust to rapid illumination changes and camera jitter. Besides, it can be used on a new video sequence without the need of ground truth background/foreground masks for training. Experiment evaluations on challenging sequences in SBMnet data set demonstrate promising results over state-of-the-art methods.
Ye Tao 0006, Petar Palasek, Zhihao Ling, Ioannis Patras
AVSS4
2017 Fusing Multilabel Deep Networks for Facial Action Unit Detection
abstract
The automatic detection of the activation of facial muscles, i.e. the detection of the so called facial Action Units (AUs), has received significant attention due to the application of facial expression analysis/recognition in areas such as affect recognition or behavior analysis. However, the recognition of subtle expressions is a challenging task that requires a multimodal approach where several sources of information are used. In this paper, we follow such an approach and propose a novel Deep Learning architecture that fuses information from several specialized Deep Neural Networks (DNNs) each of which models a different aspect of the problem in question. At the core of our approach is a novel dynamic adaptation of the Deep Network cost function so as to deal with the data imbalances that are inherent in multilabel classification problems - this allows crossdatabase training.We show the benefits of the proposed training approach and how different architectures are more suitable for particular AUs. Extensive experimental results show that our multi-modal approach outperform the state of the art by a considerable margin.
Mina Bishay, Ioannis Patras
FG2
2017 Deep Refinement Convolutional Networks for Human Pose Estimation
abstract
This work introduces a novel Convolutional Network architecture (ConvNet) for the task of human pose estimation, that is the localization of body joints in a single static image. The proposed coarse to fine architecture addresses shortcomings of the baseline architecture that stem from the fact that large inaccuracies of its coarse ConvNet cannot becorrected by the refinement ConvNet that refines the estimation with in small windows of the coarse prediction. This is achieved by (a) changes in architectural parameters that both increase the accuracy of the coarse model and make the refinement model more capable of correcting the errors of the coarse model, (b) the introduction of a Markov Random Field (MRF)-based spatial model network between the coarse and the refinement model that introduces geometric constraints and (c) a training scheme that adapts the data augmentation and the learning rate according to the difficulty of the data examples. The proposed architecture is trained in an end-to-end fashion. Experimental results show that the proposed method improves the baseline model and provides state of the art results on the FashionPose[8] and MPII benchmarks [1].
Ioannis Marras, Petar Palasek, Ioannis Patras
FG3
2017 Generic to Specific Recognition Models for Membership Analysis in Group Videos
abstract
Automatic understanding and analysis of groups has attracted increasing attention in the vision and multimedia communities in recent years. However, little attention has been paid to the automatic analysis of group membership - i.e., recognizing which group the individual in question is part of. This paper presents a novel two-phase Support Vector Machine (SVM) based specific recognition model that is learned using an optimized generic recognition model. We conduct a set of experiments using a database collected to study group analysis from multimodal cues while each group (i.e., four participants together) were watching a number of long movie segments. Our experimental results show that the proposed specific recognition model (52%) outperforms the generic recognition model trained across all different videos (35%) and the independent recognition model trained directly on each specific video (33%) using linear SVM.
Wenxuan Mou, Christos Tzelepis, Vasileios Mezaris, Hatice Gunes, Ioannis Patras
FG5
2017 Deep Globally Constrained MRFs for Human Pose Estimation
abstract
This work introduces a novel Convolutional Network architecture (ConvNet) for the task of human pose estimation, that is the localization of body joints in a single static image. We propose a coarse to fine architecture that addresses shortcomings of the baseline architecture in [26] that stem from the fact that large inaccuracies of its coarse ConvNet cannot be corrected by the refinement ConvNet that refines the estimation within small windows of the coarse prediction. We overcome this by introducing a Markov Random Field (MRF)-based spatial model network between the coarse and the refinement model that introduces geometric constraints on the relative locations of the body joints. We propose an architecture in which a) the filters that implement the message passing in the MRF inference are factored in a way that constrains them by a low dimensional pose manifold the projection to which is estimated by a separate branch of the proposed ConvNet and b) the strengths of the pairwise joint constraints are modeled by weights that are jointly estimated by the other parameters of the network. The proposed network is trained in an end-to-end fashion. Experimental results show that the proposed method improves the baseline model and provides state of the art results on very challenging benchmarks.
Ioannis Marras, Petar Palasek, Ioannis Patras
ICCV3
2017 VideoAnalysis4ALL: An On-line Tool for the Automatic Fragmentation and Concept-based Annotation, and the Interactive Exploration of Videos
abstract
This paper presents the VideoAnalysis4ALL tool that supports the automatic fragmentation and concept-based annotation of videos, and the exploration of the annotated video fragments through an interactive user interface. The developed web application decomposes the video into two different granularities, namely shots and scenes, and annotates each fragment by evaluating the existence of a number (several hundreds) of high-level visual concepts in the keyframes extracted from these fragments. Through the analysis the tool enables the identification and labeling of semantically coherent video fragments, while its user interfaces allow the discovery of these fragments with the help of human-interpretable concepts. The integrated state-of-the-art video analysis technologies perform very well and, by exploiting the processing capabilities of multi-thread / multi-core architectures, reduce the time required for analysis to approximately one third of the video's duration, thus making the analysis three times faster than real-time processing.
Chrysa Collyda, Evlampios Apostolidis, Alexandros Pournaras, Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
ICMR6
2017 Concept Language Models and Event-based Concept Number Selection for Zero-example Event Detection
abstract
Zero-example event detection is a problem where, given an event query as input but no example videos for training a detector, the system retrieves the most closely related videos. In this paper we present a fully-automatic zero-example event detection method that is based on translating the event description to a predefined set of concepts for which previously trained visual concept detectors are available. We adopt the use of Concept Language Models (CLMs), which is a method of augmenting semantic concept definition, and we propose a new concept-selection method for deciding on the appropriate number of the concepts needed to describe an event query. The proposed system achieves state-of-the-art performance in automatic zero-example event detection.
Damianos Galanopoulos, Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
ICMR4
2017 Query and Keyframe Representations for Ad-hoc Video Search
abstract
This paper presents a fully-automatic method that combines video concept detection and textual query analysis in order to solve the problem of ad-hoc video search. We present a set of NLP steps that cleverly analyse different parts of the query in order to convert it to related semantic concepts, we propose a new method for transforming concept-based keyframe and query representations into a common semantic embedding space, and we show that our proposed combination of concept-based representations with their corresponding semantic embeddings results to improved video search accuracy. Our experiments in the TRECVID AVS 2016 and the Video Search 2008 datasets show the effectiveness of the proposed method compared to other similar approaches.
Fotini Markatopoulou, Damianos Galanopoulos, Vasileios Mezaris, Ioannis Patras
ICMR4
2017 Near-Duplicate Video Retrieval by Aggregating Intermediate CNN Layers
Giorgos Kordopatis-Zilos, Symeon Papadopoulos, Ioannis Patras, Ioannis Kompatsiaris
MMM (1)3
2017 VERGE in VBS 2017
Anastasia Moumtzidou, Theodoros Mironidis, Fotini Markatopoulou, Stelios Andreadis, Ilias Gialampoukidis, Damianos Galanopoulos, Anastasia Ioannidou, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris, Ioannis Patras
MMM (2)11
2017 Comparison of Fine-Tuning and Extension Strategies for Deep Convolutional Neural Networks
Nikiforos Pittaras, Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
MMM (1)4
2017 Gaze movement-driven random forests for query clustering in automatic video annotation
Stefanos Vrochidis, Ioannis Patras, Ioannis Kompatsiaris
Multim. Tools Appl.2
2016 Unsupervised convolutional neural networks for motion estimation
abstract
Traditional methods for motion estimation estimate the motion field F between a pair of images as the one that minimizes a predesigned cost function. In this paper, we propose a direct method and train a Convolutional Neural Network (CNN) that when, at test time, is given a pair of images as input it produces a dense motion field F at its output layer. In the absence of large datasets with ground truth motion that would allow classical supervised training, we propose to train the network in an unsupervised manner. The proposed cost function that is optimized during training, is based on the classical optical flow constraint. The latter is differentiable with respect to the motion field and, therefore, allows backpropagation of the error to previous layers of the network. Our method is tested on both synthetic and real image sequences and performs similarly to the state-of-the-art methods.
Aria Ahmadi, Ioannis Patras
ICIP2
2016 Online multi-task learning for semantic concept detection in video
abstract
In this paper we propose an online multi-task learning algorithm for video concept detection. In particular, we extend the Efficient Lifelong Learning Algorithm (ELLA) in the following ways: (a) we solve the objective function of ELLA using quadratic programming instead of solving the Lasso problem, (b) we add a new label-based constraint that considers concept correlations, (c) we use linear SVMs as base learners instead of logistic regression. Experimental results show improvement over both the single-task learning methods typically used in this problem and the original ELLA algorithm.
Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
ICIP3
2016 Video aesthetic quality assessment using kernel Support Vector Machine with isotropic Gaussian sample uncertainty (KSVM-IGSU)
abstract
In this paper we propose a video aesthetic quality assessment method that combines the representation of each video according to a set of photographic and cinematographic rules, with the use of a learning method that takes the video representation's uncertainty into consideration. Specifically, our method exploits the information derived from both low- and high-level analysis of video layout, leading to a photo- and motion-based video representation scheme. Subsequently, a kernel Support Vector Machine (SVM) extension, the KSVM-iGSU, is trained to classify the videos and retrieve those of high aesthetic value. Experimental results on our large dataset verify the effectiveness of the proposed method. We also make publicly available our dataset, in order to facilitate research in the area of video aesthetic quality assessment.
Christos Tzelepis, Eftichia Mavridaki, Vasileios Mezaris, Ioannis Patras
ICIP4
2016 Minimal filtered channel features for pedestrian detection
abstract
This paper addresses the problem of efficient pedestrian detection using features that are extracted by convolving feature channels with a very small number of filters. The method uses as feature channels low level features such as LUV colour and HOG, and trains a boosted decision forest on top of the learned features. The feature selection is guided by a greedy search or by an exhaustive search on a few number of scaled versions of simple horizontal, vertical and uniform filters. Extensive results on the challenging Caltech dataset show that with only 3 filters we obtain state-of-the-art results achieving a 18.3% miss rate (MR). Using optical flow as an additional input, we further improve the results and obtain a 15.5% MR.
Yoshiki Kuranuki, Ioannis Patras
ICPR2
2016 Deep Multi-task Learning with Label Correlation Constraint for Video Concept Detection
abstract
In this work we propose a method that integrates multi-task learning (MTL) and deep learning. Our method appends a MTL-like loss to a deep convolutional neural network, in order to learn the relations between tasks together at the same time, and also incorporates the label correlations between pairs of tasks. We apply the proposed method on a transfer learning scenario, where our objective is to fine-tune the parameters of a network that has been originally trained on a large-scale image dataset for concept detection, so that it be applied on a target video dataset and a corresponding new set of target concepts. We evaluate the proposed method for the video concept detection problem on the TRECVID 2013 Semantic Indexing dataset. Our results show that the proposed algorithm leads to better concept-based video annotation than existing state-of-the-art methods.
Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
ACM Multimedia3
2016 Alone versus In-a-group: A Comparative Analysis of Facial Affect Recognition
abstract
Automatic affect analysis and understanding has become a well established research area in the last two decades. Recent works have started moving from individual to group scenarios. However, little attention has been paid to comparing the affect expressed in individual and group settings. This paper presents a framework to investigate the differences in affect recognition models along arousal and valence dimensions in individual and group settings. We analyse how a model trained on data collected from an individual setting performs on test data collected from a group setting, and vice versa. A third model combining data from both individual and group settings is also investigated. A set of experiments is conducted to predict the affective states along both arousal and valence dimensions on two newly collected databases that contain sixteen participants watching affective movie stimuli in individual and group settings, respectively. The experimental results show that (1) the affect model trained with group data performs better on individual test data than the model trained with individual data tested on group data, indicating that facial behaviours expressed in a group setting capture more variation than in an individual setting; and (2) the combined model does not show better performance than the affect model trained with a specific type of data (i.e., individual or group), but proves a good compromise. These results indicate that in settings where multiple affect models trained with different types of data are not available, using the affect model trained with group data is a viable solution.
Wenxuan Mou, Hatice Gunes, Ioannis Patras
ACM Multimedia3
2016 Ordering of Visual Descriptors in a Classifier Cascade Towards Improved Video Concept Detection
Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
MMM (1)3
2016 VERGE: A Multimodal Interactive Search Engine for Video Browsing and Retrieval
Anastasia Moumtzidou, Theodoros Mironidis, Evlampios Apostolidis, Fotini Markatopoulou, Anastasia Ioannidou, Ilias Gialampoukidis, Konstantinos Avgerinakis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris, Ioannis Patras
MMM (2)11
2016 Video Event Detection Using Kernel Support Vector Machine with Isotropic Gaussian Sample Uncertainty (KSVM-iGSU)
Christos Tzelepis, Vasileios Mezaris, Ioannis Patras
MMM (1)3
2016 Special Issue on Individual and Group Activities in Video Event Analysis
Liang Wang 0001, Ioannis Patras, Jian Zhang 0002, Greg Mori, Larry Davis 0001
Comput. Vis. Image Underst.2
2016 Action recognition using saliency learned from recorded human gaze
Daria Stefic, Ioannis Patras
Image Vis. Comput.2
2016 Learning to detect video events from zero or very few video examples
Christos Tzelepis, Damianos Galanopoulos, Vasileios Mezaris, Ioannis Patras
Image Vis. Comput.4
2015 Concept Detection in Multimedia Web Resources About Home Made Explosives
abstract
This work investigates the effectiveness of a state-of-the-art concept detection framework for the automatic classification of multimedia content, namely images and videos, embedded in publicly available Web resources containing recipes for the synthesis of Home Made Explosives (HMEs), to a set of predefined semantic concepts relevant to the HME domain. The concept detection framework employs advanced methods for video (shot) segmentation, visual feature extraction (using SIFT, SURF, and their variations), and classification based on machine learning techniques (logistic regression). The evaluation experiments are performed using an annotated collection of multimedia HME content discovered on the Web, and a set of concepts, which emerged both from an empirical study, and were also provided by domain experts and interested stakeholders, including Law Enforcement Agencies personnel. The experiments demonstrate the satisfactory performance of our framework, which in turn indicates the significant potential of the adopted approaches on the HME domain.
George Kalpakis, Theodora Tsikrika, Fotini Markatopoulou, Nikiforos Pittaras, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Patras, Ioannis Kompatsiaris
ARES7
2015 Identifying valence and arousal levels via connectivity between EEG channels
abstract
Implicit emotion tagging is a central theme in the area of affective computing. To this end, Several physiological signals acquired from subjects can be employed, for example, electroencephalography (EEG) and functional magnetic resonance imaging (fMRI) from brain, electrocardiography (ECG) from cardiac activities, and other peripheral physiological signals, such as galvanic skin resistance, electromyogram (EMG), blood volume pressure etc. Brain is regarded as the place where emotional activities evoke. Determining affective states by observing brain activities directly is of therefore great interest. There are several published works that use EEG signals to identify affective states in different aspects with various stimuli, e.s., images, musics and videos. In this paper, we propose to adopt EEG connectivity between electrodes to identify subjects' affective levels in both valence and arousal space during video stimuli presentation. Three catagories of connectivity are adopted in magnitude and phase domains. One open accessed affective database, DEAP, is used as benchmark. We will show that with the proposed connectivity-based representation, the accuracy of affective levels identification tasks are higher than the same tasks in existing works based on same database.
Junwei Han 0001, Lei Guo 0002, Ioannis Patras
ACII5
2015 Face Alignment Assisted by Head Pose Estimation
abstract
In this paper we propose a supervised initialization scheme for cascaded face alignment based on explicit head pose estimation. We first investigate the failure cases of most state of the art face alignment approaches and observe that these failures often share one common global property, i.e. the head pose variation is usually large. Inspired by this, we propose a deep convolutional network model for reliable and accurate head pose estimation. Instead of using a mean face shape, or randomly selected shapes for cascaded face alignment initialisation, we propose two schemes for generating initialisation: the first one relies on projecting a mean 3D face shape (represented by 3D facial landmarks) onto 2D image under the estimated head pose; the second one searches nearest neighbour shapes from the training set according to head pose distance. By doing so, the initialisation gets closer to the actual shape, which enhances the possibility of convergence and in turn improves the face alignment performance. We demonstrate the proposed method on the benchmark 300W dataset and show very competitive performance in both head pose estimation and face alignment.
Heng Yang 0001, Wenxuan Mou, Ioannis Patras, Hatice Gunes, Peter Robinson 0001
BMVC4
2015 Mirror, mirror on the wall, tell me, is the error small?
abstract
Do object part localization methods produce bilaterally symmetric results on mirror images? Surprisingly not, even though state of the art methods augment the training set with mirrored images. In this paper we take a closer look into this issue. We first introduce the concept of mirrorability as the ability of a model to produce symmetric results in mirrored images and introduce a corresponding measure, namely the mirror error that is defined as the difference between the detection result on an image and the mirror of the detection result on its mirror image. We evaluate the mirrorability of several state of the art algorithms in two of the most intensively studied problems, namely human pose estimation and face alignment. Our experiments lead to several interesting findings: 1) Most of state of the art methods struggle to preserve the mirror symmetry, despite the fact that they do have very similar overall performance on the original and mirror images; 2) the low mirrorability is not caused by training or testing sample bias - all algorithms are trained on both the original images and their mirrored versions; 3) the mirror error is strongly correlated to the localization/alignment error (with correlation coefficients around 0.7). Since the mirror error is calculated without knowledge of the ground truth, we show two interesting applications - in the first it is used to guide the selection of difficult samples and in the second to give feedback in a popular Cascaded Pose Regression method for face alignment.
Heng Yang 0001, Ioannis Patras
CVPR2
2015 Cascade of classifiers based on binary, non-binary and deep convolutional network descriptors for video concept detection
abstract
In this paper we propose a cascade architecture that can be used to train and combine different visual descriptors (local binary, local non-binary and Deep Convolutional Neural Network-based) for video concept detection. The proposed architecture is computationally more efficient than typical state-of-the-art video concept detection systems, without affecting the detection accuracy. In addition, this work presents a detailed study on combining descriptors based on Deep Convolutional Neural Networks with other popular local descriptors, both within a cascade and when using different late-fusion schemes. We evaluate our methods on the extensive video dataset of the 2013 TRECVID Semantic Indexing Task.
Fotini Markatopoulou, Vasileios Mezaris, Ioannis Patras
ICIP3
2015 A Study on the Use of a Binary Local Descriptor and Color Extensions of Local Descriptors for Video Concept Detection
Fotini Markatopoulou, Nikiforos Pittaras, Olga Papadopoulou, Vasileios Mezaris, Ioannis Patras
MMM (1)5
2015 VERGE: A Multimodal Interactive Video Search Engine
Anastasia Moumtzidou, Konstantinos Avgerinakis, Evlampios Apostolidis, Fotini Markatopoulou, Konstantinos Apostolidis, Theodoros Mironidis, Stefanos Vrochidis, Vasileios Mezaris, Ioannis Kompatsiaris, Ioannis Patras
MMM (2)10
2015 Cascade of forests for face alignment
abstract
In this study, we propose a regression forests‐based cascaded method for face alignment. We build on the cascaded pose regression (CPR) framework and propose to use the regression forest as a primitive regressor. The regression forests are easier to train and naturally handle the over‐fitting problem via averaging the outputs of the trees at each stage. We address the fact that the CPR approaches are sensitive to the shape initialisation; in contrast to using a number of blind initialisations and selecting the median values, we propose an intelligent shape initialisation scheme. More specifically, a large number of initialisations are propagated to a few early stages in the cascade, then only a proportion of them are propagated to the remaining cascades according to their convergence measurement. We evaluate the performance of the proposed approach on the challenging face alignment in the wild database and obtain superior or comparable performance with the state‐of‐the‐art, in spite of the fact that we have utilised only the freely available public training images. More importantly, we show that the intelligent initialisation scheme makes the CPR framework more robust to unreliable initialisations that are typically produced by different face detections.
Heng Yang 0001, Changqing Zou, Ioannis Patras
IET Comput. Vis.3
2015 Random Subspace Supervised Descent Method for Regression Problems in Computer Vision
abstract
Supervised Descent Method (SDM) has shown good performance in solving non-linear least squares problems in computer vision, giving state of the art results for the problem of face alignment. However, when SDM learns the generic descent maps, it is very difficult to avoid over-fitting due to the high dimensionality of the input features. In this paper we propose a Random Subspace SDM (RSSDM) that maintains the high accuracy on the training data and improves the generalization accuracy. Instead of using all the features for descent learning at each iteration, we randomly select sub-sets of the features and learn an ensemble of descent maps in the corresponding subspaces, one in each subspace. Then, we average the ensemble of descents to calculate the update of the iteration. We test the proposed methods on two representative regression problems, namely, 3D pose estimation and face alignment and show that RSSDM consistently outperforms SDM in both tasks in terms of accuracy (e.g., RSSDM is able to localize 4% more landmarks at error level of 0.1 on the challenging iBug dataset). RSSDM also holds several useful generalization properties: 1) it is more effective when the number of training samples is small-with 3 Monte-Carlo permutations RSSDM can achieve similar performance to SDM with 9 Monte-Carlo permutations; 2) it is less sensitive to the changes of the strength of the regularization-when the regularization parameter is changed to 10 times larger, the mean error increases 9.0% for SDM vs. 3.4% for RSSDM.
Heng Yang 0001, Xuhui Jia, Ioannis Patras, Kwok-Ping Chan
IEEE Signal Process. Lett.3
2015 DECAF: MEG-Based Multimodal Database for Decoding Affective Physiological Responses
abstract
In this work, we present DECAF-a multimodal data set for decoding user physiological responses to affective multimedia content. Different from data sets such as DEAP [15] and MAHNOB-HCI [31], DECAF contains (1) brain signals acquired using the Magnetoencephalogram (MEG) sensor, which requires little physical contact with the user's scalp and consequently facilitates naturalistic affective response, and (2) explicit and implicit emotional responses of 30 participants to 40 one-minute music video segments used in [15] and 36 movie clips, thereby enabling comparisons between the EEG versus MEG modalities as well as movie versus music stimuli for affect recognition. In addition to MEG data, DECAF comprises synchronously recorded near-infra-red (NIR) facial videos, horizontal Electrooculogram (hEOG), Electrocardiogram (ECG), and trapezius-Electromyogram (tEMG) peripheral physiological responses. To demonstrate DECAF's utility, we present (i) a detailed analysis of the correlations between participants' self-assessments and their physiological responses and (ii) single-trial classification results for valence, arousal and dominance, with performance evaluation against existing data sets. DECAF also contains time-continuous emotion annotations for movie clips from seven users, which we use to demonstrate dynamic emotion prediction.
Mojtaba Khomami Abadi, Subramanian Ramanathan, Seyed Mostafa Kia, Paolo Avesani, Ioannis Patras, Nicu Sebe
IEEE Trans. Affect. Comput.5
2015 Privileged Information-Based Conditional Structured Output Regression Forest for Facial Point Detection
abstract
This paper introduces a regression method called Privileged Information (PI)-based Conditional Structured Output Regression Forest (RF) for facial point detection. To train RF more efficiently, the method utilizes both PI, that is side information that is available only during training, such as head pose or gender, and shape constraints on the location of the facial points. We propose selection of the test functions at some randomly chosen internal tree nodes according to the information gain calculated on the PI. In this way, the training patches that arrive at leaves tend to have low variance both in terms of their displacements in relation to the facial points and in terms of the PI. At each leaf node, we learn three models: first, a probabilistic model of the pdf of the PI; second, a probabilistic regression model for the locations of the facial points; and third, shape models that model the interdependencies of the locations of neighboring facial points in a predefined structure graph. The latter two are conditioned on the PI. During testing, the marginal probability of the PI is estimated and the facial point locations are estimated using the appropriate conditional regression and shape models. The proposed method is validated and compared with very recent methods, especially that use Regression Forests, on datasets recorded in controlled and uncontrolled environments, namely, the BioID, the Labeled Faces in the Wild, the Labeled Face Parts in the Wild, and the Annotated Facial Landmarks in the Wild.
Heng Yang 0001, Ioannis Patras
IEEE Trans. Circuits Syst. Video Technol.2
2015 Robust Face Alignment Under Occlusion via Regional Predictive Power Estimation
abstract
Face alignment has been well studied in recent years, however, when a face alignment model is applied on facial images with heavy partial occlusion, the performance deteriorates significantly. In this paper, instead of training an occlusion-aware model with visibility annotation, we address this issue via a model adaptation scheme that uses the result of a local regression forest (RF) voting method. In the proposed scheme, the consistency of the votes of the local RF in each of several oversegmented regions is used to determine the reliability of predicting the location of the facial landmarks. The latter is what we call regional predictive power (RPP). Subsequently, we adapt a holistic voting method (cascaded pose regression based on random ferns) by putting weights on the votes of each fern according to the RPP of the regions used in the fern tests. The proposed method shows superior performance over existing face alignment models in the most challenging data sets (COFW and 300-W). Moreover, it can also estimate with high accuracy (72.4% overlap ratio) which image areas belong to the face or nonface objects, on the heavily occluded images of the COFW data set, without explicit occlusion modeling.
Heng Yang 0001, Xuming He 0001, Xuhui Jia, Ioannis Patras
IEEE Trans. Image Process.4
2015 Fine-Tuning Regression Forests Votes for Object Alignment in the Wild
abstract
In this paper, we propose a object alignment method that detects the landmarks of an object in 2D images. In the regression forests (RFs) framework, observations (patches) that are extracted at several image locations cast votes for the localization of several landmarks. We propose to refine the votes before accumulating them into the Hough space, by sieving and/or aggregating. In order to filter out false positive votes, we pass them through several sieves, each associated with a discrete or continuous latent variable. The sieves filter out votes that are not consistent with the latent variable in question, something that implicitly enforces global constraints. In order to aggregate the votes when necessary, we adjusts on-the-fly a proximity threshold by applying a classifier on middle-level features extracted from voting maps for the object landmark in question. Moreover, our method is able to predict the unreliability of an individual object landmark. This information can be useful for subsequent object analysis like object recognition. Our contributions are validated for two object alignment tasks, face alignment and car alignment, on data sets with challenging images collected in the wild, i.e., the Labeled Face in the Wild, the Annotated Facial Landmarks in the Wild, and the street scene car data set. We show that with the proposed approach, and without explicitly introducing shape models, we obtain performance superior or close to the state of the art for both tasks.
Heng Yang 0001, Ioannis Patras
IEEE Trans. Image Process.2
2014 Structured Semi-supervised Forest for Facial Landmarks Localization with Face Mask Reasoning
Xuhui Jia, Heng Yang 0001, Kwok-Ping Chan, Ioannis Patras
BMVC4
2014 Learning visual saliency using topographic independent component analysis
abstract
Understanding the underlying mechanisms that drive human visual attention is a topic of immense interest. Most of the work is focused on extracting manually selected features that might resemble the human visual processing pathway and using a combination of those features to train a classifier that learns to predict where humans look. In contrast, we will learn the features using a generalization of Independent Component Analysis (ICA), namely the topographic Independent Component Analysis (tICA). We will show that those learned features in combination with linear SVM outperform the hand-crafted ones. In addition, we propose a novel optimization scheme, which jointly optimizes for linear SVM and tICA pooling weights and show that it further improves the results.
Daria Stefic, Ioannis Patras
ICIP2
2014 Multimodal random forest based tensor regression
abstract
This study presents a method, called random forest based tensor regression, for real‐time head pose estimation using both depth and intensity data. The method builds on random forests and proposes to train and use tensor regressors at each leaf node of the trees of the forest. The tensor regressors are trained using both intensity and depth data and their votes are fused. The proposed method is shown to outperform current state of the art approaches in terms of accuracy when applied to the publicly available Biwi Kinect head pose dataset.
Sertan Kaymak, Ioannis Patras
IET Comput. Vis.2
2014 Face Sketch Landmarks Localization in the Wild
abstract
In this letter, we propose a method for facial landmarks localization in face sketch images. As recent approaches and the corresponding datasets are designed for ordinary face photos, the performance of such models drop significantly when they are applied on face sketch images. We first propose a scheme to synthesize face sketches from face photos based on random-forests edge detection and local face region enhancement. Then we jointly train a Cascaded Pose Regression based method for facial landmarks localization for both face photos and sketches. We build an evaluation dataset, called Face Sketches in the Wild (FSW), with 450 face sketch images collected from the Internet and with the manual annotation of 68 facial landmark locations on each face sketch. The proposed multi-modality facial landmark localization method shows competitive performance on both face sketch images (the FSW dataset) and face photo images (the Labeled Face Parts in the Wild dataset), despite the fact that we do not use extra annotation of face sketches for model building.
Heng Yang 0001, Changqing Zou, Ioannis Patras
IEEE Signal Process. Lett.3
2013 Sieving Regression Forest Votes for Facial Feature Detection in the Wild
abstract
In this paper we propose a method for the localization of multiple facial features on challenging face images. In the regression forests (RF) framework, observations (patches) that are extracted at several image locations cast votes for the localization of several facial features. In order to filter out votes that are not relevant, we pass them through two types of sieves, that are organised in a cascade, and which enforce geometric constraints. The first sieve filters out votes that are not consistent with a hypothesis for the location of the face center. Several sieves of the second type, one associated with each individual facial point, filter out distant votes. We propose a method that adjusts on-the-fly the proximity threshold of each second type sieve by applying a classifier which, based on middle-level features extracted from voting maps for the facial feature in question, makes a sequence of decisions on whether the threshold should be reduced or not. We validate our proposed method on two challenging datasets with images collected from the Internet in which we obtain state of the art results without resorting to explicit facial shape models. We also show the benefits of our method for proximity threshold adjustment especially on 'difficult' face images.
Heng Yang 0001, Ioannis Patras
ICCV2
2013 Semi-supervised visual recognition with constrained graph regularized non negative matrix factorization
abstract
This paper proposes a semi-supervised nonnegative matrix factorization algorithm for face and gait recognition. The proposed algorithm imposes hard constraints on the labelled data points, such that the data points that belong to the same class are projected to the same lower dimensional point. In addition, it introduces a graph Laplacian regularization term that preserves the local geometry structure of the data by penalising large distances between the projections of points that are close in the original space. This results in a constrained optimization problem, that is solved using block coordinate descent with multiplicative update rules. Experimental results on several publicly available datasets demonstrate that proposed method performs in par or considerably better than state of the art methods.
Weiwei Guo, Weidong Hu, Nikolaos V. Boulgouris, Ioannis Patras
ICIP4
2013 Fusion of facial expressions and EEG for implicit affective tagging
Sander Koelstra, Ioannis Patras
Image Vis. Comput.2
2013 Coupled Gaussian Processes for Pose-Invariant Facial Expression Recognition
abstract
We propose a method for head-pose invariant facial expression recognition that is based on a set of characteristic facial points. To achieve head-pose invariance, we propose the Coupled Scaled Gaussian Process Regression (CSGPR) model for head-pose normalization. In this model, we first learn independently the mappings between the facial points in each pair of (discrete) nonfrontal poses and the frontal pose, and then perform their coupling in order to capture dependences between them. During inference, the outputs of the coupled functions from different poses are combined using a gating function, devised based on the head-pose estimation for the query points. The proposed model outperforms state-of-the-art regression-based approaches to head-pose normalization, 2D and 3D Point Distribution Models (PDMs), and Active Appearance Models (AAMs), especially in cases of unknown poses and imbalanced training data. To the best of our knowledge, the proposed method is the first one that is able to deal with expressive faces in the range from -45° to +45° pan rotation and -30° to +30° tilt rotation, and with continuous changes in head pose, despite the fact that training was conducted on a small set of discrete poses. We evaluate the proposed method on synthetic and real images depicting acted and spontaneously displayed facial expressions.
Ognjen Rudovic, Maja Pantic, Ioannis Patras
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 High order pLSA for indexing tagged images
Spiros Nikolopoulos, Stefanos Zafeiriou, Ioannis Patras, Ioannis Kompatsiaris
Signal Process.3
2012 Exploring the Similarities of Neighboring Spatiotemporal Points for Action Pair Matching
Irene Kotsia, Ioannis Patras
ACCV (3)2
2012 Face Parts Localization Using Structured-Output Regression Forests
Heng Yang 0001, Ioannis Patras
ACCV (2)2
2012 Support tensor action spotting
abstract
In this paper we address the action spotting problem, that is the spatiotemporal detection and localization of an action. We first calculate a novel objective function between an input video sequence and an action's weights tensor, as acquired from a Support Tensor Machine classifier. We subsequently search for an appropriate transformation that maximizes the objective function, calculated as the multiplication of the original input tensor with the weights tensor. The proposed algorithm is very fast, as the above mentioned multiplication involves a set of a separable filters applied along each mode. We demonstrate the effectiveness of our method with experiments in a publicly available database where we show that our method outperforms existing techniques in terms of spatiotemporal action localization.
Irene Kotsia, Ioannis Patras
ICIP2
2012 A simple and effective extrinsic calibration method of a camera and a single line scanning lidar
Heng Yang 0001, Ioannis Patras
ICPR3
2012 Coupled 3D tracking and pose optimization of rigid objects using particle filter
Heng Yang 0001, Yueqiang Zhang, Ioannis Patras
ICPR4
2012 Max-margin Non-negative Matrix Factorization
Irene Kotsia, Ioannis Patras
Image Vis. Comput.3
2012 Leveraging social media for scalable object detection
Elisavet Chatzilari, Spiros Nikolopoulos, Ioannis Patras, Ioannis Kompatsiaris
Pattern Recognit.3
2012 Higher rank Support Tensor Machines for visual recognition
Irene Kotsia, Weiwei Guo, Ioannis Patras
Pattern Recognit.3
2012 DEAP: A Database for Emotion Analysis ;Using Physiological Signals
abstract
We present a multimodal data set for the analysis of human affective states. The electroencephalogram (EEG) and peripheral physiological signals of 32 participants were recorded as each watched 40 one-minute long excerpts of music videos. Participants rated each video in terms of the levels of arousal, valence, like/dislike, dominance, and familiarity. For 22 of the 32 participants, frontal face video was also recorded. A novel method for stimuli selection is proposed using retrieval by affective tags from the last.fm website, video highlight detection, and an online assessment tool. An extensive analysis of the participants' ratings during the experiment is presented. Correlates between the EEG signal frequencies and the participants' ratings are investigated. Methods and results are presented for single-trial classification of arousal, valence, and like/dislike ratings using the modalities of EEG, peripheral physiological signals, and multimedia content analysis. Finally, decision fusion of the classification results from different modalities is performed. The data set is made publicly available and we encourage other researchers to use it for testing their own affective state estimation methods.
Sander Koelstra, Christian Mühl, Mohammad Soleymani 0001, Jong-Seok Lee, Ashkan Yazdani, Touradj Ebrahimi, Thierry Pun, Anton Nijholt, Ioannis Patras
IEEE Trans. Affect. Comput.9
2012 Tensor Learning for Regression
abstract
In this paper, we exploit the advantages of tensorial representations and propose several tensor learning models for regression. The model is based on the canonical/parallel-factor decomposition of tensors of multiple modes and allows the simultaneous projections of an input tensor to more than one direction along each mode. Two empirical risk functions are studied, namely, the square loss and ε -insensitive loss functions. The former leads to higher rank tensor ridge regression (TRR), and the latter leads to higher rank support tensor regression (STR), both formulated using the Frobenius norm for regularization. We also use the group-sparsity norm for regularization, favoring in that way the low rank decomposition of the tensorial weight. In that way, we achieve the automatic selection of the rank during the learning process and obtain the optimal-rank TRR and STR. Experiments conducted for the problems of head-pose, human-age, and 3-D body-pose estimations using real data from publicly available databases, verified not only the superiority of tensors over their vector counterparts but also the efficiency of the proposed algorithms.
Weiwei Guo, Irene Kotsia, Ioannis Patras
IEEE Trans. Image Process.3
2012 Tree-Structured Feature Extraction Using Mutual Information
abstract
One of the most informative measures for feature extraction (FE) is mutual information (MI). In terms of MI, the optimal FE creates new features that jointly have the largest dependency on the target class. However, obtaining an accurate estimate of a high-dimensional MI as well as optimizing with respect to it is not always easy, especially when only small training sets are available. In this paper, we propose an efficient tree-based method for FE in which at each step a new feature is created by selecting and linearly combining two features such that the MI between the new feature and the class is maximized. Both the selection of the features to be combined and the estimation of the coefficients of the linear transform rely on estimating 2-D MIs. The estimation of the latter is computationally very efficient and robust. The effectiveness of our method is evaluated on several real-world data sets. The results show that the classification accuracy obtained by the proposed method is higher than that achieved by other FE methods.
Farid Oveisi, Shahrzad Oveisi, Abbas Erfanian, Ioannis Patras
IEEE Trans. Neural Networks Learn. Syst.4
2011 Max-Margin Semi-NMF
abstract
In this paper, we propose a maximum-margin framework for classification using Non-negative Matrix Factorization. In contrast to previous approaches where the classification and matrix factorization are separated, we incorporate the maximum margin constraints within the NMF formuation i.e. we solve for a base matrix that maximizes the margin of the classifier in the low dimensional feature space. This results in a non-convex constrained optimization problem with respect to the bases, the projection coefficients and the separating hyperplane, which we propose to solve in an iterative way, solving at each iteration a set of convex sub-problems with respect to subsets of the unknown variables. The resulting basis matrix is used to extract features that maximize the margin of the resulting classifier. The performance of the proposed algorithm is evaluated on several publicly available datasets where it is shown to consistently outperform Discriminative NMF and SVM classifiers that use features extracted by semi-NMF.
Ioannis Patras, Irene Kotsia
BMVC2
2011 Support tucker machines
abstract
In this paper we address the two-class classification problem within the tensor-based framework, by formulating the Support Tucker Machines (STuMs). More precisely, in the proposed STuMs the weights parameters are regarded to be a tensor, calculated according to the Tucker tensor decomposition as the multiplication of a core tensor with a set of matrices, one along each mode. We further extend the proposed STuMs to the Σ/ΣwSTuMs, in order to fully exploit the information offered by the total or the within-class covariance matrix and whiten the data, thus providing in-variance to affine transformations in the feature space. We formulate the two above mentioned problems in such a way that they can be solved in an iterative manner, where at each iteration the parameters corresponding to the projections along a single tensor mode are estimated by solving a typical Support Vector Machine-type problem. The superiority of the proposed methods in terms of classification accuracy is illustrated on the problems of gait and action recognition.
Irene Kotsia, Ioannis Patras
CVPR2
2011 Continuous emotion detection in response to music videos
abstract
Viewers' preference for multimedia selection depends highly on their emotional experience. In this paper, we present an emotion detection method for music videos using central and peripheral nervous system physiological signals as well as multimedia content analysis. A set of 40 music clips eliciting a broad range of emotions were first selected. After extracting the one minute long emotional highlight of each video, they were shown to 32 participants while their physiological responses were recorded. Participants self-reported their felt emotions after watching each clip by means of arousal, valence, dominance, and liking ratings. The physiological signals included electroencephalogram, galvanic skin response, respiration pattern, skin temperature, electromyograms and blood volume pulse using plethysmograph. Emotional features were extracted from the signals and the multimedia content. The emotional features were used to train a linear ridge regressor to detect emotions for each participant using a leave-one-out cross-validation strategy. The performance of the personalized emotion detection is shown to be significantly superior to a random regressor.
Mohammad Soleymani 0001, Sander Koelstra, Ioannis Patras, Thierry Pun
FG3
2011 An eye-tracking-based approach to facilitate interactive video search
abstract
This paper investigates the role of gaze movements as implicit user feedback during interactive video retrieval tasks. In this context, we use a content-based video search engine to perform an interactive video retrieval experiment, during which, we record the user gaze movements with the aid of an eye-tracking device and generate features for each video shot based on aggregated past user eye fixation and pupil dilation data. Then, we employ support vector machines, in order to train a classifier that could identify shots marked as relevant to a new query topic submitted by new users. The positive results provided by the classifier are used as recommendations for future users, who search for similar topics. The evaluation shows that important information can be extracted from aggregated gaze movements during video retrieval tasks, while the involvement of pupil dilation data improves the performance of the system and facilitates interactive video search.
Stefanos Vrochidis, Ioannis Patras, Ioannis Kompatsiaris
ICMR2
2011 Spatiotemporal Localization and Categorization of Human Actions in Unsegmented Image Sequences
abstract
In this paper we address the problem of localization and recognition of human activities in unsegmented image sequences. The main contribution of the proposed method is the use of an implicit representation of the spatiotemporal shape of the activity which relies on the spatiotemporal localization of characteristic ensembles of feature descriptors. Evidence for the spatiotemporal localization of the activity is accumulated in a probabilistic spatiotemporal voting scheme. The local nature of the proposed voting framework allows us to deal with multiple activities taking place in the same scene, as well as with activities in the presence of clutter and occlusion. We use boosting in order to select characteristic ensembles per class. This leads to a set of class specific codebooks where each codeword is an ensemble of features. During training, we store the spatial positions of the codeword ensembles with respect to a set of reference points, as well as their temporal positions with respect to the start and end of the action instance. During testing, each activated codeword ensemble casts votes concerning the spatiotemporal position and extend of the action, using the information that was stored during training. Mean Shift mode estimation in the voting space provides the most probable hypotheses concerning the localization of the subjects at each frame, as well as the extend of the activities depicted in the image sequences. We present classification and localization results for a number of publicly available datasets, and for a number of sequences where there is a significant amount of clutter and occlusion.
Antonios Oikonomopoulos, Ioannis Patras, Maja Pantic
IEEE Trans. Image Process.2
2011 Evidence-Driven Image Interpretation by Combining Implicit and Explicit Knowledge in a Bayesian Network
abstract
Computer vision techniques have made considerable progress in recognizing object categories by learning models that normally rely on a set of discriminative features. However, in contrast to human perception that makes extensive use of logic-based rules, these models fail to benefit from knowledge that is explicitly provided. In this paper, we propose a framework that can perform knowledge-assisted analysis of visual content. We use ontologies to model the domain knowledge and a set of conditional probabilities to model the application context. Then, a Bayesian network is used for integrating statistical and explicit knowledge and performing hypothesis testing using evidence-driven probabilistic inference. In addition, we propose the use of a focus-of-attention (FoA) mechanism that is based on the mutual information between concepts. This mechanism selects the most prominent hypotheses to be verified/tested by the BN, hence removing the need to exhaustively test all possible combinations of the hypotheses set. We experimentally evaluate our framework using content from three domains and for the following three tasks: 1) image categorization; 2) localized region labeling; and 3) weak annotation of video shot keyframes. The results obtained demonstrate the improvement in performance compared to a set of baseline concept classifiers that are not aware of any context or domain knowledge. Finally, we also demonstrate the ability of the proposed FoA mechanism to significantly reduce the computational cost of visual inference while obtaining results comparable to the exhaustive case.
Spiros Nikolopoulos, Georgios Th. Papadopoulos, Ioannis Kompatsiaris, Ioannis Patras
IEEE Trans. Syst. Man Cybern. Part B4
2010 Coupled Gaussian Process Regression for Pose-Invariant Facial Expression Recognition
Ognjen Rudovic, Ioannis Patras, Maja Pantic
ECCV (2)2
2010 Multiplicative Update Rules for Multilinear Support Tensor Machines
abstract
In this paper, we formulate the Multilinear Support Tensor Machines (MSTMs) problem in a similar to the Non-negative Matrix Factorization (NMF) algorithm way. A novel set of simple and robust multiplicative update rules are proposed in order to find the multilinear classifier. Updates rules are provided for both hard and soft margin MSTMs and the existence of a bias term is also investigated. We present results on standard gait and action datasets and report faster convergence of equivalent classification performance in comparison to standard MSTMs.
Irene Kotsia, Ioannis Patras
ICPR2
2010 Pyramidal Model for Image Semantic Segmentation
abstract
We present a new hierarchical model applied to the problem of image semantic segmentation, that is, the association to each pixel in an image with a category label (e.g. tree, cow, building, ...). This problem is usually addressed with a combination of an appearance-based pixel classification and a pixel context model. In our proposal, the images are initially over-segmented in dense patches. The proposed pyramidal model naturally embeds the compositional nature of a scene to achieve a multi-scale contextualisation of patches. This is obtained by imposing an order on the patches aggregation operations towards the final scene. The nodes of the pyramid (that is, a dendrogram) thus represent patch clusters, or super-patches. The probabilistic model favours the homogeneous labelling of super-patches that are likely to contain a single object instance, modelling the uncertainty in identifying such super-patches. The proposed model has several advantages, including the computational efficiency, as well as the expandability. Initial results place the model in line with other works in the recent literature.
Giuseppe Passino, Ioannis Patras, Ebroul Izquierdo
ICPR2
2010 Regression-Based Multi-view Facial Expression Recognition
abstract
We present a regression-based scheme for multi-view facial expression recognition based on 2D geometric features. We address the problem by mapping facial points (e.g. mouth corners) from non-frontal to frontal view where further recognition of the expressions can be performed using a state-of-the-art facial expression recognition method. To learn the mapping functions we investigate four regression models: Linear Regression (LR), Support Vector Regression (SVR), Relevance Vector Regression (RVR) and Gaussian Process Regression (GPR). Our extensive experiments on the CMU Multi-PIE facial expression database show that the proposed scheme outperforms view-specific classifiers by utilizing considerably less training data.
Ognjen Rudovic, Ioannis Patras, Maja Pantic
ICPR2
2010 A Dynamic Texture-Based Approach to Recognition of Facial Actions and Their Temporal Models
abstract
In this work, we propose a dynamic texture-based approach to the recognition of facial Action Units (AUs, atomic facial gestures) and their temporal models (i.e., sequences of temporal segments: neutral, onset, apex, and offset) in near-frontal-view face videos. Two approaches to modeling the dynamics and the appearance in the face region of an input video are compared: an extended version of Motion History Images and a novel method based on Nonrigid Registration using Free-Form Deformations (FFDs). The extracted motion representation is used to derive motion orientation histogram descriptors in both the spatial and temporal domain. Per AU, a combination of discriminative, frame-based GentleBoost ensemble learners and dynamic, generative Hidden Markov Models detects the presence of the AU in question and its temporal segments in an input image sequence. When tested for recognition of all 27 lower and upper face AUs, occurring alone or in combination in 264 sequences from the MMI facial expression database, the proposed method achieved an average event recognition accuracy of 89.2 percent for the MHI method and 94.3 percent for the FFD method. The generalization performance of the FFD method has been tested using the Cohn-Kanade database. Finally, we also explored the performance on spontaneous expressions in the Sensitive Artificial Listener data set.
Sander Koelstra, Maja Pantic, Ioannis Patras
IEEE Trans. Pattern Anal. Mach. Intell.3
2010 Coupled Prediction Classification for Robust Visual Tracking
abstract
This paper addresses the problem of robust template tracking in image sequences. Our work falls within the discriminative framework in which the observations at each frame yield direct probabilistic predictions of the state of the target. Our primary contribution is that we explicitly address the problem that the prediction accuracy for different observations varies, and in some cases, can be very low. To this end, we couple the predictor to a probabilistic classifier which, when trained, can determine the probability that a new observation can accurately predict the state of the target (that is, determine the "relevance" or "reliability" of the observation in question). In the particle filtering framework, we derive a recursive scheme for maintaining an approximation of the posterior probability of the state in which multiple observations can be used and their predictions moderated by their corresponding relevance. In this way, the predictions of the "relevant" observations are emphasized, while the predictions of the "irrelevant" observations are suppressed. We apply the algorithm to the problem of 2D template tracking and demonstrate that the proposed scheme outperforms classical methods for discriminative tracking both in the case of motions which are large in magnitude and also for partial occlusions.
Ioannis Patras, Edwin R. Hancock
IEEE Trans. Pattern Anal. Mach. Intell.1
2009 Latent Semantics Local Distribution for CRF-based Image Semantic Segmentation
abstract
Semantic image segmentation is the task of assigning a semantic label to every pixel of an image. This task is posed as a supervised learning problem in which the appearance of areas that correspond to a number of semantic categories are learned from a dataset of manually labelled images. This paper proposes a method that combines a region-based probabilistic graphical model that builds on the recent success of Conditional Random Fields (CRFs) in the problem of semantic segmentation, with a salient-points-based bagsof-words paradigm. In a first stage, the image is oversegmented into patches. Then, in a CRF-based formulation we learn both the appearance for each semantic category and the neighbouring relations between patches. In addition to patch features, we also consider information extracted on salient points that are detected in the patch’s vicinity. A visual word is associated to each salient point. Two different types of information are used. First, we consider the local weighted distribution of visual words. Using local (i.e. centred at each patch) word histograms enriches the classical global bags-of-word representation with positional information on word distributions. Second, we consider the un-normalised local distribution of a set of latent topics that are obtained by probabilistic Latent Semantic Analysis (pLSA). This distribution is obtained by the weighted accumulation of the latent topic distributions that are associated to the visual words in the area. The advantage of this second approach lays in the separate representation of the semantic content for each visual word. This allows us to consider the word contributions as independent in the CRF formulation without introducing too strong simplification assumptions. Tests on a publicly available dataset demonstrate the validity of the proposed salient point integration strategies. The results obtained with different configurations show an advance compared to other leading works in the area.
Giuseppe Passino, Ioannis Patras, Ebroul Izquierdo
BMVC2
2009 Sparse B-spline polynomial descriptors for human activity recognition
Antonios Oikonomopoulos, Maja Pantic, Ioannis Patras
Image Vis. Comput.3
2008 Incremental salient point detection
abstract
In this paper, we investigate an approach that computes salient points, i.e. areas of natural images that contain corners or edges, incrementally. We focus on the popular Harris corner detector and demonstrate how such an approach can operate when the image samples are refined in a bitwise manner, i.e. the image bitplanes are received one-by-one from the image sensor. This has the advantage that the image sensing and the salient point detection can be terminated at any input image precision (e.g. at a bound set by the sensory equipment or by computation, or by the salient point accuracy required by the application) and the obtained salient points under this precision are readily available. We estimate the required energy for image sensing as well as the computation required for the salient point detection and compare them against the conventional salient point detector realization that operates directly on each source precision and cannot refine the result. Our experiments demonstrate the feasibility of incremental approaches for salient point detection in various classes of natural images. In addition, a first comparison between the results obtained by the intermediate detectors is presented.
Ioannis Patras, Yiannis Andreopoulos
ICASSP1
2008 On the role of structure in part-based object detection
abstract
Part-based approaches in image analysis aim at exploiting the considerable discriminative power embedded in relations among image parts. Nonetheless, learning structural information is not always possible without the availability of a training set of classified parts, and taking into account this additional information can even degrade the performance of the system. In this paper, a discriminative graphical model for object detection is introduced and used in order to analyse and report results on the role of structural information in image classification tasks.
Giuseppe Passino, Ioannis Patras, Ebroul Izquierdo
ICIP2
2008 Incremental Refinement of Image Salient-Point Detection
abstract
Low-level image analysis systems typically detect "points of interest", i.e., areas of natural images that contain corners or edges. Most of the robust and computationally efficient detectors proposed for this task use the autocorrelation matrix of the localized image derivatives. Although the performance of such detectors and their suitability for particular applications has been studied in relevant literature, their behavior under limited input source (image) precision or limited computational or energy resources is largely unknown. All existing frameworks assume that the input image is readily available for processing and that sufficient computational and energy resources exist for the completion of the result. Nevertheless, recent advances in incremental image sensors or compressed sensing, as well as the demand for low-complexity scene analysis in sensor networks now challenge these assumptions. In this paper, we investigate an approach to compute salient points of images incrementally, i.e., the salient point detector can operate with a coarsely quantized input image representation and successively refine the result (the derived salient points) as the image precision is successively refined by the sensor. This has the advantage that the image sensing and the salient point detection can be terminated at any input image precision (e.g., bound set by the sensory equipment or by computation, or by the salient point accuracy required by the application) and the obtained salient points under this precision are readily available. We focus on the popular detector proposed by Harris and Stephens and demonstrate how such an approach can operate when the image samples are refined in a bitwise manner, i.e., the image bitplanes are received one-by-one from the image sensor. We estimate the required energy for image sensing as well as the computation required for the salient point detection based on stochastic source modeling. The computation and energy required by the proposed incremental refinement approach is compared against the conventional salient-point detector realization that operates directly on each source precision and cannot refine the result. Our experiments demonstrate the feasibility of incremental approaches for salient point detection in various classes of natural images. In addition, a first comparison between the results obtained by the intermediate detectors is presented and a novel application for adaptive low-energy image sensing based on points of saliency is presented.
Yiannis Andreopoulos, Ioannis Patras
IEEE Trans. Image Process.2
2007 Regression tracking with data relevance determination
abstract
This paper addresses the problem of efficient visual 2D template tracking in image sequences. We adopt a discriminative approach in which the observations at each frame yield direct predictions of a parametrisation of the state (e.g. position/scale/rotation) of the tracked target. To this end, a Bayesian mixture of experts (BME) is trained on a dataset of image patches that are generated by applying artificial transformations to the template at the first frame. In contrast to other methods in the literature, we explicitly address the problem that the prediction accuracy can deteriorate drastically for observations that are not similar to the ones in the training set; such observations are common in case of partial occlusions or of fast motion. To do so, we couple the BME with a probabilistic kernel-based classifier which, when trained, can determine the probability that a new/unseen observation can accurately predict the state of the target (the 'relevance' of the observation in question). In addition, in the particle filtering framework, we derive a recursive scheme for maintaining an approximation of the posterior probability of the target's state in which the probabilistic predictions of multiple observations are moderated by their corresponding relevance. We apply the algorithm in the problem of 2D template tracking and demonstrate that the proposed scheme outperforms classical methods for discriminative tracking in case of motions large in magnitude and of partial occlusions.
Ioannis Patras, Edwin R. Hancock
CVPR1
2007 Template Trackingwith Observation Relevance Determination
abstract
This paper addresses the problem of template tracking in the presence of occlusions, clutter and rapid motion. We adopt a learning approach, using a Bayesian Mixture of Experts (BME), in which observations at each frame yield direct predictions of the state (e.g. position / scale) of the tracked target. In contrast to other methods in the literature, we explicitly address the problem that the prediction accuracy can deteriorate drastically for observations that are not similar to the ones in the training set; such observations are common in case of partial occlusions or of fast motion. To do so, we couple the BME with a probabilistic kernel-based classifier which, when trained, can determine the probability that a new/unseen observation can accurately predict the state of the target (the 'relevance' of the observation in question). In addition, in the particle filtering framework, we derive a recursive scheme for maintaining an approximation of the posterior probability of the target's state in which the probabilistic predictions of multiple observations are moderated by their corresponding relevance. We apply the algorithm in the problem of 2D template tracking and demonstrate that the proposed scheme outperforms classical methods for discriminative tracking in case of motions large in magnitude and of partial occlusions.
Ioannis Patras, Edwin R. Hancock
ICIP (1)1
2007 Probabilistic Confidence Measures for Block Matching Motion Estimation
abstract
This paper addresses the problem of deriving measures that express the degree of the reliability of motion vectors estimated by a block-matching motion estimation method. We express the block matching motion estimation scheme in the probabilistic framework as a maximum-likelihood estimation scheme. Subsequently, we derive the confidence measures in terms of the a posteriori probabilities and the likelihoods of the estimated vectors. The assumptions about the type of the likelihood, that is, about the underlying conditional probability distribution of the motion compensated intensity differences (e.g., Laplacian) are derived from the objective criterion of the block-matching estimator. All parameters are estimated from data that are derived as by-product of the motion estimation scheme and our method, practically, introduces no additional computational cost. The derivation of the confidence measures is incorporated in a multiscale scheme. Experimental results are presented for image sequences with known ground-truth motion.
Ioannis Patras, Emile A. Hendriks, Reginald L. Lagendijk
IEEE Trans. Circuits Syst. Video Technol.1
2006 Combining color and shape information for illumination-viewpoint invariant object recognition
abstract
In this paper, we propose a new scheme that merges color- and shape-invariant information for object recognition. To obtain robustness against photometric changes, color-invariant derivatives are computed first. Color invariance is an important aspect of any object recognition scheme, as color changes considerably with the variation in illumination, object pose, and camera viewpoint. These color invariant derivatives are then used to obtain similarity invariant shape descriptors. Shape invariance is equally important as, under a change in camera viewpoint and object pose, the shape of a rigid object undergoes a perspective projection on the image plane. Then, the color and shape invariants are combined in a multidimensional color-shape context which is subsequently used as an index. As the indexing scheme makes use of a color-shape invariant context, it provides a high-discriminative information cue robust against varying imaging conditions. The matching function of the color-shape context allows for fast recognition, even in the presence of object occlusion and cluttering. From the experimental results, it is shown that the method recognizes rigid objects with high accuracy in 3-D complex scenes and is robust against changing illumination, camera viewpoint, object pose, and noise.
Aristeidis Diplaros, Theo Gevers, Ioannis Patras
IEEE Trans. Image Process.3
2006 Spatiotemporal salient points for visual recognition of human actions
abstract
This paper addresses the problem of human-action recognition by introducing a sparse representation of image sequences as a collection of spatiotemporal events that are localized at points that are salient both in space and time. The spatiotemporal salient points are detected by measuring the variations in the information content of pixel neighborhoods not only in space but also in time. An appropriate distance metric between two collections of spatiotemporal salient points is introduced, which is based on the chamfer distance and an iterative linear time-warping technique that deals with time expansion or time-compression issues. A classification scheme that is based on relevance vector machines and on the proposed distance measure is proposed. Results on real image sequences from a small database depicting people performing 19 aerobic exercises are presented.
Antonios Oikonomopoulos, Ioannis Patras, Maja Pantic
IEEE Trans. Syst. Man Cybern. Part B2
2006 Dynamics of Facial Expression: Recognition of Facial Actions and Their Temporal Segments From Face Profile Image Sequences
abstract
Automatic analysis of human facial expression is a challenging problem with many applications. Most of the existing automated systems for facial expression analysis attempt to recognize a few prototypic emotional expressions, such as anger and happiness. Instead of representing another approach to machine analysis of prototypic facial expressions of emotion, the method presented in this paper attempts to handle a large range of human facial behavior by recognizing facial muscle actions that produce expressions. Virtually all of the existing vision systems for facial muscle action detection deal only with frontal-view face images and cannot handle temporal dynamics of facial actions. In this paper, we present a system for automatic recognition of facial action units (AUs) and their temporal models from long, profile-view face image sequences. We exploit particle filtering to track 15 facial points in an input face-profile sequence, and we introduce facial-action-dynamics recognition from continuous video input using temporal rules. The algorithm performs both automatic segmentation of an input video into facial expressions pictured and recognition of temporal segments (i.e., onset, apex, offset) of 27 AUs occurring alone or in a combination in the input face-profile video. A recognition rate of 87% is achieved.
Maja Pantic, Ioannis Patras
IEEE Trans. Syst. Man Cybern. Part B2
2005 Spatiotemporal saliency for human action recognition
abstract
This paper addresses the problem of human action recognition by introducing a sparse representation of image sequences as a collection of spatiotemporal events that are localized at points that are salient both in space and time. We detect the spatiotemporal salient points by measuring changes in the information content of pixel neighborhoods not only in space but also in time. We introduce an appropriate distance metric between two collections of spatiotemporal salient points that is based on the Chamfer distance and an iterative linear time warping technique that deals with time expansion or time compression issues. We propose a classification scheme that is based on relevance vector machines and on the proposed distance measure. We present results on real image sequences from a small database depicting people performing 19 aerobic exercises.
Antonios Oikonomopoulos, Ioannis Patras, Maja Pantic
ICME2
2005 Detecting facial actions and their temporal segments in nearly frontal-view face image sequences
abstract
The recognition of facial expressions in image sequences is a difficult problem with many applications in human-machine interaction. Facial expression analyzers achieve good recognition rates, but virtually all of them deal only with prototypic facial expressions of emotions and cannot handle temporal dynamics of facial displays. The method presented here attempts to handle a large range of human facial behavior by recognizing facial action units (AUs) and their temporal segments (i.e., onset, apex, offset) that produce expressions. We exploit particle filtering to track 20 facial points in an input face video and we introduce AU-dynamics recognition using temporal rules. When tested on Cohn-Kanade and MMI facial expression databases, the proposed method achieved a recognition rate of 90% when detecting 27 AUs occurring alone or in a combination in an input face image sequence.
Maja Pantic, Ioannis Patras
SMC2
2005 Tracking deformable motion
abstract
This paper addresses the problem of template-based tracking of non rigid objects. We use the well-known framework of auxiliary particle filtering and propose an observation model that explicitly addresses appearance changes that are caused by local deformations of the tracked object. In addition, by adopting a colour difference that is invariant to local changes in the illumination, the proposed observation model can deal with changing lighting conditions and shadows. Experimental results with real image sequences demonstrate the efficiency of the proposed method in tracking facial features, such as mouth and eye corners
Ioannis Patras, Maja Pantic
SMC1
2004 Temporal modeling of facial actions from face profile image sequences
abstract
The recognition of facial action units (AUs) in image sequences is a challenging problem. AU detectors achieve good recognition rates, but virtually all of them deal only with frontal-view face images and cannot handle the temporal dynamics of AUs. We report on a system for automatic recognition of temporal models of AUs from long, profile-view, face image sequences. We exploit particle filtering to track 15 facial points in an input face-profile video sequence and we introduce facial-behavior temporal-dynamics recognition from continuous video input using temporal rules. The utilized algorithm performs both automatic segmentation and recognition of temporal segments (i.e., onset, apex, offset) of 23 AUs occurring alone or in a combination in an input face-profile video sequence. A recognition rate of 88% is achieved.
Maja Pantic, Ioannis Patras
ICME2
2004 Dense motion estimation using regularization constraints on local parametric models
abstract
This paper presents a method for dense optical flow estimation in which the motion field within patches that result from an initial intensity segmentation is parametrized with models of different order. We propose a novel formulation which introduces regularization constraints between the model parameters of neighboring patches. In this way, we provide the additional constraints for very small patches and for patches whose intensity variation cannot sufficiently constrain the estimation of their motion parameters. In order to preserve motion discontinuities, we use robust functions as a regularization mean. We adopt a three-frame approach and control the balance between the backward and forward constraints by a real-valued direction field on which regularization constraints are applied. An iterative deterministic relaxation method is employed in order to solve the corresponding optimization problem. Experimental results show that the proposed method deals successfully with motions large in magnitude, motion discontinuities, and produces accurate piecewise-smooth motion fields.
Ioannis Patras, Marcel Worring, Rein van den Boomgaard
IEEE Trans. Image Process.1
2003 Semi-automatic object-based video segmentation with labeling of color segments
Ioannis Patras, Emile A. Hendriks, Reginald L. Lagendijk
Signal Process. Image Commun.1
2002 Confidence measures for block matching motion estimation
abstract
This paper addresses the problem of deriving measures that express the degree of the reliability of motion vectors estimated by a block matching motion estimation method. We express the block matching motion estimation scheme in the probabilistic framework and derive the confidence measures in terms of the a-posteriori probabilities of the estimated vectors. The type of the underlying conditional probability distribution of the motion compensated intensity differences (e.g. Laplacian) is derived from the objective criterion of the block matching estimator. All parameters are estimated from data that are derived as a by-product of the motion estimation scheme. The derivation is incorporated in a multiscale scheme. Experimental results are presented for image sequences with known ground-truth motion.
Ioannis Patras, Emile A. Hendriks, Reginald L. Lagendijk
ICIP (2)1
2002 Facial action recognition in face profile image sequences
abstract
A robust way to discern facial gestures in images of faces, insensitive to scale, pose, and occlusion, is still the key research challenge in the automatic facial-expression analysis domain. A practical method recognized as the most promising one for addressing this problem is through a facial-gesture analysis of multiple views of the face. Yet, current systems for automatic facial-gesture analysis utilize mainly portrait or nearly frontal views of faces. To advance the existing technological framework upon which research on automatic facial-gesture analysis from multiple facial views can be based, we developed an automatic system as to analyze subtle changes in facial expressions based on profile-contour reference points in a profile-view video. A probabilistic classification method based on statistical modeling of the color and motion properties of the profile in the scene is proposed for tracking the profile face. From the segmented profile face, we extract the profile contour and from it, we extract 10 profile-contour reference points. Based on these, 20 individual facial muscle actions occurring alone or in a combination are recognized by a rule-based method. A recognition rate of 85% is achieved.
Maja Pantic, Ioannis Patras, Léon J. M. Rothkrantz
ICME (1)2
2001 Video Segmentation by MAP Labeling of Watershed Segments
abstract
This paper addresses the problem of spatio-temporal segmentation of video sequences. An initial intensity segmentation method (watershed segmentation) provides a number of initial segments which are subsequently labeled, with a known number of labels, according to motion information. The label field is modeled as a Markov random field where the statistical spatial and and temporal interactions are expressed on the basis of the initial watershed segments. The labeling criterion is the maximization of the conditional a posteriori probability of the label field given the motion hypotheses, the estimate of the label field of the previous frame, and the image intensities. For the optimization, an iterative motion estimation-labeling algorithm is proposed and experimental results are presented.
Ioannis Patras, Emile A. Hendriks, Reginald L. Lagendijk
IEEE Trans. Pattern Anal. Mach. Intell.1
1998 An Iterative Motion Estimation-Segmentation Method using Watershed Segments
abstract
This paper addresses the problem of segmentation of image sequences into moving objects. The segmentation (label) field is modeled by applying a Markov random field on regions provided by a watershed algorithm. The algorithm iterates between a motion estimation and a segmentation step. Temporal consistency is introduced by tracking the segmentation in time.
Ioannis Patras, Emile A. Hendriks, Reginald L. Lagendijk
ICIP (2)1
1996 Joint disparity and motion field estimation in stereoscopic image sequences
abstract
This work aims at determining, given two stereoscopic image sequences, at any time instant two dense velocity fields, for the left and the right sequence, and the disparity field. The disparity field of the previous stereoscopic pair is considered as known. Thus at the initial time instant the disparity field of the first stereoscopic pair is estimated. For both problems multiscale iterative relaxation algorithms are used. Results are given with real stereoscopic data.
Ioannis Patras, Nicolas Alvertos, Georgios Tziritas
ICPR1