Panagiotis Paraskevas Filntisis

dblp:210/1927 · also Panayiotis Paraskevas Filntisis · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0002-2042-245XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Power in Unity: Combining in-Domain and out-of-Domain Pre-Training Strategies for EEG-Based Person Identification
abstract
We present the NTUA-IRAL team’s solution for the Person Identification track of the Signal Processing EEG-Music Emotion Recognition Grand Challenge, hosted at ICASSP. Our approach employs an ensemble of three CNNs, each pretrained using a distinct strategy: contrastive pre-training, traditional ImageNet pre-training, and task-specific pre-training on a publicly available EEG dataset. This diverse pre-training regimen enabled our models to achieve a test set accuracy of 100%, earning third place in the challenge subtrack.
Christos Garoufis, Marios Glytsos, Ioanna Chourdaki, Panagiotis Paraskevas Filntisis, Petros Maragos
ICASSP4
2025 Towards Open-Ended Robotic Exploration Using Vision-Inspired Similarity and Foundation Models
abstract
In the domain of robotics, achieving Lifelong Open-ended Learning Autonomy (LOLA) represents a significant milestone, especially in contexts where autonomous agents must adapt to unforeseen environmental variations and evolving objectives. This paper introduces VISOR (VisionSimilarity for Open-ended Robotic exploration), a vision-based framework designed to assist robotic agents in autonomously exploring and learning from new environments and objects, whether through guided or random exploration, without reliance on predefined design considerations. In that direction, VISOR acts as a perception mediator, classifying everything a robot encounters in a scene as either known or unknown. It further identifies potential distractors (e.g., background elements), known categories, or objects specified through text seeds. By leveraging recent advancements in vision foundation models, VISOR operates in a training-free manner. It begins by segmenting a scene into its constituent entities, regardless of familiarity, and then extracts robust visual representations for each one. These representations are compared against an adaptive memory system that evolves over time; unknown objects are assigned unique IDs and added to this memory as new classes, enriching the robot's understanding of its environment. We argue that this evolving memory can facilitate guided exploration through prior knowledge, enhancing the efficiency of robotic exploration, and validate this by designing two exploration scenarios and running both simulated and real-world experiments.
Panagiotis Paraskevas Filntisis, Efthymios Tsaprazlis, Paraskevas Oikonomou, Francesco Mattioli 0003, Vieri G. Santucci, George Retsinas, Petros Maragos
ICRA1
2025 Category-Level 6D Object Pose Estimation in Agricultural Settings Using a Lattice-Deformation Framework and Diffusion-Augmented Synthetic Data
abstract
Accurate 6D object pose estimation is essential for robotic grasping and manipulation, particularly in agriculture, where fruits and vegetables exhibit high intra-class variability in shape, size, and texture. The vast majority of existing methods rely on instance-specific CAD models or require depth sensors to resolve geometric ambiguities, making them impractical for real-world agricultural applications. In this work, we introduce PLANTPose, a novel framework for category-level 6D pose estimation that operates purely on RGB input. PLANT-Pose predicts both the 6D pose and deformation parameters relative to a base mesh, allowing a single category-level CAD model to adapt to unseen instances. This enables accurate pose estimation across varying shapes without relying on instance-specific data. To enhance realism and improve generalization, we also leverage Stable Diffusion to refine synthetic training images with realistic texturing, mimicking variations due to ripeness and environmental factors and bridging the domain gap between synthetic data and the real world. Our evaluations on a challenging benchmark that includes bananas of various shapes, sizes, and ripeness status demonstrate the effectiveness of our framework in handling large intraclass variations while maintaining accurate 6D pose predictions, significantly outperforming the state-of-the-art RGB-based approach MegaPose. Our code, data, and models are publicly available at https://github.com/mariosgly/PLANTPose.
Marios Glytsos, Panagiotis Paraskevas Filntisis, George Retsinas, Petros Maragos
IROS2
2025 Instance-Level Composed Image Retrieval
abstract
The progress of composed image retrieval (CIR), a popular research direction in image retrieval, where a combined visual and textual query is used, is held back by the absence of high-quality training and evaluation data. We introduce a new evaluation dataset, i-CIR, which, unlike existing datasets, focuses on an instance-level class definition. The goal is to retrieve images that contain the same particular object as the visual query, presented under a variety of modifications defined by textual queries. Its design and curation process keep the dataset compact to facilitate future research, while maintaining its challenge—comparable to retrieval among more than 40M random distractors—through a semi-automated selection of hard negatives. To overcome the challenge of obtaining clean, diverse, and suitable training data, we leverage pre-trained vision-and-language models (VLMs) in a training-free approach called BASIC. The method separately estimates query-image-to-image and query-text-to-image similarities, performing late fusion to upweight images that satisfy both queries, while down-weighting those that exhibit high similarity with only one of the two. Each individual similarity is further improved by a set of components that are simple and intuitive. BASIC sets a new state of the art on i-CIR but also on existing CIR datasets that follow a semantic-level class definition. Project page: https://vrg.fel.cvut.cz/icir/.
Bill Psomas, George Retsinas, Nikos Efthymiadis, Panagiotis Paraskevas Filntisis, Yannis Avrithis, Petros Maragos, Ondrej Chum, Giorgos Tolias
NeurIPS4
2024 3D Facial Expressions through Analysis-by-Neural-Synthesis
abstract
While existing methods for 3D face reconstruction from in-the-wild images excel at recovering the overall face shape, they commonly miss subtle, extreme, asymmetric, or rarely observed expressions. We improve upon these meth-ods with SMIRK (Spatial Modeling for Image-based Reconstruction of Kinesics), which faithfully reconstructs expres-sive 3D faces from images. We identify two key limitations in existing methods: shortcomings in their self-supervised training formulation, and a lack of expression diversity in the training images. For training, most methods employ differentiable rendering to compare a predicted face mesh with the input image, along with a plethora of additional loss functions. This differentiable rendering loss not only has to provide supervision to optimize for 3D face geom-etry, camera, albedo, and lighting, which is an ill-posed optimization problem, but the domain gap between ren-dering and input image further hinders the learning pro-cess. Instead, SMIRK replaces the differentiable rendering with a neural rendering module that, given the ren-dered predicted mesh geometry, and sparsely sampled pix-els of the input image, generates a face image. As the neural rendering gets color information from sampled im-age pixels, supervising with neural rendering-based reconstruction loss can focus solely on the geometry. Further it enables us to generate images of the input identity with varying expressions while training. These are then utilized as input to the reconstruction model and used as supervision with ground truth geometry. This effectively augments the training data and enhances the generalization for di-verse expressions. Our qualitative, quantitative and partic-ularly our perceptual evaluations demonstrate that SMIRK achieves the new state-of-the art performance on accurate expression reconstruction. For our method's source code, demo video and more, please visit our project webpage: https://georgeretsi.github.io/smirk/.
George Retsinas, Panagiotis Paraskevas Filntisis, Radek Danecek, Victoria Fernández Abrevaya, Anastasios Roussos, Timo Bolkart, Petros Maragos
CVPR2
2024 Augmenting Transformer Autoencoders with Phenotype Classification for Robust Detection of Psychotic Relapses
abstract
Recently, deep autoencoder architectures have received attention for the problem of unsupervised anomaly detection. Detecting psychotic relapses in mental health patients is a crucial challenge, often framed as anomaly detection, given the limited availability of data during relapsing states. In this paper, motivated by the fact that during relapses patients tend to undergo behavioral changes, we augment the classical autoencoder architecture with extra patient identification components. We show that formulating the problem as one of both signal reconstruction and patient identification largely improves the overall precision and robustness of relapse detection and significantly outperforms previous methods with a relative improvement of 15%. In addition, we also explore multiple ways to fuse the identification and reconstruction errors into a unified anomaly score that outperforms the results achieved by each error in isolation.
Niki Efthymiou, George Retsinas, Panagiotis Paraskevas Filntisis, Petros Maragos
ICASSP3
2023 Multimodal Recognition of Valence, Arousal and Dominance via Late-Fusion of Text, Audio and Facial Expressions
abstract
We present an approach for the prediction of valence, arousal, and dominance of people communicating via text/audio/video streams for a translation from and to sign languages.The approach consists of the fusion of the output of three CNN-based models dedicated to the analysis of text, audio, and facial expressions.Our experiments show that any combination of two or three modalities increases prediction performance for valence and arousal.
Fabrizio Nunnari, Annette Rios, Uwe D. Reichel, Chirag Bhuvaneshwara, Panagiotis Paraskevas Filntisis, Petros Maragos, Felix Burkhardt, Florian Eyben, Björn W. Schuller, Sarah Ebling
ESANN5
2023 Relapse Prediction from Long-Term Wearable Data Using Self-Supervised Learning and Survival Analysis
abstract
The introduction of biometric signal analysis in psychiatry could potentially reshape the field by making it more accurate, proactive and personalized. Such biosignals usually acquired from wearables encompass the quantification of human behavior and traits. In this study, we use long-term data acquired from commercial smartwatches, including kinetic and physiological signals, to extract information-thick descriptors that are used for the prediction of subsequent relapses in patients in the psychotic spectrum. Specifically, we propose a novel combination of methods based on Self-Supervised Learning and Survival Analysis that operates on unlabeled and censored data. When combined with other static features that describe the past course of the patient’s health, the proposed methodology yields promising predictive results in terms of two standard survival analysis metrics.
E. Fekas, Athanasia Zlatintsi, Panagiotis Paraskevas Filntisis, Christos Garoufis, Niki Efthymiou, Petros Maragos
ICASSP3
2023 Newton-Based Trainable Learning Rate
abstract
Selecting an appropriate learning rate for efficiently training deep neural networks is a difficult process that can be affected by numerous parameters, such as the dataset, the model architecture or even the batch size. In this work, we propose an algorithm for automatically adjusting the learning rate during the training process, assuming a gradient descent formulation. The rationale behind our approach is to train the learning rate along with the model weights. Specifically, we formulate first and second-order gradients w.r.t. the learning rate as functions of consecutive weight gradients, leading to a cost-effective implementation. Our extensive experimental evaluation validates the effectiveness of the proposed method for a plethora of different settings. The proposed method has proven to be robust to both the initial learning rate and the batch size, making it ideal for an off-the-shelf optimizing scheme.
George Retsinas, Giorgos Sfikas, Panagiotis Paraskevas Filntisis, Petros Maragos
ICASSP3
2023 E-Prevention: The ICASSP-2023 Challenge on Person Identification and Relapse Detection from Continuous Recordings of Biosignals
abstract
The e-Prevention challenge concerns the analysis and processing of long-term continuous recordings of biosignals recorded from wearable sensors, i.e., accelerometers, gyroscopes and heart rate monitors embedded in smartwatches, as well as sleep information and daily step count, in order to extract high-level representations of the wearer’s activity and behavior, termed as digital phenotypes. The ability of these digital phenotypes to quantify behavioral patterns and traits will be evaluated in two different tasks: 1) Person Identification, and 2) Relapse Detection in patients in the psychotic spectrum. The long-term data that will be used in this challenge have been acquired during the course of the e-Prevention project, an innovative integrated system for medical support that facilitates effective monitoring and relapse prevention in patients with mental disorders (i.e, schizophrenia and bipolar disorder). Specifically, the data were continuously collected from patients for a monitoring period of up to 2.5 years, while from the control subgroup for a period of 3 months, constituting one of the largest of its kind ever recorded.
Athanasia Zlatintsi, Panagiotis Paraskevas Filntisis, Niki Efthymiou, Christos Garoufis, George Retsinas, Thomas Sounapoglou, Ilias Maglogiannis, Panayiotis Tsanakas, Nikolaos Smyrnis, Petros Maragos
ICASSP2
2022 Neural Emotion Director: Speech-preserving semantic control of facial expressions in "in-the-wild" videos
abstract
In this paper, we introduce a novel deep learning method for photo-realistic manipulation of the emotional state of actors in “in-the-wild” videos. The proposed method is based on a parametric 3D face representation of the actor in the input scene that offers a reliable disentanglement of the facial identity from the head pose and facial expressions. It then uses a novel deep domain translation framework that alters the facial expressions in a consistent and plausible manner, taking into account their dynamics. Finally, the altered facial expressions are used to photo-realistically manipulate the facial region in the input scene based on an especially-designed neural face renderer. To the best of our knowledge, our method is the first to be capable of controlling the actor's facial expressions by even using as a sole input the semantic labels of the manipulated emotions, while at the same time preserving the speech-related lip movements. We conduct extensive qualitative and quantitative evaluations and comparisons, which demonstrate the effectiveness of our approach and the especially promising results that we obtain. Our method opens a plethora of new possibilities for useful applications of neural rendering technologies, ranging from movie post-production and video games to photo-realistic affective avatars.
Foivos Paraperas Papantoniou, Panagiotis Paraskevas Filntisis, Petros Maragos, Anastasios Roussos
CVPR2
2021 Exploiting Emotional Dependencies with Graph Convolutional Networks for Facial Expression Recognition
abstract
Over the past few years, deep learning methods have shown remarkable results in many face-related tasks including automatic facial expression recognition (FER) in-the-wild. Meanwhile, numerous models describing the human emotional states have been proposed by the psychology community. However, we have no clear evidence as to which representation is more appropriate and the majority of FER systems use either the categorical or the dimensional model of affect. Inspired by recent work in multi-label classification, this paper proposes a novel multi-task learning (MTL) framework that exploits the dependencies between these two models using a Graph Convolutional Network (GCN) to recognize facial expressions in-the-wild. Specifically, a shared feature representation is learned for both discrete and continuous recognition in a MTL setting. Moreover, the facial expression classifiers and the valence-arousal regressors are learned through a GCN that explicitly captures the dependencies between them. To evaluate the performance of our method under real-world conditions we perform extensive experiments on the AffectNet and Aff-Wild2 datasets. The results of our experiments show that our method is capable of improving the performance across different datasets and backbone architectures. Finally, we also surpass the previous state-of-the-art methods on the categorical model of AffectNet.
Panagiotis Antoniadis, Panagiotis Paraskevas Filntisis, Petros Maragos
FG2
2021 Leveraging Semantic Scene Characteristics and Multi-Stream Convolutional Architectures in a Contextual Approach for Video-Based Visual Emotion Recognition in the Wild
abstract
In this work we tackle the task of video-based visual emotion recognition in the wild. Standard methodologies that rely solely on the extraction of bodily and facial features often fall short of accurate emotion prediction in cases where the aforementioned sources of affective information are inaccessible due to head/body orientation, low resolution and poor illumination. We aspire to alleviate this problem by leveraging visual context in the form of scene characteristics and attributes, as part of a broader emotion recognition framework. Temporal Segment Networks (TSN) constitute the backbone of our proposed model. Apart from the RGB input modality, we make use of dense Optical Flow, following an intuitive multi-stream approach for a more effective encoding of motion. Furthermore, we shift our attention towards skeleton-based learning and leverage action-centric data as means of pretraining a Spatial-Temporal Graph Convolutional Network (ST-GCN) for the task of emotion recognition. Our extensive experiments on the challenging Body Language Dataset (BoLD) verify the superiority of our methods over existing approaches, while by properly incorporating all of the aforementioned modules in a network ensemble, we manage to surpass the previous best published recognition scores, by a large margin.
Ioannis Pikoulis, Panagiotis Paraskevas Filntisis, Petros Maragos
FG2
2020 Person Identification Using Deep Convolutional Neural Networks on Short-Term Signals from Wearable Sensors
abstract
In this work, we explore the discriminating ability of short-term signal patterns (e.g. few minutes long) with respect to the person identification task. We focus on signals recorded by simple wearable devices, such as smart watches, which can measure movements (accelerometer and gyroscope sensors) and biosignals (heart rate monitor). To address the person identification problem, we develop a deep neural network, based on one-dimensional convolutions, which receives raw signals from three different smartwatch sensors and predicts the person wearing the smartwatch. Experimental results indicate that even with signals from wearable sensors collected at intervals of only 10 minutes, different users can be identified with notably high accuracy, revealing the existence of distinct short-term patterns of movement and heart rate between different persons.
George Retsinas, Panagiotis Paraskevas Filntisis, Niki Efthymiou, Emmanouil Theodosis, Athanasia Zlatintsi, Petros Maragos
ICASSP2
2018 Far-Field Audio-Visual Scene Perception of Multi-Party Human-Robot Interaction for Children and Adults
abstract
Human-robot interaction (HRI) is a research area of growing interest with a multitude of applications for both children and adult user groups, as, for example, in edutainment and social robotics. Crucial, however, to its wider adoption remains the robust perception of HRI scenes in natural, untethered, and multi-party interaction scenarios, across user groups. Towards this goal, we investigate three focal HRI perception modules operating on data from multiple audio-visual sensors that observe the HRI scene from the far-field, thus bypassing limitations and platform-dependency of contemporary robotic sensing. In particular, the developed modules fuse intra- and/or inter-modality data streams to perform: (i) audio-visual speaker localization; (ii) distant speech recognition; and (iii) visual recognition of hand-gestures. Emphasis is also placed on ensuring high speech and gesture recognition rates for both children and adults. Development and objective evaluation of the three modules is conducted on a corpus of both user groups, collected by our far-field multisensory setup, for an interaction scenario of a question-answering “guess-the-object” collaborative HRI game with a “Furhat” robot. In addition, evaluation of the game incorporating the three developed modules is reported. Our results demonstrate robust far-field audio-visual perception of the multi-party HRI scene.
Antigoni Tsiami, Panagiotis Paraskevas Filntisis, Niki Efthymiou, Petros Koutras, Gerasimos Potamianos, Petros Maragos
ICASSP2
2018 Multi- View Fusion for Action Recognition in Child-Robot Interaction
abstract
Answering the challenge of leveraging computer vision methods in order to enhance Human Robot Interaction (HRI) experience, this work explores methods that can expand the capabilities of an action recognition system in such tasks. A multi-view action recognition system is proposed for integration in HRI scenarios with special users, such as children, in which there is limited data for training and many state-of-the-art techniques face difficulties. Different feature extraction approaches, encoding methods and fusion techniques are combined and tested in order to create an efficient system that recognizes children pantomime actions. This effort culminates in the integration of a robotic platform and is evaluated under an alluring Children Robot Interaction scenario.
Niki Efthymiou, Petros Koutras, Panagiotis Paraskevas Filntisis, Gerasimos Potamianos, Petros Maragos
ICIP3
2018 Multi3: Multi-Sensory Perception System for Multi-Modal Child Interaction with Multiple Robots
abstract
Child-robot interaction is an interdisciplinary research area that has been attracting growing interest, primarily focusing on edutainment applications. A crucial factor to the successful deployment and wide adoption of such applications remains the robust perception of the child's multi-modal actions, when interacting with the robot in a natural and untethered fashion. Since robotic sensory and perception capabilities are platform-dependent and most often rather limited, we propose a multiple Kinect-based system to perceive the child-robot interaction scene that is robot-independent and suitable for indoors interaction scenarios. The audio-visual input from the Kinect sensors is fed into speech, gesture, and action recognition modules, appropriately developed in this paper to address the challenging nature of child-robot interaction. For this purpose, data from multiple children are collected and used for module training or adaptation. Further, information from the multiple sensors is fused to enhance module performance. The perception system is integrated in a modular multi-robot architecture demonstrating its flexibility and scalability with different robotic platforms. The whole system, called Multi3, is evaluated, both objectively at the module level and subjectively in its entirety, under appropriate child-robot interaction scenarios containing several carefully designed games between children and robots.
Antigoni Tsiami, Petros Koutras, Niki Efthymiou, Panagiotis Paraskevas Filntisis, Gerasimos Potamianos, Petros Maragos
ICRA4
2017 Photorealistic adaptation and interpolation of facial expressions using HMMS and AAMS for audio-visual speech synthesis
abstract
In this paper, motivated by the continuously increasing presence of intelligent agents in everyday life, we address the problem of expressive photorealistic audio-visual speech synthesis, with a strong focus on the visual modality. Emotion constitutes one of the main driving factors of social life and it is expressed mainly through facial expressions. Synthesis of a talking head capable of expressive audio-visual speech is challenging due to the data overhead that arises when considering the vast number of emotions we would like the talking head to express. In order to tackle this challenge, we propose the usage of two methods, namely Hidden Markov Model (HMM) adaptation and interpolation, with HMMs modeling visual parameters via an Active Appearance Model (AAM) of the face. We show that through HMM adaptation we can successfully adapt a “neutral” talking head to a target emotion with a small amount of adaptation data, as well as that through HMM interpolation we can robustly achieve different levels of intensity for an emotion.
Panagiotis Paraskevas Filntisis, Athanasios Katsamanis, Petros Maragos
ICIP1
2017 Demonstration of an HMM-based photorealistic expressive audio-visual speech synthesis system
abstract
Summary form only given. The usage of conversational agents is rapidly increasing in everyday life (cortana, siri, etc.). It has been shown that the inclusion of a talking face, increases the intelligibility of speech and the naturalness of human-computer interaction. Furthermore, an agent capable of expressing emotions has a stronger appeal to the human party and affects the interlocutor's emotional state. The proposed demonstration is a Hidden Markov Model (HMM) based photorealistic audio-visual speech synthesis system, capable of expressing emotions [1, 2]. The system is capable of generating a talking head speaking in three emotions: happiness, anger, and sadness, plus in neutral speaking style. Further capabilities of the system include 1) the usage of HMM interpolation [3] in order to generate speech with mixtures of the original emotions (e.g., both anger and happiness), and speech with different levels of expressiveness (by mixing with the neutral emotion), 2) the usage of HMM adaptation [4], in order to adapt to a target emotion using only a few number of sentences. Equipment In order to showcase our system we will use a laptop and speakers. The system will run fully on the laptop. Demonstration Experience During the demonstration, viewers will have the opportunity to: 1. Watch videos of the talking head speaking in 3 different emotions (plus neutral) and see how the expressive talking head feels more natural compared to the talking head speaking in neutral style. 2. Watch the talking head speaking in two or more emotions at the same time, and see how the weights assigned to each emotion affects the outcome. It will also be of great interest to see which emotion each viewer perceives. In addition, through interpolation with the neutral emotion, viewers will be able to watch the talking head speak in different expressiveness levels for each emotion. 3. See how the neutral talking head can be adapted to speak in another emotion using only a few sentences, and how the number of sentences used affects the expressiveness of the resulting talking head.
Panagiotis Paraskevas Filntisis, Athanasios Katsamanis, Petros Maragos
ICIP1
2017 Video-realistic expressive audio-visual speech synthesis for the Greek language
Panagiotis Paraskevas Filntisis, Athanasios Katsamanis, Pirros Tsiakoulis, Petros Maragos
Speech Commun.1