EDBT 2026 Demo / reviewers in the wild / expert
Petros Koutras
dblp:146/6070
· DBLP profile ↗
20ranked-venue papers
4as first author
1since 2021 · last 2024
0000-0003-3183-2726ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 7 · 1 since 2021Systems, architecture and hardware · 4
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
3D vision · 44% Face, body and person analysis · 25% Learning paradigms · 19% | |
| Human-computer interaction and pervasive computing
2 papers |
Human-robot interaction · 42% Wearable and physiological sensing · 17% Health and well-being technologies · 15% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 61% Multimedia analysis and retrieval · 30% Audio and music processing · 9% |
Topics — the 16 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision
3d human reconstruction |
0.8 | 1 | 2024 | MeshPose: Unifying DensePose and 3D Body Mesh reconstruction · CVPR 2024 |
Computer vision › 3D vision
human mesh recovery |
0.8 | 1 | 2024 | MeshPose: Unifying DensePose and 3D Body Mesh reconstruction · CVPR 2024 |
Multimedia analysis and retrieval › audio-visual learning
audio-visual saliency |
0.4 | 1 | 2020 | STAViS: Spatio-Temporal AudioVisual Saliency Network · CVPR 2020 |
Image and video processing
saliency detection |
0.4 | 1 | 2020 | STAViS: Spatio-Temporal AudioVisual Saliency Network · CVPR 2020 |
Image and video processing › saliency detection
video saliency |
0.4 | 1 | 2020 | STAViS: Spatio-Temporal AudioVisual Saliency Network · CVPR 2020 |
Human-robot interaction
assistive robotics |
0.4 | 1 | 2019 | LSTM-based Network for Human Gait Stability Prediction in an Intelligent Robotic Rollator · ICRA 2019 |
Human-robot interaction › assistive robotics
robotic rollator |
0.4 | 1 | 2019 | LSTM-based Network for Human Gait Stability Prediction in an Intelligent Robotic Rollator · ICRA 2019 |
Machine learning › Learning paradigms
multiple instance learning |
0.3 | 1 | 2018 | Multimodal Visual Concept Learning With Weakly Supervised Techniques · CVPR 2018 |
Machine learning › Learning paradigms
weakly supervised learning |
0.3 | 1 | 2018 | Multimodal Visual Concept Learning With Weakly Supervised Techniques · CVPR 2018 |
Haptics and multimodal interaction › multimodal fusion
audiovisual fusion |
0.3 | 1 | 2018 | Multi3: Multi-Sensory Perception System for Multi-Modal Child Interaction with Multiple Robots · ICRA 2018 |
Human-robot interaction
child-robot interaction |
0.3 | 1 | 2018 | Multi3: Multi-Sensory Perception System for Multi-Modal Child Interaction with Multiple Robots · ICRA 2018 |
Interaction techniques and input › input sensing
gesture recognition |
0.3 | 1 | 2018 | Multi3: Multi-Sensory Perception System for Multi-Modal Child Interaction with Multiple Robots · ICRA 2018 |
Wearable and physiological sensing
sensor fusion |
0.3 | 1 | 2018 | Multi3: Multi-Sensory Perception System for Multi-Modal Child Interaction with Multiple Robots · ICRA 2018 |
Wearable and physiological sensing
gait analysis |
0.1 | 1 | 2019 | LSTM-based Network for Human Gait Stability Prediction in an Intelligent Robotic Rollator · ICRA 2019 |
Computer vision › Video understanding and tracking
action recognition |
0.1 | 1 | 2018 | Multimodal Visual Concept Learning With Weakly Supervised Techniques · CVPR 2018 |
Computer vision › Face, body and person analysis
face recognition |
0.1 | 1 | 2018 | Multimodal Visual Concept Learning With Weakly Supervised Techniques · CVPR 2018 |
Methods — techniques the papers use, named apart from their topics
weak supervision · 0.8end-to-end training · 0.8multimodal fusion · 0.4deep neural network · 0.4unscented kalman filter · 0.4sequence-to-sequence model · 0.4deep learning · 0.4LSTM · 0.4speech recognition · 0.3sensor fusion · 0.3semantic similarity · 0.3probabilistic labels · 0.3kinect · 0.3fuzzy sets · 0.3convex optimization · 0.3action recognition · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MeshPose: Unifying DensePose and 3D Body Mesh reconstructionabstractDensePose provides a pixel-accurate association of images with 3D mesh coordinates, but does not provide a 3D mesh, while Human Mesh Reconstruction (HMR) systems have high 2D reprojection error, as measured by DensePose localization metrics. In this work we introduce MeshPose to jointly tackle DensePose and HMR. For this we first introduce new losses that allow us to use weak DensePose supervision to accurately localize in 2D a subset of the mesh vertices (‘VertexPose’). We then lift these vertices to 3D, yielding a low-poly body mesh (‘MeshPose’). Our system is trained in an end -to-end manner and is the first HMR method to attain competitive DensePose accuracy, while also being lightweight and amenable to efficient inference, making it suitable for real-time AR applications. Eric-Tuan Le, Antonis Kakolyris, Petros Koutras, Himmy Tam, Efstratios Skordos, George Papandreou, Riza Alp Güler, Iasonas Kokkinos |
CVPR | 3 |
| 2020 | STAViS: Spatio-Temporal AudioVisual Saliency NetworkabstractWe introduce STAViS, a spatio-temporal audiovisual saliency network that combines spatio-temporal visual and auditory information in order to efficiently address the problem of saliency estimation in videos. Our approach employs a single network that combines visual saliency and auditory features and learns to appropriately localize sound sources and to fuse the two saliencies in order to obtain a final saliency map. The network has been designed, trained end-to-end, and evaluated on six different databases that contain audiovisual eye-tracking data of a large variety of videos. We compare our method against 8 different state-of-the-art visual saliency models. Evaluation results across databases indicate that our STAViS model outperforms our visual only variant as well as the other state-of-the-art models in the majority of cases. Also, the consistently good performance it achieves for all databases indicates that it is appropriate for estimating saliency "in-the-wild". The code is available at https://github.com/atsiami/STAViS. Antigoni Tsiami, Petros Koutras, Petros Maragos |
CVPR | 2 |
| 2019 | Deeply Supervised Multimodal Attentional Translation Embeddings for Visual Relationship DetectionabstractDetecting visual relationships, i.e.triplets, has been a challenging Scene Understanding task approached in the past via linguistic priors or spatial information in a single feature branch. We introduce a new deeply supervised two-branch architecture, the Multimodal Attentional Translation Embeddings, where the visual features of each branch are driven by a multimodal attentional mechanism that exploits spatio-linguistic similarities in a low-dimensional space. We present a variety of experiments comparing against all related approaches in the literature, as well as by re-implementing and fine-tuning several of them. Results on the commonly employed VRD dataset [1] show that the proposed method clearly outperforms all others, while we also justify our claims both quantitatively and qualitatively. Nikolaos Gkanatsios, Vassilis Pitsikalis, Petros Koutras, Athanasia Zlatintsi, Petros Maragos |
ICIP | 3 |
| 2019 | Showcasing Deeply Supervised Multimodal Attentional Translation Embeddings: a Demo for Visual Relationship DetectionabstractWe address the task of Visual Relationship Detection, i.e. the detection oftriplets in an image, introducing Multimodal Attentional Translation Embeddings (ICIP 2019 paper, id 3642). Motivated by the need of visualization and interpretation of the results, as well as the lack of other tools for online predictions on this task, we design and implement the first architecture for live inference of visual relationships on video streams and wild images, including research and engineering extensions, ablation models and a lightweight CPU-version. The code is available at https://bitbucket.org/deeplabai/vrd. Nikolaos Gkanatsios, Vassilis Pitsikalis, Petros Koutras, Athanasia Zlatintsi, Petros Maragos |
ICIP | 3 |
| 2019 | Video Processing and Learning in Assistive Robotic ApplicationsabstractThe integration of visual perception to robotic systems is a key research area in recent years. Advances in modern computer vision techniques along with the development of faster and more accurate visual sensors led to the emergence of new methods for robotic visual perception [1]. One area of research is the development of robotic assistive vision for Human-Robot Interaction (HRI) systems [2], [3]. The rapid increase of people with special needs, such as the elderly population, and the simultaneous reduction of personal care staff, reinforce the need for robotic assistants [4], [5]. There are many challenges in this area including the familiarity of these users with new technologies and the domain specific datasets, which are required for training user oriented models. Nowadays, modern assistive and social human-robot interaction requires the multimodal communication with speech, gestures and human movements so as to enhance the classic interaction with only spoken commands. Petros Koutras, Georgia Chalvatzaki, Antigoni Tsiami, Alexandros Nikolakakis, Costas S. Tzafestas, Petros Maragos |
ICIP | 1 |
| 2019 | LSTM-based Network for Human Gait Stability Prediction in an Intelligent Robotic RollatorabstractIn this work, we present a novel framework for on-line human gait stability prediction of the elderly users of an intelligent robotic rollator using Long Short Term Memory (LSTM) networks, fusing multimodal RGB-D and Laser Range Finder (LRF) data from non-wearable sensors. A Deep Learning (DL) based approach is used for the upper body pose estimation. The detected pose is used for estimating the body Center of Mass (CoM) using Unscented Kalman Filter (UKF). An Augmented Gait State Estimation framework exploits the LRF data to estimate the legs' positions and the respective gait phase. These estimates are the inputs of an encoder-decoder sequence to sequence model which predicts the gait stability state as Safe or Fall Risk walking. It is validated with data from real patients, by exploring different network architectures, hyperparameter settings and by comparing the proposed method with other baselines. The presented LSTM-based human gait stability predictor is shown to provide robust predictions of the human stability state, and thus has the potential to be integrated into a general user-adaptive control architecture as a fall-risk alarm. Georgia Chalvatzaki, Petros Koutras, Jack Hadfield, Xanthi S. Papageorgiou, Costas S. Tzafestas, Petros Maragos |
ICRA | 2 |
| 2019 | A Deep Learning Approach for Multi-View Engagement Estimation of Children in a Child-Robot Joint Attention TaskabstractIn this work, we tackle the problem of child engagement estimation while children freely interact with a robot in a friendly, room-like environment. We propose a deep learning-based multi-view solution that takes advantage of recent developments in human pose detection. We extract the child's pose from different RGB-D cameras placed regularly in the room, fuse the results and feed them to a deep Neural Network (NN) trained for classifying engagement levels. The deep network contains a recurrent layer, in order to exploit the rich temporal information contained in the pose data. The resulting method outperforms a number of baseline classifiers and provides a promising tool for better automatic understanding of a child's attitude, interest and attention while cooperating with a robot. The goal is to integrate this model in next-generation social robots as an attention monitoring tool during various Child Robot Interaction (CRI) tasks both for Typically Developed (TD) children and children affected by autism (ASD). Jack Hadfield, Georgia Chalvatzaki, Petros Koutras, Mehdi Khamassi, Costas S. Tzafestas, Petros Maragos |
IROS | 3 |
| 2019 | A behaviorally inspired fusion approach for computational audiovisual saliency modeling
Antigoni Tsiami, Petros Koutras, Athanasios Katsamanis, Argiro Vatakis, Petros Maragos |
Signal Process. Image Commun. | 2 |
| 2018 | Multimodal Visual Concept Learning With Weakly Supervised TechniquesabstractDespite the availability of a huge amount of video data accompanied by descriptive texts, it is not always easy to exploit the information contained in natural language in order to automatically recognize video concepts. Towards this goal, in this paper we use textual cues as means of supervision, introducing two weakly supervised techniques that extend the Multiple Instance Learning (MIL) framework: the Fuzzy Sets Multiple Instance Learning (FSMIL) and the Probabilistic Labels Multiple Instance Learning (PLMIL). The former encodes the spatio-temporal imprecision of the linguistic descriptions with Fuzzy Sets, while the latter models different interpretations of each description's semantics with Probabilistic Labels, both formulated through a convex optimization algorithm. In addition, we provide a novel technique to extract weak labels in the presence of complex semantics, that consists of semantic similarity computations. We evaluate our methods on two distinct problems, namely face and action recognition, in the challenging and realistic setting of movies accompanied by their screenplays, contained in the COGNIMUSE database. We show that, on both tasks, our method considerably outperforms a state-of-the-art weakly supervised approach, as well as other baselines. Giorgos Bouritsas, Petros Koutras, Athanasia Zlatintsi, Petros Maragos |
CVPR | 2 |
| 2018 | Far-Field Audio-Visual Scene Perception of Multi-Party Human-Robot Interaction for Children and AdultsabstractHuman-robot interaction (HRI) is a research area of growing interest with a multitude of applications for both children and adult user groups, as, for example, in edutainment and social robotics. Crucial, however, to its wider adoption remains the robust perception of HRI scenes in natural, untethered, and multi-party interaction scenarios, across user groups. Towards this goal, we investigate three focal HRI perception modules operating on data from multiple audio-visual sensors that observe the HRI scene from the far-field, thus bypassing limitations and platform-dependency of contemporary robotic sensing. In particular, the developed modules fuse intra- and/or inter-modality data streams to perform: (i) audio-visual speaker localization; (ii) distant speech recognition; and (iii) visual recognition of hand-gestures. Emphasis is also placed on ensuring high speech and gesture recognition rates for both children and adults. Development and objective evaluation of the three modules is conducted on a corpus of both user groups, collected by our far-field multisensory setup, for an interaction scenario of a question-answering “guess-the-object” collaborative HRI game with a “Furhat” robot. In addition, evaluation of the game incorporating the three developed modules is reported. Our results demonstrate robust far-field audio-visual perception of the multi-party HRI scene. Antigoni Tsiami, Panagiotis Paraskevas Filntisis, Niki Efthymiou, Petros Koutras, Gerasimos Potamianos, Petros Maragos |
ICASSP | 4 |
| 2018 | Multimodal Signal Processing and Learning Aspects of Human-Robot Interaction for an Assistive Bathing RobotabstractWe explore new aspects of assistive living on smart human-robot interaction (HRI) that involve automatic recognition and online validation of speech and gestures in a natural interface, providing social features for HRI. We introduce a whole framework and resources of a real-life scenario for elderly subjects supported by an assistive bathing robot, addressing health and hygiene care issues. We contribute a new dataset and a suite of tools used for data acquisition and a state-of-the-art pipeline for multimodal learning within the framework of the I -Support bathing robot, with emphasis on audio and RGB- D visual streams. We consider privacy issues by evaluating the depth visual stream along with the RGB, using Kinect sensors. The audio-gestural recognition task on this new dataset yields up to 84.5%, while the online validation of the I-Support system on elderly users accomplishes up to 84% when the two modalities are fused together. The results are promising enough to support further research in the area of multimodal recognition for assistive social HRI, considering the difficulties of the specific task. Athanasia Zlatintsi, Isidoros Rodomagoulakis, Petros Koutras, Athanasios Dometios, Vassilis Pitsikalis, Costas S. Tzafestas, Petros Maragos |
ICASSP | 3 |
| 2018 | Multi- View Fusion for Action Recognition in Child-Robot InteractionabstractAnswering the challenge of leveraging computer vision methods in order to enhance Human Robot Interaction (HRI) experience, this work explores methods that can expand the capabilities of an action recognition system in such tasks. A multi-view action recognition system is proposed for integration in HRI scenarios with special users, such as children, in which there is limited data for training and many state-of-the-art techniques face difficulties. Different feature extraction approaches, encoding methods and fusion techniques are combined and tested in order to create an efficient system that recognizes children pantomime actions. This effort culminates in the integration of a robotic platform and is evaluated under an alluring Children Robot Interaction scenario. Niki Efthymiou, Petros Koutras, Panagiotis Paraskevas Filntisis, Gerasimos Potamianos, Petros Maragos |
ICIP | 2 |
| 2018 | Multi3: Multi-Sensory Perception System for Multi-Modal Child Interaction with Multiple RobotsabstractChild-robot interaction is an interdisciplinary research area that has been attracting growing interest, primarily focusing on edutainment applications. A crucial factor to the successful deployment and wide adoption of such applications remains the robust perception of the child's multi-modal actions, when interacting with the robot in a natural and untethered fashion. Since robotic sensory and perception capabilities are platform-dependent and most often rather limited, we propose a multiple Kinect-based system to perceive the child-robot interaction scene that is robot-independent and suitable for indoors interaction scenarios. The audio-visual input from the Kinect sensors is fed into speech, gesture, and action recognition modules, appropriately developed in this paper to address the challenging nature of child-robot interaction. For this purpose, data from multiple children are collected and used for module training or adaptation. Further, information from the multiple sensors is fused to enhance module performance. The perception system is integrated in a modular multi-robot architecture demonstrating its flexibility and scalability with different robotic platforms. The whole system, called Multi3, is evaluated, both objectively at the module level and subjectively in its entirety, under appropriate child-robot interaction scenarios containing several carefully designed games between children and robots. Antigoni Tsiami, Petros Koutras, Niki Efthymiou, Panagiotis Paraskevas Filntisis, Gerasimos Potamianos, Petros Maragos |
ICRA | 2 |
| 2018 | Object Assembly Guidance in Child-Robot Interaction using RGB-D based 3D TrackingabstractThis work examines how and to what benefit an autonomous humanoid robot can supervise a child in an object assembly task. In order to understand the child's actions, a novel 3D object tracking algorithm for RGB-D data is employed. The tracker consists of two stages: the first performs a tracking-by-detection scheme on the color stream, to locate the objects on the image plane, while the second uses a particle filter that operates on the depth data stream to refine the first stage output and infer the objects' rotations. Given the six degrees-of-freedom of the assembly part poses, the system is able to recognize which connections have been completed at any given time. This information is then used to select an appropriate verbal or gestural response for the robot. Experimental results show that (a) the tracking algorithm is accurate, fast and robust to severe occlusions and fast movements, (b) the proposed method of assembly state estimation is indeed effective, and (c) the resulting Child-Robot Interaction scenario is educational and enjoyable for the children involved. Jack Hadfield, Petros Koutras, Niki Efthymiou, Gerasimos Potamianos, Costas S. Tzafestas, Petros Maragos |
IROS | 2 |
| 2016 | FMRI-based perceptual validation of a computational model for visual and auditory saliency in videosabstractIn this study, we make use of brain activation data to investigate the perceptual plausibility of a visual and an auditory model for visual and auditory saliency in video processing. These models have already been successfully employed in a number of applications. In addition, we experiment with parameters, modifications and suitable fusion schemes. As part of this work, fMRI data from complex video stimuli were collected, on which we base our analysis and results. The core part of the analysis involves the use of well-established methods for the manipulation of fMRI data and the examination of variability across brain responses of different individuals. Our results indicate a success in confirming the value of these saliency models in terms of perceptual plausibility. Georgia Panagiotaropoulou, Petros Koutras, Athanasios Katsamanis, Petros Maragos, Athanasia Zlatintsi, Athanassios Protopapas, Eustratios Karavasilis, Nikolaos Smyrnis |
ICIP | 2 |
| 2015 | Max-product dynamical systems and applications to audio-visual salient event detection in videosabstractThis paper introduces a theory for max-product systems by analyzing them as discrete-time nonlinear dynamical systems that obey a superposition of a weighted maximum type and evolve on nonlinear spaces which we call complete weighted lattices. Special cases of such systems have found applications in speech recognition as weighted finite-state transducers and in belief propagation on graphical models. Our theoretical approach establishes their representation in state and input-output spaces using monotone lattice operators, finds analytically their state and output responses using nonlinear convolutions, studies their stability, and provides optimal solutions to solving max-product matrix equations. Further, we apply these systems to extend the Viterbi algorithm in HMMs by adding control inputs and model cognitive processes such as detecting audio and visual salient events in multimodal video streams, which shows good performance as compared to human attention. Petros Maragos, Petros Koutras |
ICASSP | 2 |
| 2015 | Estimation of eye gaze direction angles based on active appearance modelsabstractIn this paper we demonstrate efficient methods for continuous estimation of eye gaze angles with application to sign language videos. The difficulty of the task lies on the fact that those videos contain images with low face resolution since they are recorded from distance. First, we proceed to the modeling of face and eyes region by training and fitting Global and Local Active Appearance Models (LAAM). Next, we propose a system for eye gaze estimation based on a machine learning approach. In the first stage of our method, we classify gaze into discrete classes using GMMs that are based either on the parameters of the LAAM, or on HOG descriptors for the eyes region. We also propose a method for computing gaze direction angles from GMM log-likelihoods. We qualitatively and quantitatively evaluate our methods on two sign language databases and compare with a state of the art geometric model of the eye based on LAAM landmarks, which provides an estimate in direction angles. Finally, we further evaluate our framework by getting ground truth data from an eye tracking system Our proposed methods, and especially the GMMs using LAAM parameters, demonstrate high accuracy and robustness even in challenging tasks. Petros Koutras, Petros Maragos |
ICIP | 1 |
| 2015 | Predicting audio-visual salient events based on visual, audio and text modalities for movie summarizationabstractIn this paper, we present a new and improved synergistic approach to the problem of audio-visual salient event detection and movie summarization based on visual, audio and text modalities. Spatio-temporal visual saliency is estimated through a perceptually inspired frontend based on 3D (space, time) Gabor filters and frame-wise features are extracted from the saliency volumes. For the auditory salient event detection we extract features based on Teager-Kaiser Energy Operator, while text analysis incorporates part-of-speech tagging and affective modeling of single words on the movie subtitles. For the evaluation of the proposed system, we employ an elementary and non-parametric classification technique like KNN. Detection results are reported on the MovSum database, using objective evaluations against ground-truth denoting the perceptually salient events, and human evaluations of the movie summaries. Our evaluation verifies the appropriateness of the proposed methods compared to our baseline system. Finally, our newly proposed summarization algorithm produces summaries that consist of salient and meaningful events, also improving the comprehension of the semantics. Petros Koutras, Athanasia Zlatintsi, Elias Iosif, Athanasios Katsamanis, Petros Maragos, Alexandros Potamianos |
ICIP | 1 |
| 2015 | A perceptually based spatio-temporal computational framework for visual saliency estimation
Petros Koutras, Petros Maragos |
Signal Process. Image Commun. | 1 |
| 2014 | Advances on action recognition in videos using an interest point detector based on multiband spatio-temporal energiesabstractThis paper proposes a new visual framework for action recognition in videos, that consists of an energy detector coupled with a carefully designed multiband energy based filterbank. The tracking of video energy is performed using perceptually inspired 3D Gabor filters combined with ideas from Dominant Energy Analysis. Within this framework, we utilize different alternatives such as non-linear energy operators where actions are implicitly considered as manifestations of spatio-temporal oscillations in the dynamic visual stream. Texture and motion decomposition of actions through multiband filtering is the basis of our approach. This new energy-based saliency measure of action videos leads to the extraction of local spatio-temporal interest points that give promising results for the task of action recognition. Such interest points are processed further in order to formulate a robust representation of an action in a video. Theoretical formulation is supported by evaluation in two popular action databases, in which our method seems to outperform the state of the art. Kevis-Kokitsi Maninis, Petros Koutras, Petros Maragos |
ICIP | 2 |