VLDB 2026 Research / reviewers in the wild / expert
Konstantinos E. Papoutsakis
dblp:08/8738
· DBLP profile ↗
13ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-2467-8727ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Metapath-Driven Embeddings for Zero-Shot Object State Classification
Filippos Gouidis, Konstantinos E. Papoutsakis, Theodore Patkos, Antonis A. Argyros, Dimitris Plexousakis |
ICPR (5) | 2 |
| 2026 | Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?abstractAnticipating actions before they occur is a core challenge in action understanding research. While conventional methods rely on extracting and aggregating temporal information from videos, as humans we can often predict upcoming actions by observing a single moment from a scene, when given sufficient context. Can a model achieve this competence? The short answer is yes, although its effectiveness depends on the complexity of the task. In this work, we investigate to what extent video aggregation can be replaced with alternative modalities. To this end, based on recent advances in visual feature extraction and language-based reasoning, we introduce AAG, a method for Action Anticipation at a Glimpse. AAG combines RGB features with depth cues from a single frame for enhanced spatial reasoning, and incorporates prior action information to provide long-term context. This context is obtained either through textual summaries from Vision-Language Models, or from predictions generated by a single-frame action recognizer. Our results demonstrate that multimodal single-frame action anticipation using AAG can perform competitively compared to both temporally aggregated video baselines and state-of-the-art methods across three instructional activity datasets: IKEA-ASM, Meccano, and Assembly101. Manuel Benavent-Lledó, Konstantinos Bacharidis, Victoria Manousaki, Konstantinos E. Papoutsakis, Antonis A. Argyros, José García Rodríguez 0001 |
WACV | 4 |
| 2026 | A vision-based framework and dataset for human behavior understanding in industrial assembly linesabstractThis paper introduces a vision-based framework and dataset for capturing and understanding human behavior in industrial assembly lines, focusing on car door manufacturing. The framework leverages advanced computer vision techniques to estimate workers’ locations and 3D poses and analyze work postures, actions, and task progress. A key contribution is the introduction of the CarDA dataset, which contains domain-relevant assembly actions captured in a realistic setting to support the analysis of the framework for human pose and action analysis. The dataset comprises time-synchronized multi-camera RGB-D videos, motion capture data recorded in a real car manufacturing environment, and annotations for EAWS-based ergonomic risk scores and assembly activities. Experimental results demonstrate the effectiveness of the proposed approach in classifying worker postures and robust performance in monitoring assembly task progress. Konstantinos E. Papoutsakis, Nikolaos Bakalos, Athena Zacharia, Konstantinos Fragkoulis, Georgia Kapetadimitri, Maria Pateraki |
Comput. Vis. Image Underst. | 1 |
| 2025 | An End-to-End Class-Aware and Attention-Guided Model for Object State ClassificationabstractObject State Classification (OSC) is a critical task in computer vision, enabling systems to understand the functional state of objects. This work proposes a novel end-to-end architecture for OSC that leverages the inherent relationship between object classification and state recognition. Our approach first classifies the object and then uses object-specific attention mechanisms to focus on relevant features for state classification. This two-stage design allows the model to effectively capture object-state dependencies while maintaining modularity and flexibility. We conduct an extensive ablation study to analyze the impact of key parameters, such as attention mechanisms and loss weighting, and evaluate our method against three baselines across four benchmark datasets. Experimental results demonstrate that our approach outperforms competing methods by a significant margin, achieving state-of-the-art performance. Filippos Gouidis, Konstantinos E. Papoutsakis, Theodore Patkos, Antonis A. Argyros, Dimitris Plexousakis |
VCIP | 2 |
| 2025 | Recognizing Unseen States of Unknown Objects by Leveraging Knowledge GraphsabstractWe investigate the problem of Object State Classification (OSC) in the context of zero-shot learning. Specifically, we propose the first method for Zero-shot Object-agnostic State Classification (OaSC) that, given an image, infers the state of a single object without relying on the knowledge or the estimation of the object class. In that direction, we capitalize on Knowledge Graphs (KGs) for structuring and organizing external knowledge, which, in combination with visual information, enable effective inference of the states of objects that have not been encountered in the training set. Having this unique property, a significant strength of our method is that it can handle an Open Set of object classes. We investigate the performance of OaSC in various datasets and settings, against several hypotheses and in comparison with state-of-the-art approaches for object attribute classification. OaSC outperforms these methods significantly across all benchmarks.1 Filippos Gouidis, Konstantinos E. Papoutsakis, Theodore Patkos, Antonis A. Argyros, Dimitris Plexousakis |
WACV | 2 |
| 2021 | Action Prediction During Human-Object Interaction Based on DTW and Early Fusion of Human and Object Representations
Victoria Manousaki, Konstantinos E. Papoutsakis, Antonis A. Argyros |
ICVS | 2 |
| 2020 | Results of Field Trials with a Mobile Service Robot for Older Adults in 16 Private HouseholdsabstractIn this article, we present results obtained from field trials with the Hobbit robotic platform, an assistive, social service robot aiming at enabling prolonged independent living of older adults in their own homes. Our main contribution lies within the detailed results on perceived safety, usability, and acceptance from field trials with autonomous robots in real homes of older users. In these field trials, we studied how 16 older adults (75 plus) lived with autonomously interacting service robots over multiple weeks. Robots have been employed for periods of months previously in home environments for older people, and some have been tested with manipulation abilities, but this is the first time a study has tested a robot in private homes that provided the combination of manipulation abilities, autonomous navigation, and non-scheduled interaction for an extended period of time. This article aims to explore how older adults interact with such a robot in their private homes. Our results show that all users interacted with Hobbit daily, rated most functions as well working, and reported that they believe that Hobbit will be part of future elderly care. We show that Hobbit’s adaptive behavior approach towards the user increasingly eased the interaction between the users and the robot. Our trials reveal the necessity to move into actual users’ homes, as only there, we encounter real-world challenges and demonstrate issues such as misinterpretation of actions during non-scripted human-robot interaction. Markus Bajones, David Fischinger, Astrid Weiss, Paloma de la Puente, Daniel Wolf, Markus Vincze, Tobias Körtner, Markus Weninger, Konstantinos E. Papoutsakis, Damien Michel, Ammar Qammaz, Paschalis Panteleris, Michalis Foukarakis, Ilia Adami, Danae Ioannidi, Asterios Leonidis, Margherita Antona, Antonis A. Argyros, Peter Mayer 0002, Paul Panek, Håkan Eftring, Susanne Frennert |
ACM Trans. Hum. Robot Interact. | 9 |
| 2019 | Unsupervised and Explainable Assessment of Video Similarity
Konstantinos E. Papoutsakis, Antonis A. Argyros |
BMVC | 1 |
| 2018 | A graph-based approach for detecting common actions in motion capture data and videos
Costas Panagiotakis, Konstantinos E. Papoutsakis, Antonis A. Argyros |
Pattern Recognit. | 2 |
| 2017 | Temporal Action Co-Segmentation in 3D Motion Capture Data and VideosabstractGiven two action sequences, we are interested in spotting/co-segmenting all pairs of sub-sequences that represent the same action. We propose a totally unsupervised solution to this problem. No a-priori model of the actions is assumed to be available. The number of common sub-sequences may be unknown. The sub-sequences can be located anywhere in the original sequences, may differ in duration and the corresponding actions may be performed by a different person, in different style. We treat this type of temporal action co-segmentation as a stochastic optimization problem that is solved by employing Particle Swarm Optimization (PSO). The objective function that is minimized by PSO capitalizes on Dynamic Time Warping (DTW) to compare two action sub-sequences. Due to the generic problem formulation and solution, the proposed method can be applied to motion capture (i.e., 3D skeletal) data or to conventional RGB videos acquired in the wild. We present extensive quantitative experiments on standard data sets as well as on data sets we introduced in this paper. The obtained results demonstrate that the proposed method achieves a remarkable increase in co-segmentation quality compared to all tested state of the art methods. Konstantinos E. Papoutsakis, Costas Panagiotakis, Antonis A. Argyros |
CVPR | 1 |
| 2017 | A Framework for Online Segmentation and Classification of Modeled Actions Performed in the Context of Unmodeled OnesabstractIn this paper, we propose a discriminative framework for online simultaneous segmentation and classification of modeled visual actions that can be performed in the context of other unknown actions. To this end, we employ Hough transform to vote in a 3D space for the begin point, the end point, and the label of the segmented part of the input stream. A support vector machine is used to model each class and to suggest putative labeled segments on the timeline. To identify the most plausible segments among the putative ones, we apply a dynamic programming algorithm, which maximizes the likelihood for label assignment in linear time. The performance of our method is evaluated on synthetic as well as on real data (Weizmann, TUM Kitchen, UTKAD, and Berkeley Multimodal Human Action databases). Extensive quantitative results obtained on a number of standard data sets demonstrate that the proposed approach is of comparable accuracy with the state-of-the-art approaches for online stream segmentation and classification when all performed actions are known, and performs considerably better in the presence of unmodeled actions. Dimitrios I. Kosmopoulos, Konstantinos E. Papoutsakis, Antonis A. Argyros |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Segmentation and classification of modeled actions in the context of unmodeled ones
Dimitrios I. Kosmopoulos, Konstantinos E. Papoutsakis, Antonis A. Argyros |
BMVC | 2 |
| 2013 | Integrating tracking with fine object segmentation
Konstantinos E. Papoutsakis, Antonis A. Argyros |
Image Vis. Comput. | 1 |