Antonis A. Argyros

dblp:81/3715 · DBLP profile ↗
← Back
107ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0001-8230-3192ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 79 · 7 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 64 · 6 first-author · 11 since 2021Systems, architecture and hardware · 10 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3Computer networks · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Semantic Consistency in Hierarchical Classification via Probabilistic Logic Constraints
Filippos Gouidis, Antonis A. Argyros, Dimitris Plexousakis
ICAART (1)2
2026 Metapath-Driven Embeddings for Zero-Shot Object State Classification
Filippos Gouidis, Konstantinos E. Papoutsakis, Theodore Patkos, Antonis A. Argyros, Dimitris Plexousakis
ICPR (5)4
2026 Combining Facial Videos and Biosignals for Stress Estimation During Driving
Paraskevi Valergaki, Vassilis C. Nicodemou, Iasonas Oikonomidis, Antonis A. Argyros, Anastasios Roussos
ICPR (13)4
2026 Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
abstract
Anticipating actions before they occur is a core challenge in action understanding research. While conventional methods rely on extracting and aggregating temporal information from videos, as humans we can often predict upcoming actions by observing a single moment from a scene, when given sufficient context. Can a model achieve this competence? The short answer is yes, although its effectiveness depends on the complexity of the task. In this work, we investigate to what extent video aggregation can be replaced with alternative modalities. To this end, based on recent advances in visual feature extraction and language-based reasoning, we introduce AAG, a method for Action Anticipation at a Glimpse. AAG combines RGB features with depth cues from a single frame for enhanced spatial reasoning, and incorporates prior action information to provide long-term context. This context is obtained either through textual summaries from Vision-Language Models, or from predictions generated by a single-frame action recognizer. Our results demonstrate that multimodal single-frame action anticipation using AAG can perform competitively compared to both temporally aggregated video baselines and state-of-the-art methods across three instructional activity datasets: IKEA-ASM, Meccano, and Assembly101.
Manuel Benavent-Lledó, Konstantinos Bacharidis, Victoria Manousaki, Konstantinos E. Papoutsakis, Antonis A. Argyros, José García Rodríguez 0001
WACV5
2026 Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors
abstract
We revisit the role of texture in monocular 3D hand reconstruction, not as an afterthought for photorealism, but as a dense, spatially grounded cue that can actively support pose and shape estimation. Our observation is simple: even in high-performing models, the overlay between predicted hand geometry and image appearance is often imperfect, suggesting that texture alignment may be an underused supervisory signal. We propose a lightweight texture module that embeds per-pixel observations into UV texture space and enables a novel dense alignment loss between predicted and observed hand appearances. Our approach assumes access to a differentiable rendering pipeline and a model that maps images to 3D hand meshes with known topology, allowing us to back-project a textured hand onto the image and perform pixel-based alignment. The module is self-contained and easily pluggable into existing reconstruction pipelines. To isolate and highlight the value of texture-guided supervision, we augment HaMeR [47], a high-performing yet unadorned transformer architecture for 3D hand pose estimation. The resulting system improves both accuracy and realism, demonstrating the value of appearance-guided alignment in hand reconstruction. Code, models and weights are available at https://gkarv.github.io/hand-texture-module
Giorgos Karvounas, Nikolaos Kyriazis, Iasonas Oikonomidis, Georgios Pavlakos, Antonis A. Argyros
WACV5
2026 Vision-based mistake analysis in procedural activities: A review of advances and challenges
abstract
Mistake analysis in procedural activities is a critical area of research with applications spanning industrial automation, physical rehabilitation, education and human–robot collaboration. This paper reviews vision-based methods for detecting and predicting mistakes in structured tasks, focusing on procedural and executional errors. By leveraging advancements in computer vision, including action recognition, anticipation and activity understanding, vision-based systems can identify deviations in task execution, such as incorrect sequencing, use of improper techniques, or timing errors. We provide a comprehensive overview of existing datasets, evaluation metrics and state-of-the-art methods, categorizing approaches based on their use of procedural structure, supervision levels and learning strategies. Open challenges, such as distinguishing permissible variations from true mistakes and modeling error propagation are discussed alongside future directions, including neuro-symbolic reasoning and counterfactual state modeling. This work aims to establish a unified perspective on vision-based mistake analysis in procedural activities, highlighting its potential across diverse domains and aspects.
Konstantinos Bacharidis, Antonis A. Argyros
Comput. Vis. Image Underst.2
2025 Enact: Entropy-Based Clustering of Attention Input for Reducing the Computational Needs of Object Detection Transformers
abstract
Transformers demonstrate competitive performance in terms of precision on the problem of vision-based object detection. However, they require considerable computational re-sources due to the quadratic size of the attention weights. In this work, we propose to cluster the transformer input on the basis of its entropy, due to its similarity between same object pixels. This is expected to reduce GPU usage during training, while maintaining reasonable accuracy. This idea is realized with an implemented module that is called ENtropy-based Attention Clustering for detection Trans-formers (ENACT), which serves as a plug-in to any multi-head self-attention based transformer network. Experiments on the COCO object detection dataset and three detection transformers demonstrate that the requirements on memory are reduced, while the detection accuracy is degraded only slightly. The code of the ENACT module is available at https://github.com/GSavathrakis/ENACT.
Giorgos Savathrakis, Antonis A. Argyros
ICIP2
2025 An End-to-End Class-Aware and Attention-Guided Model for Object State Classification
abstract
Object State Classification (OSC) is a critical task in computer vision, enabling systems to understand the functional state of objects. This work proposes a novel end-to-end architecture for OSC that leverages the inherent relationship between object classification and state recognition. Our approach first classifies the object and then uses object-specific attention mechanisms to focus on relevant features for state classification. This two-stage design allows the model to effectively capture object-state dependencies while maintaining modularity and flexibility. We conduct an extensive ablation study to analyze the impact of key parameters, such as attention mechanisms and loss weighting, and evaluate our method against three baselines across four benchmark datasets. Experimental results demonstrate that our approach outperforms competing methods by a significant margin, achieving state-of-the-art performance.
Filippos Gouidis, Konstantinos E. Papoutsakis, Theodore Patkos, Antonis A. Argyros, Dimitris Plexousakis
VCIP4
2025 Recognizing Unseen States of Unknown Objects by Leveraging Knowledge Graphs
abstract
We investigate the problem of Object State Classification (OSC) in the context of zero-shot learning. Specifically, we propose the first method for Zero-shot Object-agnostic State Classification (OaSC) that, given an image, infers the state of a single object without relying on the knowledge or the estimation of the object class. In that direction, we capitalize on Knowledge Graphs (KGs) for structuring and organizing external knowledge, which, in combination with visual information, enable effective inference of the states of objects that have not been encountered in the training set. Having this unique property, a significant strength of our method is that it can handle an Open Set of object classes. We investigate the performance of OaSC in various datasets and settings, against several hypotheses and in comparison with state-of-the-art approaches for object attribute classification. OaSC outperforms these methods significantly across all benchmarks.1
Filippos Gouidis, Konstantinos E. Papoutsakis, Theodore Patkos, Antonis A. Argyros, Dimitris Plexousakis
WACV4
2025 Enhancing action recognition by leveraging the hierarchical structure of actions and textual context
abstract
We propose a novel approach to improve action recognition by exploiting the hierarchical organization of actions and by incorporating contextualized textual information, including location and previous actions, to reflect the action’s temporal context . To achieve this, we introduce a transformer architecture tailored for action recognition that employs both visual and textual features. Visual features are obtained from RGB and optical flow data, while text embeddings represent contextual information. Furthermore, we define a joint loss function to simultaneously train the model for both coarse- and fine-grained action recognition, effectively exploiting the hierarchical nature of actions. To demonstrate the effectiveness of our method, we extend the Toyota Smarthome Untrimmed (TSU) dataset by incorporating action hierarchies, resulting in the Hierarchical TSU dataset , a hierarchical dataset designed for monitoring activities of the elderly in home environments. An ablation study assesses the performance impact of different strategies for integrating contextual and hierarchical data. Experimental results demonstrate that the proposed method consistently outperforms SOTA methods on the Hierarchical TSU dataset, Assembly101 and IkeaASM, achieving over a 17% improvement in top-1 accuracy. • We propose a vision-language transformer model that improves action recognition using contextual information and action hierarchies, outperforming state-of-the-art visual-only models on the Hierarchical TSU, Assembly101 and IkeaASM action recognition benchmarks. • We introduce the Hierarchical TSU dataset for contextual action recognition with structured action hierarchies for activities of daily living. • We conduct an extensive ablation study on integrating contextual and hierarchical data to enhance action recognition performance.
Manuel Benavent-Lledó, David Mulero-Pérez, David Ortiz-Perez, José García Rodríguez 0001, Antonis A. Argyros
Comput. Vis. Image Underst.5
2024 Noise-robust person re-identification through nearest-neighbor sample filtering
abstract
Person re-identification is an important component of vision-based surveillance systems. The robustness of these components depends heavily on mechanisms that safeguard them from noisy observations which may affect negatively the quality of the required training data. In this work we present an approach that capitalizes of the KNN algorithm for identifying noisy observations and either relabeling them or excluding them from further consideration. We evaluate our approach on standard datasets and on various noise types and contamination levels. The performed experiments demonstrate that on two classical noise types, our approach performs on par to the state of the art. However, in a third noise type that is especially common in the person re-identification use cases, our method is superior to the state of the art by a great margin. In all cases, our approach results in considerable computational savings.
George Galanakis, Xenophon Zabulis, Antonis A. Argyros
AVSS3
2022 A Review on Deep Learning Techniques for Video Prediction
abstract
The ability to predict, anticipate and reason about future outcomes is a key component of intelligent decision-making systems. In light of the success of deep learning in computer vision, deep-learning-based video prediction emerged as a promising research direction. Defined as a self-supervised learning task, video prediction represents a suitable framework for representation learning, as it demonstrated potential capabilities for extracting meaningful representations of the underlying patterns in natural videos. Motivated by the increasing interest in this task, we provide a review on the deep learning methods for prediction in video sequences. We first define the video prediction fundamentals, as well as mandatory background concepts and the most used datasets. Next, we carefully analyze existing video prediction models organized according to a proposed taxonomy, highlighting their contributions and their significance in the field. The summary of the datasets and methods is accompanied with experimental results that facilitate the assessment of the state of the art on a quantitative basis. The paper is summarized by drawing some general conclusions, identifying open research challenges and by pointing out future research directions.
Sergiu Ovidiu-Oprea, Pablo Martinez-Gonzalez, Alberto Garcia-Garcia, John Alejandro Castro-Vargas, Sergio Orts, José García Rodríguez 0001, Antonis A. Argyros
IEEE Trans. Pattern Anal. Mach. Intell.7
2021 Multi-view Image-based Hand Geometry Refinement using Differentiable Monte Carlo Ray Tracing
Giorgos Karvounas, Nikolaos Kyriazis, Iasonas Oikonomidis, Aggeliki Tsoli, Antonis A. Argyros
BMVC5
2021 Towards Holistic Real-time Human 3D Pose Estimation using MocapNETs
Ammar Qammaz, Antonis A. Argyros
BMVC2
2021 Action Prediction During Human-Object Interaction Based on DTW and Early Fusion of Human and Object Representations
Victoria Manousaki, Konstantinos E. Papoutsakis, Antonis A. Argyros
ICVS3
2021 Multi-GPU SNN Simulation with Static Load Balancing
abstract
We present a SNN simulator which scales to millions of neurons, billions of synapses, and 8 GPUs. This is made possible by 1) a novel, cache-aware spike transmission algorithm 2) a model parallel multi-GPU distribution scheme and 3) a static, yet very effective load balancing strategy. The simulator further features an easy to use API and the ability to create custom models. We compare the proposed simulator against two state of the art ones on a series of benchmarks using three well-established models. We find that our simulator is faster, consumes less memory, and scales linearly with the number of GPUs.
Dennis Bautembach, Iasonas Oikonomidis, Antonis A. Argyros
IJCNN3
2021 H-GAN: the power of GANs in your Hands
abstract
We present HandGAN (H-GAN), a cycle-consistent adversarial learning approach implementing multi-scale perceptual discriminators. It is designed to translate synthetic images of hands to the real domain. Synthetic hands provide complete ground-truth annotations, yet they are not representative of the target distribution of real-world data. We strive to provide the perfect blend of a realistic hand appearance with synthetic annotations. Relying on image-to-image translation, we improve the appearance of synthetic hands to approximate the statistical distribution underlying a collection of real images of hands. H-GAN tackles not only the cross-domain tone mapping but also structural differences in localized areas such as shading discontinuities. Results are evaluated on a qualitative and quantitative basis improving previous works. Furthermore, we relied on the hand classification task to claim our generated hands are statistically similar to the real domain of hands.
Sergiu Ovidiu-Oprea, Giorgos Karvounas, Pablo Martinez-Gonzalez, Nikolaos Kyriazis, Sergio Orts, Iasonas Oikonomidis, Alberto Garcia-Garcia, Aggeliki Tsoli, José García Rodríguez 0001, Antonis A. Argyros
IJCNN10
2020 Towards a visual Sign Language dataset for home care services
abstract
We present our work towards creating a dataset, which is intended to be used for the implementation of a home care services system for the deaf. The dataset includes recorded realistic scenarios of interactions between deaf patients and mental health experts in their native sign language. The scenarios allow for contextualized representations, in contrast to typical datasets presenting isolated signs or sentences. It includes continuous videos in RGB and depth, which are challenging to analyze and closely resemble real-life scenarios. The research on representation of signs is supported by providing the hand shapes and trajectories for every video using hand and skeleton models, as well as facial features. Furthermore, the dataset may be used for studying the emotional context, since such conversations are typically emotionally charged.
Dimitrios I. Kosmopoulos, Iasonas Oikonomidis, Constantinos Constantinopoulos, Nikolaos Arvanitis, Klimis Antzakas, Aristeidis Bifis, Georgios Lydakis, Anastasios Roussos, Antonis A. Argyros
FG9
2020 Extracting Action Hierarchies from Action Labels and their Use in Deep Action Recognition
abstract
Human activity recognition is a fundamental and challenging task in computer vision. Its solution can support multiple and diverse applications in areas including but not limited to smart homes, surveillance, daily living assistance, Human-Robot Collaboration (HRC), etc. In realistic conditions, the complexity of human activities ranges from simple coarse actions, such as siting or standing up, to more complex activities that consist of multiple actions with subtle variations in appearance and motion patterns. A large variety of existing datasets target specific action classes, with some of them being coarse and others being fine-grained. In all of them, a description of the action and its complexity is manifested in the action label sentence. As the action/activity complexity increases, so is the label sentence size and the amount of action-related semantic information contained in this description. In this paper, we propose an approach to exploit the information content of these action labels to formulate a coarse-to-fine action hierarchy based on linguistic label associations, and investigate the potential benefits and drawbacks. Moreover, in a series of quantitative and qualitative experiments, we show that the exploitation of this hierarchical organization of action classes in different levels of granularity improves the learning speed and overall performance of a range of baseline and mid-range deep architectures for human action recognition (HAR).
Konstantinos Bacharidis, Antonis A. Argyros
ICPR2
2020 Occlusion-tolerant and personalized 3D human pose estimation in RGB images
abstract
We introduce a real-time method that estimates the 3D human pose directly in the popular Bio Vision Hierarchy (BVH) format, given estimations of the 2D body joints originating from monocular color images. Our contributions include: (a) A novel and compact 2D pose representation. (b) A human body orientation classifier and an ensemble of orientation-tuned neural networks that regress the 3D human pose by also allowing for the decomposition of the body to an upper and lower kinematic hierarchy. This permits the recovery of the human pose even in the case of significant occlusions. (c) An efficient Inverse Kinematics solver that refines the neural-network-based solution providing 3D human pose estimations that are consistent with the limb sizes of a target person (if known). All the above yield a 33% accuracy improvement on the Human 3.6 Million (H3.6M) dataset compared to the baseline method (MocapNET) while maintaining real-time performance (70 fps in CPU-only execution).
Ammar Qammaz, Antonis A. Argyros
ICPR2
2020 Improving Deep Learning Approaches for Human Activity Recognition based on Natural Language Processing of Action Labels
abstract
Human activity recognition has always been an appealing research topic in computer vision due its theoretic interest and vast range of applications. In recent years, machine learning has dominated computer vision and human activity recognition research. Supervised learning methods and especially deep learning-based ones are considered to provide the best solutions for this task, achieving state-of-the art results. However, the performance of deep learning-based approaches depends greatly on the modelling capabilities of the spatio-temporal neural network architecture and the learning goals of the training process. Moreover, the design complexity is task-depended. In this paper, we show that we can exploit the information contained in the label description of action classes (action labels) to extract information regarding their similarity which can then be used to steer the learning process and improve the activity recognition performance. Moreover, we experimentally verify that the adopted strategy can be useful in both single and multi-stream architectures, providing better scalability on the training of the network in more complex datasets featuring activity classes with larger intra- and inter-class similarities.
Konstantinos Bacharidis, Antonis A. Argyros
IJCNN2
2020 Faster and Simpler SNN Simulation with Work Queues
abstract
We present a clock-driven Spiking Neural Network simulator which is up to 3x faster than the state of the art while, at the same time, being more general and requiring less programming effort on both the user's and maintainer's side. This is made possible by designing our pipeline around "work queues" which act as interfaces between stages and greatly reduce implementation complexity. We evaluate our work using three well-established SNN models on a series of benchmarks.
Dennis Bautembach, Iasonas Oikonomidis, Nikolaos Kyriazis, Antonis A. Argyros
IJCNN4
2020 Learning to Infer the Depth Map of a Hand from its Color Image
abstract
We present the first direct approach targeted explicitly on human hands that infers depth from monocular RGB images. We achieve this with a Convolutional Neural Network (CNN) that employs a stacked hourglass model as its main building block. Intermediate supervision is used in several outputs of the proposed architecture in a staged approach. To aid the process of training and inference, hand segmentation masks are also estimated in such intermediate supervision steps, and used to guide the subsequent depth estimation process. In order to train and evaluate the proposed method we compile and make publicly available HandRGBD, a new dataset of 20,601 views of hands, each consisting of an RGB image and an aligned depth map. Based on HandRGBD, we explore variants of the proposed approach in an ablative study and determine the most accurate one. The results of an extensive experimental evaluation demonstrate that hand depth estimation from a single RGB frame can be achieved with an accuracy of 22mm, which is comparable to the accuracy achieved by contemporary low-cost depth cameras. Such a 3D reconstruction of hands based on RGB information is valuable as a final result on its own right, but also as an input to several other hand analysis and perception algorithms that require depth input. In this context, the proposed approach bridges the gap between RGB and RGBD, by making all existing RGBD-based methods applicable to RGB input.
Vassilis C. Nicodemou, Iasonas Oikonomidis, Georgios Tzimiropoulos, Antonis A. Argyros
IJCNN4
2020 Region-based Fitting of Overlapping Ellipses and its application to cells segmentation
Costas Panagiotakis, Antonis A. Argyros
Image Vis. Comput.2
2020 Single-shot 3D hand pose estimation using radial basis function networks trained on synthetic data
Vassilis C. Nicodemou, Iasonas Oikonomidis, Antonis A. Argyros
Pattern Anal. Appl.3
2020 Results of Field Trials with a Mobile Service Robot for Older Adults in 16 Private Households
abstract
In this article, we present results obtained from field trials with the Hobbit robotic platform, an assistive, social service robot aiming at enabling prolonged independent living of older adults in their own homes. Our main contribution lies within the detailed results on perceived safety, usability, and acceptance from field trials with autonomous robots in real homes of older users. In these field trials, we studied how 16 older adults (75 plus) lived with autonomously interacting service robots over multiple weeks. Robots have been employed for periods of months previously in home environments for older people, and some have been tested with manipulation abilities, but this is the first time a study has tested a robot in private homes that provided the combination of manipulation abilities, autonomous navigation, and non-scheduled interaction for an extended period of time. This article aims to explore how older adults interact with such a robot in their private homes. Our results show that all users interacted with Hobbit daily, rated most functions as well working, and reported that they believe that Hobbit will be part of future elderly care. We show that Hobbit’s adaptive behavior approach towards the user increasingly eased the interaction between the users and the robot. Our trials reveal the necessity to move into actual users’ homes, as only there, we encounter real-world challenges and demonstrate issues such as misinterpretation of actions during non-scripted human-robot interaction.
Markus Bajones, David Fischinger, Astrid Weiss, Paloma de la Puente, Daniel Wolf, Markus Vincze, Tobias Körtner, Markus Weninger, Konstantinos E. Papoutsakis, Damien Michel, Ammar Qammaz, Paschalis Panteleris, Michalis Foukarakis, Ilia Adami, Danae Ioannidi, Asterios Leonidis, Margherita Antona, Antonis A. Argyros, Peter Mayer 0002, Paul Panek, Håkan Eftring, Susanne Frennert
ACM Trans. Hum. Robot Interact.18
2020 REMPE: Registration of Retinal Images Through Eye Modelling and Pose Estimation
abstract
OBJECTIVE: In-vivo assessment of small vessels can promote accurate diagnosis and monitoring of diseases related to vasculopathy, such as hypertension and diabetes. The eye provides a unique, open, and accessible window for directly imaging small vessels in the retina with non-invasive techniques, such as fundoscopy. In this context, accurate registration of retinal images is of paramount importance in the comparison of vessel measurements from original and follow-up examinations, which is required for monitoring the disease and its treatment. At the same time, retinal registration exhibits a range of challenges due to the curved shape of the retina and the modification of imaged tissue across examinations. Thereby, the objective is to improve the state-of-the-art in the accuracy of retinal image registration. METHOD: In this work, a registration framework that simultaneously estimates eye pose and shape is proposed. Corresponding points in the retinal images are utilized to solve the registration as a 3D pose estimation. RESULTS: The proposed framework is evaluated quantitatively and shown to outperform state-of-the-art methods in retinal image registration for fundoscopy images. CONCLUSION: Retinal image registration methods based on eye modelling allow to perform more accurate registration than conventional methods. SIGNIFICANCE: This is the first method to perform retinal image registration combined with eye modelling. The method improves the state-of-the-art in accuracy of retinal registration for fundoscopy images, quantitatively evaluated in benchmark datasets annotated with ground truth. The implementation of registration method has been made publicly available.
Carlos Hernandez-Matas, Xenophon Zabulis, Antonis A. Argyros
IEEE J. Biomed. Health Informatics3
2020 3D Hand Tracking in the Presence of Excessive Motion Blur
abstract
We present a sensor-fusion method that exploits a depth camera and a gyroscope to track the articulation of a hand in the presence of excessive motion blur. In case of slow and smooth hand motions, the existing methods estimate the hand pose fairly accurately and robustly, despite challenges due to the high dimensionality of the problem, self-occlusions, uniform appearance of hand parts, etc. However, the accuracy of hand pose estimation drops considerably for fast-moving hands because the depth image is severely distorted due to motion blur. Moreover, when hands move fast, the actual hand pose is far from the one estimated in the previous frame, therefore the assumption of temporal continuity on which tracking methods rely, is not valid. In this paper, we track fast-moving hands with the combination of a gyroscope and a depth camera. As a first step, we calibrate a depth camera and a gyroscope attached to a hand so as to identify their time and pose offsets. Following that, we fuse the rotation information of the calibrated gyroscope with model-based hierarchical particle filter tracking. A series of quantitative and qualitative experiments demonstrate that the proposed method performs more accurately and robustly in the presence of motion blur, when compared to state of the art algorithms, especially in the case of very fast hand rotations.
Gabyong Park, Antonis A. Argyros, Woontack Woo
IEEE Trans. Vis. Comput. Graph.2
2019 Unsupervised and Explainable Assessment of Video Similarity
Konstantinos E. Papoutsakis, Antonis A. Argyros
BMVC2
2019 MocapNET: Ensemble of SNN Encoders for 3D Human Pose Estimation in RGB Images
Ammar Qammaz, Antonis A. Argyros
BMVC2
2019 A Two-Stage Approach for Commonality-Based Temporal Localization of Periodic Motions
Costas Panagiotakis, Antonis A. Argyros
ICVS2
2019 3D Hand Tracking by Employing Probabilistic Principal Component Analysis to Model Action Priors
Emmanouil Oulof Porfyrakis, Alexandros Makris, Antonis A. Argyros
ICVS3
2018 Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future Goals
abstract
In this paper, we strive to answer two questions: What is the current state of 3D hand pose estimation from depth images? And, what are the next challenges that need to be tackled? Following the successful Hands In the Million Challenge (HIM2017), we investigate the top 10 state-of-the-art methods on three tasks: single frame 3D pose estimation, 3D hand tracking, and hand pose estimation during object interaction. We analyze the performance of different CNN structures with regard to hand shape, joint visibility, view point and articulation distributions. Our findings include: (1) isolated 3D hand pose estimation achieves low mean errors (10 mm) in the view point range of [70, 120] degrees, but it is far from being solved for extreme view points; (2) 3D volumetric representations outperform 2D CNNs, better capturing the spatial structure of the depth data; (3) Discriminative methods still generalize poorly to unseen hand shapes; (4) While joint occlusions pose a challenge for most methods, explicit modeling of structure constraints can significantly narrow the gap between errors on visible and occluded joints.
Shanxin Yuan, Guillermo Garcia-Hernando, Björn Stenger, Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee, Pavlo Molchanov 0001, Jan Kautz, Sina Honari, Liuhao Ge, Junsong Yuan 0001, Xinghao Chen 0001, Guijin Wang, Fan Yang 0032, Kai Akiyama, Yang Wu 0001, Qingfu Wan, Meysam Madadi, Sergio Escalera, Shile Li, Dongheui Lee, Iasonas Oikonomidis, Antonis A. Argyros, Tae-Kyun Kim 0001
CVPR23
2018 Joint 3D Tracking of a Deformable Object in Interaction with a Hand
Aggeliki Tsoli, Antonis A. Argyros
ECCV (14)2
2018 Cell Segmentation Via Region-Based Ellipse Fitting
abstract
We present a region based method for segmenting and splitting images of cells in an automatic and unsupervised manner. The detection of cell nuclei is based on the Bradley's method. False positives are automatically identified and rejected based on shape and intensity features. Additionally, the proposed method is able to automatically detect and split touching cells. To do so, we employ a variant of a region based multi-ellipse fitting method that makes use of constraints on the area of the split cells. The quantitative assessment of the proposed method has been based on two challenging public datasets. This experimental study shows that the proposed method outperforms clearly existing methods for segmenting fluorescence microscopy images.
Costas Panagiotakis, Antonis A. Argyros
ICIP2
2018 Unsupervised Detection of Periodic Segments in Videos
abstract
We present a solution to the problem of discovering all periodic segments of a video and of estimating their period in a completely unsupervised manner. These segments may be located anywhere in the video, may differ in duration, speed, period and may represent unseen motion patterns of any type of objects (e.g., humans, animals, machines, etc). The proposed method capitalizes on earlier research on the problem of detecting common actions in videos, also known as commonality detection or video co-segmentation. The proposed method has been evaluated quantitatively and in comparison to a baseline, power-spectrum-based approach, on two ground-truth-annotated datasets (MHAD202-v, PERTUBE). From those, PERTUBE has been compiled specifically for the purposes of this study and includes a collection of you tube videos that have been shot in the wild, with several periodic segments. The results of this evaluation demonstrate that the propose method outperforms the baseline considerably, especially in the more challenging PERTUBE dataset.
Costas Panagiotakis, Giorgos Karvounas, Antonis A. Argyros
ICIP3
2018 Distributed Real-Time Generative 3D Hand Tracking using Edge GPGPU Acceleration
abstract
This work demonstrates a real-time 3D hand tracking application that runs via computation offloading. The proposed framework enables the application to run on low-end mobile devices such as laptops and tablets, despite the fact that they lack the sufficient hardware to perform the required computations locally. The network connection takes the place of a GPGPU accelerator and sharing resources with a larger workstation becomes the acceleration mechanism. The unique properties of a generative optimizer are examined and constitute a challenging use-case, since the requirement for real-time performance makes it very latency-sensitive.
Ammar Qammaz, Sokol Kosta, Nikolaos Kyriazis, Antonis A. Argyros
MobiSys4
2018 Using a Single RGB Frame for Real Time 3D Hand Pose Estimation in the Wild
abstract
We present a method for the real-time estimation of the full 3D pose of one or more human hands using a single commodity RGB camera. Recent work in the area has displayed impressive progress using RGBD input. However, since the introduction of RGBD sensors, there has been little progress for the case of monocular color input. We capitalize on the latest advancements of deep learning, combining them with the power of generative hand pose estimation techniques to achieve real-time monocular 3D hand pose estimation in unrestricted scenarios. More specifically, given an RGB image and the relevant camera calibration information, we employ a state-of-the-art detector to localize hands. Given a crop of a hand in the image, we run the pretrained network of OpenPose for hands to estimate the 2D location of hand joints. Finally, non-linear least-squares minimization fits a 3D model of the hand to the estimated 2D joint positions, recovering the 3D hand pose. Extensive experimental results provide comparison to the state of the art as well as qualitative assessment of the method in the wild.
Paschalis Panteleris, Iasonas Oikonomidis, Antonis A. Argyros
WACV3
2018 A Hybrid Method for 3D Pose Estimation of Personalized Human Body Models
abstract
We propose a new hybrid method for 3D human body pose estimation based on RGBD data. We treat this as an optimization problem that is solved using a stochastic optimization technique. The solution to the optimization problem is the pose parameters of a human model that register it to the available observations. Our method can make use of any skinned, articulated human body model. However, we focus on personalized models that can be acquired easily and automatically based on existing human scanning and mesh rigging techniques. Observations consist of the 3D structure of the human (measured by the RGBD camera) and the body joints locations (computed based on a dis-criminative, CNN-based component). A series of quantitative and qualitative experiments demonstrate the accuracy and the benefits of the proposed approach. In particular, we show that the proposed approach achieves state of the art results compared to competitive methods and that the use of personalized body models improve significantly the accuracy in 3D human pose estimation.
Ammar Qammaz, Damien Michel, Antonis A. Argyros
WACV3
2018 Hand-Object Contact Force Estimation from Markerless Visual Tracking
abstract
We consider the problem of estimating realistic contact forces during manipulation, backed with ground-truth measurements, using vision alone. Interaction forces are usually measured by mounting force transducers onto the manipulated objects or the hands. Those are costly, cumbersome, and alter the objects' physical properties and their perception by the human sense of touch. Our work establishes that interaction forces can be estimated in a cost-effective, reliable, non-intrusive way using vision. This is a complex and challenging problem. Indeed, in multi-contact, a given motion can generally be caused by an infinity of possible force distributions. To alleviate the limitations of traditional models based on inverse optimization, we collect and release the first large-scale dataset on manipulation kinodynamics as 3.2 hours of synchronized force and motion measurements under 193 object-grasp configurations. We learn a mapping between high-level kinematic features based on the equations of motion and the underlying manipulation forces using recurrent neural networks (RNN). The RNN predictions are consistently refined using physics-based optimization through second-order cone programming (SOCP). We show that our method can successfully capture interaction forces compatible with both the observations and the way humans intuitively manipulate objects, using a single RGB-D camera.
Tu-Hoa Pham, Nikolaos Kyriazis, Antonis A. Argyros, Abderrahmane Kheddar
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 A graph-based approach for detecting common actions in motion capture data and videos
Costas Panagiotakis, Konstantinos E. Papoutsakis, Antonis A. Argyros
Pattern Recognit.3
2017 Generative 3D Hand Tracking with Spatially Constrained Pose Sampling
Konstantinos Roditakis, Alexandros Makris, Antonis A. Argyros
BMVC3
2017 Temporal Action Co-Segmentation in 3D Motion Capture Data and Videos
abstract
Given two action sequences, we are interested in spotting/co-segmenting all pairs of sub-sequences that represent the same action. We propose a totally unsupervised solution to this problem. No a-priori model of the actions is assumed to be available. The number of common sub-sequences may be unknown. The sub-sequences can be located anywhere in the original sequences, may differ in duration and the corresponding actions may be performed by a different person, in different style. We treat this type of temporal action co-segmentation as a stochastic optimization problem that is solved by employing Particle Swarm Optimization (PSO). The objective function that is minimized by PSO capitalizes on Dynamic Time Warping (DTW) to compare two action sub-sequences. Due to the generic problem formulation and solution, the proposed method can be applied to motion capture (i.e., 3D skeletal) data or to conventional RGB videos acquired in the wild. We present extensive quantitative experiments on standard data sets as well as on data sets we introduced in this paper. The obtained results demonstrate that the proposed method achieves a remarkable increase in co-segmentation quality compared to all tested state of the art methods.
Konstantinos E. Papoutsakis, Costas Panagiotakis, Antonis A. Argyros
CVPR3
2017 A Framework for Online Segmentation and Classification of Modeled Actions Performed in the Context of Unmodeled Ones
abstract
In this paper, we propose a discriminative framework for online simultaneous segmentation and classification of modeled visual actions that can be performed in the context of other unknown actions. To this end, we employ Hough transform to vote in a 3D space for the begin point, the end point, and the label of the segmented part of the input stream. A support vector machine is used to model each class and to suggest putative labeled segments on the timeline. To identify the most plausible segments among the putative ones, we apply a dynamic programming algorithm, which maximizes the likelihood for label assignment in linear time. The performance of our method is evaluated on synthetic as well as on real data (Weizmann, TUM Kitchen, UTKAD, and Berkeley Multimodal Human Action databases). Extensive quantitative results obtained on a number of standard data sets demonstrate that the proposed approach is of comparable accuracy with the state-of-the-art approaches for online stream segmentation and classification when all performed actions are known, and performs considerably better in the presence of unmodeled actions.
Dimitrios I. Kosmopoulos, Konstantinos E. Papoutsakis, Antonis A. Argyros
IEEE Trans. Circuits Syst. Video Technol.3
2016 Tracking Deformable Surfaces That Undergo Topological Changes Using an RGB-D Camera
abstract
We present a method for 3D tracking of deformable surfaces with dynamic topology, for instance a paper that undergoes cutting or tearing. Existing template-based methods assume a template of fixed topology. Thus, they fail in tracking deformable objects that undergo topological changes. In our work, we employ a dynamic template (3D mesh) whose topology evolves based on the topological changes of the observed geometry. Our tracking framework deforms the defined template based on three types of constraints: (a) the surface of the template has to be registered to the 3D shape of the tracked surface, (b) the template deformation should respect feature (SIFT) correspondences between selected pairs of frames, and (c) the lengths of the template edges should be preserved. The latter constraint is relaxed when an edge is found to lie on a "geometric gap", that is, when a significant depth discontinuity is detected along this edge. The topology of the template is updated on the fly by removing overstretched edges that lie on a geometric gap. The proposed method has been evaluated quantitatively and qualitatively in both synthetic and real sequences of monocular RGB-D views of surfaces that undergo various types of topological changes. The obtained results show that our approach tracks effectively objects with evolving topology and outperforms state of the art methods in tracking accuracy.
Aggeliki Tsoli, Antonis A. Argyros
3DV2
2016 Localizing Periodicity in Time Series and Videos
Giorgos Karvounas, Iasonas Oikonomidis, Antonis A. Argyros
BMVC3
2016 Parameter-free modelling of 2D shapes with ellipses
Costas Panagiotakis, Antonis A. Argyros
Pattern Recognit.2
2015 Model-based 3D Hand Tracking with on-line Shape Adaptation
abstract
One of the shortcomings of the existing model-based 3D hand tracking methods is the fact that they consider a fixed hand model, i.e. one with fixed shape parameters. In this work we propose an online model-based method that tackles jointly the hand pose tracking and the hand shape estimation problems. The hand pose is estimated using a hierarchical particle filter. The hand shape is estimated by fitting the shape model parameters over the observations in a frame history. The candidate shapes required by the fitting framework are obtained by optimizing the shape parameters independently in each frame. Extensive experiments demonstrate that the proposed method tracks the pose of the hand and estimates its shape parameters accurately, even under heavy noise and inaccurate shape initialization.
Alexandros Makris, Antonis A. Argyros
BMVC2
2015 3D Tracking of Human Hands in Interaction with Unknown Objects
abstract
The analysis and the understanding of object manipulation scenarios based on computer vision techniques can be greatly facilitated if we can gain access to the full articulation of the manipulating hands and the 3D pose of the manipulated objects. Currently, there exist methods for tracking hands in interaction with objects whose 3D models are known [2]. There are also methods that can reconstruct 3D models of objects that are partially observable in each frame of a sequence [3]. However, no method can track hands in interaction with unknown objects, ie objects whose 3D model is not known a priori. In this paper we propose a novel approach that can track human hands in interaction with unknown objects. As illustrated in Fig.1, the input to the method is a sequence of RGBD frames showing the interaction of one or two hands with an unknown object. Starting with the raw depth map (left) we perform a pre-processing step and compute the scene point cloud. We employ an appropriately modified model based hand tracker [4] and temporal information to track the hand 3D positions and posture (middle bottom). In this process, a progressively built object model is also taken into account to cope with hand-object occlusions. We use the estimated fingertip positions of the hand to segment the manipulated object from the rest of the scene (middle top). The segmented object points are used to update the object position and orientation in the current frame and are integrated into the object 3D representation (right). More specifically, the work flow of the proposed approach consists of five main components linked together as shown in Fig. 2. At a first, preprocessing stage, the raw depth information from the sensor is prepared to enter the pipeline. A point cloud is computed along with the normals for each vertex. Then, the user’s hands are tracked in the scene. An articulated model for the left and right hands, with 26 degrees of freedom each, is fit to the pre-processed depth input. The current, possibly incomplete (or even empty, for the first frame) object model is incorporated to hand tracking to assist in handling hand/object occlusions. Using the computed 3D location of the user’s hands as well as the last position of the (possibly incomplete) object model, the region of the object is segmented in the input depth map. The hands are masked-out from the observation, by comparing it to the rendered hand models. Object tracking is achieved using a mutli-scale ICP [1]. The segmented object depth is used for a coarse to fine alignment with the (partially reconstructed) object model. Finally, the segmented and aligned depth data of the object with the current, partial 3D model are merged. The object’s 3D model is maintained in a voxel grid with a Truncated Signed Distance Function (TSDF) [3] representation. Experiment Proposed [2], GT model [2], Scanned model mean/median error mean/median error mean/median error Single hand, cat 0.42 / 0.39 0.47 / 0.43 0.45 / 0.43 Single hand, spray 0.65 / 0.63 0.70 / 0.53 0.63 / 0.47 Two hands, cat 0.38 / 0.34 0.33 / 0.31 0.44 / 0.39 Two hands, spray 0.59 / 0.44 0.51 / 0.38 0.62 / 0.41 Table 1: Hand tracking accuracy (in cm) measured on the synthetic datasets. The accuracy of the method is close to that of [2], although the latter assumes that the object model is known a priori.
Paschalis Panteleris, Nikolaos Kyriazis, Antonis A. Argyros
BMVC3
2015 Hybrid One-Shot 3D Hand Pose Estimation by Exploiting Uncertainties
abstract
Model-based approaches to 3D hand tracking have been shown to perform well in a wide range of scenarios. However, they require initialisation and cannot recover easily from tracking failures that occur due to fast hand motions. Data-driven approaches, on the other hand, can quickly deliver a solution, but the results often suffer from lower accuracy or missing anatomical validity compared to those obtained from model-based approaches. In this work we propose a hybrid approach for hand pose estimation from a single depth image. First, a learned regressor is employed to deliver multiple initial hypotheses for the 3D position of each hand joint. Subsequently, the kinematic parameters of a 3D hand model are found by deliberately exploiting the inherent uncertainty of the inferred joint proposals. This way, the method provides anatomically valid and accurate solutions without requiring manual initialisation or suffering from track losses. Quantitative results on several standard datasets demonstrate that the proposed method outperforms state-of-the-art representatives of the model-based, data-driven and hybrid paradigms.
Georg Poier, Konstantinos Roditakis, Samuel Schulter, Damien Michel, Horst Bischof, Antonis A. Argyros
BMVC6
2015 Boosting the Performance of Model-based 3D Tracking by Employing Low Level Motion Cues
abstract
3D tracking of objects and hands in an object manipulation scenario is a very interesting computer vision problem with a wide variety of applications ranging from consumer electronics to robotics and medicine. Recent advances in this research topic allow for 3D tracking of complex scenarios involving bi-manual manipulation of several rigid objects using commodity hardware and with high accuracy. The problem with these approaches is that they treat tracking as a search problem whose dimensionality increases with the number of objects in the scene. This fact typically limits the number of the tracked objects and/or the processing framerate. In this paper we present a method that utilizes simple low level motion cues for dynamically assigning computational resources to parts of the scene where they are actually required. In a series of experiments, we show that this simple idea improves tracking performance dramatically at a cost of only a minor degradation of tracking accuracy. The works that are most related to ours are the approaches by Kyriazis and Argyros on top-down 3D tracking of multiple active objects from RGBD input [1, 2]. The methodological part of our contribution can be briefly described as an extra processing node in the pipeline of [2] which it extends. The tracking approach in [2], the Ensemble of Collaborative Trackers (ECT), regards a set of semi-independent trackers. Each tracker is associated with a distinct object in the scene. For an object to be tracked, a separate optimization problem is solved, one for each frame and each object. Each optimization problem is numerically solved using a black box optimizer, i.e. a variant of the Particle Swarm Optimization (PSO) algorithm. PSO treats the objective function as an oracle and queries it on purposefully evolved “guesses” in the search space, in order to find the optimum. The harder the problem is, the more guesses (budget, which is the product of PSO particles and PSO generations that need to be computed) are required for adequate accuracy to be achieved. For example, tracking the pose and the articulation of the hand amounts to solving a 27-parameter optimization problem. This is much harder than solving for the 3D pose of a rigid object (6 parameters). ECT [2] assigns a fixed amount of computational resources to each tracking sub-problem which depends on this notion of complexity and which is empirically estimated. In this work, we quantify at run time, how hard the tracking of an object should be, not only based on its intrinsic complexity, but also based on its observed dynamics. An object that appears static in the recent temporal window requires less resources compared to an object whose state is more dynamic. Thus, change detection is performed on the image space of color intensities and depth measurements of the RGBD input. In more detail, for the next tracking frame, and the ith tracker, a value mi is computed, which takes the value of 0 if the corresponding tracked object appears to be relatively static, and takes the value of 1 otherwise. For mi to be 1, any of the following need to be true: • the mean value of the pixel-wise differences between the observations of object i and the back-projection of the last estimated configuration of object i is high enough, • the kinetic energy of object i, as computed from so far tracked velocities, suggests a moving object, • the amount of missing depth measurements changes substantially, from one frame to the next, which might me attributed to change in the slant of object i, or • any of the above was satisfied in the recent past (damping). Budget has the trivial minimum value of 0. Moreover, as it has been shown in [3, 4], a PSO budget of 64 particles and 64 generations suffices to track a hand, even in interaction with another hand. No more budget should be required to track simpler structures such as rigid objects. The proposed dynamic budget allocation policy assigns a minimum budget Bmin of 64 particles running for 4 generations to the objects that are static and a maximum budget Bmax that never exceeds the aforementioned (a) Pouring pancake mix (b) Waiting for the mix to cook
Ammar Qammaz, Nikolaos Kyriazis, Antonis A. Argyros
BMVC3
2015 Towards force sensing from vision: Observing hand-object interactions to infer manipulation forces
abstract
We present a novel, non-intrusive approach for estimating contact forces during hand-object interactions relying solely on visual input provided by a single RGB-D camera. We consider a manipulated object with known geometrical and physical properties. First, we rely on model-based visual tracking to estimate the object's pose together with that of the hand manipulating it throughout the motion. Following this, we compute the object's first and second order kinematics using a new class of numerical differentiation operators. The estimated kinematics is then instantly fed into a second-order cone program that returns a minimal force distribution explaining the observed motion. However, humans typically apply more forces than mechanically required when manipulating objects. Thus, we complete our estimation method by learning these excessive forces and their distribution among the fingers in contact. We provide a full validity analysis of the proposed method by evaluating it based on ground truth data from additional sensors such as accelerometers, gyroscopes and pressure sensors. Experimental results show that force sensing from vision (FSV) is indeed feasible.
Tu-Hoa Pham, Abderrahmane Kheddar, Ammar Qammaz, Antonis A. Argyros
CVPR4
2015 Quantifying the Effect of a Colored Glove in the 3D Tracking of a Human Hand
Konstantinos Roditakis, Antonis A. Argyros
ICVS2
2015 Tracking the articulated motion of the human body with two RGBD cameras
Damien Michel, Costas Panagiotakis, Antonis A. Argyros
Mach. Vis. Appl.3
2014 Segmentation and classification of modeled actions in the context of unmodeled ones
Dimitrios I. Kosmopoulos, Konstantinos E. Papoutsakis, Antonis A. Argyros
BMVC3
2014 Scalable 3D Tracking of Multiple Interacting Objects
abstract
We consider the problem of tracking multiple interacting objects in 3D, using RGBD input and by considering a hypothesize-and-test approach. Due to their interaction, objects to be tracked are expected to occlude each other in the field of view of the camera observing them. A naive approach would be to employ a Set of Independent Trackers (SIT) and to assign one tracker to each object. This approach scales well with the number of objects but fails as occlusions become stronger due to their disjoint consideration. The solution representing the current state of the art employs a single Joint Tracker (JT) that accounts for all objects simultaneously. This directly resolves ambiguities due to occlusions but has a computational complexity that grows geometrically with the number of tracked objects. We propose a middle ground, namely an Ensemble of Collaborative Trackers (ECT), that combines best traits from both worlds to deliver a practical and accurate solution to the multi-object 3D tracking problem. We present quantitative and qualitative experiments with several synthetic and real world sequences of diverse complexity. Experiments demonstrate that ECT manages to track far more complex scenes than JT at a computational time that is only slightly larger than that of SIT.
Nikolaos Kyriazis, Antonis A. Argyros
CVPR2
2014 Evolutionary Quasi-Random Search for Hand Articulations Tracking
abstract
We present a new method for tracking the 3D position, global orientation and full articulation of human hands. Following recent advances in model-based, hypothesize-and-test methods, the high-dimensional parameter space of hand configurations is explored with a novel evolutionary optimization technique specifically tailored to the problem. The proposed method capitalizes on the fact that samples from quasi-random sequences such as the Sobol have low discrepancy and exhibit a more uniform coverage of the sampled space compared to random samples obtained from the uniform distribution. The method has been tested for the problems of tracking the articulation of a single hand (27D parameter space) and two hands (54D space). Extensive experiments have been carried out with synthetic and real data, in comparison with state of the art methods. The quantitative evaluation shows that for cases of limited computational resources, the new approach achieves a speed-up of four (single hand tracking) and eight (two hands tracking) without compromising tracking accuracy. Interestingly, the proposed method is preferable compared to the state of the art either in the case of limited computational resources or in the case of more complex (i.e., higher dimensional) problems, thus improving the applicability of the method in a number of application domains.
Iasonas Oikonomidis, Manolis I. A. Lourakis, Antonis A. Argyros
CVPR3
2014 Temporal Segmentation and Seamless Stitching of Motion Patterns for Synthesizing Novel Animations of Periodic Dances
abstract
In this paper, we present an efficient algorithm for synthesizing novel, arbitrarily long animations of periodic dances. The input to the proposed method is motion capture data acquired from markeless visual observations of a human performing a periodic dance. The provided human motion capture data are temporally segmented into the constituent periodic motion patterns. These are further organized in a motion graph that also represents possible transitions among them. Finally, an efficient algorithm exploits this representation to come up with a previously unseen sequence of motion patterns that are stitched seamlessly into a novel, realistic dance animation. Several experiments have been conducted with real recordings of Greek folk dances. The obtained results are very promising and indicate the efficacy of the proposed approach, as well as its tolerance to dynamic and noisy human motion capture input.
Costas Panagiotakis, Antonis A. Argyros, Damien Michel
ICPR2
2014 Shape from interaction
Damien Michel, Xenophon Zabulis, Antonis A. Argyros
Mach. Vis. Appl.3
2013 Physically Plausible 3D Scene Tracking: The Single Actor Hypothesis
abstract
In several hand-object(s) interaction scenarios, the change in the objects' state is a direct consequence of the hand's motion. This has a straightforward representation in Newtonian dynamics. We present the first approach that exploits this observation to perform model-based 3D tracking of a table-top scene comprising passive objects and an active hand. Our forward modelling of 3D hand-object(s) interaction regards both the appearance and the physical state of the scene and is parameterized over the hand motion (26 DoFs) between two successive instants in time. We demonstrate that our approach manages to track the 3D pose of all objects and the 3D pose and articulation of the hand by only searching for the parameters of the hand motion. In the proposed framework, covert scene state is inferred by connecting it to the overt state, through the incorporation of physics. Thus, our tracking approach treats a variety of challenging observability issues in a principled manner, without the need to resort to heuristics.
Nikolaos Kyriazis, Antonis A. Argyros
CVPR2
2013 Multicamera tracking of multiple humans based on colored visual hulls
abstract
Detecting, localizing and tracking humans within an industrial environment are three tasks which are of central importance towards achieving automation in workplaces and intelligent environments. This is because unobtrusive, real-time and reliable person tracking provides valuable input to solving problems such as workplace surveillance and event/activity recognition and, also, contributes to safety and optimized use of resources. This paper presents a passive approach to the problem of person tracking that is based on a network of conventional color cameras. The proposed approach exhibits robustness to challenging conditions that are encountered in industrial environments due to illumination artifacts, occlusions and the highly dynamic nature of the observed scenes. The multiple views of the environment that the system employs are used to obtain a volumetric representation of the humans within it, in real-time. Although human tracking can be achieved based solely on such a volumetric representation, in demanding scenes, this information is not enough to recover from tracking failures. Thus, in this work, we collect and update a representation of the color appearance of the persons in the environment. The combination of volumetric and color information reinforces tracking robustness, even when a person is not visible by any of the cameras for extended time intervals. The proposed approach has been extensively evaluated in comparison with an existing state of the art method and pertinent results are reported.
Pashalis Padeleris, Xenophon Zabulis, Antonis A. Argyros
ETFA3
2013 ChaLearn multi-modal gesture recognition 2013: grand challenge and workshop summary
abstract
We organized a Grand Challenge and Workshop on Multi-Modal Gesture Recognition.
Sergio Escalera, Jordi Gonzàlez 0001, Xavier Baró, Miguel Reyes, Isabelle Guyon, Vassilis Athitsos, Hugo Jair Escalante, Leonid Sigal, Antonis A. Argyros, Cristian Sminchisescu, Richard Bowden, Stan Sclaroff
ICMI9
2013 Language for learning complex human-object interactions
abstract
In this paper we use a Hierarchical Hidden Markov Model (HHMM) to represent and learn complex activities/task performed by humans/robots in everyday life. Action primitives are used as a grammar to represent complex human behaviour and learn the interactions and behaviour of human/robots with different objects. The main contribution is the use of a probabilistic model capable of representing behaviours at multiple levels of abstraction to support the proposed hypothesis. The hierarchical nature of the model allows decomposition of the complex task into simple action primitives. The framework is evaluated with data collected for tasks of everyday importance performed by a human user.
Carl Henrik Ek, Nikolaos Kyriazis, Antonis A. Argyros, Jaime Valls Miró, Danica Kragic
ICRA4
2013 Predicting human intention in visual observations of hand/object interactions
abstract
The main contribution of this paper is a probabilistic method for predicting human manipulation intention from image sequences of human-object interaction. Predicting intention amounts to inferring the imminent manipulation task when human hand is observed to have stably grasped the object. Inference is performed by means of a probabilistic graphical model that encodes object grasping tasks over the 3D state of the observed scene. The 3D state is extracted from RGB-D image sequences by a novel vision-based, markerless hand-object 3D tracking framework. To deal with the high-dimensional state-space and mixed data types (discrete and continuous) involved in grasping tasks, we introduce a generative vector quantization method using mixture models and self-organizing maps. This yields a compact model for encoding of grasping actions, able of handling uncertain and partial sensory data. Experimentation showed that the model trained on simulated data can provide a potent basis for accurate goal-inference with partial and noisy observations of actual real-world demonstrations. We also show a grasp selection process, guided by the inferred human intention, to illustrate the use of the system for goal-directed grasp imitation.
Dan Song 0002, Nikolaos Kyriazis, Iasonas Oikonomidis, Chavdar Papazov, Antonis A. Argyros, Darius Burschka, Danica Kragic
ICRA5
2013 Dimensionality Reduction for Efficient Single Frame Hand Pose Estimation
Petros Douvantzis, Iasonas Oikonomidis, Nikolaos Kyriazis, Antonis A. Argyros
ICVS4
2013 Integrating tracking with fine object segmentation
Konstantinos E. Papoutsakis, Antonis A. Argyros
Image Vis. Comput.2
2013 Multicamera human detection and tracking supporting natural interaction with large-scale displays
Xenophon Zabulis, Dimitris Grammenos, Thomas Sarmis, Konstantinos Tzevanidis, Pashalis Padeleris, Panayotis Koutlemanis, Antonis A. Argyros
Mach. Vis. Appl.7
2012 Tracking the articulated motion of two strongly interacting hands
abstract
We propose a method that relies on markerless visual observations to track the full articulation of two hands that interact with each-other in a complex, unconstrained manner. We formulate this as an optimization problem whose 54-dimensional parameter space represents all possible configurations of two hands, each represented as a kinematic structure with 26 Degrees of Freedom (DoFs). To solve this problem, we employ Particle Swarm Optimization (PSO), an evolutionary, stochastic optimization method with the objective of finding the two-hands configuration that best explains observations provided by an RGB-D sensor. To the best of our knowledge, the proposed method is the first to attempt and achieve the articulated motion tracking of two strongly interacting hands. Extensive quantitative and qualitative experiments with simulated and real world image sequences demonstrate that an accurate and efficient solution of this problem is indeed feasible.
Iasonas Oikonomidis, Nikolaos Kyriazis, Antonis A. Argyros
CVPR3
2011 Binding Computer Vision to Physics Based Simulation: The Case Study of a Bouncing Ball
Nikolaos Kyriazis, Iasonas Oikonomidis, Antonis A. Argyros
BMVC3
2011 Efficient model-based 3D tracking of hand articulations using Kinect
abstract
We present a novel solution to the problem of recovering and tracking the 3D position, orientation and full articulation of a human hand from markerless visual observations obtained by a Kinect sensor. We treat this as an optimization problem, seeking for the hand model parameters that minimize the discrepancy between the appearance and 3D structure of hypothesized instances of a hand model and actual hand observations. This optimization problem is effectively solved using a variant of Particle Swarm Optimization (PSO). The proposed method does not require special markers and/or a complex image acquisition setup. Being model based, it provides continuous solutions to the problem of tracking hand articulations. Extensive experiments with a prototype GPU-based implementation of the proposed method demonstrate that accurate and robust 3D tracking of hand articulations can be achieved in near real-time (15Hz).
Iasonas Oikonomidis, Nikolaos Kyriazis, Antonis A. Argyros
BMVC3
2011 Full DOF tracking of a hand interacting with an object by modeling occlusions and physical constraints
abstract
Due to occlusions, the estimation of the full pose of a human hand interacting with an object is much more challenging than pose recovery of a hand observed in isolation. In this work we formulate an optimization problem whose solution is the 26-DOF hand pose together with the pose and model parameters of the manipulated object. Optimization seeks for the joint hand-object model that (a) best explains the incompleteness of observations resulting from occlusions due to hand-object interaction and (b) is physically plausible in the sense that the hand does not share the same physical space with the object. The proposed method is the first that solves efficiently the continuous, full-DOF, joint hand-object tracking problem based solely on markerless multicamera input. Additionally, it is the first to demonstrate how hand-object interaction can be exploited as a context that facilitates hand pose estimation, instead of being considered as a complicating factor. Extensive quantitative and qualitative experiments with simulated and real world image sequences as well as a comparative evaluation with a state-of-the-art method for pose estimation of isolated hands, support the above findings.
Iasonas Oikonomidis, Nikolaos Kyriazis, Antonis A. Argyros
ICCV3
2011 PaperView: augmenting physical surfaces with location-aware digital information
abstract
A frequent need of museums is to provide visitors with context-sensitive information about exhibits in the form of maps, or scale models. This paper suggests an augmented-reality approach for supplementing physical surfaces with digital information, through the use of pieces of plain paper that act as personal, location-aware, interactive screens. The technologies employed are presented, along with the interactive behavior of the system, which was instantiated and tested in the form of two prototype setups: a wooden table covered with a printed map and a glass case containing a scale model. The paper also discusses key issues stemming from experience and observations in the course of qualitative evaluation sessions.
Dimitris Grammenos, Damien Michel, Xenophon Zabulis, Antonis A. Argyros
TEI4
2011 Unsupervised learning of background modeling parameters in multicamera systems
Konstantinos Tzevanidis, Antonis A. Argyros
Comput. Vis. Image Underst.2
2011 Scale invariant and deformation tolerant partial shape matching
Damien Michel, Iasonas Oikonomidis, Antonis A. Argyros
Image Vis. Comput.3
2010 Markerless and Efficient 26-DOF Hand Pose Recovery
Iasonas Oikonomidis, Nikolaos Kyriazis, Antonis A. Argyros
ACCV (3)3
2010 Horizon matching for localizing unordered panoramic images
Damien Michel, Antonis A. Argyros, Manolis I. A. Lourakis
Comput. Vis. Image Underst.2
2010 Multiple objects tracking in the presence of long-term occlusions
Vasilis Papadourakis, Antonis A. Argyros
Comput. Vis. Image Underst.2
2009 3D Head Pose Estimation from Multiple Distant Views
abstract
3D head pose estimation constitutes a special problem of human motion modeling. An accurate and robust solution to this problem is of particular interest, because the 3D head pose of a human conveys important information on his/her behavior. Significant advances have been achieved in human head pose estimation for relatively close-range images, but the related available methods are not directly applicable in wider-range imaging conditions. The proposed method is overviewed in Fig. 1. The visual hull of a person is obtained from images acquired synchronously from multiple viewpoints. While moving, the person’s head is tracked in 3D employing a variant of the Mean-Shift algorithm and a spherical kernel. The texture on the surface of the hull is collected from multiple views and projected on a hypothetical sphere S that is concentric to the person’s head. This gives rise to spherical image Is within which face detection is simplified, because exactly one frontal face is guaranteed to appear in it at a known spatial scale.
Xenophon Zabulis, Thomas Sarmis, Antonis A. Argyros
BMVC3
2009 Visual homing for undulatory robotic locomotion
abstract
This paper addresses the problem of vision-based closed-loop control for undulatory robots. We present an image-based visual servoing scheme, which drives the robot to a desired location specified by a target image, without explicitly estimating its pose. Instead, the control relies on the computation of the epipolar geometry between the current and target images. We analyze controllability and stability of the proposed control scheme, which is validated by simulation studies using the SIMUUN computational tools. Preliminary experiments, involving the Nereisbot undulatory robotic prototype, are also presented.
Gonzalo López-Nicolás, Michael Sfakiotakis, Dimitris P. Tsakiris, Antonis A. Argyros, Carlos Sagüés, Josechu J. Guerrero
ICRA4
2009 Integrated vision system for the semantic interpretation of activities where a person handles objects
Markus Vincze, Michael Zillich, Wolfgang Ponweiser, Václav Hlavác, Jiri Matas, Stepán Obdrzálek, Hilary Buxton, A. Jonathan Howell, Kingsley Sage, Antonis A. Argyros, Christof Eberst, Gerald Umgeher
Comput. Vis. Image Underst.10
2009 SBA: A software package for generic sparse bundle adjustment
abstract
Bundle adjustment constitutes a large, nonlinear least-squares problem that is often solved as the last step of feature-based structure and motion estimation computer vision algorithms to obtain optimal estimates. Due to the very large number of parameters involved, a general purpose least-squares algorithm incurs high computational and memory storage costs when applied to bundle adjustment. Fortunately, the lack of interaction among certain subgroups of parameters results in the corresponding Jacobian being sparse, a fact that can be exploited to achieve considerable computational savings. This article presents sba, a publicly available C/C++ software package for realizing generic bundle adjustment with high efficiency and flexibility regarding parameterization.
Manolis I. A. Lourakis, Antonis A. Argyros
ACM Trans. Math. Softw.2
2008 Shading models for illumination and reflectance invariant shape detectors
abstract
Many objects have smooth surfaces of a fairly uniform color, thereby exhibiting shading patterns that reveal information about its shape, an important clue to the nature of the object. This papers explores extracting this information from images, by creating shape detectors based on shading.
Peter Nillius, Josephine Sullivan, Antonis A. Argyros
CVPR3
2008 Dynamic time warping for binocular hand tracking and reconstruction
abstract
We show how matching and reconstruction of contour points can be performed using dynamic time warping (DTW) for the purpose of 3D hand contour tracking. We evaluate the performance of the proposed algorithm in object manipulation activities and perform comparison with the iterative closest point (ICP) method.
Javier Romero 0002, Danica Kragic, Ville Kyrki, Antonis A. Argyros
ICRA4
2008 Tracking of Human Hands and Faces through Probabilistic Fusion of Multiple Visual Cues
Haris Baltzakis, Antonis A. Argyros, Manolis I. A. Lourakis, Panos E. Trahanias
ICVS2
2008 Lumen detection for capsule endoscopy
abstract
In this paper, two visual cues are proposed, to be exploited for the navigation of active endoscopic capsules within the gastrointestinal (GI) tract. These cues consist of the detection and tracking of the lumen and of an illumination highlight in capsule endoscopy (CE) images. The proposed approach aims at developing vision algorithms which are robust with respect to the challenging imaging conditions encountered in the GI tract and the great variability of the acquired images. Cases where no or more than one lumens exists, are also detected. The proposed approach extends the state-of-the-art in lumen detection, and is demonstrated for in-vivo video sequences acquired from endoscopic capsules.
Xenophon Zabulis, Antonis A. Argyros, Dimitris P. Tsakiris
IROS2
2008 Learning temporal structure for task based control
Kingsley Sage, A. Jonathan Howell, Hilary Buxton, Antonis A. Argyros
Image Vis. Comput.4
2007 Localizing Unordered Panoramic Images Using the Levenshtein Distance
abstract
This paper proposes a feature-based method for recovering the relative positions of the viewpoints of a set of panoramic images for which no a priori order information is available, along with certain structure information regarding the imaged environment. The proposed approach operates incrementally, employing the Levenshtein distance to deduce the spatial proximity of image viewpoints and thus determine the order in which images should be processed. The Levenshtein distance also provides matches between images, from which their underlying environment points can be recovered. Recovered points that are visible in multiple views permit the localization of more views which in turn allow the recovery of more points. The process repeats until all views have been localized. Periodic refinement of the reconstruction with the aid of bundle adjustment, distributes the reconstruction errors among images. The method is demonstrated on several unordered sets of panoramic images obtained in an indoor environment.
Damien Michel, Antonis A. Argyros, Manolis I. A. Lourakis
ICCV2
2005 Tracking Multiple Colored Blobs with a Moving Camera
abstract
This paper concerns a method for tracking multiple blobs exhibiting certain color distributions in images acquired by a possibly moving camera. The method encompasses a collection of techniques that enable modeling and detecting the blobs possessing the desired color distribution(s), as well as inferring their temporal association across image sequences. Appropriately colored blobs are detected with a Bayesian classifier, which is bootstrapped with a small set of training data. Then, an online iterative training procedure is employed to refine the classifier using additional training images. Online adaptation of color probabilities is used to enable the classifier to cope with illumination changes. Tracking over time is realized through a novel technique, which can handle multiple colored blobs. Such blobs may move in complex trajectories and occlude each other in the field of view of a possibly moving camera, while their number may vary over time. A prototype implementation of the developed system running on a conventional Pentium IV processor at 2.5 GHz operates on 320/spl times/240 live video in real time (30Hz). It is worth pointing out that currently, the cycle time of the tracker is determined by the maximum acquisition frame rate that is supported by our IEEE 1394 camera, rather than the latency introduced by the computational overhead for tracking blobs.
Antonis A. Argyros, Manolis I. A. Lourakis
CVPR (2)1
2005 Camera Matchmoving in Unprepared, Unknown Environments
abstract
y are being acquired. Furthermore, it does not rely upon the presence in the environment of fiducial markers or special calibration objects. A brief overview of our approach is given in the next section. More detailed descriptions can be found in [1] and online at http://www. ics.forth.gr/lourakis/camtrack/. THE APPROACH Our method is based on a feature-based 3D plane tracking technique, which permits the estimation of the homographies induced by a 3D plane between successive image pairs. At the core of plane tracking lies a homography "chaining" operation that is applied to triplets of consecutive images through a sliding time window and exploits the constraint that all images of a planar surface acquired by a rigidly moving observer depend upon the same 3D geometry. Since the tracked plane is not required to be physically present in the scene, a virtual one can be used instead. Plane tracking is achieved by matching between images the 2D projections of points from all over the scen
Manolis I. A. Lourakis, Antonis A. Argyros
CVPR (2)2
2005 Is Levenberg-Marquardt the Most Efficient Optimization Algorithm for Implementing Bundle Adjustment?
abstract
In order to obtain optimal 3D structure and viewing parameter estimates, bundle adjustment is often used as the last step of feature-based structure and motion estimation algorithms. Bundle adjustment involves the formulation of a large scale, yet sparse minimization problem, which is traditionally solved using a sparse variant of the Levenberg-Marquardt optimization algorithm that avoids storing and operating on zero entries. This paper argues that considerable computational benefits can be gained by substituting the sparse Levenberg-Marquardt algorithm in the implementation of bundle adjustment with a sparse variant of Powell's dog leg non-linear least squares technique. Detailed comparative experimental results provide strong evidence supporting this claim.
Manolis I. A. Lourakis, Antonis A. Argyros
ICCV2
2005 Fast trifocal tensor estimation using virtual parallax
abstract
We present a computationally efficient method for estimating the trifocal tensor corresponding to three images acquired by a freely moving camera. The proposed method represents projective space through a "plane + parallax" decomposition and employs a novel technique for estimating the homographies induced by a virtual 3D plane between successive image pairs. Knowledge of these homographies allows the corresponding camera projection matrices to be expressed in a common projective frame and, therefore, to be recovered directly. The trifocal tensor can then be recovered in a straightforward manner from the estimated projection matrices. Sample experimental results demonstrate that the method performs considerably faster compared to a state of the art method, without a serious loss in accuracy.
Manolis I. A. Lourakis, Antonis A. Argyros
ICIP (2)2
2005 Efficient, causal camera tracking in unprepared environments
Manolis I. A. Lourakis, Antonis A. Argyros
Comput. Vis. Image Underst.2
2004 Vision-Based Camera Motion Recovery for Augmented Reality
abstract
We address the problem of tracking the 3D position and orientation of a camera, using the images it acquires while moving freely in unmodeled, arbitrary environments. This task has a broad spectrum of useful applications in domains such as augmented reality and video post production. Most of the existing methods for vision-based camera tracking are designed to operate in a batch, off-line mode, assuming that the whole video sequence to be tracked is available before tracking commences. Typically, such methods operate noncausally, processing video frames backwards and forwards in time as they see fit. Furthermore, they resort to optimization in very high dimensional spaces, a process that is computationally intensive. For these reasons, batch methods are inapplicable to tracking in online, time-critical applications such as video see-through augmented reality. This paper puts forward a novel feature-based approach for camera tracking. The proposed approach operates on images continuously as they are acquired, has realistic computational requirements and does not require modifications of the environment. Sample experimental results demonstrating the feasibility of the approach on video images are also provided
Manolis I. A. Lourakis, Antonis A. Argyros
Computer Graphics International2
2004 Real-Time Tracking of Multiple Skin-Colored Objects with a Possibly Moving Camera
Antonis A. Argyros, Manolis I. A. Lourakis
ECCV (3)1
2004 Angle-based Methods for Mobile Robot Navigation: Reaching the Entire Plane
abstract
Popular approaches for mobile robot navigation involve range information and metric maps of the workspace. For many sensors, however, such as cameras and wireless hardware, the angle between two features or beacons is easier to measure. With these sensors' features in mind, we initially present a control law, which allows a robot with an omni-directional sensor to reach a subset of the plane by monitoring the angles of only three landmarks. By analyzing the law's properties, a second law has been developed that reaches the complementary set of points. The two methods are then combined in a path planning framework that reaches any possible goal configuration in a planar obstacle-free workspace with three landmarks. The proposed framework could be used together with other techniques, such as obstacle avoidance and topological maps to improve the efficiency of autonomous navigation. Experiments have been conducted on a robotic platform using a panoramic camera that exhibits the effectiveness and accuracy of the proposed techniques. This work provides evidence that navigational tasks can be performed using only a small number of primitive sensor cues and without the explicit computation of range information.
Kostas E. Bekris, Antonis A. Argyros, Lydia E. Kavraki
ICRA2
2003 Fusion of laser and visual data for robot motion planning and collision avoidance
Haris Baltzakis, Antonis A. Argyros, Panos E. Trahanias
Mach. Vis. Appl.2
2003 Feature Transfer and Matching in Disparate Stereo Views through the Use of Plane Homographies
abstract
Many vision tasks rely upon the identification of sets of corresponding features among different images. This paper presents a method that, given some corresponding features in two stereo images, matches them with features extracted from a second stereo pair captured from a distant viewpoint. The proposed method is based on the assumption that the viewed scene contains two planar surfaces and exploits geometric constraints that are imposed by the existence of these planes to first transfer and then match image features between the two stereo pairs. The resulting scheme handles point and line features in a unified manner and is capable of successfully matching features extracted from stereo pairs that are acquired from considerably different viewpoints. Experimental results are presented, which demonstrate that the performance of the proposed method compares favorably to that of epipolar and tensor-based approaches.
Manolis I. A. Lourakis, Stavros V. Tzurbakis, Antonis A. Argyros, Stelios C. Orphanoudakis
IEEE Trans. Pattern Anal. Mach. Intell.3
2002 Detecting Planes In An Uncalibrated Image Pair
abstract
Plane detection is a prerequisite to a wide variety of vision tasks. This paper proposes a novel method that exploits results from projective geometry to automatically detect planes using two images. Using a set of point and line features that have been matched between images, the method exploits the fact that every pair of a 3D line and a 3D point defines a plane and utilizes an iterative voting scheme for identifying coplanar subsets of the employed feature set. The method does not require camera calibration, circumvents the 3D reconstruction problem, is robust to the existence of mismatched features and is applicable either to stereo or motion sequence images. Sample results from the application of the proposed method to real imagery are also provided. 1
Manolis I. A. Lourakis, Antonis A. Argyros, Stelios C. Orphanoudakis
BMVC2
2002 Fast positioning of limited-visibility guards for the inspection of 2D workspaces
abstract
This paper presents a novel method for deciding the locations of "guards" required to visually inspect a given 2D workspace. The decided guard positions can then be used as control points in the path of a mobile robot that autonomously inspects a workspace. It is assumed that each of the guards (or the mobile robot that visits the guard positions in some order) is equipped with a panoramic camera of 360 degrees field of view. However, the camera has limited visibility, in the sense that it can observe with sufficient detail objects that are not further than a predefined visibility range. The method seeks to efficiently produce solutions that contain the smaller possible number of guards. Experimental results demonstrate that the proposed method is computationally efficient and that, although suboptimal, decides a small number of guards.
Giorgos D. Kazazakis, Antonis A. Argyros
IROS2
2001 Robot Homing based on Corner Tracking in a Sequence of Panoramic Images
abstract
In robotics, homing can be defined as that behavior which enables a robot to return to its initial (home) position, after traveling a certain distance along an arbitrary path. Odometry has traditionally been used for the implementation of such a behavior, but it has been shown to be an unreliable source of information. In this work, a novel method for visual homing is proposed, based on a panoramic camera. As the robot departs from its initial position, it tracks characteristic features of the environment (corners). As soon as homing is activated, the robot selects intermediate target positions on the original path. These intermediate positions (IPs) are then visited sequentially, until the home position is reached. For the robot to move between two consecutive IPs, it is only required to establish correspondence among at least three corners. This correspondence is obtained through a feature tracking mechanism. The proposed homing scheme is based on the extraction of very low-level sensory information, namely the bearing angles of corners, and has been implemented on a robotic platform. Experimental results show that the proposed scheme achieves homing with a remarkable accuracy, which is not affected by the distance traveled by the robot.
Antonis A. Argyros, Kostas E. Bekris, Stelios C. Orphanoudakis
CVPR (2)1
2000 Using Geometric Constraints for Matching Disparate Stereo Views of 3D Scenes Containing Planes
abstract
Several vision tasks rely upon the availability of sets of corresponding features among images. This paper presents a method which, given some corresponding features in two stereo images, addresses the problem of matching them with features extracted from a second stereo pair captured from a distant viewpoint. The proposed method is based on the assumption that the viewed scene contains two planar surfaces and exploits geometric constraints that are imposed by the existence of these planes to predict the location of image features in the second stereo pair. The resulting scheme handles point and line features in a unified manner and is capable of successfully matching features extracted from stereo pairs acquired from considerably different viewpoints. Experimental results from a prototype implementation demonstrate the effectiveness of the approach.
Manolis I. A. Lourakis, Stavros V. Tzurbakis, Antonis A. Argyros, Stelios C. Orphanoudakis
ICPR3
1999 Combining Central and Peripheral Vision for Reactive Robot Navigation
abstract
In this paper we present a new method for vision-based, reactive robot navigation that enables a robot to move in the middle of the free space by exploiting both central and peripheral vision. The robot employs a forward-looking camera for central vision and two side-looking cameras for sensing the periphery of its visual field. The developed method combines the information acquired by this trinocular vision system and produces low-level motor commands that keep the robot in the middle of the free space. The approach follows the purposive vision paradigm in the sense that vision is not studied in isolation but in the context of the behaviors that the system is engaged as well as the environment and the robot's motor capabilities. It is demonstrated that by taking into account these issues, vision processing can be drastically simplified still giving rise to quite complex behaviors. The proposed method does not make strict assumptions about the environment, requires very low level information to be extracted from the images, produces a robust robot behavior and is computationally efficient. Results obtained by bath simulations and from a prototype on-line implementation demonstrate the effectiveness of the method.
Antonis A. Argyros, Fredrik Bergholm
CVPR1
1998 Independent 3D Motion Detection Using Residual Parallax Normal Flow Fields
abstract
This paper considers a specific problem of visual perception of motion, tamely the problem of visual detection of independent 3D motion. Most of the existing techniques for solving this problem rely on restrictive assumptions about the environment, the observer's motion, or both. Moreover, they are based on the computation of a dense optical flow field, which amounts to solving the ill-posed correspondence problem. In this work independent motion detection is formulated as a problem of robust parameter estimation applied to the visual input acquired by a rigidly moving observer. The proposed method automatically selects a planar surface in the scene and the residual planar parallax normal flow field with respect to the motion of this surface is computed at two successive time! instants. The two resulting normal flow fields are then combined in a linear model. The parameters of this model are related to the parameters of self-motion (ego-motion) and their robust estimation leads to a segmentation of the scene based on 3D motion. The method avoids a complete solution to the correspondence problem by selectively matching subsets of image points and by employing normal flow fields. Experimental results demonstrate the effectiveness of the proposed method in detecting independent motion in scenes with large depth variations and unrestricted observer motion.
Manolis I. A. Lourakis, Antonis A. Argyros, Stelios C. Orphanoudakis
ICCV2
1997 Independent 3D Motion Detection Based on Depth Elimination in Normal Flow Fields
abstract
This paper considers a specific problem of visual perception of motion, namely the problem of visual detection of independent 3D motion. Most of the existing techniques for solving this problem rely on restrictive assumptions about the environment, the observer's motion, or both. Moreover they are based on the computation of optical flow, which amounts to solving the ill-posed correspondence problem. In this work, independent motion detection is formulated as robust parameter estimation applied to the visual input acquired by a binocular rigidly moving observer. Depth and motion measurements are combined in a linear model. The parameters of this model are related to the parameters of self-motion (egomotion) and the parameters of the stereoscopic configuration of the observer. The robust estimation of this model leads to a segmentation of the scene based on 3D motion. The method avoids the correspondence problem by employing only normal flow fields. Experimental results demonstrate the effectiveness of this method in detecting independent motion in scenes with large depth variations, without any constraints imposed on observer motion.
Antonis A. Argyros, Stelios C. Orphanoudakis
CVPR1
1997 Navigational support for robotic wheelchair platforms: an approach that combines vision and range sensors
abstract
An approach towards providing advanced navigational support to robotic wheelchair platforms is presented. In order to avoid any modifications to the environment, we propose an approach that employs computer vision techniques which facilitate space perception and navigation. Computer vision has not been introduced to date in rehabilitation robotics, since the former is not mature enough to meet the needs of this sensitive application. However, in the proposed approach, stable techniques are exploited that facilitate reliable, automatic navigation to any point in the visible environment. Preliminary results obtained from its implementation on a laboratory robotic platform indicate its usefulness and flexibility.
Panos E. Trahanias, Manolis I. A. Lourakis, Antonis A. Argyros, Stelios C. Orphanoudakis
ICRA3
1996 Independent 3D Motion Detection through Robust Regression in Depth Layers
abstract
This paper presents a methodology for the detection of objects that move independently of the observer in a 3D dynamic environment. Independent 3D motion detection is formulated as a problem of robust regression applied to visual input acquired by a binocular, rigidly moving observer. The qualitative analysis of images acquired by a parallel stereo configuration yields a segmentation of a scene into depth layers. A depth layer consists of points of the 3D space with almost constant depth from the observer. Robust regression in the form of Least Median of Squares estimation is applied within each depth layer in order to segment the latter into coherently moving regions. Finally, a combination stage is applied across all layers in order to come up with an integrated view of independent motion in the whole 3D scene. In contrast to other existing approaches for independent motion detection which are based on the ill-posed problem of optical flow computation, the proposed methodology relies...
Antonis A. Argyros, Manolis I. A. Lourakis, Panos E. Trahanias, Stelios C. Orphanoudakis
BMVC1
1996 Qualitative detection of 3D motion discontinuities
abstract
This paper presents a method for the detection of objects that move independently of the observer in a 3D dynamic scene, Independent motion detection is achieved through processing of stereoscopic image sequences acquired by a binocular, rigidly moving observer. A weak assumption is made about the observer's motion (egomotion), namely that the direction of the translational and rotational components of egomotion are constant in small image patches. This assumption facilitates the extraction of qualitative information about depth from motion, while additional qualitative depth information is independently computed from image stereo pairs acquired by the binocular vision system. Robust regression in the form of least median of squares estimation is applied within each image patch to test for consistency between the depth functions computed from motion and stereo. Possible inconsistencies signal the presence of independently moving objects. In contrast to other existing approaches for independent motion detection, which are based on the ill-posed problem of optical flow computation, the proposed method relies on normal flow fields for both stereo and motion processing. By exploiting local constraints of qualitative nature, the problem of independent motion detection is approached directly, without relying on a solution to the general structure from motion problem. Experimental results indicate that the proposed method is both effective and robust.
Antonis A. Argyros, Manolis I. A. Lourakis, Panos E. Trahanias, Stelios C. Orphanoudakis
IROS1