Aphrodite Galata

dblp:31/4420 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
3since 2021 · last 2023
0000-0002-9229-7811ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Probabilistic and Bayesian machine learning · 38% Speech recognition and synthesis · 27% Face, body and person analysis · 12%
Computer graphics and multimedia
2 papers
Computer animation and physical simulation · 100%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis
speech-driven animation
0.212013
Visual Speech Synthesis Using a Variable-Order Switching Shared Gaussian Process Dynamical Model · IEEE Trans. Multim. 2013
Computer animation and physical simulation › audio-driven animation
speech-driven animation
0.212013
Visual Speech Synthesis Using a Variable-Order Switching Shared Gaussian Process Dynamical Model · IEEE Trans. Multim. 2013
Computer animation and physical simulation › facial animation
speech-driven facial animation
0.212013
Visual Speech Synthesis Using a Variable-Order Switching Shared Gaussian Process Dynamical Model · IEEE Trans. Multim. 2013
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
mixture model learning
0.112008
Robust estimation of gaussian mixtures from noisy input data · CVPR 2008
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.112008
Robust estimation of gaussian mixtures from noisy input data · CVPR 2008
Computer vision › 3D vision › motion capture
articulated body tracking
0.112007
Real-time Body Tracking Using a Gaussian Process Latent Variable Model · ICCV 2007
Computer vision › Face, body and person analysis › human pose estimation
human pose tracking
0.112007
Real-time Body Tracking Using a Gaussian Process Latent Variable Model · ICCV 2007
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.112007
Real-time Body Tracking Using a Gaussian Process Latent Variable Model · ICCV 2007
Robotics › Motion planning and robot control
robot learning
0.112007
Real-time Body Tracking Using a Gaussian Process Latent Variable Model · ICCV 2007
Data mining
clustering
0.012008
Robust estimation of gaussian mixtures from noisy input data · CVPR 2008
Data mining › clustering › robust clustering
noisy clustering
0.012008
Robust estimation of gaussian mixtures from noisy input data · CVPR 2008

Methods — techniques the papers use, named apart from their topics

switching state space model · 0.3shared gaussian process dynamical model · 0.3forced phonetic alignment · 0.3variable length markov model · 0.2variable-length markov model · 0.2variational bayes · 0.2uncertainty modeling · 0.2gaussian process latent variable model · 0.1EM clustering · 0.1stochastic tracking · 0.0statistical behavior model · 0.0
YearPublicationVenuePosition
2023 Signs of Language: Embodied Sign Language Fingerspelling Acquisition from Demonstrations for Human-Robot Interaction
abstract
Learning fine-grained movements is a challenging topic in robotics, particularly in the context of robotic hands. One specific instance of this challenge is the acquisition of fingerspelling sign language in robots. In this paper, we propose an approach for learning dexterous motor imitation from video examples without additional information. To achieve this, we first build a URDF model of a robotic hand with a single actuator for each joint. We then leverage pre-trained deep vision models to extract the 3D pose of the hand from RGB videos. Next, using state-of-the-art reinforcement learning algorithms for motion imitation (namely, proximal policy optimization and soft actor-critic), we train a policy to reproduce the movement extracted from the demonstrations. We identify the optimal set of hyperparameters for imitation based on a reference motion. Finally, we demonstrate the generalizability of our approach by testing it on six different tasks, corresponding to fingerspelled letters. Our results show that our approach is able to successfully imitate these fine-grained movements without additional information, highlighting its potential for real-world applications in robotics.
Federico Tavella, Aphrodite Galata, Angelo Cangelosi
RO-MAN2
2022 Phonology Recognition in American Sign Language
abstract
Inspired by recent developments in natural language processing, we propose a novel approach to sign language processing based on phonological properties validated by American Sign Language users. By taking advantage of datasets composed of phonological data and people speaking sign language, we use a pretrained deep model based on mesh reconstruction to extract the 3D coordinates of the signers keypoints. Then, we train standard statistical and deep machine learning models in order to assign phonological classes to each temporal sequence of coordinates.Our paper introduces the idea of exploiting the phonological properties manually assigned by sign language users to classify videos of people performing signs by regressing a 3D mesh. We establish a new baseline for this problem based on the statistical distribution of 725 different signs. Our best-performing models achieve a micro-averaged F1-score of 58% for the major location class and 70% for the sign type using statistical and deep learning algorithms, compared to their corresponding baselines of 35% and 39%.
Federico Tavella, Aphrodite Galata, Angelo Cangelosi
ICASSP2
2021 ChoiceNet: CNN learning through choice of multiple feature map representations
abstract
Abstract We introduce a new architecture called ChoiceNet where each layer of the network is highly connected with skip connections and channelwise concatenations. This enables the network to alleviate the problem of vanishing gradients, reduces the number of parameters without sacrificing performance and encourages feature reuse. We evaluate our proposed architecture on three independent tasks: classification, segmentation and facial landmark localisation. For this, we use benchmark datasets such as ImageNet, CIFAR-10, CIFAR-100, SVHN CamVid and 300W.
Farshid Rayhan, Aphrodite Galata, Timothy F. Cootes
Pattern Anal. Appl.2
2020 Not all points are created equal - an anisotropic cost function for facial landmark location
Farshid Rayhan, Aphrodite Galata, Timothy F. Cootes
BMVC2
2020 Hand tracking from monocular RGB with dense semantic labels
abstract
Recent years have seen a renewed interest in RGB-based hand tracking, as opposed to the depth-based tracking that has dominated the field since the introduction of commodity depth cameras. This trend has been driven by the ability of convolutional neural networks to process large quantities of image data. In this paper, we propose an approach to hand tracking that operates on sets of dense semantic labels. A full pipeline for RGB-based hand tracking is presented. This pipeline uses convolutional neural networks to produce a per-pixel semantic map of the scene before optimising the state of a kinematic model according to this semantic map using a tracking algorithm based on Different Evolution (DE). This technique allows us to simultaneously localise the hand in 3D space and recover the pose, and requires only monocular RGB input. We apply our technique to a benchmark dataset, reporting semantic segmentation and 3D pose tracking results, which we compare to the current state of the art. We also compare our DE-based algorithm to an equivalent one based on Particle Swarm Optimisation (PSO) and show that it is superior.
Aphrodite Galata
FG2
2017 3D Hand-Object Pose Estimation from Depth with Convolutional Neural Networks
abstract
Estimating the 3D pose of a hand interacting with an object is a challenging task, harder than hand-only pose estimation as the object can cause heavy occlusion on the hand. We present a two stage discriminative approach using convolutional neural networks (CNN). The first stage classifies and segments the object pixels from a depth image containing the hand and object. This processed image is used to aid the second stage in estimating hand-object pose as it contains information regarding the object location and object occlusion. To the best of our knowledge, this is the first attempt at discriminative one shot hand-object pose estimation. We show that this approach outperforms the current state-of-the-art and that the inclusion of a segmentation stage to learned discriminative single stage systems improves their performance.
Duncan Goudie, Aphrodite Galata
FG2
2013 Mixtures of Gaussian process models for human pose estimation
Martin Fergie, Aphrodite Galata
Image Vis. Comput.2
2013 Visual Speech Synthesis Using a Variable-Order Switching Shared Gaussian Process Dynamical Model
abstract
In this paper, we present a novel approach to speech- driven facial animation using a non-parametric switching state space model based on Gaussian processes. The model is an extension of the shared Gaussian process dynamical model, augmented with switching states. Two talking head corpora are processed by extracting visual and audio data from the sequences followed by a parameterization of both data streams. Phonetic labels are obtained by performing forced phonetic alignment on the audio. The switching states are found using a variable length Markov model trained on the labelled phonetic data. The audio and visual data corresponding to phonemes matching each switching state are extracted and modelled together using a shared Gaussian process dynamical model. We propose a synthesis method that takes into account both previous and future phonetic context, thus accounting for forward and backward coarticulation in speech. Both objective and subjective evaluation results are presented. The quantitative results demonstrate that the proposed method outperforms other state-of-the-art methods in visual speech synthesis and the qualitative results reveal that the synthetic videos are comparable to ground truth in terms of visual perception and intelligibility.
Salil Deena, Shaobo Hou, Aphrodite Galata
IEEE Trans. Multim.3
2012 Dynamical Pose Filtering for Mixtures of Gaussian Processes
abstract
In this paper we propose a novel method for discriminative monocular human pose tracking using a mixture of Gaussian processes and a dynamic programming algorithm for selecting the optimal expert at each frame. The proposed tracking mechanism incorporates a dynamical model into the predictive distribution which is combined with the appearance model in a principled manner. This model is able to give a smoother predicted pose and resolves ambiguities in the image to pose mapping. We introduce a mixture of Gaussian processes model which optimises the size and location of each expert ensuring that each expert models a coherent region of the dataset resulting in an accurate predictive density. We compare our method to other state of the art methods on 2D and 3D monocular pose estimation on ballet and sign language data sets.
Martin Fergie, Aphrodite Galata
BMVC2
2010 Local Gaussian Processes for Pose Recognition from Noisy Inputs
abstract
Gaussian processes have been widely used as a method for inferring the pose of articulated bodies directly from image data. While able to model complex non-linear functions, they are limited due to their inability to model multi-modality caused by ambiguities and varying noise in the data set. For this reason techniques employing mixtures of local Gaussian processes have been proposed to allow multi-modal functions to be predicted accurately [11]. These techniques rely on the calculation of nearest neighbours in the input space to make accurate predictions. However, this becomes a limiting factor when image features are noisy due to changing backgrounds. In this paper we propose a novel method that overcomes this limitation by learning a logistic regression model over the input space to select between the local Gaussian processes. Our proposed method is more robust to a noisy input space than a nearest neighbour approach and provides a better prior over each Gaussian process prediction. Results are demonstrated using synthetic and real data from a sign language data set and HumanEva [9]. 1
Martin Fergie, Aphrodite Galata
BMVC2
2008 Robust estimation of gaussian mixtures from noisy input data
abstract
We propose a variational bayes approach to the problem of robust estimation of gaussian mixtures from noisy input data. The proposed algorithm explicitly takes into account the uncertainty associated with each data point, makes no assumptions about the structure of the covariance matrices and is able to automatically determine the number of the gaussian mixture components. Through the use of both synthetic and real world data examples, we show that by incorporating uncertainty information into the clustering algorithm, we get better results at recovering the true distribution of the training data compared to other variational bayesian clustering algorithms.
Shaobo Hou, Aphrodite Galata
CVPR2
2008 Real-time 3-D human body tracking using learnt models of behaviour
Fabrice Caillette, Aphrodite Galata, Toby Howard
Comput. Vis. Image Underst.2
2007 Real-time Body Tracking Using a Gaussian Process Latent Variable Model
abstract
In this paper, we present a tracking framework for capturing articulated human motions in real-time, without the need for attaching markers onto the subject's body. This is achieved by first obtaining a low dimensional representation of the training motion data, using a nonlinear dimensionality reduction technique called back-constrained GPLVM. A prior dynamics model is then learnt from this low dimensional representation by partitioning the motion sequences into elementary movements using an unsupervised EM clustering algorithm. The temporal dependencies between these elementary movements are efficiently captured by a Variable Length Markov Model. The learnt dynamics model is used to bias the propagation of candidate pose feature vectors in the low dimensional space. By combining this with an efficient volumetric reconstruction algorithm, our framework can quickly evaluate each candidate pose against image evidence captured from multiple views. We present results that show our system can accurately track complex structured activities such as ballet dancing in real-time.
Shaobo Hou, Aphrodite Galata, Fabrice Caillette, Neil A. Thacker, Paul A. Bromiley
ICCV2
2007 A real-time hand tracker using variable-length Markov models of behaviour
Nikolay Stefanov, Aphrodite Galata, Roger J. Hubbold
Comput. Vis. Image Underst.2
2005 Real-Time 3-D Human Body Tracking using Variable Length Markov Models
abstract
In this paper, we introduce a 3-D human-body tracker capable of handling fast and complex motions in real-time. The parameter space, augmented with first order derivatives, is automatically partitioned into Gaussian clusters each representing an elementary motion: hypothesis propagation inside each cluster is therefore accurate and efficient. The transitions between clusters use the predictions of a Variable Length Markov Model which can explain highlevel behaviours over a long history. Using Monte-Carlo methods, evaluation of model candidates is critical for both speed and robustness. We present a new evaluation scheme based on volumetric reconstruction and blobs-fitting, where appearance models and image evidences are represented by Gaussian mixtures. We demonstrate the application of our tracker to long video sequences exhibiting rapid and diverse movements. 1
Fabrice Caillette, Aphrodite Galata, Toby Howard
BMVC2
2002 Modeling Interaction Using Learnt Qualitative Spatio-Temporal Relations and Variable Length Markov Models
Aphrodite Galata, Anthony G. Cohn 0001, Derek R. Magee, David C. Hogg
ECAI1
2001 Learning Variable-Length Markov Models of Behavior
Aphrodite Galata, Neil Johnson 0001, David C. Hogg
Comput. Vis. Image Underst.1
1999 Learning Behaviour Models of Human Activities
abstract
In recent years there has been an increased interest in the modelling and recognition of human activities involving highly structured and semantically rich behaviour such as dance, aerobics, and sign language. A novel approach is presented for automatically acquiring stochastic models of the high-level structure of an activity without the assumption of any prior knowledge. The process involves temporal segmentation intoplausible atomic behaviour com-ponents and the use of variable length Markov models for the efficient rep-resentation of behaviours. Experimental results are presented which demon-strate the generation of realistic sample behaviours and evaluate the perfor-mance of models for long-term temporal prediction. 1
Aphrodite Galata, Neil Johnson 0001, David C. Hogg
BMVC1
1998 The Acquisition and Use of Interaction Behavior Models
abstract
Providing a machine with the ability to learn and use models of natural interaction is a challenging and largely unaddressed problem. A framework is developed enabling both the acquisition of interaction behaviours from the observation of humans, and the use of the acquired behaviour models to simulate a plausible partner during interaction. Statistically based interaction behaviour models are acquired automatically from the observation of interacting humans. Interaction with a virtual human is achieved using the model together with a stochastic tracking algorithm. Experimental results demonstrate the generation and use of the model for a simple human interaction.
Neil Johnson 0001, Aphrodite Galata, David C. Hogg
CVPR2