VLDB 2026 Research / reviewers in the wild / expert
Sarah Taylor
dblp:151/0026
· DBLP profile ↗
15ranked-venue papers
2as first author
8since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Llanimation: Llama Driven Gesture AnimationabstractAbstract Co‐speech gesturing is an important modality in conversation, providing context and social cues. In character animation, appropriate and synchronised gestures add realism, and can make interactive agents more engaging. Historically, methods for automatically generating gestures were predominantly audio‐driven, exploiting the prosodic and speech‐related content that is encoded in the audio signal. In this paper we instead experiment with using Large‐Language Model (LLM) features for gesture generation that are extracted from text using L lama 2. We compare against audio features, and explore combining the two modalities in both objective tests and a user study. Surprisingly, our results show that L lama 2 features on their own perform significantly better than audio features and that including both modalities yields no significant difference to using L lama 2 features in isolation. We demonstrate that the L lama 2 based model can generate both beat and semantic gestures without any audio input, suggesting LLMs can provide rich encodings that are well suited for gesture generation. Jonathan Windle, Iain A. Matthews, Sarah Taylor |
Comput. Graph. Forum | 3 |
| 2022 | Self-distillation and Uncertainty Boosting Self-supervised Monocular Depth Estimation
Sarah Taylor, David Greenwood 0001, Michal Mackiewicz |
BMVC | 2 |
| 2022 | UEA Digital Humans entry to the GENEA Challenge 2022abstractThis paper describes our entry to the GENEA (Generation and Evaluation of Non-verbal Behaviour for Embodied Agents) challenge 2022. The challenge aims to further the scientific knowledge using a large-scale, joint subjective evaluation of many gesture generation systems. We present two models to the challenge. A Bi-Directional LSTM for the full-body tier and a BDLSTM multi-decoder to produce body-section specific experts. We develop a loss function using both rotations and positions for training our models. We also introduce PASE+ features to the task of pose prediction, along with FastText word embeddings. Our models performed competitively regarding human likeness, and our multiple decoder system performed in the top two submissions for appropriateness of gesture. Jonathan Windle, David Greenwood 0001, Sarah Taylor |
ICMI | 3 |
| 2022 | Pose augmentation: mirror the right wayabstractWe demonstrate an effective method of augmenting speech animation data, and show comparable performance to double the quantity of real data. We investigate the effect of lateral mirroring as a means of data augmentation for 3D poses in multi-speaker, speech-to-motion modelling. Our approach uses a bi-directional LSTM to generate 3D joint positions from audio features extracted using problem-agnostic speech encoder (PASE+) [7]. We demonstrate that naive mirroring for augmentation has a detrimental effect on model performance. We show our method of providing a virtual speaker identity embedding improved performance over no augmentation and was competitive with a model trained on an equal number of samples of real data. Jonathan Windle, Sarah Taylor, David Greenwood 0001, Iain A. Matthews |
IVA | 2 |
| 2022 | Arm motion symmetry in conversationabstractData-driven synthesis of human motion during conversational speech is an active research area with applications that include character animation, computer gaming and conversational agents. Natural looking motion is key to both perceived realism and understanding of any synthesised animation. Multi-modal speech and body-motion data is scarce and limited, so it is common to augment real motion data by mirroring the body pose to double the number of training samples. This augmentation is based on the assumption that a person’s gesturing is not affected by handedness and that the reflected pose is plausible. In this study, we explore the validity of this assumption by evaluating the reflective symmetry of a speaker’s arms during conversational exchanges. We analyse the left and right arm motion of 36 subjects during dyadic conversation and present the per-frame symmetry of the arm gestures. To identify temporal offsets caused by the presence of a leading hand, we compute the time lag between movements of the left and right arms. We perform a nearest neighbour search to test the validity of any mirrored pose. We also consider information theory to examine the information gain from mirroring the data. We implement a speech-to-gesture generative model to determine the efficacy of lateral mirroring techniques for data augmentation. Our findings suggest that both positional symmetry and left–right motion offsets vary from speaker to speaker. We conclude that data augmentation by mirroring is valid in certain cases when considering the mirrored pose as a new virtual identity, but that it should be carefully considered as a generic approach if the gesturing style and handedness of the original speaker is to be maintained. Jonathan Windle, Sarah Taylor, David Greenwood 0001, Iain A. Matthews |
Speech Commun. | 2 |
| 2022 | Speaker-Independent Speech Animation Using Perceptual Loss Functions and Synthetic DataabstractWe propose a real-time speaker-independent speech-to-facial animation system that predicts lip and jaw movements on a reference face for audio speech taken from any speaker. Our approach is motivated by two key observations; 1) Speakerindependent facial animation can be generated from phoneme labels, but to perform this automatically a speech recogniser is needed which, due to contextual look-ahead, introduces too much time lag. 2) Audio-driven speech animation can be performed in real-time but requires large, multi-speaker audio-visual speech datasets of which there are few. We adopt a novel threestage training procedure that leverages the advantages of each approach. First we train aphoneme-to-visual speech model from a large single-speaker audio-visual dataset. Next, we use this model to generate the synthetic visual component of a large multi-speaker audio dataset of which the video is not available. Finally, we learn anaudio-to-visual speech mapping using the synthetic visual features as the target. Furthermore, we increase the realism of the predicted facial animation by introducing two perceptually-based loss functions that aim to improve mouth closures and openings. The proposed method and loss functions are evaluated objectively using mean square error, global variance and a new metric that measures the extent of mouth opening. Subjective tests show that our approach produces facial animation comparable to those produced from phoneme sequences and that improved mouth closures, particularly for bilabial closures, are achieved. Danny Websdale, Sarah Taylor, Ben P. Milner |
IEEE Trans. Multim. | 2 |
| 2021 | Self-Supervised Monocular Depth Estimation with Internal Feature Fusion
David Greenwood 0001, Sarah Taylor |
BMVC | 3 |
| 2021 | Importance of Parasagittal Sensor Information in Tongue Motion Capture Through a Diphonic AnalysisabstractOur study examines the information obtained by adding two parasagittal sensors to the standard midsagittal configuration of an Electromagnetic Articulography (EMA) observation of lingual articulation. In this work, we present a large and phonetically balanced corpus obtained from an EMA recording session of a single English native speaker reading 1899 sentences from the Harvard and TIMIT corpora. According to a statistical analysis of the diphones produced during the recording session, the motion captured by the parasagittal sensors has a low correlation to the midsagittal sensors in the mediolateral direction. We perform a geometric analysis of the lateral tongue by the measure of its width and using a proxy of the tongue’s curvature that is computed using the Menger curvature. To provide a better understanding of the tongue sensor motion we present dynamic visualizations of all diphones. Finally, we present a summary of the velocity information computed from the tongue sensor information. Sarah Taylor, Mark K. Tiede, Alex Hauptmann 0001, Iain A. Matthews |
Interspeech | 2 |
| 2019 | Synthesising visual speech using dynamic visemes and deep learning architectures
Ausdang Thangthai, Ben P. Milner, Sarah Taylor |
Comput. Speech Lang. | 3 |
| 2018 | The Effect of Real-Time Constraints on Automatic Speech AnimationabstractMachine learning has previously been applied successfully to speech-driven facial animation. To account for carry-over and anticipatory coarticulation a common approach is to predict the facial pose using a symmetric window of acoustic speech that includes both past and future context. Using future context limits this approach for animating the faces of characters in real-time and networked applications, such as online gaming. An acceptable latency for conversational speech is 200ms and typically network transmission times will consume a significant part of this. Consequently, we consider asymmetric windows by investigating the extent to which decreasing the future context effects the quality of predicted animation using both deep neural networks (DNNs) and bi-directional LSTM recurrent neural networks (BiLSTMs). Specifically we investigate future contexts from 170ms (fully-symmetric) to 0ms (fullyasymmetric … Danny Websdale, Sarah Taylor, Ben P. Milner |
INTERSPEECH | 2 |
| 2018 | Towards a data archiving solution for learning analyticsabstractData solutions in the teaching and learning space are in need of pro-active innovations in data management, to ensure that systems for learning analytics can scale up to match the size of datasets now available. Here, we illustrate the scale at which a Learning Management System (LMS) accumulates data, and discuss the barriers to using this data for in-depth analyses. We illustrate the exponential growth of our LMS data to represent a single example dataset, and highlight the broader need for taking a pro-active approach to dimensional modelling in learning analytics, anticipating that common learning analytics questions will be computationally expensive, and that the most useful data structures for learning analytics will not necessarily follow those of the source dataset. Sarah Taylor, Pablo Munguia |
LAK | 1 |
| 2018 | Time Series Classification with HIVE-COTE: The Hierarchical Vote Collective of Transformation-Based EnsemblesabstractA recent experimental evaluation assessed 19 time series classification (TSC) algorithms and found that one was significantly more accurate than all others: the Flat Collective of Transformation-based Ensembles (Flat-COTE). Flat-COTE is an ensemble that combines 35 classifiers over four data representations. However, while comprehensive, the evaluation did not consider deep learning approaches. Convolutional neural networks (CNN) have seen a surge in popularity and are now state of the art in many fields and raises the question of whether CNNs could be equally transformative for TSC. We implement a benchmark CNN for TSC using a common structure and use results from a TSC-specific CNN from the literature. We compare both to Flat-COTE and find that the collective is significantly more accurate than both CNNs. These results are impressive, but Flat-COTE is not without deficiencies. We significantly improve the collective by proposing a new hierarchical structure with probabilistic voting, defining and including two novel ensemble classifiers built in existing feature spaces, and adding further modules to represent two additional transformation domains. The resulting classifier, the Hierarchical Vote Collective of Transformation-based Ensembles (HIVE-COTE), encapsulates classifiers built on five data representations. We demonstrate that HIVE-COTE is significantly more accurate than Flat-COTE (and all other TSC algorithms that we are aware of) over 100 resamples of 85 TSC problems and is the new state of the art for TSC. Further analysis is included through the introduction and evaluation of 3 new case studies and extensive experimentation on 1,000 simulated datasets of 5 different types. Jason Lines, Sarah Taylor, Anthony J. Bagnall |
ACM Trans. Knowl. Discov. Data | 2 |
| 2016 | HIVE-COTE: The Hierarchical Vote Collective of Transformation-Based Ensembles for Time Series ClassificationabstractThere have been many new algorithms proposed over the last five years for solving time series classification (TSC) problems. A recent experimental comparison of the leading TSC algorithms has demonstrated that one approach is significantly more accurate than all others over 85 datasets. That approach, the Flat Collective of Transformation-based Ensembles (Flat-COTE), achieves superior accuracy through combining predictions of 35 individual classifiers built on four representations of the data into a flat hierarchy. Outside of TSC, deep learning approaches such as convolutional neural networks (CNN) have seen a recent surge in popularity and are now state of the art in many fields. An obvious question is whether CNNs could be equally transformative in the field of TSC. To test this, we implement a common CNN structure and compare performance to Flat-COTE and a recently proposed time series-specific CNN implementation. We find that Flat-COTE is significantly more accurate than both deep learning approaches on 85 datasets. These results are impressive, but Flat-COTE is not without deficiencies. We improve the collective by adding new components and proposing a modular hierarchical structure with a probabilistic voting scheme that allows us to encapsulate the classifiers built on each transformation. We add two new modules representing dictionary and interval-based classifiers, and significantly improve upon the existing frequency domain classifiers with a novel spectral ensemble. The resulting classifier, the Hierarchical Vote Collective of Transformation-based Ensembles (HIVE-COTE) is significantly more accurate than Flat-COTE and represents a new state of the art for TSC. HIVE-COTE captures more sources of possible discriminatory features in time series and has a more modular, intuitive structure. Jason Lines, Sarah Taylor, Anthony J. Bagnall |
ICDM | 2 |
| 2016 | Audio-to-Visual Speech Conversion Using Deep Neural NetworksabstractWe study the problem of mapping from acoustic to visual speech with the goal of generating accurate, perceptually natural speech animation automatically from an audio speech signal. We present a sliding window deep neural network that learns a mapping from a window of acoustic features to a window of visual features from a large audio-visual speech dataset. Overlapping visual predictions are averaged to generate continuous, smoothly varying speech animation. We outperform a baseline HMM inversion approach in both objective and subjective evaluations and perform a thorough analysis of our results. Sarah Taylor, Akihiro Kato, Iain A. Matthews, Ben P. Milner |
INTERSPEECH | 1 |
| 2016 | Visual Speech Synthesis Using Dynamic Visemes, Contextual Features and DNNsabstractThis paper examines methods to improve visual speech synthesis from a text input using a deep neural network (DNN). Two representations of the input text are considered, namely into phoneme sequences or dynamic viseme sequences. From these sequences, contextual features are extracted that include information at varying linguistic levels, from frame level down to the utterance level. These are extracted from a broad sliding window that captures context and produces features that are input into the DNN to estimate visual features. Experiments first compare the accuracy of these visual features against an HMM baseline method which establishes that both the phoneme and dynamic viseme systems perform better with best performance obtained by a combined phoneme-dynamic viseme system. An investigation into the features then reveals the importance of the frame level information which is able to avoid discontinuities in the visual feature sequence and produces a smooth and realistic output. Ausdang Thangthai, Ben P. Milner, Sarah Taylor |
INTERSPEECH | 3 |