VLDB 2026 Research / reviewers in the wild / expert
Iain A. Matthews
dblp:58/813
· DBLP profile ↗
82ranked-venue papers
8as first author
6since 2021 · last 2024
0009-0000-5004-2397ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 57 · 7 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 51 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 since 2021Databases, data management, data science and information retrieval · 5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | PhISANet: Phonetically Informed Speech Animation NetworkabstractRealistic animation is crucial for immersive and seamless human-avatar interactions as digital avatars become more prevalent. This work presents PhISANet, an encoder-decoder model that realistically animates the face and tongue solely from speech. PhISANet leverages neural audio representations trained on vast amounts of speech to map the speech signal into animation parameters that control the lower face and tongue of realistic 3D models. By integrating a novel multi-task learning strategy during the training phase, PhISANet reincorporates the phonetic information from the input speech, improving articulation in the generated animations. A thorough quantitative and qualitative study validates this improvement, and it determines that WavLM and Whisper features are ideal for training a generalizable speech-animation model regardless of gender, age, and language. Sarah L. Taylor, Carsten Stoll, Gareth Edwards, Alex Hauptmann 0001, Shinji Watanabe 0001, Iain A. Matthews |
ICASSP | 7 |
| 2024 | Llanimation: Llama Driven Gesture AnimationabstractAbstract Co‐speech gesturing is an important modality in conversation, providing context and social cues. In character animation, appropriate and synchronised gestures add realism, and can make interactive agents more engaging. Historically, methods for automatically generating gestures were predominantly audio‐driven, exploiting the prosodic and speech‐related content that is encoded in the audio signal. In this paper we instead experiment with using Large‐Language Model (LLM) features for gesture generation that are extracted from text using L lama 2. We compare against audio features, and explore combining the two modalities in both objective tests and a user study. Surprisingly, our results show that L lama 2 features on their own perform significantly better than audio features and that including both modalities yields no significant difference to using L lama 2 features in isolation. We demonstrate that the L lama 2 based model can generate both beat and semantic gestures without any audio input, suggesting LLMs can provide rich encodings that are well suited for gesture generation. Jonathan Windle, Iain A. Matthews, Sarah Taylor |
Comput. Graph. Forum | 2 |
| 2022 | Speech Driven Tongue AnimationabstractAdvances in speech driven animation techniques allow the creation of convincing animations for virtual characters solely from audio data. Many existing approaches focus on facial and lip motion and they often do not provide realistic animation of the inner mouth. This paper addresses the problem of speech-driven inner mouth animation. Obtaining performance capture data of the tongue and jaw from video alone is difficult because the inner mouth is only partially observable during speech. In this work, we introduce a large-scale speech and mocap dataset that focuses on capturing tongue, jaw, and lip motion. This dataset enables research using data-driven techniques to generate realistic inner mouth animation from speech. We then propose a deep-learning based method for accurate and generalizable speech to tongue and jaw animation, and evaluate several encoder-decoder network architectures and audio feature encoders. We find that recent self-supervised deep learning based audio feature encoders are robust, generalize well to unseen speakers and content, and work best for our task. To demonstrate the practical application of our approach, we show animations on high-quality parametric 3D face models driven by the landmarks generated from our speech-to-tongue animation method. Denis Tomè, Carsten Stoll, Mark K. Tiede, Kevin Munhall, Alex Hauptmann 0001, Iain A. Matthews |
CVPR | 7 |
| 2022 | Pose augmentation: mirror the right wayabstractWe demonstrate an effective method of augmenting speech animation data, and show comparable performance to double the quantity of real data. We investigate the effect of lateral mirroring as a means of data augmentation for 3D poses in multi-speaker, speech-to-motion modelling. Our approach uses a bi-directional LSTM to generate 3D joint positions from audio features extracted using problem-agnostic speech encoder (PASE+) [7]. We demonstrate that naive mirroring for augmentation has a detrimental effect on model performance. We show our method of providing a virtual speaker identity embedding improved performance over no augmentation and was competitive with a model trained on an equal number of samples of real data. Jonathan Windle, Sarah Taylor, David Greenwood 0001, Iain A. Matthews |
IVA | 4 |
| 2022 | Arm motion symmetry in conversationabstractData-driven synthesis of human motion during conversational speech is an active research area with applications that include character animation, computer gaming and conversational agents. Natural looking motion is key to both perceived realism and understanding of any synthesised animation. Multi-modal speech and body-motion data is scarce and limited, so it is common to augment real motion data by mirroring the body pose to double the number of training samples. This augmentation is based on the assumption that a person’s gesturing is not affected by handedness and that the reflected pose is plausible. In this study, we explore the validity of this assumption by evaluating the reflective symmetry of a speaker’s arms during conversational exchanges. We analyse the left and right arm motion of 36 subjects during dyadic conversation and present the per-frame symmetry of the arm gestures. To identify temporal offsets caused by the presence of a leading hand, we compute the time lag between movements of the left and right arms. We perform a nearest neighbour search to test the validity of any mirrored pose. We also consider information theory to examine the information gain from mirroring the data. We implement a speech-to-gesture generative model to determine the efficacy of lateral mirroring techniques for data augmentation. Our findings suggest that both positional symmetry and left–right motion offsets vary from speaker to speaker. We conclude that data augmentation by mirroring is valid in certain cases when considering the mirrored pose as a new virtual identity, but that it should be carefully considered as a generic approach if the gesturing style and handedness of the original speaker is to be maintained. Jonathan Windle, Sarah Taylor, David Greenwood 0001, Iain A. Matthews |
Speech Commun. | 4 |
| 2021 | Importance of Parasagittal Sensor Information in Tongue Motion Capture Through a Diphonic AnalysisabstractOur study examines the information obtained by adding two parasagittal sensors to the standard midsagittal configuration of an Electromagnetic Articulography (EMA) observation of lingual articulation. In this work, we present a large and phonetically balanced corpus obtained from an EMA recording session of a single English native speaker reading 1899 sentences from the Harvard and TIMIT corpora. According to a statistical analysis of the diphones produced during the recording session, the motion captured by the parasagittal sensors has a low correlation to the midsagittal sensors in the mediolateral direction. We perform a geometric analysis of the lateral tongue by the measure of its width and using a proxy of the tongue’s curvature that is computed using the Menger curvature. To provide a better understanding of the tongue sensor motion we present dynamic visualizations of all diphones. Finally, we present a summary of the velocity information computed from the tongue sensor information. Sarah Taylor, Mark K. Tiede, Alex Hauptmann 0001, Iain A. Matthews |
Interspeech | 5 |
| 2020 | Learning Physics-Guided Face Relighting Under Directional LightabstractRelighting is an essential step in realistically transferring objects from a captured image into another environment. For example, authentic telepresence in Augmented Reality requires faces to be displayed and relit consistent with the observer's scene lighting. We investigate end-to-end deep learning architectures that both de-light and relight an image of a human face. Our model decomposes the input image into intrinsic components according to a diffuse physics-based image formation model. We enable non-diffuse effects including cast shadows and specular highlights by predicting a residual correction to the diffuse render. To train and evaluate our model, we collected a portrait database of 21 subjects with various expressions and poses. Each sample is captured in a controlled light stage setup with 32 individual light sources. Our method creates precise and believable relighting results and generalizes to complex illumination conditions and challenging poses, including when the subject is not looking straight at the camera. Thomas Nestmeyer, Jean-François Lalonde, Iain A. Matthews, Andreas M. Lehrmann |
CVPR | 3 |
| 2019 | Panoptic Studio: A Massively Multiview System for Social Interaction CaptureabstractWe present an approach to capture the 3D motion of a group of people engaged in a social interaction. The core challenges in capturing social interactions are: (1) occlusion is functional and frequent; (2) subtle motion needs to be measured over a space large enough to host a social group; (3) human appearance and configuration variation is immense; and (4) attaching markers to the body may prime the nature of interactions. The Panoptic Studio is a system organized around the thesis that social interactions should be measured through the integration of perceptual analyses over a large variety of view points. We present a modularized system designed around this principle, consisting of integrated structural, hardware, and software innovations. The system takes, as input, 480 synchronized video streams of multiple people engaged in social activities, and produces, as output, the labeled time-varying 3D structure of anatomical landmarks on individuals in the space. Our algorithm is designed to fuse the "weak" perceptual processes in the large number of views by progressively generating skeletal proposals from low-level appearance cues, and a framework for temporal refinement is also presented by associating body parts to reconstructed dense 3D trajectory stream. Our system and method are the first in reconstructing full body motion of more than five people engaged in social interactions without using markers. We also empirically demonstrate the impact of the number of views in achieving this goal. Hanbyul Joo, Tomas Simon, Xulong Li 0001, Hao Liu 0125, Sean Banerjee, Timothy Godisart, Bart C. Nabbe, Iain A. Matthews, Takeo Kanade, Shohei Nobuhara, Yaser Sheikh |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2019 | Estimating Audience Engagement to Predict Movie RatingsabstractWhile watching movies, audience members exhibit both subtle and coarse gestures (e.g., smiles, head-pose change, fidgeting, stretching) which convey sentiment (i.e., engaged or disengaged) during feature length movies. Noticing these behaviors using computer vision systems is a very challenging problem-especially in a movie theatre environment. The environment is dark and contains views of people at different scales and viewpoints. Feature length movies typically run 80-120 minutes, and tracking people uninterrupted for this duration is still an unsolved problem. Facial expressions of audience members are subtle, short, and sparse; making it difficult to detect and recognize activities. Finally, annotating audience sentiment at the frame-level is prohibitively time consuming. To circumvent these issues, we use an infrared illuminated test-bed to obtain a visually uniform input of audiences watching feature length movies. We present a method which can automatically detect the change in behavior (key-gestures) using “key-frames”, which can convey audience sentiment. As the number of key-frames are many orders of magnitudes lower than the number of frames, the annotation problem is reduced to assigning a sentiment label for each key-frame. Using these discovered key-gestures, we create a movie rating classifier from crowd-sourced ratings and demonstrate its predictive capability. Our dataset consists of over 50 hours of audience behavior collected across 237 subjects. Rajitha Navarathna, Peter Carr 0001, Patrick Lucey, Iain A. Matthews |
IEEE Trans. Affect. Comput. | 4 |
| 2018 | Joint Learning of Facial Expression and Head Pose from SpeechabstractNatural movement plays a significant role in realistic speech animation, and numerous studies have demonstrated the contribution visual cues make to the degree human observers find an animation acceptable.Natural, expressive, emotive, and prosodic speech exhibits motion patterns that are difficult to predict with considerable variation in visual modalities.Recently, there have been some impressive demonstrations of face animation derived in some way from the speech signal.Each of these methods have taken unique approaches, but none have included rigid head pose in their predicted output.We observe a high degree of correspondence with facial activity and rigid head pose during speech, and exploit this observation to jointly learn full face animation and head pose rotation and translation combined.From our own corpus, we train Deep Bi-Directional LSTMs (BLSTM) capable of learning long-term structure in language to model the relationship that speech has with the complex activity of the face.We define a model architecture to encourage learning of rigid head motion via the latent space of the speaker's facial activity.The result is a model that can predict lip sync and other facial motion along with rigid head motion directly from audible speech. David Greenwood 0001, Iain A. Matthews, Stephen D. Laycock |
INTERSPEECH | 2 |
| 2018 | From Faces to Outdoor Light ProbesabstractAbstract Image‐based lighting has allowed the creation of photo‐realistic computer‐generated content. However, it requires the accurate capture of the illumination conditions, a task neither easy nor intuitive, especially to the average digital photography enthusiast. This paper presents an approach to directly estimate an HDR light probe from a single LDR photograph, shot outdoors with a consumer camera, without specialized calibration targets or equipment. Our insight is to use a person's face as an outdoor light probe. To estimate HDR light probes from LDR faces we use an inverse rendering approach which employs data‐driven priors to guide the estimation of realistic, HDR lighting. We build compact, realistic representations of outdoor lighting both parametrically and in a data‐driven way, by training a deep convolutional autoencoder on a large dataset of HDR sky environment maps. Our approach can recover high‐frequency, extremely high dynamic range lighting environments. For quantitative evaluation of lighting estimation accuracy and relighting accuracy, we also contribute a new database of face photographs with corresponding HDR light probes. We show that relighting objects with HDR light probes estimated by our method yields realistic results in a wide variety of settings. Dan Andrei Calian, Jean-François Lalonde, Paulo F. U. Gotardo, Tomas Simon, Iain A. Matthews, Kenny Mitchell |
Comput. Graph. Forum | 5 |
| 2017 | Factorized Variational Autoencoders for Modeling Audience Reactions to MoviesabstractMatrix and tensor factorization methods are often used for finding underlying low-dimensional patterns from noisy data. In this paper, we study non-linear tensor factorization methods based on deep variational autoencoders. Our approach is well-suited for settings where the relationship between the latent representation to be learned and the raw data representation is highly complex. We apply our approach to a large dataset of facial expressions of movie-watching audiences (over 16 million faces). Our experiments show that compared to conventional linear factorization methods, our method achieves better reconstruction of the data, and further discovers interpretable latent factors. Zhiwei Deng, Rajitha Navarathna, Peter Carr 0001, Stephan Mandt, Yisong Yue, Iain A. Matthews, Greg Mori |
CVPR | 6 |
| 2017 | Hand Keypoint Detection in Single Images Using Multiview BootstrappingabstractWe present an approach that uses a multi-camera system to train fine-grained detectors for keypoints that are prone to occlusion, such as the joints of a hand. We call this procedure multiview bootstrapping: first, an initial keypoint detector is used to produce noisy labels in multiple views of the hand. The noisy detections are then triangulated in 3D using multiview geometry or marked as outliers. Finally, the reprojected triangulations are used as new labeled training data to improve the detector. We repeat this process, generating more labeled data in each iteration. We derive a result analytically relating the minimum number of views to achieve target true and false positive rates for a given detector. The method is used to train a hand keypoint detector for single images. The resulting keypoint detector runs in realtime on RGB images and has accuracy comparable to methods that use depth sensors. The single view detector, triangulated over multiple views, enables 3D markerless hand motion capture with complex object interactions. Tomas Simon, Hanbyul Joo, Iain A. Matthews, Yaser Sheikh |
CVPR | 3 |
| 2017 | Predicting Head Pose from Speech with a Conditional Variational AutoencoderabstractNatural movement plays a significant role in realistic speech animation. Numerous studies have demonstrated the contribution visual cues make to the degree we, as human observers, find an animation acceptable. Rigid head motion is one visual mode that universally co-occurs with speech, and so it is a reasonable strategy to seek a transformation from the speech mode to predict the head pose. Several previous authors have shown that prediction is possible, but experiments are typically confined to rigidly produced dialogue. Natural, expressive, emotive and prosodic speech exhibit motion patterns that are far more difficult to predict with considerable variation in expected head pose. Recently, Long Short Term Memory (LSTM) networks have become an important tool for modelling speech and natural language tasks. We employ Deep Bi-Directional LSTMs (BLSTM) capable of learning long-term structure in language, to model the relationship that speech has with rigid head motion. We then extend our model by conditioning with prior motion. Finally, we introduce a generative head motion model, conditioned on audio features using a Conditional Variational Autoencoder (CVAE). Each approach mitigates the problems of the one to many mapping that a speech to head pose model must accommodate David Greenwood 0001, Stephen D. Laycock, Iain A. Matthews |
INTERSPEECH | 3 |
| 2017 | Predicting Head Pose in Dyadic Conversation
David Greenwood 0001, Stephen D. Laycock, Iain A. Matthews |
IVA | 3 |
| 2017 | Kronecker-Markov Prior for Dynamic 3D ReconstructionabstractRecovering dynamic 3D structures from 2D image observations is highly under-constrained because of projection and missing data, motivating the use of strong priors to constrain shape deformation. In this paper, we empirically show that the spatiotemporal covariance of natural deformations is dominated by a Kronecker pattern. We demonstrate that this pattern arises as the limit of a spatiotemporal autoregressive process, and derive a Kronecker Markov Random Field as a prior distribution over dynamic structures. This distribution unifies shape and trajectory models of prior art and has the individual models as its marginals. The key assumption of the Kronecker MRF is that the spatiotemporal covariance is separable into the product of a temporal and a shape covariance, and can therefore be modeled using the matrix normal distribution. Analysis on motion capture data validates that this distribution is an accurate approximation with significantly fewer free parameters. Using the trace-norm, we present a convex method to estimate missing data from a single sequence when the marginal shape distribution is unknown. The Kronecker-Markov distribution, fit to a single sequence, outperforms state-of-the-art methods at inferring missing 3D data, and additionally provides covariance estimates of the uncertainty. Tomas Simon, Jack Valmadre, Iain A. Matthews, Yaser Sheikh |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | A deep learning approach for generalized speech animationabstractWe introduce a simple and effective deep learning approach to automatically generate natural looking speech animation that synchronizes to input speech. Our approach uses a sliding window predictor that learns arbitrary nonlinear mappings from phoneme label input sequences to mouth movements in a way that accurately captures natural motion and visual coarticulation effects. Our deep learning approach enjoys several attractive properties: it runs in real-time, requires minimal parameter tuning, generalizes well to novel input speech sequences, is easily edited to create stylized and emotional speech, and is compatible with existing animation retargeting approaches. One important focus of our work is to develop an effective approach for speech animation that can be easily integrated into existing production pipelines. We provide a detailed description of our end-to-end approach, including machine learning design decisions. Generalized speech animation results are demonstrated over a wide range of animation clips on a variety of characters and voices, including singing and foreign language input. Our approach can also generate on-demand speech animation in real-time from user speech input. Sarah L. Taylor, Yisong Yue, Moshe Mahler, James Krahe, Anastasio Garcia Rodriguez, Jessica K. Hodgins, Iain A. Matthews |
ACM Trans. Graph. | 8 |
| 2016 | Synthetic Prior Design for Real-Time Face TrackingabstractReal-time facial performance capture has recently been gaining popularity in virtual film production, driven by advances in machine learning, which allows for fast inference of facial geometry from video streams. These learning-based approaches are significantly influenced by the quality and amount of labelled training data. Tedious construction of training sets from real imagery can be replaced by rendering a facial animation rig under on-set conditions expected at runtime. We learn a synthetic actor-specific prior by adapting a state-of-the-art facial tracking method. Synthetic training significantly reduces the capture and annotation burden and in theory allows generation of an arbitrary amount of data. But practical realities such as training time and compute resources still limit the size of any training set. We construct better and smaller training sets by investigating which facial image appearances are crucial for tracking accuracy, covering the dimensions of expression, viewpoint and illumination. A reduction of training data in 1-2 orders of magnitude is demonstrated whilst tracking accuracy is retained for challenging on-set footage. Steven McDonagh 0001, Martin Klaudiny, Derek Bradley, Thabo Beeler, Iain A. Matthews, Kenny Mitchell |
3DV | 5 |
| 2016 | Audio-to-Visual Speech Conversion Using Deep Neural NetworksabstractWe study the problem of mapping from acoustic to visual speech with the goal of generating accurate, perceptually natural speech animation automatically from an audio speech signal. We present a sliding window deep neural network that learns a mapping from a window of acoustic features to a window of visual features from a large audio-visual speech dataset. Overlapping visual predictions are averaged to generate continuous, smoothly varying speech animation. We outperform a baseline HMM inversion approach in both objective and subjective evaluations and perform a thorough analysis of our results. Sarah Taylor, Akihiro Kato, Iain A. Matthews, Ben P. Milner |
INTERSPEECH | 3 |
| 2016 | Chalkboarding: A New Spatiotemporal Query Paradigm for Sports Play RetrievalabstractThe recent explosion of sports tracking data has dramatically increased the interest in effective data processing and access of sports plays (i.e., short trajectory sequences of players and the ball). And while there exist systems that offer improved categorizations of sports plays (e.g., into relatively coarse clusters), to the best of our knowledge there does not exist any retrieval system that can effectively search for the most relevant plays given a specific input query. One significant design challenge is how best to phrase queries for multi-agent spatiotemporal trajectories such as sports plays.We have developed a novel query paradigm and retrieval system, which we call Chalkboarding, that allows the user to issue queries by drawing a play of interest (similar to how coaches draw up plays). Our system utilizes effective alignment, templating, and hashing techniques tailored to multi-agent trajectories, and achieves accurate play retrieval at interactive speeds.We showcase the efficacy of our approach in a user study, where we demonstrate orders-of-magnitude improvements in search quality compared to baseline systems. Long Sha, Patrick Lucey, Yisong Yue, Peter Carr 0001, Charlie Rohlf, Iain A. Matthews |
IUI | 6 |
| 2016 | Discovering Team Structures in Soccer from Spatiotemporal DataabstractIn team sports like soccer, utilizing tracking data for analysis is challenging due to the dynamic and multi-agent nature of the data. The biggest issue surrounds the changing of positions or “roles” between players on a frame-to-frame basis, which causes misalignment of the data and makes it difficult to perform team analysis. In this paper, we present an unsupervised method to learn a formation template which allows us to “align” the tracking data at the frame level. Not only does this approach give important contextual information to facilitate large-scale analysis (e.g., we know when a player is in the left-wing position compared to left-back), it also yields the team structure or “formation” which serves as a strong descriptor for identifying a team's style. The utility of the approach is demonstrated on a full season of player and ball tracking data from a professional soccer league consisting of over 21.5 million frames of player tracking data. Alina Bialkowski, Patrick Lucey, Peter Carr 0001, Iain A. Matthews, Sridha Sridharan, Clinton Fookes |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2015 | A mouth full of words: Visually consistent acoustic redubbingabstractThis paper introduces a method for automatic redubbing of video that exploits the many-to-many mapping of phoneme sequences to lip movements modelled as dynamic visemes (Taylor et al., 2012). For a given utterance, the corresponding dynamic viseme sequence is sampled to construct a graph of possible phoneme sequences that synchronize with the video. When composed with a pronunciation dictionary and language model, this produces a vast number of word sequences that are in sync with the original video, literally putting plausible words into the mouth of the speaker. We demonstrate that traditional, many-to-one, static visemes lack flexibility for this application as they produce significantly fewer word sequences. This work explores the natural ambiguity in visual speech and offers insight for automatic speech recognition and the importance of language modeling. Sarah L. Taylor, Barry-John Theobald, Iain A. Matthews |
ICASSP | 3 |
| 2015 | Photogeometric Scene Flow for High-Detail Dynamic 3D ReconstructionabstractPhotometric stereo (PS) is an established technique for high-detail reconstruction of 3D geometry and appearance. To correct for surface integration errors, PS is often combined with multiview stereo (MVS). With dynamic objects, PS reconstruction also faces the problem of computing optical flow (OF) for image alignment under rapid changes in illumination. Current PS methods typically compute optical flow and MVS as independent stages, each one with its own limitations and errors introduced by early regularization. In contrast, scene flow methods estimate geometry and motion, but lack the fine detail from PS. This paper proposes photogeometric scene flow (PGSF) for high-quality dynamic 3D reconstruction. PGSF performs PS, OF, and MVS simultaneously. It is based on two key observations: (i) while image alignment improves PS, PS allows for surfaces to be relit to improve alignment, (ii) PS provides surface gradients that render the smoothness term in MVS unnecessary, leading to truly data-driven, continuous depth estimates. This synergy is demonstrated in the quality of the resulting RGB appearance, 3D geometry, and 3D motion. Paulo F. U. Gotardo, Tomas Simon, Yaser Sheikh, Iain A. Matthews |
ICCV | 4 |
| 2015 | Panoptic Studio: A Massively Multiview System for Social Motion CaptureabstractWe present an approach to capture the 3D structure and motion of a group of people engaged in a social interaction. The core challenges in capturing social interactions are: (1) occlusion is functional and frequent, (2) subtle motion needs to be measured over a space large enough to host a social group, and (3) human appearance and configuration variation is immense. The Panoptic Studio is a system organized around the thesis that social interactions should be measured through the perceptual integration of a large variety of view points. We present a modularized system designed around this principle, consisting of integrated structural, hardware, and software innovations. The system takes, as input, 480 synchronized video streams of multiple people engaged in social activities, and produces, as output, the labeled time-varying 3D structure of anatomical landmarks on individuals in the space. The algorithmic contributions include a hierarchical approach for generating skeletal trajectory proposals, and an optimization framework for skeletal reconstruction with trajectory re-association. Hanbyul Joo, Hao Liu 0125, Bart C. Nabbe, Iain A. Matthews, Takeo Kanade, Shohei Nobuhara, Yaser Sheikh |
ICCV | 6 |
| 2015 | A Decision Tree Framework for Spatiotemporal Sequence PredictionabstractWe study the problem of learning to predict a spatiotemporal output sequence given an input sequence. In contrast to conventional sequence prediction problems such as part-of-speech tagging (where output sequences are selected using a relatively small set of discrete labels), our goal is to predict sequences that lie within a high-dimensional continuous output space. We present a decision tree framework for learning an accurate non-parametric spatiotemporal sequence predictor. Our approach enjoys several attractive properties, including ease of training, fast performance at test time, and the ability to robustly tolerate corrupted training data using a novel latent variable approach. We evaluate on several datasets, and demonstrate substantial improvements over existing decision tree based sequence learning frameworks such as SEARN and DAgger. Yisong Yue, Sarah L. Taylor, Iain A. Matthews |
KDD | 4 |
| 2015 | 3D Trajectory Reconstruction under Perspective Projection
Hyun Soo Park, Takaaki Shiratori, Iain A. Matthews, Yaser Sheikh |
Int. J. Comput. Vis. | 3 |
| 2014 | Lighting Estimation in Outdoor Image CollectionsabstractLarge scale structure-from-motion (SfM) algorithms have recently enabled the reconstruction of highly detailed 3-D models of our surroundings simply by taking photographs. In this paper, we propose to leverage these reconstruction techniques to automatically estimate the outdoor illumination conditions for each image in a SfM photo collection. We introduce a novel dataset of outdoor photo collections, where the ground truth lighting conditions are known at each image. We also present an inverse rendering approach that recovers a high dynamic range estimate of the lighting conditions for each low dynamic range input image. Our novel database is used to quantitatively evaluate the performance of our algorithm. Results show that physically plausible lighting estimates can faithfully be recovered, both in terms of light direction and intensity. Jean-François Lalonde, Iain A. Matthews |
3DV | 2 |
| 2014 | Separable Spatiotemporal Priors for Convex Reconstruction of Time-Varying 3D Point Clouds
Tomas Simon, Jack Valmadre, Iain A. Matthews, Yaser Sheikh |
ECCV (3) | 3 |
| 2014 | The effect of speaking rate on audio and visual speechabstractThe speed that an utterance is spoken affects both the duration of the speech and the position of the articulators. Consequently, the sounds that are produced are modified, as are the position and appearance of the lips, teeth, tongue and other visible articulators. We describe an experiment designed to measure the effect of variable speaking rate on audio and visual speech by comparing sequences of phonemes and dynamic visemes appearing in the same sentences spoken at different speeds. We find that both audio and visual speech production are affected by varying the rate of speech, however, the effect is significantly more prominent in visual speech. Sarah L. Taylor, Barry-John Theobald, Iain A. Matthews |
ICASSP | 3 |
| 2014 | Large-Scale Analysis of Soccer Matches Using Spatiotemporal Tracking DataabstractAlthough the collection of player and ball tracking data is fast becoming the norm in professional sports, large-scale mining of such spatiotemporal data has yet to surface. In this paper, given an entire season's worth of player and ball tracking data from a professional soccer league (≈400,000,000 data points), we present a method which can conduct both individual player and team analysis. Due to the dynamic, continuous and multi-player nature of team sports like soccer, a major issue is aligning player positions over time. We present a "role-based" representation that dynamically updates each player's relative role at each frame and demonstrate how this captures the short-term context to enable both individual player and team analysis. We discover role directly from data by utilizing a minimum entropy data partitioning method and show how this can be used to accurately detect and visualize formations, as well as analyze individual player behavior. Alina Bialkowski, Patrick Lucey, Peter Carr 0001, Yisong Yue, Sridha Sridharan, Iain A. Matthews |
ICDM | 6 |
| 2014 | Learning Fine-Grained Spatial Models for Dynamic Sports Play PredictionabstractWe consider the problem of learning predictive models for in-game sports play prediction. Focusing on basketball, we develop models for anticipating near-future events given the current game state. We employ a latent factor modeling approach, which leads to a compact data representation that enables efficient prediction given raw spatiotemporal tracking data. We validate our approach using tracking data from the 2012-2013 NBA season, and show that our model can make accurate in-game predictions. We provide a detailed inspection of our learned factors, and show that our model is interpretable and corresponds to known intuitions of basketball game play. Yisong Yue, Patrick Lucey, Peter Carr 0001, Alina Bialkowski, Iain A. Matthews |
ICDM | 5 |
| 2014 | Predicting movie ratings from audience behaviorsabstractWe propose a method of representing audience behavior through facial and body motions from a single video stream, and use these features to predict the rating for feature-length movies. This is a very challenging problem as: i) the movie viewing environment is dark and contains views of people at different scales and viewpoints; ii) the duration of feature-length movies is long (80-120 mins) so tracking people uninterrupted for this length of time is still an unsolved problem; and iii) expressions and motions of audience members are subtle, short and sparse making labeling of activities unreliable. To circumvent these issues, we use an infrared illuminated test-bed to obtain a visually uniform input. We then utilize motion-history features which capture the subtle movements of a person within a pre-defined volume, and then form a group representation of the audience by a histogram of pair-wise correlations over a small-window of time. Using this group representation, we learn our movie rating classifier from crowd-sourced ratings collected by rottentomatoes.com and show our prediction capability on audiences from 30 movies across 250 subjects (> 50 hrs). Rajitha Navarathna, Patrick Lucey, Peter Carr 0001, Elizabeth J. Carter, Sridha Sridharan, Iain A. Matthews |
WACV | 6 |
| 2013 | Representing and Discovering Adversarial Team Behaviors Using Player RolesabstractIn this paper, we describe a method to represent and discover adversarial group behavior in a continuous domain. In comparison to other types of behavior, adversarial behavior is heavily structured as the location of a player (or agent) is dependent both on their teammates and adversaries, in addition to the tactics or strategies of the team. We present a method which can exploit this relationship through the use of a spatiotemporal basis model. As players constantly change roles during a match, we show that employing a "role-based" representation instead of one based on player "identity" can best exploit the playing structure. As vision-based systems currently do not provide perfect detection/tracking (e.g. missed or false detections), we show that our compact representation can effectively "denoise" erroneous detections as well as enabling temporal analysis, which was previously prohibitive due to the dimensionality of the signal. To evaluate our approach, we used a fully instrumented field-hockey pitch with 8 fixed high-definition (HD) cameras and evaluated our approach on approximately 200,000 frames of data from a state-of-the-art real-time player detector and compare it to manually labelled data. Patrick Lucey, Alina Bialkowski, Peter Carr 0001, Stuart Morgan, Iain A. Matthews, Yaser Sheikh |
CVPR | 5 |
| 2013 | Assessing team strategy using spatiotemporal dataabstractThe "Moneyball" revolution coincided with a shift in the way professional sporting organizations handle and utilize data in terms of decision making processes. Due to the demand for better sports analytics and the improvement in sensor technology, there has been a plethora of ball and player tracking information generated within professional sports for analytical purposes. However, due to the continuous nature of the data and the lack of associated high-level labels to describe it - this rich set of information has had very limited use especially in the analysis of a team's tactics and strategy. In this paper, we give an overview of the types of analysis currently performed mostly with hand-labeled event data and highlight the problems associated with the influx of spatiotemporal data. By way of example, we present an approach which uses an entire season of ball tracking data from the English Premier League (2010-2011 season) to reinforce the common held belief that teams should aim to "win home games and draw away ones". We do this by: i) forming a representation of team behavior by chunking the incoming spatiotemporal signal into a series of quantized bins, and ii) generate an expectation model of team behavior based on a code-book of past performances. We show that home advantage in soccer is partly due to the conservative strategy of the away team. We also show that our approach can flag anomalous team behavior which has many potential applications. Patrick Lucey, Dean Oliver, Peter Carr 0001, Joe Roth, Iain A. Matthews |
KDD | 5 |
| 2013 | Hybrid robotic/virtual pan-tilt-zom cameras for autonomous event recordingabstractWe present a method to generate aesthetic video from a robotic camera by incorporating a virtual camera operating on a delay, and a hybrid controller which uses feedback from both the robotic and virtual cameras. Our strategy employs a robotic camera to follow a coarse region-of-interest identified by a realtime computer vision system, and then resamples the captured images to synthesize the video that would have been recorded along a smooth, aesthetic camera trajectory. The smooth motion trajectory is obtained by operating the virtual camera on a short delay so that perfect knowledge of immediate future events is known. Previous autonomous camera installations have employed either robotic cameras or stationary wide-angle cameras with subregion cropping. Robotic cameras track the subject using realtime sensor data, and regulate a smoothness-latency trade-off through control gains. Fixed cameras post-process the data and suffer significant reductions in image resolution when the subject moves freely over a large area. Peter Carr 0001, Michael N. Mistry, Iain A. Matthews |
ACM Multimedia | 3 |
| 2013 | One-man-band: a touch screen interface for producing live multi-camera sports broadcastsabstractGenerating live broadcasts of sporting events requires a coordinated crew of camera operators, directors, and technical personnel to control and switch between multiple cameras to tell the evolving story of a game. In this paper, we present an unimodal interface concept that allows one person to cover live sporting action by controlling multiple cameras and and determining which view to broadcast. The interface exploits the structure of sports broadcasts which typically switch between a zoomed out game-camera view (which records the strategic team-level play), and a zoomed in iso-camera view (which captures the animated adversarial relations between opposing players). The operator simultaneously controls multiple pan-tilt-zoom cameras by pointing at a location on the touch screen, and selects which camera to broadcast using one or two points of contact. The image from the selected camera is superimposed on top of a wide-angle view captured from a context-camera which provides the operator with periphery information (which is useful for ensuring good framing while controlling the camera). We show that by unifying directorial and camera operation functions, we can achieve comparable broadcast quality to a multi-person crew, while reducing cost, logistical, and communication complexities. Eric Foote, Peter Carr 0001, Patrick Lucey, Yaser Sheikh, Iain A. Matthews |
ACM Multimedia | 5 |
| 2012 | Characterizing Multi-Agent Team Behavior from Partial Team Tracings: Evidence from the English Premier LeagueabstractReal-world AI systems have been recently deployed which can automatically analyze the plan and tactics of tennis players. As the game-state is updated regularly at short intervals (i.e. point-level), a library of successful and unsuccessful plans of a player can be learnt over time. Given the relative strengths and weaknesses of a player’s plans, a set of proven plans or tactics from the library that characterize a player can be identified. For low-scoring, continuous team sports like soccer, such analysis for multi-agent teams does not exist as the game is not segmented into “discretized” plays (i.e. plans), making it difficult to obtain a library that characterizes a team’s behavior. Additionally, as player tracking data is costly and difficult to obtain, we only have partial team tracings in the form of ball actions which makes this problem even more difficult. In this paper, we propose a method to overcome these issues by representing team behavior via play-segments, which are spatio-temporal descriptions of ball movement over fixed windows of time. Using these representations we can characterize team behavior from entropy maps, which give a measure of predictability of team behaviors across the field. We show the efficacy and applicability of our method on the 2010-2011 English Premier League soccer data. Patrick Lucey, Alina Bialkowski, Peter Carr 0001, Eric Foote, Iain A. Matthews |
AAAI | 5 |
| 2012 | Monocular Object Detection Using 3D Geometric Primitives
Peter Carr 0001, Yaser Sheikh, Iain A. Matthews |
ECCV (1) | 3 |
| 2012 | Point-less calibration: Camera parameters from gradient-based alignment to edge imagesabstractPoint-based targets, such as checkerboards, are often not practical for outdoor camera calibration, as cameras are usually at significant heights requiring extremely large calibration patterns on the ground. Fortunately, it is possible to make use of existing non-point landmarks in the scene by formulating camera calibration in terms of image alignment. In this paper, we simultaneously estimate the camera intrinsic, extrinsic and lens distortion parameters directly by aligning to a planar schematic of the scene. For cameras with square pixels and known principal point, finding the parameters to such an image warp is equivalent to calibrating the camera. Overhead schematics of many environments resemble edge images. Edge images are difficult to align using image-based algorithms because both the image and its gradient are sparse. We employ a `long range' gradient which enables informative parameter updates at each iteration while maintaining a precise alignment measure. As a result, we are able to calibrate our camera models robustly using regular gradient-based image alignment, given an initial ground to image homography estimate. Peter Carr 0001, Yaser Sheikh, Iain A. Matthews |
WACV | 3 |
| 2012 | Painful monitoring: Automatic pain monitoring using the UNBC-McMaster shoulder pain expression archive database
Patrick Lucey, Jeffrey F. Cohn, Kenneth M. Prkachin, Patricia E. Solomon, Sien W. Chew, Iain A. Matthews |
Image Vis. Comput. | 6 |
| 2012 | Relating Objective and Subjective Performance Measures for AAM-Based Visual Speech SynthesisabstractWe compare two approaches for synthesizing visual speech using active appearance models (AAMs): one that utilizes acoustic features as input, and one that utilizes a phonetic transcription as input. Both synthesizers are trained using the same data and the performance is measured using both objective and subjective testing. We investigate the impact of likely sources of error in the synthesized visual speech by introducing typical errors into real visual speech sequences and subjectively measuring the perceived degradation. When only a small region (e.g., a single syllable) of ground-truth visual speech is incorrect we find that the subjective score for the entire sequence is subjectively lower than sequences generated by our synthesizers. This observation motivates further consideration of an often ignored issue, which is to what extent are subjective measures correlated with objective measures of performance? Significantly, we find that the most commonly used objective measures of performance are not necessarily the best indicator of viewer perception of quality. We empirically evaluate alternatives and show that the cost of a dynamic time warp of synthesized visual speech parameters to the respective ground-truth parameters is a better indicator of subjective quality. Barry-John Theobald, Iain A. Matthews |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Bilinear spatiotemporal basis modelsabstractA variety of dynamic objects, such as faces, bodies, and cloth, are represented in computer graphics as a collection of moving spatial landmarks. Spatiotemporal data is inherent in a number of graphics applications including animation, simulation, and object and camera tracking. The principal modes of variation in the spatial geometry of objects are typically modeled using dimensionality reduction techniques, while concurrently, trajectory representations like splines and autoregressive models are widely used to exploit the temporal regularity of deformation. In this article, we present the bilinear spatiotemporal basis as a model that simultaneously exploits spatial and temporal regularity while maintaining the ability to generalize well to new sequences. This factorization allows the use of analytical, predefined functions to represent temporal variation (e.g., B-Splines or the Discrete Cosine Transform) resulting in efficient model representation and estimation. The model can be interpreted as representing the data as a linear combination of spatiotemporal sequences consisting of shape modes oscillating over time at key frequencies. We apply the bilinear model to natural spatiotemporal phenomena, including face, body, and cloth motion data, and compare it in terms of compaction, generalization ability, predictive precision, and efficiency to existing models. We demonstrate the application of the model to a number of graphics tasks including labeling, gap-filling, denoising, and motion touch-up. Ijaz Akhter, Tomas Simon, Sohaib Khan, Iain A. Matthews, Yaser Sheikh |
ACM Trans. Graph. | 4 |
| 2012 | In the Pursuit of Effective Affective Computing: The Relationship Between Features and RegistrationabstractFor facial expression recognition systems to be applicable in the real world, they need to be able to detect and track a previously unseen person's face and its facial movements accurately in realistic environments. A highly plausible solution involves performing a "dense" form of alignment, where 60-70 fiducial facial points are tracked with high accuracy. The problem is that, in practice, this type of dense alignment had so far been impossible to achieve in a generic sense, mainly due to poor reliability and robustness. Instead, many expression detection methods have opted for a "coarse" form of face alignment, followed by an application of a biologically inspired appearance descriptor such as the histogram of oriented gradients or Gabor magnitudes. Encouragingly, recent advances to a number of dense alignment algorithms have demonstrated both high reliability and accuracy for unseen subjects [e.g., constrained local models (CLMs)]. This begs the question: Aside from countering against illumination variation, what do these appearance descriptors do that standard pixel representations do not? In this paper, we show that, when close to perfect alignment is obtained, there is no real benefit in employing these different appearance-based representations (under consistent illumination conditions). In fact, when misalignment does occur, we show that these appearance descriptors do work well by encoding robustness to alignment error. For this work, we compared two popular methods for dense alignment-subject-dependent active appearance models versus subject-independent CLMs-on the task of action-unit detection. These comparisons were conducted through a battery of experiments across various publicly available data sets (i.e., CK+, Pain, M3, and GEMEP-FERA). We also report our performance in the recent 2011 Facial Expression Recognition and Analysis Challenge for the subject-independent task. Sien W. Chew, Patrick Lucey, Simon Lucey, Jason M. Saragih, Jeffrey F. Cohn, Iain A. Matthews, Sridha Sridharan |
IEEE Trans. Syst. Man Cybern. Part B | 6 |
| 2011 | Painful data: The UNBC-McMaster shoulder pain expression archive databaseabstractA major factor hindering the deployment of a fully functional automatic facial expression detection system is the lack of representative data. A solution to this is to narrow the context of the target application, so enough data is available to build robust models so high performance can be gained. Automatic pain detection from a patient's face represents one such application. To facilitate this work, researchers at McMaster University and University of Northern British Columbia captured video of participant's faces (who were suffering from shoulder pain) while they were performing a series of active and passive range-of-motion tests to their affected and unaffected limbs on two separate occasions. Each frame of this data was AU coded by certified FACS coders, and self-report and observer measures at the sequence level were taken as well. This database is called the UNBC-McMaster Shoulder Pain Expression Archive Database. To promote and facilitate research into pain and augment current datasets, we have publicly made available a portion of this database which includes: (1) 200 video sequences containing spontaneous facial expressions, (2) 48,398 FACS coded frames, (3) associated pain frame-by-frame scores and sequence-level self-report and observer measures, and (4) 66-point AAM landmarks. This paper documents this data distribution in addition to describing baseline results of our AAM/SVM system. This data will be available for distribution in March 2011. Patrick Lucey, Jeffrey F. Cohn, Kenneth M. Prkachin, Patricia E. Solomon, Iain A. Matthews |
FG | 5 |
| 2011 | Facial Expression Transfer with Input-Output Temporal Restricted Boltzmann MachinesabstractWe present a type of Temporal Restricted Boltzmann Machine that defines a probability distribution over an output sequence conditional on an input sequence. It shares the desirable properties of RBMs: efficient exact inference, an exponentially more expressive latent state than HMMs, and the ability to model nonlinear structure and dynamics. We apply our model to a challenging real-world graphics problem: facial expression transfer. Our results demonstrate improved performance over several baselines modeling high-dimensional 2D and 3D data. Matthew D. Zeiler, Graham W. Taylor, Leonid Sigal, Iain A. Matthews, Rob Fergus |
NIPS | 4 |
| 2011 | Modeling and animating eye blinksabstractFacial animation often falls short in conveying the nuances present in the facial dynamics of humans. In this article, we investigate the subtleties of the spatial and temporal aspects of eye blinks. Conventional methods for eye blink animation generally employ temporally and spatially symmetric sequences; however, naturally occurring blinks in humans show a pronounced asymmetry on both dimensions. We present an analysis of naturally occurring blinks that was performed by tracking data from high-speed video using active appearance models. Based on this analysis, we generate a set of key-frame parameters that closely match naturally occurring blinks. We compare the perceived naturalness of blinks that are animated based on real data to those created using textbook animation curves. The eye blinks are animated on two characters, a photorealistic model and a cartoon model, to determine the influence of character style. We find that the animated blinks generated from the human data model with fully closing eyelids are consistently perceived as more natural than those created using the various types of blink dynamics proposed in animation textbooks. Laura C. Trutoiu, Elizabeth J. Carter, Iain A. Matthews, Jessica K. Hodgins |
ACM Trans. Appl. Percept. | 3 |
| 2011 | Interactive region-based linear 3D face modelsabstractLinear models, particularly those based on principal component analysis (PCA), have been used successfully on a broad range of human face-related applications. Although PCA models achieve high compression, they have not been widely used for animation in a production environment because their bases lack a semantic interpretation. Their parameters are not an intuitive set for animators to work with. In this paper we present a linear face modelling approach that generalises to unseen data better than the traditional holistic approach while also allowing click-and-drag interaction for animation. Our model is composed of a collection of PCA sub-models that are independently trained but share boundaries. Boundary consistency and user-given constraints are enforced in a soft least mean squares sense to give flexibility to the model while maintaining coherence. Our results show that the region-based model generalises better than its holistic counterpart when describing previously unseen motion capture data from multiple subjects. The decomposition of the face into several regions, which we determine automatically from training data, gives the user localised manipulation control. This feature allows to use the model for face posing and animation in an intuitive style. J. Rafael Tena, Fernando De la Torre, Iain A. Matthews |
ACM Trans. Graph. | 3 |
| 2011 | Automatically Detecting Pain in Video Through Facial Action UnitsabstractIn a clinical setting, pain is reported either through patient self-report or via an observer. Such measures are problematic as they are: 1) subjective, and 2) give no specific timing information. Coding pain as a series of facial action units (AUs) can avoid these issues as it can be used to gain an objective measure of pain on a frame-by-frame basis. Using video data from patients with shoulder injuries, in this paper, we describe an active appearance model (AAM)-based system that can automatically detect the frames in video in which a patient is in pain. This pain data set highlights the many challenges associated with spontaneous emotion detection, particularly that of expression and head movement due to the patient's reaction to pain. In this paper, we show that the AAM can deal with these movements and can achieve significant improvements in both the AU and pain detection performance compared to the current-state-of-the-art approaches which utilize similarity-normalized appearance features only. Patrick Lucey, Jeffrey F. Cohn, Iain A. Matthews, Simon Lucey, Sridha Sridharan, Jessica Howlett, Kenneth M. Prkachin |
IEEE Trans. Syst. Man Cybern. Part B | 3 |
| 2010 | Motion fields to predict play evolution in dynamic sport scenesabstractVideos of multi-player team sports provide a challenging domain for dynamic scene analysis. Player actions and interactions are complex as they are driven by many factors, such as the short-term goals of the individual player, the overall team strategy, the rules of the sport, and the current context of the game. We show that constrained multi-agent events can be analyzed and even predicted from video. Such analysis requires estimating the global movements of all players in the scene at any time, and is needed for modeling and predicting how the multi-agent play evolves over time on the field. To this end, we propose a novel approach to detect the locations of where the play evolution will proceed, e.g. where interesting events will occur, by tracking player positions and movements over time. We start by extracting the ground level sparse movement of players in each time-step, and then generate a dense motion field. Using this field we detect locations where the motion converges, implying positions towards which the play is evolving. We evaluate our approach by analyzing videos of a variety of complex soccer plays. Matthias Grundmann 0002, Ariel Shamir, Iain A. Matthews, Jessica K. Hodgins, Irfan A. Essa |
CVPR | 4 |
| 2010 | 3D Reconstruction of a Moving Point from a Series of 2D Projections
Hyun Soo Park, Takaaki Shiratori, Iain A. Matthews, Yaser Sheikh |
ECCV (3) | 3 |
| 2010 | Multi-PIE
Ralph Gross, Iain A. Matthews, Jeffrey F. Cohn, Takeo Kanade, Simon Baker |
Image Vis. Comput. | 2 |
| 2010 | Perceptually motivated guidelines for voice synchronization in filmabstractWe consume video content in a multitude of ways, including in movie theaters, on television, on DVDs and Blu-rays, online, on smart phones, and on portable media players. For quality control purposes, it is important to have a uniform viewing experience across these various platforms. In this work, we focus on voice synchronization, an aspect of video quality that is strongly affected by current post-production and transmission practices. We examined the synchronization of an actor's voice and lip movements in two distinct scenarios. First, we simulated the temporal mismatch between the audio and video tracks that can occur during dubbing or during broadcast. Next, we recreated the pitch changes that result from conversions between formats with different frame rates. We show, for the first time, that these audio visual mismatches affect viewer enjoyment. When temporal synchronization is noticeably absent, there is a decrease in the perceived performance quality and the perceived emotional intensity of a performance. For pitch changes, we find that higher pitch voices are not preferred, especially for male actors. Based on our findings, we advise that mismatched audio and video signals negatively affect viewer experience. Elizabeth J. Carter, Lavanya Sharan, Laura C. Trutoiu, Iain A. Matthews, Jessica K. Hodgins |
ACM Trans. Appl. Percept. | 4 |
| 2008 | Increasing the density of Active Appearance ModelsabstractActive appearance models (AAMs) typically only use 50-100 mesh vertices because they are usually constructed from a set of training images with the vertices hand-labeled on them. In this paper, we propose an algorithm to increase the density of an AAM. Our algorithm operates by iteratively building the AAM, refitting the AAM to the training data, and refining the AAM.We compare our algorithm with the state of the art in optical flow algorithms and find it to be significantly more accurate. We also show that dense AAMs can be fit more robustly than sparse ones. Finally, we show how our algorithm can be used to construct AAMs automatically, starting with a single affine model that is subsequently refined to model non-planarity and non-rigidity. Krishnan Ramnath, Simon Baker, Iain A. Matthews, Deva Ramanan |
CVPR | 3 |
| 2008 | Multi-PIEabstractA close relationship exists between the advancement of face recognition algorithms and the availability of face databases varying factors that affect facial appearance in a controlled manner. The CMU PIE database has been very influential in advancing research in face recognition across pose and illumination. Despite its success the PIE database has several shortcomings: a limited number of subjects, a single recording session and only few expressions captured. To address these issues we collected the CMU Multi-PIE database. It contains 337 subjects, imaged under 15 view points and 19 illumination conditions in up to four recording sessions. In this paper we introduce the database and describe the recording procedure. We furthermore present results from baseline experiments using PCA and LDA classifiers to highlight similarities and differences between PIE and Multi-PIE. Ralph Gross, Iain A. Matthews, Jeffrey F. Cohn, Takeo Kanade, Simon Baker |
FG | 2 |
| 2008 | Comparing text-driven and speech-driven visual speech synthesisers
Barry-John Theobald, Gavin C. Cawley, J. Andrew Bangham, Iain A. Matthews, Nicholas Wilkinson |
INTERSPEECH | 4 |
| 2008 | Multi-View AAM Fitting and Construction
Krishnan Ramnath, Seth Koterba, Jing Xiao 0006, Changbo Hu, Iain A. Matthews, Simon Baker, Jeffrey F. Cohn, Takeo Kanade |
Int. J. Comput. Vis. | 5 |
| 2007 | Real-time expression cloning using appearance modelsabstractActive Appearance Models (AAMs) are generative parametric models commonly used to track, recognise and synthesise faces in images and video sequences. In this paper we describe a method for transferring dynamic facial gestures between subjects in real-time. The main advantages of our approach are that: 1) the mapping is computed automatically and does not require high-level semantic information describing facial expressions or visual speech gestures. 2) The mapping is simple and intuitive, allowing expressions to be transferred and rendered in real-time. 3) The mapped expression can be constrained to have the appearance of the target producing the expression, rather than the source expression imposed onto the target face. 4) Near-videorealistic talking faces for new subjects can be created without the cost of recording and processing a complete training corpus for each. Our system enables face-to-face interaction with an avatar driven by an AAM of an actual person in real-time and we show examples of arbitrary expressive speech frames cloned across different subjects. Barry-John Theobald, Iain A. Matthews, Jeffrey F. Cohn, Steven M. Boker |
ICMI | 2 |
| 2007 | 2D vs. 3D Deformable Face Models: Representational Power, Construction, and Real-Time Fitting
Iain A. Matthews, Jing Xiao 0006, Simon Baker |
Int. J. Comput. Vis. | 1 |
| 2006 | Active appearance models with occlusion
Ralph Gross, Iain A. Matthews, Simon Baker |
Image Vis. Comput. | 2 |
| 2005 | Multi-View AAM Fitting and Camera CalibrationabstractIn this paper, we study the relationship between multi-view active appearance model (AAM) fitting and camera calibration. In the first part of the paper we propose an algorithm to calibrate the relative orientation of a set of N > 1 cameras by fitting an AAM to sets of N images. In essence, we use the human face as a (non-rigid) calibration grid. Our algorithm calibrates a set of 2 /spl times/ 3 weak-perspective camera projection matrices, protections of the world coordinate system origin into the images, depths of the world coordinate system origin, and focal lengths. We demonstrate that the performance of this algorithm is comparable to a standard algorithm using a calibration grid. In the second part of the paper, we show how calibrating the cameras improves tile performance of multi-view AAM fitting. Seth Koterba, Simon Baker, Iain A. Matthews, Changbo Hu, Jing Xiao 0006, Jeffrey F. Cohn, Takeo Kanade |
ICCV | 3 |
| 2005 | Generic vs. person specific active appearance models
Ralph Gross, Iain A. Matthews, Simon Baker |
Image Vis. Comput. | 2 |
| 2004 | Generic vs. Person Specific Active Appearance ModelsabstractActive Appearance Models (AAMs) are generative parametric models that have been successfully used in the past to model faces. Anecdotal evidence, however, suggests that the performance of an AAM built to model the variation in appearance of a single person across pose, illumination, and expression (a Person Specific AAM) is substantially better than the performance of an AAM built to model the variation in appearance of many faces, including unseen subjects not in the training set (a Generic AAM). In this paper, we present an empirical evaluation that shows that Person Specific AAMs are, as expected, both easier to build and more robust to fit than Generic AAMs. Moreover, we show that: (1) building a generic shape model is far easier than building a generic appearance model, and (2) the shape component is the main cause of the reduced fitting robustness of Generic AAMs. We then proceed to describe two refinements to Generic AAMs to improve their performance: (1) a refitting procedure to improve the quality of the ground-truth data used to build the AAM and (2) a new fitting algorithm. For both refinements we demonstrate dramatically improved fitting performance. Finally, we evaluate the effect of these improvements on a combined model construction and fitting task. Ralph Gross, Iain A. Matthews, Simon Baker |
BMVC | 2 |
| 2004 | Fitting a Single Active Appearance Model Simultaneously to Multiple ImagesabstractActive Appearance Models (AAMs) are a well studied 2D deformable model. One recently proposed extension of AAMs to multiple images is the Coupled-View AAM. Coupled-View AAMs model the 2D shape and appearance of a face in two or more views simultaneously. The major limitation of Coupled-View AAMs, however, is that they are specific to a particular set of cameras, both in geometry and the photometric responses. In this paper, we describe how a single AAM can be fit to multiple images, captured simultaneously by cameras with arbitrary geometry and response functions. Our algorithm retains the major benefits of Coupled-View AAMs: the integration of information from multiple images into a single model, and improved fitting robustness. 1 Changbo Hu, Jing Xiao 0006, Iain A. Matthews, Simon Baker, Jeffrey F. Cohn, Takeo Kanade |
BMVC | 3 |
| 2004 | Real-Time Combined 2D+3D Active Appearance Models
Jing Xiao 0006, Simon Baker, Iain A. Matthews, Takeo Kanade |
CVPR (2) | 3 |
| 2004 | Lucas-Kanade 20 Years On: A Unifying Framework
Simon Baker, Iain A. Matthews |
Int. J. Comput. Vis. | 2 |
| 2004 | Active Appearance Models Revisited
Iain A. Matthews, Simon Baker |
Int. J. Comput. Vis. | 1 |
| 2004 | Automatic Construction of Active Appearance Models as an Image Coding ProblemabstractThe automatic construction of Active Appearance Models (AAMs) is usually posed as finding the location of the base mesh vertices in the input training images. In this paper, we repose the problem as an energy-minimizing image coding problem and propose an efficient gradient-descent algorithm to solve it. Simon Baker, Iain A. Matthews, Jeff G. Schneider |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | Appearance-Based Face Recognition and Light-FieldsabstractArguably the most important decision to be made when developing an object recognition algorithm is selecting the scene measurements or features on which to base the algorithm. In appearance-based object recognition, the features are chosen to be the pixel intensity values in an image of the object. These pixel intensities correspond directly to the radiance of light emitted from the object along certain rays in space. The set of all such radiance values over all possible rays is known as the plenoptic function or light-field. In this paper, we develop a theory of appearance-based object recognition from light-fields. This theory leads directly to an algorithm for face recognition across pose that uses as many images of the face as are available, from one upwards. All of the pixels, whichever image they come from, are treated equally and used to estimate the (eigen) light-field of the object. The eigen light-field is then used as the set of features on which to base recognition, analogously to how the pixel intensities are used in appearance-based face and object recognition. Ralph Gross, Iain A. Matthews, Simon Baker |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | The Template Update ProblemabstractTemplate tracking dates back to the 1981 Lucas-Kanade algorithm. One question that has received very little attention, however, is how to update the template so that it remains a good model of the tracked object. We propose a template update algorithm that avoids the "drifting" inherent in the naive algorithm. Iain A. Matthews, Takahiro Ishikawa, Simon Baker |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Near-videorealistic synthetic talking faces: implementation and evaluation
Barry-John Theobald, J. Andrew Bangham, Iain A. Matthews, Gavin C. Cawley |
Speech Commun. | 3 |
| 2003 | The Template Update ProblemabstractTemplate tracking dates back to the 1981 Lucas-Kanade algorithm. One question that has received very little attention, however, is how to update the template so that it remains a good model of the tracked object. We propose a template update algorithm that avoids the "drifting" inherent in the naive algorithm. Iain A. Matthews, Takahiro Ishikawa, Simon Baker |
BMVC | 1 |
| 2003 | 2.5D Visual Speech Synthesis Using Appearance ModelsabstractTwo dimensional (2D) shape and appearance models are applied to the problem of creating a near-videorealistic talking head. A speech corpus of a talker uttering a set of phonetically balanced training sentences is analysed using a generative model of the human face. Segments of original parameter trajectories, corresponding to the synthesis unit (e.g.~triphone), are extracted from a codebook, then normalised, blended, concatenated and smoothed before being applied to the model to give natural, realistic animations of novel utterances. The system provides a 2D image sequence corresponding to the face of a talker. It is also used to animate the face of a 3D avatar by displacing the mesh according to movements of points in the shape model and dynamically texturing the face polygons using the appearance model. Barry-John Theobald, J. Andrew Bangham, Iain A. Matthews, John R. W. Glauert, Gavin C. Cawley |
BMVC | 3 |
| 2003 | Near-videorealistic synthetic visual speech using non-rigid appearance modelsabstractWe present work towards videorealistic synthetic visual speech using non-rigid appearance models. These models are used to track a talking face enunciating a set of training sentences. The resultant parameter trajectories are used in a concatenative synthesis scheme, where samples of original data are extracted from a corpus and concatenated to form new unseen sequences. Here we explore the effect on the synthesiser output of blending several synthesis units considered similar to the desired unit. We present preliminary subjective and objective results used to judge the realism of the system. Barry-John Theobald, Gavin C. Cawley, Iain A. Matthews, J. Andrew Bangham |
ICASSP (5) | 3 |
| 2002 | Towards video realistic synthetic visual speechabstractIn this paper we present initial work towards a video-realistic visual speech synthesiser based on statistical models of shape and appearance. A synthesised image sequence corresponding to an utterance is formed by concatenation of synthesis units (in this case phonemes) from a pre-recorded corpus of training data. A smoothing spline is applied to the concatenated parameters to ensure smooth transitions between frames and the resultant parameters applied to the model—early results look promising. Barry-John Theobald, J. Andrew Bangham, Iain A. Matthews, Gavin C. Cawley |
ICASSP | 3 |
| 2002 | Extraction of Visual Features for LipreadingabstractThe multimodal nature of speech is often ignored in human-computer interaction, but lip deformations and other body motion, such as those of the head, convey additional information. We integrate speech cues from many sources and this improves intelligibility, especially when the acoustic signal is degraded. The paper shows how this additional, often complementary, visual speech information can be used for speech recognition. Three methods for parameterizing lip image sequences for recognition using hidden Markov models are compared. Two of these are top-down approaches that fit a model of the inner and outer lip contours and derive lipreading features from a principal component analysis of shape or shape and appearance, respectively. The third, bottom-up, method uses a nonlinear scale-space analysis to form features directly from the pixel intensity. All methods are compared on a multitalker visual speech recognition task of isolated letters. Iain A. Matthews, Timothy F. Cootes, J. Andrew Bangham, Stephen J. Cox, Richard W. Harvey |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2001 | Equivalence and Efficiency of Image Alignment AlgorithmsabstractThere are two major formulations of image alignment using gradient descent. The first estimates an additive increment to the parameters (the additive approach), the second an incremental warp (the compositional approach). We first prove that these two formulations are equivalent. A very efficient algorithm was proposed by Hager and Belhumeur (1998) using the additive approach that unfortunately can only be applied to a very restricted class of warps. We show that using the compositional approach an equally efficient algorithm (the inverse compositional algorithm) can be derived that can be applied to any set of warps which form a group. While most warps used in computer vision form groups, there are a certain warps that do not. Perhaps most notable is the set of piecewise affine warps used in flexible appearance models (FAMs). We end this paper by extending the inverse compositional algorithm to apply to FAMs. Simon Baker, Iain A. Matthews |
CVPR (1) | 2 |
| 2001 | Tele-Graffiti: A Pen and Paper-Based Remote Sketching SystemabstractTele-Graffiti is a system allowing two.or more users to communicate remotely via hand-drawn sketches. What one person writes at one site is captured using a video camera, transmitted to the other site(s), and displayed there using an LCD projector. The advantage of our system over other intelligent desktops and white-boards is that the users are free to move the pieces of paper on which they are writing. In Tele-Graffiti, paper detection and tracking is based on real-time paper boundary detection. Naoya Takao, Jianbo Shi, Simon Baker, Iain A. Matthews, Bart C. Nabbe |
ICCV | 4 |
| 2001 | A Comparison Of Model And Transform-Based Visual Features For Audio-Visual LVCSRabstractFour different visual speech parameterisation methods are compared on a large vocabulary, continuous, audio-visual speech recognition task using the IBM ViaVoice TM audio-visual speech database. Three are direct mouth image region based transforms; discrete cosine and wavelet transforms, and principal component analysis. The fourth uses a statistical model of shape and appearance called an active appearance model, to track and obtain model parameters describing the entire face. All parameterisations are compared experimentally using hidden Markov models (HMM's) in a speaker independent test. Visualonly HMM's are used to rescore lattices obtained from audio models trained in noisy conditions. 1. Iain A. Matthews, Gerasimos Potamianos, Chalapathy Neti, Jürgen Lüttin |
ICME | 1 |
| 2001 | Large-vocabulary audio-visual speech recognition: a summary of the Johns Hopkins Summer 2000 WorkshopabstractWe report a summary of the Johns Hopkins Summer 2000 Workshop on audio-visual automatic speech recognition (ASR) in the large-vocabulary, continuous speech domain. Two problems of audio-visual ASR were mainly addressed: visual feature extraction and audio-visual information fusion. First, image transform and model-based visual features were considered, obtained by means of the discrete cosine transform (DCT) and active appearance models, respectively. The former were demonstrated to yield superior automatic speech reading. Subsequently, a number of feature fusion and decision fusion techniques for combining the DCT visual features with traditional acoustic ones were implemented and compared. Hierarchical discriminant feature fusion and asynchronous decision fusion by means of the multi-stream hidden Markov model consistently improved ASR for both clean and noisy speech. Compared to an equivalent audio-only recognizer, introducing the visual modality reduced ASR word error rate by 7% relative in clean speech, and by 27% relative at an 8.5 dB SNR audio condition. Chalapathy Neti, Gerasimos Potamianos, Jürgen Lüttin, Iain A. Matthews, Hervé Glotin, Dimitra Vergyri |
MMSP | 4 |
| 1998 | A Comparison of Active Shape Model and Scale Decomposition Based Features for Visual Speech Recognition
Iain A. Matthews, J. Andrew Bangham, Richard W. Harvey, Stephen J. Cox |
ECCV (2) | 1 |
| 1997 | Lip reading from scale-space measurementsabstractSystems that attempt to recover the spoken word from image sequences usually require complicated models of the mouth and its motions. Here we describe a new approach based on a fast mathematical morphology transform called the sieve. We form statistics of scale measurements in one and two dimensions and these are used as a feature vector for standard Hidden Markov Models (HMMs). Richard W. Harvey, Iain A. Matthews, J. Andrew Bangham, Stephen J. Cox |
CVPR | 2 |
| 1996 | Audiovisual speech recognition using multiscale nonlinear image decomposition
Iain A. Matthews, J. Andrew Bangham, Stephen J. Cox |
ICSLP | 1 |