EDBT 2026 Demo / reviewers in the wild / expert
Tadas Baltrusaitis
dblp:23/9606
· DBLP profile ↗
49ranked-venue papers
7as first author
12since 2021 · last 2025
0000-0001-7923-8780ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 27 · 5 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 15 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GASP: Gaussian Avatars with Synthetic PriorsabstractGaussian Splatting has changed the game for real-time photo-realistic rendering. One of the most popular applications of Gaussian Splatting is to create animatable avatars, known as Gaussian Avatars. Recent works have pushed the boundaries of quality and rendering efficiency but suffer from two main limitations. Either they require expensive multi-camera rigs to produce avatars with free-viewpoint rendering, or they can be trained with a single camera but only rendered at high quality from this fixed viewpoint. An ideal model would be trained using a short monocular video or image from available hardware, such as a webcam, and rendered from any view. To this end, we propose GASP: Gaussian Avatars with Synthetic Priors. To overcome the limitations of existing datasets, we exploit the pixel-perfect nature of synthetic data to train a Gaussian Avatar prior. By fitting this prior model to a single photo or video and fine-tuning it, we get a high-quality Gaussian Avatar, which supports 360° rendering. Our prior is only required for fitting, not inference, enabling real-time applications. Through our method, we obtain high-quality, animatable Avatars from limited data which can be animated and rendered at 70fps on commercial hardware. Jack R. Saunders, Charlie Hewitt, Yanan Jian, Marek Kowalski, Tadas Baltrusaitis, Yiye Chen, Darren Cosker, Virginia Estellers, Nicholas Gyde, Vinay P. Namboodiri, Ben Lundell |
CVPR | 5 |
| 2025 | DAViD: Data-Efficient and Accurate Vision Models from Synthetic Data DAViD also references Michelangelo's David - an iconic symbol of anatomical precision-and the David vs. Goliath story, reflecting our small yet powerful dataset and models
Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Charlie Hewitt, Lohit Petikam, Xian Xiao, Antonio Criminisi, Thomas J. Cashman 0001, Tadas Baltrusaitis |
ICCV | 8 |
| 2024 | SimpleEgo: Predicting Probabilistic Body Pose from Egocentric CamerasabstractOur work addresses the problem of egocentric human pose estimation from downwards-facing cameras on head-mounted devices (HMD). This presents a challenging scenario, as parts of the body often fall outside of the image or are occluded. Previous solutions minimize this problem by using fish-eye camera lenses to capture a wider view, but these can present hardware design issues. They also predict 2D heat-maps per joint and lift them to 3D space to deal with self-occlusions, but this requires large network architectures which are impractical to deploy on resource-constrained HMDs. We predict pose from images captured with conventional rectilinear camera lenses. This resolves hardware design issues, but means body parts are often out of frame. As such, we directly regress probabilistic joint rotations represented as matrix Fisher distributions for a parameterized body model. This allows us to quantify pose uncertainties and explain out-of-frame or occluded joints. This also removes the need to compute 2D heat-maps and allows for simplified DNN architectures which require less compute. Given the lack of egocentric datasets using rectilinear camera lenses, we introduce the SynthEgo dataset1, a synthetic dataset with 60K stereo images containing high diversity of pose, shape, clothing and skin tone. Our approach achieves state-of-the-art results for this challenging configuration, reducing mean per-joint position error by 23% overall and 58% for the lower body. Our architecture also has eight times fewer parameters and runs twice as fast as the current state-of-the-art. Experiments show that training on our synthetic dataset leads to good generalization to real world images without fine-tuning.1Available at https://microsoft.github.io/SimpleEgo. Hanz Cuevas-Velasquez, Charlie Hewitt, Mohammad Sadegh Ali Akbarian, Tadas Baltrusaitis |
3DV | 4 |
| 2024 | Hairmony: Fairness-aware hairstyle classificationabstractWe present a method for prediction of a person's hairstyle from a single image. Despite growing use cases in user digitization and enrollment for virtual experiences, available methods are limited, particularly in the range of hairstyles they can capture. Human hair is extremely diverse and lacks any universally accepted description or categorization, making this a challenging task. Most current methods rely on parametric models of hair at a strand level. These approaches, while very promising, are not yet able to represent short, frizzy, coily hair and gathered hairstyles. We instead choose a classification approach which can represent the diversity of hairstyles required for a truly robust and inclusive system. Previous classification approaches have been restricted by poorly labeled data that lacks diversity, imposing constraints on the usefulness of any resulting enrollment system. We use only synthetic data to train our models. This allows for explicit control of diversity of hairstyle attributes, hair colors, facial appearance, poses, environments and other parameters. It also produces noise-free ground-truth labels. We introduce a novel hairstyle taxonomy developed in collaboration with a diverse group of domain experts which we use to balance our training data, supervise our model, and directly measure fairness. We annotate our synthetic training data and a real evaluation dataset using this taxonomy and release both to enable comparison of future hairstyle prediction approaches. We employ an architecture based on a pre-trained feature extraction network in order to improve generalization of our method to real data and predict taxonomy attributes as an auxiliary task to improve accuracy. Results show our method to be significantly more robust for challenging hairstyles than recent parametric approaches. Givi Meishvili, James Clemoes, Charlie Hewitt, Zafiirah Hosenie, Xian Xiao, Martin de La Gorce, Tadas Baltrusaitis, Antonio Criminisi, Chyna McRae, Nina Jablonski, Marta Wilczkowiak |
SIGGRAPH Asia | 8 |
| 2024 | Look Ma, no markers: holistic performance capture without the hassleabstractWe tackle the problem of highly-accurate, holistic performance capture for the face, body and hands simultaneously. Motion-capture technologies used in film and game production typically focus only on face, body or hand capture independently, involve complex and expensive hardware and a high degree of manual intervention from skilled operators. While machine-learning-based approaches exist to overcome these problems, they usually only support a single camera, often operate on a single part of the body, do not produce precise world-space results, and rarely generalize outside specific contexts. In this work, we introduce the first technique for markerfree, high-quality reconstruction of the complete human body, including eyes and tongue, without requiring any calibration, manual intervention or custom hardware. Our approach produces stable world-space results from arbitrary camera rigs as well as supporting varied capture environments and clothing. We achieve this through a hybrid approach that leverages machine learning models trained exclusively on synthetic data and powerful parametric models of human shape and motion. We evaluate our method on a number of body, face and hand reconstruction benchmarks and demonstrate state-of-the-art results that generalize on diverse datasets. Charlie Hewitt, Fatemehsadat Saleh, Mohammad Sadegh Ali Akbarian, Lohit Petikam, Shideh Rezaeifar, Louis Florentin, Zafiirah Hosenie, Thomas J. Cashman 0001, Julien Valentin, Darren Cosker, Tadas Baltrusaitis |
ACM Trans. Graph. | 11 |
| 2023 | RODIN: A Generative Model for Sculpting 3D Digital Avatars Using DiffusionabstractThis paper presents a 3D diffusion model that automatically generates 3D digital avatars represented as neural radiance fields (NeRFs). A significant challenge for 3D diffusion is that the memory and processing costs are prohibitive for producing high-quality results with rich details. To tackle this problem, we propose the roll-out diffusion network (RODIN), which takes a 3D NeRF model represented as multiple 2D feature maps and rolls out them onto a single 2D feature plane within which we perform 3D-aware diffusion. The RODIN model brings much-needed computational efficiency while preserving the integrity of 3D diffusion by using 3D-aware convolution that attends to projected features in the 2D plane according to their original relationships in 3D. We also use latent conditioning to orchestrate the feature generation with global coherence, leading to high-fidelity avatars and enabling semantic editing based on text prompts. Finally, we use hierarchical synthesis to further enhance details. The 3D avatars generated by our model compare favorably with those produced by existing techniques. We can generate highly detailed avatars with realistic hairstyles and facial hair. We also demonstrate 3D avatar generation from image or text, as well as text-guided editability. Tengfei Wang 0002, Bo Zhang 0025, Ting Zhang 0002, Shuyang Gu, Jianmin Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen 0003, Fang Wen 0001, Qifeng Chen 0001, Baining Guo |
CVPR | 6 |
| 2023 | HiFace: High-Fidelity 3D Face Reconstruction by Learning Static and Dynamic Detailsabstract3D Morphable Models (3DMMs) demonstrate great potential for reconstructing faithful and animatable 3D facial surfaces from a single image. The facial surface is influenced by the coarse shape, as well as the static detail (e.g., person-specific appearance) and dynamic detail (e.g., expression-driven wrinkles). Previous work struggles to decouple the static and dynamic details through image-level supervision, leading to reconstructions that are not realistic. In this paper, we aim at high-fidelity 3D face reconstruction and propose HiFace to explicitly model the static and dynamic details. Specifically, the static detail is modeled as the linear combination of a displacement basis, while the dynamic detail is modeled as the linear interpolation of two displacement maps with polarized expressions. We exploit several loss functions to jointly learn the coarse shape and fine details with both synthetic and real-world datasets, which enable HiFace to reconstruct high-fidelity 3D shapes with animatable details. Extensive quantitative and qualitative experiments demonstrate that HiFace presents state-of-the-art reconstruction quality and faithfully recovers both the static and dynamic details. Our project page: https://project-hiface.github.io. Zenghao Chai, Tianke Zhang, Tianyu He, Xu Tan 0003, Tadas Baltrusaitis, HsiangTao Wu, Runnan Li, Sheng Zhao 0002, Chun Yuan 0003, Jiang Bian 0002 |
ICCV | 5 |
| 2023 | DigiFace-1M: 1 Million Digital Face Images for Face RecognitionabstractState-of-the-art face recognition models show impressive accuracy, achieving over 99.8% on Labeled Faces in the Wild (LFW) dataset. Such models are trained on large-scale datasets that contain millions of real human face images collected from the internet. Web-crawled face images are severely biased (in terms of race, lighting, makeup, etc) and often contain label noise. More importantly, the face images are collected without explicit consent, raising ethical concerns. To avoid such problems, we introduce a large-scale synthetic dataset for face recognition, obtained by rendering digital faces using a computer graphics pipeline1. We first demonstrate that aggressive data augmentation can significantly reduce the synthetic-to-real domain gap. Having full control over the rendering pipeline, we also study how each attribute (e.g., variation in facial pose, accessories and textures) affects the accuracy. Compared to Syn-Face, a recent method trained on GAN-generated synthetic faces, we reduce the error rate on LFW by 52.5% (accuracy from 91.93% to 96.17%). By fine-tuning the network on a smaller number of real face images that could reason-ably be obtained with consent, we achieve accuracy that is comparable to the methods trained on millions of real face images. Gwangbin Bae, Martin de La Gorce, Tadas Baltrusaitis, Charlie Hewitt, Dong Chen 0003, Julien P. C. Valentin, Roberto Cipolla, Jingjing Shen |
WACV | 3 |
| 2023 | Mesh-Tension Driven Expression-Based Wrinkles for Synthetic FacesabstractRecent advances in synthesizing realistic faces have shown that synthetic training data can replace real data for various face-related computer vision tasks. A question arises: how important is realism? Is the pursuit of photorealism excessive? In this work, we show otherwise. We boost the realism of our synthetic faces by introducing dynamic skin wrinkles in response to facial expressions, and observe significant performance improvements in downstream computer vision tasks. Previous approaches for producing such wrinkles either required prohibitive artist effort to scale across identities and expressions, or were not capable of reconstructing high-frequency skin details with sufficient fidelity. Our key contribution is an approach that produces realistic wrinkles across a large and diverse population of digital humans. Concretely, we formalize the concept of mesh-tension and use it to aggregate possible wrinkles from high-quality expression scans into albedo and displacement texture maps. At synthesis, we use these maps to produce wrinkles even for expressions not represented in the source scans. Additionally, to provide a more nuanced indicator of model performance under deformations resulting from com-pressed expressions, we introduce the 300W-winks evaluation subset and the Pexels dataset of closed eyes and winks. Chirag Raman, Charlie Hewitt, Erroll Wood, Tadas Baltrusaitis |
WACV | 4 |
| 2022 | 3D Face Reconstruction with Dense Landmarks
Erroll Wood, Tadas Baltrusaitis, Charlie Hewitt, Matthew Johnson 0003, Jingjing Shen, Nikola Milosavljevic, Daniel Wilde, Stephan J. Garbin, Toby Sharp, Ivan Stojiljkovic, Thomas J. Cashman 0001, Julien P. C. Valentin |
ECCV (13) | 2 |
| 2022 | SCAMPS: Synthetics for Camera Measurement of Physiological SignalsabstractThe use of cameras and computational algorithms for noninvasive, low-cost and scalable measurement of physiological (e.g., cardiac and pulmonary) vital signs is very attractive. However, diverse data representing a range of environments, body motions, illumination conditions and physiological states is laborious, time consuming and expensive to obtain. Synthetic data have proven a valuable tool in several areas of machine learning, yet are not widely available for camera measurement of physiological states. Synthetic data offer "perfect" labels (e.g., without noise and with precise synchronization), labels that may not be possible to obtain otherwise (e.g., precise pixel level segmentation maps) and provide a high degree of control over variation and diversity in the dataset. We present SCAMPS, a dataset of synthetics containing 2,800 videos (1.68M frames) with aligned cardiac and respiratory signals and facial action intensities. The RGB frames are provided alongside segmentation maps and precise descriptive statistics about the underlying waveforms, including inter-beat interval, heart rate variability, and pulse arrival time. Finally, we present baseline results training on these synthetic data and testing on real-world datasets to illustrate generalizability. Daniel McDuff, Miah Wander, Xin Liu 0034, Brian L. Hill, Javier Hernandez, Jonathan Lester, Tadas Baltrusaitis |
NeurIPS | 7 |
| 2021 | Fake it till you make it: face analysis in the wild using synthetic data aloneabstractWe demonstrate that it is possible to perform face-related computer vision in the wild using synthetic data alone. The community has long enjoyed the benefits of synthesizing training data with graphics, but the domain gap between real and synthetic data has remained a problem, especially for human faces. Researchers have tried to bridge this gap with data mixing, domain adaptation, and domain-adversarial training, but we show that it is possible to synthesize data with minimal domain gap, so that models trained on synthetic data generalize to real in-the-wild datasets. We describe how to combine a procedurally-generated parametric 3D face model with a comprehensive library of hand-crafted assets to render training images with unprecedented realism and diversity. We train machine learning systems for face-related tasks such as landmark localization and face parsing, showing that synthetic data can both match real data in accuracy as well as open up new approaches where manual labeling would be impossible. Erroll Wood, Tadas Baltrusaitis, Charlie Hewitt, Sebastian Dziadzio, Thomas J. Cashman 0001, Jamie Shotton |
ICCV | 2 |
| 2020 | CONFIG: Controllable Neural Face Image Generation
Marek Kowalski, Stephan J. Garbin, Virginia Estellers, Tadas Baltrusaitis, Matthew Johnson 0003, Jamie Shotton |
ECCV (11) | 4 |
| 2020 | Attended End-to-End Architecture for Age Estimation From Facial Expression VideosabstractThe main challenges of age estimation from facial expression videos lie not only in the modeling of the static facial appearance, but also in the capturing of the temporal facial dynamics. Traditional techniques to this problem focus on constructing handcrafted features to explore the discriminative information contained in facial appearance and dynamics separately. This relies on sophisticated feature-refinement and framework-design. In this paper, we present an end-toend architecture for age estimation, called Spatially-Indexed Attention Model (SIAM), which is able to simultaneously learn both the appearance and dynamics of age from raw videos of facial expressions. Specifically, we employ convolutional neural networks to extract effective latent appearance representations and feed them into recurrent networks to model the temporal dynamics. More importantly, we propose to leverage attention models for salience detection in both the spatial domain for each single image and the temporal domain for the whole video as well. We design a specific spatially-indexed attention mechanism among the convolutional layers to extract the salient facial regions in each individual image, and a temporal attention layer to assign attention weights to each frame. This two-pronged approach not only improves the performance by allowing the model to focus on informative frames and facial areas, but it also offers an interpretable correspondence between the spatial facial regions as well as temporal frames, and the task of age estimation. We demonstrate the strong performance of our model in experiments on a large, gender-balanced database with 400 subjects with ages spanning from 8 to 76 years. Experiments reveal that our model exhibits significant superiority over the state-of-the-art methods given sufficient training data. Wenjie Pei, Hamdi Dibeklioglu, Tadas Baltrusaitis, David M. J. Tax |
IEEE Trans. Image Process. | 3 |
| 2019 | Multimodal Machine Learning: A Survey and TaxonomyabstractOur experience of the world is multimodal - we see objects, hear sounds, feel texture, smell odors, and taste flavors. Modality refers to the way in which something happens or is experienced and a research problem is characterized as multimodal when it includes multiple such modalities. In order for Artificial Intelligence to make progress in understanding the world around us, it needs to be able to interpret such multimodal signals together. Multimodal machine learning aims to build models that can process and relate information from multiple modalities. It is a vibrant multi-disciplinary field of increasing importance and with extraordinary potential. Instead of focusing on specific multimodal applications, this paper surveys the recent advances in multimodal machine learning itself and presents them in a common taxonomy. We go beyond the typical early and late fusion categorization and identify broader challenges that are faced by multimodal machine learning, namely: representation, translation, alignment, fusion, and co-learning. This new taxonomy will enable researchers to better understand the state of the field and identify directions for future research. Tadas Baltrusaitis, Chaitanya Ahuja, Louis-Philippe Morency |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | The ASC-Inclusion Perceptual Serious Gaming Platform for Autistic Childrenabstract“Serious games” are becoming extremely relevant to individuals who have specific needs, such as children with an autism spectrum condition (ASC). Often, individuals with an ASC have difficulties in interpreting verbal and nonverbal communication cues during social interactions. The ASC-Inclusion EU-FP7 funded project aims to provide children who have an ASC with a platform to learn emotion expression and recognition, through play in the virtual world. In particular, the ASC-Inclusion platform focuses on the expression of emotion via facial, vocal, and bodily gestures. The platform combines multiple analysis tools, using onboard microphone and webcam capabilities. The platform utilizes these capabilities via training games, text-based communication, animations, video, and audio clips. This paper introduces current findings and evaluations of the ASC-Inclusion platform and provides detailed description for the different modalities. Erik Marchi, Tadas Baltrusaitis, Andra Adams, Marwa Mahmoud, Ofer Golan, Shimrit Fridenson-Hayo, Shahar Tal, Shai Newman, Noga Meir-Goren, Antonio Camurri, Stefano Piana, Björn W. Schuller, Sven Bölte, Tevfik Metin Sezgin, Nese Alyüz, Agnieszka Rynkiewicz, Aurelie Baranger, Alice Baird, Simon Baron-Cohen, Amandine Lassalle, Helen O'Reilly, Delia Pigat, Peter Robinson 0001, Ian Davies |
IEEE Trans. Games | 2 |
| 2018 | OpenFace 2.0: Facial Behavior Analysis ToolkitabstractOver the past few years, there has been an increased interest in automatic facial behavior analysis and understanding. We present OpenFace 2.0 - a tool intended for computer vision and machine learning researchers, affective computing community and people interested in building interactive applications based on facial behavior analysis. OpenFace 2.0 is an extension of OpenFace toolkit and is capable of more accurate facial landmark detection, head pose estimation, facial action unit recognition, and eye-gaze estimation. The computer vision algorithms which represent the core of OpenFace 2.0 demonstrate state-of-the-art results in all of the above mentioned tasks. Furthermore, our tool is capable of real-time performance and is able to run from a simple webcam without any specialist hardware. Finally, unlike a lot of modern approaches or toolkits, OpenFace 2.0 source code for training models and running them is freely available for research purposes. Tadas Baltrusaitis, Amir Zadeh 0001, Yao Chong Lim, Louis-Philippe Morency |
FG | 1 |
| 2018 | Toward Visual Behavior Markers of Suicidal IdeationabstractSuicide is an increasingly present issue in our society whose eradication could be greatly aided by decision support technologies that can objectively identify behavior markers of suicidal ideation. In this paper, we examine the predictive ability of a variety of smiling and eye gaze behaviors in categorizing hospital patients by mental health status: patients with suicidal ideation, patients with other mental illnesses such as depression, or control group without suicidal ideation or mental illness. We study three main research questions related to suicide behavior markers: (1) Do people with suicidal ideation smile with different dynamics (e.g. genuine vs fake smile)? (2) Do smiles while speaking, listening, and laughing show different levels of occurrence between the three groups? (3) Is gaze aversion (e.g. looking down) also a useful behavior marker? To answer these questions, we created new behavioral annotations on 74 semi-structured interviews from hospital patients, each of them within one of the three mental health conditions. Our data analysis identified behavior markers of mental health status from both smiling and eye gaze behaviors. Using these behavioral features, we created predictive models that show promising results when distinguishing between these three mental health conditions, especially when differentiating suicidal from non-suicidal patients. Naomi Eigbe, Tadas Baltrusaitis, Louis-Philippe Morency, John Pestian |
FG | 2 |
| 2018 | Edge Convolutional Network for Facial Action Intensity EstimationabstractIn this paper, we propose a novel convolutional neural architecture for facial action unit intensity estimation. While Convolutional Neural Networks (CNNs) have shown great promise in a wide range of computer vision tasks, these achievements have not translated as well to facial expression analysis, with hand crafted features (e.g. the Histogram of Orientated Gradient) still being very competitive. We introduce a novel Edge Convolutional Network (ECN) that is able to capture subtle changes in facial appearance. Our model is able to learn edge-like detectors that can capture subtle wrinkles and facial muscle contours at multiple orientations and frequencies. The core novelty of our ECN model is in its first layer which integrates three main components: an edge filter generator, a receptive gate and a filter rotator. All the components are differentiable and our ECN model is end-to-end trainable and learns the important edge detectors for facial expression analysis. Experiments on two facial action unit datasets show that the proposed ECN outperforms state-of-the-art methods for both AU intensity estimation tasks. Liandong Li, Tadas Baltrusaitis, Bo Sun 0006, Louis-Philippe Morency |
FG | 2 |
| 2018 | GazeDirector: Fully Articulated Eye Gaze Redirection in VideoabstractAbstract We present GazeDirector, a new approach for eye gaze redirection that uses model‐fitting. Our method first tracks the eyes by fitting a multi‐part eye region model to video frames using analysis‐by‐synthesis, thereby recovering eye region shape, texture, pose, and gaze simultaneously. It then redirects gaze by 1) warping the eyelids from the original image using a model‐derived flow field, and 2) rendering and compositing synthesized 3D eyeballs onto the output image in a photorealistic manner. GazeDirector allows us to change where people are looking without person‐specific training data, and with full articulation, i.e. we can precisely specify new gaze directions in 3D. Quantitatively, we evaluate both model‐fitting and gaze synthesis, with experiments for gaze estimation and redirection on the Columbia gaze dataset. Qualitatively, we compare GazeDirector against recent work on gaze redirection, showing better results especially for large redirection angles. Finally, we demonstrate gaze redirection on YouTube videos by introducing new 3D gaze targets and by manipulating visual behavior. Erroll Wood, Tadas Baltrusaitis, Louis-Philippe Morency, Peter Robinson 0001, Andreas Bulling |
Comput. Graph. Forum | 2 |
| 2017 | Integrating Verbal and Nonvebval Input into a Dynamic Response Spoken Dialogue SystemabstractIn this work, we present a dynamic response spoken dialogue system (DRSDS). It is capable of understanding the verbal and nonverbal language of users and making instant, situation-aware response. Incorporating with two external systems, MultiSense and email summarization, we built an email reading agent on mobile device to show the functionality of DRSDS. Ting-Yao Hu, Chirag Raman, Salvador Medina Maza, Liangke Gui, Tadas Baltrusaitis, Robert E. Frederking, Louis-Philippe Morency, Alan W. Black, Maxine Eskénazi |
AAAI | 5 |
| 2017 | Local-global ranking for facial expression intensity estimationabstractFacial action units provide an objective characterization of facial muscle movements. Automatic estimation of facial action unit intensities is a challenging problem given individual differences in neutral face appearances and the need to generalize across different pose, illumination and datasets. In this paper, we introduce the Local-Global Ranking method as a novel alternative to direct prediction of facial action unit intensities. Our method takes advantage of the additional information present in videos and image collections of the same person (e.g. a photo album). Instead of trying to estimate facial expression intensities independently for each image, our proposed method performs a two-stage ranking: a local pair-wise ranking followed by a global ranking. The local ranking is designed to be accurate and robust by making a simple 3-class comparison (higher, equal, or lower) between randomly sampled pairs of images. We use a Bayesian model to integrate all these pair-wise rankings and construct a global ranking. Our Local-Global Ranking method shows state-of-the-art performance on two publicly-available datasets. Our cross-dataset experiments also show better generalizability. Tadas Baltrusaitis, Liandong Li, Louis-Philippe Morency |
ACII | 1 |
| 2017 | Hand2Face: Automatic synthesis and recognition of hand over face occlusionsabstractA person's face discloses important information about their affective state. Although there has been extensive research on recognition of facial expressions, the performance of existing approaches is challenged by facial occlusions. Facial occlusions are often treated as noise and discarded in recognition of affective states. However, hand over face occlusions can provide additional information for recognition of some affective states such as curiosity, frustration and boredom. One of the reasons that this problem has not gained attention is the lack of naturalistic occluded faces that contain hand over face occlusions as well as other types of occlusions. Traditional approaches for obtaining affective data are time demanding and expensive, which limits researchers in affective computing to work on small datasets. This limitation affects the generalizability of models and deprives researchers from taking advantage of recent advances in deep learning that have shown great success in many fields but require large volumes of data. In this paper, we first introduce a novel framework for synthesizing naturalistic facial occlusions from an initial dataset of non-occluded faces and separate images of hands, reducing the costly process of data collection and annotation. We then propose a model for facial occlusion type recognition to differentiate between hand over face occlusions and other types of occlusions such as scarves, hair, glasses and objects. Finally, we present a model to localize hand over face occlusions and identify the occluded regions of the face. Behnaz Nojavanasghari, Charles E. Hughes, Tadas Baltrusaitis, Louis-Philippe Morency |
ACII | 3 |
| 2017 | Visual attention in schizophrenia: Eye contact and gaze aversion during clinical interactionsabstractMany of the essential clues to the psychiatric condition of an individual lie within the nonverbal and communicative behavior patterns they express during social interactions. Unfortunately, these behaviors are particularly difficult to assess subjectively in a time-constrained environment, to which clinicians are often limited in realistic settings. The present analysis examines quantified patterns of gaze aversion across a set of persons recently admitted to an inpatient psychotic disorder unit at a major psychiatric hospital. These patterns are used to inform the development of discriminative models with the task of predicting schizophrenic symptom severity from both a typological and a dimensional assessment perspective. The results expose a novel set of gaze aversion behaviors distinguishing between positive subtype schizophrenia, characterized by excessive behaviors such as hallucinations and grandiosity, and negative subtype schizophrenia, characterized by diminished behaviors such as blunted affect and emotional withdrawal. The predictive models constitute a significant step toward the development of automated tools to aid medical professionals in the diagnosis of psychotic disorders. Alexandria K. Vail, Tadas Baltrusaitis, Luciana Pennant, Elizabeth S. Liebson, Justin T. Baker, Louis-Philippe Morency |
ACII | 2 |
| 2017 | Temporal Attention-Gated Model for Robust Sequence ClassificationabstractTypical techniques for sequence classification are designed for well-segmented sequences which have been edited to remove noisy or irrelevant parts. Therefore, such methods cannot be easily applied on noisy sequences expected in real-world applications. In this paper, we present the Temporal Attention-Gated Model (TAGM) which integrates ideas from attention models and gated recurrent networks to better deal with noisy or unsegmented sequences. Specifically, we extend the concept of attention model to measure the relevance of each observation (time step) of a sequence. We then use a novel gated recurrent network to learn the hidden representation for the final prediction. An important advantage of our approach is interpretability since the temporal attention weights provide a meaningful value for the salience of each time step in the sequence. We demonstrate the merits of our TAGM approach, both for prediction accuracy and interpretability, on three different tasks: spoken digit recognition, text-based sentiment analysis and visual event recognition. Wenjie Pei, Tadas Baltrusaitis, David M. J. Tax, Louis-Philippe Morency |
CVPR | 2 |
| 2017 | Curriculum Learning for Facial Expression RecognitionabstractOver the past few years, there has been an increased interest in machine understanding and recognition of affective states based on facial expressions. While great progress has been made, there are still a lot of challenges facing automatic emotion recognition, namely: generalizability of models across datasets, accounting for individual differences, and recognition of subtle expressions. While deep learning techniques enabled a large amount of progress in many areas of computer vision, this progress has not yet been fully translated to emotion recognition. Our work attempts to partly address that by presenting a novel learning technique for deep learning methods that leads to better generalization for emotion recognition from facial expressions. Liangke Gui, Tadas Baltrusaitis, Louis-Philippe Morency |
FG | 2 |
| 2017 | Investigating Facial Behavior Indicators of Suicidal IdeationabstractSuicide is the deliberate self-inflicted act with the intent to end one’s own life. It reflects both profound personal suffering and societal failure. While certain suicide risk factors are well understood, predicting suicide attempts remains a very challenging problem. In this paper, we investigate non-verbal facial behaviors to discriminate among control, mentally ill, and suicidal patients. For this task, we used a balanced corpus containing interviews of male and female patients with and without suicide ideation and/or mental health disorders from 3 different hospitals. In our experiments, we explored smiling, frowning, eyebrow raising, and head motion behaviors. We investigated both the occurrence of these behaviors and also how they were conducted. We found that facial behavior descriptors such as the percentage of smiles involving the contraction of the orbicularis oculi muscles (Duchenne smiles) had statistically significant differences between the suicidal and nonsuicidal groups. Our experiments also demonstrated that the stage of the interview in which these facial behaviors occur impacts their discriminative power. Eugene Laksana, Tadas Baltrusaitis, Louis-Philippe Morency, John Pestian |
FG | 2 |
| 2017 | Constrained Ensemble Initialization for Facial Landmark Tracking in VideoabstractAccurate and robust facial landmark tracking is a crucial step for face recognition and affect analysis systems. We often want to not only detect facial landmarks in images but to be able to track them reliably and consistently over time. Recently there has been an increase in research interest in facial landmark detection, especially in cascaded regression based methods such as the Supervised Descent Method (SDM). However, while facial landmark detection in images has improved significantly, comparably very little attention has been given to the task of landmark detection/tracking in videos. In our work we present a novel initialization procedure that can help with cascaded regression based facial landmark detection and tracking. Our initialization technique exploits the fact that cascaded regression is sensitive to initialization noise, especially in the presence of out-of-plane head pose variation, e.g. when a person is looking down when reading or during fast head motion. Our approach allows to learn good candidates for initialization, that we exploit in our tracking framework. We evaluate our technique on 300VW dataset – a large publicly available corpus of in-the-wild videos and demonstrate its effectiveness for a number of cascaded-regression landmark detection approaches. Christy Li, Tadas Baltrusaitis, Louis-Philippe Morency |
FG | 2 |
| 2017 | Automatically predicting human knowledgeability through non-verbal cuesabstractHumans possess an incredible ability to transmit and decode ``metainformation" through non-verbal actions in daily communication amongst each other. One communicative phenomena that is transmitted through these subtle cues is knowledgeability. In this work, we conduct two experiments. First, we analyze which non-verbal features are important for identifying knowledgeable people when responding to a question. Next, we train a model to predict the knowledgeability of speakers in a game show setting. We achieve results that surpass chance and human performance using a multimodal approach fusing prosodic and visual features. We believe computer systems that can incorporate emotional reasoning of this level can greatly improve human-computer communication and interaction. Abdelwahab Bourai, Tadas Baltrusaitis, Louis-Philippe Morency |
ICMI | 2 |
| 2017 | Multimodal sentiment analysis with word-level fusion and reinforcement learningabstractWith the increasing popularity of video sharing websites such as YouTube and Facebook, multimodal sentiment analysis has received increasing attention from the scientific community. Contrary to previous works in multimodal sentiment analysis which focus on holistic information in speech segments such as bag of words representations and average facial expression intensity, we propose a novel deep architecture for multimodal sentiment analysis that is able to perform modality fusion at the word level. In this paper, we propose the Gated Multimodal Embedding LSTM with Temporal Attention (GME-LSTM(A)) model that is composed of 2 modules. The Gated Multimodal Embedding allows us to alleviate the difficulties of fusion when there are noisy modalities. The LSTM with Temporal Attention can perform word level fusion at a finer fusion resolution between the input modalities and attends to the most important time steps. As a result, the GME-LSTM(A) is able to better model the multimodal structure of speech through time and perform better sentiment comprehension. We demonstrate the effectiveness of this approach on the publicly-available Multimodal Corpus of Sentiment Intensity and Subjectivity Analysis (CMU-MOSI) dataset by achieving state-of-the-art sentiment classification and regression results. Qualitative analysis on our model emphasizes the importance of the Temporal Attention Layer in sentiment prediction because the additional acoustic and visual modalities are noisy. We also demonstrate the effectiveness of the Gated Multimodal Embedding in selectively filtering these noisy modalities out. These results and analysis open new areas in the study of sentiment analysis in human communication and provide new models for multimodal fusion. Minghai Chen, Paul Pu Liang, Tadas Baltrusaitis, Amir Zadeh 0001, Louis-Philippe Morency |
ICMI | 4 |
| 2017 | Computational Analysis of Acoustic Descriptors in Psychotic Patients
Torsten Wörtwein, Tadas Baltrusaitis, Eugene Laksana, Luciana Pennant, Elizabeth S. Liebson, Dost Öngür, Justin T. Baker, Louis-Philippe Morency |
INTERSPEECH | 2 |
| 2016 | Holistically Constrained Local Model: Going Beyond Frontal Poses for Facial Landmark Detection
KangGeon Kim, Tadas Baltrusaitis, Amir Zadeh 0001, Louis-Philippe Morency, Gérard G. Medioni |
BMVC | 2 |
| 2016 | Extending Long Short-Term Memory for Multi-View Structured Learning
Shyam Sundar Rajagopalan, Louis-Philippe Morency, Tadas Baltrusaitis, Roland Göcke |
ECCV (7) | 3 |
| 2016 | A 3D Morphable Eye Region Model for Gaze Estimation
Erroll Wood, Tadas Baltrusaitis, Louis-Philippe Morency, Peter Robinson 0001, Andreas Bulling |
ECCV (1) | 2 |
| 2016 | Learning an appearance-based gaze estimator from one million synthesised imagesabstractLearning-based methods for appearance-based gaze estimation achieve state-of-the-art performance in challenging real-world settings but require large amounts of labelled training data. Learning-by-synthesis was proposed as a promising solution to this problem but current methods are limited with respect to speed, appearance variability, and the head pose and gaze angle distribution they can synthesize. We present UnityEyes, a novel method to rapidly synthesize large amounts of variable eye region images as training data. Our method combines a novel generative 3D model of the human eye region with a real-time rendering framework. The model is based on high-resolution 3D face scans and uses real-time approximations for complex eyeball materials and structures as well as anatomically inspired procedural geometry methods for eyelid animation. We show that these synthesized images can be used to estimate gaze in difficult in-the-wild scenarios, even for extreme gaze angles or in cases in which the pupil is fully occluded. We also demonstrate competitive gaze estimation results on a benchmark in-the-wild dataset, despite only using a light-weight nearest-neighbor algorithm. We are making our UnityEyes synthesis framework available online for the benefit of the research community. Erroll Wood, Tadas Baltrusaitis, Louis-Philippe Morency, Peter Robinson 0001, Andreas Bulling |
ETRA | 2 |
| 2016 | EmoReact: a multimodal approach and dataset for recognizing emotional responses in childrenabstractAutomatic emotion recognition plays a central role in the technologies underlying social robots, affect-sensitive human computer interaction design and affect-aware tutors. Although there has been a considerable amount of research on automatic emotion recognition in adults, emotion recognition in children has been understudied. This problem is more challenging as children tend to fidget and move around more than adults, leading to more self-occlusions and non-frontal head poses. Also, the lack of publicly available datasets for children with annotated emotion labels leads most researchers to focus on adults. In this paper, we introduce a newly collected multimodal emotion dataset of children between the ages of four and fourteen years old. The dataset contains 1102 audio-visual clips annotated for 17 different emotional states: six basic emotions, neutral, valence and nine complex emotions including curiosity, uncertainty and frustration. Our experiments compare unimodal and multimodal emotion recognition baseline models to enable future research on this topic. Finally, we present a detailed analysis of the most indicative behavioral cues for emotion recognition in children. Behnaz Nojavanasghari, Tadas Baltrusaitis, Charles E. Hughes, Louis-Philippe Morency |
ICMI | 2 |
| 2016 | Deep multimodal fusion for persuasiveness predictionabstractPersuasiveness is a high-level personality trait that quantifies the influence a speaker has on the beliefs, attitudes, intentions, motivations, and behavior of the audience. With social multimedia becoming an important channel in propagating ideas and opinions, analyzing persuasiveness is very important. In this work, we use the publicly available Persuasive Opinion Multimedia (POM) dataset to study persuasion. One of the challenges associated with this problem is the limited amount of annotated data. To tackle this challenge, we present a deep multimodal fusion architecture which is able to leverage complementary information from individual modalities for predicting persuasiveness. Our methods show significant improvement in performance over previous approaches. Behnaz Nojavanasghari, Deepak Gopinath, Jayanth Koushik, Tadas Baltrusaitis, Louis-Philippe Morency |
ICMI | 4 |
| 2016 | OpenFace: An open source facial behavior analysis toolkitabstractOver the past few years, there has been an increased interest in automatic facial behavior analysis and understanding. We present OpenFace - an open source tool intended for computer vision and machine learning researchers, affective computing community and people interested in building interactive applications based on facial behavior analysis. OpenFace is the first open source tool capable of facial landmark detection, head pose estimation, facial action unit recognition, and eye-gaze estimation. The computer vision algorithms which represent the core of OpenFace demonstrate state-of-the-art results in all of the above mentioned tasks. Furthermore, our tool is capable of real-time performance and is able to run from a simple webcam without any specialist hardware. Finally, OpenFace allows for easy integration with other applications and devices through a lightweight messaging system. Tadas Baltrusaitis, Peter Robinson 0001, Louis-Philippe Morency |
WACV | 1 |
| 2016 | Automatic Analysis of Naturalistic Hand-Over-Face GesturesabstractOne of the main factors that limit the accuracy of facial analysis systems is hand occlusion. As the face becomes occluded, facial features are lost, corrupted, or erroneously detected. Hand-over-face occlusions are considered not only very common but also very challenging to handle. However, there is empirical evidence that some of these hand-over-face gestures serve as cues for recognition of cognitive mental states. In this article, we present an analysis of automatic detection and classification of hand-over-face gestures. We detect hand-over-face occlusions and classify hand-over-face gesture descriptors in videos of natural expressions using multi-modal fusion of different state-of-the-art spatial and spatio-temporal features. We show experimentally that we can successfully detect face occlusions with an accuracy of 83%. We also demonstrate that we can classify gesture descriptors ( hand shape , hand action , and facial region occluded ) significantly better than a naïve baseline. Our detailed quantitative analysis sheds some light on the challenges of automatic classification of hand-over-face gestures in natural expressions. Marwa Mahmoud, Tadas Baltrusaitis, Peter Robinson 0001 |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2015 | Empirical analysis of continuous affectabstractAutomatic analysis of affect from facial expressions has been extensively studied, but most work has considered only a small set of discrete emotions, typically Ekman's six basic emotions, or a small number of continuous measures, typically valence and arousal. We have developed a system that accommodates a much larger vocabulary of discrete emotions and links them with continuous measures that have been aggregated over a few seconds. The approach has a sound theoretical basis in multi-dimensional statistics, making it both principled and robust, while a graphical presentation makes it easy to understand. Peter Robinson 0001, Tadas Baltrusaitis |
ACII | 2 |
| 2015 | Decoupling facial expressions and head motions in complex emotionsabstractPerception of emotion through facial expressions and head motion is of interest to both psychology and affective computing researchers. However, very little is known about the importance of each modality individually, as they are often treated together rather than separately. We present a study which isolates the effect of head motion from facial expression in the perception of complex emotions in videos. We demonstrate that head motions carry emotional information that is complementary rather than redundant to the emotion content in facial expressions. Finally, we show that emotional expressivity in head motion is not limited to nods and shakes and that additional gestures (such as head tilts, raises and general amount of motion) could be beneficial to automated recognition systems. Andra Adams, Marwa Mahmoud, Tadas Baltrusaitis, Peter Robinson 0001 |
ACII | 3 |
| 2015 | Rendering of Eyes for Eye-Shape Registration and Gaze EstimationabstractImages of the eye are key in several computer vision problems, such as shape registration and gaze estimation. Recent large-scale supervised methods for these problems require time-consuming data collection and manual annotation, which can be unreliable. We propose synthesizing perfectly labelled photo-realistic training data in a fraction of the time. We used computer graphics techniques to build a collection of dynamic eye-region models from head scan geometry. These were randomly posed to synthesize close-up eye images for a wide range of head poses, gaze directions, and illumination conditions. We used our model's controllability to verify the importance of realistic illumination and shape variations in eye-region training data. Finally, we demonstrate the benefits of our synthesized training data (SynthesEyes) by out-performing state-of-the-art methods for eye-shape registration as well as cross-dataset appearance-based gaze estimation in the wild. Erroll Wood, Tadas Baltrusaitis, Xucong Zhang, Yusuke Sugano, Peter Robinson 0001, Andreas Bulling |
ICCV | 2 |
| 2014 | Continuous Conditional Neural Fields for Structured Regression
Tadas Baltrusaitis, Peter Robinson 0001, Louis-Philippe Morency |
ECCV (4) | 1 |
| 2014 | Automatic Detection of Naturalistic Hand-over-Face Gesture DescriptorsabstractOne of the main factors that limit the accuracy of facial analysis systems is hand occlusion. As the face becomes occluded, facial features are either lost, corrupted or erroneously detected. Hand-over-face occlusions are considered not only very common but also very challenging to handle. Moreover, there is empirical evidence that some of these hand-over-face gestures serve as cues for recognition of cognitive mental states. In this paper, we detect hand-over-face occlusions and classify hand-over-face gesture descriptors in videos of natural expressions using multi-modal fusion of different state-of-the-art spatial and spatio-temporal features. We show experimentally that we can successfully detect face occlusions with an accuracy of 83%. We also demonstrate that we can classify gesture descriptors (hand shape, hand action and facial region occluded) significantly higher than a naive baseline. To our knowledge, this work is the first attempt to automatically detect and classify hand-over-face gestures in natural expressions. Marwa Mahmoud, Tadas Baltrusaitis, Peter Robinson 0001 |
ICMI | 2 |
| 2013 | What Really Matters? A Study into People's Instinctive Evaluation Metrics for Continuous Emotion Prediction in MusicabstractContinuous emotion prediction in the arousal-valence space is now being used in various modalities: music, facial expressions, gestures, text, etc. In order to be able to compare the work of different research groups effectively, we believe it is necessary to set certain guidelines for how to conduct research-the choice of evaluation metrics of emotion recognition algorithms in particular. In this paper we focus on the field of musical emotion recognition and describe a study designed to discover people's instinctive preference among the most commonly used evaluation techniques. We gather strong evidence that root mean squared error or Kullback-Leibler divergence should be used for regression based approaches. The raw study data we collected is made publicly available. Vaiva Imbrasaite, Tadas Baltrusaitis, Peter Robinson 0001 |
ACII | 2 |
| 2012 | 3D Constrained Local Model for rigid and non-rigid facial trackingabstractWe present 3D Constrained Local Model (CLM-Z) for robust facial feature tracking under varying pose. Our approach integrates both depth and intensity information in a common framework. We show the benefit of our CLM-Z method in both accuracy and convergence rates over regular CLM formulation through experiments on publicly available datasets. Additionally, we demonstrate a way to combine a rigid head pose tracker with CLM-Z that benefits rigid head tracking. We show better performance than the current state-of-the-art approaches in head pose tracking with our extension of the generalised adaptive view-based appearance model (GAVAM). Tadas Baltrusaitis, Peter Robinson 0001, Louis-Philippe Morency |
CVPR | 1 |
| 2011 | 3D Corpus of Spontaneous Complex Mental States
Marwa Mahmoud, Tadas Baltrusaitis, Peter Robinson 0001, Laurel D. Riek |
ACII (1) | 2 |
| 2011 | Modeling Latent Discriminative Dynamic of Multi-dimensional Affective Signals
Geovany A. Ramírez, Tadas Baltrusaitis, Louis-Philippe Morency |
ACII (2) | 2 |
| 2011 | Real-time inference of mental states from facial expressions and upper body gesturesabstractWe present a real-time system for detecting facial action units and inferring emotional states from head and shoulder gestures and facial expressions. The dynamic system uses three levels of inference on progressively longer time scales. Firstly, facial action units and head orientation are identified from 22 feature points and Gabor filters. Secondly, Hidden Markov Models are used to classify sequences of actions into head and shoulder gestures. Finally, a multi level Dynamic Bayesian Network is used to model the unfolding emotional state based on probabilities of different gestures. The most probable state over a given video clip is chosen as the label for that clip. The average F1 score for 12 action units (AUs 1, 2, 4, 6, 7, 10, 12, 15, 17, 18, 25, 26), labelled on a frame by frame basis, was 0.461. The average classification rate for five emotional states (anger, fear, joy, relief, sadness) was 0.440. Sadness had the greatest rate, 0.64, anger the smallest, 0.11. Tadas Baltrusaitis, Daniel McDuff, Ntombikayise Banda, Marwa Mahmoud, Rana El Kaliouby, Peter Robinson 0001, Rosalind W. Picard |
FG | 1 |