EDBT 2026 Demo / reviewers in the wild / expert
Fernando De la Torre
dblp:d/FernandoDelaTorre · also Fernando De la Torre Frade
· DBLP profile ↗
181ranked-venue papers
22as first author
50since 2021 · last 2026
0000-0002-7086-8572ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 139 · 19 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 120 · 15 first-author · 39 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3Computer networks · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Blurry to Believable: Enhancing Low-Quality Talking Heads with 3D Generative PriorsabstractCreating high-fidelity, animatable 3D talking heads is crucial for immersive applications, yet often hindered by the prevalence of low-quality image or video sources, which yield poor 3D reconstructions. In this paper, we introduce SuperHead, a novel framework for enhancing lowresolution, animatable 3D head avatars. The core challenge lies in synthesizing high-quality geometry and textures, while ensuring both 3D and temporal consistency during animation and preserving subject identity. Despite recent progress in image, video and 3D-based superresolution ($S R$), existing$S R$techniques are ill-equipped to handle dynamic 3D inputs. To address this, SuperHead leverages the rich priors from pre-trained 3D generative models via a novel dynamics-aware 3D inversion scheme. This process optimizes the latent representation of the generative model to produce a super-resolved 3D Gaussian Splatting (3DGS) head model, which is subsequently rigged to an underlying parametric head model (e.g., FLAME) for animation. The inversion is jointly supervised using a sparse collection of upscaled 2D face renderings and corresponding depth maps, captured from diverse facial expressions and camera viewpoints, to ensure realism under$d y$namic facial motions. Experiments demonstrate that SuperHead generates avatars with fine-grained facial details under dynamic motions, significantly outperforming baseline methods in visual quality. Ding-Jiun Huang, Yuanhao Wang 0011, Shao-Ji Yuan, Albert Mosella-Montoro, Francisco Vicente 0001, Cheng Zhang 0014, Fernando De la Torre |
3DV | 7 |
| 2026 | GarmentCrafter: Progressive Novel View Synthesis for Single-View 3D Garment Reconstruction and EditingabstractWe introduce GarmentCrafter, a new approach to enable non-professional users to create and modify 3D garments from a single-view image. While recent advances in image generation have facilitated$2 D$garment design, creating and editing 3D garments remains challenging for nonprofessional users. Existing methods for single-view 3D reconstruction often rely on pre-trained generative models to hallucinate novel views conditioning on the reference image and camera pose, yet they lack cross-view consistency, failing to capture the internal relationships across different views. In this paper, we tackle this challenge through progressive depth prediction and image warping to approximate novel views. Subsequently, we train a multi-view diffusion model to complete occluded and unknown clothing regions, informed by the evolving camera pose. By jointly inferring RGB and depth, GarmentCrafter enforces inter-view coherence and reconstructs precise geometries and fine details. Extensive experiments demonstrate that our method achieves superior visual fidelity and inter-view coherence compared to state-of-the-art single-view 3D garment reconstruction methods. Our code will be publicly available. Yuanhao Wang 0011, Cheng Zhang 0014, Gonçalo Frazão, Alexandru Eugen Ichim, Thabo Beeler, Fernando De la Torre |
3DV | 7 |
| 2026 | LLA MADRS: Evaluating Open-Source LLMs on Real Clinical Interviews - To Reason or Not to Reason?abstractAbstract Large language models (LLMs) excel on many NLP benchmarks, but their behavior on real-world, semi-structured prediction remains underexplored. We present LLAMADRS, a benchmark for structured clinical assessment from dialogue built on the CAMI corpus of psychiatric interviews, comprising 5,804 expert annotations across 541 sessions. We evaluate 25 open-source models (standard and reasoning-augmented; 0.6B–400B parameters) and generate over 400,000 predictions. Our results demonstrate that strong open-source LLMs achieve item-level accuracy with residual error below clinically substantial thresholds. Additionally, an Item-then-Sum (ITS) strategy, assessing symptoms individually through discrete LLM calls before synthesizing final scores, significantly reduces error relative to Direct Total Score (DTS) prediction across most model architectures and scales, despite reasoning models attempting similar decomposition in the reasoning traces of their DTS predictions. In fact, we find that performance gains attributed to “reasoning” depend fundamentally on prompt design: standard models equipped with structured task definitions and examples match reasoning-augmented counterparts. Among the latter, longer reasoning traces correlate with reduced error; while higher model scale does across both architectures. Our results clarify when and why reasoning helps and offer actionable guidance for deploying LLMs in semi-structured clinical assessment. Gaoussou Youssouf Kebe, Jeffrey M. Girard, Einat Liebenthal, Justin T. Baker, Fernando De la Torre, Louis-Philippe Morency |
Trans. Assoc. Comput. Linguistics | 5 |
| 2025 | MaterialFusion: Enhancing Inverse Rendering with Material Diffusion PriorsabstractRecent works in inverse rendering have shown promise in using multi-view images of an object to recover shape, albedo, and materials. However, the recovered components often fail to render accurately under new lighting conditions due to the intrinsic challenge of disentangling albedo and material properties from input images. To address this challenge, we introduce MaterialFusion, an enhanced conventional 3D inverse rendering pipeline that incorporates a 2D prior on texture and material properties. We present StableMaterial, a 2D diffusion model prior that refines multilit data to estimate the most likely albedo and material from given input appearances. This model is trained on albedo, material, and relit image data derived from a curated dataset of approximately ∼12K artist-designed synthetic Blender objects called BlenderVault. We incorporate this diffusion prior with an inverse rendering framework where we use score distillation sampling (SDS) to guide the optimization of the albedo and materials, improving relighting performance in comparison with previous work. We validate MaterialFusion's relighting performance on 4 datasets of synthetic and real objects under diverse illumination conditions, showing our diffusion-aided approach significantly improves the appearance of reconstructed objects under novel lighting conditions. We intend to publicly release our BlenderVault dataset to support further research in this field. Yehonathan Litman, Or Patashnik, Kangle Deng, Aviral Agrawal, Rushikesh Zawar, Fernando De la Torre, Shubham Tulsiani |
3DV | 6 |
| 2025 | Custom Condition Generation for Zero-Shot Human-Scene Interactions SynthesisabstractExisting methods for creating human interactions within scenes show promise for common interactions, but often fail with less frequent ones. To overcome this, we introduce a new approach that creates tailored conditions for generating these interactions without previously seen examples. This method leverages the strengths of both large language models (LLMs) and vision-language models (VLMs). Unlike the GenZI, the current state-of-the-art approach, which struggles with rare interactions due to its reliance on VLM inpainting, our method follows a three-step process: first, we generate a preliminary human posture using VLMs and then estimate this posture in three dimensions. Next, we refine the conditions to fit the specific scene and interaction by analyzing the inputs with both LLMs and VLMs. Finally, we fine-tune the placement, orientation, and posture of the human figure using specific optimization techniques. Our experimental results show that this method performs well across a wide range of interactions, including those that are less common. Ryosuke Kawamura, Zoltán Ádám Milacski, Fernando De la Torre, László A. Jeni, Koichiro Niinuma |
FG | 3 |
| 2025 | OASIS: Object-guided Attention for Text-conditional Diffusion Synthesis of Human Interaction SequencesabstractAnalyzing and synthesizing human-object interaction is crucial for advancing intelligent systems that engage with the physical environment. However, simultaneous tracking of human and object data presents inherent challenges, resulting in limitations in dataset scale, diversity, and annotation quality within this domain, thereby hindering the generalization ability of trained models. This study introduces OASIS, a novel framework that extends pretrained text-conditional human motion diffusion models to address the complex task of fullbody 3D hand-object interaction generation. Specifically, we freeze the parameters of the pretrained motion diffusion model, while incorporating additional object-guided attention layers, which we train to adapt the human motion latents to match the input object motion sequence and the text. Our method can be understood as a ControlNet [38] for interaction. Through extensive experimentation, we demonstrate the effectiveness and robustness of our framework in generating realistic handobject interactions from textual descriptions. Our method surpasses the state-of-the-art performance in FID and accuracy interaction fidelity metrics compared to the prior best method IMoS [10], with improvements of 0.08 in FID and $2 \%$ in accuracy for body motion synthesis, and 0.15 in FID and $10 \%$ in accuracy for hand motion synthesis. Chih-Chun Yang 0005, Tianhui Cai, Zoltán Ádám Milacski, Aayush Prakash, Shingo Takagi 0001, Daeil Kim, Fernando De la Torre |
FG | 7 |
| 2025 | Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak SupervisionabstractDetecting vehicles in aerial imagery is a critical task with applications in traffic monitoring, urban planning, and defense intelligence. Deep learning methods have provided state-of-the-art (SOTA) results for this application. However, a significant challenge arises when models trained on data from one geographic region fail to generalize effectively to other areas. Variability in factors such as environmental conditions, urban layouts, road networks, vehicle types, and image acquisition parameters (e.g., resolution, lighting, and angle) leads to domain shifts that degrade model performance. This paper proposes a novel method that uses generative AI to synthesize high-quality aerial images and their labels, improving detector training through data augmentation. Our key contribution is the development of a multi-stage, multi-modal knowledge transfer framework utilizing fine-tuned latent diffusion models (LDMs) to mitigate the distribution gap between the source and target environments. Extensive experiments across diverse aerial imagery domains show consistent performance improvements in AP50 over supervised learning on source domain data, weakly supervised adaptation methods, unsupervised domain adaptation methods, and open-set object detectors by 4-23%, 6-10%, 7-40%, and more than 50%, respectively. Furthermore, we introduce two newly annotated aerial datasets from New Zealand and Utah to support further research in this field. Project page is available at: https://humansensinglab.github.io/AGenDA Minhyek Jeon, Shuowen Hu, Zheyang Qin, Shayok Chakraborty, Stanislav Panev, Celso de Melo, Fernando De la Torre |
ICCV | 8 |
| 2025 | Teleportraits: Training-Free People Insertion Into Any SceneabstractThe task of realistically inserting a human from a reference image into a background scene is highly challenging, requiring the model to (1) determine the correct location and poses of the person and (2) perform high-quality personalization conditioned on the background. Previous approaches often treat them as separate problems, overlooking their interconnections, and typically rely on training to achieve high performance. In this work, we introduce a unified training-free pipeline that leverages pre-trained text-to-image diffusion models. We show that diffusion models inherently possess the knowledge to place people in complex scenes without requiring task-specific training. By combining inversion techniques with classifier-free guidance, our method achieves affordance-aware global editing, seamlessly inserting people into scenes. Furthermore, our proposed mask-guided self-attention mechanism ensures high-quality personalization, preserving the subject's identity, clothing, and body features from just a single reference image. To the best of our knowledge, we are the first to perform realistic human insertions into scenes in a training-free manner and achieve state-of-the-art results in diverse composite scene images with excellent identity preservation in backgrounds and subjects. Jialu Gao, K. J. Joseph, Fernando De la Torre |
ICCV | 3 |
| 2025 | LightSwitch: Multi-View Relighting with Material-Guided DiffusionabstractRecent approaches for 3D relighting have shown promise in integrating 2D image relighting generative priors to alter the appearance of a 3D representation while preserving the underlying structure. Nevertheless, generative priors used for 2D relighting that directly relight from an input image do not take advantage of intrinsic properties of the subject that can be inferred or cannot consider multi-view data at scale, leading to subpar relighting. In this paper, we propose Lightswitch, a novel finetuned material-relighting diffusion framework that efficiently relights an arbitrary number of input images to a target lighting condition while incorporating cues from inferred intrinsic properties. By using multi-view and material information cues together with a scalable denoising scheme, our method consistently and efficiently relights dense multi-view data of objects with diverse material compositions. We show that our 2D relighting prediction quality exceeds previous state-of-the-art relighting priors that directly relight from images. We further demonstrate that LightSwitch matches or outperforms state-of-the-art diffusion inverse rendering methods in relighting synthetic and real objects in as little as 2 minutes. Yehonathan Litman, Fernando De la Torre, Shubham Tulsiani |
ICCV | 2 |
| 2025 | GAS: Generative Avatar Synthesis from a Single ImageabstractWe present a unified and generalizable framework for synthesizing view-consistent and temporally coherent avatars from a single image, addressing the challenging task of single-image avatar generation. Existing diffusion-based methods often condition on sparse human templates (e.g., depth or normal maps), which leads to multi-view and temporal inconsistencies due to the mismatch between these signals and the true appearance of the subject. Our approach bridges this gap by combining the reconstruction power of regression-based 3D human reconstruction with the generative capabilities of a diffusion model. In a first step, an initial 3D reconstructed human through a generalized NeRF provides comprehensive conditioning, ensuring high-quality synthesis faithful to the reference appearance and structure. Subsequently, the derived geometry and appearance from the generalized NeRF serve as input to a video-based diffusion model. This strategic integration is pivotal for enforcing both multi-view and temporal consistency throughout the avatar's generation. Empirical results underscore the superior generalization ability of our proposed method, demonstrating its effectiveness across diverse in-domain and out-of-domain in-the-wild datasets. Yixing Lu, Junting Dong, Youngjoong Kwon, Bo Dai 0002, Fernando De la Torre |
ICCV | 6 |
| 2025 | Improving Noise Efficiency in Privacy-Preserving Dataset DistillationabstractModern machine learning models heavily rely on large datasets that often include sensitive and private information, raising serious privacy concerns. Differentially private (DP) data generation offers a solution by creating synthetic datasets that limit the leakage of private information within a predefined privacy budget; however, it requires a substantial amount of data to achieve performance comparable to models trained on the original data. To mitigate the significant expense incurred with synthetic data generation, Dataset Distillation (DD) stands out for its remarkable training and storage efficiency. This efficiency is particularly advantageous when integrated with DP mechanisms, curating compact yet informative synthetic datasets without compromising privacy. However, current state-of-the-art private DD methods suffer from a synchronized sampling-optimization process and the dependency on noisy training signals from randomly initialized networks. This results in the inefficient utilization of private information due to the addition of excessive noise. To address these issues, we introduce a novel framework that decouples sampling from optimization for better convergence and improves signal quality by mitigating the impact of DP noise through matching in an informative subspace. On CIFAR-10, our method achieves a \textbf{10.0\%} improvement with 50 images per class and \textbf{8.3\%} increase with just \textbf{one-fifth} the distilled set size of previous state-of-the-art methods, demonstrating significant potential to advance privacy-preserving DD. Runkai Zheng, Vishnu Asutosh Dasu, Yinong Wang 0001, Haohan Wang, Fernando De la Torre |
ICCV | 5 |
| 2025 | Texture- and Shape-Based Adversarial Attacks for Overhead Image Vehicle DetectionabstractDetecting vehicles in aerial images is difficult due to complex backgrounds, small object sizes, shadows, and occlusions. Although recent deep learning advancements have improved object detection, these models remain susceptible to adversarial attacks (AAs), challenging their reliability. Traditional AA strategies often ignore practical implementation constraints. Our work proposes realistic and practical constraints on texture (lowering resolution, limiting modified areas, and color ranges) and analyzes the impact of shape modifications on attack performance. We conducted extensive experiments with three object detector architectures, demonstrating the performance-practicality trade-off: more practical modifications tend to be less effective, and vice versa. We release both code and data to support reproducibility at https://github.com/humansensinglab/texture-shape-adversarial-attacks. Mikael Yeghiazaryan, Sai Abhishek Si Namburu, Emily Kim, Stanislav Panev, Celso de Melo, Fernando De la Torre, Jessica K. Hodgins |
ICIP | 6 |
| 2025 | GHOST: Grounded Human Motion Generation with Open Vocabulary Scene-and-Text ContextsabstractThe connection between our 3D surroundings and the descriptive language that characterizes them would be well-suited for localizing and generating human motion in context but for one problem. The complexity introduced by multiple modalities makes capturing this connection challenging with a fixed set of descriptors. Specifically, closed vocabulary scene encoders, which require learning text-scene associations from scratch, have been favored in the literature, often resulting in inaccurate motion grounding. In this paper, we propose a method that integrates an open vocabulary scene encoder into the architecture, establishing a robust connection between text and scene. Our two-step approach starts with pretraining the scene encoder through knowledge distillation from an existing open vocabulary semantic image segmentation model, ensuring a shared text-scene feature space. Subsequently, the scene encoder is fine-tuned for conditional motion generation, incorporating two novel regularization losses that regress the category and size of the goal object. Our methodology achieves up to a 30% reduction in the goal object distance metric compared to the prior state-of-the-art baseline model on the HUMANISE dataset. This improvement is demonstrated through evaluations conducted using three implementations of our framework, a perceptual study, and an open vocabulary experiment. Additionally, our method is designed to accommodate future 2D open vocabulary segmentation methods for distillation in a plug-and-play manner. Zoltán Ádám Milacski, Koichiro Niinuma, Ryosuke Kawamura, Fernando De la Torre, László A. Jeni |
WACV | 4 |
| 2025 | Exploring image and skeleton-based action recognition approaches for clinical in-bed classification of simulated epileptic seizure movementsabstractEpileptic seizure classification based on seizure semiology requires automated, quantitative approaches to support the diagnosis of epilepsy, which affects 1% of the world’s population. Current approaches address the problem on a seizure level, neglecting the detailed evaluation of the classification of the underlying action features, also known as Movements of Interest (MOIs), which are critical for epileptologists in determining their classifications. Moreover, it hinders objective comparison of these approaches and attribution of performance differences due to datasets, intra-dataset MOI distribution, or architecture variations. Objective evaluation of action recognition techniques is crucial, with MOIs serving as foundational elements of semiology for clinical in-bed applications to facilitate epileptic seizure classification. However, until now, there were no MOI datasets available nor benchmarks comparing different action recognition approaches for this clinical problem. Therefore, as a pilot, we introduced a novel, simulated seizure semiology dataset carried out by 8 experienced epileptologists in an EMU bed, consisting of 7 MOI classes. We compare several computer vision methods for MOI classification, two image-based (I3D and Uniformerv2), and two skeleton-based (ST-GCN++ and PoseC3D) action recognition approaches. This study emphasizes the advantages of a 2-stage skeleton-based action recognition approach in a transfer learning setting (4 classes) and the multi-scale challenge of MOI classification (7 classes), advocating for the integration of skeleton-based methods with hand gesture recognition technologies in the future. The study’s controlled MOI simulation dataset provides us with the opportunity to advance the development of automated epileptic seizure classification systems, paving the way for enhancing their performance and having the potential to contribute to improved patient care. Tamás Karácsony, Nicholas Fearns, Denise Birk, Selina Denise Trapp, Katharina Ernst, Christian Vollmar, Jan Rémi, László A. Jeni, Fernando De la Torre, João Paulo da Silva Cunha |
Expert Syst. Appl. | 9 |
| 2025 | Echoes of the Coliseum: Towards 3D Live streaming of Sports EventsabstractHuman-centered live events have always played a pivotal role in shaping culture and fostering social connections. Traditional 2D live transmissions fail to replicate the immersive quality of physical attendance. Addressing this gap, this paper proposes LiveSplats , a framework towards real-time, photo-realistic 3D reconstructions of live events using high-performance 3D Gaussian Splatting. Our solution capitalizes on strong geometric priors to optimize through distributed processing and load balancing, enabling interactive, freely explorable 3D experiences. By dividing scene reconstruction into actor-centric and environment-specific tasks, we employ hierarchical coarse-to-fine optimization to rapidly and accurately reconstruct human actors based on pose data, refining their geometry and appearance with photometric loss. For static environments, we focus on view-dependent appearance changes, streamlining rendering efficiency and maximizing GPU performance. To facilitate evaluation, we introduce (and distribute) a synthetic benchmark dataset of basketball games, offering high visual fidelity as ground truth. In both our synthetic benchmark and publicly available benchmarks, LiveSplats consistently outperforms existing approaches. The dataset is available at https://humansensinglab.github.io/basket-multiview. Junkai Huang 0004, Saswat Subhajyoti Mallick, Alejandro Amat, Marc Ruiz Olle, Albert Mosella-Montoro, Bernhard Kerbl, Francisco Vicente 0001, Fernando De la Torre |
ACM Trans. Graph. | 8 |
| 2024 | Domain Gap Embeddings for Generative Dataset AugmentationabstractThe performance of deep learning models is intrinsically tied to the quality, volume, and relevance of their training data. Gathering ample data for production scenarios of-ten demands significant time and resources. Among various strategies, data augmentation circumvents exhaustive data collection by generating new data points from existing ones. However, traditional augmentation techniques can be less effective amidst a shift in training and testing distributions. This paper explores the potential of synthetic data by leveraging large pre-trained models for data augmentation, especially when confronted with distribution shifts. Al-though recent advancements in generative models have en-abled several prior works in cross-distribution data gener-ation, they require model fine-tuning and a complex setup. To bypass these shortcomings, we introduce Domain Gap Embeddings (DoGE), a plug-and-play semantic data aug-mentation framework in a cross-distribution few-shot set-ting. Our method extracts disparities between source and desired data distributions in a latent form, and subsequently steers a generative process to supplement the training set with endless diverse synthetic samples. Our evaluations, conducted on a subpopulation shift and three domain adap-tation scenarios under afew-shot paradigm, reveal that our versatile method improves performance across tasks with-out needing hands-on intervention or intricate fine-tuning. DoGE paves the way to effortlessly generate realistic, con-trollable synthetic datasets following the test distributions, bolstering real-world efficacy for downstream task models. Yinong Wang 0001, Younjoon Chung, Chen Henry Wu, Fernando De la Torre |
CVPR | 4 |
| 2024 | POET: Prompt Offset Tuning for Continual Human Action Adaptation
Prachi Garg, K. J. Joseph, Vineeth N. Balasubramanian, Necati Cihan Camgöz, Chengde Wan, Kenrick Kin, Weiguang Si, Shugao Ma, Fernando De la Torre |
ECCV (64) | 9 |
| 2024 | FAMOUS: High-Fidelity Monocular 3D Human Digitization Using View Synthesis
Vishnu Mani Hema, Shubhra Aich, Christian Häne, Jean-Charles Bazin, Fernando De la Torre |
ECCV (82) | 5 |
| 2024 | Generalizable Human Gaussians for Sparse View Synthesis
Youngjoong Kwon, Baole Fang, Yixing Lu, Haoye Dong, Cheng Zhang 0014, Francisco Vicente 0001, Albert Mosella-Montoro, Jianjin Xu, Shingo Takagi 0001, Daeil Kim, Aayush Prakash, Fernando De la Torre |
ECCV (78) | 12 |
| 2024 | Visual Data Diagnosis and Debiasing with Concept GraphsabstractThe widespread success of deep learning models today is owed to the curation of extensive datasets significant in size and complexity. However, such models frequently pick up inherent biases in the data during the training process, leading to unreliable predictions. Diagnosing and debiasing datasets is thus a necessity to ensure reliable model performance. In this paper, we present ConBias, a novel framework for diagnosing and mitigating Concept co-occurrence Biases in visual datasets. ConBias represents visual datasets as knowledge graphs of concepts, enabling meticulous analysis of spurious concept co-occurrences to uncover concept imbalances across the whole dataset. Moreover, we show that by employing a novel clique-based concept balancing strategy, we can mitigate these imbalances, leading to enhanced performance on downstream tasks. Extensive experiments show that data augmentation based on a balanced concept distribution augmented by ConBias improves generalization performance across multiple datasets compared to state-of-the-art methods. Rwiddhi Chakraborty, Yinong Wang 0001, Jialu Gao, Runkai Zheng, Cheng Zhang 0014, Fernando De la Torre |
NeurIPS | 6 |
| 2024 | Doubly Hierarchical Geometric Representations for Strand-based Human Hairstyle GenerationabstractWe introduce a doubly hierarchical generative representation for strand-based 3D hairstyle geometry that progresses from coarse, low-pass filtered guide hair to densely populated hair strands rich in high-frequency details. We employ the Discrete Cosine Transform (DCT) to separate low-frequency structural curves from high-frequency curliness and noise, avoiding the Gibbs' oscillation issues associated with the standard Fourier transform in open curves. Unlike the guide hair sampled from the scalp UV map grids which may lose capturing details of the hairstyle in existing methods, our method samples optimal sparse guide strands by utilising $k$-medoids clustering centres from low-pass filtered dense strands, which more accurately retain the hairstyle's inherent characteristics. The proposed variational autoencoder-based generation network, with an architecture inspired by geometric deep learning and implicit neural representations, facilitates flexible, off-the-grid guide strand modelling and enables the completion of dense strands in any quantity and density, drawing on principles from implicit neural representations. Empirical evaluations confirm the capacity of the model to generate convincing guide hair and dense strands, complete with nuanced high-frequency details. Yunlu Chen, Francisco Vicente 0001, Christian Häne, Giljoo Nam, Jean-Charles Bazin, Fernando De la Torre |
NeurIPS | 6 |
| 2024 | Hamba: Single-view 3D Hand Reconstruction with Graph-guided Bi-Scanning Mambaabstract3D Hand reconstruction from a single RGB image is challenging due to the articulated motion, self-occlusion, and interaction with objects. Existing SOTA methods employ attention-based transformers to learn the 3D hand pose and shape, yet they do not fully achieve robust and accurate performance, primarily due to inefficiently modeling spatial relations between joints. To address this problem, we propose a novel graph-guided Mamba framework, named Hamba, which bridges graph learning and state space modeling. Our core idea is to reformulate Mamba's scanning into graph-guided bidirectional scanning for 3D reconstruction using a few effective tokens. This enables us to efficiently learn the spatial relationships between joints for improving reconstruction performance. Specifically, we design a Graph-guided State Space (GSS) block that learns the graph-structured relations and spatial sequences of joints and uses 88.5\% fewer tokens than attention-based methods. Additionally, we integrate the state space features and the global features using a fusion module. By utilizing the GSS block and the fusion module, Hamba effectively leverages the graph-guided state space features and jointly considers global and local features to improve performance. Experiments on several benchmarks and in-the-wild tests demonstrate that Hamba significantly outperforms existing SOTAs, achieving the PA-MPVPE of 5.3mm and F@15mm of 0.992 on FreiHAND. At the time of this paper's acceptance, Hamba holds the top position, Rank 1, in two competition leaderboards on 3D hand reconstruction. Haoye Dong, Aviral Chharia, Wenbo Gou, Francisco Vicente 0001, Fernando De la Torre |
NeurIPS | 5 |
| 2024 | Taming 3DGS: High-Quality Radiance Fields with Limited Resources
Saswat Subhajyoti Mallick, Rahul Goel, Bernhard Kerbl, Markus Steinberger, Francisco Vicente 0001, Fernando De la Torre |
SIGGRAPH Asia | 6 |
| 2024 | Consolidating Attention Features for Multi-view Image Editing
Or Patashnik, Rinon Gal, Daniel Cohen-Or, Jun-Yan Zhu, Fernando De la Torre |
SIGGRAPH Asia | 5 |
| 2024 | FabricDiffusion: High-Fidelity Texture Transfer for 3D Garments Generation from In-The-Wild Images
Cheng Zhang 0014, Yuanhao Wang 0011, Francisco Vicente 0001, Chenglei Wu, Thabo Beeler, Fernando De la Torre |
SIGGRAPH Asia | 7 |
| 2024 | A Sequential Learning-based Approach for Monocular Human Performance CaptureabstractHuman performance capture from RGB videos in unconstrained environments has become very popular for applications that require generating virtual avatars or digital actors. SOTA methods use neural network (NN) techniques to estimate the shape directly from photos, yielding a simplified model of the human body. While effective, NN techniques frequently fail under challenging poses and do not preserve temporal consistency. On the other hand, optimization-based methods like shape-from-silhouette can produce more precise reconstruction; however, they typically require a good initialization and are computationally more intensive than NN. To address issues of previous methods, this work proposes a learning-based approach for optimizing fine-grained shape representation from a monocular RGB video. Our main idea is to sequentially recover different shape details (i.e. average shape, clothing, wrinkles) using separate neural networks. At each level, our network takes the sparse/noisy gradients of body mesh vertices w.r.t. the shape, and predicts dense gradients to update the body shape. Despite being trained on synthetic data, these networks have surprisingly good generalization to real images. Experimental validation shows that our approach outperforms NN approaches in recovering shape details while also being an order of magnitude faster than optimization-based methods and robust across varied poses and novel views. Jianchun Chen, Jayakorn Vongkulbhisal, Fernando De la Torre |
WACV | 3 |
| 2024 | Exploring the Impact of Rendering Method and Motion Quality on Model Performance when Using Multi-view Synthetic Data for Action RecognitionabstractThis paper explores the use of synthetic data in a human action recognition (HAR) task to avoid the challenges of obtaining and labeling real-world datasets. We introduce a new dataset suite comprising five datasets, eleven common human activities, three synchronized camera views (aerial and ground) in three outdoor environments, and three visual domains (real and two synthetic). For the synthetic data, two rendering methods (standard computer graphics and neural rendering) and two sources of human motions (motion capture and video-based motion reconstruction) were employed. We evaluated each dataset type by training popular activity recognition models and comparing the performance on the real test data. Our results show that synthetic data achieve slightly lower accuracy (4–8 %) than real data. On the other hand, a model pre-trained on synthetic data and fine-tuned on limited real data surpasses the performance of either domain alone. Standard computer graphics (CG)-rendered data delivers better performance than the data generated from the neural-based rendering method. The results suggest that the quality of the human motions in the training data also affects the test results: motion capture delivers higher test accuracy. Additionally, a model trained on CG aerial view synthetic data exhibits greater robustness against camera viewpoint changes than one trained on real data. See the project page: http://humansensinglab.github.io/REMAG/ Stanislav Panev, Emily Kim, Sai Abhishek Si Namburu, Desislava Nikolova, Celso de Melo, Fernando De la Torre, Jessica K. Hodgins |
WACV | 6 |
| 2024 | Towards Realistic Generative 3D Face ModelsabstractIn recent years, there has been significant progress in 2D generative face models fueled by applications such as animation, synthetic data generation, and digital avatars. However, due to the absence of 3D information, these 2D models often struggle to accurately disentangle facial attributes like pose, expression, and illumination, limiting their editing capabilities. To address this limitation, this paper proposes a 3D controllable generative face model to produce high-quality albedo and precise 3D shapes by leveraging existing 2D generative models. By combining 2D face generative models with semantic face manipulation, this method enables editing of detailed 3D rendered faces. The proposed framework utilizes an alternating descent optimization approach over shape and albedo. Differentiable rendering is used to train high-quality shapes and albedo without 3D supervision. Moreover, this approach outperforms most state-of-the-art (SOTA) methods in the well-known NoW and REALY benchmarks for 3D face re construction. It also outperforms the SOTA reconstruction models in recovering rendered faces’ identities across novel poses. Additionally, the paper demonstrates direct control of expressions in 3D faces by exploiting latent space leading to text-based editing of 3D faces. Aashish Rai, Hiresh Gupta, Francisco Vicente 0001, Shingo Takagi 0001, Amaury Aubel, Daeil Kim, Aayush Prakash, Fernando De la Torre |
WACV | 9 |
| 2024 | MotionGPT: Human Motion Synthesis with Improved Diversity and Realism via GPT-3 PromptingabstractThere are numerous applications for human motion synthesis, including animation, gaming, robotics, or sports science. In recent years, human motion generation from natural language has emerged as a promising alternative to costly and labor-intensive data collection methods relying on motion capture or wearable sensors (e.g., suits). Despite this, generating human motion from textual descriptions remains a challenging and intricate task, primarily due to the scarcity of large-scale supervised datasets capable of capturing the full diversity of human activity.This study proposes a new approach, called MotionGPT, to address the limitations of previous text-based human motion generation methods by utilizing the extensive semantic information available in large language models (LLMs). We first pretrain a doubly text-conditional motion diffusion model on both coarse ("high-level") and detailed ("low-level") ground truth text data. Then during inference, we improve motion diversity and alignment with the training set, by zero-shot prompting GPT-3 for additional "low-level" details. Our method achieves new state-of-the-art quantitative results in terms of Fréchet Inception Distance (FID) and motion diversity metrics, and improves all considered metrics. Furthermore, it has strong qualitative performance, producing natural results. Code is available at https://github.com/humansensinglab/MotionGPT José Ribeiro-Gomes, Tianhui Cai, Zoltán Ádám Milacski, Aayush Prakash, Shingo Takagi 0001, Amaury Aubel, Daeil Kim, Alexandre Bernardino, Fernando De la Torre |
WACV | 10 |
| 2024 | Personalized Face Inpainting with Diffusion Models by Parallel Visual AttentionabstractFace inpainting is important in various applications, such as photo restoration, image editing, and virtual reality. Despite the significant advances in face generative models, ensuring that a person’s unique facial identity is maintained during the inpainting process is still an elusive goal. Current state-of-the-art techniques, exemplified by MyStyle, necessitate resource-intensive fine-tuning and a substantial number of images for each new identity. Furthermore, existing methods often fall short in accommodating user-specified semantic attributes, such as beard or expression.To improve inpainting results, and reduce the computational complexity during inference, this paper proposes the use of Parallel Visual Attention (PVA) in conjunction with diffusion models. Specifically, we insert parallel attention matrices to each cross-attention module in the denoising network, which attends to features extracted from reference images by an identity encoder. We train the added attention modules and identity encoder on CelebAHQ-IDI, a dataset proposed for identity-preserving face inpainting. Experiments demonstrate that PVA attains unparalleled identity resemblance in both face inpainting and face inpainting with language guidance tasks, in comparison to various benchmarks, including MyStyle, Paint by Example, and Custom Diffusion. Our findings reveal that PVA ensures good identity preservation while offering effective language-controllability. Additionally, in contrast to Custom Diffusion, PVA requires just 40 fine-tuning steps for each new identity, which translates to a significant speed increase of over 20 times. Jianjin Xu, Saman Motamed, Praneetha Vaddamanu, Chen Henry Wu, Christian Häne, Jean-Charles Bazin, Fernando De la Torre |
WACV | 7 |
| 2024 | D3GU: Multi-target Active Domain Adaptation via Enhancing Domain AlignmentabstractUnsupervised domain adaptation (UDA) for image classification has made remarkable progress in transferring classification knowledge from a labeled source domain to an unlabeled target domain, thanks to effective domain alignment techniques. Recently, in order to further improve performance on a target domain, many Single-Target Active Domain Adaptation (ST-ADA) methods have been proposed to identify and annotate the salient and exemplar target samples. However, it requires one model to be trained and deployed for each target domain and the domain label associated with each test sample. This largely restricts its application in the ubiquitous scenarios with multiple target domains. Therefore, we propose a Multi-Target Active Domain Adaptation (MT-ADA) framework for image classification, named D3GU, to simultaneously align different domains and actively select samples from them for annotation. This is the first research effort in this field to our best knowledge. D3GU applies Decomposed Domain Discrimination (D3) during training to achieve both source-target and target-target domain alignments. Then during active sampling, a Gradient Utility (GU) score is designed to weight every unlabeled target image by its contribution towards classification and domain alignment tasks, and is further combined with KMeans clustering to form GU-KMeans for diverse image sampling. Extensive experiments on three benchmark datasets, Office31, OfficeHome, and DomainNet, have been conducted to validate consistently superior performance of D3GU for MT-ADA1. Lin Zhang 0040, Linghan Xu, Saman Motamed, Shayok Chakraborty, Fernando De la Torre |
WACV | 5 |
| 2024 | Deep learning methods for single camera based clinical in-bed movement action recognition
Tamás Karácsony, László A. Jeni, Fernando De la Torre, João Paulo da Silva Cunha |
Image Vis. Comput. | 3 |
| 2023 | Zero-Shot Model DiagnosisabstractWhen it comes to deploying deep vision models, the behavior of these systems must be explicable to ensure confidence in their reliability and fairness. A common approach to evaluate deep learning models is to build a labeled test set with attributes of interest and assess how well it performs. However, creating a balanced test set (i.e., one that is uniformly sampled over all the important traits) is often time-consuming, expensive, and prone to mistakes. The question we try to address is: can we evaluate the sensitivity of deep learning models to arbitrary visual attributes without an annotated test set? This paper argues the case that Zero-shot Model Diagnosis (ZOOM) is possible without the need for a test set nor labeling. To avoid the need for test sets, our system relies on a generative model and CLIP. The key idea is enabling the user to select a set of prompts (relevant to the problem) and our system will automatically search for semantic counterfactual images (i.e., synthesized images that flip the prediction in the case of a binary classifier) using the generative model. We evaluate several visual tasks (classification, key-point detection, and segmentation) in multiple visual domains to demonstrate the viability of our methodology. Extensive experiments demonstrate that our method is capable of producing counterfactual images and offering sensitivity analysis for model diagnosis without the need for a test set. Jinqi Luo, Zhaoning Wang, Chen Henry Wu, Dong Huang 0007, Fernando De la Torre |
CVPR | 5 |
| 2023 | Data-Free Class-Incremental Hand Gesture RecognitionabstractThis paper investigates data-free class-incremental learning (DFCIL) for hand gesture recognition from 3D skeleton sequences. In this class-incremental learning (CIL) setting, while incrementally registering the new classes, we do not have access to the training samples (i.e. data-free) of the already known classes due to privacy. Existing DFCIL methods primarily focus on various forms of knowledge distillation for model inversion to mitigate catastrophic forgetting. Unlike SOTA methods, we delve deeper into the choice of the best samples for inversion. Inspired by the well-grounded theory of max-margin classification, we find that the best samples tend to lie close to the approximate decision boundary within a reasonable margin. To this end, we propose BOAT-MI – a simple and effective boundary-aware prototypical sampling mechanism for model inversion for DFCIL. Our sampling scheme outperforms SOTA methods significantly on two 3D skeleton gesture datasets, the publicly available SHREC 2017, and EgoGesture3D – which we extract from a publicly available RGBD dataset. Both our codebase and the EgoGesture3D skeleton dataset are publicly available: https://github.com/humansensinglab/dfcil-hgr. Shubhra Aich, Jesús Ruiz-Santaquiteria, Prachi Garg, K. J. Joseph, Alvaro Fernandez Garcia, Vineeth N. Balasubramanian, Kenrick Kin, Chengde Wan, Necati Cihan Camgöz, Shugao Ma, Fernando De la Torre |
ICCV | 12 |
| 2023 | PATMAT: Person Aware Tuning of Mask-Aware Transformer for Face inpaintingabstractGenerative models such as StyleGAN2 and Stable Diffusion have achieved state-of-the-art performance in computer vision tasks such as image synthesis, inpainting, and de-noising. However, current generative models for face inpainting often fail to preserve fine facial details and the identity of the person, despite creating aesthetically convincing image structures and textures. In this work, we propose Person Aware Tuning (PAT) of Mask-Aware Transformer (MAT) for face inpainting, which addresses this issue. Our proposed method, PATMAT1, effectively preserves identity by incorporating reference images of a subject and fine-tuning a MAT architecture trained on faces. By using ~40 reference images, PATMAT creates anchor points in MAT’s style module, and tunes the model using the fixed anchors to adapt the model to a new face identity. Moreover, PATMAT’s use of multiple images per anchor during training allows the model to use fewer reference images than competing methods. We demonstrate that PATMAT outperforms state-of-the-art models in terms of image quality, the preservation of person-specific details, and the identity of the subject. Our results suggest that PATMAT can be a promising approach for improving the quality of personalized face inpainting2. Saman Motamed, Jianjin Xu, Chen Henry Wu, Christian Häne, Jean-Charles Bazin, Fernando De la Torre |
ICCV | 6 |
| 2023 | A Latent Space of Stochastic Diffusion Models for Zero-Shot Image Editing and GuidanceabstractDiffusion models generate images by iterative denoising. Recent work has shown that by making the denoising process deterministic, one can encode real images into latent codes of the same size, which can be used for image editing. This paper explores the possibility of defining a latent space even when the denoising process remains stochastic. Recall that, in stochastic diffusion models, Gaussian noises are added in each denoising step, and we can concatenate all the noises to form a latent code. This results in a latent space of much higher dimensionality than the original image. We demonstrate that this latent space of stochastic diffusion models can be used in the same way as that of deterministic diffusion models in two applications. First, we propose CycleDiffusion, a method for zero-shot and unpaired image editing using stochastic diffusion models, which improves the performance over its deterministic counterpart. Second, we demonstrate unified, plug-and-play guidance in the latent spaces of deterministic and stochastic diffusion models.1 Chen Henry Wu, Fernando De la Torre |
ICCV | 2 |
| 2023 | ITI-Gen: Inclusive Text-to-Image GenerationabstractText-to-image generative models often reflect the biases of the training data, leading to unequal representations of underrepresented groups. This study investigates inclusive text-to-image generative models that generate images based on human-written prompts and ensure the resulting images are uniformly distributed across attributes of interest. Unfortunately, directly expressing the desired attributes in the prompt often leads to sub-optimal results due to linguistic ambiguity or model misrepresentation. Hence, this paper proposes a drastically different approach that adheres to the maxim that "a picture is worth a thousand words". We show that, for some attributes, images can represent concepts more expressively than text. For instance, categories of skin tones are typically hard to specify by text but can be easily represented by example images. Building upon these insights, we propose a novel approach, ITI-Gen1, that leverages readily available reference images for Inclusive Text-to-Image GENeration. The key idea is learning a set of prompt embeddings to generate images that can effectively represent all desired attribute categories. More importantly, ITI-Gen requires no model fine-tuning, making it computationally efficient to augment existing text-to-image models. Extensive experiments demonstrate that ITI-Gen largely improves over state-of-the-art models to generate inclusive images from a prompt. Cheng Zhang 0014, Xuanbai Chen, Siqi Chai, Chen Henry Wu, Dmitry Lagun, Thabo Beeler, Fernando De la Torre |
ICCV | 7 |
| 2023 | Controllable 3D Generative Adversarial Face Model via Disentangling Shape and Appearanceabstract3D face modeling has been an active area of research in computer vision and computer graphics, fueling applications ranging from facial expression transfer in virtual avatars to synthetic data generation. Existing 3D deep learning generative models (e.g., VAE, GANs) allow generating compact face representations (both shape and texture) that can model non-linearities in the shape and appearance space (e.g., scatter effects, specularities,..). However, they lack the capability to control the generation of subtle expressions. This paper proposes a new 3D face generative model that can decouple identity and expression and provides granular control over expressions. In particular, we propose using a pair of supervised auto-encoder and generative adversarial networks to produce high-quality 3D faces, both in terms of appearance and shape. Experimental results in the generation of 3D faces learned with holistic expression labels, or Action Unit (AU) labels, show how we can decouple identity and expression; gaining fine-control over expressions while preserving identity.1 Fariborz Taherkhani, Aashish Rai, Quankai Gao, Shaunak Srivastava, Xuanbai Chen, Fernando De la Torre, Steven Song, Aayush Prakash, Daeil Kim |
WACV | 6 |
| 2023 | SelfPose: 3D Egocentric Pose Estimation From a Headset Mounted CameraabstractWe present a new solution to egocentric 3D body pose estimation from monocular images captured from a downward looking fish-eye camera installed on the rim of a head mounted virtual reality device. This unusual viewpoint leads to images with unique visual appearance, characterized by severe self-occlusions and strong perspective distortions that result in a drastic difference in resolution between lower and upper body. We propose a new encoder-decoder architecture with a novel multi-branch decoder designed specifically to account for the varying uncertainty in 2D joint locations. Our quantitative evaluation, both on synthetic and real-world datasets, shows that our strategy leads to substantial improvements in accuracy over state of the art egocentric pose estimation approaches. To tackle the severe lack of labelled training data for egocentric 3D pose estimation we also introduced a large-scale photo-realistic synthetic dataset. xR-EgoPose offers 383K frames of high quality renderings of people with diverse skin tones, body shapes and clothing, in a variety of backgrounds and lighting conditions, performing a range of actions. Our experiments show that the high variability in our new synthetic training corpus leads to good generalization to real world footage and to state of the art results on real world datasets with ground truth. Moreover, an evaluation on the Human3.6M benchmark shows that the performance of our method is on par with top performing approaches on the more classic problem of 3D human pose from a third person viewpoint. Denis Tomè, Thiemo Alldieck, Patrick Peluse, Gerard Pons-Moll, Lourdes Agapito, Hernán Badino, Fernando De la Torre |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Robust Egocentric Photo-realistic Facial Expression Transfer for Virtual RealityabstractSocial presence, the feeling of being there with a “real” person, will fuel the next generation of communication systems driven by digital humans in virtual reality (VR). The best 3D video-realistic VR avatars that minimize the uncanny effect rely on person-specific (PS) models. However, these PS models are time-consuming to build and are typically trained with limited data variability, which results in poor generalization and robustness. Major sources of variability that affects the accuracy of facial expression transfer algorithms include using different VR headsets (e.g., camera configuration, slop of the headset), facial appearance changes over time (e.g., beard, make-up), and environmental factors (e.g., lighting, backgrounds). This is a major drawback for the scalability of these models in VR. This paper makes progress in overcoming these limitations by proposing an end-to-end multi-identity architecture (MIA) trained with specialized augmentation strategies. MIA drives the shape component of the avatar from three cameras in the VR headset (two eyes, one mouth), in untrained subjects, using minimal personalized information (i.e., neutral 3D mesh shape). Similarly, if the PS texture decoder is available, MIA is able to drive the full avatar (shape + texture) robustly outperforming PS models in challenging scenarios. Our key contribution to improve robustness and generalization, is that our method implicitly decouples, in an unsupervised manner, the facial expression from nuisance factors (e.g., headset, environment, facial appearance). We demonstrate the superior performance and robustness of the proposed method versus state-of-the-art PS approaches in a variety of experiments. Amin Jourabloo, Fernando De la Torre, Jason M. Saragih, Shih-En Wei, Stephen Lombardi, Te-Li Wang, Danielle Belko, Autumn Trimble, Hernán Badino |
CVPR | 2 |
| 2022 | Generative Visual Prompt: Unifying Distributional Control of Pre-Trained Generative ModelsabstractGenerative models (e.g., GANs, diffusion models) learn the underlying data distribution in an unsupervised manner. However, many applications of interest require sampling from a particular region of the output space or sampling evenly over a range of characteristics. For efficient sampling in these scenarios, we propose Generative Visual Prompt (PromptGen), a framework for distributional control over pre-trained generative models by incorporating knowledge of other off-the-shelf models. PromptGen defines control as energy-based models (EBMs) and samples images in a feed-forward manner by approximating the EBM with invertible neural networks, avoiding optimization at inference. Our experiments demonstrate how PromptGen can efficiently sample from several unconditional generative models (e.g., StyleGAN2, StyleNeRF, diffusion autoencoder, NVAE) in a controlled or/and de-biased manner using various off-the-shelf models: (1) with the CLIP model as control, PromptGen can sample images guided by text, (2) with image classifiers as control, PromptGen can de-bias generative models across a set of attributes or attribute combinations, and (3) with inverse graphics models as control, PromptGen can sample images of the same identity in different poses. (4) Finally, PromptGen reveals that the CLIP model shows a "reporting bias" when used as control, and PromptGen can further de-bias this controlled distribution in an iterative manner. The code is available at https://github.com/ChenWu98/Generative-Visual-Prompt. Chen Henry Wu, Saman Motamed, Shaunak Srivastava, Fernando De la Torre |
NeurIPS | 4 |
| 2022 | 3D Human Pose, Shape and Texture From Low-Resolution Images and Videosabstract3D human pose and shape estimation from monocular images has been an active research area in computer vision. Existing deep learning methods for this task rely on high-resolution input, which however, is not always available in many scenarios such as video surveillance and sports broadcasting. Two common approaches to deal with low-resolution images are applying super-resolution techniques to the input, which may result in unpleasant artifacts, or simply training one model for each resolution, which is impractical in many realistic applications. To address the above issues, this paper proposes a novel algorithm called RSC-Net, which consists of a Resolution-aware network, a Self-supervision loss, and a Contrastive learning scheme. The proposed method is able to learn 3D body pose and shape across different resolutions with one single model. The self-supervision loss enforces scale-consistency of the output, and the contrastive learning scheme enforces scale-consistency of the deep features. We show that both these new losses provide robustness when learning in a weakly-supervised manner. Moreover, we extend the RSC-Net to handle low-resolution videos and apply it to reconstruct textured 3D pedestrians from low-resolution input. Extensive experiments demonstrate that the RSC-Net can achieve consistently better results than the state-of-the-art methods for challenging low-resolution images. Xiangyu Xu 0002, Hao Chen 0102, Francesc Moreno-Noguer, László A. Jeni, Fernando De la Torre |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | High-Fidelity Face Tracking for AR/VR via Deep Lighting Adaptationabstract3D video avatars can empower virtual communications by providing compression, privacy, entertainment, and a sense of presence in AR/VR. Best 3D photo-realistic AR/VR avatars driven by video, that can minimize uncanny effects, rely on person-specific models. However, existing person-specific photo-realistic 3D models are not robust to lighting, hence their results typically miss subtle facial behaviors and cause artifacts in the avatar. This is a major drawback for the scalability of these models in communication systems (e.g., Messenger, Skype, FaceTime) and AR/VR. This paper addresses previous limitations by learning a deep learning lighting model, that in combination with a high-quality 3D face tracking algorithm, provides a method for subtle and robust facial motion transfer from a regular video to a 3D photo-realistic avatar. Extensive experimental validation and comparisons to other state-of-the-art methods demonstrate the effectiveness of the proposed framework in real-world scenarios with variability in pose, expression, and illumination. Our project page can be found at this website. Chen Cao 0001, Fernando De la Torre, Jason M. Saragih, Chenliang Xu, Yaser Sheikh |
CVPR | 3 |
| 2021 | Pixel Codec Avatars
Shugao Ma, Tomas Simon, Jason M. Saragih, Yuecheng Li, Fernando De la Torre, Yaser Sheikh |
CVPR | 6 |
| 2021 | Implicit HRTF Modeling Using Temporal Convolutional NetworksabstractEstimation of accurate head-related transfer functions (HRTFs) is crucial to achieve realistic binaural acoustic experiences. HRTFs depend on source/listener locations and are therefore expensive and cumbersome to measure; traditional approaches require listener-dependent measurements of HRTFs at thousands of distinct spatial directions in an anechoic chamber. In this work, we present a data-driven approach to learn HRTFs implicitly with a neural network that achieves state of the art results compared to traditional approaches but relies on a much simpler data capture that can be performed in arbitrary, non-anechoic rooms. Despite that simpler and less acoustically ideal data capture, our deep learning based approach learns HRTF of high quality. We show in a perceptual study that the produced binaural audio is ranked on par with traditional DSP approaches by humans and illustrate that interaural time differences (ITDs), interaural level differences (ILDs) and spectral clues are accurately estimated. Israel D. Gebru, Dejan Markovic, Alexander Richard, Steven Krenn, Gladstone Alexander Butler, Fernando De la Torre, Yaser Sheikh |
ICASSP | 6 |
| 2021 | MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementabstractThis paper presents a generic method for generating full facial 3D animation from speech. Existing approaches to audio-driven facial animation exhibit uncanny or static upper face animation, fail to produce accurate and plausible co-articulation or rely on person-specific models that limit their scalability. To improve upon existing models, we propose a generic audio-driven facial animation approach that achieves highly realistic motion synthesis results for the entire face. At the core of our approach is a categorical latent space for facial animation that disentangles audio-correlated and audio-uncorrelated information based on a novel cross-modality loss. Our approach ensures highly accurate lip motion, while also synthesizing plausible animation of the parts of the face that are uncorrelated to the audio signal, such as eye blinks and eye brow motion. We demonstrate that our approach outperforms several baselines and obtains state-of-the-art quality both qualitatively and quantitatively. A perceptual user study demonstrates that our approach is deemed more realistic than the current state-of-the-art in over 75% of cases. We recommend watching the supplemental video before reading the paper: https://github.com/facebookresearch/meshtalk Alexander Richard, Michael Zollhöfer, Yandong Wen, Fernando De la Torre, Yaser Sheikh |
ICCV | 4 |
| 2021 | Neural Synthesis of Binaural Speech From Mono Audio
Alexander Richard, Dejan Markovic, Israel D. Gebru, Steven Krenn, Gladstone Alexander Butler, Fernando De la Torre, Yaser Sheikh |
ICLR | 6 |
| 2021 | Audio- and Gaze-driven Facial Animation of Codec AvatarsabstractCodec Avatars are a recent class of learned, photorealistic face models that accurately represent the geometry and texture of a person in 3D (i.e., for virtual reality), and are almost indistinguishable from video [28]. In this paper we describe the first approach to animate these parametric models in real-time which could be deployed on commodity virtual reality hardware using audio and/or eye tracking. Our goal is to display expressive conversations between individuals that exhibit important social signals such as laughter and excitement solely from la-tent cues in our lossy input signals. To this end we collected over 5 hours of high frame rate 3D face scans across three participants including traditional neutral speech as well as expressive and conversational speech. We investigate a multimodal fusion approach that dynamically identifies which sensor encoding should animate which parts of the face at any time. See the supplemental video which demonstrates our ability to generate full face motion far beyond the typically neutral lip articulations seen in competing work: https://research.fb.com/videos/audio-and-gaze-driven-facial-animation-of-codec-avatars/. Alexander Richard, Colin Lea, Shugao Ma, Juergen Gall, Fernando De la Torre, Yaser Sheikh |
WACV | 5 |
| 2021 | Weakly-Supervised Learning of Category-Specific 3D Object ShapesabstractCategory-specific 3D object shape models have greatly boosted the recent advances in object detection, recognition and segmentation. However, even the most advanced approach for learning 3D object shapes still requires heavy manual annotations on large-scale 2D images. Such annotations include object categories, object keypoints, and figure-ground segmentation for the instances in each image. In particular, annotating figure-ground segmentation is unbearably labor-intensive and time-consuming. To address this problem, this paper devotes to learn category-specific 3D shape models under weak supervision, where only object categories and keypoints are required to be manually annotated on the training 2D images. By exploring the underlying relationship between two tasks: object segmentation and category-specific 3D shape reconstruction, we propose a novel weakly-supervised learning framework to jointly address these two tasks and combine them to boost the final performance of the learned 3D shape models. Moreover, learning without using figure-ground segmentation leads to ambiguous solutions. To this end, we develop the confidence weighting schemes in the viewpoint estimation and 3D shape learning procedure. These schemes effectively reduce the confusion caused by the noisy data and thus increase the chances for recovering more reliable 3D object shapes. Comprehensive experiments on the challenging PASCAL VOC benchmark show that our framework achieves comparable performance with the state-of-the-art methods that use expensive manual segmentation-level annotations. In addition, our experiments also demonstrate that our 3D shape models improve object segmentation performance. Junwei Han 0001, Yang Yang 0009, Dingwen Zhang, Dong Huang 0007, Dong Xu 0001, Fernando De la Torre |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Real-time 3D neural facial animation from binocular videoabstractWe present a method for performing real-time facial animation of a 3D avatar from binocular video. Existing facial animation methods fail to automatically capture precise and subtle facial motions for driving a photo-realistic 3D avatar "in-the-wild" (i.e., variability in illumination, camera noise). The novelty of our approach lies in a light-weight process for specializing a personalized face model to new environments that enables extremely accurate real-time face tracking anywhere. Our method uses a pre-trained high-fidelity personalized model of the face that we complement with a novel illumination model to account for variations due to lighting and other factors often encountered in-the-wild (e.g., facial hair growth, makeup, skin blemishes). Our approach comprises two steps. First, we solve for our illumination model's parameters by applying analysis-by-synthesis on a short video recording. Using the pairs of model parameters (rigid, non-rigid) and the original images, we learn a regression for real-time inference from the image space to the 3D shape and texture of the avatar. Second, given a new video, we fine-tune the real-time regression model with a few-shot learning strategy to adapt the regression model to the new environment. We demonstrate our system's ability to precisely capture subtle facial motions in unconstrained scenarios, in comparison to competing methods, on a diverse collection of identities, expressions, and real-world environments. Chen Cao 0001, Vasu Agrawal, Fernando De la Torre, Jason M. Saragih, Tomas Simon, Yaser Sheikh |
ACM Trans. Graph. | 3 |
| 2020 | Expressive Telepresence via Modular Codec Avatars
Hang Chu, Shugao Ma, Fernando De la Torre, Sanja Fidler, Yaser Sheikh |
ECCV (12) | 3 |
| 2020 | 3D Human Shape and Pose from a Single Low-Resolution Image with Self-Supervised Learning
Xiangyu Xu 0002, Hao Chen 0102, Francesc Moreno-Noguer, László A. Jeni, Fernando De la Torre |
ECCV (9) | 5 |
| 2019 | Editorial: Special Issue on Deep Learning for Face Analysis
Chen Change Loy, Xiaoming Liu 0002, Tae-Kyun Kim 0001, Fernando De la Torre, Rama Chellappa |
Int. J. Comput. Vis. | 4 |
| 2019 | Learning facial action units with spatiotemporal cues and multi-label sampling
Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn |
Image Vis. Comput. | 2 |
| 2019 | Discriminative Optimization: Theory and Applications to Computer VisionabstractMany computer vision problems are formulated as the optimization of a cost function. This approach faces two main challenges: designing a cost function with a local optimum at an acceptable solution, and developing an efficient numerical method to search for this optimum. While designing such functions is feasible in the noiseless case, the stability and location of local optima are mostly unknown under noise, occlusion, or missing data. In practice, this can result in undesirable local optima or not having a local optimum in the expected place. On the other hand, numerical optimization algorithms in high-dimensional spaces are typically local and often rely on expensive first or second order information to guide the search. To overcome these limitations, we propose Discriminative Optimization (DO), a method that learns search directions from data without the need of a cost function. DO explicitly learns a sequence of updates in the search space that leads to stationary points that correspond to the desired solutions. We provide a formal analysis of DO and illustrate its benefits in the problem of 3D registration, camera pose estimation, and image denoising. We show that DO outperformed or matched state-of-the-art algorithms in terms of accuracy, robustness, and computational efficiency. Jayakorn Vongkulbhisal, Fernando De la Torre, João Paulo Costeira |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Road Curb Detection and Localization With Monocular Forward-View Vehicle CameraabstractWe propose a robust method for estimating road curb 3-D parameters (size, location, and orientation) using a calibrated monocular camera equipped with a fisheye lens. Automatic curb detection and localization is particularly important in the context of an advanced driver assistance system, i.e., to prevent possible collision and damage to the vehicle's bumper during perpendicular and diagonal parking maneuvers. Combining 3-D geometric reasoning with advanced vision-based detection methods, our approach is able to estimate the vehicle to curb distance in real time with a mean accuracy of more than 90%, as well as its orientation, height, and depth. Our approach consists of two distinct components-curb detection in each individual video frame and temporal analysis. The first part is comprised of sophisticated curb edges extraction and parameterized 3-D curb template fitting. Using a few assumptions regarding the real-world geometry, we can thus retrieve the curb's height and its relative position with respect to the moving vehicle on which the camera is mounted. Support vector machine classifier fed with histograms of oriented gradients is used for appearance-based filtering out outliers. In the second part, the detected curb regions are tracked in the temporal domain, so as to perform a second pass of false positives rejection. We have validated our approach on a newly collected database of 11 videos under different conditions. We have used point-wise LIDAR measurements and manual exhaustive labels as a ground truth. Stanislav Panev, Francisco Vicente 0001, Fernando De la Torre, Véronique Prinet |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2018 | Inverse Composition Discriminative Optimization for Point Cloud RegistrationabstractRigid Point Cloud Registration (PCReg) refers to the problem of finding the rigid transformation between two sets of point clouds. This problem is particularly important due to the advances in new 3D sensing hardware, and it is challenging because neither the correspondence nor the transformation parameters are known. Traditional local PCReg methods (e.g., ICP) rely on local optimization algorithms, which can get trapped in bad local minima in the presence of noise, outliers, bad initializations, etc. To alleviate these issues, this paper proposes Inverse Composition Discriminative Optimization (ICDO), an extension of Discriminative Optimization (DO), which learns a sequence of update steps from synthetic training data that search the parameter space for an improved solution. Unlike DO, ICDO is object-independent and generalizes even to unseen shapes. We evaluated ICDO on both synthetic and real data, and show that ICDO can match the speed and outperform the accuracy of state-of-the-art PCReg algorithms. Jayakorn Vongkulbhisal, Beñat Irastorza Ugalde, Fernando De la Torre, João Paulo Costeira |
CVPR | 3 |
| 2018 | Error-Correcting FactorizationabstractError Correcting Output Codes (ECOC) is a successful technique in multi-class classification, which is a core problem in Pattern Recognition and Machine Learning. A major advantage of ECOC over other methods is that the multi-class problem is decoupled into a set of binary problems that are solved independently. However, literature defines a general error-correcting capability for ECOCs without analyzing how it distributes among classes, hindering a deeper analysis of pair-wise error-correction. To address these limitations this paper proposes an Error-Correcting Factorization (ECF) method. Our contribution is three fold: (I) We propose a novel representation of the error-correction capability, called the design matrix, that enables us to build an ECOC on the basis of allocating correction to pairs of classes. (II) We derive the optimal code length of an ECOC using rank properties of the design matrix. (III) ECF is formulated as a discrete optimization problem, and a relaxed solution is found using an efficient constrained block coordinate descent approach. (IV) Enabled by the flexibility introduced with the design matrix we propose to allocate the error-correction on classes that are prone to confusion. Experimental results in several databases show that when allocating the error-correction to confusable classes ECF outperforms state-of-the-art approaches. Miguel Ángel Bautista 0001, Oriol Pujol, Fernando De la Torre, Sergio Escalera |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | A Functional Regression Approach to Facial Landmark TrackingabstractLinear regression is a fundamental building block in many face detection and tracking algorithms, typically used to predict shape displacements from image features through a linear mapping. This paper presents a Functional Regression solution to the least squares problem, which we coin Continuous Regression, resulting in the first real-time incremental face tracker. Contrary to prior work in Functional Regression, in which B-splines or Fourier series were used, we propose to approximate the input space by its first-order Taylor expansion, yielding a closed-form solution for the continuous domain of displacements. We then extend the continuous least squares problem to correlated variables, and demonstrate the generalisation of our approach. We incorporate Continuous Regression into the cascaded regression framework, and show its computational benefits for both training and testing. We then present a fast approach for incremental learning within Cascaded Continuous Regression, coined iCCR, and show that its complexity allows real-time face tracking, being 20 times faster than the state of the art. To the best of our knowledge, this is the first incremental face tracker that is shown to operate in real-time. We show that iCCR achieves state-of-the-art performance on the 300-VW dataset, the most recent, large-scale benchmark for face tracking. Enrique Sánchez-Lozano, Georgios Tzimiropoulos, Brais Martínez, Fernando De la Torre, Michel F. Valstar |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Soft-Margin Mixture of Regressions
Dong Huang 0007, Longfei Han, Fernando De la Torre |
CVPR | 3 |
| 2017 | Additive Component AnalysisabstractPrincipal component analysis (PCA) is one of the most versatile tools for unsupervised learning with applications ranging from dimensionality reduction to exploratory data analysis and visualization. While much effort has been devoted to encouraging meaningful representations through regularization (e.g. non-negativity or sparsity), underlying linearity assumptions can limit their effectiveness. To address this issue, we propose Additive Component Analysis (ACA), a novel nonlinear extension of PCA. Inspired by multivariate nonparametric regression with additive models, ACA fits a smooth manifold to data by learning an explicit mapping from a low-dimensional latent space to the input space, which trivially enables applications like denoising. Furthermore, ACA can be used as a drop-in replacement in many algorithms that use linear component analysis methods as a subroutine via the local tangent space of the learned manifold. Unlike many other nonlinear dimensionality reduction techniques, ACA can be efficiently applied to large datasets since it does not require computing pairwise similarities or storing training data during testing. Multiple ACA layers can also be composed and learned jointly with essentially the same procedure for improved representational power, demonstrating the encouraging potential of nonparametric deep learning. We evaluate ACA on a variety of datasets, showing improved robustness, reconstruction performance, and interpretability. Calvin Murdock, Fernando De la Torre |
CVPR | 2 |
| 2017 | Discriminative Optimization: Theory and Applications to Point Cloud RegistrationabstractMany computer vision problems are formulated as the optimization of a cost function. This approach faces two main challenges: (1) designing a cost function with a local optimum at an acceptable solution, and (2) developing an efficient numerical method to search for one (or multiple) of these local optima. While designing such functions is feasible in the noiseless case, the stability and location of local optima are mostly unknown under noise, occlusion, or missing data. In practice, this can result in undesirable local optima or not having a local optimum in the expected place. On the other hand, numerical optimization algorithms in high-dimensional spaces are typically local and often rely on expensive first or second order information to guide the search. To overcome these limitations, this paper proposes Discriminative Optimization (DO), a method that learns search directions from data without the need of a cost function. Specifically, DO explicitly learns a sequence of updates in the search space that leads to stationary points that correspond to desired solutions. We provide a formal analysis of DO and illustrate its benefits in the problem of 2D and 3D point cloud registration both in synthetic and range-scan data. We show that DO outperforms state-of-the-art algorithms by a large margin in terms of accuracy, robustness to perturbations, and computational efficiency. Jayakorn Vongkulbhisal, Fernando De la Torre, João Paulo Costeira |
CVPR | 2 |
| 2017 | Learning Spatial and Temporal Cues for Multi-Label Facial Action Unit DetectionabstractFacial action units (AU) are the fundamental units to decode human facial expressions. At least three aspects affect performance of automated AU detection: spatial representation, temporal modeling, and AU correlation. Unlike most studies that tackle these aspects separately, we propose a hybrid network architecture to jointly model them. Specifically, spatial representations are extracted by a Convolutional Neural Network (CNN), which, as analyzed in this paper, is able to reduce person-specific biases caused by hand-crafted descriptors (e.g., HOG and Gabor). To model temporal dependencies, Long Short-Term Memory (LSTMs) are stacked on top of these representations, regardless of the lengths of input videos. The outputs of CNNs and LSTMs are further aggregated into a fusion network to produce per-frame prediction of 12 AUs. Our network naturally addresses the three issues together, and yields superior performance compared to existing methods that consider these issues independently. Extensive experiments were conducted on two large spontaneous datasets, GFT and BP4D, with more than 400,000 frames coded with 12 AUs. On both datasets, we report improvements over a standard multi-label CNN and feature-based state-of-the-art. Finally, we provide visualization of the learned AU models, which, to our best knowledge, reveal how machines see AUs for the first time. Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn |
FG | 2 |
| 2017 | Approximate Grassmannian Intersections: Subspace-Valued Subspace Learning
Calvin Murdock, Fernando De la Torre |
ICCV | 2 |
| 2017 | Deep multi-task learning for gait-based biometricsabstractThe task of identifying people by the way they walk is known as `gait recognition'. Although gait is mainly used for identification, additional tasks as gender recognition or age estimation may be addressed based on gait as well. In such cases, traditional approaches consider those tasks as independent ones, defining separated task-specific features and models for them. This paper shows that by training jointly more than one gait-based tasks, the identification task converges faster than when it is trained independently, and the recognition performance of multi-task models is equal or superior to more complex single-task ones. Our model is a multi-task CNN that receives as input a fixed-length sequence of optical flow channels and outputs several biometric features (identity, gender and age). Manuel J. Marín-Jiménez, Francisco M. Castro, Nicolás Guil, Fernando De la Torre, Rafael Medina Carnicer |
ICIP | 4 |
| 2017 | A Branch-and-Bound Framework for Unsupervised Common Event Discovery
Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn, Daniel S. Messinger |
Int. J. Comput. Vis. | 2 |
| 2017 | Subspace Procrustes Analysis
Xavier Perez-Sala, Fernando De la Torre, Laura Igual, Sergio Escalera, Cecilio Angulo |
Int. J. Comput. Vis. | 2 |
| 2017 | Selective Transfer Machine for Personalized Facial Expression AnalysisabstractAutomatic facial action unit (AU) and expression detection from videos is a long-standing problem. The problem is challenging in part because classifiers must generalize to previously unknown subjects that differ markedly in behavior and facial morphology (e.g., heavy versus delicate brows, smooth versus deeply etched wrinkles) from those on which the classifiers are trained. While some progress has been achieved through improvements in choices of features and classifiers, the challenge occasioned by individual differences among people remains. Person-specific classifiers would be a possible solution but for a paucity of training data. Sufficient training data for person-specific classifiers typically is unavailable. This paper addresses the problem of how to personalize a generic classifier without additional labels from the test subject. We propose a transductive learning method, which we refer to as a Selective Transfer Machine (STM), to personalize a generic classifier by attenuating person-specific mismatches. STM achieves this effect by simultaneously learning a classifier and re-weighting the training samples that are most relevant to the test subject. We compared STM to both generic classifiers and cross-domain learning methods on four benchmarks: CK+ [44], GEMEP-FERA [67], RUFACS [4] and GFT [57]. STM outperformed generic classifiers in all. Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Motion from Structure (MfS): Searching for 3D Objects in Cluttered Point TrajectoriesabstractObject detection has been a long standing problem in computer vision, and state-of-the-art approaches rely on the use of sophisticated features and/or classifiers. However, these learning-based approaches heavily depend on the quality and quantity of labeled data, and do not generalize well to extreme poses or textureless objects. In this work, we explore the use of 3D shape models to detect objects in videos in an unsupervised manner. We call this problem Motion from Structure (MfS): given a set of point trajectories and a 3D model of the object of interest, find a subset of trajectories that correspond to the 3D model and estimate its alignment (i.e., compute the motion matrix). MfS is related to Structure from Motion (SfM) and motion segmentation problems: unlike SfM, the structure of the object is known but the correspondence between the trajectories and the object is unknown, unlike motion segmentation, the MfS problem incorporates 3D structure, providing robustness to tracking mismatches and outliers. Experiments illustrate how our MfS algorithm outperforms alternative approaches in both synthetic data and real videos extracted from YouTube. Jayakorn Vongkulbhisal, Ricardo Silveira Cabral, Fernando De la Torre, João Paulo Costeira |
CVPR | 3 |
| 2016 | Cascade of Tasks for facial expression analysis
Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn |
Image Vis. Comput. | 3 |
| 2016 | Robust RegressionabstractDiscriminative methods (e.g., kernel regression, SVM) have been extensively used to solve problems such as object recognition, image alignment and pose estimation from images. These methods typically map image features ( X) to continuous (e.g., pose) or discrete (e.g., object category) values. A major drawback of existing discriminative methods is that samples are directly projected onto a subspace and hence fail to account for outliers common in realistic training sets due to occlusion, specular reflections or noise. It is important to notice that existing discriminative approaches assume the input variables X to be noise free. Thus, discriminative methods experience significant performance degradation when gross outliers are present. Despite its obvious importance, the problem of robust discriminative learning has been relatively unexplored in computer vision. This paper develops the theory of robust regression (RR) and presents an effective convex approach that uses recent advances on rank minimization. The framework applies to a variety of problems in computer vision including robust linear discriminant analysis, regression with missing data, and multi-label classification. Several synthetic and real examples with applications to head pose estimation from images, image and video classification and facial attribute classification with missing data are used to illustrate the benefits of RR. Dong Huang 0007, Ricardo Silveira Cabral, Fernando De la Torre |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | Generalized Canonical Time WarpingabstractTemporal alignment of human motion has been of recent interest due to its applications in animation, tele-rehabilitation and activity recognition. This paper presents generalized canonical time warping (GCTW), an extension of dynamic time warping (DTW) and canonical correlation analysis (CCA) for temporally aligning multi-modal sequences from multiple subjects performing similar activities. GCTW extends previous work on DTW and CCA in several ways: (1) it combines CCA with DTW to align multi-modal data (e.g., video and motion capture data); (2) it extends DTW by using a linear combination of monotonic functions to represent the warping path, providing a more flexible temporal warp. Unlike exact DTW, which has quadratic complexity, we propose a linear time algorithm to minimize GCTW. (3) GCTW allows simultaneous alignment of multiple sequences. Experimental results on aligning multi-modal data, facial expressions, motion capture data and video illustrate the benefits of GCTW. The code is available at http://humansensing.cs.cmu.edu/ctw. Feng Zhou 0002, Fernando De la Torre |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Spatio-Temporal Matching for Human Pose Estimation in VideoabstractDetection and tracking humans in videos have been long-standing problems in computer vision. Most successful approaches (e.g., deformable parts models) heavily rely on discriminative models to build appearance detectors for body joints and generative models to constrain possible body configurations (e.g., trees). While these 2D models have been successfully applied to images (and with less success to videos), a major challenge is to generalize these models to cope with camera views. In order to achieve view-invariance, these 2D models typically require a large amount of training data across views that is difficult to gather and time-consuming to label. Unlike existing 2D models, this paper formulates the problem of human detection in videos as spatio-temporal matching (STM) between a 3D motion capture model and trajectories in videos. Our algorithm estimates the camera view and selects a subset of tracked trajectories that matches the motion of the 3D model. The STM is efficiently solved with linear programming, and it is robust to tracking mismatches, occlusions and outliers. To the best of our knowledge this is the first paper that solves the correspondence between video and 3D motion capture data for human pose detection. Experiments on the CMU motion capture, Human3.6M, Berkeley MHAD and CMU MAD databases illustrate the benefits of our method over state-of-the-art approaches. Feng Zhou 0002, Fernando De la Torre |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Factorized Graph MatchingabstractGraph matching (GM) is a fundamental problem in computer science, and it plays a central role to solve correspondence problems in computer vision. GM problems that incorporate pairwise constraints can be formulated as a quadratic assignment problem (QAP). Although widely used, solving the correspondence problem through GM has two main limitations: (1) the QAP is NP-hard and difficult to approximate; (2) GM algorithms do not incorporate geometric constraints between nodes that are natural in computer vision problems. To address aforementioned problems, this paper proposes factorized graph matching (FGM). FGM factorizes the large pairwise affinity matrix into smaller matrices that encode the local structure of each graph and the pairwise affinity between edges. Four are the benefits that follow from this factorization: (1) There is no need to compute the costly (in space and time) pairwise affinity matrix; (2) The factorization allows the use of a path-following optimization algorithm, that leads to improved optimization strategies and matching performance; (3) Given the factorization, it becomes straight-forward to incorporate geometric transformations (rigid and non-rigid) to the GM problem. (4) Using a matrix formulation for the GM problem and the factorization, it is easy to reveal commonalities and differences between different GM methods. The factorization also provides a clean connection with other matching algorithms such as iterative closest point; Experimental results on synthetic and real databases illustrate how FGM outperforms state-of-the-art algorithms for GM. The code is available at http://humansensing.cs.cmu.edu/fgm. Feng Zhou 0002, Fernando De la Torre |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Feature and Region Selection for Visual LearningabstractVisual learning problems, such as object classification and action recognition, are typically approached using extensions of the popular bag-of-words (BoWs) model. Despite its great success, it is unclear what visual features the BoW model is learning. Which regions in the image or video are used to discriminate among classes? Which are the most discriminative visual words? Answering these questions is fundamental for understanding existing BoW models and inspiring better models for visual recognition. To answer these questions, this paper presents a method for feature selection and region selection in the visual BoW model. This allows for an intermediate visualization of the features and regions that are important for visual learning. The main idea is to assign latent weights to the features or regions, and jointly optimize these latent variables with the parameters of a classifier (e.g., support vector machine). There are four main benefits of our approach: 1) our approach accommodates non-linear additive kernels, such as the popular χ(2) and intersection kernel; 2) our approach is able to handle both regions in images and spatio-temporal regions in videos in a unified way; 3) the feature selection problem is convex, and both problems can be solved using a scalable reduced gradient method; and 4) we point out strong connections with multiple kernel learning and multiple instance learning approaches. Experimental results in the PASCAL VOC 2007, MSR Action Dataset II and YouTube illustrate the benefits of our approach. Ji Zhao 0001, Liantao Wang, Ricardo Silveira Cabral, Fernando De la Torre |
IEEE Trans. Image Process. | 4 |
| 2016 | Confidence Preserving Machine for Facial Action Unit DetectionabstractFacial action unit (AU) detection from video has been a long-standing problem in the automated facial expression analysis. While progress has been made, accurate detection of facial AUs remains challenging due to ubiquitous sources of errors, such as inter-personal variability, pose, and low-intensity AUs. In this paper, we refer to samples causing such errors as hard samples, and the remaining as easy samples. To address learning with the hard samples, we propose the confidence preserving machine (CPM), a novel two-stage learning framework that combines multiple classifiers following an "easy-to-hard" strategy. During the training stage, CPM learns two confident classifiers. Each classifier focuses on separating easy samples of one class from all else, and thus preserves confidence on predicting each class. During the test stage, the confident classifiers provide "virtual labels" for easy test samples. Given the virtual labels, we propose a quasi-semi-supervised (QSS) learning strategy to learn a person-specific classifier. The QSS strategy employs a spatio-temporal smoothness that encourages similar predictions for samples within a spatio-temporal neighborhood. In addition, to further improve detection performance, we introduce two CPM extensions: iterative CPM that iteratively augments training samples to train the confident classifiers, and kernel CPM that kernelizes the original CPM model to promote nonlinearity. Experiments on four spontaneous data sets GFT, BP4D, DISFA, and RU-FACS illustrate the benefits of the proposed CPM models over baseline methods and the state-of-the-art semi-supervised learning and transfer learning methods. Jiabei Zeng, Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn, Zhang Xiong 0001 |
IEEE Trans. Image Process. | 3 |
| 2016 | Joint Patch and Multi-label Learning for Facial Action Unit and Holistic Expression RecognitionabstractMost action unit (AU) detection methods use one-versus-all classifiers without considering dependences between features or AUs. In this paper, we introduce a joint patch and multi-label learning (JPML) framework that models the structured joint dependence behind features, AUs, and their interplay. In particular, JPML leverages group sparsity to identify important facial patches, and learns a multi-label classifier constrained by the likelihood of co-occurring AUs. To describe such likelihood, we derive two AU relations, positive correlation and negative competition, by statistically analyzing more than 350,000 video frames annotated with multiple AUs. To the best of our knowledge, this is the first work that jointly addresses patch learning and multi-label learning for AU detection. In addition, we show that JPML can be extended to recognize holistic expressions by learning common and specific patches, which afford a more compact representation than the standard expression recognition methods. We evaluate JPML on three benchmark datasets CK+, BP4D, and GFT, using within-and cross-dataset scenarios. In four of five experiments, JPML achieved the highest averaged F1 scores in comparison with baseline and alternative methods that use either patch learning or multi-label learning alone. Kaili Zhao, Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn, Honggang Zhang 0002 |
IEEE Trans. Image Process. | 3 |
| 2015 | Global supervised descent methodabstractMathematical optimization plays a fundamental role in solving many problems in computer vision (e.g., camera calibration, image alignment, structure from motion). It is generally accepted that second order descent methods are the most robust, fast, and reliable approaches for nonlinear optimization of a general smooth function. However, in the context of computer vision, second order descent methods have two main drawbacks: 1) the function might not be analytically differentiable and numerical approximations are impractical, and 2) the Hessian may be large and not positive definite. Recently, Supervised Descent Method (SDM), a method that learns the “weighted averaged gradients” in a supervised manner has been proposed to solve these issues. However, SDM is a local algorithm and it is likely to average conflicting gradient directions. This paper proposes Global SDM (GSDM), an extension of SDM that divides the search space into regions of similar gradient directions. GSDM provides a better and more efficient strategy to minimize non-linear least squares functions in computer vision problems. We illustrate the effectiveness of GSDM in two problems: non-rigid image alignment and extrinsic camera calibration. Xuehan Xiong, Fernando De la Torre |
CVPR | 2 |
| 2015 | Joint patch and multi-label learning for facial action unit detectionabstractThe face is one of the most powerful channel of nonverbal communication. The most commonly used taxonomy to describe facial behaviour is the Facial Action Coding System (FACS). FACS segments the visible effects of facial muscle activation into 30+ action units (AUs). AUs, which may occur alone and in thousands of combinations, can describe nearly all-possible facial expressions. Most existing methods for automatic AU detection treat the problem using one-vs-all classifiers and fail to exploit dependencies among AU and facial features. We introduce joint-patch and multi-label learning (JPML) to address these issues. JPML leverages group sparsity by selecting a sparse subset of facial patches while learning a multi-label classifier. In four of five comparisons on three diverse datasets, CK+, GFT, and BP4D, JPML produced the highest average F1 scores in comparison with state-of-the art. Kaili Zhao, Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn, Honggang Zhang 0002 |
CVPR | 3 |
| 2015 | Unsupervised Synchrony Discovery in Human InteractionabstractPeople are inherently social. Social interaction plays an important and natural role in human behavior. Most computational methods focus on individuals alone rather than in social context. They also require labelled training data. We present an unsupervised approach to discover interpersonal synchrony, referred as to two or more persons preforming common actions in overlapping video frames or segments. For computational efficiency, we develop a branch-and-bound (B&B) approach that affords exhaustive search while guaranteeing a globally optimal solution. The proposed method is entirely general. It takes from two or more videos any multi-dimensional signal that can be represented as a histogram. We derive three novel bounding functions and provide efficient extensions, including multi-synchrony detection and accelerated search, using a warm-start strategy and parallelism. We evaluate the effectiveness of our approach in multiple databases, including human actions using the CMU Mocap dataset [1], spontaneous facial behaviors using group-formation task dataset [37] and parent-infant interaction dataset [28]. Wen-Sheng Chu, Jiabei Zeng, Fernando De la Torre, Jeffrey F. Cohn, Daniel S. Messinger |
ICCV | 3 |
| 2015 | Semantic Component AnalysisabstractUnsupervised and weakly-supervised visual learning in large image collections are critical in order to avoid the time-consuming and error-prone process of manual labeling. Standard approaches rely on methods like multiple-instance learning or graphical models, which can be computationally intensive and sensitive to initialization. On the other hand, simpler component analysis or clustering methods usually cannot achieve meaningful invariances or semantic interpretability. To address the issues of previous work, we present a simple but effective method called Semantic Component Analysis (SCA), which provides a decomposition of images into semantic components. Unsupervised SCA decomposes additive image representations into spatially-meaningful visual components that naturally correspond to object categories. Using an overcomplete representation that allows for rich instance-level constraints and spatial priors, SCA gives improved results and more interpretable components in comparison to traditional matrix factorization techniques. If weakly-supervised information is available in the form of image-level tags, SCA factorizes a set of images into semantic groups of superpixels. We also provide qualitative connections to traditional methods for component analysis (e.g. Grassmann averages, PCA, and NMF). The effectiveness of our approach is validated through synthetic data and on the MSRC2 and Sift Flow datasets, demonstrating competitive results in unsupervised and weakly-supervised semantic segmentation. Calvin Murdock, Fernando De la Torre |
ICCV | 2 |
| 2015 | Confidence Preserving Machine for Facial Action Unit DetectionabstractVaried sources of error contribute to the challenge of facial action unit detection. Previous approaches address specific and known sources. However, many sources are unknown. To address the ubiquity of error, we propose a Confident Preserving Machine (CPM) that follows an easy-to-hard classification strategy. During training, CPM learns two confident classifiers. A confident positive classifier separates easily identified positive samples from all else, a confident negative classifier does same for negative samples. During testing, CPM then learns a person-specific classifier using "virtual labels" provided by confident classifiers. This step is achieved using a quasi-semi-supervised (QSS) approach. Hard samples are typically close to the decision boundary, and the QSS approach disambiguates them using spatio-temporal constraints. To evaluate CPM, we compared it with a baseline single-margin classifier and state-of-the-art semi-supervised learning, transfer learning, and boosting methods in three datasets of spontaneous facial behavior. With few exceptions, CPM outperformed baseline and state-of-the art methods. Jiabei Zeng, Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn, Zhang Xiong 0001 |
ICCV | 3 |
| 2015 | Matrix Completion for Weakly-Supervised Multi-Label Image ClassificationabstractIn the last few years, image classification has become an incredibly active research topic, with widespread applications. Most methods for visual recognition are fully supervised, as they make use of bounding boxes or pixelwise segmentations to locate objects of interest. However, this type of manual labeling is time consuming, error prone and it has been shown that manual segmentations are not necessarily the optimal spatial enclosure for object classifiers. This paper proposes a weakly-supervised system for multi-label image classification. In this setting, training images are annotated with a set of keywords describing their contents, but the visual concepts are not explicitly segmented in the images. We formulate the weakly-supervised image classification as a low-rank matrix completion problem. Compared to previous work, our proposed framework has three advantages: (1) Unlike existing solutions based on multiple-instance learning methods, our model is convex. We propose two alternative algorithms for matrix completion specifically tailored to visual data, and prove their convergence. (2) Unlike existing discriminative methods, our algorithm is robust to labeling errors, background noise and partial occlusions. (3) Our method can potentially be used for semantic segmentation. Experimental validation on several data sets shows that our method outperforms state-of-the-art classification algorithms, while effectively capturing each class appearance. Ricardo Silveira Cabral, Fernando De la Torre, João Paulo Costeira, Alexandre Bernardino |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Pareto models for discriminative multiclass linear dimensionality reduction
Karim T. Abou-Moustafa, Fernando De la Torre, Frank P. Ferrie |
Pattern Recognit. | 2 |
| 2015 | Estimating smile intensity: A better way
Jeffrey M. Girard, Jeffrey F. Cohn, Fernando De la Torre |
Pattern Recognit. Lett. | 3 |
| 2015 | Driver Gaze Tracking and Eyes Off the Road Detection SystemabstractDistracted driving is one of the main causes of vehicle collisions in the United States. Passively monitoring a driver's activities constitutes the basis of an automobile safety system that can potentially reduce the number of accidents by estimating the driver's focus of attention. This paper proposes an inexpensive vision-based system to accurately detect Eyes Off the Road (EOR). The system has three main components: 1) robust facial feature tracking; 2) head pose and gaze estimation; and 3) 3-D geometric reasoning to detect EOR. From the video stream of a camera installed on the steering wheel column, our system tracks facial features from the driver's face. Using the tracked landmarks and a 3-D face model, the system computes head pose and gaze direction. The head pose estimation algorithm is robust to nonrigid face deformations due to changes in expressions. Finally, using a 3-D geometric analysis, the system reliably detects EOR. Francisco Vicente 0001, Zehua Huang, Xuehan Xiong, Fernando De la Torre, Wende Zhang, Dan Levi |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2014 | Multi-label Discriminative Weakly-Supervised Human Activity Recognition and Localization
Ehsan Adeli-Mosabbeb, Ricardo Silveira Cabral, Fernando De la Torre, Mahmood Fathy |
ACCV (5) | 3 |
| 2014 | Complex Non-rigid Motion 3D Reconstruction by Union of SubspacesabstractThe task of estimating complex non-rigid 3D motion through a monocular camera is of increasing interest to the wider scientific community. Assuming one has the 2D point tracks of the non-rigid object in question, the vision community refers to this problem as Non-Rigid Structure from Motion (NRSfM). In this paper we make two contributions. First, we demonstrate empirically that the current state of the art approach to NRSfM (i.e. Dai et al. [5]) exhibits poor reconstruction performance on complex motion (i.e motions involving a sequence of primitive actions such as walk, sit and stand involving a human object). Second, we propose that this limitation can be circumvented by modeling complex motion as a union of subspaces. This does not naturally occur in Dai et al.'s approach which instead makes a less compact summation of subspaces assumption. Experiments on both synthetic and real videos illustrate the benefits of our approach for the complex nonrigid motion analysis. Yingying Zhu 0004, Dong Huang 0007, Fernando De la Torre, Simon Lucey |
CVPR | 3 |
| 2014 | Sequential Max-Margin Event Detectors
Dong Huang 0007, Shitong Yao, Fernando De la Torre |
ECCV (3) | 4 |
| 2014 | Motion Words for Videos
Ekaterina H. Taralova, Fernando De la Torre, Martial Hebert |
ECCV (1) | 2 |
| 2014 | Spatio-temporal Matching for Human Detection in Video
Feng Zhou 0002, Fernando De la Torre |
ECCV (6) | 2 |
| 2014 | Optimal no-intersection multi-label binary localization for time series using totally unimodular linear programmingabstractWe propose a new model for simultaneously localizing different classes in the same media, casting it as an integer optimization problem. Our model subsumes into a single formulation previous single and multi-class localization methods, as well as allows us to exploit optimal relaxations to the linear domain. We apply our model to the problem of multi-label multiple instance learning for tagging video collections. Given weakly labeled training samples, where tags for actions in video and objects in images are known but not their locations, our aim is to train classifiers for both detection and localization of said classes on new data. Experimental results demonstrate our approach obtains similar performances when compared to fully supervised methods. Ricardo Silveira Cabral, João Paulo Costeira, Alexandre Bernardino, Fernando De la Torre |
ICIP | 4 |
| 2014 | Multiple instance learning via Gaussian processes
Minyoung Kim 0001, Fernando De la Torre |
Data Min. Knowl. Discov. | 2 |
| 2014 | Max-Margin Early Event Detectors
Minh Hoai, Fernando De la Torre |
Int. J. Comput. Vis. | 2 |
| 2014 | Continuous Generalized Procrustes analysis
Laura Igual, Xavier Perez-Sala, Sergio Escalera, Cecilio Angulo, Fernando De la Torre |
Pattern Recognit. | 5 |
| 2014 | Learning discriminative localization from weakly labeled data
Minh Hoai, Lorenzo Torresani, Fernando De la Torre, Carsten Rother |
Pattern Recognit. | 3 |
| 2013 | Facing Imbalanced Data-Recommendations for the Use of Performance MetricsabstractRecognizing facial action units (AUs) is important for situation analysis and automated video annotation. Previous work has emphasized face tracking and registration and the choice of features classifiers. Relatively neglected is the effect of imbalanced data for action unit detection. While the machine learning community has become aware of the problem of skewed data for training classifiers, little attention has been paid to how skew may bias performance metrics. To address this question, we conducted experiments using both simulated classifiers and three major databases that differ in size, type of FACS coding, and degree of skew. We evaluated influence of skew on both threshold metrics (Accuracy, F-score, Cohen's kappa, and Krippendorf's alpha) and rank metrics (area under the receiver operating characteristic (ROC) curve and precision-recall curve). With exception of area under the ROC curve, all were attenuated by skewed distributions, in many cases, dramatically so. While ROC was unaffected by skew, precision-recall curves suggest that ROC may mask poor performance. Our findings suggest that skew is a critical factor in evaluating performance metrics. To avoid or minimize skew-biased estimates of performance, we recommend reporting skew-normalized scores along with the obtained ones. László A. Jeni, Jeffrey F. Cohn, Fernando De la Torre |
ACII | 3 |
| 2013 | Selective Transfer Machine for Personalized Facial Action Unit DetectionabstractAutomatic facial action unit (AFA) detection from video is a long-standing problem in facial expression analysis. Most approaches emphasize choices of features and classifiers. They neglect individual differences in target persons. People vary markedly in facial morphology (e.g., heavy versus delicate brows, smooth versus deeply etched wrinkles) and behavior. Individual differences can dramatically influence how well generic classifiers generalize to previously unseen persons. While a possible solution would be to train person-specific classifiers, that often is neither feasible nor theoretically compelling. The alternative that we propose is to personalize a generic classifier in an unsupervised manner (no additional labels for the test subjects are required). We introduce a transductive learning method, which we refer to Selective Transfer Machine (STM), to personalize a generic classifier by attenuating person-specific biases. STM achieves this effect by simultaneously learning a classifier and re-weighting the training samples that are most relevant to the test subject. To evaluate the effectiveness of STM, we compared STM to generic classifiers and to cross-domain learning methods in three major databases: CK+ [20], GEMEP-FERA [32] and RU-FACS [2]. STM outperformed generic classifiers in all. Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn |
CVPR | 2 |
| 2013 | Supervised Descent Method and Its Applications to Face AlignmentabstractMany computer vision problems (e.g., camera calibration, image alignment, structure from motion) are solved through a nonlinear optimization method. It is generally accepted that 2nd order descent methods are the most robust, fast and reliable approaches for nonlinear optimization of a general smooth function. However, in the context of computer vision, 2nd order descent methods have two main drawbacks: (1) The function might not be analytically differentiable and numerical approximations are impractical. (2) The Hessian might be large and not positive definite. To address these issues, this paper proposes a Supervised Descent Method (SDM) for minimizing a Non-linear Least Squares (NLS) function. During training, the SDM learns a sequence of descent directions that minimizes the mean of NLS functions sampled at different points. In testing, SDM minimizes the NLS objective using the learned descent directions without computing the Jacobian nor the Hessian. We illustrate the benefits of our approach in synthetic and real examples, and show how SDM achieves state-of-the-art performance in the problem of facial feature detection. The code is available at www.humansensing.cs. cmu.edu/intraface. Xuehan Xiong, Fernando De la Torre |
CVPR | 2 |
| 2013 | Deformable Graph MatchingabstractGraph matching (GM) is a fundamental problem in computer science, and it has been successfully applied to many problems in computer vision. Although widely used, existing GM algorithms cannot incorporate global consistence among nodes, which is a natural constraint in computer vision problems. This paper proposes deformable graph matching (DGM), an extension of GM for matching graphs subject to global rigid and non-rigid geometric constraints. The key idea of this work is a new factorization of the pair-wise affinity matrix. This factorization decouples the affinity matrix into the local structure of each graph and the pair-wise affinity edges. Besides the ability to incorporate global geometric transformations, this factorization offers three more benefits. First, there is no need to compute the costly (in space and time) pair-wise affinity matrix. Second, it provides a unified view of many GM methods and extends the standard iterative closest point algorithm. Third, it allows to use the path-following optimization algorithm that leads to improved optimization strategies and matching performance. Experimental results on synthetic and real databases illustrate how DGM outperforms state-of-the-art algorithms for GM. The code is available at http://humansensing.cs.cmu.edu/fgm. Feng Zhou 0002, Fernando De la Torre |
CVPR | 2 |
| 2013 | Robust Principal Component Analysis for Brain Imaging
Petia Georgieva, Fernando De la Torre |
ICANN | 2 |
| 2013 | Unifying Nuclear Norm and Bilinear Factorization Approaches for Low-Rank Matrix DecompositionabstractLow rank models have been widely used for the representation of shape, appearance or motion in computer vision problems. Traditional approaches to fit low rank models make use of an explicit bilinear factorization. These approaches benefit from fast numerical methods for optimization and easy kernelization. However, they suffer from serious local minima problems depending on the loss function and the amount/type of missing data. Recently, these low-rank models have alternatively been formulated as convex problems using the nuclear norm regularizer, unlike factorization methods, their numerical solvers are slow and it is unclear how to kernelize them or to impose a rank a priori. This paper proposes a unified approach to bilinear factorization and nuclear norm regularization, that inherits the benefits of both. We analyze the conditions under which these approaches are equivalent. Moreover, based on this analysis, we propose a new optimization algorithm and a "rank continuation'' strategy that outperform state-of-the-art approaches for Robust PCA, Structure from Motion and Photometric Stereo with outliers and missing data. Ricardo Silveira Cabral, Fernando De la Torre, João Paulo Costeira, Alexandre Bernardino |
ICCV | 2 |
| 2013 | Facial Action Unit Event Detection by Cascade of TasksabstractAutomatic facial Action Unit (AU) detection from video is a long-standing problem in facial expression analysis. AU detection is typically posed as a classification problem between frames or segments of positive examples and negative ones, where existing work emphasizes the use of different features or classifiers. In this paper, we propose a method called Cascade of Tasks (CoT) that combines the use of different tasks (i.e., frame, segment and transition) for AU event detection. We train CoT in a sequential manner embracing diversity, which ensures robustness and generalization to unseen data. In addition to conventional frame-based metrics that evaluate frames independently, we propose a new event-based metric to evaluate detection performance at event-level. We show how the CoT method consistently outperforms state-of-the-art approaches in both frame-based and event-based metrics, across three public datasets that differ in complexity: CK+, FERA and RU-FACS. Wen-Sheng Chu, Fernando De la Torre, Jeffrey F. Cohn |
ICCV | 3 |
| 2013 | Robust Matrix Factorization with Unknown NoiseabstractMany problems in computer vision can be posed as recovering a low-dimensional subspace from high-dimensional visual data. Factorization approaches to low-rank subspace estimation minimize a loss function between the observed measurement matrix and a bilinear factorization. Most popular loss functions include the L1and L2losses. While L1is optimal for Laplacian distributed noise, L2is optimal for Gaussian noise. However, real data is often corrupted by an unknown noise distribution, which is unlikely to be purely Gaussian or Laplacian. To address this problem, this paper proposes a low-rank matrix factorization problem with a Mixture of Gaussians (MoG) noise. The MoG model is a universal approximator for any continuous distribution, and hence is able to model a wider range of real noise distributions. The parameters of the MoG model can be estimated with a maximum likelihood method, while the subspace is computed with standard approaches. We illustrate the benefits of our approach in extensive synthetic, structure from motion, face modeling and background subtraction experiments. Deyu Meng, Fernando De la Torre |
ICCV | 2 |
| 2013 | Learning probability distributions over partially-ordered human everyday activitiesabstractWe propose a method to learn the partially-ordered structure inherent in human everyday activities from observations by exploiting variability in the data. Using statistical relational learning, the system extracts a full-joint probability distribution over the actions that form a task, their (partial) ordering, and their properties. Relevant action properties and relations among actions are learned as those that are consistent among the observations. The models can be used for classifying action sequences, for determining which actions are relevant for a task, which objects are usually manipulated, and which action properties are typical for a person. We evaluate the approach on synthetic data sampled from partial-order trees as well as two real-world data sets of humans activities: the TUM kitchen data set and the CMU MMAC data set. The results show that our approach outperforms sequence-based models like Conditional Random Fields for classifying observations of activities that allow a large amount of variation. Moritz Tenorth, Fernando De la Torre, Michael Beetz |
ICRA | 2 |
| 2013 | Canonical locality preserving Latent Variable Model for discriminative pose inference
Leonid Sigal, Fernando De la Torre, Yonghua Jia |
Image Vis. Comput. | 3 |
| 2013 | Hierarchical Aligned Cluster Analysis for Temporal Clustering of Human MotionabstractTemporal segmentation of human motion into plausible motion primitives is central to understanding and building computational models of human motion. Several issues contribute to the challenge of discovering motion primitives: the exponential nature of all possible movement combinations, the variability in the temporal scale of human actions, and the complexity of representing articulated motion. We pose the problem of learning motion primitives as one of temporal clustering, and derive an unsupervised hierarchical bottom-up framework called hierarchical aligned cluster analysis (HACA). HACA finds a partition of a given multidimensional time series into m disjoint segments such that each segment belongs to one of k clusters. HACA combines kernel k-means with the generalized dynamic time alignment kernel to cluster time series data. Moreover, it provides a natural framework to find a low-dimensional embedding for time series. HACA is efficiently optimized with a coordinate descent strategy and dynamic programming. Experimental results on motion capture and video data demonstrate the effectiveness of HACA for segmenting complex motions and as a visualization tool. We also compare the performance of HACA to state-of-the-art algorithms for temporal clustering on data of a honey bee dance. The HACA code is available online. Feng Zhou 0002, Fernando De la Torre, Jessica K. Hodgins |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Fixed-rank representation for unsupervised visual learningabstractSubspace clustering and feature extraction are two of the most commonly used unsupervised learning techniques in computer vision and pattern recognition. State-of-the-art techniques for subspace clustering make use of recent advances in sparsity and rank minimization. However, existing techniques are computationally expensive and may result in degenerate solutions that degrade clustering performance in the case of insufficient data sampling. To partially solve these problems, and inspired by existing work on matrix factorization, this paper proposes fixed-rank representation (FRR) as a unified framework for unsupervised visual learning. FRR is able to reveal the structure of multiple subspaces in closed-form when the data is noiseless. Furthermore, we prove that under some suitable conditions, even with insufficient observations, FRR can still reveal the true subspace memberships. To achieve robustness to outliers and noise, a sparse regularizer is introduced into the FRR framework. Beyond subspace clustering, FRR can be used for unsupervised feature extraction. As a non-trivial byproduct, a fast numerical solver is developed for FRR. Experimental results on both synthetic data and real applications validate our theoretical analysis and demonstrate the benefits of FRR for unsupervised visual learning. Risheng Liu, Zhouchen Lin, Fernando De la Torre, Zhixun Su |
CVPR | 3 |
| 2012 | Max-margin early event detectorsabstractThe need for early detection of temporal events from sequential data arises in a wide spectrum of applications ranging from human-robot interaction to video security. While temporal event detection has been extensively studied, early detection is a relatively unexplored problem. This paper proposes a maximum-margin framework for training temporal event detectors to recognize partial events, enabling early detection. Our method is based on Structured Output SVM, but extends it to accommodate sequential data. Experiments on datasets of varying complexity, for detecting facial expressions, hand gestures, and human activities, demonstrate the benefits of our approach. To the best of our knowledge, this is the first paper in the literature of computer vision that proposes a learning formulation for early event detection. Minh Hoai, Fernando De la Torre |
CVPR | 2 |
| 2012 | Factorized graph matchingabstractGraph matching plays a central role in solving correspondence problems in computer vision. Graph matching problems that incorporate pair-wise constraints can be cast as a quadratic assignment problem (QAP). Unfortunately, QAP is NP-hard and many algorithms have been proposed to solve different relaxations. This paper presents factorized graph matching (FGM), a novel framework for interpreting and optimizing graph matching problems. In this work we show that the affinity matrix can be factorized as a Kronecker product of smaller matrices. There are three main benefits of using this factorization in graph matching: (1) There is no need to compute the costly (in space and time) pair-wise affinity matrix; (2) The factorization provides a taxonomy for graph matching and reveals the connection among several methods; (3) Using the factorization we derive a new approximation of the original problem that improves state-of-the-art algorithms in graph matching. Experimental results in synthetic and real databases illustrate the benefits of FGM. The code is available at http://humansensing.cs.cmu.edu/fgm. Feng Zhou 0002, Fernando De la Torre |
CVPR | 2 |
| 2012 | Generalized time warping for multi-modal alignment of human motionabstractTemporal alignment of human motion has been a topic of recent interest due to its applications in animation, telerehabilitation and activity recognition among others. This paper presents generalized time warping (GTW), an extension of dynamic time warping (DTW) for temporally aligning multi-modal sequences from multiple subjects performing similar activities. GTW solves three major drawbacks of existing approaches based on DTW: (1) GTW provides a feature weighting layer to adapt different modalities (e.g., video and motion capture data), (2) GTW extends DTW by allowing a more flexible time warping as combination of monotonic functions, (3) unlike DTW that typically incurs in quadratic cost, GTW has linear complexity. Experimental results demonstrate that GTW can efficiently solve the multi-modal temporal alignment problem and outperforms state-of-the-art DTW methods for temporal alignment of time series within the same modality. Feng Zhou 0002, Fernando De la Torre |
CVPR | 2 |
| 2012 | Unsupervised Temporal Commonality Discovery
Wen-Sheng Chu, Feng Zhou 0002, Fernando De la Torre |
ECCV (4) | 3 |
| 2012 | Robust Regression
Dong Huang 0007, Ricardo Silveira Cabral, Fernando De la Torre |
ECCV (4) | 3 |
| 2012 | Facial Action Transfer with Personalized Bilinear Regression
Dong Huang 0007, Fernando De la Torre |
ECCV (2) | 2 |
| 2012 | Continuous Regression for Non-rigid Image Alignment
Enrique Sánchez-Lozano, Fernando De la Torre, Daniel González-Jiménez |
ECCV (7) | 2 |
| 2012 | Multimodal feature analysis for quantitative performance evaluation of endotracheal intubation (ETI)abstractEndotracheal intubation (ETI) is a crucial medical procedure performed on critically ill patients. It involves insertion of a breathing tube into the trachea i.e. the windpipe connecting the larynx and the lungs. Often, this procedure is performed by the paramedics (aka providers) under challenging prehospital settings e.g. roadside, ambulances or helicopters. Successful intubations could be lifesaving, whereas, failed intubation could potentially be fatal. Under prehospital environments, ETI success rates among the paramedics are surprisingly low and this necessitates better training and performance evaluation of ETI skills. Currently, few objective metrics exist to quantify the differences in ETI techniques between providers. In this pilot study, we develop a quantitative framework for discriminating the kinematic characteristics of providers with different experience levels. The system utilizes statistical analysis on spatio-temporal multimodal features extracted from optical motion capture, accelerometers and electromyography (EMG) sensors. Our experiments involved three individuals performing intubations on a dummy, each with different levels of training. Quantitative performance analysis on multimodal features revealed distinctive differences among different skill levels. In future work, the feedback from these analysis could potentially be harnessed for enhanced ETI training. Samarjit Das, Jestin N. Carlson, Fernando De la Torre, Paul E. Phrampus, Jessica K. Hodgins |
ICASSP | 3 |
| 2012 | A Least-Squares Framework for Component AnalysisabstractOver the last century, Component Analysis (CA) methods such as Principal Component Analysis (PCA), Linear Discriminant Analysis (LDA), Canonical Correlation Analysis (CCA), Locality Preserving Projections (LPP), and Spectral Clustering (SC) have been extensively used as a feature extraction step for modeling, classification, visualization, and clustering. CA techniques are appealing because many can be formulated as eigen-problems, offering great potential for learning linear and nonlinear representations of data in closed-form. However, the eigen-formulation often conceals important analytic and computational drawbacks of CA techniques, such as solving generalized eigen-problems with rank deficient matrices (e.g., small sample size problem), lacking intuitive interpretation of normalization factors, and understanding commonalities and differences between CA methods. This paper proposes a unified least-squares framework to formulate many CA methods. We show how PCA, LDA, CCA, LPP, SC, and its kernel and regularized extensions correspond to a particular instance of least-squares weighted kernel reduced rank regression (LS--WKRRR). The LS-WKRRR formulation of CA methods has several benefits: 1) provides a clean connection between many CA techniques and an intuitive framework to understand normalization factors; 2) yields efficient numerical schemes to solve CA techniques; 3) overcomes the small sample size problem; 4) provides a framework to easily extend CA methods. We derive weighted generalizations of PCA, LDA, SC, and CCA, and several new CA techniques. Fernando De la Torre |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Fast-FACS: A Computer-Assisted System to Increase Speed and Reliability of Manual FACS Coding
Fernando De la Torre, Tomas Simon, Zara Ambadar, Jeffrey F. Cohn |
ACII (1) | 1 |
| 2011 | Supervised local subspace learning for continuous head pose estimationabstractHead pose estimation from images has recently attracted much attention in computer vision due to its diverse applications in face recognition, driver monitoring and human computer interaction. Most successful approaches to head pose estimation formulate the problem as a nonlinear regression between image features and continuous 3D angles (i.e. yaw, pitch and roll). However, regression-like methods suffer from three main drawbacks: (1) They typically lack generalization and overfit when trained using a few samples. (2) They fail to get reliable estimates over some regions of the output space (angles) when the training set is not uniformly sampled. For instance, if the training data contains under-sampled areas for some angles. (3) They are not robust to image noise or occlusion. To address these problems, this paper presents Supervised Local Subspace Learning (SL2), a method that learns a local linear model from a sparse and non-uniformly sampled training set. SL2learns a mixture of local tangent spaces that is robust to under-sampled regions, and due to its regularization properties it is also robust to over-fitting. Moreover, because SL2is a generative model, it can deal with image noise. Experimental results on the CMU Multi-PIE and BU-3DFE database show the effectiveness of our approach in terms of accuracy and computational complexity. Dong Huang 0007, Markus Storer, Fernando De la Torre, Horst Bischof |
CVPR | 3 |
| 2011 | Local isomorphism to solve the pre-image problem in kernel methodsabstractKernel methods have been popular over the last decade to solve many computer vision, statistics and machine learning problems. An important, both theoretically and practically, open problem in kernel methods is the pre-image problem. The pre-image problem consists of finding a vector in the input space whose mapping is known in the feature space induced by a kernel. To solve the pre-image problem, this paper proposes a framework that computes an isomorphism between local Gram matrices in the input and feature space. Unlike existing methods that rely on analytic properties of kernels, our framework derives closed-form solutions to the pre-image problem in the case of non-differentiable and application-specific kernels. Experiments on the pre-image problem for visualizing cluster centers computed by kernel k-means and denoising high-dimensional images show that our algorithm outperforms state-of-the-art methods. Dong Huang 0007, Yuandong Tian, Fernando De la Torre |
CVPR | 3 |
| 2011 | Joint segmentation and classification of human actions in videoabstractAutomatic video segmentation and action recognition has been a long-standing problem in computer vision. Much work in the literature treats video segmentation and action recognition as two independent problems; while segmentation is often done without a temporal model of the activity, action recognition is usually performed on pre-segmented clips. In this paper we propose a novel method that avoids the limitations of the above approaches by jointly performing video segmentation and action recognition. Unlike standard approaches based on extensions of dynamic Bayesian networks, our method is based on a discriminative temporal extension of the spatial bag-of-words model that has been very popular in object recognition. The classification is performed robustly within a multi-class SVM framework whereas the inference over the segments is done efficiently with dynamic programming. Experimental results on honeybee, Weizmann, and Hollywood datasets illustrate the benefits of our approach compared to state-of-the-art methods. Minh Hoai, Zhen-Zhong Lan, Fernando De la Torre |
CVPR | 3 |
| 2011 | Active conditional modelsabstractMatching images with large geometric and iconic changes (e.g. faces under different poses and facial expressions) is an open research problem in computer vision. There are two fundamental approaches to solve the correspondence problem in images: Feature-based matching and model-based matching. Feature-based matching relies on the assumption that features are stable across view-points and iconic changes, and it uses some unary, pair-wise or higher-order constraints as a measure of correspondence. On the other hand, model-based approaches such as Active Shape Models (ASMs) align appearance features with respect to a model. The model is learned from hand-labeled samples. However, model-based approaches typically suffer from lack of generalization to untrained situations. This paper proposes Active Conditional Models (ACM) that combines the benefits of both approaches. ACM learns the conditional relation (both in shape and appearance) between a reference view of the object and other view-points or iconic changes. The ACM model generalizes better to untrained situations, because it has less number of parameters (less prone to overfitting) and directly learns variations w.r.t a reference image (similar to feature-based methods). Several examples in the context of facial feature matching across pose and expression illustrate the benefits of ACMs. Fernando De la Torre |
FG | 2 |
| 2011 | Hierarchical CRF with product label spaces for parts-based modelsabstractNon-rigid object detection is a challenging open research problem in computer vision. It is a critical part in many applications such as image search, surveillance, human-computer interaction or image auto-annotation. Most successful approaches to non-rigid object detection make use of part-based models. In particular, Conditional Random Fields (CRF) have been successfully embedded into a discriminative parts-based model framework due to its effectiveness for learning and inference (usually based on a tree structure). However, CRF-based approaches do not incorporate global constraints and only model pairwise interactions. This is especially important when modeling object classes that may have complex parts interactions (e.g. facial features or body articulations), because neglecting them yields an oversimplified model with suboptimal performance. To overcome this limitation, this paper proposes a novel hierarchical CRF (HCRF). The main contribution is to build a hierarchy of part combinations by extending the label set to a hierarchy of product label spaces. In order to keep the inference computation tractable, we propose an effective method to reduce the new label set. We test our method on two applications: facial feature detection on the Multi-PIE database and human pose estimation on the Buffy dataset. Gemma Roig, Xavier Boix, Fernando De la Torre, Joan Serrat 0002, Carles Vilella |
FG | 3 |
| 2011 | Source constrained clusteringabstractWe consider the problem of quantizing data generated from disparate sources, e.g. subjects performing actions with different styles, movies with particular genre bias, various conditions in which images of objects are taken, etc. These are scenarios where unsupervised clustering produces inadequate codebooks because algorithms like K-means tend to cluster samples based on data biases (e.g. cluster subjects), rather than cluster similar samples across sources (e.g. cluster actions). We propose a new quantization technique, Source Constrained Clustering (SCC), which extends the K-means algorithm by enforcing clusters to group samples from multiple sources. We evaluate the method in the context of activity recognition from videos in an unconstrained environment. Experiments on several tasks and features show that using source information improves classification performance. Ekaterina H. Taralova, Fernando De la Torre, Martial Hebert |
ICCV | 2 |
| 2011 | Fast incremental method for matrix completion: An application to trajectory correctionabstractWe address the problem of incrementally recovering a matrix of tracked image points, based on partial observations of their trajectories. Besides partial observability, we assume the existence of gross, but sparse, noise on the known entries. This problem has obvious applications in real-time tracking and structure from motion, where observations are plagued by self-occlusion and outliers. Recently, work in the optimization community has spun optimal methods for matrix completion when this matrix is known to be low rank by minimizing the nuclear norm, the sum of its singular values. Despite exhibiting several optimality properties, no available algorithms perform this minimization incrementally. In this paper, we build upon the Nuclear Norm Robust PCA method and SPectrally Optimal Completion to propose a fast and incremental algorithm which is able to cope with outliers. We present experiments showing the competitive speed of our method while maintaining performance comparable to the state-of-the-art. Ricardo Silveira Cabral, João Paulo Costeira, Fernando De la Torre, Alexandre Bernardino |
ICIP | 3 |
| 2011 | Matrix Completion for Multi-label Image ClassificationabstractRecently, image categorization has been an active research topic due to the urgent need to retrieve and browse digital images via semantic keywords. This paper formulates image categorization as a multi-label classification problem using recent advances in matrix completion. Under this setting, classification of testing data is posed as a problem of completing unknown label entries on a data matrix that concatenates training and testing features with training labels. We propose two convex algorithms for matrix completion based on a Rank Minimization criterion specifically tailored to visual data, and prove its convergence properties. A major advantage of our approach w.r.t. standard discriminative classification methods for image categorization is its robustness to outliers, background noise and partial occlusions both in the feature and label space. Experimental validation on several datasets shows how our method outperforms state-of-the-art algorithms, while effectively capturing semantic concepts of classes. Ricardo Silveira Cabral, Fernando De la Torre, João Paulo Costeira, Alexandre Bernardino |
NIPS | 2 |
| 2011 | Face recognition using Histograms of Oriented Gradients
Oscar Déniz-Suárez, Gloria Bueno García, Jesús Salido, Fernando De la Torre |
Pattern Recognit. Lett. | 4 |
| 2011 | Fast and Robust Circular Object Detection With Probabilistic Pairwise VotingabstractAccurate and efficient detection of circular objects in images is a challenging computer vision problem. Existing circular object detection methods can be broadly classified into two categories: voting based and maximum likelihood estimation (MLE) based. The former is robust to noise, however its computational complexity and memory requirement are high. On the other hand, MLE based methods (e.g., robust least squares fitting) are more computationally efficient but sensitive to noise, and can not detect multiple circles. This letter proposes Probabilistic Pairwise Voting (PPV), a fast and robust algorithm for circular object detection based on an extension of Hough Transform. The main contributions are threefold. 1) We formulate the problem of circular object detection as finding the intersection of lines in the three dimensional parameter space (i.e., center and radius of the circle). 2) We propose a probabilistic pairwise voting scheme to robustly discover circular objects under occlusion, image noise and moderate shape deformations. 3) We use a mode-finding algorithm to efficiently find multiple circular objects. We demonstrate the benefits of our approach on two real-world problems: 1) detecting circular objects in natural images, and 2) localizing iris in face images. Lili Pan 0001, Wen-Sheng Chu, Jason M. Saragih, Fernando De la Torre, Mei Xie |
IEEE Signal Process. Lett. | 4 |
| 2011 | Dynamic Cascades with Bidirectional Bootstrapping for Action Unit Detection in Spontaneous Facial BehaviorabstractAutomatic facial action unit detection from video is a long-standing problem in facial expression analysis. Research has focused on registration, choice of features, and classifiers. A relatively neglected problem is the choice of training images. Nearly all previous work uses one or the other of two standard approaches. One approach assigns peak frames to the positive class and frames associated with other actions to the negative class. This approach maximizes differences between positive and negative classes, but results in a large imbalance between them, especially for infrequent AUs. The other approach reduces imbalance in class membership by including all target frames from onsets to offsets in the positive class. However, because frames near onsets and offsets often differ little from those that precede them, this approach can dramatically increase false positives. We propose a novel alternative, dynamic cascades with bidirectional bootstrapping (DCBB), to select training samples. Using an iterative approach, DCBB optimally selects positive and negative samples in the training data. Using Cascade Adaboost as basic classifier, DCBB exploits the advantages of feature selection, efficiency, and robustness of Cascade Adaboost. To provide a real-world test, we used the RU-FACS (a.k.a. M3) database of nonposed behavior recorded during interviews. For most tested action units, DCBB improved AU detection relative to alternative approaches. Yunfeng Zhu, Fernando De la Torre, Jeffrey F. Cohn, Yu-Jin Zhang |
IEEE Trans. Affect. Comput. | 2 |
| 2011 | Interactive region-based linear 3D face modelsabstractLinear models, particularly those based on principal component analysis (PCA), have been used successfully on a broad range of human face-related applications. Although PCA models achieve high compression, they have not been widely used for animation in a production environment because their bases lack a semantic interpretation. Their parameters are not an intuitive set for animators to work with. In this paper we present a linear face modelling approach that generalises to unseen data better than the traditional holistic approach while also allowing click-and-drag interaction for animation. Our model is composed of a collection of PCA sub-models that are independently trained but share boundaries. Boundary consistency and user-given constraints are enforced in a soft least mean squares sense to give flexibility to the model while maintaining coherence. Our results show that the region-based model generalises better than its holistic counterpart when describing previously unseen motion capture data from multiple subjects. The decomposition of the face into several regions, which we determine automatically from training data, gives the user localised manipulation control. This feature allows to use the model for face posing and animation in an intuitive style. J. Rafael Tena, Fernando De la Torre, Iain A. Matthews |
ACM Trans. Graph. | 2 |
| 2010 | Latent Gaussian Mixture Regression for Human Pose Estimation
Leonid Sigal, Hernán Badino, Fernando De la Torre, Yong Liu 0027 |
ACCV (3) | 4 |
| 2010 | Pareto discriminant analysisabstractLinear Discriminant Analysis (LDA) is a popular tool for multiclass discriminative dimensionality reduction. However, LDA suffers from two major problems: (1) It only optimizes the Bayes error for the case of unimodal Gaussian classes with equal covariances (assuming full rank matrices) and, (2) The multiclass extension maximizes the sum of pairwise distances between the classes, and does not “simultaneously” maximize each pairwise distance between the classes. This typically results in serious overlapping in the projected space between classes that are “close” in the input space. To solve these two problems, this paper proposes Pareto Discriminant Analysis (PARDA). Firstly, PARDA explicitly models each of the classes as a multidimensional Gaussian with a sample covariance. Secondly, PARDA decomposes the multiclass problem to a set of pairwise objective functions representing the pairwise distance between different classes. Unlike existing extensions of Fisher discriminant analysis (FDA) to multiclass problems, that typically maximize the sum of pairwise distances between classes, PARDA simultaneously maximizes each pairwise distance, thus encouraging the case that all classes are equidistant from each other in the lower dimensional space. Solving PARDA is a multiobjective optimization problem - simultaneously optimizing more than one, possibly conflicting, objective functions - and the resulting solution is known to be “Pareto Optimal”. Experimental results on synthetic data, several image data sets and data sets from the UCI repository show positive and encouraging results in favor of PARDA when compared with standard and state-of-the-art multiclass extensions of LDA. Karim T. Abou-Moustafa, Fernando De la Torre, Frank P. Ferrie |
CVPR | 2 |
| 2010 | Action unit detection with segment-based SVMsabstractAutomatic facial action unit (AU) detection from video is a long-standing problem in computer vision. Two main approaches have been pursued: (1) static modeling - typically posed as a discriminative classification problem in which each video frame is evaluated independently; (2) temporal modeling - frames are segmented into sequences and typically modeled with a variant of dynamic Bayesian networks. We propose a segment-based approach, kSeg-SVM, that incorporates benefits of both approaches and avoids their limitations. kSeg-SVM is a temporal extension of the spatial bag-of-words. kSeg-SVM is trained within a structured output SVM framework that formulates AU detection as a problem of detecting temporal events in a time series of visual features. Each segment is modeled by a variant of the BoW representation with soft assignment of the words based on similarity. Our framework has several benefits for AU detection: (1) both dependencies between features and the length of action units are modeled; (2) all possible segments of the video may be used for training; and (3) no assumptions are required about the underlying structure of the action unit events (e.g., i.i.d.). Our algorithm finds the best k-or-fewer segments that maximize the SVM score. Experimental results suggest that the proposed method outperforms state-of-the-art static methods for AU detection. Tomas Simon, Minh Hoai, Fernando De la Torre, Jeffrey F. Cohn |
CVPR | 3 |
| 2010 | Unsupervised discovery of facial eventsabstractAutomatic facial image analysis has been a long standing research problem in computer vision. A key component in facial image analysis, largely conditioning the success of subsequent algorithms (e.g. facial expression recognition), is to define a vocabulary of possible dynamic facial events. To date, that vocabulary has come from the anatomically-based Facial Action Coding System (FACS) or more subjective approaches (i.e. emotion-specified expressions). The aim of this paper is to discover facial events directly from video of naturally occurring facial behavior, without recourse to FACS or other labeling schemes. To discover facial events, we propose a temporal clustering algorithm, Aligned Cluster Analysis (ACA), and a multi-subject correspondence algorithm for matching expressions. We use a variety of video sources: posed facial behavior (Cohn-Kanade database), unscripted facial behavior (RU-FACS database) and some video in infants. Accuracy of (unsupervised) ACA approached that of a supervised version, achieved moderate intersystem agreement with FACS, and proved informative as a visualization/summarization tool. Feng Zhou 0002, Fernando De la Torre, Jeffrey F. Cohn |
CVPR | 2 |
| 2010 | Bilinear Kernel Reduced Rank Regression for Facial Expression Synthesis
Dong Huang 0007, Fernando De la Torre |
ECCV (2) | 2 |
| 2010 | Local Minima Embedding
Minyoung Kim 0001, Fernando De la Torre |
ICML | 2 |
| 2010 | Gaussian Processes Multiple Instance Learning
Minyoung Kim 0001, Fernando De la Torre |
ICML | 2 |
| 2010 | Unsupervised summarization of rushes videosabstractThis paper proposes a new framework to formulate summarization of rushes video as an unsupervised learning problem. We pose the problem of video summarization as one of time-series clustering, and proposed Constrained Aligned Cluster Analysis (CACA). CACA combines kernel k-means, Dynamic Time Alignment Kernel (DTAK), and unlike previous work, CACA jointly optimizes video segmentation and shot clustering. CACA is effciently solved via dynamic programming. Experimental results on the TRECVID 2007 and 2008 BBC rushes video summarization databases validate the accuracy and effectiveness of CACA. Yang Liu 0007, Feng Zhou 0002, Wei Liu 0220, Fernando De la Torre, Yan Liu 0004 |
ACM Multimedia | 4 |
| 2010 | Metric Learning for Image Alignment
Minh Hoai, Fernando De la Torre |
Int. J. Comput. Vis. | 2 |
| 2010 | Learning a generic 3D face model from 2D image databases using incremental Structure-from-Motion
Jose Gonzalez-Mora, Fernando De la Torre, Nicolás Guil, Emilio L. Zapata |
Image Vis. Comput. | 2 |
| 2010 | Optimal feature selection for support vector machines
Minh Hoai, Fernando De la Torre |
Pattern Recognit. | 2 |
| 2009 | Efficient image alignment using linear appearance modelsabstractVisual tracking is a key component in many computer vision applications. Linear subspace techniques (e.g. eigen-tracking) are one of the most popular approaches to align templates with appearance variations (e.g. illumination, iconic changes). A number of well known tracking algorithms have been proposed in the last years to accurately fit these models to images. Computational efficiency is an important limitation in object tracking algorithms and different efficient techniques, such as the “projected-out” optimization, have been proposed. They reduce the computational cost using an efficient formulation in which many of the involved operations can be precomputed. On the other hand, alternative “simultaneous” algorithms jointly optimize pose and appearance parameters, providing better performance but increasing the computational cost. In this paper, we propose an algorithm for efficient linear appearance model fitting based on the inverse compositional simultaneous optimization of pose and appearance. We introduce a novel formulation which reduces the required computational time while maintaining similar convergence properties of previous “simultaneous” approaches. Experimental results illustrate the capabilities of this algorithm in face tracking. Jose Gonzalez-Mora, Nicolás Guil, Emilio L. Zapata, Fernando De la Torre |
CVPR | 4 |
| 2009 | Canonical Time Warping for Alignment of Human BehaviorabstractAlignment of time series is an important problem to solve in many scientific disciplines. In particular, temporal alignment of two or more subjects performing similar activities is a challenging problem due to the large temporal scale difference between human actions as well as the inter/intra subject variability. In this paper we present canonical time warping (CTW), an extension of canonical correlation analysis (CCA) for spatio-temporal alignment of the behavior between two subjects. CTW extends previous work on CCA in two ways: (i) it combines CCA with dynamic time warping for temporal alignment; and (ii) it extends CCA to allow local spatial deformations. We show CTWs effectiveness in three experiments: alignment of synthetic data, alignment of motion capture data of two subjects performing similar actions, and alignment of two people with similar facial expressions. Our results demonstrate that CTW provides both visually and qualitatively better alignment than state-of-the-art techniques based on dynamic time warping. Feng Zhou 0002, Fernando De la Torre |
NIPS | 2 |
| 2009 | Emphatic Visual Speech SynthesisabstractThe synthesis of talking heads has been a flourishing research area over the last few years. Since human beings have an uncanny ability to read people's faces, most related applications (e.g., advertising, video-teleconferencing) require absolutely realistic photometric and behavioral synthesis of faces. This paper proposes a person-specific facial synthesis framework that allows high realism and includes a novel way to control visual emphasis (e.g., level of exaggeration of visible articulatory movements of the vocal tract). There are three main contributions: a geodesic interpolation with visual unit selection, a parameterization of visual emphasis, and the design of minimum size corpora. Perceptual tests with human subjects reveal high realism properties, achieving similar perceptual scores as real samples. Furthermore, the visual emphasis level and two communication styles show a statistical interaction relationship. Javier Melenchón, Elisa Martínez Marroquín, Fernando De la Torre, José Antonio Montero |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | Model-Based De-Identification of Facial Images
Ralph Gross, Latanya Sweeney, Jeffrey F. Cohn, Fernando De la Torre, Simon Baker |
AMIA | 4 |
| 2008 | Semi-supervised learning of multi-factor models for face de-identificationabstractWith the emergence of new applications centered around the sharing of image data, questions concerning the protection of the privacy of people visible in the scene arise. Recently, formal methods for the de-identification of images have been proposed which would benefit from multi-factor coding to separate identity and non-identity related factors. However, existing multi-factor models require complete labels during training which are often not available in practice. In this paper we propose a new multi-factor framework which unifies linear, bilinear, and quadratic models. We describe a new fitting algorithm which jointly estimates all model parameters and show that it outperforms the standard alternating algorithm. We furthermore describe how to avoid overfitting the model and how to train the model in a semi-supervised manner. In experiments on a large expression-variant face database we show that data coded using our multi-factor model leads to improved data utility while providing the same privacy protection. Ralph Gross, Latanya Sweeney, Fernando De la Torre, Simon Baker |
CVPR | 3 |
| 2008 | Local minima free Parameterized Appearance ModelsabstractParameterized Appearance Models (PAMs) (e.g. Eigentracking, Active Appearance Models, Morphable Models) are commonly used to model the appearance and shape variation of objects in images. While PAMs have numerous advantages relative to alternate approaches, they have at least two drawbacks. First, they are especially prone to local minima in the fitting process. Second, often few if any of the local minima of the cost function correspond to acceptable solutions. To solve these problems, this paper proposes a method to learn a cost function by explicitly optimizing that the local minima occur at and only at the places corresponding to the correct fitting parameters. To the best of our knowledge, this is the first paper to address the problem of learning a cost function to explicitly model local properties of the error surface to fit PAMs. Synthetic and real examples show improvement in alignment performance in comparison with traditional approaches. Minh Hoai, Fernando De la Torre |
CVPR | 2 |
| 2008 | Parameterized Kernel Principal Component Analysis: Theory and applications to supervised and unsupervised image alignmentabstractParameterized Appearance Models (PAMs) (e.g. eigen-tracking, active appearance models, morphable models) use Principal Component Analysis (PCA) to model the shape and appearance of objects in images. Given a new image with an unknown appearance/shape configuration, PAMs can detect and track the object by optimizing the model’s parameters that best match the image. While PAMs have numerous advantages for image registration relative to alternative approaches, they suffer from two major limitations: First, PCA cannot model non-linear structure in the data. Second, learning PAMs requires precise manually labeled training data. This paper proposes Parameterized Kernel Principal Component Analysis (PKPCA), an extension of PAMs that uses Kernel PCA (KPCA) for learning a non-linear appearance model invariant to rigid and/or non-rigid deformations. We demonstrate improved performance in supervised and unsupervised image registration, and present a novel application to improve the quality of manual landmarks in faces. In addition, we suggest a clean and effective matrix formulation for PKPCA. Fernando De la Torre, Minh Hoai |
CVPR | 1 |
| 2008 | Facial feature detection with optimal pixel reduction SVMabstractAutomatic facial feature localization has been a long-standing challenge in the field of computer vision for several decades. This can be explained by the large variation a face in an image can have due to factors such as position, facial expression, pose, illumination, and background clutter. Support Vector Machines (SVMs) have been a popular statistical tool for facial feature detection. Traditional SVM approaches to facial feature detection typically extract features from images (e.g. multiband filter, SIFT features) and learn the SVM parameters. Independently learning features and SVM parameters might result in a loss of information related to the classification process. This paper proposes an energy-based framework to jointly perform relevant feature weighting and SVM parameter learning. Preliminary experiments on standard face databases have shown significant improvement in speed with our approach. Minh Hoai, Joan Perez, Fernando De la Torre |
FG | 3 |
| 2008 | Learning image alignment without local minima for face detection and trackingabstractActive appearance models (AAMs) have been extensively used for face alignment during the last 20 years. While AAMs have numerous advantages relative to alternate approaches, they suffer from two major drawbacks: (i) AAMs are especially prone to local minima in the fitting process; (ii) few if any of the local minima of the cost function correspond to acceptable solutions. To minimize these problems, this paper proposes a method to learn the fitting cost function that explicitly optimizes that the local minima occur at and only at the places corresponding to the correct fitting parameters. The paper explores two methods to parameterize the cost function: pixel weighting and subspace learning. Experiments on synthetic and real data show the effectiveness of our approach for face alignment. Minh Hoai, Fernando De la Torre |
FG | 2 |
| 2008 | Aligned Cluster Analysis for temporal segmentation of human motionabstractTemporal segmentation of human motion into actions is a crucial step for understanding and building computational models of human motion. Several issues contribute to the challenge of this task. These include the large variability in the temporal scale and periodicity of human actions, as well as the exponential nature of all possible movement combinations. We formulate the temporal segmentation problem as an extension of standard clustering algorithms. In particular, this paper proposes aligned cluster analysis (ACA), a robust method to temporally segment streams of motion capture data into actions. ACA extends standard kernel k-means clustering in two ways: (1) the cluster means contain a variable number of features, and (2) a dynamic time warping (DTW) kernel is used to achieve temporal invariance. Experimental results, reported on synthetic data and the Carnegie Mellon Motion Capture database, demonstrate its effectiveness. Feng Zhou 0002, Fernando De la Torre, Jessica K. Hodgins |
FG | 2 |
| 2008 | Robust Kernel Principal Component AnalysisabstractKernel Principal Component Analysis (KPCA) is a popular generalization of linear PCA that allows non-linear feature extraction. In KPCA, data in the input space is mapped to higher (usually) dimensional feature space where the data can be linearly modeled. The feature space is typically induced implicitly by a kernel function, and linear PCA in the feature space is performed via the kernel trick. However, due to the implicitness of the feature space, some extensions of PCA such as robust PCA cannot be directly generalized to KPCA. This paper presents a technique to overcome this problem, and extends it to a unified framework for treating noise, missing data, and outliers in KPCA. Our method is based on a novel cost function to perform inference in KPCA. Extensive experiments, in both synthetic and real data, show that our algorithm outperforms existing methods. Minh Hoai, Fernando De la Torre |
NIPS | 2 |
| 2008 | Image-based ShavingabstractAbstract Many categories of objects, such as human faces, can be naturally viewed as a composition of several different layers. For example, a bearded face with glasses can be decomposed into three layers: a layer for glasses, a layer for the beard and a layer for other permanent facial features. While modeling such a face with a linear subspace model could be very difficult, layer separation allows for easy modeling and modification of some certain structures while leaving others unchanged. In this paper, we present a method for automatic layer extraction and its applications to face synthesis and editing. Layers are automatically extracted by utilizing the differences between subspaces and modeled separately. We show that our method can be used for tasks such beard removal (virtual shaving), beard synthesis, and beard transfer, among others. Minh Hoai, Jean-François Lalonde, Alexei A. Efros, Fernando De la Torre |
Comput. Graph. Forum | 4 |
| 2007 | Filtered Component Analysis to Increase Robustness to Local Minima in Appearance ModelsabstractAppearance models (AM) are commonly used to model appearance and shape variation of objects in images. In particular, they have proven useful to detection, tracking, and synthesis of people's faces from video. While AM have numerous advantages relative to alternative approaches, they have at least two important drawbacks. First, they are especially prone to local minima in fitting; this problem becomes increasingly problematic as the number of parameters to estimate grows. Second, often few if any of the local minima correspond to the correct location of the model error. To address these problems, we propose filtered component analysis (FCA), an extension of traditional principal component analysis (PCA). FCA learns an optimal set of filters with which to build a multi-band representation of the object. FCA representations were found to be more robust than either grayscale or Gabor filters to problems of local minima. The effectiveness and robustness of the proposed algorithm is demonstrated in both synthetic and real data. Fernando De la Torre, Alvaro Collet, Manuel Quero, Jeffrey F. Cohn, Takeo Kanade |
CVPR | 1 |
| 2007 | Learning Kernel Expansions for Image ClassificationabstractKernel machines (e.g. SVM, KLDA) have shown state-of-the-art performance in several visual classification tasks. The classification performance of kernel machines greatly depends on the choice of kernels and its parameters. In this paper, we propose a method to search over a space of parameterized kernels using a gradient-descent based method. Our method effectively learns a non-linear representation of the data useful for classification and simultaneously performs dimensionality reduction. In addition, we suggest a new matrix formulation that simplifies and unifies previous approaches. The effectiveness and robustness of the proposed algorithm is demonstrated in both synthetic and real examples of pedestrian and mouth detection in images. Fernando De la Torre, Oriol Vinyals |
CVPR | 1 |
| 2007 | Bilinear Active Appearance ModelsabstractAppearance Models have been applied to model the space of human faces over the last two decades. In particular, Active Appearance Models (AAMs) have been successfully used for face tracking, synthesis and recognition, and they are one of the state-of-the-art approaches due to its efficiency and representational power. Although widely employed, AAMs suffer from a few drawbacks, such as the inability to isolate pose, identity and expression changes. This paper proposes Bilinear Active Appearance Models (BAAMs), an extension of AAMs, that effectively decouple changes due to pose and expression/identity. We derive a gradient-descent algorithm to efficiently fit BAAMs to new images. Experimental results show how BAAMs improve generalization and convergence with respect to the linear model. In addition, we illustrate decoupling benefits of BAAMs in face recognition across pose. We show how the pose normalization provided by BAAMs increase the recognition performance of commercial systems. Jose Gonzalez-Mora, Fernando De la Torre, Rajesh Murthi, Nicolás Guil, Emilio L. Zapata |
ICCV | 2 |
| 2007 | Temporal Segmentation of Facial BehaviorabstractTemporal segmentation of facial gestures in spontaneous facial behavior recorded in real-world settings is an important, unsolved, and relatively unexplored problem in facial image analysis. Several issues contribute to the challenge of this task. These include non-frontal pose, moderate to large out-of-plane head motion, large variability in the temporal scale of facial gestures, and the exponential nature of possible facial action combinations. To address these challenges, we propose a two-step approach to temporally segment facial behavior. The first step uses spectral graph techniques to cluster shape and appearance features invariant to some geometric transformations. The second step groups the clusters into temporally coherent facial gestures. We evaluated this method in facial behavior recorded during face-to- face interactions. The video data were originally collected to answer substantive questions in psychology without concern for algorithm development. The method achieved moderate convergent validity with manual FACS (Facial Action Coding System) annotation. Further, when used to preprocess video for manual FACS annotation, the method significantly improves productivity, thus addressing the need for ground-truth data for facial image analysis. Moreover, we were also able to detect unusual facial behavior. Fernando De la Torre, Joan Campoy, Zara Ambadar, Jeff F. Conn |
ICCV | 1 |
| 2007 | Multimodal DiariesabstractTime management is an important aspect of a successful professional life. In order to have a better understanding of where our time goes, we propose a system that summarizes the user's daily activity (e.g. sleeping, walking, working on the pc, talking, ...) using all-day multimodal data recordings. Two main novelties are proposed: (i) a system that combines both physical and contextual awareness hardware and software. It records synchronized audio, video, body sensors, GPS and computer monitoring data. (ii) A semi-supervised temporal clustering (SSTC) algorithm that accurately and efficiently groups large amounts of multimodal data into different activities. The effectiveness and accuracy of our SSTC is demonstrated in synthetic and real examples of activity segmentation from multimodal data gathered over long periods of time. Fernando De la Torre, Carlos Agell |
ICME | 1 |
| 2007 | Indoor people tracking based on dynamic weighted multidimensional scalingabstractAccurate location of people in indoor environments is a key aspect of many applications such as resource management or security. In this paper, we explore the use of short-range radio technologies to track people indoors. The network consists of two kind of radio nodes: static nodes (anchors) and mobile nodes (people). From a set of sparse connectivity matrices (people vs. people and people vs. anchors) at each time instant and people's dynamics, we infer people's trajectories. To combine connectivity and dynamic information, we propose an extension of Multidimensional Scaling(MDS), Dynamic Weighted MDS (DWMDS), that finds an embedding of people's trajectories (x and y coordinates of people through time). DWMDS has proven to be more accurate and effective, especially for low connectivity degree networks (i.e. sparse networks), compared to existing location algorithms. Extensive simulations show the effectiveness and robustness of the proposed algorithm. José María Cabero, Fernando De la Torre, Aritz Sanchez, Iñigo Arizaga |
MSWiM | 2 |
| 2006 | Automatic Clustering of Faces in MeetingsabstractMeetings are an integral part of business life for any organization. In previous work, we have developed a physical awareness system called CAMEO (camera assisted meeting event observer) to record and process the audio/visual information of a meeting. An important task in meeting understanding is to know who and how many people are attending the meeting. In this paper, we present an automatic approach to detect, track, and cluster people's faces in long video sequences. This is a challenging problem due to the appearance variability of people's faces (illumination, expression, pose,...). Two main novelties are presented: a robust real-time adaptive subspace face tracker which combines color and appearance. A temporal subspace clustering algorithm. The effectiveness and robustness of the proposed system is demonstrated over a data set of long videos (i.e. 1 hour). Carlos Vallespí, Fernando De la Torre, Manuela M. Veloso, Takeo Kanade |
ICIP | 2 |
| 2006 | Discriminative cluster analysisabstractClustering is one of the most widely used statistical tools for data analysis. Among all existing clustering techniques, k-means is a very popular method because of its ease of programming and because it accomplishes a good trade-off between achieved performance and computational complexity. However, k-means is prone to local minima problems, and it does not scale too well with high dimensional data sets. A common approach to dealing with high dimensional data is to cluster in the space spanned by the principal components (PC). In this paper, we show the benefits of clustering in a low dimensional discriminative space rather than in the PC space (generative). In particular, we propose a new clustering algorithm called Discriminative Cluster Analysis (DCA). DCA jointly performs dimensionality reduction and clustering. Several toy and real examples show the benefits of DCA versus traditional PCA+k-means clustering. Additionally, a new matrix formulation is proposed and connections with related techniques such as spectral graph methods and linear discriminant analysis are provided. Fernando De la Torre, Takeo Kanade |
ICML | 1 |
| 2005 | Representational Oriented Component Analysis (ROCA) for Face Recognition with One Sample Image per Training ClassabstractSubspace methods such as PCA, LDA, ICA have become a standard tool to perform visual learning and recognition. In this paper we propose representational oriented component analysis (ROCA), an extension of OCA, to perform face recognition when just one sample per training class is available. Several novelties are introduced in order to improve generalization and efficiency: (1) combining several OCA classifiers based on different image representations of the unique training sample is shown to greatly improve the recognition performance. (2) To improve generalization and to account for small misregistration effect, a learned subspace is added to constrain the OCA solution, (3) a stable/efficient generalized eigenvector algorithm that solves the small size sample problem and avoids overfitting. Preliminary experiments in the FRGC Ver 1.0 dataset show that ROCA outperforms existing linear techniques (PCA, OCA) and some commercial systems. Fernando De la Torre, Ralph Gross, Simon Baker, B. V. K. Vijaya Kumar |
CVPR (2) | 1 |
| 2005 | Multimodal oriented discriminant analysisabstractLinear discriminant analysis (LDA) has been an active topic of research during the last century. However, the existing algorithms have several limitations when applied to visual data. LDA is only optimal for Gaussian distributed classes with equal covariance matrices, and only classes-1 features can be extracted. On the other hand, LDA does not scale well to high dimensional data (overfitting), and it cannot handle optimally multimodal distributions. In this paper, we introduce Multimodal Oriented Discriminant Analysis (MODA), a LDA extension which can overcome these drawbacks. A new formulation and several novelties are proposed:• An optimal dimensionality reduction for multimodal Gaussian classes with different covariances is derived. The new criteria allows for extracting more than classes-1 features.• A covariance approximation is introduced to improve generalization and avoid over-fitting when dealing with high dimensional data.• A linear time iterative majorization method is suggested in order to find a local optimum.Several synthetic and real experiments on face recognition show that MODA outperform existing linear techniques. Fernando De la Torre, Takeo Kanade |
ICML | 1 |
| 2005 | Learning to Track Multiple People in Omnidirectional VideoabstractMeetings are a very important part of everyday life for professionals working in universities, companies or governmental institutions. We have designed a physical awareness system called CAMEO (Camera Assisted Meeting Event Observer), a hardware/software system to record and monitor people's activities in meetings. CAMEO captures a high resolution omnidirectional view of the meeting by stitching images coming from almost concentric cameras. Besides recording capability, CAMEO automatically detects people and learns a person-specific facial appearance model (PS-FAM) for each of the participants. The PSFAMs allow more robust/reliable tracking and identification. In this paper, we describe the video-capturing device, photometric/geometric autocalibration process, and the multiple people tracking system. The effectiveness and robustness of the proposed system is demonstrated over several real-time experiments and a large data set of videos. Fernando De la Torre, Carlos Vallespí, Paul E. Rybski, Manuela M. Veloso, Takeo Kanade |
ICRA | 1 |
| 2004 | CAMEO: Modeling Human Activity in Formal Meeting Situations
Paul E. Rybski, Fernando De la Torre, Raju Patil, Carlos Vallespí, Manuela M. Veloso, Brett Browning |
AAAI | 2 |
| 2004 | Oriented Discriminant Analysis
Fernando De la Torre, Takeo Kanade |
BMVC | 1 |
| 2004 | Segmentation and classification of meetings using multiple information streamsabstractWe present a meeting recorder infrastructure used to record and annotate events that occur in meetings. Multiple data streams are recorded and analyzed in order to infer a higher-level state of the group’s activities. We describe the hardware and software systems used to capture people’s activities as well as the methods used to characterize them. Paul E. Rybski, Satanjeev Banerjee, Fernando De la Torre, Carlos Vallespí, Alexander I. Rudnicky, Manuela M. Veloso |
ICMI | 3 |
| 2004 | CAMEO: Camera Assisted Meeting Event ObserverabstractStatic cameras are pervasive in a variety of environments. However it remains a challenging problem to extract and reason about high-level features from real-time and continuous observation of an environment. In this paper, we present CAMEO, the Camera Assisted Meeting Event Observer, which is a physical awareness system designed for use by an agent-based electronic assistant. CAMEO is an inexpensive high-resolution omnidirectional vision system designed to be used in meeting environments. The multiple camera design achieves the desired high image resolution and lower cost that can be achieved when compared to traditional omnicameras that make use of a single camera and mirror solution. Paul E. Rybski, Fernando De la Torre, Raju Patil, Carlos Vallespí, Manuela M. Veloso, Brett Browning |
ICRA | 2 |
| 2004 | Robust normalization of silhouettes for recognition applications
Javier Cortadellas, Josep Amat, Fernando De la Torre |
Pattern Recognit. Lett. | 3 |
| 2003 | Text to visual synthesis with appearance modelsabstractThis paper presents a new method named text to visual synthesis with appearance models (TEVISAM) for generating videorealistic talking heads. In a first step, the system learns a person-specific facial appearance model (PSFAM) automatically. PSFAM allows modeling all facial components (e.g. eyes, mouth, etc) independently and it will be used to animate the face from the input text dynamically. As reported by other researches, one of the key aspects in visual synthesis is the coarticulation effect. To solve such a problem, we introduce a new interpolation method in the high dimensional space of appearance allowing to create photorealistic and videorealistic avatars. In this work, preliminary experiments synthesizing virtual avatars from text are reported. Summarizing, in this paper we introduce three novelties: first, we make use of color PSFAM to animate virtual avatars; second, we introduce a nonlinear high dimensional interpolation to achieve videorealistic animations; finally, this method allows to generate new expressions modeling the different facial elements. Javier Melenchón, Fernando De la Torre, Ignasi Iriondo Sanz, Francesc Alías, Elisa Martínez Marroquín, Lluís Vicent |
ICIP (1) | 2 |
| 2003 | Subspace eyetracking for driver warningabstractDriver's fatigue/distraction is one of the most common causes of traffic accidents. The aim of this paper is to develop a real time system to detect anomalous situations while driving. In a learning stage, the user will sit in front of the camera and the system will learn a person-specific facial appearance model (PSFAM) in an automatic manner. The PSFAM will be used to perform gaze detection and eye-activity recognition in a real time based on subspace constraints. Preliminary experiments measuring the PERCLOS index (average time that the eyes are closed) under a variety of conditions are reported. Fernando De la Torre, Carlos Javier Garcia Rubio, Elisa Martínez Marroquín |
ICIP (3) | 1 |
| 2003 | Robust parameterized component analysis: theory and applications to 2D facial appearance models
Fernando De la Torre, Michael J. Black |
Comput. Vis. Image Underst. | 1 |
| 2003 | A Framework for Robust Subspace Learning
Fernando De la Torre, Michael J. Black |
Int. J. Comput. Vis. | 1 |
| 2002 | Robust Parameterized Component Analysis
Fernando De la Torre, Michael J. Black |
ECCV (4) | 1 |
| 2001 | Dynamic Coupled Component AnalysisabstractWe present a method for simultaneously learning linear models of multiple high dimensional data sets and the dependencies between them. For example, we learn asymmetrically coupled linear models for the faces of two different people and show how these models can be used to animate one face given a video sequence of the other. We pose the problem as a form of Asymmetric Coupled Component Analysis (ACCA) in which we simultaneously learn the subspaces for reducing the dimensionality of each dataset while coupling the parameters of the low dimensional representations. Additionally, a dynamic form of ACCA is proposed, that extends this work to model temporal dependencies in the data sets. To account for outliers and missing data, we formulate the problem in a statistically robust estimation framework. We review connections with previous work and illustrate the method with examples of synthesized dancing and the animation of facial avatars. Fernando De la Torre, Michael J. Black |
CVPR (2) | 1 |
| 2001 | Robust Principal Component Analysis for Computer Vision
Fernando De la Torre, Michael J. Black |
ICCV | 1 |
| 2000 | A Framework for Modeling the Appearance of 3D Articulated FiguresabstractThis paper describes a framework for constructing a linear subspace model of image appearance for complex articulated 3D figures such as humans and other animals. A commercial motion capture system provides 3D data that is aligned with images of subjects performing various activities. Portions of a limb's image appearance are seen from multiple views and for multiple subjects. From these partial views, weighted principal component analysis is used to construct a linear subspace representation of the "unwrapped" image appearance of each limb. The linear subspaces provide a generative model of the object appearance that is exploited in a Bayesian particle filtering tracking system. Results of tracking single limbs and walking humans are presented. Hedvig Kjellström, Fernando De la Torre, Michael J. Black |
FG | 2 |
| 2000 | A Probabilistic Framework for Rigid and Non-Rigid Appearance Based Tracking and RecognitionabstractThis paper describes an unified probabilistic framework for appearance-based tracking of rigid and non-rigid objects. A spatio-temporal dependent shape-texture eigenspace and mixture of diagonal Gaussians are learned in a hidden Markov model (HMM)-like structure to better constrain the model and for recognition purposes. Particle filtering is used to track the object while switching between different shape/texture models. This framework allows recognition and temporal segmentation of activities. Additionally an automatic stochastic initialization is proposed, the number of states in the HMM are selected based on the Akaike information criterion and comparison with deterministic tracking for 2D models is discussed. Preliminary results of eye tracking, lip tracking and temporal segmentation of mouth events are presented. Fernando De la Torre, Yaser Yacoob, Larry Davis 0001 |
FG | 1 |
| 2000 | Eigenfiltering for Flexible Eigentracking (EFE)abstractTraditional techniques for tracking nonrigid objects such as optical flow, correlation, active contours or color, cannot deal with situations where image changes are not due to motion but appearance (e.g. tracking the lips when the teeth appear). Two main contributions for appearance tracking of flexible objects are proposed. The first one is a flexible generalization of eigentracking within the same robust continuous optimization framework. The second one is a generalization of traditional graylevel eigenspaces, constructing a multiple channel "eigenspace" using filter responses to give robustness against variations in the training conditions such as illumination changes. Additionally, 3D geometric transformations are incorporated, a regularization term is added for numerical stability reasons and the optimization problem is solved in closed form. Experiments on lip tracking are reported. Fernando De la Torre, Javier Melenchón, Jordi Vitrià, Petia Radeva |
ICPR | 1 |
| 1998 | View-Based Adaptive Affine Tracking
Fernando De la Torre, Shaogang Gong, Stephen J. McKenna |
ECCV (1) | 1 |
| 1998 | View Alignment with Dynamically Updated Affine Tracking
Fernando De la Torre, Shaogang Gong, Stephen J. McKenna |
FG | 1 |