EDBT 2026 Demo / reviewers in the wild / expert
Xiang Yu 0002
dblp:19/2453-2
· DBLP profile ↗
39ranked-venue papers
7as first author
14since 2021 · last 2025
0000-0003-2765-2749ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 5 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Revisiting Source-Free Domain Adaptation: Insights into Representativeness, Generalization, and VarietyabstractDomain adaptation addresses the challenge where the distribution of target inference data differs from that of the source training data. Recently, data privacy has become a significant constraint, limiting access to the source domain. To mitigate this issue, Source-Free Domain Adaptation (SFDA) methods bypass source domain data by generating source-like data or pseudo-labeling the unlabeled target domain. However, these approaches often lack theoretical grounding. In this work, we provide a theoretical analysis of the SFDA problem, focusing on the general empirical risk of the unlabeled target domain. Our analysis offers a comprehensive understanding of how representativeness, generalization, and variety contribute to controlling the upper bound of target domain empirical risk in SFDA settings. We further explore how to balance this trade-off from three perspectives: sample selection, semantic domain alignment, and a progressive learning framework. These insights inform the design of novel algorithms. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on three benchmark datasets—Office-Home, DomainNet, and VisDA-C—yielding relative improvements of 3.2%, 9.1%, and 7.5%, respectively, over the representative SFDA method, SHOT. Ronghang Zhu, Mengxuan Hu, Weiming Zhuang, Lingjuan Lyu, Xiang Yu 0002, Sheng Li 0001 |
CVPR | 5 |
| 2024 | A Survey of Trustworthy Representation Learning Across DomainsabstractAs AI systems have obtained significant performance to be deployed widely in our daily lives and human society, people both enjoy the benefits brought by these technologies and suffer many social issues induced by these systems. To make AI systems good enough and trustworthy, plenty of researches have been done to build guidelines for trustworthy AI systems. Machine learning is one of the most important parts of AI systems, and representation learning is the fundamental technology in machine learning. How to make representation learning trustworthy in real-world application, e.g., cross domain scenarios, is very valuable and necessary for both machine learning and AI system fields. Inspired by the concepts in trustworthy AI, we proposed the first trustworthy representation learning across domains framework, which includes four concepts, i.e., robustness, privacy, fairness, and explainability, to give a comprehensive literature review on this research direction. Specifically, we first introduce the details of the proposed trustworthy framework for representation learning across domains. Second, we provide basic notions and comprehensively summarize existing methods for the trustworthy framework from four concepts. Finally, we conclude this survey with insights and discussions on future research directions. Ronghang Zhu, Dongliang Guo 0002, Daiqing Qi, Zhixuan Chu, Xiang Yu 0002, Sheng Li 0001 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2023 | Q: How to Specialize Large Vision-Language Models to Data-Scarce VQA Tasks? A: Self-Train on Unlabeled Images!abstractFinetuning a large vision language model (VLM) on a target dataset after large scale pretraining is a dominant paradigm in visual question answering (VQA). Datasets for specialized tasks such as knowledge-based VQA or VQA in non natural-image domains are orders of magnitude smaller than those for general-purpose VQA. While collecting additional labels for specialized tasks or domains can be challenging, unlabeled images are often available. We introduce SelTDA (Self-Taught Data Augmentation), a strategy for finetuning large VLMs on small-scale VQA datasets. SelTDA uses the VLM and target dataset to build a teacher model that can generate question-answer pseudolabels directly conditioned on an image alone, allowing us to pseudolabel unlabeled images. SelTDA then finetunes the initial VLM on the original dataset augmented with freshly pseudolabeled images. We describe a series of experiments showing that our self-taught data augmentation increases robustness to adversarially searched questions, counterfactual examples and rephrasings, improves domain generalization, and results in greater retention of numerical reasoning skills. The proposed strategy requires no additional annotations or architectural modifications, and is compatible with any modern encoder-decoder multimodal transformer. Code available at https://github.com/codezakh/SelTDA. Zaid Khan 0001, Samuel Schulter, Xiang Yu 0002, Yun Fu 0001, Manmohan Krishna Chandraker |
CVPR | 4 |
| 2023 | DeFormer: Integrating Transformers with Deformable Models for 3D Shape Abstraction from a Single ImageabstractAccurate 3D shape abstraction from a single 2D image is a long-standing problem in computer vision and graphics. By leveraging a set of primitives to represent the target shape, recent methods have achieved promising results. However, these methods either use a relatively large number of primitives or lack geometric flexibility due to the limited expressibility of the primitives. In this paper, we propose a novel bi-channel Transformer architecture, integrated with parameterized deformable models, termed DeFormer, to simultaneously estimate the global and local deformations of primitives. In this way, DeFormer can abstract complex object shapes while using a small number of primitives which offer a broader geometry coverage and finer details. Then, we introduce a force-driven dynamic fitting and a cycle-consistent re-projection loss to optimize the primitive parameters. Extensive experiments on ShapeNet across various settings show that DeFormer achieves better reconstruction accuracy over the state-of-the-art, and visualizes with consistent semantic correspondences for improved interpretability. Di Liu 0003, Xiang Yu 0002, Meng Ye 0003, Qilong Zhangli, Zhuowei Li 0002, Dimitris N. Metaxas |
ICCV | 2 |
| 2023 | Domain Generalization Guided by Gradient Signal to Noise Ratio of ParametersabstractOverfitting to the source domain is a common issue in gradient-based training of deep neural networks. To compensate for the over-parameterized models, numerous regularization techniques have been introduced such as those based on dropout. While these methods achieve significant improvements on classical benchmarks such as ImageNet, their performance diminishes with the introduction of domain shift in the test set i.e. when the unseen data comes from a significantly different distribution. In this paper, we move away from the classical approach of Bernoulli sampled dropout mask construction and propose to base the selection on gradient-signal-to-noise ratio (GSNR) of network’s parameters. Specifically, at each training step, parameters with high GSNR will be discarded. Furthermore, we alleviate the burden of manually searching for the optimal dropout ratio by leveraging a meta-learning approach. We evaluate our method on standard domain generalization benchmarks and achieve competitive results on classification and face anti-spoofing problems. Mateusz Michalkiewicz, Masoud Faraki, Xiang Yu 0002, Manmohan Krishna Chandraker, Mahsa Baktash |
ICCV | 3 |
| 2023 | Progressive Mix-Up for Few-Shot Supervised Multi-Source Domain Transfer
Ronghang Zhu, Xiang Yu 0002, Sheng Li 0001 |
ICLR | 2 |
| 2023 | Split to Learn: Gradient Split for Multi-Task Human Image AnalysisabstractThis paper presents an approach to train a unified deep network that simultaneously solves multiple human-related tasks. A multi-task framework is favorable for sharing information across tasks under restricted computational resources. However, tasks not only share information but may also compete for resources and conflict with each other, making the optimization of shared parameters difficult and leading to suboptimal performance. We propose a simple but effective training scheme called GradSplit that alleviates this issue by utilizing asymmetric inter-task relations. Specifically, at each convolution module, it splits features into T groups for T tasks and trains each group only using the gradient back-propagated from the task losses with which it does not have conflicts. During training, we apply GradSplit to a series of convolution modules. As a result, each module is trained to generate a set of task-specific features using the shared features from the previous module. This enables a network to use complementary information across tasks while circumventing gradient conflicts. Experimental results show that GradSplit achieves a better accuracy-efficiency trade-off than existing methods. It minimizes accuracy drop caused by task conflicts while significantly saving compute resources in terms of both FLOPs and memory at inference. We further show that GradSplit achieves higher cross-dataset accuracy compared to single-task and other multi-task networks. Weijian Deng, Yumin Suh, Xiang Yu 0002, Masoud Faraki, Liang Zheng 0001, Manmohan Krishna Chandraker |
WACV | 3 |
| 2022 | Learning to Learn across Diverse Data Biases in Deep Face RecognitionabstractConvolutional Neural Networks have achieved remarkable success in face recognition, in part due to the abundant availability of data. However, the data used for training CNNs is often imbalanced. Prior works largely focus on the long-tailed nature of face datasets in data volume per identity, or focus on single bias variation. In this paper, we show that many bias variations such as ethnicity, head pose, occlusion and blur can jointly affect the accuracy significantly. We propose a sample level weighting approach termed Multi-variation Cosine Margin (MvCoM), to simultaneously consider the multiple variation factors, which orthogonally enhances the face recognition losses to incorporate the importance of training samples. Further, we leverage a learning to learn approach, guided by a held-out meta learning set and use an additive modeling to predict the MvCoM. Extensive experiments on challenging face recognition benchmarks demonstrate the advantages of our method in jointly handling imbalances due to multiple variations. Chang Liu 0022, Xiang Yu 0002, Yi-Hsuan Tsai, Masoud Faraki, Ramin Moslemi, Manmohan Krishna Chandraker, Yun Fu 0001 |
CVPR | 2 |
| 2022 | Controllable Dynamic Multi-Task ArchitecturesabstractMulti-task learning commonly encounters competition for resources among tasks, specifically when model capac-ity is limited. This challenge motivates models which al-low control over the relative importance of tasks and total compute cost during inference time. In this work, we pro-pose such a controllable multi-task network that dynami-cally adjusts its architecture and weights to match the de-sired task preference as well as the resource constraints. In contrast to the existing dynamic multi-task approaches that adjust only the weights within a fixed architecture, our approach affords the flexibility to dynamically control the total computational cost and match the user-preferred task importance better. We propose a disentangled training of two hype rnetwo rks, by exploiting task affinity and a novel branching regularized loss, to take input prefer-ences and accordingly predict tree-structured models with adapted weights. Experiments on three multi-task bench-marks, namely PASCAL-Context, NYU-v2, and CIFAR-100, show the efficacy of our approach. Project page is available at https://www.nec-labs.com/-mas/DYMU. Dripta S. Raychaudhuri, Yumin Suh, Samuel Schulter, Xiang Yu 0002, Masoud Faraki, Amit K. Roy-Chowdhury, Manmohan Krishna Chandraker |
CVPR | 4 |
| 2022 | On Generalizing Beyond Domains in Cross-Domain Continual LearningabstractHumans have the ability to accumulate knowledge of new tasks in varying conditions, but deep neural networks of-ten suffer from catastrophic forgetting of previously learned knowledge after learning a new task. Many recent methods focus on preventing catastrophic forgetting under the assumption of train and test data following similar distributions. In this work, we consider a more realistic scenario of continual learning under domain shifts where the model must generalize its inference to an unseen domain. To this end, we encourage learning semantically meaningful features by equipping the classifier with class similarity metrics as learning parameters which are obtained through Mahalanobis similarity computations. Learning of the backbone representation along with these extra parameters is done seamlessly in an end-to-end manner. In addition, we propose an approach based on the exponential moving average of the parameters for better knowledge distillation. We demonstrate that, to a great extent, existing continual learning algorithms fail to handle the forgetting issue under multiple distributions, while our proposed approach learns new tasks under domain shift with accuracy boosts up to 10% on challenging datasets such as DomainNet and OfficeHome. Christian Simon, Masoud Faraki, Yi-Hsuan Tsai, Xiang Yu 0002, Samuel Schulter, Yumin Suh, Mehrtash Harandi, Manmohan Krishna Chandraker |
CVPR | 4 |
| 2022 | Single-Stream Multi-level Alignment for Vision-Language Pretraining
Zaid Khan 0001, Xiang Yu 0002, Samuel Schulter, Manmohan Krishna Chandraker, Yun Fu 0001 |
ECCV (36) | 3 |
| 2022 | Learning Phase Mask for Privacy-Preserving Passive Depth Estimation
Zaid Tasneem, Giovanni Milione, Yi-Hsuan Tsai, Xiang Yu 0002, Ashok Veeraraghavan, Manmohan Krishna Chandraker, Francesco Pittaluga |
ECCV (7) | 4 |
| 2021 | Cross-Domain Similarity Learning for Face Recognition in Unseen DomainsabstractFace recognition models trained under the assumption of identical training and test distributions often suffer from poor generalization when faced with unknown variations, such as a novel ethnicity or unpredictable individual make-ups during test time. In this paper, we introduce a novel cross-domain metric learning loss, which we dub Cross-Domain Triplet (CDT) loss, to improve face recognition in unseen domains. The CDT loss encourages learning semantically meaningful features by enforcing compact feature clusters of identities from one domain, where the compactness is measured by underlying similarity metrics that belong to another training domain with different statistics. Intuitively, it discriminatively correlates explicit metrics derived from one domain, with triplet samples from another domain in a unified loss function to be minimized within a network, which leads to better alignment of the training domains. The network parameters are further enforced to learn generalized features under domain shift, in a model-agnostic learning pipeline. Unlike the recent work of Meta Face Recognition [18], our method does not require careful hard-pair sample mining and filtering strategy during training. Extensive experiments on various face recognition benchmarks show the superiority of our method in handling variations, compared to baseline and the state-of-the-art methods. Masoud Faraki, Xiang Yu 0002, Yi-Hsuan Tsai, Yumin Suh, Manmohan Krishna Chandraker |
CVPR | 2 |
| 2021 | Learning Cross-Modal Contrastive Features for Video Domain AdaptationabstractLearning transferable and domain adaptive feature representations from videos is important for video-relevant tasks such as action recognition. Existing video domain adaptation methods mainly rely on adversarial feature alignment, which has been derived from the RGB image space. However, video data is usually associated with multi-modal information, e.g., RGB and optical flow, and thus it remains a challenge to design a better method that considers the cross-modal inputs under the cross-domain adaptation setting. To this end, we propose a unified framework for video domain adaptation, which simultaneously regularizes cross-modal and cross-domain feature representations. Specifically, we treat each modality in a domain as a view and leverage the contrastive learning technique with properly designed sampling strategies. As a result, our objectives regularize feature spaces, which originally lack the connection across modalities or have less alignment across domains. We conduct experiments on domain adaptive action recognition benchmark datasets, i.e., UCF, HMDB, and EPIC-Kitchens, and demonstrate the effectiveness of our components against state-of-the-art algorithms. Donghyun Kim 0006, Yi-Hsuan Tsai, Bingbing Zhuang, Xiang Yu 0002, Stan Sclaroff, Kate Saenko, Manmohan Krishna Chandraker |
ICCV | 4 |
| 2020 | Private-kNN: Practical Differential Privacy for Computer VisionabstractWith increasing ethical and legal concerns on privacy for deep models in visual recognition, differential privacy has emerged as a mechanism to disguise membership of sensitive data in training datasets. Recent methods like Private Aggregation of Teacher Ensembles (PATE) leverage a large ensemble of teacher models trained on disjoint subsets of private data, to transfer knowledge to a student model with privacy guarantees. However, labeled vision data is often expensive and datasets, when split into many disjoint training sets, lead to significantly sub-optimal accuracy and thus hardly sustain good privacy bounds. We propose a practically data-efficient scheme based on private release of k-nearest neighbor (kNN) queries, which altogether avoids splitting the training dataset. Our approach allows the use of privacy-amplification by subsampling and iterative refinement of the kNN feature embedding. We rigorously analyze the theoretical properties of our method and demonstrate strong experimental performance on practical computer vision datasets for face attribute recognition and person reidentification. In particular, we achieve comparable or better accuracy than PATE while reducing more than 90% of the privacy loss, thereby providing the “most practical method to-date” for private deep learning in computer vision. Yuqing Zhu 0005, Xiang Yu 0002, Manmohan Krishna Chandraker, Yu-Xiang Wang 0003 |
CVPR | 2 |
| 2020 | Towards Universal Representation Learning for Deep Face RecognitionabstractRecognizing wild faces is extremely hard as they appear with all kinds of variations. Traditional methods either train with specifically annotated variation data from target domains, or by introducing unlabeled target variation data to adapt from the training data. Instead, we propose a universal representation learning framework that can deal with larger variation unseen in the given training data without leveraging target domain knowledge. We firstly synthesize training data alongside some semantically meaningful variations, such as low resolution, occlusion and head pose. However, directly feeding the augmented data for training will not converge well as the newly introduced samples are mostly hard examples. We propose to split the feature embedding into multiple sub-embeddings, and associate different confidence values for each sub-embedding to smooth the training procedure. The sub-embeddings are further decorrelated by regularizing variation classification loss and variation adversarial loss on different partitions of them. Experiments show that our method achieves top performance on general face recognition datasets such as LFW and MegaFace, while significantly better on extreme benchmarks such as TinyFace and IJB-S. Yichun Shi, Xiang Yu 0002, Kihyuk Sohn, Manmohan Krishna Chandraker, Anil K. Jain 0001 |
CVPR | 2 |
| 2020 | Improving Face Recognition by Clustering Unlabeled Faces in the Wild
Aruni Roy Chowdhury, Xiang Yu 0002, Kihyuk Sohn, Erik G. Learned-Miller, Manmohan Krishna Chandraker |
ECCV (24) | 2 |
| 2020 | DAVID: Dual-Attentional Video DeblurringabstractBlind video deblurring restores sharp frames from a blurry sequence without any prior. It is a challenging task because the blur due to camera shake, object movement and defocusing is heterogeneous in both temporal and spatial dimensions. Traditional methods train on datasets synthesized with a single level of blur, and thus do not generalize well across levels of blurriness. To address this challenge, we propose a dual attention mechanism to dynamically aggregate temporal cues for deblurring with an end-to-end trainable network structure. Specifically, an internal attention module adaptively selects the optimal temporal scales for restoring the sharp center frame. An external attention module adaptively aggregates and refines multiple sharp frame estimates, from several internal attention modules designed for different blur levels. To train and evaluate on more diverse blur severity levels, we propose a Challenging DVD dataset generated from the raw DVD video set by pooling frames with different temporal windows. Our framework achieves consistently better performance on this more challenging dataset while obtaining strongly competitive results on the original DVD benchmark. Extensive ablative studies and qualitative visualizations further demonstrate the advantage of our method in handling real video blur. Xiang Yu 0002, Ding Liu 0001, Manmohan Krishna Chandraker, Zhangyang Wang |
WACV | 2 |
| 2019 | Feature Transfer Learning for Face Recognition With Under-Represented DataabstractDespite the large volume of face recognition datasets, there is a significant portion of subjects, of which the samples are insufficient and thus under-represented. Ignoring such significant portion results in insufficient training data. Training with under-represented data leads to biased classifiers in conventionally-trained deep networks. In this paper, we propose a center-based feature transfer framework to augment the feature space of under-represented subjects from the regular subjects that have sufficiently diverse samples. A Gaussian prior of the variance is assumed across all subjects and the variance from regular ones are transferred to the under-represented ones. This encourages the under-represented distribution to be closer to the regular distribution. Further, an alternating training regimen is proposed to simultaneously achieve less biased classifiers and a more discriminative feature representation. We conduct ablative study to mimic the under-represented datasets by varying the portion of under-represented classes on the MS-Celeb-1M dataset. Advantageous results on LFW, IJB-A and MS-Celeb-1M demonstrate the effectiveness of our feature transfer and training strategy, compared to both general baselines and state-of-the-art methods. Moreover, our feature transfer successfully presents smooth visual interpolation, which conducts disentanglement to preserve identity of a class while augmenting its feature space with non-identity variations such as pose and lighting. Xi Yin 0001, Xiang Yu 0002, Kihyuk Sohn, Xiaoming Liu 0002, Manmohan Krishna Chandraker |
CVPR | 2 |
| 2019 | Gotta Adapt 'Em All: Joint Pixel and Feature-Level Domain Adaptation for Recognition in the WildabstractRecent developments in deep domain adaptation have allowed knowledge transfer from a labeled source domain to an unlabeled target domain at the level of intermediate features or input pixels. We propose that advantages may be derived by combining them, in the form of different insights that lead to a novel design and complementary properties that result in better performance. At the feature level, inspired by insights from semi-supervised learning, we propose a classification-aware domain adversarial neural network that brings target examples into more classifiable regions of source domain. Next, we posit that computer vision insights are more amenable to injection at the pixel level. In particular, we use 3D geometry and image synthesis based on a generalized appearance flow to preserve identity across pose transformations, while using an attribute-conditioned CycleGAN to translate a single source into multiple target images that differ in lower-level properties such as lighting. Besides standard UDA benchmark, we validate on a novel and apt problem of car recognition in unlabeled surveillance images using labeled images from the web, handling explicitly specified, nameable factors of variation through pixel-level and implicit, unspecified factors through feature-level adaptation. Luan Tran, Kihyuk Sohn, Xiang Yu 0002, Xiaoming Liu 0002, Manmohan Krishna Chandraker |
CVPR | 3 |
| 2019 | Unsupervised Domain Adaptation for Distance Metric Learning
Kihyuk Sohn, Wenling Shang, Xiang Yu 0002, Manmohan Krishna Chandraker |
ICLR (Poster) | 3 |
| 2019 | A coupled encoder-decoder network for joint face detection and landmark localization
Lezi Wang, Xiang Yu 0002, Thirimachos Bourlai, Dimitris N. Metaxas |
Image Vis. Comput. | 2 |
| 2019 | Deep Supervision with Intermediate ConceptsabstractRecent data-driven approaches to scene interpretation predominantly pose inference as an end-to-end black-box mapping, commonly performed by a Convolutional Neural Network (CNN). However, decades of work on perceptual organization in both human and machine vision suggest that there are often intermediate representations that are intrinsic to an inference task, and which provide essential structure to improve generalization. In this work, we explore an approach for injecting prior domain structure into neural network training by supervising hidden layers of a CNN with intermediate concepts that normally are not observed in practice. We formulate a probabilistic framework which formalizes these notions and predicts improved generalization via this deep supervision method. One advantage of this approach is that we are able to train only from synthetic CAD renderings of cluttered scenes, where concept values can be extracted, but apply the results to real images. Our implementation achieves the state-of-the-art performance of 2D/3D keypoint localization and image classification on real image benchmarks including KITTI, PASCAL VOC, PASCAL3D+, IKEA, and CIFAR100. We provide additional evidence that our approach outperforms alternative forms of supervision, such as multi-task networks. M. Zeeshan Zia, Quoc-Huy Tran, Xiang Yu 0002, Gregory D. Hager, Manmohan Krishna Chandraker |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Deep Supervision with Shape Concepts for Occlusion-Aware 3D Object ParsingabstractMonocular 3D object parsing is highly desirable in various scenarios including occlusion reasoning and holistic scene interpretation. We present a deep convolutional neural network (CNN) architecture to localize semantic parts in 2D image and 3D space while inferring their visibility states, given a single RGB image. Our key insight is to exploit domain knowledge to regularize the network by deeply supervising its hidden layers, in order to sequentially infer intermediate concepts associated with the final task. To acquire training data in desired quantities with ground truth 3D shape and relevant concepts, we render 3D object CAD models to generate large-scale synthetic data and simulate challenging occlusion configurations between objects. We train the network only on synthetic data and demonstrate state-of-the-art performances on real image benchmarks including an extended version of KITTI, PASCAL VOC, PASCAL3D+ and IKEA for 2D and 3D keypoint localization and instance segmentation. The empirical results substantiate the utility of our deep supervision scheme by demonstrating effective transfer of knowledge from synthetic data to real images, resulting in less overfitting compared to standard end-to-end training. M. Zeeshan Zia, Quoc-Huy Tran, Xiang Yu 0002, Gregory D. Hager, Manmohan Krishna Chandraker |
CVPR | 4 |
| 2017 | A Coupled Encoder-Decoder Network for Joint Face Detection and Landmark LocalizationabstractFace detection and landmark localization have been extensively investigated and are the prerequisite for many face applications, such as face recognition and 3D face reconstruction. Most existing methods achieve success on only one of the two problems. In this paper, we propose a coupled encoder-decoder network to jointly detect faces and localize facial key points. The encoder and decoder generate response maps for facial landmark localization. Moreover, we observe that the intermediate feature maps from the encoder and decoder have strong power in describing facial regions, which motivates us to build a unified framework by coupling the feature maps for multi-scale cascaded face detection. Experiments on face detection show strongly competitive results against the existing methods on two public benchmarks. The landmark localization further shows consistently better accuracy than state-of-the-arts on three face-in-the-wild databases. Lezi Wang, Xiang Yu 0002, Dimitris N. Metaxas |
FG | 2 |
| 2017 | Reconstruction-Based Disentanglement for Pose-Invariant Face RecognitionabstractDeep neural networks (DNNs) trained on large-scale datasets have recently achieved impressive improvements in face recognition. But a persistent challenge remains to develop methods capable of handling large pose variations that are relatively under-represented in training data. This paper presents a method for learning a feature representation that is invariant to pose, without requiring extensive pose coverage in training data. We first propose to generate non-frontal views from a single frontal face, in order to increase the diversity of training data while preserving accurate facial details that are critical for identity discrimination. Our next contribution is to seek a rich embedding that encodes identity features, as well as non-identity ones such as pose and landmark locations. Finally, we propose a new feature reconstruction metric learning to explicitly disentangle identity and pose, by demanding alignment between the feature reconstructions through various combinations of identity and pose features, which is obtained from two images of the same subject. Experiments on both controlled and in-the-wild face datasets, such as MultiPIE, 300WLP and the profile view database CFP, show that our method consistently outperforms the state-of-the-art, especially on images with large head pose variations. Xi Peng 0005, Xiang Yu 0002, Kihyuk Sohn, Dimitris N. Metaxas, Manmohan Krishna Chandraker |
ICCV | 2 |
| 2017 | Unsupervised Domain Adaptation for Face Recognition in Unlabeled VideosabstractDespite rapid advances in face recognition, there remains a clear gap between the performance of still image-based face recognition and video-based face recognition, due to the vast difference in visual quality between the domains and the difficulty of curating diverse large-scale video datasets. This paper addresses both of those challenges, through an image to video feature-level domain adaptation approach, to learn discriminative video frame representations. The framework utilizes large-scale unlabeled video data to reduce the gap between different domains while transferring discriminative knowledge from large-scale labeled still images. Given a face recognition network that is pretrained in the image domain, the adaptation is achieved by (i) distilling knowledge from the network to a video adaptation network through feature matching, (ii) performing feature restoration through synthetic data augmentation and (iii) learning a domain-invariant feature through a domain adversarial discriminator. We further improve performance through a discriminator-guided feature fusion that boosts high-quality frames while eliminating those degraded by video domain-specific factors. Experiments on the YouTube Faces and IJB-A datasets demonstrate that each module contributes to our feature-level domain adaptation framework and substantially improves video face recognition performance to achieve state-of-the-art accuracy. We demonstrate qualitatively that the network learns to suppress diverse artifacts in videos such as pose, illumination or occlusion without being explicitly trained for them. Kihyuk Sohn, Sifei Liu, Guangyu Zhong, Xiang Yu 0002, Ming-Hsuan Yang 0001, Manmohan Krishna Chandraker |
ICCV | 4 |
| 2017 | Towards Large-Pose Face Frontalization in the Wild
Xi Yin 0001, Xiang Yu 0002, Kihyuk Sohn, Xiaoming Liu 0002, Manmohan Krishna Chandraker |
ICCV | 2 |
| 2017 | Learning Efficient Object Detection Models with Knowledge DistillationabstractDespite significant accuracy improvement in convolutional neural networks (CNN) based object detectors, they often require prohibitive runtimes to process an image for real-time applications. State-of-the-art models often use very deep networks with a large number of floating point operations. Efforts such as model compression learn compact models with fewer number of parameters, but with much reduced accuracy. In this work, we propose a new framework to learn compact and fast ob- ject detection networks with improved accuracy using knowledge distillation [20] and hint learning [34]. Although knowledge distillation has demonstrated excellent improvements for simpler classification setups, the complexity of detection poses new challenges in the form of regression, region proposals and less voluminous la- bels. We address this through several innovations such as a weighted cross-entropy loss to address class imbalance, a teacher bounded loss to handle the regression component and adaptation layers to better learn from intermediate teacher distribu- tions. We conduct comprehensive empirical evaluation with different distillation configurations over multiple datasets including PASCAL, KITTI, ILSVRC and MS-COCO. Our results show consistent improvement in accuracy-speed trade-offs for modern multi-class detection models. Guobin Chen, Wongun Choi, Xiang Yu 0002, Tony X. Han, Manmohan Krishna Chandraker |
NIPS | 3 |
| 2016 | Deep Deformation Network for Object Landmark Localization
Xiang Yu 0002, Manmohan Krishna Chandraker |
ECCV (5) | 1 |
| 2016 | Nonlinear Hierarchical Part-Based Regression for Unconstrained Face Alignment
Xiang Yu 0002, Zhe Lin 0001, Shaoting Zhang 0001, Dimitris N. Metaxas |
IJCAI | 1 |
| 2016 | Customized expression recognition for performance-driven cutout character animationabstractPerformance-driven character animation enables users to create expressive results by performing the desired motion of the character with their face and/or body. However, for cutout animations where continuous motion is combined with discrete artwork replacements, supporting a performance-driven workflow has some unique requirements. To trigger the appropriate artwork replacements, the system must reliably detect a wide range of customized facial expressions that are challenging for existing recognition methods, which focus on a few canonical expressions (e.g., angry, disgusted, scared, happy, sad and surprised). Also, real usage scenarios require the system to work in realtime with minimal training. In this paper, we propose a novel customized expression recognition technique that meets all of these requirements. We first use a set of handcrafted features combining geometric features derived from facial landmarks and patch-based appearance features through group sparsity-based facial component learning. To improve discrimination and generalization, these handcrafted features are integrated into a custom-designed Deep Convolutional Neural Network (CNN) structure trained from publicly available facial expression datasets. The combined features are fed to an online ensemble of SVMs designed for the few training sample problem and performs in realtime. To improve temporal coherence, we also apply a Hidden Markov Model (HMM) to smooth the recognition results. Our system achieves state-of-the-art performance on canonical expression datasets and promising results on our collected dataset of customized expressions. Xiang Yu 0002, Jianchao Yang, Linjie Luo, Wilmot Li, Jonathan Brandt, Dimitris N. Metaxas |
WACV | 1 |
| 2016 | Face Landmark Fitting via Optimized Part Mixtures and Cascaded Deformable ModelabstractThis paper addresses the problem of facial landmark localization and tracking from a single camera. We present a two-stage cascaded deformable shape model to effectively and efficiently localize facial landmarks with large head pose variations. In initialization stage, we propose a group sparse optimized mixture model to automatically select the most salient facial landmarks. By introducing 3D face shape model, we apply procrustes analysis to provide pose-aware landmark initialization. In landmark localization stage, the first step uses mean-shift local search with constrained local model to rapidly approach the global optimum. The second step uses component-wise active contours to discriminatively refine the subtle shape variation. Our framework simultaneously handles face detection, pose-robust landmark localization and tracking in real time. Extensive experiments are conducted on both laboratory environmental databases and face-in-the-wild databases. The results reveal that our approach consistently outperforms state-of-the-art methods for face alignment and tracking. Xiang Yu 0002, Junzhou Huang, Shaoting Zhang 0001, Dimitris N. Metaxas |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Is Interactional Dissynchrony a Clue to Deception? Insights From Automated Analysis of Nonverbal Visual CuesabstractDetecting deception in interpersonal dialog is challenging since deceivers take advantage of the give-and-take of interaction to adapt to any sign of skepticism in an interlocutor's verbal and nonverbal feedback. Human detection accuracy is poor, often with no better than chance performance. In this investigation, we consider whether automated methods can produce better results and if emphasizing the possible disruption in interactional synchrony can signal whether an interactant is truthful or deceptive. We propose a data-driven and unobtrusive framework using visual cues that consists of face tracking, head movement detection, facial expression recognition, and interactional synchrony estimation. Analysis were conducted on 242 video samples from an experiment in which deceivers and truth-tellers interacted with professional interviewers either face-to-face or through computer mediation. Results revealed that the framework is able to automatically track head movements and expressions of both interlocutors to extract normalized meaningful synchrony features and to learn classification models for deception recognition. Further experiments show that these features reliably capture interactional synchrony and efficiently discriminate deception from truth. Xiang Yu 0002, Shaoting Zhang 0001, Zhennan Yan, Fei Yang 0001, Junzhou Huang, Norah E. Dunbar, Matthew L. Jensen, Judee K. Burgoon, Dimitris N. Metaxas |
IEEE Trans. Cybern. | 1 |
| 2014 | Consensus of Regression for Occlusion-Robust Facial Feature Localization
Xiang Yu 0002, Zhe Lin 0001, Jonathan Brandt, Dimitris N. Metaxas |
ECCV (4) | 1 |
| 2014 | 3D Face Tracking and Multi-Scale, Spatio-temporal Analysis of Linguistically Significant Facial Expressions and Head Positions in ASL
Bo Liu 0005, Jingjing Liu 0001, Xiang Yu 0002, Dimitris N. Metaxas, Carol Neidle |
LREC | 3 |
| 2013 | Pose-Free Facial Landmark Fitting via Optimized Part Mixtures and Cascaded Deformable Shape ModelabstractThis paper addresses the problem of facial landmark localization and tracking from a single camera. We present a two-stage cascaded deformable shape model to effectively and efficiently localize facial landmarks with large head pose variations. For face detection, we propose a group sparse learning method to automatically select the most salient facial landmarks. By introducing 3D face shape model, we use procrustes analysis to achieve pose-free facial landmark initialization. For deformation, the first step uses mean-shift local search with constrained local model to rapidly approach the global optimum. The second step uses component-wise active contours to discriminatively refine the subtle shape variation. Our framework can simultaneously handle face detection, pose-free landmark localization and tracking in real time. Extensive experiments are conducted on both laboratory environmental face databases and face-in-the-wild databases. All results demonstrate that our approach has certain advantages over state-of-the-art methods in handling pose variations. Xiang Yu 0002, Junzhou Huang, Shaoting Zhang 0001, Wang Yan, Dimitris N. Metaxas |
ICCV | 1 |
| 2012 | Robust face tracking with a consumer depth cameraabstractWe address the problem of tracking human faces under various poses and lighting conditions. Reliable face tracking is a challenging task. The shapes of the faces may change dramatically with various identities, poses and expressions. Moreover, poor lighting conditions may cause a low contrast image or cast shadows on faces, which will significantly degrade the performance of the tracking system. In this paper, we develop a framework to track face shapes by using both color and depth information. Since the faces in various poses lie on a nonlinear manifold, we build piecewise linear face models, each model covering a range of poses. The low-resolution depth image is captured by using Microsoft Kinect, and is used to predict head pose and generate extra constraints at the face boundary. Our experiments show that, by exploiting the depth information, the performance of the tracking system is significantly improved. Fei Yang 0001, Junzhou Huang, Xiang Yu 0002, Xinyi Cui, Dimitris N. Metaxas |
ICIP | 3 |
| 2012 | Robust eyelid tracking for fatigue detectionabstractWe develop a non-intrusive system for monitoring fatigue by tracking eyelids with a single web camera. Tracking slow eyelid closures is one of the most reliable ways to monitor fatigue during critical performance tasks. The challenges come from arbitrary head movement, occlusion, reflection of glasses, motion blurs, etc. We model the shape of eyes using a pair of parameterized parabolic curves, and fit the model in each frame to maximize the total likelihood of the eye regions. Our system is able to track face movement and fit eyelids reliably in real time. We test our system with videos captured from both alert and drowsy subjects. The experiment results prove the effectiveness of our system. Fei Yang 0001, Xiang Yu 0002, Junzhou Huang, Peng Yang 0001, Dimitris N. Metaxas |
ICIP | 2 |