Xi Peng 0005

dblp:149/7762-5 · DBLP profile ↗
← Back
55ranked-venue papers
10as first author
26since 2021 · last 2025
0000-0002-7772-001XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 10 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 30 · 6 first-author · 11 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Beyond Accuracy: On the Effects of Fine-Tuning Towards Vision-Language Model's Prediction Rationality
abstract
Vision-Language Models (VLMs), such as CLIP, have already seen widespread applications. Researchers actively engage in further fine-tuning VLMs in safety-critical domains. In these domains, prediction rationality is crucial: the prediction should be correct and based on valid evidence. Yet, for VLMs, the impact of fine-tuning on prediction rationality is seldomly investigated. To study this problem, we proposed two new metrics called Prediction Trustworthiness and Inference Reliability. We conducted extensive experiments on various settings and observed some interesting phenomena. On the one hand, we found that the well-adopted fine-tuning methods led to more correct predictions based on invalid evidence. This potentially undermines the trustworthiness of correct predictions from fine-tuned VLMs. On the other hand, having identified valid evidence of target objects, fine-tuned VLMs were more likely to make correct predictions. Moreover, the findings are also consistent under distributional shifts and across various experimental settings. We hope our research offer fresh insights to VLM fine-tuning.
Qitong Wang 0001, Tang Li 0005, Kien X. Nguyen 0001, Xi Peng 0005
AAAI4
2025 Interpretable Failure Detection with Human-Level Concepts
abstract
Reliable failure detection holds paramount importance in safety-critical applications. Yet, neural networks are known to produce overconfident predictions for misclassified samples. As a result, it remains a problematic matter as existing confidence score functions rely on category-level signals, the logits, to detect failures. This research introduces an innovative strategy, leveraging human-level concepts for a dual purpose: to reliably detect when a model fails and to transparently interpret why. By integrating a nuanced array of signals for each category, our method enables a finer-grained assessment of the model's confidence. We present a simple yet highly effective approach based on the ordinal ranking of concept activation to the input image. Without bells and whistles, our method is able to significantly reduce the false positive rate across diverse real-world image classification benchmarks, specifically by 3.7% on ImageNet and 9.0% on EuroSAT.
Kien X. Nguyen 0001, Tang Li 0005, Xi Peng 0005
AAAI3
2025 "Why Is There a Tumor?": Tell Me the Reason, Show Me the Evidence
abstract
Medical AI models excel at tumor detection and segmentation. However, their latent representations often lack explicit ties to clinical semantics, producing outputs less trusted in clinical practice. Most of the existing models generate either segmentation masks/labels (localizing where without why) or textual justifications (explaining why without where), failing to ground clinical concepts in spatially localized evidence. To bridge this gap, we propose to develop models that can justify the segmentation or detection using clinically relevant terms and point to visual evidence. We address two core challenges: First, we curate a rationale dataset to tackle the lack of paired images, annotations, and textual rationales for training. The dataset includes 180K image-mask-rationale triples with quality evaluated by expert radiologists. Second, we design rationale-informed optimization that disentangles and localizes fine-grained clinical concepts in a self-supervised manner without requiring pixel-level concept annotations. Experiments across medical benchmarks show our model demonstrates superior performance in segmentation, detection, and beyond. The anonymous link to our code.
Mengmeng Ma 0002, Tang Li 0005, Yunxiang Peng 0002, Volkan Beylergil, Binsheng Zhao, Oguz Akin, Xi Peng 0005
ICML8
2025 Structure-informed Risk Minimization for Robust Ensemble Learning
abstract
Ensemble learning is a powerful approach for improving generalization under distribution shifts, but its effectiveness heavily depends on how individual models are combined. Existing methods often optimize ensemble weights based on validation data, which may not represent unseen test distributions, leading to suboptimal performance in out-of-distribution (OoD) settings. Inspired by Distributionally Robust Optimization (DRO), we propose Structure-informed Risk Minimization (SRM), a principled framework that learns robust ensemble weights without access to test data. Unlike standard DRO, which defines uncertainty sets based on divergence metrics alone, SRM incorporates structural information of training distributions, ensuring that the uncertainty set aligns with plausible real-world shifts. This approach mitigates the over-pessimism of traditional worst-case optimization while maintaining robustness. We introduce a computationally efficient optimization algorithm with theoretical guarantees and demonstrate that SRM achieves superior OoD generalization compared to existing ensemble combination strategies across diverse benchmarks. Code is available at: https://github.com/deep-real/SRM.
Fengchun Qiao, Xi Peng 0005
ICML3
2024 Adaptive Cascading Network for Continual Test-Time Adaptation
Kien X. Nguyen 0001, Fengchun Qiao, Xi Peng 0005
CIKM3
2024 DEAL: Disentangle and Localize Concept-Level Explanations for VLMs
Tang Li 0005, Mengmeng Ma 0002, Xi Peng 0005
ECCV (39)3
2024 Beyond the Federation: Topology-aware Federated Learning for Generalization to Unseen Clients
abstract
Federated Learning is widely employed to tackle distributed sensitive data. Existing methods primarily focus on addressing in-federation data heterogeneity. However, we observed that they suffer from significant performance degradation when applied to unseen clients for out-of-federation (OOF) generalization. The recent attempts to address generalization to unseen clients generally struggle to scale up to large-scale distributed settings due to high communication or computation costs. Moreover, methods that scale well often demonstrate poor generalization capability. To achieve OOF-resiliency in a scalable manner, we propose Topology-aware Federated Learning (TFL) that leverages client topology - a graph representing client relationships - to effectively train robust models against OOF data. We formulate a novel optimization problem for TFL, consisting of two key modules: Client Topology Learning, which infers the client relationships in a privacy-preserving manner, and Learning on Client Topology, which leverages the learned topology to identify influential clients and harness this information into the FL optimization process to efficiently build robust models. Empirical evaluation on a variety of real-world datasets verifies TFL's superior OOF robustness and scalability.
Mengmeng Ma 0002, Tang Li 0005, Xi Peng 0005
ICML3
2024 Ensemble Pruning for Out-of-distribution Generalization
abstract
Ensemble of deep neural networks has achieved great success in hedging against single-model failure under distribution shift. However, existing techniques suffer from producing redundant models, limiting predictive diversity and yielding compromised generalization performance. Existing ensemble pruning methods can only guarantee predictive diversity for in-distribution data, which may not transfer well to out-of-distribution (OoD) data. To address this gap, we propose a principled optimization framework for ensemble pruning under distribution shifts. Since the annotations of test data are not available, we explore relationships between prediction distributions of the models, encapsulated in a topology graph. By incorporating this topology into a combinatorial optimization framework, complementary models with high predictive diversity are selected with theoretical guarantees. Our approach is model-agnostic and can be applied on top of a broad spectrum of off-the-shelf ensembling methods for improved generalization performance. Experiments on common benchmarks demonstrate the superiority of our approach in both multi- and single-source OoD generalization. The source codes are publicly available at: https://github.com/joffery/TEP.
Fengchun Qiao, Xi Peng 0005
ICML2
2024 SeafloorAI: A Large-scale Vision-Language Dataset for Seafloor Geological Survey
abstract
A major obstacle to the advancements of machine learning models in marine science, particularly in sonar imagery analysis, is the scarcity of AI-ready datasets. While there have been efforts to make AI-ready sonar image dataset publicly available, they suffer from limitations in terms of environment setting and scale. To bridge this gap, we introduce $\texttt{SeafloorAI}$, the first extensive AI-ready datasets for seafloor mapping across 5 geological layers that is curated in collaboration with marine scientists. We further extend the dataset to $\texttt{SeafloorGenAI}$ by incorporating the language component in order to facilitate the development of both $\textit{vision}$- and $\textit{language}$-capable machine learning models for sonar imagery. The dataset consists of 62 geo-distributed data surveys spanning 17,300 square kilometers, with 696K sonar images, 827K annotated segmentation masks, 696K detailed language descriptions and approximately 7M question-answer pairs. By making our data processing source code publicly available, we aim to engage the marine science community to enrich the data pool and inspire the machine learning community to develop more robust models. This collaborative approach will enhance the capabilities and applications of our datasets within both fields.
Kien X. Nguyen 0001, Fengchun Qiao, Arthur Trembanis, Xi Peng 0005
NeurIPS4
2024 Beyond Accuracy: Ensuring Correct Predictions With Correct Rationales
abstract
Large pretrained foundation models demonstrate exceptional performance and, in some high-stakes applications, even surpass human experts. However, most of these models are currently evaluated primarily on prediction accuracy, overlooking the validity of the rationales behind their accurate predictions. For the safe deployment of foundation models, there is a pressing need to ensure *double-correct predictions*, *i.e.*, correct prediction backed by correct rationales. To achieve this, we propose a two-phase scheme: First, we curate a new dataset that offers structured rationales for visual recognition tasks. Second, we propose a rationale-informed optimization method to guide the model in disentangling and localizing visual evidence for each rationale, without requiring manual annotations. Extensive experiments and ablation studies demonstrate that our model outperforms state-of-the-art models by up to 10.1\% in prediction accuracy across a wide range of tasks. Furthermore, our method significantly improves the model's rationale correctness, improving localization by 7.5\% and disentanglement by 36.5\%. Our dataset, source code, and pretrained weights: https://github.com/deep-real/DCP
Tang Li 0005, Mengmeng Ma 0002, Xi Peng 0005
NeurIPS3
2024 Out-of-Domain Generalization From a Single Source: An Uncertainty Quantification Approach
abstract
We are concerned with a worst-case scenario in model generalization, in the sense that a model aims to perform well on many unseen domains while there is only one single domain available for training. We propose Meta-Learning based Adversarial Domain Augmentation to solve this Out-of-Domain generalization problem. The key idea is to leverage adversarial training to create "fictitious" yet "challenging" populations, from which a model can learn to generalize with theoretical guarantees. To facilitate fast and desirable domain augmentation, we cast the model training in a meta-learning scheme and use a Wasserstein Auto-Encoder to relax the widely used worst-case constraint. We further improve our method by integrating uncertainty quantification for efficient domain generalization. Extensive experiments on multiple benchmark datasets indicate its superior performance in tackling single domain generalization.
Xi Peng 0005, Fengchun Qiao, Long Zhao 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Semi-Identical Twins Variational AutoEncoder for Few-Shot Learning
abstract
Data augmentation is a popular way for few-shot learning (FSL). It generates more samples as supplements and then transforms the FSL task into a common supervised learning problem for a solution. However, most data-augmentation-based FSL approaches only consider the prior visual knowledge for feature generation, thereby leading to low diversity and poor quality of generated data. In this study, we attempt to address this issue by incorporating both prior visual and prior semantic knowledge to condition the feature generation process. Inspired by some genetic characteristics of semi-identical twins, a novel multimodal generative FSL approach was developed named semi-identical twins variational autoencoder (STVAE) to better exploit the complementarity of these modality information by considering the multimodal conditional feature generation process as a process that semi-identical twins are born and collaborate to simulate their father. STVAE conducts feature synthesis by pairing two conditional variational autoencoders (CVAEs) with the same seed but different modality conditions. Subsequently, the generated features of two CVAEs are considered as semi-identical twins and adaptively combined to yield the final feature, which is considered as their fake father. STVAE requires that the final feature can be converted back into its paired conditions while ensuring these conditions remain consistent with the original in both representation and function. Moreover, STVAE is able to work in the partial modality-absence case due to the adaptive linear feature combination strategy. STVAE essentially provides a novel idea to exploit the complementarity of different modality prior information inspired by genetics in FSL. Extensive experimental results demonstrate that our work achieves promising performances in comparison to the recent state-of-the-art approaches, as well as validate its effectiveness on FSL under various modality settings.
Yi Zhang 0113, Sheng Huang 0001, Xi Peng 0005, Dan Yang 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Adapting Wireless Network Configuration From Simulation to Reality via Deep Learning-Based Domain Adaptation
abstract
Today, wireless mesh networks (WMNs) are deployed globally to support various applications, such as industrial automation, military operations, and smart energy. Significant efforts have been made in the literature to facilitate their deployments and optimize their performance. However, configuring a WMN well is challenging because the network configuration is a complex process, which involves theoretical computation, simulation, and field testing, among other tasks. Our study shows that the models for network configuration prediction learned from simulations may not work well in physical networks because of the simulation-to-reality gap. In this paper, we employ deep learning-based domain adaptation to close the gap and leverage a teacher-student neural network and a physical sampling method to transfer the network configuration knowledge learned from a simulated network to its corresponding physical network. Experimental results show that our method effectively closes the gap and increases the accuracy of predicting a good network configuration that allows the network to meet performance requirements from 30.10% to 70.24% by learning robust machine learning models from a large amount of inexpensive simulation data and a few costly field testing measurements.
Junyang Shi, Aitian Ma, Xia Cheng, Mo Sha 0001, Xi Peng 0005
IEEE/ACM Trans. Netw.5
2023 Are Data-Driven Explanations Robust Against Out-of-Distribution Data?
abstract
As black-box models increasingly power high-stakes applications, a variety of data-driven explanation methods have been introduced. Meanwhile, machine learning models are constantly challenged by distributional shifts. A question naturally arises: Are data-driven explanations robust against out-of-distribution data? Our empirical results show that even though predict correctly, the model might still yield unreliable explanations under distributional shifts. How to develop robust explanations against out-of-distribution data? To address this problem, we propose an end-to-end model-agnostic learning framework Distributionally Robust Explanations (DRE). The key idea is, inspired by self-supervised learning, to fully utilizes the inter-distribution information to provide supervisory signals for the learning of explanations without human annotation. Can robust explanations benefit the model's generalization capability? We conduct extensive experiments on a wide range of tasks and data types, including classification and regression on image and scientific tabular data. Our results demonstrate that the proposed method significantly improves the model's performance in terms of explanation and prediction robustness against distributional shifts.
Tang Li 0005, Fengchun Qiao, Mengmeng Ma 0002, Xi Peng 0005
CVPR4
2023 Learning from Semantic Alignment between Unpaired Multiviews for Egocentric Video Recognition
abstract
We are concerned with a challenging scenario in unpaired multiview video learning. In this case, the model aims to learn comprehensive multiview representations while the cross-view semantic information exhibits variations. We propose Semantics-based Unpaired Multiview Learning (SUM-L) to tackle this unpaired multiview learning problem. The key idea is to build cross-view pseudopairs and do view-invariant alignment by leveraging the semantic information of videos. To facilitate the data efficiency of multiview learning, we further perform video-text alignment for first-person and third-person videos, to fully leverage the semantic knowledge to improve video representations. Extensive experiments on multiple benchmark datasets verify the effectiveness of our framework. Our method also outperforms multiple existing view-alignment methods, under the more challenging scenario than typical paired or unpaired multimodal or multiview learning. Our code is available at https://github.com/wqtwjt1996/SUM-L.
Qitong Wang 0001, Long Zhao 0003, Liangzhe Yuan, Ting Liu 0005, Xi Peng 0005
ICCV5
2023 Topology-aware Robust Optimization for Out-of-Distribution Generalization
Fengchun Qiao, Xi Peng 0005
ICLR2
2023 Deep learning-based estimation of whole-body kinematics from multi-view images
Kien X. Nguyen 0001, Liying Zheng, Ashley L. Hawke, Robert E. Carey, Scott P. Breloff, Kang Li 0004, Xi Peng 0005
Comput. Vis. Image Underst.7
2023 Region-Aware Arbitrary-Shaped Text Detection With Progressive Fusion
abstract
Segmentation-based text detectors are flexible to capture arbitrary-shaped text regions. Due to large geometry variance, it is necessary to construct effective and robust representations to identify text regions with various shapes and scales. In this paper, we focus on designing effective multi-scale contextual features for locating text instances. Specially, we develop a Region Context Module (RCM) to summarize the semantic response and adaptively extract text-region-aware information in a limited local area. To construct complementary multi-scale contextual representations, multiple RCM branches with different scales are employed and integrated via Progressive Fusion Module (PFM). Our proposed RCM and PFM serve as the plug-and-play modules which can be incorporated into existing scene text detection platforms to further boost detection performance. Extensive experiments show that our methods achieve state-of-the-art performances on Total-Text, SCUT-CTW1500 and MSRA-TD500 datasets. The code with models will become publicly available athttps://github.com/wqtwjt1996/RP-Text.
Qitong Wang 0001, Ming Li 0010, Junjun He, Xi Peng 0005, Yu Qiao 0001
IEEE Trans. Multim.5
2022 Are Multimodal Transformers Robust to Missing Modality?
abstract
Multimodal data collected from the real world are often imperfect due to missing modalities. Therefore multimodal models that are robust against modal-incomplete data are highly preferred. Recently, Transformer models have shown great success in processing multimodal data. However, existing work has been limited to either architecture designs or pre-training strategies; whether Transformer models are naturally robust against missing-modal data has rarely been investigated. In this paper, we present the first-of-its-kind work to comprehensively investigate the behavior of Transformers in the presence of modal-incomplete data. Unsurprising, we find Transformer models are sensitive to missing modalities while different modal fusion strategies will significantly affect the robustness. What surprised us is that the optimal fusion strategy is dataset dependent even for the same Transformer model; there does not exist a universal strategy that works in general cases. Based on these findings, we propose a principle method to improve the robustness of Transformer models byautomatically searching for an optimal fusion strategy regarding input data. Experimental validations on three benchmarks support the superior performance of the proposed method.
Mengmeng Ma 0002, Jian Ren 0005, Long Zhao 0003, Davide Testuggine, Xi Peng 0005
CVPR5
2022 Self-Guidance: Improve Deep Neural Network Generalization via Knowledge Distillation
abstract
We present Self-Guidance, a simple way to train deep neural networks via knowledge distillation. The basic idea is to train sub-network to match the prediction of the full network, so-called "Self-Guidance". Under the "teacher-student" framework, we construct both teacher and student within the same target network. Student network is the sub-networks that randomly skip some portions of the full network. The teacher network is the full network, can be considered as the ensemble of all possible student networks. The training process is performed in a closed-loop: (1) Forward prediction contains two passes that generate student and teacher predictions. (2) Backward distillation allows knowledge transfer from the teacher back to students. Comprehensive evaluations show that our approach improves the generalization ability of deep neural networks to a significant margin. The results prove our superior performance in both image classification on CIFAR10, CIFAR100, and facial expression recognition on FER-2013 and RAF.
Zhenzhu Zheng, Xi Peng 0005
WACV2
2021 SMIL: Multimodal Learning with Severely Missing Modality
abstract
A common assumption in multimodal learning is the completeness of training data, i.e., full modalities are available in all training examples. Although there exists research endeavor in developing novel methods to tackle the incompleteness of testing data, e.g., modalities are partially missing in testing examples, few of them can handle incomplete training modalities. The problem becomes even more challenging if considering the case of severely missing, e.g., ninety percent of training examples may have incomplete modalities. For the first time in the literature, this paper formally studies multimodal learning with missing modality in terms of flexibility (missing modalities in training, testing, or both) and efficiency (most training data have incomplete modality). Technically, we propose a new method named SMIL that leverages Bayesian meta-learning in uniformly achieving both objectives. To validate our idea, we conduct a series of experiments on three popular benchmarks: MM-IMDb, CMU-MOSI, and avMNIST. The results prove the state-of-the-art performance of SMIL over existing methods and generative baselines including autoencoders and generative adversarial networks.
Mengmeng Ma 0002, Jian Ren 0005, Long Zhao 0003, Sergey Tulyakov, Cathy H. Wu, Xi Peng 0005
AAAI6
2021 Understanding the factors related to the opioid epidemic using machine learning
abstract
In recent years, the US has experienced an opioid epidemic with an unprecedented number of drugs overdose deaths. Research finds such overdose deaths are linked to neighborhood-level traits, thus providing opportunity to identify effective interventions. Typically, techniques such as Ordinary Least Squares (OLS) or Maximum Likelihood Estimation (MLE) are used to document neighborhood-level factors significant in explaining such adverse outcomes. These techniques are, however, less equipped to ascertain non-linear relationships between confounding factors. Hence, in this study we apply machine learning based techniques to identify opioid risks of neighborhoods in Delaware and explore the correlation of these factors using Shapley Additive explanations (SHAP). We discovered that the factors related to neighborhoods’ environment, followed by education and then crime, were highly correlated with higher opioid risk. We also explored the change in these correlations over the years to understand the changing dynamics of the epidemic. Furthermore, we discovered that, as the epidemic has shifted from legal (i.e., prescription opioids) to illegal (e.g., heroin and fentanyl) drugs in recent years, the correlation of environment, crime and health related variables with the opioid risk has increased significantly while the correlation of economic and socio-demographic variables has decreased. The correlation of education related factors has been higher from the start and has increased slightly in recent years suggesting a need for increased awareness about the opioid epidemic.
Sachin Gavali, Chuming Chen, Julie Cowart, Xi Peng 0005, Cathy H. Wu, Tammy Anderson
BIBM4
2021 Learning View-Disentangled Human Pose Representation by Contrastive Cross-View Mutual Information Maximization
abstract
We introduce a novel representation learning method to disentangle pose-dependent as well as view-dependent factors from 2D human poses. The method trains a network using cross-view mutual information maximization (CV-MIM) which maximizes mutual information of the same pose performed from different viewpoints in a contrastive learning manner. We further propose two regularization terms to ensure disentanglement and smoothness of the learned representations. The resulting pose representations can be used for cross-view action recognition.To evaluate the power of the learned representations, in addition to the conventional fully-supervised action recognition settings, we introduce a novel task called single-shot cross-view action recognition. This task trains models with actions from only one single viewpoint while models are evaluated on poses captured from all possible viewpoints. We evaluate the learned representations on standard benchmarks for action recognition, and show that (i) CV-MIM performs competitively compared with the state-of-the-art models in the fully-supervised scenarios; (ii) CV-MIM outperforms other competing methods by a large margin in the single-shot cross-view setting; (iii) and the learned representations can significantly boost the performance when reducing the amount of supervised training data. Our code is made publicly available at https://github.com/google-research/google-research/tree/master/poem.
Long Zhao 0003, Yuxiao Wang 0001, Jiaping Zhao, Liangzhe Yuan, Jennifer J. Sun, Florian Schroff, Hartwig Adam, Xi Peng 0005, Dimitris N. Metaxas, Ting Liu 0005
CVPR8
2021 Uncertainty-Guided Model Generalization to Unseen Domains
abstract
We study a worst-case scenario in generalization: Out-of-domain generalization from a single source. The goal is to learn a robust model from a single source and expect it to generalize over many unknown distributions. This challenging problem has been seldom investigated while existing solutions suffer from various limitations. In this paper, we propose a new solution. The key idea is to augment the source capacity in both input and label spaces, while the augmentation is guided by uncertainty assessment. To the best of our knowledge, this is the first work to (1) access the generalization uncertainty from a single source and (2) leverage it to guide both input and label augmentation for robust generalization. The model training and deployment are effectively organized in a Bayesian meta-learning framework. We conduct extensive comparisons and ablation study to validate our approach. The results prove our superior performance in a wide scope of tasks including image classification, semantic segmentation, text classification, and speech recognition.
Fengchun Qiao, Xi Peng 0005
CVPR2
2021 A Good Image Generator Is What You Need for High-Resolution Video Synthesis
Yu Tian 0003, Jian Ren 0005, Menglei Chai, Kyle Olszewski, Xi Peng 0005, Dimitris N. Metaxas, Sergey Tulyakov
ICLR5
2021 Adapting Wireless Mesh Network Configuration from Simulation to Reality via Deep Learning based Domain Adaptation
Junyang Shi, Mo Sha 0001, Xi Peng 0005
NSDI3
2020 Knowledge As Priors: Cross-Modal Knowledge Generalization for Datasets Without Superior Knowledge
abstract
Cross-modal knowledge distillation deals with transferring knowledge from a model trained with superior modalities (Teacher) to another model trained with weak modalities (Student). Existing approaches require paired training examples exist in both modalities. However, accessing the data from superior modalities may not always be feasible. For example, in the case of 3D hand pose estimation, depth maps, point clouds, or stereo images usually capture better hand structures than RGB images, but most of them are expensive to be collected. In this paper, we propose a novel scheme to train the Student in a Target dataset where the Teacher is unavailable. Our key idea is to generalize the distilled cross-modal knowledge learned from a Source dataset, which contains paired examples from both modalities, to the Target dataset by modeling knowledge as priors on parameters of the Student. We name our method "Cross-Modal Knowledge Generalization" and demonstrate that our scheme results in competitive performance for 3D hand pose estimation on standard benchmark datasets.
Long Zhao 0003, Xi Peng 0005, Yuxiao Chen 0002, Mubbasir Kapadia, Dimitris N. Metaxas
CVPR2
2020 Learning to Learn Single Domain Generalization
abstract
We are concerned with a worst-case scenario in model generalization, in the sense that a model aims to perform well on many unseen domains while there is only one single domain available for training. We propose a new method named adversarial domain augmentation to solve this Out-of-Distribution (OOD) generalization problem. The key idea is to leverage adversarial training to create "fictitious" yet "challenging" populations, from which a model can learn to generalize with theoretical guarantees. To facilitate fast and desirable domain augmentation, we cast the model training in a meta-learning scheme and use a Wasserstein Auto-Encoder (WAE) to relax the widely used worst-case constraint. Detailed theoretical analysis is provided to testify our formulation, while extensive experiments on multiple benchmark datasets indicate its superior performance in tackling single domain generalization.
Fengchun Qiao, Long Zhao 0003, Xi Peng 0005
CVPR3
2020 Maximum-Entropy Adversarial Data Augmentation for Improved Generalization and Robustness
abstract
Adversarial data augmentation has shown promise for training robust deep neural networks against unforeseen data shifts or corruptions. However, it is difficult to define heuristics to generate effective fictitious target distributions containing "hard" adversarial perturbations that are largely different from the source distribution. In this paper, we propose a novel and effective regularization term for adversarial data augmentation. We theoretically derive it from the information bottleneck principle, which results in a maximum-entropy formulation. Intuitively, this regularization term encourages perturbing the underlying source distribution to enlarge predictive uncertainty of the current model, so that the generated "hard" adversarial perturbations can improve the model robustness during training. Experimental results on three standard benchmarks demonstrate that our method consistently outperforms the existing state of the art by a statistically significant margin.
Long Zhao 0003, Ting Liu 0005, Xi Peng 0005, Dimitris N. Metaxas
NeurIPS3
2020 Towards Image-to-Video Translation: A Structure-Aware Approach via Multi-stage Generative Adversarial Networks
Long Zhao 0003, Xi Peng 0005, Yu Tian 0003, Mubbasir Kapadia, Dimitris N. Metaxas
Int. J. Comput. Vis.2
2020 Towards Efficient U-Nets: A Coupled and Quantized Approach
abstract
In this paper, we propose to couple stacked U-Nets for efficient visual landmark localization. The key idea is to globally reuse features of the same semantic meanings across the stacked U-Nets. The feature reuse makes each U-Net light-weighted. Specially, we propose an order- K coupling design to trim off long-distance shortcuts, together with an iterative refinement and memory sharing mechanism. To further improve the efficiency, we quantize the parameters, intermediate features, and gradients of the coupled U-Nets to low bit-width numbers. We validate our approach in two tasks: human pose estimation and facial landmark localization. The results show that our approach achieves state-of-the-art localization accuracy but using ∼ 70% fewer parameters, ∼ 30% less inference time, ∼ 98% less model size, and saving ∼ 75% training memory compared with benchmark localizers.
Zhiqiang Tang 0001, Xi Peng 0005, Kang Li 0004, Dimitris N. Metaxas
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Construct Dynamic Graphs for Hand Gesture Recognition via Spatial-Temporal Attention
Yuxiao Chen 0002, Long Zhao 0003, Xi Peng 0005, Dimitris N. Metaxas
BMVC3
2019 Semantic Graph Convolutional Networks for 3D Human Pose Regression
abstract
In this paper, we study the problem of learning Graph Convolutional Networks (GCNs) for regression. Current architectures of GCNs are limited to the small receptive field of convolution filters and shared transformation matrix for each node. To address these limitations, we propose Semantic Graph Convolutional Networks (SemGCN), a novel neural network architecture that operates on regression tasks with graph-structured data. SemGCN learns to capture semantic information such as local and global node relationships, which is not explicitly represented in the graph. These semantic relationships can be learned through end-to-end training from the ground truth without additional supervision or hand-crafted rules. We further investigate applying SemGCN to 3D human pose regression. Our formulation is intuitive and sufficient since both 2D and 3D human poses can be represented as a structured graph encoding the relationships between joints in the skeleton of a human body. We carry out comprehensive studies to validate our method. The results prove that SemGCN outperforms state of the art while using 90% fewer parameters.
Long Zhao 0003, Xi Peng 0005, Yu Tian 0003, Mubbasir Kapadia, Dimitris N. Metaxas
CVPR2
2019 AdaTransform: Adaptive Data Transformation
abstract
Data augmentation is widely used to increase data variance in training deep neural networks. However, previous methods require either comprehensive domain knowledge or high computational cost. Can we learn data transformation automatically and efficiently with limited domain knowledge? Furthermore, can we leverage data transformation to improve not only network training but also network testing? In this work, we propose adaptive data transformation to achieve the two goals. The AdaTransform can increase data variance in training and decrease data variance in testing. Experiments on different tasks prove that it can improve generalization performance.
Zhiqiang Tang 0001, Xi Peng 0005, Tingfeng Li, Yizhe Zhu, Dimitris N. Metaxas
ICCV2
2019 Scalable Global Alignment Graph Kernel Using Random Features: From Node Embedding to Graph Embedding
abstract
Graph kernels are widely used for measuring the similarity between graphs. Many existing graph kernels, which focus on local patterns within graphs rather than their global properties, suffer from significant structure information loss when representing graphs. Some recent global graph kernels, which utilizes the alignment of geometric node embeddings of graphs, yield state-of-the-art performance. However, these graph kernels are not necessarily positive-definite. More importantly, computing the graph kernel matrix will have at least quadratic time complexity in terms of the number and the size of the graphs. In this paper, we propose a new family of global alignment graph kernels, which take into account the global properties of graphs by using geometric node embeddings and an associated node transportation based on earth mover's distance. Compared to existing global kernels, the proposed kernel is positive-definite. Our graph kernel is obtained by defining a distribution over random graphs, which can naturally yield random feature approximations. The random feature approximations lead to our graph embeddings, which is named as "random graph embeddings" (RGE). In particular, RGE is shown to achieve (quasi-)linear scalability with respect to the number and the size of the graphs. The experimental results on nine benchmark datasets demonstrate that RGE outperforms or matches twelve state-of-the-art graph classification algorithms.
Lingfei Wu 0001, Ian En-Hsu Yen, Zhen Zhang 0007, Kun Xu 0005, Liang Zhao 0002, Xi Peng 0005, Yinglong Xia, Charu C. Aggarwal
KDD6
2019 Rethinking Kernel Methods for Node Representation Learning on Graphs
abstract
Graph kernels are kernel methods measuring graph similarity and serve as a standard tool for graph classification. However, the use of kernel methods for node classification, which is a related problem to graph representation learning, is still ill-posed and the state-of-the-art methods are heavily based on heuristics. Here, we present a novel theoretical kernel-based framework for node classification that can bridge the gap between these two representation learning problems on graphs. Our approach is motivated by graph kernel methodology but extended to learn the node representations capturing the structural information in a graph. We theoretically show that our formulation is as powerful as any positive semidefinite kernels. To efficiently learn the kernel, we propose a novel mechanism for node feature aggregation and a data-driven similarity metric employed during the training phase. More importantly, our framework is flexible and complementary to other graph-based deep learning models, e.g., Graph Convolutional Networks (GCNs). We empirically evaluate our approach on a number of standard node classification benchmarks, and demonstrate that our model sets the new state of the art.
Yu Tian 0003, Long Zhao 0003, Xi Peng 0005, Dimitris N. Metaxas
NeurIPS3
2019 Semantic-Guided Multi-Attention Localization for Zero-Shot Learning
abstract
Zero-shot learning extends the conventional object classification to the unseen class recognition by introducing semantic representations of classes. Existing approaches predominantly focus on learning the proper mapping function for visual-semantic embedding, while neglecting the effect of learning discriminative visual features. In this paper, we study the significance of the discriminative region localization. We propose a semantic-guided multi-attention localization model, which automatically discovers the most discriminative parts of objects for zero-shot learning without any human annotations. Our model jointly learns cooperative global and local features from the whole object as well as the detected parts to categorize objects based on semantic descriptions. Moreover, with the joint supervision of embedding softmax loss and class-center triplet loss, the model is encouraged to learn features with high inter-class dispersion and intra-class compactness. Through comprehensive experiments on three widely used zero-shot learning benchmarks, we show the efficacy of the multi-attention localization and our proposed approach improves the state-of-the-art results by a considerable margin.
Yizhe Zhu, Jianwen Xie, Zhiqiang Tang 0001, Xi Peng 0005, Ahmed M. Elgammal
NeurIPS4
2019 Cartoonish sketch-based face editing in videos using identity deformation transfer
Long Zhao 0003, Fangda Han, Xi Peng 0005, Mubbasir Kapadia, Vladimir Pavlovic 0001, Dimitris N. Metaxas
Comput. Graph.3
2019 Predicting 3-D Lower Back Joint Load in Lifting: A Deep Pose Estimation Approach
abstract
Goal: Lifting is a common manual material handling task performed in the workplaces. It is considered as one of the main risk factors for work-related musculoskeletal disorders. An important criterion to identify the unsafe lifting task is the values of the net force and moment at L5/S1 joint. These values are mainly calculated in a laboratory environment, which utilizes marker-based sensors to collect three-dimensional (3-D) information and force plates to measure the external forces and moments. However, this method is usually expensive to set up, time-consuming in process, and sensitive to the surrounding environment. In this study, we propose a deep neural network (DNN)-based framework for 3-D pose estimation, which addresses the aforementioned limitations, and we employ the results for L5/S1 moment and force calculation. Methods: At the first step of the proposed framework, full body 3-D pose is captured using a DNN, then at the second step, estimated 3-D body pose along with the subject's anthropometric information is utilized to calculate L5/S1 join's kinetic by a top-down inverse dynamic algorithm. Results: To fully evaluate our approach, we conducted experiments using a lifting dataset consisting of 12 subjects performing various types of lifting tasks. The results are validated against a marker-based motion capture system as a reference. The grand mean ± SD of the total moment/force absolute errors across all the dataset was 9.06 ± 7.60 N·m/4.85 ± 4.85 N. Conclusion: The proposed method provides a reliable tool for assessment of the lower back kinetics during lifting and can be an alternative when the use of marker-based motion capture systems is not possible.
Rahil Mehrizi, Xi Peng 0005, Dimitris N. Metaxas, Shaoting Zhang 0001, Kang Li 0004
IEEE Trans. Hum. Mach. Syst.2
2018 CU-Net: Coupled U-Nets
Zhiqiang Tang 0001, Xi Peng 0005, Shijie Geng, Yizhe Zhu, Dimitris N. Metaxas
BMVC2
2018 Jointly Optimize Data Augmentation and Network Training: Adversarial Data Augmentation in Human Pose Estimation
abstract
Random data augmentation is a critical technique to avoid overfitting in training deep models. Yet, data augmentation and network training are often two isolated processes in most settings, yielding to a suboptimal training. Why not jointly optimize the two? We propose adversarial data augmentation to address this limitation. The key idea is to design a generator (e.g. an augmentation network) that competes against a discriminator (e.g. a target network) by generating hard examples online. The generator explores weaknesses of the discriminator, while the discriminator learns from hard augmentations to achieve better performance. A reward/penalty strategy is also proposed for efficient joint training. We investigate human pose estimation and carry out comprehensive ablation studies to validate our method. The results prove that our method can effectively improve state-of-the-art models without additional data effort.
Xi Peng 0005, Zhiqiang Tang 0001, Fei Yang 0001, Rogério Feris, Dimitris N. Metaxas
CVPR1
2018 A Generative Adversarial Approach for Zero-Shot Learning From Noisy Texts
abstract
Most existing zero-shot learning methods consider the problem as a visual semantic embedding one. Given the demonstrated capability of Generative Adversarial Networks(GANs) to generate images, we instead leverage GANs to imagine unseen categories from text descriptions and hence recognize novel classes with no examples being seen. Specifically, we propose a simple yet effective generative model that takes as input noisy text descriptions about an unseen class (e.g. Wikipedia articles) and generates synthesized visual features for this class. With added pseudo data, zero-shot learning is naturally converted to a traditional classification problem. Additionally, to preserve the inter-class discrimination of the generated features, a visual pivot regularization is proposed as an explicit supervision. Unlike previous methods using complex engineered regularizers, our approach can suppress the noise well without additional regularization. Empirically, we show that our method consistently outperforms the state of the art on the largest available benchmarks on Text-based Zero-shot Learning.
Yizhe Zhu, Xi Peng 0005, Ahmed M. Elgammal
CVPR4
2018 Quantized Densely Connected U-Nets for Efficient Landmark Localization
Zhiqiang Tang 0001, Xi Peng 0005, Shijie Geng, Lingfei Wu 0001, Shaoting Zhang 0001, Dimitris N. Metaxas
ECCV (3)2
2018 Learning to Forecast and Refine Residual Motion for Image-to-Video Generation
Long Zhao 0003, Xi Peng 0005, Yu Tian 0003, Mubbasir Kapadia, Dimitris N. Metaxas
ECCV (15)2
2018 Toward Marker-Free 3D Pose Estimation in Lifting: A Deep Multi-View Solution
abstract
Lifting is a common manual material handling task performed in the workplaces. It is considered as one of the main risk factors for Work-related Musculoskeletal Disorders. To improve work place safety, it is necessary to assess musculoskeletal and biomechanical risk exposures associated with these tasks, which requires very accurate 3D pose. Existing approaches mainly utilize marker-based sensors to collect 3D information. However, these methods are usually expensive to setup, timeconsuming in process, and sensitive to the surrounding environment. In this study, we propose a multi-view based deep perceptron approach to address aforementioned limitations. Our approach consists of two modules: a "view-specific perceptron" network extracts rich information independently from the image of view, which includes both 2D shape and hierarchical texture information; while a "multi-view integration" network synthesizes information from all available views to predict accurate 3D pose. To fully evaluate our approach, we carried out comprehensive experiments to compare different variants of our design. The results prove that our approach achieves comparable performance with former marker-based methods, i.e. an average error of 14:72 ± 2:96 mm on the lifting dataset. The results are also compared with state-of-the-art methods on HumanEva-I dataset [1], which demonstrates the superior performance of our approach.
Rahil Mehrizi, Xi Peng 0005, Zhiqiang Tang 0001, Dimitris N. Metaxas, Kang Li 0004
FG2
2018 CR-GAN: Learning Complete Representations for Multi-view Generation
abstract
Generating multi-view images from a single-view input is an important yet challenging problem. It has broad applications in vision, graphics, and robotics. Our study indicates that the widely-used generative adversarial network (GAN) may learn ?incomplete? representations due to the single-pathway framework: an encoder-decoder network followed by a discriminator network.We propose CR-GAN to address this problem. In addition to the single reconstruction path, we introduce a generation sideway to maintain the completeness of the learned embedding space. The two learning paths collaborate and compete in a parameter-sharing manner, yielding largely improved generality to ?unseen? dataset. More importantly, the two-pathway framework makes it possible to combine both labeled and unlabeled data for self-supervised learning, which further enriches the embedding space for realistic generations. We evaluate our approach on a wide range of datasets. The results prove that CR-GAN significantly outperforms state-of-the-art methods, especially when generating from ?unseen? inputs in wild conditions.
Yu Tian 0003, Xi Peng 0005, Long Zhao 0003, Shaoting Zhang 0001, Dimitris N. Metaxas
IJCAI2
2018 RED-Net: A Recurrent Encoder-Decoder Network for Video-Based Face Alignment
Xi Peng 0005, Rogério Feris, Xiaoyu Wang 0002, Dimitris N. Metaxas
Int. J. Comput. Vis.1
2018 Toward Personalized Modeling: Incremental and Ensemble Alignment for Sequential Faces in the Wild
Xi Peng 0005, Shaoting Zhang 0001, Yang Yu 0010, Dimitris N. Metaxas
Int. J. Comput. Vis.1
2017 Reconstruction-Based Disentanglement for Pose-Invariant Face Recognition
abstract
Deep neural networks (DNNs) trained on large-scale datasets have recently achieved impressive improvements in face recognition. But a persistent challenge remains to develop methods capable of handling large pose variations that are relatively under-represented in training data. This paper presents a method for learning a feature representation that is invariant to pose, without requiring extensive pose coverage in training data. We first propose to generate non-frontal views from a single frontal face, in order to increase the diversity of training data while preserving accurate facial details that are critical for identity discrimination. Our next contribution is to seek a rich embedding that encodes identity features, as well as non-identity ones such as pose and landmark locations. Finally, we propose a new feature reconstruction metric learning to explicitly disentangle identity and pose, by demanding alignment between the feature reconstructions through various combinations of identity and pose features, which is obtained from two images of the same subject. Experiments on both controlled and in-the-wild face datasets, such as MultiPIE, 300WLP and the profile view database CFP, show that our method consistently outperforms the state-of-the-art, especially on images with large head pose variations.
Xi Peng 0005, Xiang Yu 0002, Kihyuk Sohn, Dimitris N. Metaxas, Manmohan Krishna Chandraker
ICCV1
2016 Track Facial Points in Unconstrained Videos
Xi Peng 0005, Qiong Hu 0001, Junzhou Huang, Dimitris N. Metaxas
BMVC1
2016 A Recurrent Encoder-Decoder Network for Sequential Face Alignment
Xi Peng 0005, Rogério Feris, Xiaoyu Wang 0002, Dimitris N. Metaxas
ECCV (1)1
2015 PIEFA: Personalized Incremental and Ensemble Face Alignment
abstract
Face alignment, especially on real-time or large-scale sequential images, is a challenging task with broad applications. Both generic and joint alignment approaches have been proposed with varying degrees of success. However, many generic methods are heavily sensitive to initializations and usually rely on offline-trained static models, which limit their performance on sequential images with extensive variations. On the other hand, joint methods are restricted to offline applications, since they require all frames to conduct batch alignment. To address these limitations, we propose to exploit incremental learning for personalized ensemble alignment. We sample multiple initial shapes to achieve image congealing within one frame, which enables us to incrementally conduct ensemble alignment by group-sparse regularized rank minimization. At the same time, personalized modeling is obtained by subspace adaptation under the same incremental framework, while correction strategy is used to alleviate model drifting. Experimental results on multiple controlled and in-the-wild databases demonstrate the superior performance of our approach compared with state-of-the-arts in terms of fitting accuracy and efficiency.
Xi Peng 0005, Shaoting Zhang 0001, Dimitris N. Metaxas
ICCV1
2015 From circle to 3-sphere: Head pose estimation by instance parameterization
Xi Peng 0005, Junzhou Huang, Qiong Hu 0001, Shaoting Zhang 0001, Ahmed M. Elgammal, Dimitris N. Metaxas
Comput. Vis. Image Underst.1
2014 Robust Multi-pose Facial Expression Recognition
abstract
Previous research on facial expression recognition mainly focuses on near frontal face images, while in realistic interactive scenarios, the interested subjects may appear in arbitrary non-frontal poses. In this paper, we propose a framework to recognize six prototypical facial expressions, namely, anger, disgust, fear, joy, sadness and surprise, in an arbitrary head pose. We build a multi-pose training set by rendering 3D face scans from the BU-4DFE dynamic facial expression database [17] at 49 different viewpoints. We extract Local Binary Pattern (LBP) descriptors and further utilize multiple instance learning to mitigate the influence of inaccurate alignment in this challenging task. Experimental results demonstrate the power and validate the effectiveness of the proposed multi-pose facial expression recognition framework.
Qiong Hu 0001, Xi Peng 0005, Peng Yang 0001, Fei Yang 0001, Dimitris N. Metaxas
ICPR2
2014 Head Pose Estimation by Instance Parameterization
abstract
Head pose estimation from images is a challenging task with extensive applications. It has been attracting research attentions over decades and numerous approaches have been proposed. Among them, manifold embedding based methods, which assume that the pose variations lie on a low-dimensional manifold embedded in the high-dimensional feature space, have achieved great success. However, previous manifold embedding based methods have two drawbacks: first, they lack the capability to simultaneously deal with multiple pose-unrelated factors in a uniform way, second, they suffer from limited representation ability for out-of-sample testing inputs. In this paper we propose a novel head pose estimation method to address these problems. By learning the mapping from a uniform geometry representation to individual instance manifolds, this approach allows us to parameterize various pose-unrelated factors under a uniform framework. Our approach is a generative model which guarantees the reasonable and effective representation of new testing input. Besides, by employing eigen instance bases instead of full instance bases arisen from the training data, we can effectively compact the size of the trained model and significantly simplify the computational complexity of the testing process. Experiments on public databases such as CMU-MultiPIE and BU-4DFE, and quantitative comparisons with other state-of-the-art methods show the effectiveness of our approach.
Xi Peng 0005, Junzhou Huang, Qiong Hu 0001, Shaoting Zhang 0001, Dimitris N. Metaxas
ICPR1