Xinmei Tian 0001

dblp:03/5204-1 · DBLP profile ↗
← Back
139ranked-venue papers
18as first author
58since 2021 · last 2026
0000-0002-5952-8753ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 81 · 2 first-author · 49 since 2021Graphics, computer vision, multimedia, augmented reality and games · 76 · 14 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 1 since 2021Computer networks · 6 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 1 since 2021
YearPublicationVenuePosition
2026 Bridging the Language Gap: Uncovering and Aligning Shared Circuits for Multi-Hop Reasoning in Multilingual LLMs
abstract
Large language models (LLMs) present a paradox: they can correctly answer a multi-hop factual query in a high-resource language like English, yet fail on the identical query in another language. This raises a fundamental question about the nature of multilingual knowledge: are facts missing, or merely inaccessible? The underlying mechanisms for this knowledge gap have remained largely unexplored. In this work, we resolve this question by introducing a mechanistic interpretability framework that traces the causal pathways of multi-hop knowledge reasoning. Our analysis reveals a core, non-obvious finding: cross-lingual inconsistencies do not stem from a knowledge deficit. Instead, factual knowledge is robustly stored in a set of **shared, language-agnostic semantic neurons**. The failure originates from **misaligned attention pathways**, where a common set of critical attention heads fails to correctly route information along the reasoning chain to the appropriate knowledge neurons in lower-resource languages. This mechanistic diagnosis motivates a targeted alignment strategy: a surgical fine-tuning of only these critical heads. Experiments demonstrate that our method achieves significant improvements in multilingual multi-hop factuality—with positive cross-lingual transfer—while uniquely preserving general model capabilities, offering a scalable and mechanistically-grounded approach to building more reliable multilingual models.
Zhen Huang 0007, Yonggang Zhang 0003, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye
AAAI4
2026 Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality
abstract
Contrastively trained vision-language models like CLIP, have made remarkable progress in learning joint image-text representations, but still face challenges in compositional understanding.They often exhibit a "bag-of-words" behavior-struggling to capture the object relations, attribute-object bindings, and word order dependencies.This limitation arises not only from the reliance on global, single-vector representations for optimization, but also from the insufficient exploitation and modeling of the rich compositional information inherently present in paired image text data.In this work, we propose MACCO (MAsked Compositional Concept MOdeling), a framework that masks compositional concepts in one modality and reconstructs them conditioned on the full contextual information from the other, enabling the model to capture and align cross-modal compositional structures more effectively.To facilitate this process, we introduce two auxiliary objectives that jointly align and regularize masked features both inter-modally and intramodally.Extensive experiments on five compositional benchmarks, along with in-depth analyses, demonstrate that our approach not only significantly enhances compositionality in VLMs but also improves their ability to capture syntactic structure and linguistic information.Additionally, the improved compositionality also benefits text-to-image generation and multimodal large language model.
Wei Li 0317, Zhen Huang 0007, Xinmei Tian 0001
ACL (1)3
2025 A Similarity Paradigm Through Textual Regularization Without Forgetting
abstract
Prompt learning has emerged as a promising method for adapting pre-trained visual-language models (VLMs) to a range of downstream tasks. While optimizing the context can be effective for improving performance on specific tasks, it can often lead to poor generalization performance on unseen classes or datasets sampled from different distributions. It may be attributed to the fact that textual prompts tend to overfit downstream data distributions, leading to the forgetting of generalized knowledge derived from hand-crafted prompts. In this paper, we propose a novel method called Similarity Paradigm with Textual Regularization (SPTR) for prompt learning without forgetting. SPTR is a two-pronged design based on hand-crafted prompts that is an inseparable framework. 1) To avoid forgetting general textual knowledge, we introduce the optimal transport as a textual regularization to finely ensure approximation with hand-crafted features and tuning textual features. 2) In order to continuously unleash the general ability of multiple hand-crafted prompts, we propose a similarity paradigm for natural alignment score and adversarial alignment score to improve model robustness for generalization. Both modules share a common objective in addressing generalization issues, aiming to maximize the generalization capability derived from multiple hand-crafted prompts. Four representative tasks (i.e., non-generalization few-shot learning, base-to-novel generalization, cross-dataset generalization, domain generalization) across 11 datasets demonstrate that SPTR outperforms existing prompt learning methods.
Fangming Cui, Jan Fong, Rongfei Zeng, Xinmei Tian 0001, Jun Yu 0002
AAAI4
2025 Visual Evidence Prompting Mitigates Hallucinations in Large Vision-Language Models
abstract
Wei Li, Zhen Huang, Houqiang Li, Le Lu, Yang Lu, Xinmei Tian, Xu Shen, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Wei Li 0317, Zhen Huang 0007, Houqiang Li, Le Lu 0001, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye
ACL (1)6
2025 Interpret and Improve In-Context Learning via the Lens of Input-Label Mappings
abstract
Chenghao Sun, Zhen Huang, Yonggang Zhang, Le Lu, Houqiang Li, Xinmei Tian, Xu Shen, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhen Huang 0007, Yonggang Zhang 0003, Le Lu 0001, Houqiang Li, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye
ACL (1)6
2025 Leveraging Submodule Linearity Enhances Task Arithmetic Performance in LLMs
abstract
Task arithmetic is a straightforward yet highly effective strategy for model merging, enabling the resultant model to exhibit multi-task capabilities. Recent research indicates that models demonstrating linearity enhance the performance of task arithmetic. In contrast to existing methods that rely on the global linearization of the model, we argue that this linearity already exists within the model's submodules. In particular, we present a statistical analysis and show that submodules (e.g., layers, self-attentions, and MLPs) exhibit significantly higher linearity than the overall model. Based on these findings, we propose an innovative model merging strategy that independently merges these submodules. Especially, we derive a closed-form solution for optimal merging weights grounded in the linear properties of these submodules. Experimental results demonstrate that our method consistently outperforms the standard task arithmetic approach and other established baselines across different model scales and various tasks. This result highlights the benefits of leveraging the linearity of submodules and provides a new perspective for exploring solutions for effective and practical multi-task model merging.
Rui Dai 0005, Sile Hu, Xu Shen 0001, Yonggang Zhang 0003, Xinmei Tian 0001, Jieping Ye
ICLR5
2025 A Theoretical Perspective: How to Prevent Model Collapse in Self-consuming Training Loops
abstract
High-quality data is essential for training large generative models, yet the vast reservoir of real data available online has become nearly depleted. Consequently, models increasingly generate their own data for further training, forming Self-consuming Training Loops (STLs). However, the empirical results have been strikingly inconsistent: some models degrade or even collapse, while others successfully avoid these failures, leaving a significant gap in theoretical understanding to explain this discrepancy. This paper introduces the intriguing notion of *recursive stability* and presents the first theoretical generalization analysis, revealing how both model architecture and the proportion between real and synthetic data influence the success of STLs. We further extend this analysis to transformers in in-context learning, showing that even a constant-sized proportion of real data ensures convergence, while also providing insights into optimal synthetic data sizing.
Shi Fu, Xinmei Tian 0001, Dacheng Tao
ICLR4
2025 Enhancing Target-unspecific Tasks through a Features Matrix
abstract
Recent developments in prompt learning of large Vision-Language Models (VLMs) have significantly improved performance in target-specific tasks. However, these prompting methods often struggle to tackle the target-unspecific or generalizable tasks effectively. It may be attributed to the fact that overfitting training causes the model to forget its general knowledge. The general knowledge has a strong promotion on target-unspecific tasks. To alleviate this issue, we propose a novel Features Matrix (FM) approach designed to enhance these models on target-unspecific tasks. Our method extracts and leverages general knowledge, shaping a Features Matrix (FM). Specifically, the FM captures the semantics of diverse inputs from a deep and fine perspective, preserving essential general knowledge, which mitigates the risk of overfitting. Representative evaluations demonstrate that: 1) the FM is compatible with existing frameworks as a generic and flexible module, and 2) the FM significantly showcases its effectiveness in enhancing target-unspecific tasks (base-to-novel generalization, domain generalization, and cross-dataset generalization), achieving state-of-the-art performance.
Fangming Cui, Yonggang Zhang 0003, Xinmei Tian 0001, Jun Yu 0002
ICML4
2025 Towards Generalizable Detector for Generated Image
abstract
The effective detection of generated images is crucial to mitigate potential risks associated with their misuse. Despite significant progress, a fundamental challenge remains: ensuring the generalizability of detectors. To address this, we propose a novel perspective on understanding and improving generated image detection, inspired by the human cognitive process: Humans identify an image as unnatural based on specific patterns because these patterns lie outside the space spanned by those of natural images. This is intrinsically related to out-of-distribution (OOD) detection, which identifies samples whose semantic patterns (i.e., labels) lie outside the semantic pattern space of in-distribution (ID) samples. By treating patterns of generated images as OOD samples, we demonstrate that models trained merely over natural images bring guaranteed generalization ability under mild assumptions. This transforms the generalization challenge of generated image detection into the problem of fitting natural image patterns. Based on this insight, we propose a generalizable detection method through the lens of ID energy. Theoretical results capture the generalization risk of the proposed method. Experimental results across multiple benchmarks demonstrate the effectiveness of our approach.
Qianshu Cai, Chao Wu 0001, Yonggang Zhang 0003, Jun Yu 0002, Xinmei Tian 0001
NeurIPS5
2025 An Effective Levelling Paradigm for Unlabeled Scenarios
abstract
Advancements in direct-integration fine-tuning frameworks have underscored their potential to enhance the performance of labeled scenarios and tasks. To enhance the generalization of different categories in the same dataset, some methods have added visual loss to these frameworks for unlabeled scenarios. However, the performance of these methods through visual loss does not improve significantly in domain generalization and cross-dataset generalization tasks. This may be attributed to the uncoordinated learning of the two-modalities alignment and visual loss. To mitigate this issue of uncoordinated learning, we propose a novel method called Levelling Paradigm (LePa) to improve performance for unlabeled tasks or scenarios. The proposed LePa, designed as a plug-in module, dynamically constrains and coordinates multiple objective functions, thereby improving the generalization of these baseline methods. Comprehensive experiments have shown that our design can effectively address generalized scenarios and tasks.
Fangming Cui, Yuqiang Ren, Liang Xiao 0007, Xinmei Tian 0001
NeurIPS6
2025 Epistemic Uncertainty for Generated Image Detection
abstract
We introduce a novel framework for AI-generated image detection through epistemic uncertainty, aiming to address critical security concerns in the era of generative models. Our key insight stems from the observation that distributional discrepancies between training and testing data manifest distinctively in the epistemic uncertainty space of machine learning models. In this context, the distribution shift between natural and generated images leads to elevated epistemic uncertainty in models trained on natural images when evaluating generated ones. Hence, we exploit this phenomenon by using epistemic uncertainty as a proxy for detecting generated images. This converts the challenge of generated image detection into the problem of uncertainty estimation, underscoring the generalization performance of the model used for uncertainty estimation. Fortunately, advanced large-scale vision models pre-trained on extensive natural images have shown excellent generalization performance for various scenarios. Thus, we utilize these pre-trained models to estimate the epistemic uncertainty of images and flag those with high uncertainty as generated. Extensive experiments demonstrate the efficacy of our method.
Jun Nie, Yonggang Zhang 0003, Tongliang Liu, Yiu-Ming Cheung, Bo Han 0003, Xinmei Tian 0001
NeurIPS6
2025 Detecting Generated Images by Fitting Natural Image Distributions
abstract
The increasing realism of generated images has raised significant concerns about their potential misuse, necessitating robust detection methods. Current approaches mainly rely on training binary classifiers, which depend heavily on the quantity and quality of available generated images. In this work, we propose a novel framework that exploits geometric differences between the data manifolds of natural and generated images. To exploit this difference, we employ a pair of functions engineered to yield consistent outputs for natural images but divergent outputs for generated ones, leveraging the property that their gradients reside in mutually orthogonal subspaces. This design enables a simple yet effective detection method: an image is identified as generated if a transformation along its data manifold induces a significant change in the loss value of a self-supervised model pre-trained on natural images. Further more, to address diminishing manifold disparities in advanced generative models, we leverage normalizing flows to amplify detectable differences by extruding generated images away from the natural image manifold. Extensive experiments demonstrate the efficacy of this method.
Yonggang Zhang 0003, Jun Nie, Xinmei Tian 0001, Mingming Gong, Kun Zhang 0001, Bo Han 0003
NeurIPS3
2025 Out-of-Distribution Detection with Virtual Outlier Smoothing
abstract
Abstract Detecting out-of-distribution (OOD) inputs plays a crucial role in guaranteeing the reliability of deep neural networks (DNNs) when deployed in real-world scenarios. However, DNNs typically exhibit overconfidence in OOD samples, which is attributed to the similarity in patterns between OOD and in-distribution (ID) samples. To mitigate this overconfidence, advanced approaches suggest the incorporation of auxiliary OOD samples during model training, where the outliers are assigned with an equal likelihood of belonging to any category. However, identifying outliers that share patterns with ID samples poses a significant challenge. To address the challenge, we propose a novel method, V irtual O utlier S m o othing (VOSo), which constructs auxiliary outliers using ID samples, thereby eliminating the need to search for OOD samples. Specifically, VOSo creates these virtual outliers by perturbing the semantic regions of ID samples and infusing patterns from other ID samples. For instance, a virtual outlier might consist of a cat’s face with a dog’s nose, where the cat’s face serves as the semantic feature for model prediction. Meanwhile, VOSo adjusts the labels of virtual OOD samples based on the extent of semantic region perturbation, aligning with the notion that virtual outliers may contain ID patterns. Extensive experiments are conducted on diverse OOD detection benchmarks, demonstrating the effectiveness of the proposed VOSo. Our code will be available at https://github.com/junz-debug/VOSo .
Jun Nie, Yadan Luo, Shanshan Ye, Yonggang Zhang 0003, Xinmei Tian 0001, Zhen Fang 0001
Int. J. Comput. Vis.5
2025 Consistent prompt learning for vision-language models
Yonggang Zhang 0003, Xinmei Tian 0001
Knowl. Based Syst.2
2025 Boosting Fair Classifier Generalization through Adaptive Priority Reweighing
abstract
With the increasing penetration of machine learning applications in critical decision-making areas, calls for algorithmic fairness are more prominent. Although there have been various modalities to improve algorithmic fairness through learning with fairness constraints, their performance does not generalize well in the test set. A performance-promising fair algorithm with better generalizability is needed. This article proposes a novel adaptive reweighing method to eliminate the impact of the distribution shifts between training and test data on model generalizability. Most previous reweighing methods propose to assign a unified weight for each (sub)group. Rather, our method granularly models the distance from the sample predictions to the decision boundary. Our adaptive reweighing method prioritizes samples closer to the decision boundary and assigns a higher weight to improve the generalizability of fair classifiers. Extensive experiments are performed to validate the generalizability of our adaptive priority reweighing method for accuracy and fairness measures (i.e., equal opportunity, equalized odds, and demographic parity) in tabular benchmarks. We also highlight the performance of our method in improving the fairness of language and vision models. The code is available at https://github.com/che2198/APW .
Mengnan Du, Jindong Gu, Xinmei Tian 0001, Fengxiang He
ACM Trans. Knowl. Discov. Data5
2024 Sheared Backpropagation for Fine-Tuning Foundation Models
abstract
Fine-tuning is the process of extending the training of pre-trained models on specific target tasks, thereby significantly enhancing their performance across various applications. However, fine-tuning often demands large memory consumption, posing a challenge for low-memory devices that some previous memory-efficient fine-tuning methods attempted to mitigate by pruning activations for gradient computation, albeit at the cost of significant computational overhead from the pruning processes during training. To address these challenges, we introduce PreBackRazor; a novel activation pruning scheme offering both computational and memory efficiency through a sparsified back-propagation strategy, which judiciously avoids unnecessary activation pruning and storage and gradient computation. Before activation pruning, our approach samples a probability of selecting a portion of parameters to freeze, utilizing a bandit method for updates to prioritize impactful gradients on convergence. During the feed-forward pass, each model layer adjusts adaptively based on parameter activation status, obviating the need for sparsification and storage of redundant activations for subsequent backpropagation. Benchmarking on fine-tuning foundation models, our approach maintains baseline accuracy across diverse tasks, yielding over 20% speedup and around 10% memory reduction. Moreover, integrating with an advanced CUDA kernel achieves up to 60% speedup without extra memory costs or accuracy loss, significantly enhancing the efficiency of fine-tuning foundation models on memory-constrained devices.
Zhiyuan Yu 0004, Li Shen 0008, Liang Ding 0006, Xinmei Tian 0001, Yixin Chen 0001, Dacheng Tao
CVPR4
2024 Enhanced Motion-Text Alignment for Image-to-Video Transfer Learning
abstract
Extending large image-text pre-trained models (e.g., CLIP) for video understanding has made significant advancements. To enable the capability of CLIP to perceive dynamic information in videos, existing works are dedicated to equipping the visual encoder with various temporal modules. However, these methods exhibit “asymmetry” between the visual and textual sides, with neither temporal descriptions in input texts nor temporal modules in text encoder. This limitation hinders the potential of language supervision emphasized in CLIP, and restricts the learning of temporal features, as the text encoder has demonstrated limited proficiency in motion understanding. To address this issue, we propose leveraging “MoTion-Enhanced Descriptions” (MoTED) to facilitate the extraction of distinctive temporal features in videos. Specifically, we first generate discriminative motion-related descriptions via querying GPT-4 to compare easy-confusing action categories. Then, we incorporate both the visual and textual encoders with additional perception modules to process the video frames and generated descriptions, respectively. Finally, we adopt a contrastive loss to align the visual and textual motion features. Extensive experiments on five benchmarks show that MoTED surpasses state-of-the-art methods with convincing gaps, laying a solid foundation for empowering CLIP with strong temporal modeling.
Chaoqun Wan, Tongliang Liu, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye
CVPR4
2024 Interpretable Composition Attribution Enhancement for Visio-linguistic Compositional Understanding
abstract
Contrastively trained vision-language models such as CLIP have achieved remarkable progress in vision and language representation learning.Despite the promising progress, their proficiency in compositional reasoning over attributes and relations (e.g., distinguishing between "the car is underneath the person" and "the person is underneath the car") remains notably inadequate.We investigate the cause for this deficient behavior is the composition attribution issue, where the attribution scores (e.g., attention scores or GradCAM scores) for relations (e.g., underneath) or attributes (e.g., red) in the text are substantially lower than those for object terms.In this work, we show such issue is mitigated via a novel framework called CAE (Composition Attribution Enhancement).This generic framework incorporates various interpretable attribution methods to encourage the model to pay greater attention to composition words denoting relationships and attributes within the text.Detailed analysis shows that our approach enables the models to adjust and rectify the attribution of the texts.Extensive experiments across seven benchmarks reveal that our framework significantly enhances the ability to discern intricate details and construct more sophisticated interpretations of combined visual and linguistic elements.
Wei Li 0317, Zhen Huang 0007, Xinmei Tian 0001, Le Lu 0001, Houqiang Li, Xu Shen 0001, Jieping Ye
EMNLP3
2024 Convergence of Bayesian Bilevel Optimization
abstract
This paper presents the first theoretical guarantee for Bayesian bilevel optimization (BBO) that we term for the prevalent bilevel framework combining Bayesian optimization at the outer level to tune hyperparameters, and the inner-level stochastic gradient descent (SGD) for training the model. We prove sublinear regret bounds suggesting simultaneous convergence of the inner-level model parameters and outer-level hyperparameters to optimal configurations for generalization capability. A pivotal, technical novelty in the proofs is modeling the excess risk of the SGD-trained parameters as evaluation noise during Bayesian optimization. Our theory implies the inner unit horizon, defined as the number of SGD iterations, shapes the convergence behavior of BBO. This suggests practical guidance on configuring the inner unit horizon to enhance training efficiency and model performance.
Shi Fu, Fengxiang He, Xinmei Tian 0001, Dacheng Tao
ICLR3
2024 Out-of-Distribution Detection with Negative Prompts
abstract
Out-of-distribution (OOD) detection is indispensable for open-world machine learning models. Inspired by recent success in large pre-trained language-vision models, e.g., CLIP, advanced works have achieved impressive OOD detection results by matching the *similarity* between image features and features of learned prompts, i.e., positive prompts. However, existing works typically struggle with OOD samples having similar features with those of known classes. One straightforward approach is to introduce negative prompts to achieve a *dissimilarity* matching, which further assesses the anomaly level of image features by introducing the absence of specific features. Unfortunately, our experimental observations show that either employing a prompt like "not a photo of a" or learning a prompt to represent "not containing" fails to capture the dissimilarity for identifying OOD samples. The failure may be contributed to the diversity of negative features, i.e., tons of features could indicate features not belonging to a known class. To this end, we propose to learn a set of negative prompts for each class. The learned positive prompt (for all classes) and negative prompts (for each class) are leveraged to measure the similarity and dissimilarity in the feature space simultaneously, enabling more accurate detection of OOD samples. Extensive experiments are conducted on diverse OOD detection benchmarks, showing the effectiveness of our proposed method.
Jun Nie, Yonggang Zhang 0003, Zhen Fang 0001, Tongliang Liu, Bo Han 0003, Xinmei Tian 0001
ICLR6
2024 FedImpro: Measuring and Improving Client Update in Federated Learning
abstract
Federated Learning (FL) models often experience client drift caused by heterogeneous data, where the distribution of data differs across clients. To address this issue, advanced research primarily focuses on manipulating the existing gradients to achieve more consistent client models. In this paper, we present an alternative perspective on client drift and aim to mitigate it by generating improved local models. First, we analyze the generalization contribution of local training and conclude that this generalization contribution is bounded by the conditional Wasserstein distance between the data distribution of different clients. Then, we propose FedImpro, to construct similar conditional distributions for local training. Specifically, FedImpro decouples the model into high-level and low-level components, and trains the high-level portion on reconstructed feature distributions. This approach enhances the generalization contribution and reduces the dissimilarity of gradients in FL. Experimental results show that FedImpro can help FL defend against data heterogeneity and enhance the generalization performance of the model.
Zhenheng Tang, Yonggang Zhang 0003, Shaohuai Shi, Xinmei Tian 0001, Tongliang Liu, Bo Han 0003, Xiaowen Chu 0001
ICLR4
2024 Robust Training of Federated Models with Extremely Label Deficiency
abstract
Federated semi-supervised learning (FSSL) has emerged as a powerful paradigm for collaboratively training machine learning models using distributed data with label deficiency. Advanced FSSL methods predominantly focus on training a single model on each client. However, this approach could lead to a discrepancy between the objective functions of labeled and unlabeled data, resulting in gradient conflicts. To alleviate gradient conflict, we propose a novel twin-model paradigm, called **Twinsight**, designed to enhance mutual guidance by providing insights from different perspectives of labeled and unlabeled data. In particular, Twinsight concurrently trains a supervised model with a supervised objective function while training an unsupervised model using an unsupervised objective function. To enhance the synergy between these two models, Twinsight introduces a neighborhood-preserving constraint, which encourages the preservation of the neighborhood relationship among data features extracted by both models. Our comprehensive experiments on four benchmark datasets provide substantial evidence that Twinsight can significantly outperform state-of-the-art methods across various experimental settings, demonstrating the efficacy of the proposed Twinsight.
Yonggang Zhang 0003, Zhiqin Yang, Xinmei Tian 0001, Nannan Wang 0001, Tongliang Liu, Bo Han 0003
ICLR3
2024 From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint Tuning
abstract
Large Language Models (LLMs) tend to prioritize adherence to user prompts over providing veracious responses, leading to the sycophancy issue. When challenged by users, LLMs tend to admit mistakes and provide inaccurate responses even if they initially provided the correct answer. Recent works propose to employ supervised fine-tuning (SFT) to mitigate the sycophancy issue, while it typically leads to the degeneration of LLMs' general capability. To address the challenge, we propose a novel supervised pinpoint tuning (SPT), where the region-of-interest modules are tuned for a given objective. Specifically, SPT first reveals and verifies a small percentage (<5%) of the basic modules, which significantly affect a particular behavior of LLMs. i.e., sycophancy. Subsequently, SPT merely fine-tunes these identified modules while freezing the rest. To verify the effectiveness of the proposed SPT, we conduct comprehensive experiments, demonstrating that SPT significantly mitigates the sycophancy issue of LLMs (even better than SFT). Moreover, SPT introduces limited or even no side effects on the general capability of LLMs. Our results shed light on how to precisely, effectively, and efficiently explain and improve the targeted ability of LLMs.
Wei Chen 0005, Zhen Huang 0007, Liang Xie 0003, Binbin Lin 0001, Houqiang Li, Le Lu 0001, Xinmei Tian 0001, Deng Cai 0001, Yonggang Zhang 0003, Wenxiao Wang 0001, Xu Shen 0001, Jieping Ye
ICML7
2024 Towards Theoretical Understandings of Self-Consuming Generative Models
abstract
This paper tackles the emerging challenge of training generative models within a self-consuming loop, wherein successive generations of models are recursively trained on mixtures of real and synthetic data from previous generations. We construct a theoretical framework to rigorously evaluate how this training procedure impacts the data distributions learned by future models, including parametric and non-parametric models. Specifically, we derive bounds on the total variation (TV) distance between the synthetic data distributions produced by future models and the original real data distribution under various mixed training scenarios for diffusion models with a one-hidden-layer neural network score function. Our analysis demonstrates that this distance can be effectively controlled under the condition that mixed training dataset sizes or proportions of real data are large enough. Interestingly, we further unveil a phase transition induced by expanding synthetic data amounts, proving theoretically that while the TV distance exhibits an initial ascent, it declines beyond a threshold point. Finally, we present results for kernel density estimation, delivering nuanced insights such as the impact of mixed data training on error propagation.
Shi Fu, Sen Zhang 0006, Yingjie Wang 0007, Xinmei Tian 0001, Dacheng Tao
ICML4
2024 Interpreting and Improving Large Language Models in Arithmetic Calculation
abstract
Large language models (LLMs) have demonstrated remarkable potential across numerous applications and have shown an emergent ability to tackle complex reasoning tasks, such as mathematical computations. However, even for the simplest arithmetic calculations, the intrinsic mechanisms behind LLMs remains mysterious, making it challenging to ensure reliability. In this work, we delve into uncovering a specific mechanism by which LLMs execute calculations. Through comprehensive experiments, we find that LLMs frequently involve a small fraction ($<$5%) of attention heads, which play a pivotal role in focusing on operands and operators during calculation processes. Subsequently, the information from these operands is processed through multi-layer perceptrons (MLPs), progressively leading to the final solution. These pivotal heads/MLPs, though identified on a specific dataset, exhibit transferability across different datasets and even distinct tasks. This insight prompted us to investigate the potential benefits of selectively fine-tuning these essential heads/MLPs to boost the LLMs’ computational performance. We empirically find that such precise tuning can yield notable enhancements on mathematical prowess, without compromising the performance on non-mathematical tasks. Our work serves as a preliminary exploration into the arithmetic calculation abilities inherent in LLMs, laying a solid foundation to reveal more intricate mathematical tasks.
Chaoqun Wan, Yonggang Zhang 0003, Yiu-Ming Cheung, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye
ICML5
2024 Advancing Prompt Learning through an External Layer
abstract
Prompt learning represents a promising method for adapting pre-trained vision-language models (VLMs) to various downstream tasks by learning a set of text embeddings. One challenge inherent to these methods is the poor generalization performance due to the invalidity of the learned text embeddings for unseen tasks. A straightforward approach to bridge this gap is to freeze the text embeddings in prompts, which results in a lack of capacity to adapt VLMs for downstream tasks. To address this dilemma, we propose a paradigm called EnPrompt with a novel External Layer (EnLa). Specifically, we propose a textual external layer and learnable visual embeddings for adapting VLMs to downstream tasks. The learnable external layer is built upon valid embeddings of pre-trained CLIP. This design considers the balance of learning capabilities between the two branches. To align the textual and visual features, we propose a novel two-pronged approach: i) we introduce the optimal transport as the discrepancy metric to align the vision and text modalities, and ii) we introduce a novel strengthening feature to enhance the interaction between these two modalities. Four representative experiments (i.e., base-to-novel generalization, few-shot learning, cross-dataset generalization, domain shifts generalization) across 15 datasets demonstrate that our method outperforms the existing prompt learning method.
Fangming Cui, Xun Yang 0001, Chao Wu 0001, Liang Xiao 0007, Xinmei Tian 0001
ACM Multimedia5
2024 Adaptive Time-Stepping Schedules for Diffusion Models
abstract
This paper studies how to tune the stepping schedule in diffusion models, which is mostly fixed in current practice, lacking theoretical foundations and assurance of optimal performance at the chosen discretization points. In this paper, we advocate the use of adaptive time-stepping schedules and design two algorithms with an optimized sampling error bound $EB$: (1) for continuous diffusion, we treat $EB$ as the loss function to discretization points and run gradient descent to adjust them; and (2) for discrete diffusion, we propose a greedy algorithm that adjusts only one discretization point to its best position in each iteration. We conducted extensive experiments that show (1) improved generation ability in well-trained models, and (2) premature though usable generation ability in under-trained models. The code is submitted and will be released publicly.
Fengxiang He, Shi Fu, Xinmei Tian 0001, Dacheng Tao
UAI4
2024 Expert-level diagnosis of pediatric posterior fossa tumors via consistency calibration
Yonggang Zhang 0003, Xinmei Tian 0001
Knowl. Based Syst.4
2024 FedGAMMA: Federated Learning With Global Sharpness-Aware Minimization
abstract
Federated learning (FL) is a promising framework for privacy-preserving and distributed training with decentralized clients. However, there exists a large divergence between the collected local updates and the expected global update, which is known as the client drift and mainly caused by heterogeneous data distribution among clients, multiple local training steps, and partial client participation training. Most existing works tackle this challenge based on the empirical risk minimization (ERM) rule, while less attention has been paid to the relationship between the global loss landscape and the generalization ability. In this work, we propose FedGAMMA, a novel FL algorithm with Global sharpness-Aware MiniMizAtion to seek a global flat landscape with high performance. Specifically, in contrast to FedSAM which only seeks the local flatness and still suffers from performance degradation when facing the client-drift issue, we adopt a local varieties control technique to better align each client's local updates to alleviate the client drift and make each client heading toward the global flatness together. Finally, extensive experiments demonstrate that FedGAMMA can substantially outperform several existing FL baselines on various datasets, and it can well address the client-drift issue and simultaneously seek a smoother and flatter global landscape.
Rong Dai, Xun Yang 0001, Li Shen 0008, Xinmei Tian 0001, Meng Wang 0001, Yongdong Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2023 Sharper Bounds for Uniformly Stable Algorithms with Stationary Mixing Process
Shi Fu, Yunwen Lei, Qiong Cao, Xinmei Tian 0001, Dacheng Tao
ICLR4
2023 Moderately Distributional Exploration for Domain Generalization
abstract
Domain generalization (DG) aims to tackle the distribution shift between training domains and unknown target domains. Generating new domains is one of the most effective approaches, yet its performance gain depends on the distribution discrepancy between the generated and target domains. Distributionally robust optimization is promising to tackle distribution discrepancy by exploring domains in an uncertainty set. However, the uncertainty set may be overwhelmingly large, leading to low-confidence prediction in DG. It is because a large uncertainty set could introduce domains containing semantically different factors from training domains. To address this issue, we propose to perform a $\textit{mo}$derately $\textit{d}$istributional $\textit{e}$xploration (MODE) for domain generalization. Specifically, MODE performs distribution exploration in an uncertainty $\textit{subset}$ that shares the same semantic factors with the training domains. We show that MODE can endow models with provable generalization performance on unknown target domains. The experimental results show that MODE achieves competitive performance compared to state-of-the-art baselines.
Rui Dai 0005, Yonggang Zhang 0003, Zhen Fang 0001, Bo Han 0003, Xinmei Tian 0001
ICML5
2023 Structured Cooperative Learning with Graphical Model Priors
abstract
We study how to train personalized models for different tasks on decentralized devices with limited local data. We propose "Structured Cooperative Learning (SCooL)", in which a cooperation graph across devices is generated by a graphical model prior to automatically coordinate mutual learning between devices. By choosing graphical models enforcing different structures, we can derive a rich class of existing and novel decentralized learning algorithms via variational inference. In particular, we show three instantiations of SCooL that adopt Dirac distribution, stochastic block model (SBM), and attention as the prior generating cooperation graphs. These EM-type algorithms alternate between updating the cooperation graph and cooperative learning of local models. They can automatically capture the cross-task correlations among devices by only monitoring their model updating in order to optimize the cooperation graph. We evaluate SCooL and compare it with existing decentralized learning methods on an extensive set of benchmarks, on which SCooL always achieves the highest accuracy of personalized models and significantly outperforms other baselines on communication efficiency. Our code is available at https://github.com/ShuangtongLi/SCooL.
Shuangtong Li, Tianyi Zhou 0001, Xinmei Tian 0001, Dacheng Tao
ICML3
2023 Adaptive Priority Reweighing for Generalizing Fairness Improvement
abstract
With the increasing penetration of Machine-Learning (ML) applications in critical decision-making areas, calls for algorithmic fairness are more prominent. Though there have been diverse modalities to improve algorithmic fairness through training the algorithms with fairness constraints, their performance does not generalize well at the test set. A performance-promising fair algorithm with better generalizability is needed. This paper proposes a novel adaptive reweighing method to eliminate the impact of the distribution shifts between training and test data on model generalizability. Specifically, instead of assigning a unified weight for each (sub)group as most previous reweighing methods propose, we granularly model the distance from the sample predictions to the decision boundary and assign higher individual weight to the samples closer to the decision boundary in each (sub)group. Our adaptive reweighing method prioritizes the samples closer to the decision boundary and assigns a higher weight to improve the generalizability of fair classifiers. We design extensive experiments to evaluate the generalizability of our adaptive priority reweighing method for accuracy and fairness measures (i.e., equal opportunity, equalized odds, and demographic parity.) in tabular benchmarks across Adult, COMPAS, and IPUMS. We further highlight the performance of our method in improving the fairness of language and vision models. We believe that our method shows promising results in improving the fairness of any pre-trained models simply via fine-tuning.
Xinmei Tian 0001
IJCNN3
2023 Semantic-Aware Mixup for Domain Generalization
abstract
Deep neural networks (DNNs) have shown exciting performance in various tasks, yet suffer generalization failures when meeting unknown target domains. One of the most promising approaches to achieve domain generalization (DG) is generating unseen data, e.g., mixup, to cover the unknown target data. However, existing works overlook the challenges induced by the simultaneous appearance of changes in both the semantic and distribution space. Accordingly, such a challenge makes source distributions hard to fit for DNNs. To mitigate the hard-fitting issue, we propose to perform a semantic-aware mixup (SAM) for domain generalization, where whether to perform mixup depends on the semantic and domain information. The feasibility of SAM shares the same spirits with the Fourier-based mixup. Namely, the Fourier phase spectrum is expected to contain semantics information (relating to labels), while the Fourier amplitude retains other information (relating to style information). Built upon the insight, SAM applies different mixup strategies to the Fourier phase spectrum and amplitude information. For instance, SAM merely performs mixup on the amplitude spectrum when both the semantic and domain information changes. Consequently, the overwhelmingly large change can be avoided. We validate the effectiveness of SAM using image classification tasks on several DG benchmarks.
Chengchao Xu, Xinmei Tian 0001
IJCNN2
2023 3D Creation at Your Fingertips: From Text or Image to 3D Assets
abstract
We demonstrate an automatic 3D creation system, which can create realistic 3D assets solely from a text or image prompt without requiring any specialized 3D modeling skills. Users can either describe the object they envision in natural language or upload a reference image that records what they have seen with the phone. Our system will generate a high-quality 3D mesh that faithfully matches the users' input. We propose a coarse-to-fine framework to achieve this goal. Specifically, we first obtain a low-resolution mesh instantly by utilizing a pre-trained text/image conditional 3D generative model. Using such coarse mesh as the initialization, we further optimize a high-resolution textured 3D mesh with fine-grained appearance guidance from large-scale 2D diffusion models. Our system can create visually-pleasing results in minutes, which is significantly faster than existing methods. Meanwhile, the system ensures that the resulting 3D assets are precisely aligned with the input text or image prompt. With these advanced capabilities, our demonstration provides a streamlined and intuitive platform for users to incorporate 3D creation into their daily lives.
Yang Chen 0048, Jingwen Chen 0001, Yingwei Pan, Xinmei Tian 0001, Tao Mei 0001
ACM Multimedia4
2023 FedFed: Feature Distillation against Data Heterogeneity in Federated Learning
abstract
Federated learning (FL) typically faces data heterogeneity, i.e., distribution shifting among clients. Sharing clients' information has shown great potentiality in mitigating data heterogeneity, yet incurs a dilemma in preserving privacy and promoting model performance. To alleviate the dilemma, we raise a fundamental question: Is it possible to share partial features in the data to tackle data heterogeneity? In this work, we give an affirmative answer to this question by proposing a novel approach called **Fed**erated **Fe**ature **d**istillation (FedFed). Specifically, FedFed partitions data into performance-sensitive features (i.e., greatly contributing to model performance) and performance-robust features (i.e., limitedly contributing to model performance). The performance-sensitive features are globally shared to mitigate data heterogeneity, while the performance-robust features are kept locally. FedFed enables clients to train models over local and shared data. Comprehensive experiments demonstrate the efficacy of FedFed in promoting model performance.
Zhiqin Yang, Yonggang Zhang 0003, Yu Zheng 0021, Xinmei Tian 0001, Tongliang Liu, Bo Han 0003
NeurIPS4
2023 Bi-calibration Networks for Weakly-Supervised Video Representation Learning
Fuchen Long, Ting Yao 0003, Zhaofan Qiu, Xinmei Tian 0001, Jiebo Luo 0001, Tao Mei 0001
Int. J. Comput. Vis.4
2023 A heuristic multi-objective task scheduling framework for container-based clouds via actor-critic reinforcement learning
Lilu Zhu, Feng Wu 0001, Yanfeng Hu, Xinmei Tian 0001
Neural Comput. Appl.5
2023 Domain Generalization Via Encoding and Resampling in a Unified Latent Space
abstract
Domain generalization aims to generalize a network trained on multiple domains to unknown yet related domains. Operating under the assumption that invariant information generalizes well to unknown domains, previous work has aimed to minimize the discrepancies amongst distributions across given domains. However, without prior regularization of feature distributions, the network in practice overfits the invariant information in the given domains. Moreover, if there are insufficient samples in given domains, then domain generalizability is limited, as diverse domain variations are not captured. To address these two drawbacks, we propose to explicitly map features in known and unknown domains onto latent space in a fixed Gaussian mixture distribution by variational coding. As a result, features in different classes follow Gaussian distributions with different mean values. The predefined latent space narrows discrepancies between known and unknown domains and effectively separates samples into different classes. Moreover, we propose to perturb sample features with gradients from the distribution regularized loss. This perturbation generates samples beyond but near the latent space of prior distributions, which has a profound impact on domain variations. Experiments and visualizations demonstrate the effectiveness of our proposed method.
Zhiwei Xiong, Xinmei Tian 0001, Zhengjun Zha
IEEE Trans. Multim.4
2023 Domain-Class Correlation Decomposition for Generalizable Person Re-Identification
abstract
Domain generalization in person re-identification is a highly important meaningful and practical task in which a model trained with data from several source domains is expected to generalize well to unseen target domains. Domain adversarial learning is a promising domain generalization method that aims to remove domain information in the latent representation through adversarial training. However, in person re-identification, the domain and class are correlated, and we theoretically show that domain adversarial learning will lose certain information about class due to this domain-class correlation. Inspired by causal inference, we propose to perform interventions to the domain factor$d$, aiming to decompose the domain-class correlation. To achieve this goal, we proposed estimating the resulting representation$z^{*}$caused by the intervention through first- and second-order statistical characteristic matching. Specifically, we build a memory bank to restore the statistical characteristics of each domain. Then, we use the newly generated samples$\lbrace z^{*},y,d^{*}\rbrace$to compute the loss function. These samples are domain-class correlation decomposed; thus, we can learn a domain-invariant representation that can capture more class-related features. Extensive experiments show that our model outperforms the state-of-the-art methods on the large-scale domain generalization Re-ID benchmark.
Xinmei Tian 0001
IEEE Trans. Multim.2
2023 Category-Stitch Learning for Union Domain Generalization
abstract
Domain generalization aims at generalizing the network trained on multiple domains to unknown but related domains. Under the assumption that different domains share the same classes, previous works can build relationships across domains. However, in realistic scenarios, the change of domains is always followed by the change of categories, which raises a difficulty for collecting sufficient aligned categories across domains. Bearing this in mind, this article introduces union domain generalization (UDG) as a new domain generalization scenario, in which the label space varies across domains, and the categories in unknown domains belong to the union of all given domain categories. The absence of categories in given domains is the main obstacle to aligning different domain distributions and obtaining domain-invariant information. To address this problem, we propose category-stitch learning (CSL), which aims at jointly learning the domain-invariant information and completing missing categories in all domains through an improved variational autoencoder and generators. The domain-invariant information extraction and sample generation cross-promote each other to better generalizability. Additionally, we decouple category and domain information and propose explicitly regularizing the semantic information by the classification loss with transferred samples. Thus our method can breakthrough the category limit and generate samples of missing categories in each domain. Extensive experiments and visualizations are conducted on MNIST, VLCS, PACS, Office-Home, and DomainNet datasets to demonstrate the effectiveness of our proposed method.
Zhiwei Xiong, Yuning Lu, Xinmei Tian 0001, Zhengjun Zha
ACM Trans. Multim. Comput. Commun. Appl.5
2022 Learning to Collaborate in Decentralized Learning of Personalized Models
abstract
Learning personalized models for user-customized computer-vision tasks is challenging due to the limited private-data and computation available on each edge device. Decentralized learning (DL) can exploit the images distributed over devices on a network topology to train a global model but is not designed to train personalized models for different tasks or optimize the topology. Moreover, the mixing weights used to aggregate neighbors' gradient messages in DL can be suboptimal for personalization since they are not adaptive to different nodes/tasks and learning stages. In this paper, we dynamically update the mixing-weights to improve the personalized model for each node's task and meanwhile learn a sparse topology to reduce communication costs. Our first approach, “learning to collaborate (L2C) ”, directly optimizes the mixing weights to minimize the local validation loss per node for a predefined set of nodes/tasks. In order to produce mixing weights for new nodes or tasks, we further develop “meta-L2C‘, which learns an attention mechanism to automatically assign mixing weights by comparing two nodes' model updates. We evaluate both methods on diverse benchmarks and experimental settings for image classification. Thorough comparisons to both classical and recent methods for IID/non-IID decentralized and federated learning demonstrate our method's advantages in identifying collaborators among nodes, learning sparse topology, and producing better personalized models with low communication and computational cost.
Shuangtong Li, Tianyi Zhou 0001, Xinmei Tian 0001, Dacheng Tao
CVPR3
2022 Prompt Distribution Learning
abstract
We present prompt distribution learning for effectively adapting a pre-trained vision-language model to address downstream recognition tasks. Our method not only learns low-bias prompts from a few samples but also captures the distribution of diverse prompts to handle the varying visual representations. In this way, we provide high-quality task-related content for facilitating recognition. This prompt distribution learning is realized by an efficient approach that learns the output embeddings of prompts instead of the input embeddings. Thus, we can employ a Gaussian distribution to model them effectively and derive a surrogate loss for efficient training. Extensive experiments on 12 datasets demonstrate that our method consistently and significantly outperforms existing methods. For example, with 1 sample per category, it relatively improves the average result by 9.1% compared to human-crafted prompts.
Yuning Lu, Jianzhuang Liu, Yonggang Zhang 0003, Xinmei Tian 0001
CVPR5
2022 Meta Convolutional Neural Networks for Single Domain Generalization
abstract
In single domain generalization, models trained with data from only one domain are required to perform well on many unseen domains. In this paper, we propose a new model, termed meta convolutional neural network, to solve the single domain generalization problem in image recognition. The key idea is to decompose the convolutional features of images into meta features. Acting as “visual words”, meta features are defined as universal and basic visual elements for image representations (like words for documents in language). Taking meta features as reference, we propose compositional operations to eliminate irrelevant features of local convolutional features by an addressing process and then to reformulate the convolutional feature maps as a composition of related meta features. In this way, images are universally coded without biased information from the unseen domain, which can be processed by following modules trained in the source domain. The compositional operations adopt a regression analysis technique to learn the meta features in an online batch learning manner. Extensive experiments on multiple benchmark datasets verify the superiority of the proposed model in improving single domain generalization ability.
Chaoqun Wan, Xu Shen 0001, Yonggang Zhang 0003, Zhiheng Yin, Xinmei Tian 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001
CVPR5
2022 Self-Supervision Can Be a Good Few-Shot Learner
Yuning Lu, Liangjian Wen, Jianzhuang Liu, Xinmei Tian 0001
ECCV (19)5
2022 Adversarial Robustness Through the Lens of Causality
Yonggang Zhang 0003, Mingming Gong, Tongliang Liu, Gang Niu 0001, Xinmei Tian 0001, Bo Han 0003, Bernhard Schölkopf, Kun Zhang 0001
ICLR5
2022 DisPFL: Towards Communication-Efficient Personalized Federated Learning via Decentralized Sparse Training
abstract
Personalized federated learning is proposed to handle the data heterogeneity problem amongst clients by learning dedicated tailored local models for each user. However, existing works are often built in a centralized way, leading to high communication pressure and high vulnerability when a failure or an attack on the central server occurs. In this work, we propose a novel personalized federated learning framework in a decentralized (peer-to-peer) communication protocol named DisPFL, which employs personalized sparse masks to customize sparse local models on the edge. To further save the communication and computation cost, we propose a decentralized sparse training technique, which means that each local model in DisPFL only maintains a fixed number of active parameters throughout the whole local training and peer-to-peer communication process. Comprehensive experiments demonstrate that DisPFL significantly saves the communication bottleneck for the busiest node among all clients and, at the same time, achieves higher model accuracy with less computation cost and communication rounds. Furthermore, we demonstrate that our method can easily adapt to heterogeneous local clients with varying computation complexities and achieves better personalized performances.
Rong Dai, Li Shen 0008, Fengxiang He, Xinmei Tian 0001, Dacheng Tao
ICML4
2022 Identity-Disentangled Adversarial Augmentation for Self-supervised Learning
abstract
Data augmentation is critical to contrastive self-supervised learning, whose goal is to distinguish a sample’s augmentations (positives) from other samples (negatives). However, strong augmentations may change the sample-identity of the positives, while weak augmentation produces easy positives/negatives leading to nearly-zero loss and ineffective learning. In this paper, we study a simple adversarial augmentation method that can modify training data to be hard positives/negatives without distorting the key information about their original identities. In particular, we decompose a sample $x$ to be its variational auto-encoder (VAE) reconstruction $G(x)$ plus the residual $R(x)=x-G(x)$, where $R(x)$ retains most identity-distinctive information due to an information-theoretic interpretation of the VAE objective. We then adversarially perturb $G(x)$ in the VAE’s bottleneck space and adds it back to the original $R(x)$ as an augmentation, which is therefore sufficiently challenging for contrastive learning and meanwhile preserves the sample identity intact. We apply this “identity-disentangled adversarial augmentation (IDAA)” to different self-supervised learning methods. On multiple benchmark datasets, IDAA consistently improves both their efficiency and generalization performance. We further show that IDAA learned on a dataset can be transferred to other datasets. Code is available at \href{https://github.com/kai-wen-yang/IDAA}{https://github.com/kai-wen-yang/IDAA}.
Tianyi Zhou 0001, Xinmei Tian 0001, Dacheng Tao
ICML3
2022 Towards Lightweight Black-Box Attack Against Deep Neural Networks
abstract
Black-box attacks can generate adversarial examples without accessing the parameters of target model, largely exacerbating the threats of deployed deep neural networks (DNNs). However, previous works state that black-box attacks fail to mislead target models when their training data and outputs are inaccessible. In this work, we argue that black-box attacks can pose practical attacks in this extremely restrictive scenario where only several test samples are available. Specifically, we find that attacking the shallow layers of DNNs trained on a few test samples can generate powerful adversarial examples. As only a few samples are required, we refer to these attacks as lightweight black-box attacks. The main challenge to promoting lightweight attacks is to mitigate the adverse impact caused by the approximation error of shallow layers. As it is hard to mitigate the approximation error with few available samples, we propose Error TransFormer (ETF) for lightweight attacks. Namely, ETF transforms the approximation error in the parameter space into a perturbation in the feature space and alleviates the error by disturbing features. In experiments, lightweight black-box attacks with the proposed ETF achieve surprising results. For example, even if only 1 sample per category available, the attack success rate in lightweight black-box attacks is only about 3% lower than that of the black-box attacks with complete training data.
Yonggang Zhang 0003, Chaoqun Wan, Tongliang Liu, Bo Han 0003, Xinmei Tian 0001
NeurIPS8
2022 Adversarial Auto-Augment with Label Preservation: A Representation Learning Principle Guided Approach
abstract
Data augmentation is a critical contributing factor to the success of deep learning but heavily relies on prior domain knowledge which is not always available. Recent works on automatic data augmentation learn a policy to form a sequence of augmentation operations, which are still pre-defined and restricted to limited options. In this paper, we show that a prior-free autonomous data augmentation's objective can be derived from a representation learning principle that aims to preserve the minimum sufficient information of the labels. Given an example, the objective aims at creating a distant ``hard positive example'' as the augmentation, while still preserving the original label. We then propose a practical surrogate to the objective that can be optimized efficiently and integrated seamlessly into existing methods for a broad class of machine learning tasks, e.g., supervised, semi-supervised, and noisy-label learning. Unlike previous works, our method does not require training an extra generative model but instead leverages the intermediate layer representations of the end-task model for generating data augmentations. In experiments, we show that our method consistently brings non-trivial improvements to the three aforementioned learning tasks from both efficiency and final performance, either or not combined with pre-defined augmentations, e.g., on medical images when domain knowledge is unavailable and the existing augmentation techniques perform poorly. Code will be released publicly.
Yanchao Sun, Jiahao Su, Fengxiang He, Xinmei Tian 0001, Furong Huang, Tianyi Zhou 0001, Dacheng Tao
NeurIPS5
2022 DigestPath: A benchmark dataset with challenge review for the pathological detection and segmentation of digestive-system
Qian Da, Zhongyu Li 0002, Yanfei Zuo, Chenbin Zhang, Jingxin Liu 0005, Wen Chen 0001, Jiahui Li 0005, Dou Xu, Hongmei Yi, Zhe Wang 0043, Li Zhang 0040, Xianying He, Xiaofan Zhang 0002, Ke Mei, Chuang Zhu, Weizeng Lu, LinLin Shen, Jun Shi 0006, Jun Li 0106, Sreehari S, Ganapathy Krishnamurthi, Jiangcheng Yang, Tiancheng Lin 0001, Qingyu Song 0004, Xuechen Liu 0004, Simon Graham, Raja Muhammad Saad Bashir, Canqian Yang, Shaofei Qin, Xinmei Tian 0001, Jie Zhao 0014, Dimitris N. Metaxas, Hongsheng Li 0001, Chaofu Wang, Shaoting Zhang 0001
Medical Image Anal.34
2022 CRAR: Accelerating Stereo Matching with Cascaded Residual Regression and Adaptive Refinement
abstract
Dense stereo matching estimates the depth for each pixel of the referenced images. Recently, deep learning algorithms have dramatically promoted the development of stereo matching. The state-of-the-art result is achieved by models adopting deep convolutional neural networks. However, a considerable computational burden is also introduced, which slows the inference. To solve this problem, previous works down-sampled the input images to decrease the spatial size. However, down-sampling increases the error rate and its lower bound. In this article, we accelerate stereo matching algorithms through the improvement of network structure. Inspired by network compression, we conduct decomposition and sparsification to squeeze the computationally expensive cost optimization network. It is sparsified and then decomposed into smaller networks, which are designed and trained in a cascaded manner to reach the nearest possible performance of the larger network. Previous methods have utilized numerous refinement methods to adjust the coarse disparity. We integrate refinement methods to create an unified algorithm to utilize parallelism for running devices to further accelerate the inference. The extensive experiments on Kitti2015, Kitti2012, and Middlebury datasets demonstrate the efficiency of our method.
Linghua Zeng, Xinmei Tian 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2021 Revisiting Knowledge Distillation: An Inheritance and Exploration Framework
abstract
Knowledge Distillation (KD) is a popular technique to transfer knowledge from a teacher model or ensemble to a student model. Its success is generally attributed to the privileged information on similarities/consistency between the class distributions or intermediate feature representations of the teacher model and the student model. However, directly pushing the student model to mimic the probabilities/features of the teacher model to a large extent limits the student model in learning undiscovered knowledge/features. In this paper, we propose a novel inheritance and exploration knowledge distillation framework (IE-KD), in which a student model is split into two parts - inheritance and exploration. The inheritance part is learned with a similarity loss to transfer the existing learned knowledge from the teacher model to the student model, while the exploration part is encouraged to learn representations different from the inherited ones with a dis-similarity loss. Our IE-KD framework is generic and can be easily combined with existing distillation or mutual learning methods for training deep neural networks. Extensive experiments demonstrate that these two parts can jointly push the student model to learn more diversified and effective representations, and our IE-KD can be a general technique to improve the student network to achieve SOTA performance. Furthermore, by applying our IE-KD to the training of two networks, the performance of both can be improved w.r.t. deep mutual learning.
Zhen Huang 0007, Xu Shen 0001, Jun Xing, Tongliang Liu, Xinmei Tian 0001, Houqiang Li, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001
CVPR5
2021 A Style and Semantic Memory Mechanism for Domain Generalization*
abstract
Mainstream state-of-the-art domain generalization algorithms tend to prioritize the assumption on semantic in-variance across domains. Meanwhile, the inherent intra-domain style invariance is usually underappreciated and put on the shelf. In this paper, we reveal that leveraging intra-domain style invariance is also of pivotal importance in improving the efficiency of domain generalization. We verify that it is critical for the network to be informative on what domain features are invariant and shared among in-stances, so that the network sharpens its understanding and improves its semantic discriminative ability. Correspondingly, we also propose a novel “jury” mechanism, which is particularly effective in learning useful semantic feature commonalities among domains. Our complete model called STEAM can be interpreted as a novel probabilistic graphical model, for which the implementation requires convenient constructions of two kinds of memory banks: semantic feature bank and style feature bank. Empirical results show that our proposed framework surpasses the state-of-the-art methods by clear margins.
Yang Chen 0048, Yu Wang 0102, Yingwei Pan, Ting Yao 0003, Xinmei Tian 0001, Tao Mei 0001
ICCV5
2021 3D Local Convolutional Neural Networks for Gait Recognition
abstract
The goal of gait recognition is to learn the unique spatiotemporal pattern about the human body shape from its temporal changing characteristics. As different body parts behave differently during walking, it is intuitive to model the spatio-temporal patterns of each part separately. However, existing part-based methods equally divide the feature maps of each frame into fixed horizontal stripes to get local parts. It is obvious that these stripe partition-based methods cannot accurately locate the body parts. First, different body parts can appear at the same stripe (e.g., arms and the torso), and one part can appear at different stripes in different frames (e.g., hands). Second, different body parts possess different scales, and even the same part in different frames can appear at different locations and scales. Third, different parts also exhibit distinct movement patterns (e.g., at which frame the movement starts, the position change frequency, how long it lasts). To overcome these issues, we propose novel 3D local operations as a generic family of building blocks for 3D gait recognition backbones. The proposed 3D local operations support the extraction of local 3D volumes of body parts in a sequence with adaptive spatial and temporal scales, locations and lengths. In this way, the spatio-temporal patterns of the body parts are well learned from the 3D local neighborhood in partspecific scales, locations, frequencies and lengths. Experiments demonstrate that our 3D local convolutional neural networks achieve state-of-the-art performance on popular gait datasets. Code is available at: https://github.com/yellowtownhz/3DLocalCNN.
Zhen Huang 0007, Dixiu Xue, Xu Shen 0001, Xinmei Tian 0001, Houqiang Li, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ICCV4
2021 Transferrable Contrastive Learning for Visual Domain Adaptation
abstract
Self-supervised learning (SSL) has recently become the favorite among feature learning methodologies. It is therefore appealing for domain adaptation approaches to consider incorporating SSL. The intuition is to enforce instance-level feature consistency such that the predictor becomes somehow invariant across domains. However, most existing SSL methods in the regime of domain adaptation usually are treated as standalone auxiliary components, leaving the signatures of domain adaptation unattended. Actually, the optimal region where the domain gap vanishes and the instance level constraint that SSL peruses may not coincide at all. From this point, we present a particular paradigm of self-supervised learning tailored for domain adaptation, i.e., Transferrable Contrastive Learning (TCL), which links the SSL and the desired cross-domain transferability congruently. We find contrastive learning intrinsically a suitable candidate for domain adaptation, as its instance invariance assumption can be conveniently promoted to cross-domain class-level invariance favored by domain adaptation tasks. Based on particular memory bank constructions and pseudo label strategies, TCL then penalizes cross-domain intra-class domain discrepancy between source and target through a clean and novel contrastive loss. The free lunch is, thanks to the incorporation of contrastive learning, TCL relies on a moving-averaged key encoder that naturally achieves a temporally ensembled version of pseudo labels for target data, which avoids pseudo label error propagation at no extra cost. TCL therefore efficiently reduces cross-domain gaps. Through extensive experiments on benchmarks (Office-Home, VisDA-2017, Digits-five, PACS and DomainNet) for both single-source and multi-source domain adaptation tasks, TCL has demonstrated state-of-the-art performances.
Yang Chen 0048, Yingwei Pan, Yu Wang 0102, Ting Yao 0003, Xinmei Tian 0001, Tao Mei 0001
ACM Multimedia5
2021 Class-Disentanglement and Applications in Adversarial Detection and Defense
abstract
What is the minimum necessary information required by a neural net $D(\cdot)$ from an image $x$ to accurately predict its class? Extracting such information in the input space from $x$ can allocate the areas $D(\cdot)$ mainly attending to and shed novel insights to the detection and defense of adversarial attacks. In this paper, we propose ''class-disentanglement'' that trains a variational autoencoder $G(\cdot)$ to extract this class-dependent information as $x - G(x)$ via a trade-off between reconstructing $x$ by $G(x)$ and classifying $x$ by $D(x-G(x))$, where the former competes with the latter in decomposing $x$ so the latter retains only necessary information for classification in $x-G(x)$. We apply it to both clean images and their adversarial images and discover that the perturbations generated by adversarial attacks mainly lie in the class-dependent part $x-G(x)$. The decomposition results also provide novel interpretations to classification and attack models. Inspired by these observations, we propose to conduct adversarial detection and adversarial defense respectively on $x - G(x)$ and $G(x)$, which consistently outperform the results on the original $x$. In experiments, this simple approach substantially improves the detection and defense against different types of adversarial attacks.
Tianyi Zhou 0001, Yonggang Zhang 0003, Xinmei Tian 0001, Dacheng Tao
NeurIPS4
2021 Learning multi-granularity features from multi-granularity regions for person re-identification
Jiwei Yang, Xinmei Tian 0001
Neurocomputing3
2020 Learning to Localize Actions from Moments
Fuchen Long, Ting Yao 0003, Zhaofan Qiu, Xinmei Tian 0001, Jiebo Luo 0001, Tao Mei 0001
ECCV (3)4
2020 Dual-Path Distillation: A Unified Framework to Improve Black-Box Attacks
abstract
We study the problem of constructing black-box adversarial attacks, where no model information is revealed except for the feedback knowledge of the given inputs. To obtain sufficient knowledge for crafting adversarial examples, previous methods query the target model with inputs that are perturbed with different searching directions. However, these methods suffer from poor query efficiency since the employed searching directions are sampled randomly. To mitigate this issue, we formulate the goal of mounting efficient attacks as an optimization problem in which the adversary tries to fool the target model with a limited number of queries. Under such settings, the adversary has to select appropriate searching directions to reduce the number of model queries. By solving the efficient-attack problem, we find that we need to distill the knowledge in both the path of the adversarial examples and the path of the searching directions. Therefore, we propose a novel framework, dual-path distillation, that utilizes the feedback knowledge not only to craft adversarial examples but also to alter the searching directions to achieve efficient attacks. Experimental results suggest that our framework can significantly increase the query efficiency.
Yonggang Zhang 0003, Tongliang Liu, Xinmei Tian 0001
ICML4
2020 Spatio-Temporal Inception Graph Convolutional Networks for Skeleton-Based Action Recognition
abstract
Skeleton-based human action recognition has attracted much attention with the prevalence of accessible depth sensors. Recently, graph convolutional networks (GCNs) have been widely used for this task due to their powerful capability to model graph data. The topology of the adjacency graph is a key factor for modeling the correlations of the input skeletons. Thus, previous methods mainly focus on the design/learning of the graph topology. But once the topology is learned, only a single-scale feature and one transformation exist in each layer of the networks. Many insights, such as multi-scale information and multiple sets of transformations, that have been proven to be very effective in convolutional neural networks (CNNs), have not been investigated in GCNs. The reason is that, due to the gap between graph-structured skeleton data and conventional image/video data, it is very challenging to embed these insights into GCNs. To overcome this gap, we reinvent the split-transform-merge strategy in GCNs for skeleton sequence processing. Specifically, we design a simple and highly modularized graph convolutional network architecture for skeleton-based action recognition. Our network is constructed by repeating a building block that aggregates multi-granularity information from both the spatial and temporal paths. Extensive experiments demonstrate that our network outperforms state-of-the-art methods by a significant margin with only 1/5 of the parameters and 1/10 of the FLOPs.
Zhen Huang 0007, Xu Shen 0001, Xinmei Tian 0001, Houqiang Li, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ACM Multimedia3
2020 Real-World Image Denoising with Deep Boosting
abstract
We propose a Deep Boosting Framework (DBF) for real-world image denoising by integrating the deep learning technique into the boosting algorithm. The DBF replaces conventional handcrafted boosting units by elaborate convolutional neural networks, which brings notable advantages in terms of both performance and speed. We design a lightweight Dense Dilated Fusion Network (DDFN) as an embodiment of the boosting unit, which addresses the vanishing of gradients during training due to the cascading of networks while promoting the efficiency of limited parameters. The capabilities of the proposed method are first validated on several representative simulation tasks including non-blind and blind Gaussian denoising and JPEG image deblocking. We then focus on a practical scenario to tackle with the complex and challenging real-world noise. To facilitate leaning-based methods including ours, we build a new Real-world Image Denoising (RID) dataset, which contains 200 pairs of high-resolution images with diverse scene content under various shooting conditions. Moreover, we conduct comprehensive analysis on the domain shift issue for real-world denoising and propose an effective one-shot domain transfer scheme to address this issue. Comprehensive experiments on widely used benchmarks demonstrate that the proposed method significantly surpasses existing methods on the task of real-world image denoising. Code and dataset are available at https://github.com/ngchc/deepBoosting.
Chang Chen 0004, Zhiwei Xiong, Xinmei Tian 0001, Zhengjun Zha, Feng Wu 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 Accelerating Convolutional Neural Networks by Removing Interspatial and Interkernel Redundancies
abstract
Recently, the high computational resource demands of convolutional neural networks (CNNs) have hindered a wide range of their applications. To solve this problem, many previous works attempted to reduce the redundant calculations during the evaluation of CNNs. However, these works mainly focused on either interspatial or interkernel redundancy. In this paper, we further accelerate existing CNNs by removing both types of redundancies. First, we convert interspatial redundancy into interkernel redundancy by decomposing one convolutional layer to one block that we design. Then, we adopt rank-selection and pruning methods to remove the interkernel redundancy. The rank-selection method, which considerably reduces manpower, contributes to determining the number of kernels to be pruned in the pruning method. We apply a layer-wise training algorithm rather than the traditional end-to-end training to overcome the difficulty of convergence. Finally, we fine-tune the entire network to achieve better performance. Our method is applied on three widely used datasets of an image classification task. We achieve better results in terms of accuracy and compression rate compared with previous state-of-the-art methods.
Linghua Zeng, Xinmei Tian 0001
IEEE Trans. Cybern.2
2020 Principal Component Adversarial Example
abstract
Despite having achieved excellent performance on various tasks, deep neural networks have been shown to be susceptible to adversarial examples, i.e., visual inputs crafted with structural imperceptible noise. To explain this phenomenon, previous works implicate the weak capability of the classification models and the difficulty of the classification tasks. These explanations appear to account for some of the empirical observations but lack deep insight into the intrinsic nature of adversarial examples, such as the generation method and transferability. Furthermore, previous works generate adversarial examples completely rely on a specific classifier (model). Consequently, the attack ability of adversarial examples is strongly dependent on the specific classifier. More importantly, adversarial examples cannot be generated without a trained classifier. In this paper, we raise a question: what is the real cause of the generation of adversarial examples? To answer this question, we propose a new concept, called the adversarial region, which explains the existence of adversarial examples as perturbations perpendicular to the tangent plane of the data manifold. This view yields a clear explanation of the transfer property across different models of adversarial examples. Moreover, with the notion of the adversarial region, we propose a novel target-free method to generate adversarial examples via principal component analysis. We verify our adversarial region hypothesis on a synthetic dataset and demonstrate through extensive experiments on real datasets that the adversarial examples generated by our method have competitive or even strong transferability compared with model-dependent adversarial example generating methods. Moreover, our experiment shows that the proposed method is more robust to defensive methods than previous methods.
Yonggang Zhang 0003, Xinmei Tian 0001, Xinchao Wang, Dacheng Tao
IEEE Trans. Image Process.2
2020 A Multi-Organ Nucleus Segmentation Challenge
abstract
Generalized nucleus segmentation techniques can contribute greatly to reducing the time to develop and validate visual biomarkers for new digital pathology datasets. We summarize the results of MoNuSeg 2018 Challenge whose objective was to develop generalizable nuclei segmentation techniques in digital pathology. The challenge was an official satellite event of the MICCAI 2018 conference in which 32 teams with more than 80 participants from geographically diverse institutes participated. Contestants were given a training set with 30 images from seven organs with annotations of 21,623 individual nuclei. A test dataset with 14 images taken from seven organs, including two organs that did not appear in the training set was released without annotations. Entries were evaluated based on average aggregated Jaccard index (AJI) on the test set to prioritize accurate instance segmentation as opposed to mere semantic segmentation. More than half the teams that completed the challenge outperformed a previous baseline. Among the trends observed that contributed to increased accuracy were the use of color normalization as well as heavy data augmentation. Additionally, fully convolutional networks inspired by variants of U-Net, FCN, and Mask-RCNN were popularly used, typically based on ResNet or VGG base architectures. Watershed segmentation on predicted semantic segmentation maps was a popular post-processing strategy. Several of the top techniques compared favorably to an individual human annotator and can be used with confidence for nuclear morphometrics.
Neeraj Kumar 0002, Ruchika Verma, Deepak Anand, Yanning Zhou 0001, Omer Fahri Onder, Efstratios Tsougenis, Hao Chen 0011, Pheng-Ann Heng, Jiahui Li 0005, Navid Alemi Koohbanani, Mostafa Jahanifar, Neda Zamani Tajeddin, Ali Gooya, Nasir M. Rajpoot, Xuhua Ren, Sihang Zhou 0001, Qian Wang 0001, Dinggang Shen, Cheng-Kun Yang, Chi-Hung Weng, Wei-Hsiang Yu, Chao-Yuan Yeh, Shuoyu Xu, Pak-Hei Yeung, Amirreza Mahbod, Gerald Schaefer, Isabella Ellinger, Rupert Ecker, Örjan Smedby, Chunliang Wang, Benjamin Chidester, Vinh Ton-That, Minh-Triet Tran, Jian Ma 0004, Minh N. Do, Simon Graham, Quoc Dang Vu, Jin Tae Kwak, Akshaykumar Gunda, Raviteja Chunduri, Corey Hu, Dariush Lotfi, Reza Safdari, Antanas Kascenas, Alison O'Neil, Dennis Eschweiler, Johannes Stegmaier, Yanping Cui, Kailin Chen, Xinmei Tian 0001, Philipp Grüning, Erhardt Barth, Elad Arbel, Itay Remer, Amir Ben-Dor, Ekaterina Sirazitdinova, Matthias Kohl, Stefan Braunewell, Yuexiang Li, Xinpeng Xie, LinLin Shen, Jun Ma 0016, Krishanu Das Baksi, Mohammad Azam Khan, Jaegul Choo, Adrián Colomer, Valery Naranjo, Linmin Pei, Khan M. Iftekharuddin, Kaushiki Roy, Debotosh Bhattacharjee, Aníbal Pedraza, Gloria Bueno García, Sabarinathan Devanathan, Saravanan Radhakrishnan, Praveen Koduganty, Zihan Wu 0001, Guanyu Cai, Amit Sethi
IEEE Trans. Medical Imaging56
2020 Coarse-to-Fine Localization of Temporal Action Proposals
abstract
Localizing temporal action proposals from long videos is a fundamental challenge in video analysis (e.g., action detection and recognition or dense video captioning). Most existing approaches often overlook the hierarchical granularities of actions and thus fail to discriminate fine-grained action proposals (e.g., hand washing laundry or changing a tire in vehicle repair). In this paper, we propose a novel coarse-to-fine temporal proposal (CFTP) approach to localize temporal action proposals by exploring different action granularities. Our proposed CFTP consists of three stages: a coarse proposal network (CPN) to generate long action proposals, a temporal convolutional anchor network (CAN) to localize finer proposals, and a proposal reranking network (PRN) to further identify proposals from previous stages. Specifically, CPN explores three complementary actionness curves (namely pointwise, pairwise, and recurrent curves) that represent actions at different levels for generating coarse proposals, while CAN refines these proposals by a multiscale cascaded 1D-convolutional anchor network. In contrast to existing works, our coarse-to-fine approach can progressively localize fine-grained action proposals. We conduct extensive experiments on two action benchmarks (THUMOS14 and ActivityNet v1.3) and demonstrate the superior performance of our approach when compared to the state-of-the-art techniques on various video understanding tasks.
Fuchen Long, Ting Yao 0003, Zhaofan Qiu, Xinmei Tian 0001, Tao Mei 0001, Jiebo Luo 0001
IEEE Trans. Multim.4
2020 Concentrated Local Part Discovery With Fine-Grained Part Representation for Person Re-Identification
abstract
The attention mechanism for person re-identification has been widely studied with deep convolutional neural networks. This mechanism works as a good complement to the global features extracted from an image of the entire human body. However, existing works mainly focus on discovering local parts with simple feature representations, such as global average pooling. Moreover, these works either require extra supervision, such as labeling of body joints, or pay little attention to the guidance of part learning, resulting in scattered activation of learned parts. Furthermore, existing works usually extract local features from different body parts via global average pooling and then concatenate them together as good global features. We find that local features acquired in this way contribute little to the overall performance. In this paper, we argue the significance of local part description and explore the attention mechanism from both local part discovery and local part representation aspects. For local part discovery, we propose a new constrained attention module to make the activated regions concentrated and meaningful without extra supervision. For local part representation, we propose a statistical-positional-relational descriptor to represent local parts from a fine-grained viewpoint. Extensive experiments are conducted to validate the overall performance, the effectiveness of each component, and the generalization ability. We achieve a rank-1 accuracy of 95.1% on Market1501, 64.7% on CUHK03, 87.1% on DukeMTMC-ReID, and 79.9% on MSMT17, outperforming state-of-the-art methods.
Chaoqun Wan, Xinmei Tian 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001
IEEE Trans. Multim.3
2020 Graph Edge Convolutional Neural Networks for Skeleton-Based Action Recognition
abstract
Body joints, directly obtained from a pose estimation model, have proven effective for action recognition. Existing works focus on analyzing the dynamics of human joints. However, except joints, humans also explore motions of limbs for understanding actions. Given this observation, we investigate the dynamics of human limbs for skeleton-based action recognition. Specifically, we represent an edge in a graph of a human skeleton by integrating its spatial neighboring edges (for encoding the cooperation between different limbs) and its temporal neighboring edges (for achieving the consistency of movements in an action). Based on this new edge representation, we devise a graph edge convolutional neural network (CNN). Considering the complementarity between graph node convolution and edge convolution, we further construct two hybrid networks by introducing different shared intermediate layers to integrate graph node and edge CNNs. Our contributions are twofold, graph edge convolution and hybrid networks for integrating the proposed edge convolution and the conventional node convolution. Experimental results on the Kinetics and NTU-RGB+D data sets demonstrate that our graph edge convolution is effective at capturing the characteristics of actions and that our graph edge CNN significantly outperforms the existing state-of-the-art skeleton-based action recognition methods.
Xikun Zhang 0002, Chang Xu 0002, Xinmei Tian 0001, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.3
2020 Video Retrieval with Similarity-Preserving Deep Temporal Hashing
abstract
Despite the fact that remarkable progress has been made in recent years, Content-based Video Retrieval (CBVR) is still an appealing research topic due to increasing search demands in the Internet era of big data. This article aims to explore an efficient CBVR system by discriminately hashing videos into short binary codes. Existing video hashing methods usually encounter two weaknesses originating from the following sources: (1) Most works adopt the separated stages method or the frame-pooling based end-to-end architecture. However, the spatial-temporal properties of videos cannot be fully explored or kept well in the follow-up hashing step. (2) Discriminative learning based on pairwise or triplet constraints often suffers from slow convergence and poor local optimization, mainly because of the limited samples for each update. To alleviate these problems, we propose an end-to-end video retrieval framework called the Similarity-Preserving Deep Temporal Hashing (SPDTH) network. Specifically, we equip the model with the ability to capture spatial-temporal properties of videos and to generate binary codes by stacked Gated Recurrent Units (GRUs). It unifies video temporal modeling and learning to hash into one step to allow for maximum retention of information. We also introduce a deep metric learning objective called ℓ 2 All _ loss for network training by preserving intra-class similarity and inter-class separability, and a quantization loss between the real-valued outputs and the binary codes is minimized. Extensive experiments on several challenging datasets demonstrate that SPDTH can consistently outperform state-of-the-art methods.
Richang Hong, Xinmei Tian 0001, Meng Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2019 Exploring Object Relation in Mean Teacher for Cross-Domain Detection
abstract
Rendering synthetic data (e.g., 3D CAD-rendered images) to generate annotations for learning deep models in vision tasks has attracted increasing attention in recent years. However, simply applying the models learnt on synthetic images may lead to high generalization error on real images due to domain shift. To address this issue, recent progress in cross-domain recognition has featured the Mean Teacher, which directly simulates unsupervised domain adaptation as semi-supervised learning. The domain gap is thus naturally bridged with consistency regularization in a teacher-student scheme. In this work, we advance this Mean Teacher paradigm to be applicable for cross-domain detection. Specifically, we present Mean Teacher with Object Relations (MTOR) that novelly remolds Mean Teacher under the backbone of Faster R-CNN by integrating the object relations into the measure of consistency cost between teacher and student modules. Technically, MTOR firstly learns relational graphs that capture similarities between pairs of regions for teacher and student respectively. The whole architecture is then optimized with three consistency regularizations: 1) region-level consistency to align the region-level predictions between teacher and student, 2) inter-graph consistency for matching the graph structures between teacher and student, and 3) intra-graph consistency to enhance the similarity between regions of same class within the graph of student. Extensive experiments are conducted on the transfers across Cityscapes, Foggy Cityscapes, and SIM10k, and superior results are reported when comparing to state-of-the-art approaches. More remarkably, we obtain a new record of single model: 22.8% of mAP on Syn2Real detection dataset.
Yingwei Pan, Chong-Wah Ngo, Xinmei Tian 0001, Ling-Yu Duan, Ting Yao 0003
CVPR4
2019 Camera Lens Super-Resolution
abstract
Existing methods for single image super-resolution (SR) are typically evaluated with synthetic degradation models such as bicubic or Gaussian downsampling. In this paper, we investigate SR from the perspective of camera lenses, named as CameraSR, which aims to alleviate the intrinsic tradeoff between resolution (R) and field-of-view (V) in realistic imaging systems. Specifically, we view the R-V degradation as a latent model in the SR process and learn to reverse it with realistic low- and high-resolution image pairs. To obtain the paired images, we propose two novel data acquisition strategies for two representative imaging systems (i.e., DSLR and smartphone cameras), respectively. Based on the obtained City100 dataset, we quantitatively analyze the performance of commonly-used synthetic degradation models, and demonstrate the superiority of CameraSR as a practical solution to boost the performance of existing SR methods. Moreover, CameraSR can be readily generalized to different content and devices, which serves as an advanced digital zoom tool in realistic imaging systems.
Chang Chen 0004, Zhiwei Xiong, Xinmei Tian 0001, Zhengjun Zha, Feng Wu 0001
CVPR3
2019 Compact Feature Learning for Multi-Domain Image Classification
abstract
The goal of multi-domain learning is to improve the performance over multiple domains by making full use of all training data from them. However, variations of feature distributions across different domains result in a non-trivial solution of multi-domain learning. The state-of-the-art work regarding multi-domain classification aims to extract domain-invariant features and domain-specific features independently. However, they view the distributions of features from different classes as a general distribution and try to match these distributions across domains, which lead to the mixture of features from different classes across domains and degrade the performance of classification. Additionally, existing works only force the shared features among domains to be orthogonal to the features in the domain-specific network. However, redundant features between the domain-specific networks still remain, which may shrink the discriminative ability of domain-specific features. Therefore, we propose an end-to-end network to obtain the more optimal features, which we call compact features. We propose to extract the domain-invariant features by matching the joint distributions of different domains, which have dis- tinct boundaries between different classes. Moreover, we add an orthogonal constraint between the private features across domains to ensure the discriminative ability of the domain-specific space. The proposed method is validated on three landmark datasets, and the results demonstrate the effectiveness of our method.
Xinmei Tian 0001, Zhiwei Xiong, Feng Wu 0001
CVPR2
2019 Gaussian Temporal Awareness Networks for Action Localization
abstract
Temporally localizing actions in a video is a fundamental challenge in video understanding. Most existing approaches have often drawn inspiration from image object detection and extended the advances, e.g., SSD and Faster R-CNN, to produce temporal locations of an action in a 1D sequence. Nevertheless, the results can suffer from robustness problem due to the design of predetermined temporal scales, which overlooks the temporal structure of an action and limits the utility on detecting actions with complex variations. In this paper, we propose to address the problem by introducing Gaussian kernels to dynamically optimize temporal scale of each action proposal. Specifically, we present Gaussian Temporal Awareness Networks (GTAN) - a new architecture that novelly integrates the exploitation of temporal structure into an one-stage action localization framework. Technically, GTAN models the temporal structure through learning a set of Gaussian kernels, each for a cell in the feature maps. Each Gaussian kernel corresponds to a particular interval of an action proposal and a mixture of Gaussian kernels could further characterize action proposals with various length. Moreover, the values in each Gaussian curve reflect the contextual contributions to the localization of an action proposal. Extensive experiments are conducted on both THUMOS14 and ActivityNet v1.3 datasets, and superior results are reported when comparing to state-of-the-art approaches. More remarkably, GTAN achieves 1.9% and 1.1% improvements in mAP on testing set of the two datasets.
Fuchen Long, Ting Yao 0003, Zhaofan Qiu, Xinmei Tian 0001, Jiebo Luo 0001, Tao Mei 0001
CVPR4
2019 Learning Spatio-Temporal Representation With Local and Global Diffusion
abstract
Convolutional Neural Networks (CNN) have been regarded as a powerful class of models for visual recognition problems. Nevertheless, the convolutional filters in these networks are local operations while ignoring the large-range dependency. Such drawback becomes even worse particularly for video recognition, since video is an information-intensive media with complex temporal variations. In this paper, we present a novel framework to boost the spatio-temporal representation learning by Local and Global Diffusion (LGD). Specifically, we construct a novel neural network architecture that learns the local and global representations in parallel. The architecture is composed of LGD blocks, where each block updates local and global features by modeling the diffusions between these two representations. Diffusions effectively interact two aspects of information, i.e., localized and holistic, for more powerful way of representation learning. Furthermore, a kernelized classifier is introduced to combine the representations from two aspects for video recognition. Our LGD networks achieve clear improvements on the large-scale Kinetics-400 and Kinetics-600 video classification datasets against the best competitors by 3.5% and 0.7%. We further examine the generalization of both the global and local representations produced by our pre-trained LGD networks on four different benchmarks for video action recognition and spatio-temporal action detection tasks. Superior performances over several state-of-the-art techniques on these benchmarks are reported.
Zhaofan Qiu, Ting Yao 0003, Chong-Wah Ngo, Xinmei Tian 0001, Tao Mei 0001
CVPR4
2019 Quantization Networks
abstract
Although deep neural networks are highly effective, their high computational and memory costs severely hinder their applications to portable devices. As a consequence, lowbit quantization, which converts a full-precision neural network into a low-bitwidth integer version, has been an active and promising research topic. Existing methods formulate the low-bit quantization of networks as an approximation or optimization problem. Approximation-based methods confront the gradient mismatch problem, while optimizationbased methods are only suitable for quantizing weights and can introduce high computational cost during the training stage. In this paper, we provide a simple and uniform way for weights and activations quantization by formulating it as a differentiable non-linear function. The quantization function is represented as a linear combination of several Sigmoid functions with learnable biases and scales that could be learned in a lossless and end-to-end manner via continuous relaxation of the steepness of Sigmoid functions. Extensive experiments on image classification and object detection tasks show that our quantization networks outperform state-of-the-art methods. We believe that the proposed method will shed new lights on the interpretation of neural network quantization.
Jiwei Yang, Xu Shen 0001, Jun Xing, Xinmei Tian 0001, Houqiang Li, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001
CVPR4
2019 KCNN: Kernel-wise Quantization to Remarkably Decrease Multiplications in Convolutional Neural Network
abstract
Convolutional neural networks (CNNs) have demonstrated state-of-the-art performance in computer vision tasks. However, the high computational power demand of running devices of recent CNNs has hampered many of their applications. Recently, many methods have quantized the floating-point weights and activations to fixed-points or binary values to convert fractional arithmetic to integer or bit-wise arithmetic. However, since the distributions of values in CNNs are extremely complex, fixed-points or binary values lead to numerical information loss and cause performance degradation. On the other hand, convolution is composed of multiplications and accumulation, but the implementation of multiplications in hardware is more costly comparing with accumulation. We can preserve the rich information of floating-point values on dedicated low power devices by considerably decreasing the multiplications. In this paper, we quantize the floating-point weights in each kernel separately to multiple bit planes to remarkably decrease multiplications. We obtain a closed-form solution via an aggressive Lloyd algorithm and the fine-tuning is adopted to optimize the bit planes. Furthermore, we propose dual normalization to solve the pathological curvature problem during fine-tuning. Our quantized networks show negligible performance loss compared to their floating-point counterparts.
Linghua Zeng, Zhangcheng Wang, Xinmei Tian 0001
IJCAI3
2019 Mocycle-GAN: Unpaired Video-to-Video Translation
abstract
Unsupervised image-to-image translation is the task of translating an image from one domain to another in the absence of any paired training examples and tends to be more applicable to practical applications. Nevertheless, the extension of such synthesis from image-to-image to video-to-video is not trivial especially when capturing spatio-temporal structures in videos. The difficulty originates from the aspect that not only the visual appearance in each frame but also motion between consecutive frames should be realistic and consistent across transformation. This motivates us to explore both appearance structure and temporal continuity in video synthesis. In this paper, we present a new Motion-guided Cycle GAN, dubbed as Mocycle-GAN, that novelly integrates motion estimation into unpaired video translator. Technically, Mocycle-GAN capitalizes on three types of constrains: adversarial constraint discriminating between synthetic and real frame, cycle consistency encouraging an inverse translation on both frame and motion, and motion translation validating the transfer of motion between consecutive frames. Extensive experiments are conducted on video-to-labels and labels-to-video translation, and superior results are reported when comparing to state-of-the-art methods. More remarkably, we qualitatively demonstrate our Mocycle-GAN for both flower-to-flower and ambient condition transfer.
Yang Chen 0048, Yingwei Pan, Ting Yao 0003, Xinmei Tian 0001, Tao Mei 0001
ACM Multimedia4
2019 Animating Your Life: Real-Time Video-to-Animation Translation
abstract
We demonstrate a video-to-animation translator, which can transform real-world video into cartoon or ink-wash animation in real-time. When users upload a video or record what they are seeing with the phone, the video-to-animation translator renders the live streaming video with cartoon or ink-wash animation style while maintaining the original contents. We formulate this task as video-to-video translation problem in the absence of any paired training examples, since the manual labeling of such paired video-animation data is cost-expensive and even unrealistic in practice. Technically, an unified unpaired video-to-video translator is utilized to explore both appearance structure and temporal continuity in video synthesis. As such, not only the visual appearance in each frame but also motion between consecutive frames are ensured to be realistic and consistent for video translation. Based on these technologies, our demonstration can be conducted on any videos in the wild and supports live video-to-animation translation, which engages users with the animated artistic expression of their life.
Yang Chen 0048, Yingwei Pan, Ting Yao 0003, Xinmei Tian 0001, Tao Mei 0001
ACM Multimedia4
2019 RC-CNN: Representation-Consistent Convolutional Neural Networks for Achieving Transformation Invariance
abstract
Convolutional neural networks (CNNs) are powerful and have achieved state-of-the-art performance in many visual recognition tasks. Despite their impressive performance, CNNs are still unable to remain invariant while some spatial transformations are applied on images. Herein, we propose representation-consistent neural networks to solve this problem. By introducing consistent losses between the representations in different layers of transformed images, the recognition performance of transformed images is significantly improved. This model not only learns to map from the transformed images to the pre-defined labels but each layer also learns to generate invariant representations when the input images are transformed. All the characteristics of transformation invariance are embedded in the model, which means that no extra parameters or computations are introduced in the well-trained model. Comparative experiments demonstrate the superiority of our model when learning invariance to rotation, translation, and scaling on large-scale image recognition and retrieval tasks.
Anfeng He, Xinmei Tian 0001
SMC3
2019 Eigenfunction-Based Multitask Learning in a Reproducing Kernel Hilbert Space
abstract
Multitask learning aims to improve the performance on related tasks by exploring the interdependence among them. Existing multitask learning methods explore the relatedness among tasks on the basis of the input features and the model parameters. In this paper, we focus on nonparametric multitask learning and propose to measure task relatedness from a novel perspective in a reproducing kernel Hilbert space (RKHS). Past works have shown that the objective function for a given task can be approximated using the top eigenvalues and corresponding eigenfunctions of a predefined integral operator on an RKHS. In our method, we formulate our objective for multitask learning as a linear combination of two sets of eigenfunctions, common eigenfunctions shared by different tasks and unique eigenfunctions in individual tasks, such that the eigenfunctions for one task can provide additional information on another and help to improve its performance. We present both theoretical and empirical validations of our proposed approach. The theoretical analysis demonstrates that our learning algorithm is uniformly argument stable and that the convergence rate of the generalization upper bound can be improved by learning multiple tasks. Experiments on several benchmark multitask learning data sets show that our method yields promising results.
Xinmei Tian 0001, Tongliang Liu, Xinchao Wang, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.1
2019 Eigenvector-Based Distance Metric Learning for Image Classification and Retrieval
abstract
Distance metric learning has been widely studied in multifarious research fields. The mainstream approaches learn a Mahalanobis metric or learn a linear transformation. Recent related works propose learning a linear combination of base vectors to approximate the metric. In this way, fewer variables need to be determined, which is efficient when facing high-dimensional data. Nevertheless, such works obtain base vectors using additional data from related domains or randomly generate base vectors. However, obtaining base vectors from related domains requires extra time and additional data, and random vectors introduce randomness into the learning process, which requires sufficient random vectors to ensure the stability of the algorithm. Moreover, the random vectors cannot capture the rich information of the training data, leading to a degradation in performance. Considering these drawbacks, we propose a novel distance metric learning approach by introducing base vectors explicitly learned from training data. Given a specific task, we can make a sparse approximation of its objective function using the top eigenvalues and corresponding eigenvectors of a predefined integral operator on the reproducing kernel Hilbert space. Because the process of generating eigenvectors simply refers to the training data of the considered task, our proposed method does not require additional data and can reflect the intrinsic information of the input features. Furthermore, the explicitly learned eigenvectors do not result in randomness, and we can extend our method to any kernel space without changing the objective function. We only need to learn the coefficients of these eigenvectors, and the only hyperparameter that we need to determine is the number of eigenvectors that we utilize. Additionally, an optimization algorithm is proposed to efficiently solve this problem. Extensive experiments conducted on several datasets demonstrate the effectiveness of our proposed method.
Zhangcheng Wang, Richang Hong, Xinmei Tian 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2019 Multi-source Multi-level Attention Networks for Visual Question Answering
abstract
In recent years, Visual Question Answering (VQA) has attracted increasing attention due to its requirement on cross-modal understanding and reasoning of vision and language. VQA is proposed to automatically answer natural language questions with reference to a given image. VQA is challenging, because the reasoning process on a visual domain needs a full understanding of the spatial relationship, semantic concepts, as well as the common sense for a real image. However, most existing approaches jointly embed the abstract low-level visual features and high-level question features to infer answers. These works have limited reasoning ability due to the lack of modeling of the rich spatial context of regions, high-level semantics of images, and knowledge across multiple sources. To solve the challenges, we propose multi-source multi-level attention networks for visual question answering that can benefit both spatial inferences by visual attention on context-aware region representation and reasoning by semantic attention on concepts as well as external knowledge. Indeed, we learn to reason on image representation by question-guided attention at different levels across multiple sources, including region and concept level representation from image source as well as sentence level representation from the external knowledge base. First, we encode region-based middle-level outputs from Convolutional Neural Networks (CNNs) into spatially embedded representation by a multi-directional two-dimensional recurrent neural network and, further, locate the answer-related regions by Multiple Layer Perceptron as visual attention. Second, we generate semantic concepts from high-level semantics in CNNs and select those question-related concepts as concept attention. Third, we query semantic knowledge from the general knowledge base by concepts and selected question-related knowledge as knowledge attention. Finally, we jointly optimize visual attention, concept attention, knowledge attention, and question embedding by a softmax classifier to infer the final answer. Extensive experiments show the proposed approach achieved significant improvement on two very challenging VQA datasets.
Dongfei Yu, Jianlong Fu, Xinmei Tian 0001, Tao Mei 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2018 Domain Generalization via Conditional Invariant Representations
abstract
Domain generalization aims to apply knowledge gained from multiple labeled source domains to unseen target domains. The main difficulty comes from the dataset bias: training data and test data have different distributions, and the training set contains heterogeneous samples from different distributions. Let X denote the features, and Y be the class labels. Existing domain generalization methods address the dataset bias problem by learning a domain-invariant representation h(X) that has the same marginal distribution P(h(X)) across multiple source domains. The functional relationship encoded in P(Y|X) is usually assumed to be stable across domains such that P(Y|h(X)) is also invariant. However, it is unclear whether this assumption holds in practical problems. In this paper, we consider the general situation where both P(X) and P(Y|X) can change across all domains. We propose to learn a feature representation which has domain-invariant class conditional distributions P(h(X)|Y). With the conditional invariant representation, the invariance of the joint distribution P(h(X),Y) can be guaranteed if the class prior P(Y) does not change across training and test domains. Extensive experiments on both synthetic and real data demonstrate the effectiveness of the proposed method.
Mingming Gong, Xinmei Tian 0001, Tongliang Liu, Dacheng Tao
AAAI3
2018 Sequence-to-Sequence Learning via Shared Latent Representation
abstract
Sequence-to-sequence learning is a popular research area in deep learning, such as video captioning and speech recognition. Existing methods model this learning as a mapping process by first encoding the input sequence to a fixed-sized vector, followed by decoding the target sequence from the vector. Although simple and intuitive, such mapping model is task-specific, unable to be directly used for different tasks. In this paper, we propose a star-like framework for general and flexible sequence-to-sequence learning, where different types of media contents (the peripheral nodes) could be encoded to and decoded from a shared latent representation (SLR) (the central node). This is inspired by the fact that human brain could learn and express an abstract concept in different ways. The media-invariant property of SLR could be seen as a high-level regularization on the intermediate vector, enforcing it to not only capture the latent representation intra each individual media like the auto-encoders, but also their transitions like the mapping models. Moreover, the SLR model is content-specific, which means it only needs to be trained once for a dataset, while used for different tasks. We show how to train a SLR model via dropout and use it for different sequence-to-sequence tasks. Our SLR model is validated on the Youtube2Text and MSR-VTT datasets, achieving superior performance on video-to-sentence task, and the first sentence-to-video results.
Xu Shen 0001, Xinmei Tian 0001, Jun Xing, Yong Rui, Dacheng Tao
AAAI2
2018 A Twofold Siamese Network for Real-Time Object Tracking
abstract
Observing that Semantic features learned in an image classification task and Appearance features learned in a similarity matching task complement each other, we build a twofold Siamese network, named SA-Siam, for real-time object tracking. SA-Siam is composed of a semantic branch and an appearance branch. Each branch is a similaritylearning Siamese network. An important design choice in SA-Siam is to separately train the two branches to keep the heterogeneity of the two types of features. In addition, we propose a channel attention mechanism for the semantic branch. Channel-wise weights are computed according to the channel activations around the target position. While the inherited architecture from SiamFC [3] allows our tracker to operate beyond real-time, the twofold design and the attention mechanism significantly improve the tracking performance. The proposed SA-Siam outperforms all other real-time trackers by a large margin on OTB-2013/50/100 benchmarks.
Anfeng He, Chong Luo 0001, Xinmei Tian 0001, Wenjun Zeng 0001
CVPR3
2018 Deep Boosting for Image Denoising
Chang Chen 0004, Zhiwei Xiong, Xinmei Tian 0001, Feng Wu 0001
ECCV (11)3
2018 Deep Domain Generalization via Conditional Invariant Adversarial Networks
Xinmei Tian 0001, Mingming Gong, Tongliang Liu, Kun Zhang 0001, Dacheng Tao
ECCV (15)2
2018 Local Convolutional Neural Networks for Person Re-Identification
abstract
Recent works have shown that person re-identification can be substantially improved by introducing attention mechanisms, which allow learning both global and local representations. However, all these works learn global and local features in separate branches. As a consequence, the interaction/boosting of global and local information are not allowed, except in the final feature embedding layer. In this paper, we propose local operations as a generic family of building blocks for synthesizing global and local information in any layer. This building block can be inserted into any convolutional networks with only a small amount of prior knowledge about the approximate locations of local parts. For the task of person re-identification, even with only one local block inserted, our local convolutional neural networks (Local CNN) can outperform state-of-the-art methods consistently on three large-scale benchmarks, including Market-1501, CUHK03, and DukeMTMC-ReID.
Jiwei Yang, Xu Shen 0001, Xinmei Tian 0001, Houqiang Li, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ACM Multimedia3
2018 Deep Domain Adaptation Hashing with Adversarial Learning
abstract
The recent advances in deep neural networks have demonstrated high capability in a wide variety of scenarios. Nevertheless, fine-tuning deep models in a new domain still requires a significant amount of labeled data despite expensive labeling efforts. A valid question is how to leverage the source knowledge plus unlabeled or only sparsely labeled target data for learning a new model in target domain. The core problem is to bring the source and target distributions closer in the feature space. In the paper, we facilitate this issue in an adversarial learning framework, in which a domain discriminator is devised to handle domain shift. Particularly, we explore the learning in the context of hashing problem, which has been studied extensively due to its great efficiency in gigantic data. Specifically, a novel Deep Domain Adaptation Hashing with Adversarial learning (DeDAHA) architecture is presented, which mainly consists of three components: a deep convolutional neural networks (CNN) for learning basic image/frame representation followed by an adversary stream on one hand to optimize the domain discriminator, and on the other, to interact with each domain-specific hashing stream for encoding image representation to hash codes. The whole architecture is trained end-to-end by jointly optimizing two types of losses, i.e., triplet ranking loss to preserve the relative similarity ordering in the input triplets and adversarial loss to maximally fool the domain discriminator with the learnt source and target feature distributions. Extensive experiments are conducted on three domain transfer tasks, including cross-domain digits retrieval, image to image and image to video transfers, on several benchmarks. Our DeDAHA framework achieves superior results when compared to the state-of-the-art techniques.
Fuchen Long, Ting Yao 0003, Qi Dai 0001, Xinmei Tian 0001, Jiebo Luo 0001, Tao Mei 0001
SIGIR4
2018 Relative Aesthetic Quality Ranking
abstract
Given a set of images, we want to rank them according to their aesthetic quality, which is beneficial to various applications. Most existing methods consider aesthetic quality assessment to be a classification or regression problem, but describing aesthetic quality using absolute labels or scores is imprecise and unnatural. To solve this problem, some researchers have proposed listwise approaches for relative aesthetic quality ranking. However, only limited success has been achieved because 1) the ranking order of all images is employed in training although it is not reasonable to compare images with different visual content and 2) pre-defined features cannot describe the aesthetic attributes well. To address these challenges, we introduce a novel visual similarity-based approach to generate reasonable pairs in this paper. With these well-selected pairs, a ranking model can be trained to rank images according to their aesthetic quality. Furthermore, to learn more effective aesthetic features and obtain a better ranking function, we design a dual-channel deep neural network as the ranking model. The proposed approach is evaluated on the AVA dataset, and the experimental results demonstrate that our approach significantly outperforms the state-of-the-art methods.
Xinmei Tian 0001, Yujiao Long
SMC1
2018 Joints kinetic and relational features for action recognition
Xinmei Tian 0001
Signal Process.1
2018 Data mining in human activity analysis
Xinmei Tian 0001, Weifeng Liu 0001, Fionn Murtagh
Signal Process.1
2018 PageSense: Toward Stylewise Contextual Advertising via Visual Analysis of Web Pages
abstract
The Internet has emerged as the most effective and a highly popular medium for advertising. Current contextual advertising platforms need publishers to manually change the original structure of their Web pages and predefine the position and style of embedded ads. Although publishers spend significant effort optimizing their Web page layout, a large number of Web pages contain noticeable blank regions. We present an innovative stylewise advertising platform for contextual advertising, called PageSense. The “style” of Web pages refers to the visual appearance of a Web page, such as color and layout. PageSense aims to associate style-consistent ads with Web pages. It provides two advertising options: 1) If publishers predefine ad positions within Web pages, PageSense will analyze the page style and select ads, which are consistent with the Web page layout, and 2) if publishers impose no constraints for ad placement, PageSense will automatically detect blank regions, select the most nonintrusive region for ad insertion, associate color-consistent ads with the Web pages, and deliver them to blank regions without breaking the original Web page style. Our experiments have verified the effectiveness of PageSense as a complement to existing contextual advertising.
Tao Mei 0001, Lusong Li, Xinmei Tian 0001, Dacheng Tao, Chong-Wah Ngo
IEEE Trans. Circuits Syst. Video Technol.3
2018 Multigranular Event Recognition of Personal Photo Albums
abstract
People are taking more photos than ever before in recent years. To effectively organize these personal photos, the photos are usually assigned to albums according to their events. An efficient way to manage our photos would be if we could recognize the events of the albums automatically. In this paper, we study the problem of recognizing events in personal photo albums. Recognizing events in photo albums is a new challenge since the contents of photos in albums are more complicated than in traditional single-photo tasks, since not all photos in an album are relevant to the event and a single photo in an album often fails to convey the meaningful event semantic behind the album. To solve this problem, we introduce an attention network to learn the representations of photo albums. Then, we adopt a hierarchical model to recognize events from coarse to fine using multigranular features. We evaluate our model on two real-world datasets consisting of personal albums; we find that our model achieves promising results.
Cong Guo 0002, Xinmei Tian 0001, Tao Mei 0001
IEEE Trans. Multim.2
2018 On Better Exploring and Exploiting Task Relationships in Multitask Learning: Joint Model and Feature Learning
abstract
Multitask learning (MTL) aims to learn multiple tasks simultaneously through the interdependence between different tasks. The way to measure the relatedness between tasks is always a popular issue. There are mainly two ways to measure relatedness between tasks: common parameters sharing and common features sharing across different tasks. However, these two types of relatedness are mainly learned independently, leading to a loss of information. In this paper, we propose a new strategy to measure the relatedness that jointly learns shared parameters and shared feature representations. The objective of our proposed method is to transform the features of different tasks into a common feature space in which the tasks are closely related and the shared parameters can be better optimized. We give a detailed introduction to our proposed MTL method. Additionally, an alternating algorithm is introduced to optimize the nonconvex objection. A theoretical bound is given to demonstrate that the relatedness between tasks can be better measured by our proposed MTL algorithm. We conduct various experiments to verify the superiority of the proposed joint model and feature MTL method.
Xinmei Tian 0001, Tongliang Liu, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.2
2018 Continuous Dropout
abstract
Dropout has been proven to be an effective algorithm for training robust deep networks because of its ability to prevent overfitting by avoiding the co-adaptation of feature detectors. Current explanations of dropout include bagging, naive Bayes, regularization, and sex in evolution. According to the activation patterns of neurons in the human brain, when faced with different situations, the firing rates of neurons are random and continuous, not binary as current dropout does. Inspired by this phenomenon, we extend the traditional binary dropout to continuous dropout. On the one hand, continuous dropout is considerably closer to the activation characteristics of neurons in the human brain than traditional binary dropout. On the other hand, we demonstrate that continuous dropout has the property of avoiding the co-adaptation of feature detectors, which suggests that we can extract more independent feature detectors for model averaging in the test stage. We introduce the proposed continuous dropout to a feedforward neural network and comprehensively compare it with binary dropout, adaptive dropout, and DropConnect on Modified National Institute of Standards and Technology, Canadian Institute for Advanced Research-10, Street View House Numbers, NORB, and ImageNet large scale visual recognition competition-12. Thorough experiments demonstrate that our method performs better in preventing the co-adaptation of feature detectors and improves test performance.
Xu Shen 0001, Xinmei Tian 0001, Tongliang Liu, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.2
2017 Patch Reordering: A NovelWay to Achieve Rotation and Translation Invariance in Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNNs) have demonstrated state-of-the-art performance on many visual recognition tasks. However, the combination of convolution and pooling operations only shows invariance to small local location changes in meaningful objects in input. Sometimes, such networks are trained using data augmentation to encode this invariance into the parameters, which restricts the capacity of the model to learn the content of these objects. A more efficient use of the parameter budget is to encode rotation or translation invariance into the model architecture, which relieves the model from the need to learn them. To enable the model to focus on learning the content of objects other than their locations, we propose to conduct patch ranking of the feature maps before feeding them into the next layer. When patch ranking is combined with convolution and pooling operations, we obtain consistent representations despite the location of meaningful objects in input. We show that the patch ranking module improves the performance of the CNN on many benchmark tasks, including MNIST digit recognition, large-scale image recognition, and image retrieval.
Xu Shen 0001, Xinmei Tian 0001, Shaoyan Sun, Dacheng Tao
AAAI2
2017 A Small Scale Multi-Column Network for Aesthetic Classification Based on Multiple Attributes
Chaoqun Wan, Xinmei Tian 0001
ICONIP (1)2
2017 Semi-supervised Coefficient-Based Distance Metric Learning
Zhangcheng Wang, Xinmei Tian 0001
ICONIP (1)3
2017 Layer-Wise Training to Create Efficient Convolutional Neural Networks
Linghua Zeng, Xinmei Tian 0001
ICONIP (2)2
2017 Classification and Representation Joint Learning via Deep Networks
abstract
Deep learning has been proven to be effective for classification problems. However, the majority of previous works trained classifiers by considering only class label information and ignoring the local information from the spatial distribution of training samples. In this paper, we propose a deep learning framework that considers both class label information and local spatial distribution information between training samples. A two-channel network with shared weights is used to measure the local distribution. The classification performance can be improved with more detailed information provided by the local distribution, particularly when the training samples are insufficient. Additionally, the class label information can help to learn better feature representations compared with other feature learning methods that use only local distribution information between samples. The local distribution constraint between sample pairs can also be viewed as a regularization of the network, which can efficiently prevent the overfitting problem. Extensive experiments are conducted on several benchmark image classification datasets, and the results demonstrate the effectiveness of our proposed method.
Xinmei Tian 0001, Xu Shen 0001, Dacheng Tao
IJCAI2
2016 Regularized Large Margin Distance Metric Learning
abstract
Distance metric learning plays an important role in many applications, such as classification and clustering. In this paper, we propose a novel distance metric learning using two hinge losses in the objective function. One is the constraint of the pairs which makes the similar pairs (the same label) closer and the dissimilar (different labels) pairs separated as far as possible. The other one is the constraint of the triplets which makes the largest distance between pairs intra the class larger than the smallest distance between pairs inter the classes. Previous works only consider one of the two kinds of constraints. Additionally, different from the triplets used in previous works, we just need a small amount of such special triplets. This improves the efficiency of our proposed method. Consider the situation in which we might not have enough labeled samples, we extend the proposed distance metric learning into a semi-supervised learning framework. Experiments are conducted on several landmark datasets and the results demonstrate the effectiveness of our proposed method.
Xinmei Tian 0001, Dacheng Tao
ICDM2
2016 Transform-Invariant Convolutional Neural Networks for Image Classification and Search
abstract
Convolutional neural networks (CNNs) have achieved state-of-the-art results on many visual recognition tasks. However, current CNN models still exhibit a poor ability to be invariant to spatial transformations of images. Intuitively, with sufficient layers and parameters, hierarchical combinations of convolution (matrix multiplication and non-linear activation) and pooling operations should be able to learn a robust mapping from transformed input images to transform-invariant representations. In this paper, we propose randomly transforming (rotation, scale, and translation) feature maps of CNNs during the training stage. This prevents complex dependencies of specific rotation, scale, and translation levels of training images in CNN models. Rather, each convolutional kernel learns to detect a feature that is generally helpful for producing the transform-invariant answer given the combinatorially large variety of transform levels of its input feature maps. In this way, we do not require any extra training supervision or modification to the optimization process and training images. We show that random transformation provides significant improvements of CNNs on many benchmark tasks, including small-scale image recognition, large-scale image recognition, and image retrieval.
Xu Shen 0001, Xinmei Tian 0001, Anfeng He, Shaoyan Sun, Dacheng Tao
ACM Multimedia2
2016 Visual Re-ranking Through Greedy Selection and Rank Fusion
Bin Lin 0010, Ai Wei, Xinmei Tian 0001
MMM (1)3
2016 Learning Relative Aesthetic Quality with a Pairwise Approach
Xinmei Tian 0001
MMM (1)2
2016 Multi-organ plant identification with multi-column deep convolutional neural networks
abstract
Automatically identifying plants from images is a hot research topic due to its importance in production and science popularization. This process attempts to automatically identify the name of a plant with a known taxon from a given image. The majority of existing studies on automatic plant identification focus on identifying plants with a single organ, such as flower, leaf, or fruits. Plant identification using a single organ is not sufficiently reliable because different plants many have similar organs. To overcome this problem, this paper is devoted to automatically identifying plants by combining multiple organs of plants. Specifically, we propose a multi-column deep convolutional neural networks (MCDCNN) model to combine multiple organs for efficient plant identification. Extensive experiments demonstrate the effectiveness of our model, and the plant identification performance is greatly improved.
Anfeng He, Xinmei Tian 0001
SMC2
2016 Flickr group recommendation using rich social media information
Cong Guo 0002, Xinmei Tian 0001
Neurocomputing3
2016 Multi-modal and multi-scale photo collection summarization
Xu Shen 0001, Xinmei Tian 0001
Multim. Tools Appl.2
2016 Monet: A System for Reliving Your Memories by Theme-Based Photo Storytelling
abstract
With the ever-increasing use of smartphones and digital cameras, people are now able to take photos anywhere and anytime. Most of these photos simply end up stored in the cloud without further interaction. This occurs because we lack intelligent services to organize these personal photos well. Therefore, there is an urgent need for such a system to enable people to relive their memories by turning their photos into stories. This paper presents a storytelling system named Monet, which automatically creates interesting stories from personal photos by mimicking cinematic knowledge based on a set of predesigned editing styles. The system consists of two stages: photo summarization, which selects a subset of the “best” photos to represent a photo collection, and story remixing, which generates a stylish music video from the selected photos. During photo summarization, photos are grouped into events based on multimodal features (time and location). The “best” photos are then selected according to visual quality, event representativeness, and diversity. The second stage, story remixing, automatically selects an appropriate theme-dependent editing style based on the photo content. Each selected photo is converted to a video clip by applying a virtual camera with appropriate motions. A series of video effects, color filters, shapes, and transitions are then applied to the video clips according to cinematic rules. The generated video is finally multiplexed with a music clip to generate the story. Evaluations show that our system achieves superior performance to state-of-the-art photo event detection and story generation systems.
Xu Shen 0001, Tao Mei 0001, Xinmei Tian 0001, Nenghai Yu, Yong Rui
IEEE Trans. Multim.4
2015 On the selection of trending image from the web
abstract
The recommendation of trending images has become a popular feature used by commercial search engines to attract public attention. By browsing through trending images, search engine users can discover trending events at a glance. However, the selection of trending images is very challenging and remains an open issue. Most existing work is highly dependent on editorial efforts, though some preliminarily identify a few plain features for trending images. In this paper, we investigate a set of perceptual factors that can distinguish trending images from common ones. We propose a set of trending-aware features based on several common criteria, which reflect the characteristics of trending images. We further construct a manually labeled dataset based on a commercial search engine's query log over a two-week timespan. We evaluate our proposed method on this dataset and the results demonstrate its effectiveness.
Dongfei Yu, Xinmei Tian 0001, Tao Mei 0001, Yong Rui
ICME2
2015 Multi-Task Model and Feature Joint Learning
Xinmei Tian 0001, Tongliang Liu, Dacheng Tao
IJCAI2
2015 Photo Quality Assessment with DCNN that Understands Image Well
Xu Shen 0001, Houqiang Li, Xinmei Tian 0001
MMM (2)4
2015 Event recognition in personal photo collections using hierarchical model and multiple features
abstract
With the proliferation of digital cameras and mobile devices, people are taking many more photos than ever before. The explosive growth of personal photos leads to problems of photo organization and management. There is a growing need for tools to automatically manage photo collections. Recognizing events in photo collections is one efficient way to organize photos. The use of textual event labels can allow us to categorize and locate an event without browsing through an entire photo collection. Most existing research on this topic focuses on recognizing events from single photos and only a few studies have examined event recognition in personal photo collections. In this paper, we propose a hierarchical model to recognize events in personal photo collections using multiple features, including time, objects, and scenes. Since some events are more difficult to identify and categorize, ambiguous events require fine event classifiers, while the coarse categories of the events can be sufficiently organized with a coarse event classifier. We evaluate our coarse-to-fine hierarchical model on a real-world dataset consisting of personal photo collections, and our model achieves promising results.
Cong Guo 0002, Xinmei Tian 0001
MMSP2
2015 Query-Adaptive Image Search Re-ranking Using Deep Convolutional Neural Network Feature
abstract
Image search re-ranking, as an effective tool to improve the text-based image search result, has been adopted by many commercial search engines nowadays. Given a query keyword, images are first retrieved based on the textual information. Then visual features are extracted from images to reorder them by mining their visual patterns. However, the popular visual features applied in re-ranking are not informative enough. Besides, the parameters for the re-ranking models are set equally for all queries, which fails to cope with the variability of different queries. In this paper, we propose a novel re-ranking method which adopts informative visual features for image representation and adaptively re-rank the images. Specifically, we adopt a proven successful DCNN feature (deep convolutional neural network), which shows the excellent performance in many computer vision fields, to calculate the visual similarities between images. For each query, the parameters for the image search re-ranking model is adaptively determined using the QDE (query difficulty estimation) method. Experiments are conducted on the INRIA web353 dataset. The experimental results demonstrate that our method achieves significant improvement over state-of-the-art methods.
Bin Lin 0010, Xinmei Tian 0001
SMC2
2015 Multi-level photo quality assessment with multi-view features
Xinmei Tian 0001
Neurocomputing2
2015 Multi-task proximal support vector machine
Xinmei Tian 0001, Mingli Song, Dacheng Tao
Pattern Recognit.2
2015 Query difficulty estimation via relevance prediction for image retrieval
Qianghuai Jia, Xinmei Tian 0001
Signal Process.2
2015 Exploration of Image Search Results Quality Assessment
abstract
Image retrieval plays an increasingly important role in our daily lives. There are many factors which affect the quality of image search results, including chosen search algorithms, ranking functions, and indexing features. Applying different settings for these factors generates search result lists with varying levels of quality. However, no setting can always perform optimally for all queries. Therefore, given a set of search result lists generated by different settings, it is crucial to automatically determine which result list is the best in order to present it to users. This paper aims to solve this problem and makes four main innovations. First, a preference learning model is proposed to quantitatively study and formulate the best image search result list identification problem. Second, a set of valuable preference learning related features is proposed by exploring the visual characters of returned images. Third, a query-dependent preference learning model is further designed for building a more precise and query-specific model. Fourth, the proposed approach has been tested on a variety of applications including re-ranking ability assessment, optimal search engine selection, and synonymous query suggestion. Extensive experimental results on three image search datasets demonstrate the effectiveness and promising potential of the proposed method.
Xinmei Tian 0001, Yijuan Lu, Nate Stender, Linjun Yang, Dacheng Tao
IEEE Trans. Big Data1
2015 Image Search Reranking With Hierarchical Topic Awareness
abstract
With much attention from both academia and industrial communities, visual search reranking has recently been proposed to refine image search results obtained from text-based image search engines. Most of the traditional reranking methods cannot capture both relevance and diversity of the search results at the same time. Or they ignore the hierarchical topic structure of search result. Each topic is treated equally and independently. However, in real applications, images returned for certain queries are naturally in hierarchical organization, rather than simple parallel relation. In this paper, a new reranking method "topic-aware reranking (TARerank)" is proposed. TARerank describes the hierarchical topic structure of search results in one model, and seamlessly captures both relevance and diversity of the image search results simultaneously. Through a structured learning framework, relevance and diversity are modeled in TARerank by a set of carefully designed features, and then the model is learned from human-labeled training samples. The learned model is expected to predict reranking results with high relevance and diversity for testing queries. To verify the effectiveness of the proposed method, we collect an image search dataset and conduct comparison experiments on it. The experimental results demonstrate that the proposed TARerank outperforms the existing relevance-based and diversified reranking methods.
Xinmei Tian 0001, Linjun Yang, Yijuan Lu, Qi Tian 0001, Dacheng Tao
IEEE Trans. Cybern.1
2015 Query-Dependent Aesthetic Model With Deep Learning for Photo Quality Assessment
abstract
The automatic assessment of photo quality from an aesthetic perspective is a very challenging problem. Most existing research has predominantly focused on the learning of a universal aesthetic model based on hand-crafted visual descriptors . However, this research paradigm can achieve only limited success because (1) such hand-crafted descriptors cannot well preserve abstract aesthetic properties , and (2) such a universal model cannot always capture the full diversity of visual content. To address these challenges, we propose in this paper a novel query-dependent aesthetic model with deep learning for photo quality assessment. In our method, deep aesthetic abstractions are discovered from massive images , whereas the aesthetic assessment model is learned in a query- dependent manner. Our work addresses the first problem by learning mid-level aesthetic feature abstractions via powerful deep convolutional neural networks to automatically capture the underlying aesthetic characteristics of the massive training images . Regarding the second problem, because photographers tend to employ different rules of photography for capturing different images , the aesthetic model should also be query- dependent . Specifically, given an image to be assessed, we first identify which aesthetic model should be applied for this particular image. Then, we build a unique aesthetic model of this type to assess its aesthetic quality. We conducted extensive experiments on two large-scale datasets and demonstrated that the proposed query-dependent model equipped with learned deep aesthetic abstractions significantly and consistently outperforms state-of-the-art hand-crafted feature -based and universal model-based methods.
Xinmei Tian 0001, Kuiyuan Yang, Tao Mei 0001
IEEE Trans. Multim.1
2015 Query Difficulty Estimation for Image Search With Query Reconstruction Error
abstract
Current image search engines suffer from a radical variance in retrieval performance over different queries. It is therefore desirable to identify those “difficult” queries in order to handle them properly. Query difficulty estimation is an attempt to predict the performance of the search results returned by an image search system. Most existing methods for query difficulty estimation focus on investigating statistical characteristics of the returned images only, while neglecting very important information , i.e., the query and its relationship with returned images. This relationship plays a crucial role in query difficulty estimation and should be explored further. In this paper we propose a novel query difficulty estimation method with query reconstruction error. This method is proposed based on the observation that, given the images returned for an unknown query, we can easily deduce what the query is from those images if the search results are high quality (i.e., lots of relevant images returned); otherwise, it is difficult to deduce the original query. Therefore, we propose to predict the query difficulty by measuring to what extent the original query can be recovered from the image search results. Specifically, we first reconstruct a visual query from the returned images to summarize their visual theme, and then use the reconstruction error, i.e., the distance between the original textual query and the reconstructed visual query, to estimate the query difficulty. We conduct extensive experiments on two real-world Web image datasets and demonstrate the effectiveness of the proposed method.
Xinmei Tian 0001, Qianghuai Jia, Tao Mei 0001
IEEE Trans. Multim.1
2014 User specific friend recommendation in social media community
abstract
Social networks nowadays have become an important form of communication in which users can post their current status or share their lives by mobile phones or the Web. In this paper, we develop an effective and efficient model to estimate continuous tie strength between users for friend recommendation with the heterogeneous data from social media community. We categorize those multimodal data into two classes: interaction data (e.g., comments, marking favorite photos) and similarity data(e.g., common friends, groups, tags, geo, visual). We propose to use asymmetric relationship in the interaction data for tie strength estimation instead of using the conventional symmetric ones. Furthermore, by exploring the behavior of users in a social media community, we find that the tie strength between users can be approximately modeled as a linear function of their social connections. Based on this observation, we propose an effective and highly efficient user specific linear model for the tie strength estimation. The experiments on a popular social network show promising results and demonstrate the effectiveness of our proposed method.
Cong Guo 0002, Xinmei Tian 0001, Tao Mei 0001
ICME2
2014 Query difficulty estimation via pseudo relevance feedback for image search
abstract
Query difficulty estimation (QDE) attempts to automatically predict the performance of the search results returned for a given query. QDE has been widely investigated in text document retrieval for many years. However, few research works have been explored in image retrieval. State-of-the-art QDE methods in image retrieval mainly investigate the statistical characteristics (coherence, robustness, etc.) of the returned images to derive a value for indicating the query difficulty degree. To the best of our knowledge, little research has been done to directly estimate the real retrieval performance of the search results, such as average precision, instead of only an indicator. In this paper, we propose a novel query difficulty estimation approach which automatically estimate the average precision of the image search results. Specifically, we first select a set of query relevant and query irrelevant images for each query via pseudo relevance feedback. Then an efficient and effective voting scheme is proposed to estimate the relevance label of each image in the search results. Based on the images' relevance labels, the average precision of the search results returned for the given query is derived. The experimental results on a benchmark image search dataset demonstrate the effectiveness of the proposed method.
Qianghuai Jia, Xinmei Tian 0001, Tao Mei 0001
ICME2
2014 Click-through-based Subspace Learning for Image Search
abstract
One of the fundamental problems in image search is to rank image documents according to a given textual query. We address two limitations of the existing image search engines in this paper. First, there is no straightforward way of comparing textual keywords with visual image content. Image search engines therefore highly depend on the surrounding texts, which are often noisy or too few to accurately describe the image content. Second, ranking functions are trained on query-image pairs labeled by human labelers, making the annotation intellectually expensive and thus cannot be scaled~up.
Yingwei Pan, Ting Yao 0003, Xinmei Tian 0001, Houqiang Li, Chong-Wah Ngo
ACM Multimedia3
2014 Effective and efficient photo quality assessment
abstract
Automatic photo quality assessment from the perspective of visual aesthetics is a hot research topic due to its potential need in numerous applications. It tries to automatically determine whether a given image has “high” or “low” quality according to the image's visual content. Most existing researches in photo quality assessment predominantly focus on exploring hand-crafted features which may be potentially related to high-level aesthetic attributes. Most of those features are designed under the guidance of some common photography rules and prior knowledge. However, due to the subjectivity and complexity of humans' aesthetic activities, automatic image aesthetic quality assessment is very challenging. Those features are not effective enough and show varying performance on different datasets. Besides, they often require high computational cost. In this paper, we propose a set of compact aesthetic features which are not only effective but also highly efficient. We test those features on two large scale real world image datasets. The experimental results demonstrate that the proposed features achieve the best performance consistently over different datasets with a much lower computational complexity.
Xinmei Tian 0001
SMC2
2013 Semantic-Spatial Matching for image classification
abstract
Spatial Pyramid Matching (SPM) has been proven a simple but effective extension to bag-of-visual-words image representation for spatial layout information compensation. SPM describes image in coarse-to-fine scale by partitioning the image into blocks over multiple levels and the features extracted from each block are concatenated into a long vector representation. Based on the assumption that images from the same class have similar spatial configurations, SPM matches the blocks from different images according to their spatial layout, by aligning all blocks from an image in a fixed spatial order. However, target objects may appear at any location in the image with various backgrounds. Therefore, the fixed spatial matching in SPM fails to match similar objects located different locations. To solve this problem, we propose an effective and efficient block matching method, Semantic-Spatial Matching (SSM). In this method, not only the spatial layout but also the semantic content is considered for block matching. The experiments on two benchmark image classification datasets demonstrate the effectiveness of SSM.
Yupeng Yan, Xinmei Tian 0001, Linjun Yang, Yijuan Lu, Houqiang Li
ICME2
2013 Discriminative codebook learning for Web image search
Xinmei Tian 0001, Yijuan Lu
Signal Process.1
2012 Query Difficulty Prediction for Web Image Search
abstract
Image search plays an important role in our daily life. Given a query, the image search engine is to retrieve images related to it. However, different queries have different search difficulty levels. For some queries, they are easy to be retrieved (the search engine can return very good search results). While for others, they are difficult (the search results are very unsatisfactory). Thus, it is desirable to identify those “difficult” queries in order to handle them properly. Query difficulty prediction (QDP) is an attempt to predict the quality of the search result for a query over a given collection. QDP problem has been investigated for many years in text document retrieval, and its importance has been recognized in the information retrieval (IR) community. However, little effort has been conducted on the image query difficulty prediction problem for image search. Compared with QDP in document retrieval, QDP in image search is more challenging due to the noise of textual features and the well-known semantic gap of visual features. This paper aims to investigate the QDP problem in Web image search. A novel method is proposed to automatically predict the quality of image search results for an arbitrary query. This model is built based on a set of valuable features that are designed by exploring the visual characteristic of images in the search results. The experiments on two real image search datasets demonstrate the effectiveness of the proposed query difficulty prediction method. Two applications, including optimal image search engine selection and search results merging, are presented to show the promising applicability of QDP.
Xinmei Tian 0001, Yijuan Lu, Linjun Yang
IEEE Trans. Multim.1
2012 Correction to "Bayesian Visual Reranking"
abstract
In the above titled paper (ibid., vol. 13, no. 4, pp. 639-652, Aug. 2011), the first author's name appears incorrectly in the byline as "Xinmie Tian" instead of "Xinmei Tian." The name appears correctly in the biography section.
Xinmei Tian 0001, Linjun Yang, Jingdong Wang 0001, Xiuqing Wu, Xian-Sheng Hua 0001
IEEE Trans. Multim.1
2012 Sparse transfer learning for interactive video search reranking
abstract
Visual reranking is effective to improve the performance of the text-based video search. However, existing reranking algorithms can only achieve limited improvement because of the well-known semantic gap between low-level visual features and high-level semantic concepts. In this article, we adopt interactive video search reranking to bridge the semantic gap by introducing user's labeling effort. We propose a novel dimension reduction tool, termed sparse transfer learning (STL), to effectively and efficiently encode user's labeling information. STL is particularly designed for interactive video search reranking. Technically, it (a) considers the pair-wise discriminative information to maximally separate labeled query relevant samples from labeled query irrelevant ones, (b) achieves a sparse representation for the subspace to encodes user's intention by applying the elastic net penalty, and (c) propagates user's labeling information from labeled samples to unlabeled samples by using the data distribution knowledge. We conducted extensive experiments on the TRECVID 2005, 2006 and 2007 benchmark datasets and compared STL with popular dimension reduction algorithms. We report superior performance by using the proposed STL-based interactive video search reranking.
Xinmei Tian 0001, Dacheng Tao, Yong Rui
ACM Trans. Multim. Comput. Commun. Appl.1
2011 Learning to judge image search results
abstract
Given the explosive growth of the Web and the popularity of image sharing Web sites, image retrieval plays an increasingly important role in our daily lives. Search engines aim to provide beneficial image search results to users in response to queries. The quality of image search results depends on many factors: chosen search algorithms, ranking functions, indexing features, the base image database, etc. Applying different settings for these factors generates search result lists with varying levels of quality. Previous research has shown that no setting can always perform optimally for all queries. Therefore, given a set of search result lists generated by different settings, it is crucial to automatically determine which result list is the best in order to present it to users. This paper proposes a novel method to automatically identify the best search result list from a number of candidates. There are three main innovations in this paper. First, we propose a preference learning model to quantitatively study the best image search result identification problem. Second, we propose a set of valuable preference learning related features by exploring the visual characters of returned images. Third, our method shows promising potential in applications such as reranking ability assessment and optimal search engine selection. Experiments on two image search datasets show that our method achieves about 80% prediction accuracy for reranking ability assessment, and selects optimal search engine for about 70% queries correctly.
Xinmei Tian 0001, Yijuan Lu, Linjun Yang, Qi Tian 0001
ACM Multimedia1
2011 Bayesian Visual Reranking
abstract
Visual reranking has been proven effective to refine text-based video and image search results. It utilizes visual information to recover “true” ranking list from the noisy one generated by text-based search, by incorporating both textual and visual information. In this paper, we model the textual and visual information from the probabilistic perspective and formulate visual reranking as an optimization problem in the Bayesian framework, termed Bayesian visual reranking. In this method, the textual information is modeled as a likelihood, to reflect the disagreement between reranked results and text-based search results which is called ranking distance. The visual information is modeled as a conditional prior, to indicate the ranking score consistency among visually similar samples which is called visual consistency. Bayesian visual reranking derives the best reranking results by maximizing visual consistency while minimizing ranking distance. To model the ranking distance more precisely, we propose a novel pair-wise method which measure the ranking distance based on the disagreement in terms of pair-wise orders. For visual consistency, we study three different regularizers to mine the best way for its modeling. We conduct extensive experiments on both video and image search datasets. Experimental results demonstrate the effectiveness of our proposed Bayesian visual reranking.
Xinmei Tian 0001, Linjun Yang, Jingdong Wang 0001, Xiuqing Wu, Xian-Sheng Hua 0001
IEEE Trans. Multim.1
2010 Constrained Metric Learning Via Distance Gap Maximization
abstract
Vectored data frequently occur in a variety of fields, which are easy to handle since they can be mathematically abstracted as points residing in a Euclidean space. An appropriate distance metric in the data space is quite demanding for a great number of applications. In this paper, we pose robust and tractable metric learning under pairwise constraints that are expressed as similarity judgements between data pairs. The major features of our approach include: 1) it maximizes the gap between the average squared distance among dissimilar pairs and the average squared distance among similar pairs; 2) it is capable of propagating similar constraints to all data pairs; and 3) it is easy to implement in contrast to the existing approaches using expensive optimization such as semidefinite programming. Our constrained metric learning approach has widespread applicability without being limited to particular backgrounds. Quantitative experiments are performed for classification and retrieval tasks, uncovering the effectiveness of the proposed approach.
Wei Liu 0005, Xinmei Tian 0001, Dacheng Tao, Jianzhuang Liu
AAAI2
2010 Visual Reranking with Local Learning Consistency
Xinmei Tian 0001, Linjun Yang, Xiuqing Wu, Xian-Sheng Hua 0001
MMM1
2010 Active Reranking for Web Image Search
abstract
Image search reranking methods usually fail to capture the user's intention when the query term is ambiguous. Therefore, reranking with user interactions, or active reranking, is highly demanded to effectively improve the search performance. The essential problem in active reranking is how to target the user's intention. To complete this goal, this paper presents a structural information based sample selection strategy to reduce the user's labeling efforts. Furthermore, to localize the user's intention in the visual feature space, a novel local-global discriminative dimension reduction algorithm is proposed. In this algorithm, a submanifold is learned by transferring the local geometry and the discriminative information from the labelled images to the whole (global) image database. Experiments on both synthetic datasets and a real Web image search dataset demonstrate the effectiveness of the proposed active reranking scheme, including both the structural information based active sample selection strategy and the local-global discriminative dimension reduction algorithm.
Xinmei Tian 0001, Dacheng Tao, Xian-Sheng Hua 0001, Xiuqing Wu
IEEE Trans. Image Process.1
2009 Query aware visual similarity propagation for image search reranking
abstract
Image search reranking is an effective approach to refining the text-based image search result. In the reranking process, the estimation of visual similarity is critical to the performance. However, the existing measures, based on global or local features, cannot be adapted to different queries. In this paper, we propose to estimate a query aware image similarity by incorporating the global visual similarity, local visual similarity and visual word co-occurrence into an iterative propagation framework. After the propagation, a query aware image similarity combining the advantages of both global and local similarities is achieved and applied to image search reranking. The experiments on a real-world Web image dataset demonstrate that the proposed query aware similarity outperforms the global, local similarity and their linear combination, for image search reranking.
Linjun Yang, Xinmei Tian 0001
ACM Multimedia3
2008 Transductive video annotation via local learnable kernel classifier
abstract
One crucial problem in transductive video annotation is how to estimate the label from the neighboring samples. Existing methods such as graph-based Gaussian random filed only considered the pair-wise similarity and then propagated the labels based on it. In this paper, we propose a new method from the perspective of local learning, which formulate the prediction of labels from the neighbors into a learning problem. Our contributions lie in two-fold: (1) we propose a new transductive video annotation method based on local kernel classifier; (2) local learnable is proposed to measure whether a sample can be learned from the neighbors well and we employ this measure into the optimization objective. Experiments on TRECVID 2005 dataset prove that the proposed method is effective and the local learning perspective is promising for video annotation.
Xinmei Tian 0001, Linjun Yang, Jingdong Wang 0001, Xiuqing Wu, Xian-Sheng Hua 0001
ICME1
2008 Optimized video scene segmentation
abstract
In this paper, we propose an optimized video scene segmentation approach with considering both content coherence and temporally contextual dissimilarity. First, a chain structure is constructed by connecting temporally adjacent shots to represent a video. Then the chain is partitioned such that the content within a chain segment is coherent enough and the contextual similarity of temporally adjacent chain segments is small enough. This task is formulated as a ratio function of content coherence and contextual similarity. Finally, we present an effective and efficient hierarchical chain partitioning approach to find the optimal scene segmentation. Experimental results on a set of home videos and feature movies demonstrate the superiority of the proposed approach over several existing key approaches.
Jingdong Wang 0001, Xinmei Tian 0001, Linjun Yang, Zhengjun Zha, Xian-Sheng Hua 0001
ICME2
2008 Bayesian video search reranking
abstract
Content-based video search reranking can be regarded as a process that uses visual content to recover the "true" ranking list from the noisy one generated based on textual information. This paper explicitly formulates this problem in the Bayesian framework, i.e., maximizing the ranking score consistency among visually similar video shots while minimizing the ranking distance, which represents the disagreement between the objective ranking list and the initial text-based. Different from existing point-wise ranking distance measures, which compute the distance in terms of the individual scores, two new methods are proposed in this paper to measure the ranking distance based on the disagreement in terms of pair-wise orders. Specifically, hinge distance penalizes the pairs with reversed order according to the degree of the reverse, while preference strength distance further considers the preference degree. By incorporating the proposed distances into the optimization objective, two reranking methods are developed which are solved using quadratic programming and matrix computation respectively. Evaluation on TRECVID video search benchmark shows that the performance improvement up to 21% on TRECVID 2006 and 61.11% on TRECVID 2007 are achieved relative to text search baseline.
Xinmei Tian 0001, Linjun Yang, Jingdong Wang 0001, Yichen Yang 0001, Xiuqing Wu, Xian-Sheng Hua 0001
ACM Multimedia1