EDBT 2026 Demo / reviewers in the wild / expert
Xinmei Tian 0001
dblp:03/5204-1
· DBLP profile ↗
139ranked-venue papers
18as first author
58since 2021 · last 2026
0000-0002-5952-8753ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 81 · 2 first-author · 49 since 2021Graphics, computer vision, multimedia, augmented reality and games · 76 · 14 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 1 since 2021Computer networks · 6 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging the Language Gap: Uncovering and Aligning Shared Circuits for Multi-Hop Reasoning in Multilingual LLMsabstractLarge language models (LLMs) present a paradox: they can correctly answer a multi-hop factual query in a high-resource language like English, yet fail on the identical query in another language. This raises a fundamental question about the nature of multilingual knowledge: are facts missing, or merely inaccessible? The underlying mechanisms for this knowledge gap have remained largely unexplored. In this work, we resolve this question by introducing a mechanistic interpretability framework that traces the causal pathways of multi-hop knowledge reasoning. Our analysis reveals a core, non-obvious finding: cross-lingual inconsistencies do not stem from a knowledge deficit. Instead, factual knowledge is robustly stored in a set of **shared, language-agnostic semantic neurons**. The failure originates from **misaligned attention pathways**, where a common set of critical attention heads fails to correctly route information along the reasoning chain to the appropriate knowledge neurons in lower-resource languages. This mechanistic diagnosis motivates a targeted alignment strategy: a surgical fine-tuning of only these critical heads. Experiments demonstrate that our method achieves significant improvements in multilingual multi-hop factuality—with positive cross-lingual transfer—while uniquely preserving general model capabilities, offering a scalable and mechanistically-grounded approach to building more reliable multilingual models. Zhen Huang 0007, Yonggang Zhang 0003, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye |
AAAI | 4 |
| 2026 | Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic CompositionalityabstractContrastively trained vision-language models like CLIP, have made remarkable progress in learning joint image-text representations, but still face challenges in compositional understanding.They often exhibit a "bag-of-words" behavior-struggling to capture the object relations, attribute-object bindings, and word order dependencies.This limitation arises not only from the reliance on global, single-vector representations for optimization, but also from the insufficient exploitation and modeling of the rich compositional information inherently present in paired image text data.In this work, we propose MACCO (MAsked Compositional Concept MOdeling), a framework that masks compositional concepts in one modality and reconstructs them conditioned on the full contextual information from the other, enabling the model to capture and align cross-modal compositional structures more effectively.To facilitate this process, we introduce two auxiliary objectives that jointly align and regularize masked features both inter-modally and intramodally.Extensive experiments on five compositional benchmarks, along with in-depth analyses, demonstrate that our approach not only significantly enhances compositionality in VLMs but also improves their ability to capture syntactic structure and linguistic information.Additionally, the improved compositionality also benefits text-to-image generation and multimodal large language model. Wei Li 0317, Zhen Huang 0007, Xinmei Tian 0001 |
ACL (1) | 3 |
| 2025 | A Similarity Paradigm Through Textual Regularization Without ForgettingabstractPrompt learning has emerged as a promising method for adapting pre-trained visual-language models (VLMs) to a range of downstream tasks. While optimizing the context can be effective for improving performance on specific tasks, it can often lead to poor generalization performance on unseen classes or datasets sampled from different distributions. It may be attributed to the fact that textual prompts tend to overfit downstream data distributions, leading to the forgetting of generalized knowledge derived from hand-crafted prompts. In this paper, we propose a novel method called Similarity Paradigm with Textual Regularization (SPTR) for prompt learning without forgetting. SPTR is a two-pronged design based on hand-crafted prompts that is an inseparable framework. 1) To avoid forgetting general textual knowledge, we introduce the optimal transport as a textual regularization to finely ensure approximation with hand-crafted features and tuning textual features. 2) In order to continuously unleash the general ability of multiple hand-crafted prompts, we propose a similarity paradigm for natural alignment score and adversarial alignment score to improve model robustness for generalization. Both modules share a common objective in addressing generalization issues, aiming to maximize the generalization capability derived from multiple hand-crafted prompts. Four representative tasks (i.e., non-generalization few-shot learning, base-to-novel generalization, cross-dataset generalization, domain generalization) across 11 datasets demonstrate that SPTR outperforms existing prompt learning methods. Fangming Cui, Jan Fong, Rongfei Zeng, Xinmei Tian 0001, Jun Yu 0002 |
AAAI | 4 |
| 2025 | Visual Evidence Prompting Mitigates Hallucinations in Large Vision-Language ModelsabstractWei Li, Zhen Huang, Houqiang Li, Le Lu, Yang Lu, Xinmei Tian, Xu Shen, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Wei Li 0317, Zhen Huang 0007, Houqiang Li, Le Lu 0001, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye |
ACL (1) | 6 |
| 2025 | Interpret and Improve In-Context Learning via the Lens of Input-Label MappingsabstractChenghao Sun, Zhen Huang, Yonggang Zhang, Le Lu, Houqiang Li, Xinmei Tian, Xu Shen, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhen Huang 0007, Yonggang Zhang 0003, Le Lu 0001, Houqiang Li, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye |
ACL (1) | 6 |
| 2025 | Leveraging Submodule Linearity Enhances Task Arithmetic Performance in LLMsabstractTask arithmetic is a straightforward yet highly effective strategy for model merging, enabling the resultant model to exhibit multi-task capabilities. Recent research indicates that models demonstrating linearity enhance the performance of task arithmetic. In contrast to existing methods that rely on the global linearization of the model, we argue that this linearity already exists within the model's submodules. In particular, we present a statistical analysis and show that submodules (e.g., layers, self-attentions, and MLPs) exhibit significantly higher linearity than the overall model. Based on these findings, we propose an innovative model merging strategy that independently merges these submodules. Especially, we derive a closed-form solution for optimal merging weights grounded in the linear properties of these submodules. Experimental results demonstrate that our method consistently outperforms the standard task arithmetic approach and other established baselines across different model scales and various tasks. This result highlights the benefits of leveraging the linearity of submodules and provides a new perspective for exploring solutions for effective and practical multi-task model merging. Rui Dai 0005, Sile Hu, Xu Shen 0001, Yonggang Zhang 0003, Xinmei Tian 0001, Jieping Ye |
ICLR | 5 |
| 2025 | A Theoretical Perspective: How to Prevent Model Collapse in Self-consuming Training LoopsabstractHigh-quality data is essential for training large generative models, yet the vast reservoir of real data available online has become nearly depleted. Consequently, models increasingly generate their own data for further training, forming Self-consuming Training Loops (STLs). However, the empirical results have been strikingly inconsistent: some models degrade or even collapse, while others successfully avoid these failures, leaving a significant gap in theoretical understanding to explain this discrepancy. This paper introduces the intriguing notion of *recursive stability* and presents the first theoretical generalization analysis, revealing how both model architecture and the proportion between real and synthetic data influence the success of STLs. We further extend this analysis to transformers in in-context learning, showing that even a constant-sized proportion of real data ensures convergence, while also providing insights into optimal synthetic data sizing. Shi Fu, Xinmei Tian 0001, Dacheng Tao |
ICLR | 4 |
| 2025 | Enhancing Target-unspecific Tasks through a Features MatrixabstractRecent developments in prompt learning of large Vision-Language Models (VLMs) have significantly improved performance in target-specific tasks. However, these prompting methods often struggle to tackle the target-unspecific or generalizable tasks effectively. It may be attributed to the fact that overfitting training causes the model to forget its general knowledge. The general knowledge has a strong promotion on target-unspecific tasks. To alleviate this issue, we propose a novel Features Matrix (FM) approach designed to enhance these models on target-unspecific tasks. Our method extracts and leverages general knowledge, shaping a Features Matrix (FM). Specifically, the FM captures the semantics of diverse inputs from a deep and fine perspective, preserving essential general knowledge, which mitigates the risk of overfitting. Representative evaluations demonstrate that: 1) the FM is compatible with existing frameworks as a generic and flexible module, and 2) the FM significantly showcases its effectiveness in enhancing target-unspecific tasks (base-to-novel generalization, domain generalization, and cross-dataset generalization), achieving state-of-the-art performance. Fangming Cui, Yonggang Zhang 0003, Xinmei Tian 0001, Jun Yu 0002 |
ICML | 4 |
| 2025 | Towards Generalizable Detector for Generated ImageabstractThe effective detection of generated images is crucial to mitigate potential risks associated with their misuse. Despite significant progress, a fundamental challenge remains: ensuring the generalizability of detectors. To address this, we propose a novel perspective on understanding and improving generated image detection, inspired by the human cognitive process: Humans identify an image as unnatural based on specific patterns because these patterns lie outside the space spanned by those of natural images. This is intrinsically related to out-of-distribution (OOD) detection, which identifies samples whose semantic patterns (i.e., labels) lie outside the semantic pattern space of in-distribution (ID) samples.
By treating patterns of generated images as OOD samples, we demonstrate that models trained merely over natural images bring guaranteed generalization ability under mild assumptions.
This transforms the generalization challenge of generated image detection into the problem of fitting natural image patterns.
Based on this insight, we propose a generalizable detection method through the lens of ID energy. Theoretical results capture the generalization risk of the proposed method. Experimental results across multiple benchmarks demonstrate the effectiveness of our approach. Qianshu Cai, Chao Wu 0001, Yonggang Zhang 0003, Jun Yu 0002, Xinmei Tian 0001 |
NeurIPS | 5 |
| 2025 | An Effective Levelling Paradigm for Unlabeled ScenariosabstractAdvancements in direct-integration fine-tuning frameworks have underscored their potential to enhance the performance of labeled scenarios and tasks. To enhance the generalization of different categories in the same dataset, some methods have added visual loss to these frameworks for unlabeled scenarios. However, the performance of these methods through visual loss does not improve significantly in domain generalization and cross-dataset generalization tasks. This may be attributed to the uncoordinated learning of the two-modalities alignment and visual loss. To mitigate this issue of uncoordinated learning, we propose a novel method called Levelling Paradigm (LePa) to improve performance for unlabeled tasks or scenarios. The proposed LePa, designed as a plug-in module, dynamically constrains and coordinates multiple objective functions, thereby improving the generalization of these baseline methods. Comprehensive experiments have shown that our design can effectively address generalized scenarios and tasks. Fangming Cui, Yuqiang Ren, Liang Xiao 0007, Xinmei Tian 0001 |
NeurIPS | 6 |
| 2025 | Epistemic Uncertainty for Generated Image DetectionabstractWe introduce a novel framework for AI-generated image detection through epistemic uncertainty, aiming to address critical security concerns in the era of generative models. Our key insight stems from the observation that distributional discrepancies between training and testing data manifest distinctively in the epistemic uncertainty space of machine learning models.
In this context, the distribution shift between natural and generated images leads to elevated epistemic uncertainty in models trained on natural images when evaluating generated ones. Hence, we exploit this phenomenon by using epistemic uncertainty as a proxy for detecting generated images. This converts the challenge of generated image detection into the problem of uncertainty estimation, underscoring the generalization performance of the model used for uncertainty estimation. Fortunately, advanced large-scale vision models pre-trained on extensive natural images have shown excellent generalization performance for various scenarios. Thus, we utilize these pre-trained models to estimate the epistemic uncertainty of images and flag those with high uncertainty as generated.
Extensive experiments demonstrate the efficacy of our method. Jun Nie, Yonggang Zhang 0003, Tongliang Liu, Yiu-Ming Cheung, Bo Han 0003, Xinmei Tian 0001 |
NeurIPS | 6 |
| 2025 | Detecting Generated Images by Fitting Natural Image DistributionsabstractThe increasing realism of generated images has raised significant concerns about their potential misuse, necessitating robust detection methods. Current approaches mainly rely on training binary classifiers, which depend heavily on the quantity and quality of available generated images. In this work, we propose a novel framework that exploits geometric differences between the data manifolds of natural and generated images. To exploit this difference, we employ a pair of functions engineered to yield consistent outputs for natural images but divergent outputs for generated ones, leveraging the property that their gradients reside in mutually orthogonal subspaces. This design enables a simple yet effective detection method: an image is identified as generated if a transformation along its data manifold induces a significant change in the loss value of a self-supervised model pre-trained on natural images. Further more, to address diminishing manifold disparities in advanced generative models, we leverage normalizing flows to amplify detectable differences by extruding generated images away from the natural image manifold. Extensive experiments demonstrate the efficacy of this method. Yonggang Zhang 0003, Jun Nie, Xinmei Tian 0001, Mingming Gong, Kun Zhang 0001, Bo Han 0003 |
NeurIPS | 3 |
| 2025 | Out-of-Distribution Detection with Virtual Outlier SmoothingabstractAbstract Detecting out-of-distribution (OOD) inputs plays a crucial role in guaranteeing the reliability of deep neural networks (DNNs) when deployed in real-world scenarios. However, DNNs typically exhibit overconfidence in OOD samples, which is attributed to the similarity in patterns between OOD and in-distribution (ID) samples. To mitigate this overconfidence, advanced approaches suggest the incorporation of auxiliary OOD samples during model training, where the outliers are assigned with an equal likelihood of belonging to any category. However, identifying outliers that share patterns with ID samples poses a significant challenge. To address the challenge, we propose a novel method, V irtual O utlier S m o othing (VOSo), which constructs auxiliary outliers using ID samples, thereby eliminating the need to search for OOD samples. Specifically, VOSo creates these virtual outliers by perturbing the semantic regions of ID samples and infusing patterns from other ID samples. For instance, a virtual outlier might consist of a cat’s face with a dog’s nose, where the cat’s face serves as the semantic feature for model prediction. Meanwhile, VOSo adjusts the labels of virtual OOD samples based on the extent of semantic region perturbation, aligning with the notion that virtual outliers may contain ID patterns. Extensive experiments are conducted on diverse OOD detection benchmarks, demonstrating the effectiveness of the proposed VOSo. Our code will be available at https://github.com/junz-debug/VOSo . Jun Nie, Yadan Luo, Shanshan Ye, Yonggang Zhang 0003, Xinmei Tian 0001, Zhen Fang 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | Consistent prompt learning for vision-language models
Yonggang Zhang 0003, Xinmei Tian 0001 |
Knowl. Based Syst. | 2 |
| 2025 | Boosting Fair Classifier Generalization through Adaptive Priority ReweighingabstractWith the increasing penetration of machine learning applications in critical decision-making areas, calls for algorithmic fairness are more prominent. Although there have been various modalities to improve algorithmic fairness through learning with fairness constraints, their performance does not generalize well in the test set. A performance-promising fair algorithm with better generalizability is needed. This article proposes a novel adaptive reweighing method to eliminate the impact of the distribution shifts between training and test data on model generalizability. Most previous reweighing methods propose to assign a unified weight for each (sub)group. Rather, our method granularly models the distance from the sample predictions to the decision boundary. Our adaptive reweighing method prioritizes samples closer to the decision boundary and assigns a higher weight to improve the generalizability of fair classifiers. Extensive experiments are performed to validate the generalizability of our adaptive priority reweighing method for accuracy and fairness measures (i.e., equal opportunity, equalized odds, and demographic parity) in tabular benchmarks. We also highlight the performance of our method in improving the fairness of language and vision models. The code is available at https://github.com/che2198/APW . Mengnan Du, Jindong Gu, Xinmei Tian 0001, Fengxiang He |
ACM Trans. Knowl. Discov. Data | 5 |
| 2024 | Sheared Backpropagation for Fine-Tuning Foundation ModelsabstractFine-tuning is the process of extending the training of pre-trained models on specific target tasks, thereby significantly enhancing their performance across various applications. However, fine-tuning often demands large memory consumption, posing a challenge for low-memory devices that some previous memory-efficient fine-tuning methods attempted to mitigate by pruning activations for gradient computation, albeit at the cost of significant computational overhead from the pruning processes during training. To address these challenges, we introduce PreBackRazor; a novel activation pruning scheme offering both computational and memory efficiency through a sparsified back-propagation strategy, which judiciously avoids unnecessary activation pruning and storage and gradient computation. Before activation pruning, our approach samples a probability of selecting a portion of parameters to freeze, utilizing a bandit method for updates to prioritize impactful gradients on convergence. During the feed-forward pass, each model layer adjusts adaptively based on parameter activation status, obviating the need for sparsification and storage of redundant activations for subsequent backpropagation. Benchmarking on fine-tuning foundation models, our approach maintains baseline accuracy across diverse tasks, yielding over 20% speedup and around 10% memory reduction. Moreover, integrating with an advanced CUDA kernel achieves up to 60% speedup without extra memory costs or accuracy loss, significantly enhancing the efficiency of fine-tuning foundation models on memory-constrained devices. Zhiyuan Yu 0004, Li Shen 0008, Liang Ding 0006, Xinmei Tian 0001, Yixin Chen 0001, Dacheng Tao |
CVPR | 4 |
| 2024 | Enhanced Motion-Text Alignment for Image-to-Video Transfer LearningabstractExtending large image-text pre-trained models (e.g., CLIP) for video understanding has made significant advancements. To enable the capability of CLIP to perceive dynamic information in videos, existing works are dedicated to equipping the visual encoder with various temporal modules. However, these methods exhibit “asymmetry” between the visual and textual sides, with neither temporal descriptions in input texts nor temporal modules in text encoder. This limitation hinders the potential of language supervision emphasized in CLIP, and restricts the learning of temporal features, as the text encoder has demonstrated limited proficiency in motion understanding. To address this issue, we propose leveraging “MoTion-Enhanced Descriptions” (MoTED) to facilitate the extraction of distinctive temporal features in videos. Specifically, we first generate discriminative motion-related descriptions via querying GPT-4 to compare easy-confusing action categories. Then, we incorporate both the visual and textual encoders with additional perception modules to process the video frames and generated descriptions, respectively. Finally, we adopt a contrastive loss to align the visual and textual motion features. Extensive experiments on five benchmarks show that MoTED surpasses state-of-the-art methods with convincing gaps, laying a solid foundation for empowering CLIP with strong temporal modeling. Chaoqun Wan, Tongliang Liu, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye |
CVPR | 4 |
| 2024 | Interpretable Composition Attribution Enhancement for Visio-linguistic Compositional UnderstandingabstractContrastively trained vision-language models such as CLIP have achieved remarkable progress in vision and language representation learning.Despite the promising progress, their proficiency in compositional reasoning over attributes and relations (e.g., distinguishing between "the car is underneath the person" and "the person is underneath the car") remains notably inadequate.We investigate the cause for this deficient behavior is the composition attribution issue, where the attribution scores (e.g., attention scores or GradCAM scores) for relations (e.g., underneath) or attributes (e.g., red) in the text are substantially lower than those for object terms.In this work, we show such issue is mitigated via a novel framework called CAE (Composition Attribution Enhancement).This generic framework incorporates various interpretable attribution methods to encourage the model to pay greater attention to composition words denoting relationships and attributes within the text.Detailed analysis shows that our approach enables the models to adjust and rectify the attribution of the texts.Extensive experiments across seven benchmarks reveal that our framework significantly enhances the ability to discern intricate details and construct more sophisticated interpretations of combined visual and linguistic elements. Wei Li 0317, Zhen Huang 0007, Xinmei Tian 0001, Le Lu 0001, Houqiang Li, Xu Shen 0001, Jieping Ye |
EMNLP | 3 |
| 2024 | Convergence of Bayesian Bilevel OptimizationabstractThis paper presents the first theoretical guarantee for Bayesian bilevel optimization (BBO) that we term for the prevalent bilevel framework combining Bayesian optimization at the outer level to tune hyperparameters, and the inner-level stochastic gradient descent (SGD) for training the model. We prove sublinear regret bounds suggesting simultaneous convergence of the inner-level model parameters and outer-level hyperparameters to optimal configurations for generalization capability. A pivotal, technical novelty in the proofs is modeling the excess risk of the SGD-trained parameters as evaluation noise during Bayesian optimization. Our theory implies the inner unit horizon, defined as the number of SGD iterations, shapes the convergence behavior of BBO. This suggests practical guidance on configuring the inner unit horizon to enhance training efficiency and model performance. Shi Fu, Fengxiang He, Xinmei Tian 0001, Dacheng Tao |
ICLR | 3 |
| 2024 | Out-of-Distribution Detection with Negative PromptsabstractOut-of-distribution (OOD) detection is indispensable for open-world machine learning models. Inspired by recent success in large pre-trained language-vision models, e.g., CLIP, advanced works have achieved impressive OOD detection results by matching the *similarity* between image features and features of learned prompts, i.e., positive prompts. However, existing works typically struggle with OOD samples having similar features with those of known classes. One straightforward approach is to introduce negative prompts to achieve a *dissimilarity* matching, which further assesses the anomaly level of image features by introducing the absence of specific features. Unfortunately, our experimental observations show that either employing a prompt like "not a photo of a" or learning a prompt to represent "not containing" fails to capture the dissimilarity for identifying OOD samples. The failure may be contributed to the diversity of negative features, i.e., tons of features could indicate features not belonging to a known class. To this end, we propose to learn a set of negative prompts for each class. The learned positive prompt (for all classes) and negative prompts (for each class) are leveraged to measure the similarity and dissimilarity in the feature space simultaneously, enabling more accurate detection of OOD samples. Extensive experiments are conducted on diverse OOD detection benchmarks, showing the effectiveness of our proposed method. Jun Nie, Yonggang Zhang 0003, Zhen Fang 0001, Tongliang Liu, Bo Han 0003, Xinmei Tian 0001 |
ICLR | 6 |
| 2024 | FedImpro: Measuring and Improving Client Update in Federated LearningabstractFederated Learning (FL) models often experience client drift caused by heterogeneous data, where the distribution of data differs across clients. To address this issue, advanced research primarily focuses on manipulating the existing gradients to achieve more consistent client models. In this paper, we present an alternative perspective on client drift and aim to mitigate it by generating improved local models. First, we analyze the generalization contribution of local training and conclude that this generalization contribution is bounded by the conditional Wasserstein distance between the data distribution of different clients. Then, we propose FedImpro, to construct similar conditional distributions for local training. Specifically, FedImpro decouples the model into high-level and low-level components, and trains the high-level portion on reconstructed feature distributions. This approach enhances the generalization contribution and reduces the dissimilarity of gradients in FL. Experimental results show that FedImpro can help FL defend against data heterogeneity and enhance the generalization performance of the model. Zhenheng Tang, Yonggang Zhang 0003, Shaohuai Shi, Xinmei Tian 0001, Tongliang Liu, Bo Han 0003, Xiaowen Chu 0001 |
ICLR | 4 |
| 2024 | Robust Training of Federated Models with Extremely Label DeficiencyabstractFederated semi-supervised learning (FSSL) has emerged as a powerful paradigm for collaboratively training machine learning models using distributed data with label deficiency. Advanced FSSL methods predominantly focus on training a single model on each client. However, this approach could lead to a discrepancy between the objective functions of labeled and unlabeled data, resulting in gradient conflicts. To alleviate gradient conflict, we propose a novel twin-model paradigm, called **Twinsight**, designed to enhance mutual guidance by providing insights from different perspectives of labeled and unlabeled data. In particular, Twinsight concurrently trains a supervised model with a supervised objective function while training an unsupervised model using an unsupervised objective function. To enhance the synergy between these two models, Twinsight introduces a neighborhood-preserving constraint, which encourages the preservation of the neighborhood relationship among data features extracted by both models. Our comprehensive experiments on four benchmark datasets provide substantial evidence that Twinsight can significantly outperform state-of-the-art methods across various experimental settings, demonstrating the efficacy of the proposed Twinsight. Yonggang Zhang 0003, Zhiqin Yang, Xinmei Tian 0001, Nannan Wang 0001, Tongliang Liu, Bo Han 0003 |
ICLR | 3 |
| 2024 | From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint TuningabstractLarge Language Models (LLMs) tend to prioritize adherence to user prompts over providing veracious responses, leading to the sycophancy issue. When challenged by users, LLMs tend to admit mistakes and provide inaccurate responses even if they initially provided the correct answer. Recent works propose to employ supervised fine-tuning (SFT) to mitigate the sycophancy issue, while it typically leads to the degeneration of LLMs' general capability. To address the challenge, we propose a novel supervised pinpoint tuning (SPT), where the region-of-interest modules are tuned for a given objective. Specifically, SPT first reveals and verifies a small percentage (<5%) of the basic modules, which significantly affect a particular behavior of LLMs. i.e., sycophancy. Subsequently, SPT merely fine-tunes these identified modules while freezing the rest. To verify the effectiveness of the proposed SPT, we conduct comprehensive experiments, demonstrating that SPT significantly mitigates the sycophancy issue of LLMs (even better than SFT). Moreover, SPT introduces limited or even no side effects on the general capability of LLMs. Our results shed light on how to precisely, effectively, and efficiently explain and improve the targeted ability of LLMs. Wei Chen 0005, Zhen Huang 0007, Liang Xie 0003, Binbin Lin 0001, Houqiang Li, Le Lu 0001, Xinmei Tian 0001, Deng Cai 0001, Yonggang Zhang 0003, Wenxiao Wang 0001, Xu Shen 0001, Jieping Ye |
ICML | 7 |
| 2024 | Towards Theoretical Understandings of Self-Consuming Generative ModelsabstractThis paper tackles the emerging challenge of training generative models within a self-consuming loop, wherein successive generations of models are recursively trained on mixtures of real and synthetic data from previous generations. We construct a theoretical framework to rigorously evaluate how this training procedure impacts the data distributions learned by future models, including parametric and non-parametric models. Specifically, we derive bounds on the total variation (TV) distance between the synthetic data distributions produced by future models and the original real data distribution under various mixed training scenarios for diffusion models with a one-hidden-layer neural network score function. Our analysis demonstrates that this distance can be effectively controlled under the condition that mixed training dataset sizes or proportions of real data are large enough. Interestingly, we further unveil a phase transition induced by expanding synthetic data amounts, proving theoretically that while the TV distance exhibits an initial ascent, it declines beyond a threshold point. Finally, we present results for kernel density estimation, delivering nuanced insights such as the impact of mixed data training on error propagation. Shi Fu, Sen Zhang 0006, Yingjie Wang 0007, Xinmei Tian 0001, Dacheng Tao |
ICML | 4 |
| 2024 | Interpreting and Improving Large Language Models in Arithmetic CalculationabstractLarge language models (LLMs) have demonstrated remarkable potential across numerous applications and have shown an emergent ability to tackle complex reasoning tasks, such as mathematical computations. However, even for the simplest arithmetic calculations, the intrinsic mechanisms behind LLMs remains mysterious, making it challenging to ensure reliability. In this work, we delve into uncovering a specific mechanism by which LLMs execute calculations. Through comprehensive experiments, we find that LLMs frequently involve a small fraction ($<$5%) of attention heads, which play a pivotal role in focusing on operands and operators during calculation processes. Subsequently, the information from these operands is processed through multi-layer perceptrons (MLPs), progressively leading to the final solution. These pivotal heads/MLPs, though identified on a specific dataset, exhibit transferability across different datasets and even distinct tasks. This insight prompted us to investigate the potential benefits of selectively fine-tuning these essential heads/MLPs to boost the LLMs’ computational performance. We empirically find that such precise tuning can yield notable enhancements on mathematical prowess, without compromising the performance on non-mathematical tasks. Our work serves as a preliminary exploration into the arithmetic calculation abilities inherent in LLMs, laying a solid foundation to reveal more intricate mathematical tasks. Chaoqun Wan, Yonggang Zhang 0003, Yiu-Ming Cheung, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye |
ICML | 5 |
| 2024 | Advancing Prompt Learning through an External LayerabstractPrompt learning represents a promising method for adapting pre-trained vision-language models (VLMs) to various downstream tasks by learning a set of text embeddings. One challenge inherent to these methods is the poor generalization performance due to the invalidity of the learned text embeddings for unseen tasks. A straightforward approach to bridge this gap is to freeze the text embeddings in prompts, which results in a lack of capacity to adapt VLMs for downstream tasks. To address this dilemma, we propose a paradigm called EnPrompt with a novel External Layer (EnLa). Specifically, we propose a textual external layer and learnable visual embeddings for adapting VLMs to downstream tasks. The learnable external layer is built upon valid embeddings of pre-trained CLIP. This design considers the balance of learning capabilities between the two branches. To align the textual and visual features, we propose a novel two-pronged approach: i) we introduce the optimal transport as the discrepancy metric to align the vision and text modalities, and ii) we introduce a novel strengthening feature to enhance the interaction between these two modalities. Four representative experiments (i.e., base-to-novel generalization, few-shot learning, cross-dataset generalization, domain shifts generalization) across 15 datasets demonstrate that our method outperforms the existing prompt learning method. Fangming Cui, Xun Yang 0001, Chao Wu 0001, Liang Xiao 0007, Xinmei Tian 0001 |
ACM Multimedia | 5 |
| 2024 | Adaptive Time-Stepping Schedules for Diffusion ModelsabstractThis paper studies how to tune the stepping schedule in diffusion models, which is mostly fixed in current practice, lacking theoretical foundations and assurance of optimal performance at the chosen discretization points. In this paper, we advocate the use of adaptive time-stepping schedules and design two algorithms with an optimized sampling error bound $EB$: (1) for continuous diffusion, we treat $EB$ as the loss function to discretization points and run gradient descent to adjust them; and (2) for discrete diffusion, we propose a greedy algorithm that adjusts only one discretization point to its best position in each iteration. We conducted extensive experiments that show (1) improved generation ability in well-trained models, and (2) premature though usable generation ability in under-trained models. The code is submitted and will be released publicly. Fengxiang He, Shi Fu, Xinmei Tian 0001, Dacheng Tao |
UAI | 4 |
| 2024 | Expert-level diagnosis of pediatric posterior fossa tumors via consistency calibration
Yonggang Zhang 0003, Xinmei Tian 0001 |
Knowl. Based Syst. | 4 |
| 2024 | FedGAMMA: Federated Learning With Global Sharpness-Aware MinimizationabstractFederated learning (FL) is a promising framework for privacy-preserving and distributed training with decentralized clients. However, there exists a large divergence between the collected local updates and the expected global update, which is known as the client drift and mainly caused by heterogeneous data distribution among clients, multiple local training steps, and partial client participation training. Most existing works tackle this challenge based on the empirical risk minimization (ERM) rule, while less attention has been paid to the relationship between the global loss landscape and the generalization ability. In this work, we propose FedGAMMA, a novel FL algorithm with Global sharpness-Aware MiniMizAtion to seek a global flat landscape with high performance. Specifically, in contrast to FedSAM which only seeks the local flatness and still suffers from performance degradation when facing the client-drift issue, we adopt a local varieties control technique to better align each client's local updates to alleviate the client drift and make each client heading toward the global flatness together. Finally, extensive experiments demonstrate that FedGAMMA can substantially outperform several existing FL baselines on various datasets, and it can well address the client-drift issue and simultaneously seek a smoother and flatter global landscape. Rong Dai, Xun Yang 0001, Li Shen 0008, Xinmei Tian 0001, Meng Wang 0001, Yongdong Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Sharper Bounds for Uniformly Stable Algorithms with Stationary Mixing Process
Shi Fu, Yunwen Lei, Qiong Cao, Xinmei Tian 0001, Dacheng Tao |
ICLR | 4 |
| 2023 | Moderately Distributional Exploration for Domain GeneralizationabstractDomain generalization (DG) aims to tackle the distribution shift between training domains and unknown target domains. Generating new domains is one of the most effective approaches, yet its performance gain depends on the distribution discrepancy between the generated and target domains. Distributionally robust optimization is promising to tackle distribution discrepancy by exploring domains in an uncertainty set. However, the uncertainty set may be overwhelmingly large, leading to low-confidence prediction in DG. It is because a large uncertainty set could introduce domains containing semantically different factors from training domains. To address this issue, we propose to perform a $\textit{mo}$derately $\textit{d}$istributional $\textit{e}$xploration (MODE) for domain generalization. Specifically, MODE performs distribution exploration in an uncertainty $\textit{subset}$ that shares the same semantic factors with the training domains. We show that MODE can endow models with provable generalization performance on unknown target domains. The experimental results show that MODE achieves competitive performance compared to state-of-the-art baselines. Rui Dai 0005, Yonggang Zhang 0003, Zhen Fang 0001, Bo Han 0003, Xinmei Tian 0001 |
ICML | 5 |
| 2023 | Structured Cooperative Learning with Graphical Model PriorsabstractWe study how to train personalized models for different tasks on decentralized devices with limited local data. We propose "Structured Cooperative Learning (SCooL)", in which a cooperation graph across devices is generated by a graphical model prior to automatically coordinate mutual learning between devices. By choosing graphical models enforcing different structures, we can derive a rich class of existing and novel decentralized learning algorithms via variational inference. In particular, we show three instantiations of SCooL that adopt Dirac distribution, stochastic block model (SBM), and attention as the prior generating cooperation graphs. These EM-type algorithms alternate between updating the cooperation graph and cooperative learning of local models. They can automatically capture the cross-task correlations among devices by only monitoring their model updating in order to optimize the cooperation graph. We evaluate SCooL and compare it with existing decentralized learning methods on an extensive set of benchmarks, on which SCooL always achieves the highest accuracy of personalized models and significantly outperforms other baselines on communication efficiency. Our code is available at https://github.com/ShuangtongLi/SCooL. Shuangtong Li, Tianyi Zhou 0001, Xinmei Tian 0001, Dacheng Tao |
ICML | 3 |
| 2023 | Adaptive Priority Reweighing for Generalizing Fairness ImprovementabstractWith the increasing penetration of Machine-Learning (ML) applications in critical decision-making areas, calls for algorithmic fairness are more prominent. Though there have been diverse modalities to improve algorithmic fairness through training the algorithms with fairness constraints, their performance does not generalize well at the test set. A performance-promising fair algorithm with better generalizability is needed. This paper proposes a novel adaptive reweighing method to eliminate the impact of the distribution shifts between training and test data on model generalizability. Specifically, instead of assigning a unified weight for each (sub)group as most previous reweighing methods propose, we granularly model the distance from the sample predictions to the decision boundary and assign higher individual weight to the samples closer to the decision boundary in each (sub)group. Our adaptive reweighing method prioritizes the samples closer to the decision boundary and assigns a higher weight to improve the generalizability of fair classifiers. We design extensive experiments to evaluate the generalizability of our adaptive priority reweighing method for accuracy and fairness measures (i.e., equal opportunity, equalized odds, and demographic parity.) in tabular benchmarks across Adult, COMPAS, and IPUMS. We further highlight the performance of our method in improving the fairness of language and vision models. We believe that our method shows promising results in improving the fairness of any pre-trained models simply via fine-tuning. Xinmei Tian 0001 |
IJCNN | 3 |
| 2023 | Semantic-Aware Mixup for Domain GeneralizationabstractDeep neural networks (DNNs) have shown exciting performance in various tasks, yet suffer generalization failures when meeting unknown target domains. One of the most promising approaches to achieve domain generalization (DG) is generating unseen data, e.g., mixup, to cover the unknown target data. However, existing works overlook the challenges induced by the simultaneous appearance of changes in both the semantic and distribution space. Accordingly, such a challenge makes source distributions hard to fit for DNNs. To mitigate the hard-fitting issue, we propose to perform a semantic-aware mixup (SAM) for domain generalization, where whether to perform mixup depends on the semantic and domain information. The feasibility of SAM shares the same spirits with the Fourier-based mixup. Namely, the Fourier phase spectrum is expected to contain semantics information (relating to labels), while the Fourier amplitude retains other information (relating to style information). Built upon the insight, SAM applies different mixup strategies to the Fourier phase spectrum and amplitude information. For instance, SAM merely performs mixup on the amplitude spectrum when both the semantic and domain information changes. Consequently, the overwhelmingly large change can be avoided. We validate the effectiveness of SAM using image classification tasks on several DG benchmarks. Chengchao Xu, Xinmei Tian 0001 |
IJCNN | 2 |
| 2023 | 3D Creation at Your Fingertips: From Text or Image to 3D AssetsabstractWe demonstrate an automatic 3D creation system, which can create realistic 3D assets solely from a text or image prompt without requiring any specialized 3D modeling skills. Users can either describe the object they envision in natural language or upload a reference image that records what they have seen with the phone. Our system will generate a high-quality 3D mesh that faithfully matches the users' input. We propose a coarse-to-fine framework to achieve this goal. Specifically, we first obtain a low-resolution mesh instantly by utilizing a pre-trained text/image conditional 3D generative model. Using such coarse mesh as the initialization, we further optimize a high-resolution textured 3D mesh with fine-grained appearance guidance from large-scale 2D diffusion models. Our system can create visually-pleasing results in minutes, which is significantly faster than existing methods. Meanwhile, the system ensures that the resulting 3D assets are precisely aligned with the input text or image prompt. With these advanced capabilities, our demonstration provides a streamlined and intuitive platform for users to incorporate 3D creation into their daily lives. Yang Chen 0048, Jingwen Chen 0001, Yingwei Pan, Xinmei Tian 0001, Tao Mei 0001 |
ACM Multimedia | 4 |
| 2023 | FedFed: Feature Distillation against Data Heterogeneity in Federated LearningabstractFederated learning (FL) typically faces data heterogeneity, i.e., distribution shifting among clients.
Sharing clients' information has shown great potentiality in mitigating data heterogeneity, yet incurs a dilemma in preserving privacy and promoting model performance. To alleviate the dilemma, we raise a fundamental question: Is it possible to share partial features in the data to tackle data heterogeneity?
In this work, we give an affirmative answer to this question by proposing a novel approach called **Fed**erated **Fe**ature **d**istillation (FedFed).
Specifically, FedFed partitions data into performance-sensitive features (i.e., greatly contributing to model performance) and performance-robust features (i.e., limitedly contributing to model performance).
The performance-sensitive features are globally shared to mitigate data heterogeneity, while the performance-robust features are kept locally.
FedFed enables clients to train models over local and shared data. Comprehensive experiments demonstrate the efficacy of FedFed in promoting model performance. Zhiqin Yang, Yonggang Zhang 0003, Yu Zheng 0021, Xinmei Tian 0001, Tongliang Liu, Bo Han 0003 |
NeurIPS | 4 |
| 2023 | Bi-calibration Networks for Weakly-Supervised Video Representation Learning
Fuchen Long, Ting Yao 0003, Zhaofan Qiu, Xinmei Tian 0001, Jiebo Luo 0001, Tao Mei 0001 |
Int. J. Comput. Vis. | 4 |
| 2023 | A heuristic multi-objective task scheduling framework for container-based clouds via actor-critic reinforcement learning
Lilu Zhu, Feng Wu 0001, Yanfeng Hu, Xinmei Tian 0001 |
Neural Comput. Appl. | 5 |
| 2023 | Domain Generalization Via Encoding and Resampling in a Unified Latent SpaceabstractDomain generalization aims to generalize a network trained on multiple domains to unknown yet related domains. Operating under the assumption that invariant information generalizes well to unknown domains, previous work has aimed to minimize the discrepancies amongst distributions across given domains. However, without prior regularization of feature distributions, the network in practice overfits the invariant information in the given domains. Moreover, if there are insufficient samples in given domains, then domain generalizability is limited, as diverse domain variations are not captured. To address these two drawbacks, we propose to explicitly map features in known and unknown domains onto latent space in a fixed Gaussian mixture distribution by variational coding. As a result, features in different classes follow Gaussian distributions with different mean values. The predefined latent space narrows discrepancies between known and unknown domains and effectively separates samples into different classes. Moreover, we propose to perturb sample features with gradients from the distribution regularized loss. This perturbation generates samples beyond but near the latent space of prior distributions, which has a profound impact on domain variations. Experiments and visualizations demonstrate the effectiveness of our proposed method. Zhiwei Xiong, Xinmei Tian 0001, Zhengjun Zha |
IEEE Trans. Multim. | 4 |
| 2023 | Domain-Class Correlation Decomposition for Generalizable Person Re-IdentificationabstractDomain generalization in person re-identification is a highly important meaningful and practical task in which a model trained with data from several source domains is expected to generalize well to unseen target domains. Domain adversarial learning is a promising domain generalization method that aims to remove domain information in the latent representation through adversarial training. However, in person re-identification, the domain and class are correlated, and we theoretically show that domain adversarial learning will lose certain information about class due to this domain-class correlation. Inspired by causal inference, we propose to perform interventions to the domain factor$d$, aiming to decompose the domain-class correlation. To achieve this goal, we proposed estimating the resulting representation$z^{*}$caused by the intervention through first- and second-order statistical characteristic matching. Specifically, we build a memory bank to restore the statistical characteristics of each domain. Then, we use the newly generated samples$\lbrace z^{*},y,d^{*}\rbrace$to compute the loss function. These samples are domain-class correlation decomposed; thus, we can learn a domain-invariant representation that can capture more class-related features. Extensive experiments show that our model outperforms the state-of-the-art methods on the large-scale domain generalization Re-ID benchmark. Xinmei Tian 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Category-Stitch Learning for Union Domain GeneralizationabstractDomain generalization aims at generalizing the network trained on multiple domains to unknown but related domains. Under the assumption that different domains share the same classes, previous works can build relationships across domains. However, in realistic scenarios, the change of domains is always followed by the change of categories, which raises a difficulty for collecting sufficient aligned categories across domains. Bearing this in mind, this article introduces union domain generalization (UDG) as a new domain generalization scenario, in which the label space varies across domains, and the categories in unknown domains belong to the union of all given domain categories. The absence of categories in given domains is the main obstacle to aligning different domain distributions and obtaining domain-invariant information. To address this problem, we propose category-stitch learning (CSL), which aims at jointly learning the domain-invariant information and completing missing categories in all domains through an improved variational autoencoder and generators. The domain-invariant information extraction and sample generation cross-promote each other to better generalizability. Additionally, we decouple category and domain information and propose explicitly regularizing the semantic information by the classification loss with transferred samples. Thus our method can breakthrough the category limit and generate samples of missing categories in each domain. Extensive experiments and visualizations are conducted on MNIST, VLCS, PACS, Office-Home, and DomainNet datasets to demonstrate the effectiveness of our proposed method. Zhiwei Xiong, Yuning Lu, Xinmei Tian 0001, Zhengjun Zha |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | Learning to Collaborate in Decentralized Learning of Personalized ModelsabstractLearning personalized models for user-customized computer-vision tasks is challenging due to the limited private-data and computation available on each edge device. Decentralized learning (DL) can exploit the images distributed over devices on a network topology to train a global model but is not designed to train personalized models for different tasks or optimize the topology. Moreover, the mixing weights used to aggregate neighbors' gradient messages in DL can be suboptimal for personalization since they are not adaptive to different nodes/tasks and learning stages. In this paper, we dynamically update the mixing-weights to improve the personalized model for each node's task and meanwhile learn a sparse topology to reduce communication costs. Our first approach, “learning to collaborate (L2C) ”, directly optimizes the mixing weights to minimize the local validation loss per node for a predefined set of nodes/tasks. In order to produce mixing weights for new nodes or tasks, we further develop “meta-L2C‘, which learns an attention mechanism to automatically assign mixing weights by comparing two nodes' model updates. We evaluate both methods on diverse benchmarks and experimental settings for image classification. Thorough comparisons to both classical and recent methods for IID/non-IID decentralized and federated learning demonstrate our method's advantages in identifying collaborators among nodes, learning sparse topology, and producing better personalized models with low communication and computational cost. Shuangtong Li, Tianyi Zhou 0001, Xinmei Tian 0001, Dacheng Tao |
CVPR | 3 |
| 2022 | Prompt Distribution LearningabstractWe present prompt distribution learning for effectively adapting a pre-trained vision-language model to address downstream recognition tasks. Our method not only learns low-bias prompts from a few samples but also captures the distribution of diverse prompts to handle the varying visual representations. In this way, we provide high-quality task-related content for facilitating recognition. This prompt distribution learning is realized by an efficient approach that learns the output embeddings of prompts instead of the input embeddings. Thus, we can employ a Gaussian distribution to model them effectively and derive a surrogate loss for efficient training. Extensive experiments on 12 datasets demonstrate that our method consistently and significantly outperforms existing methods. For example, with 1 sample per category, it relatively improves the average result by 9.1% compared to human-crafted prompts. Yuning Lu, Jianzhuang Liu, Yonggang Zhang 0003, Xinmei Tian 0001 |
CVPR | 5 |
| 2022 | Meta Convolutional Neural Networks for Single Domain GeneralizationabstractIn single domain generalization, models trained with data from only one domain are required to perform well on many unseen domains. In this paper, we propose a new model, termed meta convolutional neural network, to solve the single domain generalization problem in image recognition. The key idea is to decompose the convolutional features of images into meta features. Acting as “visual words”, meta features are defined as universal and basic visual elements for image representations (like words for documents in language). Taking meta features as reference, we propose compositional operations to eliminate irrelevant features of local convolutional features by an addressing process and then to reformulate the convolutional feature maps as a composition of related meta features. In this way, images are universally coded without biased information from the unseen domain, which can be processed by following modules trained in the source domain. The compositional operations adopt a regression analysis technique to learn the meta features in an online batch learning manner. Extensive experiments on multiple benchmark datasets verify the superiority of the proposed model in improving single domain generalization ability. Chaoqun Wan, Xu Shen 0001, Yonggang Zhang 0003, Zhiheng Yin, Xinmei Tian 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
CVPR | 5 |
| 2022 | Self-Supervision Can Be a Good Few-Shot Learner
Yuning Lu, Liangjian Wen, Jianzhuang Liu, Xinmei Tian 0001 |
ECCV (19) | 5 |
| 2022 | Adversarial Robustness Through the Lens of Causality
Yonggang Zhang 0003, Mingming Gong, Tongliang Liu, Gang Niu 0001, Xinmei Tian 0001, Bo Han 0003, Bernhard Schölkopf, Kun Zhang 0001 |
ICLR | 5 |
| 2022 | DisPFL: Towards Communication-Efficient Personalized Federated Learning via Decentralized Sparse TrainingabstractPersonalized federated learning is proposed to handle the data heterogeneity problem amongst clients by learning dedicated tailored local models for each user. However, existing works are often built in a centralized way, leading to high communication pressure and high vulnerability when a failure or an attack on the central server occurs. In this work, we propose a novel personalized federated learning framework in a decentralized (peer-to-peer) communication protocol named DisPFL, which employs personalized sparse masks to customize sparse local models on the edge. To further save the communication and computation cost, we propose a decentralized sparse training technique, which means that each local model in DisPFL only maintains a fixed number of active parameters throughout the whole local training and peer-to-peer communication process. Comprehensive experiments demonstrate that DisPFL significantly saves the communication bottleneck for the busiest node among all clients and, at the same time, achieves higher model accuracy with less computation cost and communication rounds. Furthermore, we demonstrate that our method can easily adapt to heterogeneous local clients with varying computation complexities and achieves better personalized performances. Rong Dai, Li Shen 0008, Fengxiang He, Xinmei Tian 0001, Dacheng Tao |
ICML | 4 |
| 2022 | Identity-Disentangled Adversarial Augmentation for Self-supervised LearningabstractData augmentation is critical to contrastive self-supervised learning, whose goal is to distinguish a sample’s augmentations (positives) from other samples (negatives). However, strong augmentations may change the sample-identity of the positives, while weak augmentation produces easy positives/negatives leading to nearly-zero loss and ineffective learning. In this paper, we study a simple adversarial augmentation method that can modify training data to be hard positives/negatives without distorting the key information about their original identities. In particular, we decompose a sample $x$ to be its variational auto-encoder (VAE) reconstruction $G(x)$ plus the residual $R(x)=x-G(x)$, where $R(x)$ retains most identity-distinctive information due to an information-theoretic interpretation of the VAE objective. We then adversarially perturb $G(x)$ in the VAE’s bottleneck space and adds it back to the original $R(x)$ as an augmentation, which is therefore sufficiently challenging for contrastive learning and meanwhile preserves the sample identity intact. We apply this “identity-disentangled adversarial augmentation (IDAA)” to different self-supervised learning methods. On multiple benchmark datasets, IDAA consistently improves both their efficiency and generalization performance. We further show that IDAA learned on a dataset can be transferred to other datasets. Code is available at \href{https://github.com/kai-wen-yang/IDAA}{https://github.com/kai-wen-yang/IDAA}. Tianyi Zhou 0001, Xinmei Tian 0001, Dacheng Tao |
ICML | 3 |
| 2022 | Towards Lightweight Black-Box Attack Against Deep Neural NetworksabstractBlack-box attacks can generate adversarial examples without accessing the parameters of target model, largely exacerbating the threats of deployed deep neural networks (DNNs). However, previous works state that black-box attacks fail to mislead target models when their training data and outputs are inaccessible. In this work, we argue that black-box attacks can pose practical attacks in this extremely restrictive scenario where only several test samples are available. Specifically, we find that attacking the shallow layers of DNNs trained on a few test samples can generate powerful adversarial examples. As only a few samples are required, we refer to these attacks as lightweight black-box attacks. The main challenge to promoting lightweight attacks is to mitigate the adverse impact caused by the approximation error of shallow layers. As it is hard to mitigate the approximation error with few available samples, we propose Error TransFormer (ETF) for lightweight attacks. Namely, ETF transforms the approximation error in the parameter space into a perturbation in the feature space and alleviates the error by disturbing features. In experiments, lightweight black-box attacks with the proposed ETF achieve surprising results. For example, even if only 1 sample per category available, the attack success rate in lightweight black-box attacks is only about 3% lower than that of the black-box attacks with complete training data. Yonggang Zhang 0003, Chaoqun Wan, Tongliang Liu, Bo Han 0003, Xinmei Tian 0001 |
NeurIPS | 8 |
| 2022 | Adversarial Auto-Augment with Label Preservation: A Representation Learning Principle Guided ApproachabstractData augmentation is a critical contributing factor to the success of deep learning but heavily relies on prior domain knowledge which is not always available. Recent works on automatic data augmentation learn a policy to form a sequence of augmentation operations, which are still pre-defined and restricted to limited options. In this paper, we show that a prior-free autonomous data augmentation's objective can be derived from a representation learning principle that aims to preserve the minimum sufficient information of the labels. Given an example, the objective aims at creating a distant ``hard positive example'' as the augmentation, while still preserving the original label. We then propose a practical surrogate to the objective that can be optimized efficiently and integrated seamlessly into existing methods for a broad class of machine learning tasks, e.g., supervised, semi-supervised, and noisy-label learning. Unlike previous works, our method does not require training an extra generative model but instead leverages the intermediate layer representations of the end-task model for generating data augmentations. In experiments, we show that our method consistently brings non-trivial improvements to the three aforementioned learning tasks from both efficiency and final performance, either or not combined with pre-defined augmentations, e.g., on medical images when domain knowledge is unavailable and the existing augmentation techniques perform poorly. Code will be released publicly. Yanchao Sun, Jiahao Su, Fengxiang He, Xinmei Tian 0001, Furong Huang, Tianyi Zhou 0001, Dacheng Tao |
NeurIPS | 5 |
| 2022 | DigestPath: A benchmark dataset with challenge review for the pathological detection and segmentation of digestive-system
Qian Da, Zhongyu Li 0002, Yanfei Zuo, Chenbin Zhang, Jingxin Liu 0005, Wen Chen 0001, Jiahui Li 0005, Dou Xu, Hongmei Yi, Zhe Wang 0043, Li Zhang 0040, Xianying He, Xiaofan Zhang 0002, Ke Mei, Chuang Zhu, Weizeng Lu, LinLin Shen, Jun Shi 0006, Jun Li 0106, Sreehari S, Ganapathy Krishnamurthi, Jiangcheng Yang, Tiancheng Lin 0001, Qingyu Song 0004, Xuechen Liu 0004, Simon Graham, Raja Muhammad Saad Bashir, Canqian Yang, Shaofei Qin, Xinmei Tian 0001, Jie Zhao 0014, Dimitris N. Metaxas, Hongsheng Li 0001, Chaofu Wang, Shaoting Zhang 0001 |
Medical Image Anal. | 34 |
| 2022 | CRAR: Accelerating Stereo Matching with Cascaded Residual Regression and Adaptive RefinementabstractDense stereo matching estimates the depth for each pixel of the referenced images. Recently, deep learning algorithms have dramatically promoted the development of stereo matching. The state-of-the-art result is achieved by models adopting deep convolutional neural networks. However, a considerable computational burden is also introduced, which slows the inference. To solve this problem, previous works down-sampled the input images to decrease the spatial size. However, down-sampling increases the error rate and its lower bound. In this article, we accelerate stereo matching algorithms through the improvement of network structure. Inspired by network compression, we conduct decomposition and sparsification to squeeze the computationally expensive cost optimization network. It is sparsified and then decomposed into smaller networks, which are designed and trained in a cascaded manner to reach the nearest possible performance of the larger network. Previous methods have utilized numerous refinement methods to adjust the coarse disparity. We integrate refinement methods to create an unified algorithm to utilize parallelism for running devices to further accelerate the inference. The extensive experiments on Kitti2015, Kitti2012, and Middlebury datasets demonstrate the efficiency of our method. Linghua Zeng, Xinmei Tian 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Revisiting Knowledge Distillation: An Inheritance and Exploration FrameworkabstractKnowledge Distillation (KD) is a popular technique to transfer knowledge from a teacher model or ensemble to a student model. Its success is generally attributed to the privileged information on similarities/consistency between the class distributions or intermediate feature representations of the teacher model and the student model. However, directly pushing the student model to mimic the probabilities/features of the teacher model to a large extent limits the student model in learning undiscovered knowledge/features. In this paper, we propose a novel inheritance and exploration knowledge distillation framework (IE-KD), in which a student model is split into two parts - inheritance and exploration. The inheritance part is learned with a similarity loss to transfer the existing learned knowledge from the teacher model to the student model, while the exploration part is encouraged to learn representations different from the inherited ones with a dis-similarity loss. Our IE-KD framework is generic and can be easily combined with existing distillation or mutual learning methods for training deep neural networks. Extensive experiments demonstrate that these two parts can jointly push the student model to learn more diversified and effective representations, and our IE-KD can be a general technique to improve the student network to achieve SOTA performance. Furthermore, by applying our IE-KD to the training of two networks, the performance of both can be improved w.r.t. deep mutual learning. Zhen Huang 0007, Xu Shen 0001, Jun Xing, Tongliang Liu, Xinmei Tian 0001, Houqiang Li, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
CVPR | 5 |
| 2021 | A Style and Semantic Memory Mechanism for Domain Generalization*abstractMainstream state-of-the-art domain generalization algorithms tend to prioritize the assumption on semantic in-variance across domains. Meanwhile, the inherent intra-domain style invariance is usually underappreciated and put on the shelf. In this paper, we reveal that leveraging intra-domain style invariance is also of pivotal importance in improving the efficiency of domain generalization. We verify that it is critical for the network to be informative on what domain features are invariant and shared among in-stances, so that the network sharpens its understanding and improves its semantic discriminative ability. Correspondingly, we also propose a novel “jury” mechanism, which is particularly effective in learning useful semantic feature commonalities among domains. Our complete model called STEAM can be interpreted as a novel probabilistic graphical model, for which the implementation requires convenient constructions of two kinds of memory banks: semantic feature bank and style feature bank. Empirical results show that our proposed framework surpasses the state-of-the-art methods by clear margins. Yang Chen 0048, Yu Wang 0102, Yingwei Pan, Ting Yao 0003, Xinmei Tian 0001, Tao Mei 0001 |
ICCV | 5 |
| 2021 | 3D Local Convolutional Neural Networks for Gait RecognitionabstractThe goal of gait recognition is to learn the unique spatiotemporal pattern about the human body shape from its temporal changing characteristics. As different body parts behave differently during walking, it is intuitive to model the spatio-temporal patterns of each part separately. However, existing part-based methods equally divide the feature maps of each frame into fixed horizontal stripes to get local parts. It is obvious that these stripe partition-based methods cannot accurately locate the body parts. First, different body parts can appear at the same stripe (e.g., arms and the torso), and one part can appear at different stripes in different frames (e.g., hands). Second, different body parts possess different scales, and even the same part in different frames can appear at different locations and scales. Third, different parts also exhibit distinct movement patterns (e.g., at which frame the movement starts, the position change frequency, how long it lasts). To overcome these issues, we propose novel 3D local operations as a generic family of building blocks for 3D gait recognition backbones. The proposed 3D local operations support the extraction of local 3D volumes of body parts in a sequence with adaptive spatial and temporal scales, locations and lengths. In this way, the spatio-temporal patterns of the body parts are well learned from the 3D local neighborhood in partspecific scales, locations, frequencies and lengths. Experiments demonstrate that our 3D local convolutional neural networks achieve state-of-the-art performance on popular gait datasets. Code is available at: https://github.com/yellowtownhz/3DLocalCNN. Zhen Huang 0007, Dixiu Xue, Xu Shen 0001, Xinmei Tian 0001, Houqiang Li, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
ICCV | 4 |
| 2021 | Transferrable Contrastive Learning for Visual Domain AdaptationabstractSelf-supervised learning (SSL) has recently become the favorite among feature learning methodologies. It is therefore appealing for domain adaptation approaches to consider incorporating SSL. The intuition is to enforce instance-level feature consistency such that the predictor becomes somehow invariant across domains. However, most existing SSL methods in the regime of domain adaptation usually are treated as standalone auxiliary components, leaving the signatures of domain adaptation unattended. Actually, the optimal region where the domain gap vanishes and the instance level constraint that SSL peruses may not coincide at all. From this point, we present a particular paradigm of self-supervised learning tailored for domain adaptation, i.e., Transferrable Contrastive Learning (TCL), which links the SSL and the desired cross-domain transferability congruently. We find contrastive learning intrinsically a suitable candidate for domain adaptation, as its instance invariance assumption can be conveniently promoted to cross-domain class-level invariance favored by domain adaptation tasks. Based on particular memory bank constructions and pseudo label strategies, TCL then penalizes cross-domain intra-class domain discrepancy between source and target through a clean and novel contrastive loss. The free lunch is, thanks to the incorporation of contrastive learning, TCL relies on a moving-averaged key encoder that naturally achieves a temporally ensembled version of pseudo labels for target data, which avoids pseudo label error propagation at no extra cost. TCL therefore efficiently reduces cross-domain gaps. Through extensive experiments on benchmarks (Office-Home, VisDA-2017, Digits-five, PACS and DomainNet) for both single-source and multi-source domain adaptation tasks, TCL has demonstrated state-of-the-art performances. Yang Chen 0048, Yingwei Pan, Yu Wang 0102, Ting Yao 0003, Xinmei Tian 0001, Tao Mei 0001 |
ACM Multimedia | 5 |
| 2021 | Class-Disentanglement and Applications in Adversarial Detection and DefenseabstractWhat is the minimum necessary information required by a neural net $D(\cdot)$ from an image $x$ to accurately predict its class? Extracting such information in the input space from $x$ can allocate the areas $D(\cdot)$ mainly attending to and shed novel insights to the detection and defense of adversarial attacks. In this paper, we propose ''class-disentanglement'' that trains a variational autoencoder $G(\cdot)$ to extract this class-dependent information as $x - G(x)$ via a trade-off between reconstructing $x$ by $G(x)$ and classifying $x$ by $D(x-G(x))$, where the former competes with the latter in decomposing $x$ so the latter retains only necessary information for classification in $x-G(x)$. We apply it to both clean images and their adversarial images and discover that the perturbations generated by adversarial attacks mainly lie in the class-dependent part $x-G(x)$. The decomposition results also provide novel interpretations to classification and attack models. Inspired by these observations, we propose to conduct adversarial detection and adversarial defense respectively on $x - G(x)$ and $G(x)$, which consistently outperform the results on the original $x$. In experiments, this simple approach substantially improves the detection and defense against different types of adversarial attacks. Tianyi Zhou 0001, Yonggang Zhang 0003, Xinmei Tian 0001, Dacheng Tao |
NeurIPS | 4 |
| 2021 | Learning multi-granularity features from multi-granularity regions for person re-identification
Jiwei Yang, Xinmei Tian 0001 |
Neurocomputing | 3 |
| 2020 | Learning to Localize Actions from Moments
Fuchen Long, Ting Yao 0003, Zhaofan Qiu, Xinmei Tian 0001, Jiebo Luo 0001, Tao Mei 0001 |
ECCV (3) | 4 |
| 2020 | Dual-Path Distillation: A Unified Framework to Improve Black-Box AttacksabstractWe study the problem of constructing black-box adversarial attacks, where no model information is revealed except for the feedback knowledge of the given inputs. To obtain sufficient knowledge for crafting adversarial examples, previous methods query the target model with inputs that are perturbed with different searching directions. However, these methods suffer from poor query efficiency since the employed searching directions are sampled randomly. To mitigate this issue, we formulate the goal of mounting efficient attacks as an optimization problem in which the adversary tries to fool the target model with a limited number of queries. Under such settings, the adversary has to select appropriate searching directions to reduce the number of model queries. By solving the efficient-attack problem, we find that we need to distill the knowledge in both the path of the adversarial examples and the path of the searching directions. Therefore, we propose a novel framework, dual-path distillation, that utilizes the feedback knowledge not only to craft adversarial examples but also to alter the searching directions to achieve efficient attacks. Experimental results suggest that our framework can significantly increase the query efficiency. Yonggang Zhang 0003, Tongliang Liu, Xinmei Tian 0001 |
ICML | 4 |
| 2020 | Spatio-Temporal Inception Graph Convolutional Networks for Skeleton-Based Action RecognitionabstractSkeleton-based human action recognition has attracted much attention with the prevalence of accessible depth sensors. Recently, graph convolutional networks (GCNs) have been widely used for this task due to their powerful capability to model graph data. The topology of the adjacency graph is a key factor for modeling the correlations of the input skeletons. Thus, previous methods mainly focus on the design/learning of the graph topology. But once the topology is learned, only a single-scale feature and one transformation exist in each layer of the networks. Many insights, such as multi-scale information and multiple sets of transformations, that have been proven to be very effective in convolutional neural networks (CNNs), have not been investigated in GCNs. The reason is that, due to the gap between graph-structured skeleton data and conventional image/video data, it is very challenging to embed these insights into GCNs. To overcome this gap, we reinvent the split-transform-merge strategy in GCNs for skeleton sequence processing. Specifically, we design a simple and highly modularized graph convolutional network architecture for skeleton-based action recognition. Our network is constructed by repeating a building block that aggregates multi-granularity information from both the spatial and temporal paths. Extensive experiments demonstrate that our network outperforms state-of-the-art methods by a significant margin with only 1/5 of the parameters and 1/10 of the FLOPs. Zhen Huang 0007, Xu Shen 0001, Xinmei Tian 0001, Houqiang Li, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
ACM Multimedia | 3 |
| 2020 | Real-World Image Denoising with Deep BoostingabstractWe propose a Deep Boosting Framework (DBF) for real-world image denoising by integrating the deep learning technique into the boosting algorithm. The DBF replaces conventional handcrafted boosting units by elaborate convolutional neural networks, which brings notable advantages in terms of both performance and speed. We design a lightweight Dense Dilated Fusion Network (DDFN) as an embodiment of the boosting unit, which addresses the vanishing of gradients during training due to the cascading of networks while promoting the efficiency of limited parameters. The capabilities of the proposed method are first validated on several representative simulation tasks including non-blind and blind Gaussian denoising and JPEG image deblocking. We then focus on a practical scenario to tackle with the complex and challenging real-world noise. To facilitate leaning-based methods including ours, we build a new Real-world Image Denoising (RID) dataset, which contains 200 pairs of high-resolution images with diverse scene content under various shooting conditions. Moreover, we conduct comprehensive analysis on the domain shift issue for real-world denoising and propose an effective one-shot domain transfer scheme to address this issue. Comprehensive experiments on widely used benchmarks demonstrate that the proposed method significantly surpasses existing methods on the task of real-world image denoising. Code and dataset are available at https://github.com/ngchc/deepBoosting. Chang Chen 0004, Zhiwei Xiong, Xinmei Tian 0001, Zhengjun Zha, Feng Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Accelerating Convolutional Neural Networks by Removing Interspatial and Interkernel RedundanciesabstractRecently, the high computational resource demands of convolutional neural networks (CNNs) have hindered a wide range of their applications. To solve this problem, many previous works attempted to reduce the redundant calculations during the evaluation of CNNs. However, these works mainly focused on either interspatial or interkernel redundancy. In this paper, we further accelerate existing CNNs by removing both types of redundancies. First, we convert interspatial redundancy into interkernel redundancy by decomposing one convolutional layer to one block that we design. Then, we adopt rank-selection and pruning methods to remove the interkernel redundancy. The rank-selection method, which considerably reduces manpower, contributes to determining the number of kernels to be pruned in the pruning method. We apply a layer-wise training algorithm rather than the traditional end-to-end training to overcome the difficulty of convergence. Finally, we fine-tune the entire network to achieve better performance. Our method is applied on three widely used datasets of an image classification task. We achieve better results in terms of accuracy and compression rate compared with previous state-of-the-art methods. Linghua Zeng, Xinmei Tian 0001 |
IEEE Trans. Cybern. | 2 |
| 2020 | Principal Component Adversarial ExampleabstractDespite having achieved excellent performance on various tasks, deep neural networks have been shown to be susceptible to adversarial examples, i.e., visual inputs crafted with structural imperceptible noise. To explain this phenomenon, previous works implicate the weak capability of the classification models and the difficulty of the classification tasks. These explanations appear to account for some of the empirical observations but lack deep insight into the intrinsic nature of adversarial examples, such as the generation method and transferability. Furthermore, previous works generate adversarial examples completely rely on a specific classifier (model). Consequently, the attack ability of adversarial examples is strongly dependent on the specific classifier. More importantly, adversarial examples cannot be generated without a trained classifier. In this paper, we raise a question: what is the real cause of the generation of adversarial examples? To answer this question, we propose a new concept, called the adversarial region, which explains the existence of adversarial examples as perturbations perpendicular to the tangent plane of the data manifold. This view yields a clear explanation of the transfer property across different models of adversarial examples. Moreover, with the notion of the adversarial region, we propose a novel target-free method to generate adversarial examples via principal component analysis. We verify our adversarial region hypothesis on a synthetic dataset and demonstrate through extensive experiments on real datasets that the adversarial examples generated by our method have competitive or even strong transferability compared with model-dependent adversarial example generating methods. Moreover, our experiment shows that the proposed method is more robust to defensive methods than previous methods. Yonggang Zhang 0003, Xinmei Tian 0001, Xinchao Wang, Dacheng Tao |
IEEE Trans. Image Process. | 2 |
| 2020 | A Multi-Organ Nucleus Segmentation ChallengeabstractGeneralized nucleus segmentation techniques can contribute greatly to reducing the time to develop and validate visual biomarkers for new digital pathology datasets. We summarize the results of MoNuSeg 2018 Challenge whose objective was to develop generalizable nuclei segmentation techniques in digital pathology. The challenge was an official satellite event of the MICCAI 2018 conference in which 32 teams with more than 80 participants from geographically diverse institutes participated. Contestants were given a training set with 30 images from seven organs with annotations of 21,623 individual nuclei. A test dataset with 14 images taken from seven organs, including two organs that did not appear in the training set was released without annotations. Entries were evaluated based on average aggregated Jaccard index (AJI) on the test set to prioritize accurate instance segmentation as opposed to mere semantic segmentation. More than half the teams that completed the challenge outperformed a previous baseline. Among the trends observed that contributed to increased accuracy were the use of color normalization as well as heavy data augmentation. Additionally, fully convolutional networks inspired by variants of U-Net, FCN, and Mask-RCNN were popularly used, typically based on ResNet or VGG base architectures. Watershed segmentation on predicted semantic segmentation maps was a popular post-processing strategy. Several of the top techniques compared favorably to an individual human annotator and can be used with confidence for nuclear morphometrics. Neeraj Kumar 0002, Ruchika Verma, Deepak Anand, Yanning Zhou 0001, Omer Fahri Onder, Efstratios Tsougenis, Hao Chen 0011, Pheng-Ann Heng, Jiahui Li 0005, Navid Alemi Koohbanani, Mostafa Jahanifar, Neda Zamani Tajeddin, Ali Gooya, Nasir M. Rajpoot, Xuhua Ren, Sihang Zhou 0001, Qian Wang 0001, Dinggang Shen, Cheng-Kun Yang, Chi-Hung Weng, Wei-Hsiang Yu, Chao-Yuan Yeh, Shuoyu Xu, Pak-Hei Yeung, Amirreza Mahbod, Gerald Schaefer, Isabella Ellinger, Rupert Ecker, Örjan Smedby, Chunliang Wang, Benjamin Chidester, Vinh Ton-That, Minh-Triet Tran, Jian Ma 0004, Minh N. Do, Simon Graham, Quoc Dang Vu, Jin Tae Kwak, Akshaykumar Gunda, Raviteja Chunduri, Corey Hu, Dariush Lotfi, Reza Safdari, Antanas Kascenas, Alison O'Neil, Dennis Eschweiler, Johannes Stegmaier, Yanping Cui, Kailin Chen, Xinmei Tian 0001, Philipp Grüning, Erhardt Barth, Elad Arbel, Itay Remer, Amir Ben-Dor, Ekaterina Sirazitdinova, Matthias Kohl, Stefan Braunewell, Yuexiang Li, Xinpeng Xie, LinLin Shen, Jun Ma 0016, Krishanu Das Baksi, Mohammad Azam Khan, Jaegul Choo, Adrián Colomer, Valery Naranjo, Linmin Pei, Khan M. Iftekharuddin, Kaushiki Roy, Debotosh Bhattacharjee, Aníbal Pedraza, Gloria Bueno García, Sabarinathan Devanathan, Saravanan Radhakrishnan, Praveen Koduganty, Zihan Wu 0001, Guanyu Cai, Amit Sethi |
IEEE Trans. Medical Imaging | 56 |
| 2020 | Coarse-to-Fine Localization of Temporal Action ProposalsabstractLocalizing temporal action proposals from long videos is a fundamental challenge in video analysis (e.g., action detection and recognition or dense video captioning). Most existing approaches often overlook the hierarchical granularities of actions and thus fail to discriminate fine-grained action proposals (e.g., hand washing laundry or changing a tire in vehicle repair). In this paper, we propose a novel coarse-to-fine temporal proposal (CFTP) approach to localize temporal action proposals by exploring different action granularities. Our proposed CFTP consists of three stages: a coarse proposal network (CPN) to generate long action proposals, a temporal convolutional anchor network (CAN) to localize finer proposals, and a proposal reranking network (PRN) to further identify proposals from previous stages. Specifically, CPN explores three complementary actionness curves (namely pointwise, pairwise, and recurrent curves) that represent actions at different levels for generating coarse proposals, while CAN refines these proposals by a multiscale cascaded 1D-convolutional anchor network. In contrast to existing works, our coarse-to-fine approach can progressively localize fine-grained action proposals. We conduct extensive experiments on two action benchmarks (THUMOS14 and ActivityNet v1.3) and demonstrate the superior performance of our approach when compared to the state-of-the-art techniques on various video understanding tasks. Fuchen Long, Ting Yao 0003, Zhaofan Qiu, Xinmei Tian 0001, Tao Mei 0001, Jiebo Luo 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | Concentrated Local Part Discovery With Fine-Grained Part Representation for Person Re-IdentificationabstractThe attention mechanism for person re-identification has been widely studied with deep convolutional neural networks. This mechanism works as a good complement to the global features extracted from an image of the entire human body. However, existing works mainly focus on discovering local parts with simple feature representations, such as global average pooling. Moreover, these works either require extra supervision, such as labeling of body joints, or pay little attention to the guidance of part learning, resulting in scattered activation of learned parts. Furthermore, existing works usually extract local features from different body parts via global average pooling and then concatenate them together as good global features. We find that local features acquired in this way contribute little to the overall performance. In this paper, we argue the significance of local part description and explore the attention mechanism from both local part discovery and local part representation aspects. For local part discovery, we propose a new constrained attention module to make the activated regions concentrated and meaningful without extra supervision. For local part representation, we propose a statistical-positional-relational descriptor to represent local parts from a fine-grained viewpoint. Extensive experiments are conducted to validate the overall performance, the effectiveness of each component, and the generalization ability. We achieve a rank-1 accuracy of 95.1% on Market1501, 64.7% on CUHK03, 87.1% on DukeMTMC-ReID, and 79.9% on MSMT17, outperforming state-of-the-art methods. Chaoqun Wan, Xinmei Tian 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 3 |
| 2020 | Graph Edge Convolutional Neural Networks for Skeleton-Based Action RecognitionabstractBody joints, directly obtained from a pose estimation model, have proven effective for action recognition. Existing works focus on analyzing the dynamics of human joints. However, except joints, humans also explore motions of limbs for understanding actions. Given this observation, we investigate the dynamics of human limbs for skeleton-based action recognition. Specifically, we represent an edge in a graph of a human skeleton by integrating its spatial neighboring edges (for encoding the cooperation between different limbs) and its temporal neighboring edges (for achieving the consistency of movements in an action). Based on this new edge representation, we devise a graph edge convolutional neural network (CNN). Considering the complementarity between graph node convolution and edge convolution, we further construct two hybrid networks by introducing different shared intermediate layers to integrate graph node and edge CNNs. Our contributions are twofold, graph edge convolution and hybrid networks for integrating the proposed edge convolution and the conventional node convolution. Experimental results on the Kinetics and NTU-RGB+D data sets demonstrate that our graph edge convolution is effective at capturing the characteristics of actions and that our graph edge CNN significantly outperforms the existing state-of-the-art skeleton-based action recognition methods. Xikun Zhang 0002, Chang Xu 0002, Xinmei Tian 0001, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Video Retrieval with Similarity-Preserving Deep Temporal HashingabstractDespite the fact that remarkable progress has been made in recent years, Content-based Video Retrieval (CBVR) is still an appealing research topic due to increasing search demands in the Internet era of big data. This article aims to explore an efficient CBVR system by discriminately hashing videos into short binary codes. Existing video hashing methods usually encounter two weaknesses originating from the following sources: (1) Most works adopt the separated stages method or the frame-pooling based end-to-end architecture. However, the spatial-temporal properties of videos cannot be fully explored or kept well in the follow-up hashing step. (2) Discriminative learning based on pairwise or triplet constraints often suffers from slow convergence and poor local optimization, mainly because of the limited samples for each update. To alleviate these problems, we propose an end-to-end video retrieval framework called the Similarity-Preserving Deep Temporal Hashing (SPDTH) network. Specifically, we equip the model with the ability to capture spatial-temporal properties of videos and to generate binary codes by stacked Gated Recurrent Units (GRUs). It unifies video temporal modeling and learning to hash into one step to allow for maximum retention of information. We also introduce a deep metric learning objective called ℓ 2 All _ loss for network training by preserving intra-class similarity and inter-class separability, and a quantization loss between the real-valued outputs and the binary codes is minimized. Extensive experiments on several challenging datasets demonstrate that SPDTH can consistently outperform state-of-the-art methods. Richang Hong, Xinmei Tian 0001, Meng Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2019 | Exploring Object Relation in Mean Teacher for Cross-Domain DetectionabstractRendering synthetic data (e.g., 3D CAD-rendered images) to generate annotations for learning deep models in vision tasks has attracted increasing attention in recent years. However, simply applying the models learnt on synthetic images may lead to high generalization error on real images due to domain shift. To address this issue, recent progress in cross-domain recognition has featured the Mean Teacher, which directly simulates unsupervised domain adaptation as semi-supervised learning. The domain gap is thus naturally bridged with consistency regularization in a teacher-student scheme. In this work, we advance this Mean Teacher paradigm to be applicable for cross-domain detection. Specifically, we present Mean Teacher with Object Relations (MTOR) that novelly remolds Mean Teacher under the backbone of Faster R-CNN by integrating the object relations into the measure of consistency cost between teacher and student modules. Technically, MTOR firstly learns relational graphs that capture similarities between pairs of regions for teacher and student respectively. The whole architecture is then optimized with three consistency regularizations: 1) region-level consistency to align the region-level predictions between teacher and student, 2) inter-graph consistency for matching the graph structures between teacher and student, and 3) intra-graph consistency to enhance the similarity between regions of same class within the graph of student. Extensive experiments are conducted on the transfers across Cityscapes, Foggy Cityscapes, and SIM10k, and superior results are reported when comparing to state-of-the-art approaches. More remarkably, we obtain a new record of single model: 22.8% of mAP on Syn2Real detection dataset. Yingwei Pan, Chong-Wah Ngo, Xinmei Tian 0001, Ling-Yu Duan, Ting Yao 0003 |
CVPR | 4 |
| 2019 | Camera Lens Super-ResolutionabstractExisting methods for single image super-resolution (SR) are typically evaluated with synthetic degradation models such as bicubic or Gaussian downsampling. In this paper, we investigate SR from the perspective of camera lenses, named as CameraSR, which aims to alleviate the intrinsic tradeoff between resolution (R) and field-of-view (V) in realistic imaging systems. Specifically, we view the R-V degradation as a latent model in the SR process and learn to reverse it with realistic low- and high-resolution image pairs. To obtain the paired images, we propose two novel data acquisition strategies for two representative imaging systems (i.e., DSLR and smartphone cameras), respectively. Based on the obtained City100 dataset, we quantitatively analyze the performance of commonly-used synthetic degradation models, and demonstrate the superiority of CameraSR as a practical solution to boost the performance of existing SR methods. Moreover, CameraSR can be readily generalized to different content and devices, which serves as an advanced digital zoom tool in realistic imaging systems. Chang Chen 0004, Zhiwei Xiong, Xinmei Tian 0001, Zhengjun Zha, Feng Wu 0001 |
CVPR | 3 |
| 2019 | Compact Feature Learning for Multi-Domain Image ClassificationabstractThe goal of multi-domain learning is to improve the performance over multiple domains by making full use of all training data from them. However, variations of feature distributions across different domains result in a non-trivial solution of multi-domain learning. The state-of-the-art work regarding multi-domain classification aims to extract domain-invariant features and domain-specific features independently. However, they view the distributions of features from different classes as a general distribution and try to match these distributions across domains, which lead to the mixture of features from different classes across domains and degrade the performance of classification. Additionally, existing works only force the shared features among domains to be orthogonal to the features in the domain-specific network. However, redundant features between the domain-specific networks still remain, which may shrink the discriminative ability of domain-specific features. Therefore, we propose an end-to-end network to obtain the more optimal features, which we call compact features. We propose to extract the domain-invariant features by matching the joint distributions of different domains, which have dis- tinct boundaries between different classes. Moreover, we add an orthogonal constraint between the private features across domains to ensure the discriminative ability of the domain-specific space. The proposed method is validated on three landmark datasets, and the results demonstrate the effectiveness of our method. Xinmei Tian 0001, Zhiwei Xiong, Feng Wu 0001 |
CVPR | 2 |
| 2019 | Gaussian Temporal Awareness Networks for Action LocalizationabstractTemporally localizing actions in a video is a fundamental challenge in video understanding. Most existing approaches have often drawn inspiration from image object detection and extended the advances, e.g., SSD and Faster R-CNN, to produce temporal locations of an action in a 1D sequence. Nevertheless, the results can suffer from robustness problem due to the design of predetermined temporal scales, which overlooks the temporal structure of an action and limits the utility on detecting actions with complex variations. In this paper, we propose to address the problem by introducing Gaussian kernels to dynamically optimize temporal scale of each action proposal. Specifically, we present Gaussian Temporal Awareness Networks (GTAN) - a new architecture that novelly integrates the exploitation of temporal structure into an one-stage action localization framework. Technically, GTAN models the temporal structure through learning a set of Gaussian kernels, each for a cell in the feature maps. Each Gaussian kernel corresponds to a particular interval of an action proposal and a mixture of Gaussian kernels could further characterize action proposals with various length. Moreover, the values in each Gaussian curve reflect the contextual contributions to the localization of an action proposal. Extensive experiments are conducted on both THUMOS14 and ActivityNet v1.3 datasets, and superior results are reported when comparing to state-of-the-art approaches. More remarkably, GTAN achieves 1.9% and 1.1% improvements in mAP on testing set of the two datasets. Fuchen Long, Ting Yao 0003, Zhaofan Qiu, Xinmei Tian 0001, Jiebo Luo 0001, Tao Mei 0001 |
CVPR | 4 |
| 2019 | Learning Spatio-Temporal Representation With Local and Global DiffusionabstractConvolutional Neural Networks (CNN) have been regarded as a powerful class of models for visual recognition problems. Nevertheless, the convolutional filters in these networks are local operations while ignoring the large-range dependency. Such drawback becomes even worse particularly for video recognition, since video is an information-intensive media with complex temporal variations. In this paper, we present a novel framework to boost the spatio-temporal representation learning by Local and Global Diffusion (LGD). Specifically, we construct a novel neural network architecture that learns the local and global representations in parallel. The architecture is composed of LGD blocks, where each block updates local and global features by modeling the diffusions between these two representations. Diffusions effectively interact two aspects of information, i.e., localized and holistic, for more powerful way of representation learning. Furthermore, a kernelized classifier is introduced to combine the representations from two aspects for video recognition. Our LGD networks achieve clear improvements on the large-scale Kinetics-400 and Kinetics-600 video classification datasets against the best competitors by 3.5% and 0.7%. We further examine the generalization of both the global and local representations produced by our pre-trained LGD networks on four different benchmarks for video action recognition and spatio-temporal action detection tasks. Superior performances over several state-of-the-art techniques on these benchmarks are reported. Zhaofan Qiu, Ting Yao 0003, Chong-Wah Ngo, Xinmei Tian 0001, Tao Mei 0001 |
CVPR | 4 |
| 2019 | Quantization NetworksabstractAlthough deep neural networks are highly effective, their high computational and memory costs severely hinder their applications to portable devices. As a consequence, lowbit quantization, which converts a full-precision neural network into a low-bitwidth integer version, has been an active and promising research topic. Existing methods formulate the low-bit quantization of networks as an approximation or optimization problem. Approximation-based methods confront the gradient mismatch problem, while optimizationbased methods are only suitable for quantizing weights and can introduce high computational cost during the training stage. In this paper, we provide a simple and uniform way for weights and activations quantization by formulating it as a differentiable non-linear function. The quantization function is represented as a linear combination of several Sigmoid functions with learnable biases and scales that could be learned in a lossless and end-to-end manner via continuous relaxation of the steepness of Sigmoid functions. Extensive experiments on image classification and object detection tasks show that our quantization networks outperform state-of-the-art methods. We believe that the proposed method will shed new lights on the interpretation of neural network quantization. Jiwei Yang, Xu Shen 0001, Jun Xing, Xinmei Tian 0001, Houqiang Li, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
CVPR | 4 |
| 2019 | KCNN: Kernel-wise Quantization to Remarkably Decrease Multiplications in Convolutional Neural NetworkabstractConvolutional neural networks (CNNs) have demonstrated state-of-the-art performance in computer vision tasks. However, the high computational power demand of running devices of recent CNNs has hampered many of their applications. Recently, many methods have quantized the floating-point weights and activations to fixed-points or binary values to convert fractional arithmetic to integer or bit-wise arithmetic. However, since the distributions of values in CNNs are extremely complex, fixed-points or binary values lead to numerical information loss and cause performance degradation. On the other hand, convolution is composed of multiplications and accumulation, but the implementation of multiplications in hardware is more costly comparing with accumulation. We can preserve the rich information of floating-point values on dedicated low power devices by considerably decreasing the multiplications. In this paper, we quantize the floating-point weights in each kernel separately to multiple bit planes to remarkably decrease multiplications. We obtain a closed-form solution via an aggressive Lloyd algorithm and the fine-tuning is adopted to optimize the bit planes. Furthermore, we propose dual normalization to solve the pathological curvature problem during fine-tuning. Our quantized networks show negligible performance loss compared to their floating-point counterparts. Linghua Zeng, Zhangcheng Wang, Xinmei Tian 0001 |
IJCAI | 3 |
| 2019 | Mocycle-GAN: Unpaired Video-to-Video TranslationabstractUnsupervised image-to-image translation is the task of translating an image from one domain to another in the absence of any paired training examples and tends to be more applicable to practical applications. Nevertheless, the extension of such synthesis from image-to-image to video-to-video is not trivial especially when capturing spatio-temporal structures in videos. The difficulty originates from the aspect that not only the visual appearance in each frame but also motion between consecutive frames should be realistic and consistent across transformation. This motivates us to explore both appearance structure and temporal continuity in video synthesis. In this paper, we present a new Motion-guided Cycle GAN, dubbed as Mocycle-GAN, that novelly integrates motion estimation into unpaired video translator. Technically, Mocycle-GAN capitalizes on three types of constrains: adversarial constraint discriminating between synthetic and real frame, cycle consistency encouraging an inverse translation on both frame and motion, and motion translation validating the transfer of motion between consecutive frames. Extensive experiments are conducted on video-to-labels and labels-to-video translation, and superior results are reported when comparing to state-of-the-art methods. More remarkably, we qualitatively demonstrate our Mocycle-GAN for both flower-to-flower and ambient condition transfer. Yang Chen 0048, Yingwei Pan, Ting Yao 0003, Xinmei Tian 0001, Tao Mei 0001 |
ACM Multimedia | 4 |
| 2019 | Animating Your Life: Real-Time Video-to-Animation TranslationabstractWe demonstrate a video-to-animation translator, which can transform real-world video into cartoon or ink-wash animation in real-time. When users upload a video or record what they are seeing with the phone, the video-to-animation translator renders the live streaming video with cartoon or ink-wash animation style while maintaining the original contents. We formulate this task as video-to-video translation problem in the absence of any paired training examples, since the manual labeling of such paired video-animation data is cost-expensive and even unrealistic in practice. Technically, an unified unpaired video-to-video translator is utilized to explore both appearance structure and temporal continuity in video synthesis. As such, not only the visual appearance in each frame but also motion between consecutive frames are ensured to be realistic and consistent for video translation. Based on these technologies, our demonstration can be conducted on any videos in the wild and supports live video-to-animation translation, which engages users with the animated artistic expression of their life. Yang Chen 0048, Yingwei Pan, Ting Yao 0003, Xinmei Tian 0001, Tao Mei 0001 |
ACM Multimedia | 4 |
| 2019 | RC-CNN: Representation-Consistent Convolutional Neural Networks for Achieving Transformation InvarianceabstractConvolutional neural networks (CNNs) are powerful and have achieved state-of-the-art performance in many visual recognition tasks. Despite their impressive performance, CNNs are still unable to remain invariant while some spatial transformations are applied on images. Herein, we propose representation-consistent neural networks to solve this problem. By introducing consistent losses between the representations in different layers of transformed images, the recognition performance of transformed images is significantly improved. This model not only learns to map from the transformed images to the pre-defined labels but each layer also learns to generate invariant representations when the input images are transformed. All the characteristics of transformation invariance are embedded in the model, which means that no extra parameters or computations are introduced in the well-trained model. Comparative experiments demonstrate the superiority of our model when learning invariance to rotation, translation, and scaling on large-scale image recognition and retrieval tasks. Anfeng He, Xinmei Tian 0001 |
SMC | 3 |
| 2019 | Eigenfunction-Based Multitask Learning in a Reproducing Kernel Hilbert SpaceabstractMultitask learning aims to improve the performance on related tasks by exploring the interdependence among them. Existing multitask learning methods explore the relatedness among tasks on the basis of the input features and the model parameters. In this paper, we focus on nonparametric multitask learning and propose to measure task relatedness from a novel perspective in a reproducing kernel Hilbert space (RKHS). Past works have shown that the objective function for a given task can be approximated using the top eigenvalues and corresponding eigenfunctions of a predefined integral operator on an RKHS. In our method, we formulate our objective for multitask learning as a linear combination of two sets of eigenfunctions, common eigenfunctions shared by different tasks and unique eigenfunctions in individual tasks, such that the eigenfunctions for one task can provide additional information on another and help to improve its performance. We present both theoretical and empirical validations of our proposed approach. The theoretical analysis demonstrates that our learning algorithm is uniformly argument stable and that the convergence rate of the generalization upper bound can be improved by learning multiple tasks. Experiments on several benchmark multitask learning data sets show that our method yields promising results. Xinmei Tian 0001, Tongliang Liu, Xinchao Wang, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2019 | Eigenvector-Based Distance Metric Learning for Image Classification and RetrievalabstractDistance metric learning has been widely studied in multifarious research fields. The mainstream approaches learn a Mahalanobis metric or learn a linear transformation. Recent related works propose learning a linear combination of base vectors to approximate the metric. In this way, fewer variables need to be determined, which is efficient when facing high-dimensional data. Nevertheless, such works obtain base vectors using additional data from related domains or randomly generate base vectors. However, obtaining base vectors from related domains requires extra time and additional data, and random vectors introduce randomness into the learning process, which requires sufficient random vectors to ensure the stability of the algorithm. Moreover, the random vectors cannot capture the rich information of the training data, leading to a degradation in performance. Considering these drawbacks, we propose a novel distance metric learning approach by introducing base vectors explicitly learned from training data. Given a specific task, we can make a sparse approximation of its objective function using the top eigenvalues and corresponding eigenvectors of a predefined integral operator on the reproducing kernel Hilbert space. Because the process of generating eigenvectors simply refers to the training data of the considered task, our proposed method does not require additional data and can reflect the intrinsic information of the input features. Furthermore, the explicitly learned eigenvectors do not result in randomness, and we can extend our method to any kernel space without changing the objective function. We only need to learn the coefficients of these eigenvectors, and the only hyperparameter that we need to determine is the number of eigenvectors that we utilize. Additionally, an optimization algorithm is proposed to efficiently solve this problem. Extensive experiments conducted on several datasets demonstrate the effectiveness of our proposed method. Zhangcheng Wang, Richang Hong, Xinmei Tian 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2019 | Multi-source Multi-level Attention Networks for Visual Question AnsweringabstractIn recent years, Visual Question Answering (VQA) has attracted increasing attention due to its requirement on cross-modal understanding and reasoning of vision and language. VQA is proposed to automatically answer natural language questions with reference to a given image. VQA is challenging, because the reasoning process on a visual domain needs a full understanding of the spatial relationship, semantic concepts, as well as the common sense for a real image. However, most existing approaches jointly embed the abstract low-level visual features and high-level question features to infer answers. These works have limited reasoning ability due to the lack of modeling of the rich spatial context of regions, high-level semantics of images, and knowledge across multiple sources. To solve the challenges, we propose multi-source multi-level attention networks for visual question answering that can benefit both spatial inferences by visual attention on context-aware region representation and reasoning by semantic attention on concepts as well as external knowledge. Indeed, we learn to reason on image representation by question-guided attention at different levels across multiple sources, including region and concept level representation from image source as well as sentence level representation from the external knowledge base. First, we encode region-based middle-level outputs from Convolutional Neural Networks (CNNs) into spatially embedded representation by a multi-directional two-dimensional recurrent neural network and, further, locate the answer-related regions by Multiple Layer Perceptron as visual attention. Second, we generate semantic concepts from high-level semantics in CNNs and select those question-related concepts as concept attention. Third, we query semantic knowledge from the general knowledge base by concepts and selected question-related knowledge as knowledge attention. Finally, we jointly optimize visual attention, concept attention, knowledge attention, and question embedding by a softmax classifier to infer the final answer. Extensive experiments show the proposed approach achieved significant improvement on two very challenging VQA datasets. Dongfei Yu, Jianlong Fu, Xinmei Tian 0001, Tao Mei 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2018 | Domain Generalization via Conditional Invariant RepresentationsabstractDomain generalization aims to apply knowledge gained from multiple labeled source domains to unseen target domains. The main difficulty comes from the dataset bias: training data and test data have different distributions, and the training set contains heterogeneous samples from different distributions. Let X denote the features, and Y be the class labels. Existing domain generalization methods address the dataset bias problem by learning a domain-invariant representation h(X) that has the same marginal distribution P(h(X)) across multiple source domains. The functional relationship encoded in P(Y|X) is usually assumed to be stable across domains such that P(Y|h(X)) is also invariant. However, it is unclear whether this assumption holds in practical problems. In this paper, we consider the general situation where both P(X) and P(Y|X) can change across all domains. We propose to learn a feature representation which has domain-invariant class conditional distributions P(h(X)|Y). With the conditional invariant representation, the invariance of the joint distribution P(h(X),Y) can be guaranteed if the class prior P(Y) does not change across training and test domains. Extensive experiments on both synthetic and real data demonstrate the effectiveness of the proposed method. Mingming Gong, Xinmei Tian 0001, Tongliang Liu, Dacheng Tao |
AAAI | 3 |
| 2018 | Sequence-to-Sequence Learning via Shared Latent RepresentationabstractSequence-to-sequence learning is a popular research area in deep learning, such as video captioning and speech recognition. Existing methods model this learning as a mapping process by first encoding the input sequence to a fixed-sized vector, followed by decoding the target sequence from the vector. Although simple and intuitive, such mapping model is task-specific, unable to be directly used for different tasks. In this paper, we propose a star-like framework for general and flexible sequence-to-sequence learning, where different types of media contents (the peripheral nodes) could be encoded to and decoded from a shared latent representation (SLR) (the central node). This is inspired by the fact that human brain could learn and express an abstract concept in different ways. The media-invariant property of SLR could be seen as a high-level regularization on the intermediate vector, enforcing it to not only capture the latent representation intra each individual media like the auto-encoders, but also their transitions like the mapping models. Moreover, the SLR model is content-specific, which means it only needs to be trained once for a dataset, while used for different tasks. We show how to train a SLR model via dropout and use it for different sequence-to-sequence tasks. Our SLR model is validated on the Youtube2Text and MSR-VTT datasets, achieving superior performance on video-to-sentence task, and the first sentence-to-video results. Xu Shen 0001, Xinmei Tian 0001, Jun Xing, Yong Rui, Dacheng Tao |
AAAI | 2 |
| 2018 | A Twofold Siamese Network for Real-Time Object TrackingabstractObserving that Semantic features learned in an image classification task and Appearance features learned in a similarity matching task complement each other, we build a twofold Siamese network, named SA-Siam, for real-time object tracking. SA-Siam is composed of a semantic branch and an appearance branch. Each branch is a similaritylearning Siamese network. An important design choice in SA-Siam is to separately train the two branches to keep the heterogeneity of the two types of features. In addition, we propose a channel attention mechanism for the semantic branch. Channel-wise weights are computed according to the channel activations around the target position. While the inherited architecture from SiamFC [3] allows our tracker to operate beyond real-time, the twofold design and the attention mechanism significantly improve the tracking performance. The proposed SA-Siam outperforms all other real-time trackers by a large margin on OTB-2013/50/100 benchmarks. Anfeng He, Chong Luo 0001, Xinmei Tian 0001, Wenjun Zeng 0001 |
CVPR | 3 |
| 2018 | Deep Boosting for Image Denoising
Chang Chen 0004, Zhiwei Xiong, Xinmei Tian 0001, Feng Wu 0001 |
ECCV (11) | 3 |
| 2018 | Deep Domain Generalization via Conditional Invariant Adversarial Networks
Xinmei Tian 0001, Mingming Gong, Tongliang Liu, Kun Zhang 0001, Dacheng Tao |
ECCV (15) | 2 |
| 2018 | Local Convolutional Neural Networks for Person Re-IdentificationabstractRecent works have shown that person re-identification can be substantially improved by introducing attention mechanisms, which allow learning both global and local representations. However, all these works learn global and local features in separate branches. As a consequence, the interaction/boosting of global and local information are not allowed, except in the final feature embedding layer. In this paper, we propose local operations as a generic family of building blocks for synthesizing global and local information in any layer. This building block can be inserted into any convolutional networks with only a small amount of prior knowledge about the approximate locations of local parts. For the task of person re-identification, even with only one local block inserted, our local convolutional neural networks (Local CNN) can outperform state-of-the-art methods consistently on three large-scale benchmarks, including Market-1501, CUHK03, and DukeMTMC-ReID. Jiwei Yang, Xu Shen 0001, Xinmei Tian 0001, Houqiang Li, Jianqiang Huang 0001, Xian-Sheng Hua 0001 |
ACM Multimedia | 3 |
| 2018 | Deep Domain Adaptation Hashing with Adversarial LearningabstractThe recent advances in deep neural networks have demonstrated high capability in a wide variety of scenarios. Nevertheless, fine-tuning deep models in a new domain still requires a significant amount of labeled data despite expensive labeling efforts. A valid question is how to leverage the source knowledge plus unlabeled or only sparsely labeled target data for learning a new model in target domain. The core problem is to bring the source and target distributions closer in the feature space. In the paper, we facilitate this issue in an adversarial learning framework, in which a domain discriminator is devised to handle domain shift. Particularly, we explore the learning in the context of hashing problem, which has been studied extensively due to its great efficiency in gigantic data. Specifically, a novel Deep Domain Adaptation Hashing with Adversarial learning (DeDAHA) architecture is presented, which mainly consists of three components: a deep convolutional neural networks (CNN) for learning basic image/frame representation followed by an adversary stream on one hand to optimize the domain discriminator, and on the other, to interact with each domain-specific hashing stream for encoding image representation to hash codes. The whole architecture is trained end-to-end by jointly optimizing two types of losses, i.e., triplet ranking loss to preserve the relative similarity ordering in the input triplets and adversarial loss to maximally fool the domain discriminator with the learnt source and target feature distributions. Extensive experiments are conducted on three domain transfer tasks, including cross-domain digits retrieval, image to image and image to video transfers, on several benchmarks. Our DeDAHA framework achieves superior results when compared to the state-of-the-art techniques. Fuchen Long, Ting Yao 0003, Qi Dai 0001, Xinmei Tian 0001, Jiebo Luo 0001, Tao Mei 0001 |
SIGIR | 4 |
| 2018 | Relative Aesthetic Quality RankingabstractGiven a set of images, we want to rank them according to their aesthetic quality, which is beneficial to various applications. Most existing methods consider aesthetic quality assessment to be a classification or regression problem, but describing aesthetic quality using absolute labels or scores is imprecise and unnatural. To solve this problem, some researchers have proposed listwise approaches for relative aesthetic quality ranking. However, only limited success has been achieved because 1) the ranking order of all images is employed in training although it is not reasonable to compare images with different visual content and 2) pre-defined features cannot describe the aesthetic attributes well. To address these challenges, we introduce a novel visual similarity-based approach to generate reasonable pairs in this paper. With these well-selected pairs, a ranking model can be trained to rank images according to their aesthetic quality. Furthermore, to learn more effective aesthetic features and obtain a better ranking function, we design a dual-channel deep neural network as the ranking model. The proposed approach is evaluated on the AVA dataset, and the experimental results demonstrate that our approach significantly outperforms the state-of-the-art methods. Xinmei Tian 0001, Yujiao Long |
SMC | 1 |
| 2018 | Joints kinetic and relational features for action recognition
Xinmei Tian 0001 |
Signal Process. | 1 |
| 2018 | Data mining in human activity analysis
Xinmei Tian 0001, Weifeng Liu 0001, Fionn Murtagh |
Signal Process. | 1 |
| 2018 | PageSense: Toward Stylewise Contextual Advertising via Visual Analysis of Web PagesabstractThe Internet has emerged as the most effective and a highly popular medium for advertising. Current contextual advertising platforms need publishers to manually change the original structure of their Web pages and predefine the position and style of embedded ads. Although publishers spend significant effort optimizing their Web page layout, a large number of Web pages contain noticeable blank regions. We present an innovative stylewise advertising platform for contextual advertising, called PageSense. The “style” of Web pages refers to the visual appearance of a Web page, such as color and layout. PageSense aims to associate style-consistent ads with Web pages. It provides two advertising options: 1) If publishers predefine ad positions within Web pages, PageSense will analyze the page style and select ads, which are consistent with the Web page layout, and 2) if publishers impose no constraints for ad placement, PageSense will automatically detect blank regions, select the most nonintrusive region for ad insertion, associate color-consistent ads with the Web pages, and deliver them to blank regions without breaking the original Web page style. Our experiments have verified the effectiveness of PageSense as a complement to existing contextual advertising. Tao Mei 0001, Lusong Li, Xinmei Tian 0001, Dacheng Tao, Chong-Wah Ngo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Multigranular Event Recognition of Personal Photo AlbumsabstractPeople are taking more photos than ever before in recent years. To effectively organize these personal photos, the photos are usually assigned to albums according to their events. An efficient way to manage our photos would be if we could recognize the events of the albums automatically. In this paper, we study the problem of recognizing events in personal photo albums. Recognizing events in photo albums is a new challenge since the contents of photos in albums are more complicated than in traditional single-photo tasks, since not all photos in an album are relevant to the event and a single photo in an album often fails to convey the meaningful event semantic behind the album. To solve this problem, we introduce an attention network to learn the representations of photo albums. Then, we adopt a hierarchical model to recognize events from coarse to fine using multigranular features. We evaluate our model on two real-world datasets consisting of personal albums; we find that our model achieves promising results. Cong Guo 0002, Xinmei Tian 0001, Tao Mei 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | On Better Exploring and Exploiting Task Relationships in Multitask Learning: Joint Model and Feature LearningabstractMultitask learning (MTL) aims to learn multiple tasks simultaneously through the interdependence between different tasks. The way to measure the relatedness between tasks is always a popular issue. There are mainly two ways to measure relatedness between tasks: common parameters sharing and common features sharing across different tasks. However, these two types of relatedness are mainly learned independently, leading to a loss of information. In this paper, we propose a new strategy to measure the relatedness that jointly learns shared parameters and shared feature representations. The objective of our proposed method is to transform the features of different tasks into a common feature space in which the tasks are closely related and the shared parameters can be better optimized. We give a detailed introduction to our proposed MTL method. Additionally, an alternating algorithm is introduced to optimize the nonconvex objection. A theoretical bound is given to demonstrate that the relatedness between tasks can be better measured by our proposed MTL algorithm. We conduct various experiments to verify the superiority of the proposed joint model and feature MTL method. Xinmei Tian 0001, Tongliang Liu, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | Continuous DropoutabstractDropout has been proven to be an effective algorithm for training robust deep networks because of its ability to prevent overfitting by avoiding the co-adaptation of feature detectors. Current explanations of dropout include bagging, naive Bayes, regularization, and sex in evolution. According to the activation patterns of neurons in the human brain, when faced with different situations, the firing rates of neurons are random and continuous, not binary as current dropout does. Inspired by this phenomenon, we extend the traditional binary dropout to continuous dropout. On the one hand, continuous dropout is considerably closer to the activation characteristics of neurons in the human brain than traditional binary dropout. On the other hand, we demonstrate that continuous dropout has the property of avoiding the co-adaptation of feature detectors, which suggests that we can extract more independent feature detectors for model averaging in the test stage. We introduce the proposed continuous dropout to a feedforward neural network and comprehensively compare it with binary dropout, adaptive dropout, and DropConnect on Modified National Institute of Standards and Technology, Canadian Institute for Advanced Research-10, Street View House Numbers, NORB, and ImageNet large scale visual recognition competition-12. Thorough experiments demonstrate that our method performs better in preventing the co-adaptation of feature detectors and improves test performance. Xu Shen 0001, Xinmei Tian 0001, Tongliang Liu, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Patch Reordering: A NovelWay to Achieve Rotation and Translation Invariance in Convolutional Neural NetworksabstractConvolutional Neural Networks (CNNs) have demonstrated state-of-the-art performance on many visual recognition tasks. However, the combination of convolution and pooling operations only shows invariance to small local location changes in meaningful objects in input. Sometimes, such networks are trained using data augmentation to encode this invariance into the parameters, which restricts the capacity of the model to learn the content of these objects. A more efficient use of the parameter budget is to encode rotation or translation invariance into the model architecture, which relieves the model from the need to learn them. To enable the model to focus on learning the content of objects other than their locations, we propose to conduct patch ranking of the feature maps before feeding them into the next layer. When patch ranking is combined with convolution and pooling operations, we obtain consistent representations despite the location of meaningful objects in input. We show that the patch ranking module improves the performance of the CNN on many benchmark tasks, including MNIST digit recognition, large-scale image recognition, and image retrieval. Xu Shen 0001, Xinmei Tian 0001, Shaoyan Sun, Dacheng Tao |
AAAI | 2 |
| 2017 | A Small Scale Multi-Column Network for Aesthetic Classification Based on Multiple Attributes
Chaoqun Wan, Xinmei Tian 0001 |
ICONIP (1) | 2 |
| 2017 | Semi-supervised Coefficient-Based Distance Metric Learning
Zhangcheng Wang, Xinmei Tian 0001 |
ICONIP (1) | 3 |
| 2017 | Layer-Wise Training to Create Efficient Convolutional Neural Networks
Linghua Zeng, Xinmei Tian 0001 |
ICONIP (2) | 2 |
| 2017 | Classification and Representation Joint Learning via Deep NetworksabstractDeep learning has been proven to be effective for classification problems. However, the majority of previous works trained classifiers by considering only class label information and ignoring the local information from the spatial distribution of training samples. In this paper, we propose a deep learning framework that considers both class label information and local spatial distribution information between training samples. A two-channel network with shared weights is used to measure the local distribution. The classification performance can be improved with more detailed information provided by the local distribution, particularly when the training samples are insufficient. Additionally, the class label information can help to learn better feature representations compared with other feature learning methods that use only local distribution information between samples. The local distribution constraint between sample pairs can also be viewed as a regularization of the network, which can efficiently prevent the overfitting problem. Extensive experiments are conducted on several benchmark image classification datasets, and the results demonstrate the effectiveness of our proposed method. Xinmei Tian 0001, Xu Shen 0001, Dacheng Tao |
IJCAI | 2 |
| 2016 | Regularized Large Margin Distance Metric LearningabstractDistance metric learning plays an important role in many applications, such as classification and clustering. In this paper, we propose a novel distance metric learning using two hinge losses in the objective function. One is the constraint of the pairs which makes the similar pairs (the same label) closer and the dissimilar (different labels) pairs separated as far as possible. The other one is the constraint of the triplets which makes the largest distance between pairs intra the class larger than the smallest distance between pairs inter the classes. Previous works only consider one of the two kinds of constraints. Additionally, different from the triplets used in previous works, we just need a small amount of such special triplets. This improves the efficiency of our proposed method. Consider the situation in which we might not have enough labeled samples, we extend the proposed distance metric learning into a semi-supervised learning framework. Experiments are conducted on several landmark datasets and the results demonstrate the effectiveness of our proposed method. Xinmei Tian 0001, Dacheng Tao |
ICDM | 2 |
| 2016 | Transform-Invariant Convolutional Neural Networks for Image Classification and SearchabstractConvolutional neural networks (CNNs) have achieved state-of-the-art results on many visual recognition tasks. However, current CNN models still exhibit a poor ability to be invariant to spatial transformations of images. Intuitively, with sufficient layers and parameters, hierarchical combinations of convolution (matrix multiplication and non-linear activation) and pooling operations should be able to learn a robust mapping from transformed input images to transform-invariant representations. In this paper, we propose randomly transforming (rotation, scale, and translation) feature maps of CNNs during the training stage. This prevents complex dependencies of specific rotation, scale, and translation levels of training images in CNN models. Rather, each convolutional kernel learns to detect a feature that is generally helpful for producing the transform-invariant answer given the combinatorially large variety of transform levels of its input feature maps. In this way, we do not require any extra training supervision or modification to the optimization process and training images. We show that random transformation provides significant improvements of CNNs on many benchmark tasks, including small-scale image recognition, large-scale image recognition, and image retrieval. Xu Shen 0001, Xinmei Tian 0001, Anfeng He, Shaoyan Sun, Dacheng Tao |
ACM Multimedia | 2 |
| 2016 | Visual Re-ranking Through Greedy Selection and Rank Fusion
Bin Lin 0010, Ai Wei, Xinmei Tian 0001 |
MMM (1) | 3 |
| 2016 | Learning Relative Aesthetic Quality with a Pairwise Approach
Xinmei Tian 0001 |
MMM (1) | 2 |
| 2016 | Multi-organ plant identification with multi-column deep convolutional neural networksabstractAutomatically identifying plants from images is a hot research topic due to its importance in production and science popularization. This process attempts to automatically identify the name of a plant with a known taxon from a given image. The majority of existing studies on automatic plant identification focus on identifying plants with a single organ, such as flower, leaf, or fruits. Plant identification using a single organ is not sufficiently reliable because different plants many have similar organs. To overcome this problem, this paper is devoted to automatically identifying plants by combining multiple organs of plants. Specifically, we propose a multi-column deep convolutional neural networks (MCDCNN) model to combine multiple organs for efficient plant identification. Extensive experiments demonstrate the effectiveness of our model, and the plant identification performance is greatly improved. Anfeng He, Xinmei Tian 0001 |
SMC | 2 |
| 2016 | Flickr group recommendation using rich social media information
Cong Guo 0002, Xinmei Tian 0001 |
Neurocomputing | 3 |
| 2016 | Multi-modal and multi-scale photo collection summarization
Xu Shen 0001, Xinmei Tian 0001 |
Multim. Tools Appl. | 2 |
| 2016 | Monet: A System for Reliving Your Memories by Theme-Based Photo StorytellingabstractWith the ever-increasing use of smartphones and digital cameras, people are now able to take photos anywhere and anytime. Most of these photos simply end up stored in the cloud without further interaction. This occurs because we lack intelligent services to organize these personal photos well. Therefore, there is an urgent need for such a system to enable people to relive their memories by turning their photos into stories. This paper presents a storytelling system named Monet, which automatically creates interesting stories from personal photos by mimicking cinematic knowledge based on a set of predesigned editing styles. The system consists of two stages: photo summarization, which selects a subset of the “best” photos to represent a photo collection, and story remixing, which generates a stylish music video from the selected photos. During photo summarization, photos are grouped into events based on multimodal features (time and location). The “best” photos are then selected according to visual quality, event representativeness, and diversity. The second stage, story remixing, automatically selects an appropriate theme-dependent editing style based on the photo content. Each selected photo is converted to a video clip by applying a virtual camera with appropriate motions. A series of video effects, color filters, shapes, and transitions are then applied to the video clips according to cinematic rules. The generated video is finally multiplexed with a music clip to generate the story. Evaluations show that our system achieves superior performance to state-of-the-art photo event detection and story generation systems. Xu Shen 0001, Tao Mei 0001, Xinmei Tian 0001, Nenghai Yu, Yong Rui |
IEEE Trans. Multim. | 4 |
| 2015 | On the selection of trending image from the webabstractThe recommendation of trending images has become a popular feature used by commercial search engines to attract public attention. By browsing through trending images, search engine users can discover trending events at a glance. However, the selection of trending images is very challenging and remains an open issue. Most existing work is highly dependent on editorial efforts, though some preliminarily identify a few plain features for trending images. In this paper, we investigate a set of perceptual factors that can distinguish trending images from common ones. We propose a set of trending-aware features based on several common criteria, which reflect the characteristics of trending images. We further construct a manually labeled dataset based on a commercial search engine's query log over a two-week timespan. We evaluate our proposed method on this dataset and the results demonstrate its effectiveness. Dongfei Yu, Xinmei Tian 0001, Tao Mei 0001, Yong Rui |
ICME | 2 |
| 2015 | Multi-Task Model and Feature Joint Learning
Xinmei Tian 0001, Tongliang Liu, Dacheng Tao |
IJCAI | 2 |
| 2015 | Photo Quality Assessment with DCNN that Understands Image Well
Xu Shen 0001, Houqiang Li, Xinmei Tian 0001 |
MMM (2) | 4 |
| 2015 | Event recognition in personal photo collections using hierarchical model and multiple featuresabstractWith the proliferation of digital cameras and mobile devices, people are taking many more photos than ever before. The explosive growth of personal photos leads to problems of photo organization and management. There is a growing need for tools to automatically manage photo collections. Recognizing events in photo collections is one efficient way to organize photos. The use of textual event labels can allow us to categorize and locate an event without browsing through an entire photo collection. Most existing research on this topic focuses on recognizing events from single photos and only a few studies have examined event recognition in personal photo collections. In this paper, we propose a hierarchical model to recognize events in personal photo collections using multiple features, including time, objects, and scenes. Since some events are more difficult to identify and categorize, ambiguous events require fine event classifiers, while the coarse categories of the events can be sufficiently organized with a coarse event classifier. We evaluate our coarse-to-fine hierarchical model on a real-world dataset consisting of personal photo collections, and our model achieves promising results. Cong Guo 0002, Xinmei Tian 0001 |
MMSP | 2 |
| 2015 | Query-Adaptive Image Search Re-ranking Using Deep Convolutional Neural Network FeatureabstractImage search re-ranking, as an effective tool to improve the text-based image search result, has been adopted by many commercial search engines nowadays. Given a query keyword, images are first retrieved based on the textual information. Then visual features are extracted from images to reorder them by mining their visual patterns. However, the popular visual features applied in re-ranking are not informative enough. Besides, the parameters for the re-ranking models are set equally for all queries, which fails to cope with the variability of different queries. In this paper, we propose a novel re-ranking method which adopts informative visual features for image representation and adaptively re-rank the images. Specifically, we adopt a proven successful DCNN feature (deep convolutional neural network), which shows the excellent performance in many computer vision fields, to calculate the visual similarities between images. For each query, the parameters for the image search re-ranking model is adaptively determined using the QDE (query difficulty estimation) method. Experiments are conducted on the INRIA web353 dataset. The experimental results demonstrate that our method achieves significant improvement over state-of-the-art methods. Bin Lin 0010, Xinmei Tian 0001 |
SMC | 2 |
| 2015 | Multi-level photo quality assessment with multi-view features
Xinmei Tian 0001 |
Neurocomputing | 2 |
| 2015 | Multi-task proximal support vector machine
Xinmei Tian 0001, Mingli Song, Dacheng Tao |
Pattern Recognit. | 2 |
| 2015 | Query difficulty estimation via relevance prediction for image retrieval
Qianghuai Jia, Xinmei Tian 0001 |
Signal Process. | 2 |
| 2015 | Exploration of Image Search Results Quality AssessmentabstractImage retrieval plays an increasingly important role in our daily lives. There are many factors which affect the quality of image search results, including chosen search algorithms, ranking functions, and indexing features. Applying different settings for these factors generates search result lists with varying levels of quality. However, no setting can always perform optimally for all queries. Therefore, given a set of search result lists generated by different settings, it is crucial to automatically determine which result list is the best in order to present it to users. This paper aims to solve this problem and makes four main innovations. First, a preference learning model is proposed to quantitatively study and formulate the best image search result list identification problem. Second, a set of valuable preference learning related features is proposed by exploring the visual characters of returned images. Third, a query-dependent preference learning model is further designed for building a more precise and query-specific model. Fourth, the proposed approach has been tested on a variety of applications including re-ranking ability assessment, optimal search engine selection, and synonymous query suggestion. Extensive experimental results on three image search datasets demonstrate the effectiveness and promising potential of the proposed method. Xinmei Tian 0001, Yijuan Lu, Nate Stender, Linjun Yang, Dacheng Tao |
IEEE Trans. Big Data | 1 |
| 2015 | Image Search Reranking With Hierarchical Topic AwarenessabstractWith much attention from both academia and industrial communities, visual search reranking has recently been proposed to refine image search results obtained from text-based image search engines. Most of the traditional reranking methods cannot capture both relevance and diversity of the search results at the same time. Or they ignore the hierarchical topic structure of search result. Each topic is treated equally and independently. However, in real applications, images returned for certain queries are naturally in hierarchical organization, rather than simple parallel relation. In this paper, a new reranking method "topic-aware reranking (TARerank)" is proposed. TARerank describes the hierarchical topic structure of search results in one model, and seamlessly captures both relevance and diversity of the image search results simultaneously. Through a structured learning framework, relevance and diversity are modeled in TARerank by a set of carefully designed features, and then the model is learned from human-labeled training samples. The learned model is expected to predict reranking results with high relevance and diversity for testing queries. To verify the effectiveness of the proposed method, we collect an image search dataset and conduct comparison experiments on it. The experimental results demonstrate that the proposed TARerank outperforms the existing relevance-based and diversified reranking methods. Xinmei Tian 0001, Linjun Yang, Yijuan Lu, Qi Tian 0001, Dacheng Tao |
IEEE Trans. Cybern. | 1 |
| 2015 | Query-Dependent Aesthetic Model With Deep Learning for Photo Quality AssessmentabstractThe automatic assessment of photo quality from an aesthetic perspective is a very challenging problem. Most existing research has predominantly focused on the learning of a universal aesthetic model based on hand-crafted visual descriptors . However, this research paradigm can achieve only limited success because (1) such hand-crafted descriptors cannot well preserve abstract aesthetic properties , and (2) such a universal model cannot always capture the full diversity of visual content. To address these challenges, we propose in this paper a novel query-dependent aesthetic model with deep learning for photo quality assessment. In our method, deep aesthetic abstractions are discovered from massive images , whereas the aesthetic assessment model is learned in a query- dependent manner. Our work addresses the first problem by learning mid-level aesthetic feature abstractions via powerful deep convolutional neural networks to automatically capture the underlying aesthetic characteristics of the massive training images . Regarding the second problem, because photographers tend to employ different rules of photography for capturing different images , the aesthetic model should also be query- dependent . Specifically, given an image to be assessed, we first identify which aesthetic model should be applied for this particular image. Then, we build a unique aesthetic model of this type to assess its aesthetic quality. We conducted extensive experiments on two large-scale datasets and demonstrated that the proposed query-dependent model equipped with learned deep aesthetic abstractions significantly and consistently outperforms state-of-the-art hand-crafted feature -based and universal model-based methods. Xinmei Tian 0001, Kuiyuan Yang, Tao Mei 0001 |
IEEE Trans. Multim. | 1 |
| 2015 | Query Difficulty Estimation for Image Search With Query Reconstruction ErrorabstractCurrent image search engines suffer from a radical variance in retrieval performance over different queries. It is therefore desirable to identify those “difficult” queries in order to handle them properly. Query difficulty estimation is an attempt to predict the performance of the search results returned by an image search system. Most existing methods for query difficulty estimation focus on investigating statistical characteristics of the returned images only, while neglecting very important information , i.e., the query and its relationship with returned images. This relationship plays a crucial role in query difficulty estimation and should be explored further. In this paper we propose a novel query difficulty estimation method with query reconstruction error. This method is proposed based on the observation that, given the images returned for an unknown query, we can easily deduce what the query is from those images if the search results are high quality (i.e., lots of relevant images returned); otherwise, it is difficult to deduce the original query. Therefore, we propose to predict the query difficulty by measuring to what extent the original query can be recovered from the image search results. Specifically, we first reconstruct a visual query from the returned images to summarize their visual theme, and then use the reconstruction error, i.e., the distance between the original textual query and the reconstructed visual query, to estimate the query difficulty. We conduct extensive experiments on two real-world Web image datasets and demonstrate the effectiveness of the proposed method. Xinmei Tian 0001, Qianghuai Jia, Tao Mei 0001 |
IEEE Trans. Multim. | 1 |
| 2014 | User specific friend recommendation in social media communityabstractSocial networks nowadays have become an important form of communication in which users can post their current status or share their lives by mobile phones or the Web. In this paper, we develop an effective and efficient model to estimate continuous tie strength between users for friend recommendation with the heterogeneous data from social media community. We categorize those multimodal data into two classes: interaction data (e.g., comments, marking favorite photos) and similarity data(e.g., common friends, groups, tags, geo, visual). We propose to use asymmetric relationship in the interaction data for tie strength estimation instead of using the conventional symmetric ones. Furthermore, by exploring the behavior of users in a social media community, we find that the tie strength between users can be approximately modeled as a linear function of their social connections. Based on this observation, we propose an effective and highly efficient user specific linear model for the tie strength estimation. The experiments on a popular social network show promising results and demonstrate the effectiveness of our proposed method. Cong Guo 0002, Xinmei Tian 0001, Tao Mei 0001 |
ICME | 2 |
| 2014 | Query difficulty estimation via pseudo relevance feedback for image searchabstractQuery difficulty estimation (QDE) attempts to automatically predict the performance of the search results returned for a given query. QDE has been widely investigated in text document retrieval for many years. However, few research works have been explored in image retrieval. State-of-the-art QDE methods in image retrieval mainly investigate the statistical characteristics (coherence, robustness, etc.) of the returned images to derive a value for indicating the query difficulty degree. To the best of our knowledge, little research has been done to directly estimate the real retrieval performance of the search results, such as average precision, instead of only an indicator. In this paper, we propose a novel query difficulty estimation approach which automatically estimate the average precision of the image search results. Specifically, we first select a set of query relevant and query irrelevant images for each query via pseudo relevance feedback. Then an efficient and effective voting scheme is proposed to estimate the relevance label of each image in the search results. Based on the images' relevance labels, the average precision of the search results returned for the given query is derived. The experimental results on a benchmark image search dataset demonstrate the effectiveness of the proposed method. Qianghuai Jia, Xinmei Tian 0001, Tao Mei 0001 |
ICME | 2 |
| 2014 | Click-through-based Subspace Learning for Image SearchabstractOne of the fundamental problems in image search is to rank image documents according to a given textual query. We address two limitations of the existing image search engines in this paper. First, there is no straightforward way of comparing textual keywords with visual image content. Image search engines therefore highly depend on the surrounding texts, which are often noisy or too few to accurately describe the image content. Second, ranking functions are trained on query-image pairs labeled by human labelers, making the annotation intellectually expensive and thus cannot be scaled~up. Yingwei Pan, Ting Yao 0003, Xinmei Tian 0001, Houqiang Li, Chong-Wah Ngo |
ACM Multimedia | 3 |
| 2014 | Effective and efficient photo quality assessmentabstractAutomatic photo quality assessment from the perspective of visual aesthetics is a hot research topic due to its potential need in numerous applications. It tries to automatically determine whether a given image has “high” or “low” quality according to the image's visual content. Most existing researches in photo quality assessment predominantly focus on exploring hand-crafted features which may be potentially related to high-level aesthetic attributes. Most of those features are designed under the guidance of some common photography rules and prior knowledge. However, due to the subjectivity and complexity of humans' aesthetic activities, automatic image aesthetic quality assessment is very challenging. Those features are not effective enough and show varying performance on different datasets. Besides, they often require high computational cost. In this paper, we propose a set of compact aesthetic features which are not only effective but also highly efficient. We test those features on two large scale real world image datasets. The experimental results demonstrate that the proposed features achieve the best performance consistently over different datasets with a much lower computational complexity. Xinmei Tian 0001 |
SMC | 2 |
| 2013 | Semantic-Spatial Matching for image classificationabstractSpatial Pyramid Matching (SPM) has been proven a simple but effective extension to bag-of-visual-words image representation for spatial layout information compensation. SPM describes image in coarse-to-fine scale by partitioning the image into blocks over multiple levels and the features extracted from each block are concatenated into a long vector representation. Based on the assumption that images from the same class have similar spatial configurations, SPM matches the blocks from different images according to their spatial layout, by aligning all blocks from an image in a fixed spatial order. However, target objects may appear at any location in the image with various backgrounds. Therefore, the fixed spatial matching in SPM fails to match similar objects located different locations. To solve this problem, we propose an effective and efficient block matching method, Semantic-Spatial Matching (SSM). In this method, not only the spatial layout but also the semantic content is considered for block matching. The experiments on two benchmark image classification datasets demonstrate the effectiveness of SSM. Yupeng Yan, Xinmei Tian 0001, Linjun Yang, Yijuan Lu, Houqiang Li |
ICME | 2 |
| 2013 | Discriminative codebook learning for Web image search
Xinmei Tian 0001, Yijuan Lu |
Signal Process. | 1 |
| 2012 | Query Difficulty Prediction for Web Image SearchabstractImage search plays an important role in our daily life. Given a query, the image search engine is to retrieve images related to it. However, different queries have different search difficulty levels. For some queries, they are easy to be retrieved (the search engine can return very good search results). While for others, they are difficult (the search results are very unsatisfactory). Thus, it is desirable to identify those “difficult” queries in order to handle them properly. Query difficulty prediction (QDP) is an attempt to predict the quality of the search result for a query over a given collection. QDP problem has been investigated for many years in text document retrieval, and its importance has been recognized in the information retrieval (IR) community. However, little effort has been conducted on the image query difficulty prediction problem for image search. Compared with QDP in document retrieval, QDP in image search is more challenging due to the noise of textual features and the well-known semantic gap of visual features. This paper aims to investigate the QDP problem in Web image search. A novel method is proposed to automatically predict the quality of image search results for an arbitrary query. This model is built based on a set of valuable features that are designed by exploring the visual characteristic of images in the search results. The experiments on two real image search datasets demonstrate the effectiveness of the proposed query difficulty prediction method. Two applications, including optimal image search engine selection and search results merging, are presented to show the promising applicability of QDP. Xinmei Tian 0001, Yijuan Lu, Linjun Yang |
IEEE Trans. Multim. | 1 |
| 2012 | Correction to "Bayesian Visual Reranking"abstractIn the above titled paper (ibid., vol. 13, no. 4, pp. 639-652, Aug. 2011), the first author's name appears incorrectly in the byline as "Xinmie Tian" instead of "Xinmei Tian." The name appears correctly in the biography section. Xinmei Tian 0001, Linjun Yang, Jingdong Wang 0001, Xiuqing Wu, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 1 |
| 2012 | Sparse transfer learning for interactive video search rerankingabstractVisual reranking is effective to improve the performance of the text-based video search. However, existing reranking algorithms can only achieve limited improvement because of the well-known semantic gap between low-level visual features and high-level semantic concepts. In this article, we adopt interactive video search reranking to bridge the semantic gap by introducing user's labeling effort. We propose a novel dimension reduction tool, termed sparse transfer learning (STL), to effectively and efficiently encode user's labeling information. STL is particularly designed for interactive video search reranking. Technically, it (a) considers the pair-wise discriminative information to maximally separate labeled query relevant samples from labeled query irrelevant ones, (b) achieves a sparse representation for the subspace to encodes user's intention by applying the elastic net penalty, and (c) propagates user's labeling information from labeled samples to unlabeled samples by using the data distribution knowledge. We conducted extensive experiments on the TRECVID 2005, 2006 and 2007 benchmark datasets and compared STL with popular dimension reduction algorithms. We report superior performance by using the proposed STL-based interactive video search reranking. Xinmei Tian 0001, Dacheng Tao, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2011 | Learning to judge image search resultsabstractGiven the explosive growth of the Web and the popularity of image sharing Web sites, image retrieval plays an increasingly important role in our daily lives. Search engines aim to provide beneficial image search results to users in response to queries. The quality of image search results depends on many factors: chosen search algorithms, ranking functions, indexing features, the base image database, etc. Applying different settings for these factors generates search result lists with varying levels of quality. Previous research has shown that no setting can always perform optimally for all queries. Therefore, given a set of search result lists generated by different settings, it is crucial to automatically determine which result list is the best in order to present it to users. This paper proposes a novel method to automatically identify the best search result list from a number of candidates. There are three main innovations in this paper. First, we propose a preference learning model to quantitatively study the best image search result identification problem. Second, we propose a set of valuable preference learning related features by exploring the visual characters of returned images. Third, our method shows promising potential in applications such as reranking ability assessment and optimal search engine selection. Experiments on two image search datasets show that our method achieves about 80% prediction accuracy for reranking ability assessment, and selects optimal search engine for about 70% queries correctly. Xinmei Tian 0001, Yijuan Lu, Linjun Yang, Qi Tian 0001 |
ACM Multimedia | 1 |
| 2011 | Bayesian Visual RerankingabstractVisual reranking has been proven effective to refine text-based video and image search results. It utilizes visual information to recover “true” ranking list from the noisy one generated by text-based search, by incorporating both textual and visual information. In this paper, we model the textual and visual information from the probabilistic perspective and formulate visual reranking as an optimization problem in the Bayesian framework, termed Bayesian visual reranking. In this method, the textual information is modeled as a likelihood, to reflect the disagreement between reranked results and text-based search results which is called ranking distance. The visual information is modeled as a conditional prior, to indicate the ranking score consistency among visually similar samples which is called visual consistency. Bayesian visual reranking derives the best reranking results by maximizing visual consistency while minimizing ranking distance. To model the ranking distance more precisely, we propose a novel pair-wise method which measure the ranking distance based on the disagreement in terms of pair-wise orders. For visual consistency, we study three different regularizers to mine the best way for its modeling. We conduct extensive experiments on both video and image search datasets. Experimental results demonstrate the effectiveness of our proposed Bayesian visual reranking. Xinmei Tian 0001, Linjun Yang, Jingdong Wang 0001, Xiuqing Wu, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 1 |
| 2010 | Constrained Metric Learning Via Distance Gap MaximizationabstractVectored data frequently occur in a variety of fields, which are easy to handle since they can be mathematically abstracted as points residing in a Euclidean space. An appropriate distance metric in the data space is quite demanding for a great number of applications. In this paper, we pose robust and tractable metric learning under pairwise constraints that are expressed as similarity judgements between data pairs. The major features of our approach include: 1) it maximizes the gap between the average squared distance among dissimilar pairs and the average squared distance among similar pairs; 2) it is capable of propagating similar constraints to all data pairs; and 3) it is easy to implement in contrast to the existing approaches using expensive optimization such as semidefinite programming. Our constrained metric learning approach has widespread applicability without being limited to particular backgrounds. Quantitative experiments are performed for classification and retrieval tasks, uncovering the effectiveness of the proposed approach. Wei Liu 0005, Xinmei Tian 0001, Dacheng Tao, Jianzhuang Liu |
AAAI | 2 |
| 2010 | Visual Reranking with Local Learning Consistency
Xinmei Tian 0001, Linjun Yang, Xiuqing Wu, Xian-Sheng Hua 0001 |
MMM | 1 |
| 2010 | Active Reranking for Web Image SearchabstractImage search reranking methods usually fail to capture the user's intention when the query term is ambiguous. Therefore, reranking with user interactions, or active reranking, is highly demanded to effectively improve the search performance. The essential problem in active reranking is how to target the user's intention. To complete this goal, this paper presents a structural information based sample selection strategy to reduce the user's labeling efforts. Furthermore, to localize the user's intention in the visual feature space, a novel local-global discriminative dimension reduction algorithm is proposed. In this algorithm, a submanifold is learned by transferring the local geometry and the discriminative information from the labelled images to the whole (global) image database. Experiments on both synthetic datasets and a real Web image search dataset demonstrate the effectiveness of the proposed active reranking scheme, including both the structural information based active sample selection strategy and the local-global discriminative dimension reduction algorithm. Xinmei Tian 0001, Dacheng Tao, Xian-Sheng Hua 0001, Xiuqing Wu |
IEEE Trans. Image Process. | 1 |
| 2009 | Query aware visual similarity propagation for image search rerankingabstractImage search reranking is an effective approach to refining the text-based image search result. In the reranking process, the estimation of visual similarity is critical to the performance. However, the existing measures, based on global or local features, cannot be adapted to different queries. In this paper, we propose to estimate a query aware image similarity by incorporating the global visual similarity, local visual similarity and visual word co-occurrence into an iterative propagation framework. After the propagation, a query aware image similarity combining the advantages of both global and local similarities is achieved and applied to image search reranking. The experiments on a real-world Web image dataset demonstrate that the proposed query aware similarity outperforms the global, local similarity and their linear combination, for image search reranking. Linjun Yang, Xinmei Tian 0001 |
ACM Multimedia | 3 |
| 2008 | Transductive video annotation via local learnable kernel classifierabstractOne crucial problem in transductive video annotation is how to estimate the label from the neighboring samples. Existing methods such as graph-based Gaussian random filed only considered the pair-wise similarity and then propagated the labels based on it. In this paper, we propose a new method from the perspective of local learning, which formulate the prediction of labels from the neighbors into a learning problem. Our contributions lie in two-fold: (1) we propose a new transductive video annotation method based on local kernel classifier; (2) local learnable is proposed to measure whether a sample can be learned from the neighbors well and we employ this measure into the optimization objective. Experiments on TRECVID 2005 dataset prove that the proposed method is effective and the local learning perspective is promising for video annotation. Xinmei Tian 0001, Linjun Yang, Jingdong Wang 0001, Xiuqing Wu, Xian-Sheng Hua 0001 |
ICME | 1 |
| 2008 | Optimized video scene segmentationabstractIn this paper, we propose an optimized video scene segmentation approach with considering both content coherence and temporally contextual dissimilarity. First, a chain structure is constructed by connecting temporally adjacent shots to represent a video. Then the chain is partitioned such that the content within a chain segment is coherent enough and the contextual similarity of temporally adjacent chain segments is small enough. This task is formulated as a ratio function of content coherence and contextual similarity. Finally, we present an effective and efficient hierarchical chain partitioning approach to find the optimal scene segmentation. Experimental results on a set of home videos and feature movies demonstrate the superiority of the proposed approach over several existing key approaches. Jingdong Wang 0001, Xinmei Tian 0001, Linjun Yang, Zhengjun Zha, Xian-Sheng Hua 0001 |
ICME | 2 |
| 2008 | Bayesian video search rerankingabstractContent-based video search reranking can be regarded as a process that uses visual content to recover the "true" ranking list from the noisy one generated based on textual information. This paper explicitly formulates this problem in the Bayesian framework, i.e., maximizing the ranking score consistency among visually similar video shots while minimizing the ranking distance, which represents the disagreement between the objective ranking list and the initial text-based. Different from existing point-wise ranking distance measures, which compute the distance in terms of the individual scores, two new methods are proposed in this paper to measure the ranking distance based on the disagreement in terms of pair-wise orders. Specifically, hinge distance penalizes the pairs with reversed order according to the degree of the reverse, while preference strength distance further considers the preference degree. By incorporating the proposed distances into the optimization objective, two reranking methods are developed which are solved using quadratic programming and matrix computation respectively. Evaluation on TRECVID video search benchmark shows that the performance improvement up to 21% on TRECVID 2006 and 61.11% on TRECVID 2007 are achieved relative to text search baseline. Xinmei Tian 0001, Linjun Yang, Jingdong Wang 0001, Yichen Yang 0001, Xiuqing Wu, Xian-Sheng Hua 0001 |
ACM Multimedia | 1 |