Hong Chang 0001

dblp:02/2689-1 · DBLP profile ↗
← Back
121ranked-venue papers
10as first author
46since 2021 · last 2026
0000-0002-2668-0070ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 95 · 10 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 74 · 4 first-author · 23 since 2021Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Bilateral Transformation of Biased Pseudo-Labels under Distribution Inconsistency
Ruibing Hou, Hong Chang 0001, Minyang Hu, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
Int. J. Comput. Vis.2
2026 RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios
abstract
Human-centric perceptions play a crucial role in real-world applications. While recent human-centric works have achieved impressive progress, these efforts are often constrained to the visual domain and lack interaction with human instructions, limiting their applicability in broader scenarios such as chatbots and sports analysis. This paper introducesReferring Human Perceptions, where a referring prompt specifies the person of interest in an image. To tackle the new task, we propose RefHCM (ReferringHuman-CentricModel), a unified framework to integrate a wide range of human-centric referring tasks. Specifically, RefHCM employs sequence mergers to convert raw multimodal data—including images, text, coordinates, and parsing maps—into semantic tokens. This standardized representation enables RefHCM to reformulate diverse human-centric referring tasks into a sequence-to-sequence paradigm, solved using a plain encoder-decoder transformer architecture. Benefiting from a unified learning strategy, RefHCM effectively facilitates knowledge transfer across tasks and exhibits unforeseen capabilities in handling complex reasoning. This work represents the first attempt to address referring human perceptions with a general-purpose framework, while simultaneously establishing a corresponding benchmark that sets new standards for the field. Extensive experiments showcase RefHCM's competitive and even superior performance across multiple human-centric referring tasks. The code and data are publicly athttps://github.com/JJJYmmm/RefHCM.
Ruibing Hou, Jiahe Zhao, Hong Chang 0001, Shiguang Shan
IEEE Trans. Multim.4
2026 DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
Yinqi Li 0001, Hong Chang 0001, Ruibing Hou, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Multim.2
2025 UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
abstract
Human pose plays a crucial role in the digital age. While recent works have achieved impressive progress in understanding and generating human poses, they often support only a single modality of control signals and operate in isolation, limiting their application in real-world scenarios. This paper presents UniPose, a framework employing Large Language Models (LLMs) to comprehend, generate, and edit human poses across various modalities, including images, text, and 3D SMPL poses. Specifically, we apply a pose tokenizer to convert 3D poses into discrete pose tokens, enabling seamless integration into the LLM within a unified vocabulary. To further enhance the fine-grained pose perception capabilities, we facilitate UniPose with a mixture of visual encoders, among them a pose-specific visual encoder. Benefiting from a unified learning strategy, UniPose effectively transfers knowledge across different pose-relevant tasks, adapts to unseen tasks, and exhibits extended capabilities. This work serves as the first attempt at building a general-purpose framework for pose comprehension, generation, and editing. Extensive experiments highlight UniPose’s competitive and even superior performance across various pose-relevant tasks. Code is available at https://github.com/liyiheng23/UniPose.
Ruibing Hou, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
CVPR3
2025 G2PDiffusion: Cross-Species Genotype-to-Phenotype Prediction Via Evolutionary Diffusion
Zhangyang Gao, Hong Chang 0001, Stan Z. Li, Shiguang Shan, Xilin Chen 0001
ICCV3
2025 HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
Jiahe Zhao, Ruibing Hou, Zejie Tian, Hong Chang 0001, Shiguang Shan
ICCV4
2025 MATS: An Audio Language Model under Text-only Supervision
abstract
Large audio-language models (LALMs), built upon powerful Large Language Models (LLMs), have exhibited remarkable audio comprehension and reasoning capabilities. However, the training of LALMs demands a large corpus of audio-language pairs, which requires substantial costs in both data collection and training resources. In this paper, we propose MATS, an audio-language multimodal LLM designed to handle Multiple Audio task using solely Text-only Supervision. By leveraging pre-trained audio-language alignment models such as CLAP, we develop a text-only training strategy that projects the shared audio-language latent space into LLM latent space, endowing the LLM with audio comprehension capabilities without relying on audio data during training. To further bridge the modality gap between audio and language embeddings within CLAP, we propose the Strongly-related noisy text with audio (Santa) mechanism. Santa maps audio embeddings into CLAP language embedding space while preserving essential information from the audio input. Extensive experiments demonstrate that MATS, despite being trained exclusively on text data, achieves competitive performance compared to recent LALMs trained on large-scale audio-language pairs. The code is publicly available in https://github.com/wangwen-banban/MATS
Wen Wang 0022, Ruibing Hou, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
ICML3
2025 un2CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP
abstract
Contrastive Language-Image Pre-training (CLIP) has become a foundation model and has been applied to various vision and multimodal tasks. However, recent works indicate that CLIP falls short in distinguishing detailed differences in images and shows suboptimal performance on dense-prediction and vision-centric multimodal tasks. Therefore, this work focuses on improving existing CLIP models, aiming to capture as many visual details in images as possible. We find that a specific type of generative models, unCLIP, provides a suitable framework for achieving our goal. Specifically, unCLIP trains an image generator conditioned on the CLIP image embedding. In other words, it inverts the CLIP image encoder. Compared to discriminative models like CLIP, generative models are better at capturing image details because they are trained to learn the data distribution of images. Additionally, the conditional input space of unCLIP aligns with CLIP's original image-text embedding space. Therefore, we propose to invert unCLIP (dubbed un$^2$CLIP) to improve the CLIP model. In this way, the improved image encoder can gain unCLIP's visual detail capturing ability while preserving its alignment with the original text encoder simultaneously. We evaluate our improved CLIP across various tasks to which CLIP has been applied, including the challenging MMVP-VLM benchmark, the dense-prediction open-vocabulary segmentation task, and multimodal large language model tasks. Experiments show that un$^2$CLIP significantly improves the original CLIP and previous CLIP improvement methods. Code and models are available at https://github.com/LiYinqi/un2CLIP.
Yinqi Li 0001, Jiahe Zhao, Hong Chang 0001, Ruibing Hou, Shiguang Shan, Xilin Chen 0001
NeurIPS3
2025 Revisiting Logit Distributions for Reliable Out-of-Distribution Detection
abstract
Out-of-distribution (OOD) detection is critical for ensuring the reliability of deep learning models in open-world applications. While post-hoc methods are favored for their efficiency and ease of deployment, existing approaches often underexploit the rich information embedded in the model’s logits space. In this paper, we propose LogitGap, a novel post-hoc OOD detection method that explicitly exploits the relationship between the maximum logit and the remaining logits to enhance the separability between in-distribution (ID) and OOD samples. To further improve its effectiveness, we refine LogitGap by focusing on a more compact and informative subset of the logit space. Specifically, we introduce a training-free strategy that automatically identifies the most informative logits for scoring. We provide both theoretical analysis and empirical evidence to validate the effectiveness of our approach. Extensive experiments on both vision-language and vision-only models demonstrate that LogitGap consistently achieves state-of-the-art performance across diverse OOD detection scenarios and benchmarks.
Jiachen Liang, Ruibing Hou, Minyang Hu, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
NeurIPS4
2025 ProtInvTree: Deliberate Protein Inverse Folding with Reward-guided Tree Search
abstract
Designing protein sequences that fold into a target 3D structure—known as protein inverse folding—is a fundamental challenge in protein engineering. While recent deep learning methods have achieved impressive performance by recovering native sequences, they often overlook the one-to-many nature of the problem: multiple diverse sequences can fold into the same structure. This motivates the need for a generative model capable of designing diverse sequences while preserving structural consistency. To address this trade-off, we introduce ProtInvTree, the first reward-guided tree-search framework for protein inverse folding. ProtInvTree reformulates sequence generation as a deliberate, step-wise decision-making process, enabling the exploration of multiple design paths and exploitation of promising candidates through self-evaluation, lookahead, and backtracking. We propose a two-stage focus-and-grounding action mechanism that decouples position selection and residue generation. To efficiently evaluate intermediate states, we introduce a jumpy denoising strategy that avoids full rollouts. Built upon pretrained protein language models, ProtInvTree supports flexible test-time scaling by adjusting the search depth and breadth without retraining. Empirically, ProtInvTree outperforms state-of-the-art baselines across multiple benchmarks, generating structurally consistent yet diverse sequences, including those far from the native ground truth. The code is available at https://github.com/A4Bio/ProteinInvBench/.
Xiaoxue Cheng, Zhangyang Gao, Hong Chang 0001, Cheng Tan 0012, Shiguang Shan, Xilin Chen 0001
NeurIPS4
2025 KnowMol: Advancing Molecular Large Language Models with Multi-Level Chemical Knowledge
abstract
The molecular large language models have garnered widespread attention due to their promising potential on molecular applications. However, current molecular large language models face significant limitations in understanding molecules due to inadequate textual descriptions and suboptimal molecular representation strategies during pretraining. To address these challenges, we introduce KnowMol-100K, a large-scale dataset with 100K fine-grained molecular annotations across multiple levels, bridging the gap between molecules and textual descriptions. Additionally, we propose chemically-informative molecular representation, effectively addressing limitations in existing molecular representation strategies. Building upon these innovations, we develop KnowMol, a state-of-the-art multi-modal molecular large language model. Extensive experiments demonstrate that KnowMol achieves superior performance across molecular understanding and generation tasks.
Zaifei Yang, Hong Chang 0001, Ruibing Hou, Shiguang Shan, Xilin Chen 0001
NeurIPS2
2025 PIT: A Plug-and-Play Image Translator for Making Off-the-Shelf Models Adapt to Corruptions
abstract
Visual recognition models pretrained on clean images usually do not perform well in the presence of image corruptions, such as blurring or noise, which limits their applicability in real-world scenarios. To solve this problem, existing approaches usually design complex data augmentations to train a robust model from scratch or adapt a pretrained model to corrupted scenarios. These approaches ignore the existence of the large number of deployed models in our community, causing extensive computation and storage costs for making deployed models adapted. Based on this consideration, this paper focuses on solving a practical problem of making many clean-image-pretrained models adapt to unlabeled corrupted images through one training procedure. To this end, we aim to learn a Plug-and-play Image Translator (PIT) that can be directly combined with recognition models after training. Existing approaches, such as vanilla image translation and restoration, are not proper for solving this problem, as they are mostly based on supervised training and are not recognition-oriented. To address this issue, we propose a recognition-oriented unsupervised image translation framework to make PIT produce images with indistinguishable recognition predictions from the clean ones. We verify the effectiveness of PIT on several recognition tasks and show that PIT boosts the performance of clean-image-pretrained models significantly in the presence of image corruptions.
Yinqi Li 0001, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Inference Calibration of Vision-Language Foundation Models for Zero-Shot and Few-Shot Learning
abstract
Contrastive Language-Image Pre-training (CLIP) models exhibit impressive zero-shot performance across various downstream cross-modal tasks by simply computing the dot product between image and text features. CLIP is pre-trained on large-scale image-text pairs using the InfoNCE loss, which maximizes the cosine similarity of positive image-text pairs while minimizing the similarity of negative pairs. However, an objective mismatch exists between the downstream usage and the pre-training phase, as the inference phase fails to exploit information from negative samples. Intuitively, since the CLIP model has been optimized based on the InfoNCE loss, the downstream usage should also be in alignment. In this paper, we start from analyzing the InfoNCE loss and derive its upper bound. Our derivation reveals that the dot-product operation serves a zero-order approximation of this upper bound, while a centralization operation represents a first-order approximation. To address the objective mismatch problem, we propose a novel method, Inference Calibration (IC), which leverages the first-order and second-order moments of data distribution to calibrate features for zero-shot and few-shot scenarios. Experiments on various cross-modal tasks demonstrate the effectiveness of IC in both zero-shot and few-shot scenarios over dot-product operation and other comparative methods. • The downstream usage of CLIP model mismatches its pre-training objective. • Previous popular inference methods are the approximation of pre-training objective. • Proposed method mitigate objective mismatch problem under both zero-shot and few-shot settings.
Minyang Hu, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
Pattern Recognit. Lett.2
2025 Clothes-Changing Person Re-Identification With Feasibility-Aware Intermediary Matching
abstract
Current clothes-changing person re-identification (re-id) approaches usually perform retrieval based on clothes-irrelevant features, while neglecting the potential of clothes-relevant features. However, we observe that relying solely on clothes-irrelevant features for clothes-changing re-id is limited, since they often lack adequate identity information and suffer from large intra-class variations. On the contrary, clothes-relevant features can be used to discover same-clothes intermediaries that possess informative identity clues. Based on this observation, we propose a Feasibility-Aware Intermediary Matching (FAIM) framework to additionally utilizeclothes-relevant featuresfor retrieval. First, an Intermediary Matching (IM) module is designed to perform an intermediary-assisted matching process. This process involves using clothes-relevant features to find informative intermediates, and then using clothes-irrelevant features of these intermediates to complete the matching. Second, in order to reduce the negative effect of low-quality intermediaries, an Intermediary-Based Feasibility Weighting (IBFW) module is designed to evaluate the feasibility of intermediary matching process by assessing the quality of intermediaries. Extensive experiments demonstrate that our method outperforms state-of-the-art methods on several widely-used clothes-changing re-id benchmarks.
Jiahe Zhao, Ruibing Hou, Hong Chang 0001, Xinqian Gu, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Multim.3
2024 An Information Theoretical View for Out-of-Distribution Detection
Jinjing Hu, Wenrui Liu 0004, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
ECCV (55)3
2024 Scalable Modular Network: A Framework for Adaptive Learning via Agreement Routing
abstract
In this paper, we propose a novel modular network framework, called Scalable Modular Network (SMN), which enables adaptive learning capability and supports integration of new modules after pre-training for better adaptation. This adaptive capability comes from a novel design of router within SMN, named agreement router, which selects and composes different specialist modules through an iterative message passing process. The agreement router iteratively computes the agreements among a set of input and outputs of all modules to allocate inputs to specific module. During the iterative routing, messages of modules are passed to each other, which improves the module selection process with consideration of both local interactions (between a single module and input) and global interactions involving multiple other modules. To validate our contributions, we conduct experiments on two problems: a toy min-max game and few-shot image classification task. Our experimental results demonstrate that SMN can generalize to new distributions and exhibit sample-efficient adaptation to new tasks. Furthermore, SMN can achieve a better adaptation capability when new modules are introduced after pre-training. Our code is available at https://github.com/hu-my/ScalableModularNetwork.
Minyang Hu, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
ICLR2
2024 UMFC: Unsupervised Multi-Domain Feature Calibration for Vision-Language Models
abstract
Pre-trained vision-language models (e.g., CLIP) have shown powerful zero-shot transfer capabilities. But they still struggle with domain shifts and typically require labeled data to adapt to downstream tasks, which could be costly. In this work, we aim to leverage unlabeled data that naturally spans multiple domains to enhance the transferability of vision-language models. Under this unsupervised multi-domain setting, we have identified inherent model bias within CLIP, notably in its visual and text encoders. Specifically, we observe that CLIP’s visual encoder tends to prioritize encoding domain over discriminative category information, meanwhile its text encoder exhibits a preference for domain-relevant classes. To mitigate this model bias, we propose a training-free and label-free feature calibration method, Unsupervised Multi-domain Feature Calibration (UMFC). UMFC estimates image-level biases from domain-specific features and text-level biases from the direction of domain transition. These biases are subsequently subtracted from original image and text features separately, to render them domain-invariant. We evaluate our method on multiple settings including transductive learning and test-time adaptation. Extensive experiments show that our method outperforms CLIP and performs on par with the state-of-the-arts that need additional annotations or optimization. Our code is available at https://github.com/GIT-LJc/UMFC.
Jiachen Liang, Ruibing Hou, Minyang Hu, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
NeurIPS4
2024 M$^3$GPT: An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation
abstract
This paper presents M$^3$GPT, an advanced $\textbf{M}$ultimodal, $\textbf{M}$ultitask framework for $\textbf{M}$otion comprehension and generation. M$^3$GPT operates on three fundamental principles. The first focuses on creating a unified representation space for various motion-relevant modalities. We employ discrete vector quantization for multimodal conditional signals, such as text, music and motion/dance, enabling seamless integration into a large language model (LLM) with a single vocabulary. The second involves modeling motion generation directly in the raw motion space. This strategy circumvents the information loss associated with a discrete tokenizer, resulting in more detailed and comprehensive motion generation. Third, M$^3$GPT learns to model the connections and synergies among various motion-relevant tasks. Text, the most familiar and well-understood modality for LLMs, is utilized as a bridge to establish connections between different motion tasks, facilitating mutual reinforcement. To our knowledge, M$^3$GPT is the first model capable of comprehending and generating motions based on multiple signals. Extensive experiments highlight M$^3$GPT's superior performance across various motion-relevant tasks and its powerful zero-shot generalization capabilities for extremely challenging tasks. Project page: \url{https://github.com/luomingshuang/M3GPT}.
Mingshuang Luo, Ruibing Hou, Hong Chang 0001, Zimo Liu, Shiguang Shan
NeurIPS4
2024 Triplet Adaptation Framework for Robust Semi-Supervised Learning
abstract
Semi-supervised learning (SSL) suffers from severe performance degradation when labeled and unlabeled data come from inconsistent and imbalanced distribution. Nonetheless, there is a lack of theoretical guidance regarding a remedy for this issue. To bridge the gap between theoretical insights and practical solutions, we embark to an analysis of generalization bound of classic SSL algorithms. This analysis reveals that distribution inconsistency between unlabeled and labeled data can cause a significant generalization error bound. Motivated by this theoretical insight, we present a Triplet Adaptation Framework (TAF) to reduce the distribution divergence and improve the generalization of SSL models. TAF comprises three adapters: Balanced Residual Adapter, aiming to map the class distribution of labeled and unlabeled data to a uniform distribution for reducing class distribution divergence; Representation Adapter, aiming to map the representation distribution of unlabeled data to labeled one for reducing representation distribution divergence; and Pseudo-Label Adapter, aiming to align the predicted pseudo-labels with the class distribution of unlabeled data, thereby preventing erroneous pseudo-labels from exacerbating representation divergence. These three adapters collaborate synergistically to reduce the generalization bound, ultimately achieving a more robust and generalizable SSL model. Extensive experiments across various robust SSL scenarios validate the efficacy of our method.
Ruibing Hou, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Incorporating texture and silhouette for video-based person re-identification
Shutao Bai, Hong Chang 0001, Bingpeng Ma
Pattern Recognit.2
2024 A Comprehensive Framework for Long-Tailed Learning via Pretraining and Normalization
abstract
Data in the visual world often present long-tailed distributions. However, learning high-quality representations and classifiers for imbalanced data is still challenging for data-driven deep learning models. In this work, we aim at improving the feature extractor and classifier for long-tailed recognition via contrastive pretraining and feature normalization, respectively. First, we carefully study the influence of contrastive pretraining under different conditions, showing that current self-supervised pretraining for long-tailed learning is still suboptimal in both performance and speed. We thus propose a new balanced contrastive loss and a fast contrastive initialization scheme to improve previous long-tailed pretraining. Second, based on the motivative analysis on the normalization for classifier, we propose a novel generalized normalization classifier that consists of generalized normalization and grouped learnable scaling. It outperforms traditional inner product classifier as well as cosine classifier. Both the two components proposed can improve recognition ability on tail classes without the expense of head classes. We finally build a unified framework that achieves competitive performance compared with state of the arts on several long-tailed recognition benchmarks and maintains high efficiency.
Nan Kang, Hong Chang 0001, Bingpeng Ma, Shiguang Shan
IEEE Trans. Neural Networks Learn. Syst.2
2023 Predictive Consistency Learning for Long-Tailed Recognition
Nan Kang, Hong Chang 0001, Bingpeng Ma, Shutao Bai, Shiguang Shan, Xilin Chen 0001
BMVC2
2023 Diversity-Measurable Anomaly Detection
abstract
Reconstruction-based anomaly detection models achieve their purpose by suppressing the generalization ability for anomaly. However, diverse normal patterns are consequently not well reconstructed as well. Although some efforts have been made to alleviate this problem by modeling sample diversity, they suffer from shortcut learning due to undesired transmission of abnormal information. In this paper, to better handle the tradeoff problem, we propose Diversity-Measurable Anomaly Detection (DMAD) framework to enhance reconstruction diversity while avoid the undesired generalization on anomalies. To this end, we design Pyramid Deformation Module (PDM), which models diverse normals and measures the severity of anomaly by estimating multi-scale deformation fields from reconstructed reference to original input. Integrated with an information compression module, PDM essentially decouples deformation from prototypical embedding and makes the final anomaly score more reliable. Experimental results on both surveillance videos and industrial images demonstrate the effectiveness of our method. In addition, DMAD works equally well in front of contaminated data and anomaly-like normal samples.
Wenrui Liu 0004, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
CVPR2
2023 Understanding Few-Shot Learning: Measuring Task Relatedness and Adaptation Difficulty via Attributes
abstract
Few-shot learning (FSL) aims to learn novel tasks with very few labeled samples by leveraging experience from \emph{related} training tasks. In this paper, we try to understand FSL by exploring two key questions: (1) How to quantify the relationship between \emph{ training} and \emph{novel} tasks? (2) How does the relationship affect the \emph{adaptation difficulty} on novel tasks for different models? To answer the first question, we propose Task Attribute Distance (TAD) as a metric to quantify the task relatedness via attributes. Unlike other metrics, TAD is independent of models, making it applicable to different FSL models. To address the second question, we utilize TAD metric to establish a theoretical connection between task relatedness and task adaptation difficulty. By deriving the generalization error bound on a novel task, we discover how TAD measures the adaptation difficulty on novel tasks for different models. To validate our theoretical results, we conduct experiments on three benchmarks. Our experimental results confirm that TAD metric effectively quantifies the task relatedness and reflects the adaptation difficulty on novel tasks for various FSL methods, even if some of them do not learn attributes explicitly or human-annotated attributes are not provided. Our code is available at \href{https://github.com/hu-my/TaskAttributeDistance}{https://github.com/hu-my/TaskAttributeDistance}.
Minyang Hu, Hong Chang 0001, Zong Guo, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
NeurIPS2
2023 Generalized Semi-Supervised Learning via Self-Supervised Feature Adaptation
abstract
Traditional semi-supervised learning (SSL) assumes that the feature distributions of labeled and unlabeled data are consistent which rarely holds in realistic scenarios. In this paper, we propose a novel SSL setting, where unlabeled samples are drawn from a mixed distribution that deviates from the feature distribution of labeled samples. Under this setting, previous SSL methods tend to predict wrong pseudo-labels with the model fitted on labeled data, resulting in noise accumulation. To tackle this issue, we propose \emph{Self-Supervised Feature Adaptation} (SSFA), a generic framework for improving SSL performance when labeled and unlabeled data come from different distributions. SSFA decouples the prediction of pseudo-labels from the current model to improve the quality of pseudo-labels. Particularly, SSFA incorporates a self-supervised task into the SSL framework and uses it to adapt the feature extractor of the model to the unlabeled data. In this way, the extracted features better fit the distribution of unlabeled data, thereby generating high-quality pseudo-labels. Extensive experiments show that our proposed SSFA is applicable to various pseudo-label-based SSL learners and significantly improves performance in labeled, unlabeled, and even unseen distributions.
Jiachen Liang, Ruibing Hou, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
NeurIPS3
2023 Dual Compensation Residual Networks for Class Imbalanced Learning
abstract
Learning generalizable representation and classifier for class-imbalanced data is challenging for data-driven deep models. Most studies attempt to re-balance the data distribution, which is prone to overfitting on tail classes and underfitting on head classes. In this work, we propose Dual Compensation Residual Networks to better fit both tail and head classes. First, we propose dual Feature Compensation Module (FCM) and Logit Compensation Module (LCM) to alleviate the overfitting issue. The design of these two modules is based on the observation: an important factor causing overfitting is that there is severe feature drift between training and test data on tail classes. In details, the test features of a tail category tend to drift towards feature cloud of multiple similar head categories. So FCM estimates a multi-mode feature drift direction for each tail category and compensate for it. Furthermore, LCM translates the deterministic feature drift vector estimated by FCM along intra-class variations, so as to cover a larger effective compensation space, thereby better fitting the test features. Second, we propose a Residual Balanced Multi-Proxies Classifier (RBMC) to alleviate the under-fitting issue. Motivated by the observation that re-balancing strategy hinders the classifier from learning sufficient head knowledge and eventually causes underfitting, RBMC utilizes uniform learning with a residual path to facilitate classifier learning. Comprehensive experiments on Long-tailed and Class-Incremental benchmarks validate the efficacy of our method.
Ruibing Hou, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Person Search by a Bi-Directional Task-Consistent Learning Model
abstract
Two-stage person search methods achieve the state-of-the-art performance by separate detection and re-ID stages, but neglect the consistency needs between these two stages. The re-ID stage needs more accurate query bounding boxes and fewer boxes of distractors; The detection stage needs the re-ID stage to have robustness against unavailable detection errors. In this paper, we introduce a novel Bi-directional Task-Consistent Learning (BTCL) person search framework, including a Target-Specific Detector (TSD) and a re-ID model with Dynamic Adaptive Learning Structure (DALS). For the former consistency need, we add a verification head for predicting the similarity scores between query and proposals in parallel with the existing heads for bounding box recognition. Thus, TSD generates accurate boxes for the query-like pedestrians, which are suitable for the re-ID stage. For the re-ID robustness need, DALS dynamically generates a large number of possible detection results in line with the real distribution. By training the re-ID model on data with different types of detection errors, DLAS improves the model robustness to detection inputs. Experimental results show our framework achieves state-of-the-art performance on two widely-used person search datasets.
Cheng Wang 0043, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Multim.3
2023 Refined Knowledge Transfer for Language-Based Person Search
abstract
This paper proposes a novel method, named Refined Knowledge Transfer (RKT), for language-based person search. Existing state-of-the-art methods do not deal with knowledge imbalance between image and text. In detail, textual identity knowledge is limited, but the image contains more identity knowledge. We propose Cross-Modal Knowledge Transfer (CMKT) to enhance textual identity knowledge by image to address this problem. Besides, multiple texts of one image include more identity knowledge than a single text. Thus, we propose Intra-Modal Knowledge Transfer (IMKT) to enhance textual identity knowledge by other texts. These two types of knowledge transfer will enhance the identity knowledge in text. Additionally, by considering that identity-irrelevant knowledge is transferred to text, we propose Knowledge Refiner (KR) to refine the knowledge in text. KR is capable of preserving identity knowledge and discarding identity-irrelevant knowledge. By combining CMKT, IMKT, and KR, RKT makes textual identity knowledge more salient. Extensive experiments show the state-of-the-art performance of RKT on the CUHK-PEDES and our proposed PRW-PEDES-CN datasets. In addition, the decent generalization ability of RKT is also validated on the Flickr30K, CUB, and Flowers datasets.
Ziqiang Wu, Bingpeng Ma, Hong Chang 0001, Shiguang Shan
IEEE Trans. Multim.3
2022 Salient-to-Broad Transition for Video Person Re-identification
abstract
Due to the limited utilization of temporal relations in video re-id, the frame-level attention regions of mainstream methods are partial and highly similar. To address this problem, we propose a Salient-to-Broad Module (SBM) to enlarge the attention regions gradually. Specifically, in SBM, while the previous frames have focused on the most salient regions, the later frames tend to focus on broader regions. In this way, the additional information in broad regions can supplement salient regions, incurring more powerful video-level representations. To further improve SBM, an Integration-and-Distribution Module (IDM) is introduced to enhance frame-level representations. IDM first integrates features from the entire feature space and then distributes the integrated features to each spatial location. SBM and IDM are mutually beneficial since they enhance the representations from video-level and frame-level, respectively. Extensive experiments on four prevalent benchmarks demonstrate the effectiveness and superiority of our method. The source code is available at https://github.com/baist/SINet.
Shutao Bai, Bingpeng Ma, Hong Chang 0001, Rui Huang 0001, Xilin Chen 0001
CVPR3
2022 Clothes-Changing Person Re-identification with RGB Modality Only
abstract
The key to address clothes-changing person re-identification (re-id) is to extract clothes-irrelevant features, e.g., face, hairstyle, body shape, and gait. Most current works mainly focus on modeling body shape from multi-modality information (e.g., silhouettes and sketches), but do not make full use of the clothes-irrelevant information in the original RGB images. In this paper, we propose a Clothes-based Adversarial Loss (CAL) to mine clothes-irrelevant features from the original RGB images by penalizing the predictive power of re-id model w.r.t. clothes. Extensive experiments demonstrate that using RGB images only, CAL outperforms all state-of-the-art methods on widely-used clothes-changing person re-id benchmarks. Besides, compared with images, videos contain richer appearance and additional temporal information, which can be used to model proper spatiotemporal patterns to assist clothes-changing re-id. Since there is no publicly available clothes-changing video re-id dataset, we contribute a new dataset named CCVID and show that there exists much room for improvement in modeling spatiotemporal information. The code and new dataset are available at: h t t$p$s: //github.com/guxinqian/Simple-CCReID.
Xinqian Gu, Hong Chang 0001, Bingpeng Ma, Shutao Bai, Shiguang Shan, Xilin Chen 0001
CVPR2
2022 Joint Feature Learning and Relation Modeling for Tracking: A One-Stream Framework
Botao Ye, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
ECCV (22)2
2022 Gradual Domain Adaptation with Sample Transferability Exploitation for Person Re-Identification
abstract
In this paper, we propose a novel gradual domain adaptation method with sample transferability exploitation to tackle the unsupervised domain adaptation (UDA) for person re-identification (re-id). Due to the direct but rough adaptation scheme, existing UDA for person re-id methods usually suffer from source domain-specific characteristics. To filter out the source domain-specific characteristics, motivated by the curriculum learning strategy, we conduct gradual domain adaptation by domain-level re-weighting with polynomial weight decay. Furthermore, we exploit sample transferability via maximum mean discrepancy based sample-level re-weighting strategy to diminish the domain gap. The sample transferability exploitation spotlights samples with higher importance to the adaptation process in each domain, hence enhance the adaptation performance. By combining the gradual domain adaptation with the sample transferability exploitation, our method achieves the state-of-the-art performance on transferring between two common person re-id datasets.
Zong Guo, Bingpeng Ma, Hong Chang 0001, Xilin Chen 0001
ICME3
2022 Learning Continuous Graph Structure with Bilevel Programming for Graph Neural Networks
abstract
Learning graph structure for graph neural networks (GNNs) is crucial to facilitate the GNN-based downstream learning tasks. It is challenging due to the non-differentiable discrete graph structure and lack of ground-truth. In this paper, we address these problems and propose a novel graph structure learning framework for GNNs. Firstly, we directly model the continuous graph structure with dual-normalization, which implicitly imposes sparse constraint and reduces the influence of noisy edges. Secondly, we formulate the whole training process as a bilevel programming problem, where the inner objective is to optimize the GNNs given learned graphs, while the outer objective is to optimize the graph structure to minimize the generalization error of downstream task. Moreover, for bilevel optimization, we propose an improved Neumann-IFT algorithm to obtain an approximate solution, which is more stable and accurate than existing optimization methods. Besides, it makes the bilevel optimization process memory-efficient and scalable to large graphs. Experiments on node classification and scene graph generation show that our method can outperform related methods, especially with noisy graphs.
Minyang Hu, Hong Chang 0001, Bingpeng Ma, Shiguang Shan
IJCAI2
2022 Optimal Positive Generation via Latent Transformation for Contrastive Learning
abstract
Contrastive learning, which learns to contrast positive with negative pairs of samples, has been popular for self-supervised visual representation learning. Although great effort has been made to design proper positive pairs through data augmentation, few works attempt to generate optimal positives for each instance. Inspired by semantic consistency and computational advantage in latent space of pretrained generative models, this paper proposes to learn instance-specific latent transformations to generate Contrastive Optimal Positives (COP-Gen) for self-supervised contrastive learning. Specifically, we formulate COP-Gen as an instance-specific latent space navigator which minimizes the mutual information between the generated positive pair subject to the semantic consistency constraint. Theoretically, the learned latent transformation creates optimal positives for contrastive learning, which removes as much nuisance information as possible while preserving the semantics. Empirically, using generated positives by COP-Gen consistently outperforms other latent transformation methods and even real-image-based methods in self-supervised contrastive learning.
Yinqi Li 0001, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
NeurIPS2
2022 Extending generalized unsupervised manifold alignment
Xiaoyi Yin, Zhen Cui 0001, Hong Chang 0001, Bingpeng Ma, Shiguang Shan
Sci. China Inf. Sci.3
2022 Feature Completion for Occluded Person Re-Identification
abstract
Person re-identification (reID) plays an important role in computer vision. However, existing methods suffer from performance degradation in occluded scenes. In this work, we propose an occlusion-robust block, Region Feature Completion (RFC), for occluded reID. Different from most previous works that discard the occluded regions, RFC block can recover the semantics of occluded regions in feature space. First, a Spatial RFC (SRFC) module is developed. SRFC exploits the long-range spatial contexts from non-occluded regions to predict the features of occluded regions. The unit-wise prediction task leads to an encoder/decoder architecture, where the region-encoder models the correlation between non-occluded and occluded region, and the region-decoder utilizes the spatial correlation to recover occluded region features. Second, we introduce Temporal RFC (TRFC) module which captures the long-term temporal contexts to refine the prediction of SRFC. RFC block is lightweight, end-to-end trainable and can be easily plugged into existing CNNs to form RFCnet. Extensive experiments are conducted on occluded and commonly holistic reID benchmarks. Our method significantly outperforms existing methods on the occlusion datasets, while remains top even superior performance on holistic datasets. The source code is available at https://github.com/blue-blue272/OccludedReID-RFCnet.
Ruibing Hou, Bingpeng Ma, Hong Chang 0001, Xinqian Gu, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 SANet: Statistic Attention Network for Video-Based Person Re-Identification
abstract
Capturing long-range dependencies during feature extraction is crucial for video-based person re-identification (re-id) since it would help to tackle many challenging problems such as occlusion and dramatic pose variation. Moreover, capturing subtle differences, such as bags and glasses, is indispensable to distinguish similar pedestrians. In this paper, we propose a novel and efficacious Statistic Attention (SA) block which can capture both the long-range dependencies and subtle differences. SA block leverages high-order statistics of feature maps, which contain both long-range and high-order information. By modeling relations with these statistics, SA block can explicitly capture long-range dependencies with less time complexity. In addition, high-order statistics usually concentrate on details of feature maps and can perceive the subtle differences between pedestrians. In this way, SA block is capable of discriminating pedestrians with subtle differences. Furthermore, this lightweight block can be conveniently inserted into existing deep neural networks at any depth to form Statistic Attention Network (SANet). To evaluate its performance, we conduct extensive experiments on two challenging video re-id datasets, showing that our SANet outperforms the state-of-the-art methods. Furthermore, to show the generalizability of SANet, we evaluate it on three image re-id datasets and two more general image classification datasets, including ImageNet. The source code is available athttp://vipl.ict.ac.cn/resources/codes/code/SANet_code.zip.
Shutao Bai, Bingpeng Ma, Hong Chang 0001, Rui Huang 0001, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 PRDP: Person Reidentification With Dirty and Poor Data
abstract
In this article, we propose a novel method to simultaneously solve the data problem of dirty quality and poor quantity for person reidentification (ReID). Dirty quality refers to the wrong labels in image annotations. Poor quantity means that some identities have very few images (FewIDs). Training with these mislabeled data or FewIDs with triplet loss will lead to low generalization performance. To solve the label error problem, we propose a weighted label correction based on cross-entropy (wLCCE) strategy. Specifically, according to the influence range of the wrong labels, we first classify the mislabeled images into point label error and set label error. Then, we propose a weighted triplet loss (WTL) to correct the two label errors, respectively. To alleviate the poor quantity issue, we propose a feature simulation based on autoencoder (FSAE) method to generate some virtual samples for FewID. For the authenticity of the simulated features, we transfer the difference pattern of identities with multiple images (MultIDs) to FewIDs by training an autoencoder (AE)-based simulator. In this way, the FewIDs obtain richer expressions to distinguish from other identities. By dealing with a dirty and poor data problem, we can learn more robust ReID models using the triplet loss. We conduct extensive experiments on two public person ReID datasets: 1) Market-1501 and 2) DukeMTMC-reID, to verify the effectiveness of our approach.
Furong Xu, Bingpeng Ma, Hong Chang 0001, Shiguang Shan
IEEE Trans. Cybern.3
2022 Motion Feature Aggregation for Video-Based Person Re-Identification
abstract
Most video-based person re-identification (re-id) methods only focus on appearance features but neglect motion features. In fact, motion features can help to distinguish the target persons that are hard to be identified only by appearance features. However, most existing temporal information modeling methods cannot extract motion features effectively or efficiently for v ideo-based re-id. In this paper, we propose a more efficient Motion Feature Aggregation (MFA) method to model and aggregate motion information in the feature map level for video-based re-id. The proposed MFA consists of (i) a coarse-grained motion learning module, which extracts coarse-grained motion features based on the position changes of body parts over time, and (ii) a fine-grained motion learning module, which extracts fine-grained motion features based on the appearance changes of body parts over time. These two modules can model motion information from different granularities and are complementary to each other. It is easy to combine the proposed method with existing network architectures for end-to-end training. Extensive experiments on four widely used datasets demonstrate that the motion features extracted by MFA are crucial complements to appearance features for video-based re-id, especially for the scenario with large appearance changes. Besides, the results on LS-VID, the current largest publicly available video-based re-id dataset, surpass the state-of-the-art methods by a large margin. The code is available at: https://github.com/guxinqian/Simple-ReID.
Xinqian Gu, Hong Chang 0001, Bingpeng Ma, Shiguang Shan
IEEE Trans. Image Process.2
2022 Interactive Regression and Classification for Dense Object Detector
abstract
In object detection, enhancing feature representation using localization information has been revealed as a crucial procedure to improve detection performance. However, the localization information (i.e., regression feature and regression offset) captured by the regression branch is still not well utilized. In this paper, we propose a simple but effective method called Interactive Regression and Classification (IRC) to better utilize localization information. Specifically, we propose Feature Aggregation Module (FAM) and Localization Attention Module (LAM) to leverage localization information to the classification branch during forward propagation. Furthermore, the classifier also guides the learning of the regression branch during backward propagation, to guarantee that the localization information is beneficial to both regression and classification. Thus, the regression and classification branches are learned in an interactive manner. Our method can be easily integrated into anchor-based and anchor-free object detectors without increasing computation cost. With our method, the performance is significantly improved on many popular dense object detectors, including RetinaNet, FCOS, ATSS, PAA, GFL, GFLV2, OTA, GA-RetinaNet, RepPoints, BorderDet and VFNet. Based on ResNet-101 backbone, IRC achieves 47.2% AP on COCO test-dev, surpassing the previous state-of-the-art PAA (44.8% AP), GFL (45.0% AP) and without sacrificing the efficiency both in training and inference. Moreover, our best model (Res2Net-101-DCN) can achieve a single-model single-scale AP of 51.4%.
Linmao Zhou, Hong Chang 0001, Bingpeng Ma, Shiguang Shan
IEEE Trans. Image Process.2
2021 BiCnet-TKS: Learning Efficient Spatial-Temporal Representation for Video Person Re-Identification
abstract
In this paper, we present an efficient spatial-temporal representation for video person re-identification (reID). Firstly, we propose a Bilateral Complementary Network (BiCnet) for spatial complementarity modeling. Specifically, BiCnet contains two branches. Detail Branch processes frames at original resolution to preserve the detailed visual clues, and Context Branch with a down-sampling strategy is employed to capture long-range contexts. On each branch, BiCnet appends multiple parallel and diverse attention modules to discover divergent body parts for consecutive frames, so as to obtain an integral characteristic of target identity. Furthermore, a Temporal Kernel Selection (TKS) block is designed to capture short-term as well as long-term temporal relations by an adaptive mode. TKS can be inserted into BiCnet at any depth to construct BiCnet-TKS for spatial-temporal modeling. Experimental results on multiple benchmarks show that BiCnet-TKS outperforms state-of-the-arts with about 50% less computations. The source code is available at https://github.com/blue-blue272/BiCnet-TKS.
Ruibing Hou, Hong Chang 0001, Bingpeng Ma, Rui Huang 0001, Shiguang Shan
CVPR2
2021 Enhancing Latent Features for Unsupervised Video Anomaly Detection
Linmao Zhou, Hong Chang 0001, Nan Kang, Xiangjun Zhao, Bingpeng Ma
PRCV (2)2
2021 Learning efficient text-to-image synthesis via interstage cross-sample similarity distillation
Fengling Mao, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
Sci. China Inf. Sci.3
2021 Cross-Modal Knowledge Adaptation for Language-Based Person Search
abstract
In this paper, we present a method named Cross-Modal Knowledge Adaptation (CMKA) for language-based person search. We argue that the image and text information are not equally important in determining a person's identity. In other words, image carries image-specific information such as lighting condition and background, while text contains more modal agnostic information that is more beneficial to cross-modal matching. Based on this consideration, we propose CMKA to adapt the knowledge of image to the knowledge of text. Specially, text-to-image guidance is obtained at different levels: individuals, lists, and classes. By combining these levels of knowledge adaptation, the image-specific information is suppressed, and the common space of image and text is better constructed. We conduct experiments on the CUHK-PEDES dataset. The experimental results show that the proposed CMKA outperforms the state-of-the-art methods.
Rui Huang 0001, Hong Chang 0001, Chuanqi Tan, Bingpeng Ma
IEEE Trans. Image Process.3
2021 Location Sensitive Network for Human Instance Segmentation
abstract
Location is an important distinguishing information for instance segmentation. In this paper, we propose a novel model, called Location Sensitive Network (LSNet), for human instance segmentation. LSNet integrates instance-specific location information into one-stage segmentation framework. Specifically, in the segmentation branch, Pose Attention Module (PAM) encodes the location information into the attention regions through coordinates encoding. Based on the location information provided by PAM, the segmentation branch is able to effectively distinguish instances in feature-level. Moreover, we propose a combination operation named Keypoints Sensitive Combination (KSCom) to utilize the location information from multiple sampling points. These sampling points construct the points representation for instances via human keypoints and random points. Human keypoints provide the spatial locations and semantic information of the instances, and random points expand the receptive fields. Based on the points representation for each instance, KSCom effectively reduces the mis-classified pixels. Our method is validated by the experiments on public datasets. LSNet-5 achieves 56.2 mAP at 18.5 FPS on COCOPersons. Besides, the proposed method is significantly superior to its peers in the case of severe occlusion.
Xiangzhou Zhang, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Image Process.3
2021 IAUnet: Global Context-Aware Feature Learning for Person Reidentification
abstract
Person reidentification (reID) by convolutional neural network (CNN)-based networks has achieved favorable performance in recent years. However, most of existing CNN-based methods do not take full advantage of spatial-temporal context modeling. In fact, the global spatial-temporal context can greatly clarify local distractions to enhance the target feature representation. To comprehensively leverage the spatial-temporal context information, in this work, we present a novel block, interaction-aggregation-update (IAU), for high-performance person reID. First, the spatial-temporal IAU (STIAU) module is introduced. STIAU jointly incorporates two types of contextual interactions into a CNN framework for target feature learning. Here, the spatial interactions learn to compute the contextual dependencies between different body parts of a single frame, while the temporal interactions are used to capture the contextual dependencies between the same body parts across all frames. Furthermore, a channel IAU (CIAU) module is designed to model the semantic contextual interactions between channel features to enhance the feature representation, especially for small-scale visual cues and body parts. Therefore, the IAU block enables the feature to incorporate the globally spatial, temporal, and channel context. It is lightweight, end-to-end trainable, and can be easily plugged into existing CNNs to form IAUnet. The experiments show that IAUnet performs favorably against state of the art on both image and video reID tasks and achieves compelling results on a general object categorization task. The source code is available at https://github.com/blue-blue272/ImgReID-IAnet.
Ruibing Hou, Bingpeng Ma, Hong Chang 0001, Xinqian Gu, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Neural Networks Learn. Syst.3
2020 TCTS: A Task-Consistent Two-Stage Framework for Person Search
abstract
The state of the art person search methods separate person search into detection and re-ID stages, but ignore the consistency between these two stages. The general person detector has no special attention on the query target; The re-ID model is trained on hand-drawn bounding boxes which are not available in person search. To address the consistency problem, we introduce a Task-Consist Two-Stage (TCTS) person search framework, includes an identity-guided query (IDGQ) detector and a Detection Results Adapted (DRA) re-ID model. In the detection stage, the IDGQ detector learns an auxiliary identity branch to compute query similarity scores for proposals. With consideration of the query similarity scores and foreground score, IDGQ produces query-like bounding boxes for the re-ID stage. In the re-ID stage, we predict identity labels of detected bounding boxes, and use these examples to construct a more practical mixed train set for the DRA model. Training on the mixed train set improves the robustness of the re-ID stage to inaccurate detection. We evaluate our method on two benchmark datasets, CUHK-SYSU and PRW. Our framework achieves 93.9% of mAP and 95.1% of rank1 accuracy on CUHK-SYSU, outperforming the previous state of the art methods.
Cheng Wang 0043, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
CVPR3
2020 Appearance-Preserving 3D Convolution for Video-Based Person Re-identification
Xinqian Gu, Hong Chang 0001, Bingpeng Ma, Xilin Chen 0001
ECCV (2)2
2020 Temporal Complementary Learning for Video Person Re-identification
Ruibing Hou, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
ECCV (25)2
2020 Dynamic R-CNN: Towards High Quality Object Detection via Dynamic Training
Hong Chang 0001, Bingpeng Ma, Naiyan Wang, Xilin Chen 0001
ECCV (15)2
2020 Part alignment network for vehicle re-identification
Bingpeng Ma, Hong Chang 0001
Neurocomputing3
2020 Visual concept conjunction learning with recurrent neural networks
Kongming Liang, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
Neurocomputing2
2020 Isosceles Constraints for Person Re-Identification
abstract
In the existing works of person re-identification (ReID), batch hard triplet loss has achieved great success. However, it only cares about the hardest samples within the batch. For any probe, there are massive mismatched samples (crucial samples) outside the batch which are closer than the matched samples. To reduce the disruptive influence of crucial samples, we propose a novel isosceles contraint for triplet. Theoretically, we show that if a matched pair has equal distance to any one of mismatched sample, the matched pair should be infinitely close. Motivated by this, the isosceles constraint is designed for the two mismatched pairs of each triplet, to restrict some matched pairs with equal distance to different mismatched samples. Meanwhile, to ensure that the distance of mismatched pairs are larger than the matched pairs, margin constraints are necessary. Minimizing the isosceles and margin constraints with respect to the feature extraction network makes the matched pairs closer and the mismatched pairs farther away than the matched ones. By this way, crucial samples are effectively reduced and the performance on ReID is improved greatly. Likewise, our isosceles contraint can be applied to quadruplet as well. Comprehensive experimental evaluations on Market-1501, DukeMTMC-reID and CUHK03 datasets demonstrate the advantages of our isosceles constraint over the related state-of-the-art approaches.
Furong Xu, Bingpeng Ma, Hong Chang 0001, Shiguang Shan
IEEE Trans. Image Process.3
2019 MS-GAN: Text to Image Synthesis with Attention-Modulated Generators and Similarity-aware Discriminators
Fengling Mao, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
BMVC3
2019 Cascade RetinaNet: Maintaining Consistency for Single-Stage Object Detection
Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
BMVC2
2019 Relation-aware Multiple Attention Siamese Networks for Robust Visual Tracking
Fangyi Zhang, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
BMVC3
2019 VRSTC: Occlusion-Free Video Person Re-Identification
abstract
Video person re-identification (re-ID) plays an important role in surveillance video analysis. However, the performance of video re-ID degenerates severely under partial occlusion. In this paper, we propose a novel network, called Spatio-Temporal Completion network (STCnet), to explicitly handle partial occlusion problem. Different from most previous works that discard the occluded frames, STCnet can recover the appearance of the occluded parts. For one thing, the spatial structure of a pedestrian frame can be used to predict the occluded body parts from the unoccluded body parts of this frame. For another, the temporal patterns of pedestrian sequence provide important clues to generate the contents of occluded parts. With the spatio-temporal information, STCnet can recover the appearance for the occluded parts, which could be leveraged with those unoccluded parts for more accurate video re-ID. By combining a re-ID network with STCnet, a video re-ID framework robust to partial occlusion (VRSTC) is proposed. Experiments on three challenging video re-ID databases demonstrate that the proposed approach outperforms the state-of-the-arts.
Ruibing Hou, Bingpeng Ma, Hong Chang 0001, Xinqian Gu, Shiguang Shan, Xilin Chen 0001
CVPR3
2019 Interaction-And-Aggregation Network for Person Re-Identification
abstract
Person re-identification (reID) benefits greatly from deep convolutional neural networks (CNNs) which learn robust feature embeddings. However, CNNs are inherently limited in modeling the large variations in person pose and scale due to their fixed geometric structures. In this paper, we propose a novel network structure, Interaction-and-Aggregation (IA), to enhance the feature representation capability of CNNs. Firstly, Spatial IA (SIA) module is introduced. It models the interdependencies between spatial features and then aggregates the correlated features corresponding to the same body parts. Unlike CNNs which extract features from fixed rectangle regions, SIA can adaptively determine the receptive fields according to the input person pose and scale. Secondly, we introduce Channel IA (CIA) module which selectively aggregates channel features to enhance the feature representation, especially for small-scale visual cues. Further, IA network can be constructed by inserting IA blocks into CNNs at any depth. We validate the effectiveness of our model for person reID by demonstrating its superiority over state-of-the-art methods on three benchmark datasets.
Ruibing Hou, Bingpeng Ma, Hong Chang 0001, Xinqian Gu, Shiguang Shan, Xilin Chen 0001
CVPR3
2019 Video Prediction with Bidirectional Constraint Network
abstract
Future frame prediction in videos is promising avenue for unsupervised video representation learning. However video prediction has the huge solution space since the high-dimensionality and inherent uncertainty of the future video frames. Existing approaches impose weak constraints on the predictions, which results in motion confusion. To alleviate this problem, we propose a novel model named Bidirectional Constraint Network (BCnet). BCnet consists of forward prediction module and backward prediction module. The forward prediction module learns to predict the future sequence from the present sequence, while the backward prediction module learns to invert the task. The closed loop of the two modules allows that the backward prediction module generates informative feedback signals. The feedback signals clamp down the solution space of forward prediction module. Therefore, our approach can effectively alleviate the motion confusion. We further evaluate BCnet by fine-tuning it for a supervised learning problem: human action recognition on the UCF-101 dataset. We show that the representation help improve classification accuracy. Extensive experiments on several challenging public datasets show that our approach significantly outperforms state-of-the-art approaches, which demonstrates the effectiveness and generalization ability of our approach.
Ruibing Hou, Hong Chang 0001, Bingpeng Ma, Xilin Chen 0001
FG2
2019 Temporal Knowledge Propagation for Image-to-Video Person Re-Identification
abstract
In many scenarios of Person Re-identification (Re-ID), the gallery set consists of lots of surveillance videos and the query is just an image, thus Re-ID has to be conducted between image and videos. Compared with videos, still person images lack temporal information. Besides, the information asymmetry between image and video features increases the difficulty in matching images and videos. To solve this problem, we propose a novel Temporal Knowledge Propagation (TKP) method which propagates the temporal knowledge learned by the video representation network to the image representation network. Specifically, given the input videos, we enforce the image representation network to fit the outputs of video representation network in a shared feature space. With back propagation, temporal knowledge can be transferred to enhance the image features and the information asymmetry problem can be alleviated. With additional classification and integrated triplet losses, our model can learn expressive and discriminative image and video features for image-to-video re-identification. Extensive experiments demonstrate the effectiveness of our method and the overall results on two widely used datasets surpass the state-of-the-art methods by a large margin.
Xinqian Gu, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
ICCV3
2019 Attribute-Aware Pedestrian Image Editing
Xiaoyi Yin, Xinqian Gu, Hong Chang 0001, Bingpeng Ma, Xilin Chen 0001
ICIG (1)3
2019 Cross Attention Network for Few-shot Classification
abstract
Few-shot classification aims to recognize unlabeled samples from unseen classes given only few labeled samples. The unseen classes and low-data problem make few-shot classification very challenging. Many existing approaches extracted features from labeled and unlabeled samples independently, as a result, the features are not discriminative enough. In this work, we propose a novel Cross Attention Network to address the challenging problems in few-shot classification. Firstly, Cross Attention Module is introduced to deal with the problem of unseen classes. The module generates cross attention maps for each pair of class feature and query sample feature so as to highlight the target object regions, making the extracted feature more discriminative. Secondly, a transductive inference algorithm is proposed to alleviate the low-data problem, which iteratively utilizes the unlabeled query set to augment the support set, thereby making the class features more representative. Extensive experiments on two benchmarks show our method is a simple, effective and computationally efficient framework and outperforms the state-of-the-arts.
Ruibing Hou, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
NeurIPS2
2019 Unifying Visual Attribute Learning with Object Recognition in a Multiplicative Framework
abstract
Attributes are mid-level semantic properties of objects. Recent research has shown that visual attributes can benefit many typical learning problems in computer vision community. However, attribute learning is still a challenging problem as the attributes may not always be predictable directly from input images and the variation of visual attributes is sometimes large across categories. In this paper, we propose a unified multiplicative framework for attribute learning, which tackles the key problems. Specifically, images and category information are jointly projected into a shared feature space, where the latent factors are disentangled and multiplied to fulfil attribute prediction. The resulting attribute classifier is category-specific instead of being shared by all categories. Moreover, our model can leverage auxiliary data to enhance the predictive ability of attribute classifiers, which can reduce the effort of instance-level attribute annotation to some extent. By integrated into an existing deep learning framework, our model can both accurately predict attributes and learn efficient image representations. Experimental results show that our method achieves superior performance on both instance-level and category-level attribute prediction. For zero-shot learning based on visual attributes and human-object interaction recognition, our method can improve the state-of-the-art performance on several widely used datasets.
Kongming Liang, Hong Chang 0001, Bingpeng Ma, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 Visual Relationship Detection With Deep Structural Ranking
abstract
Visual relationship detection aims to describe the interactions between pairs of objects. Different from individual object learning tasks, the number of possible relationships are much larger, which makes it hard to explore only based on the visual appearance of objects. In addition, due to the limited human effort, the annotations for visual relationships are usually incomplete which increases the difficulty of model training and evaluation. In this paper, we propose a novel framework, called Deep Structural Ranking, for visual relationship detection. To complement the representation ability of visual appearance, we integrate multiple cues for predicting the relationships contained in an input image. Moreover, we design a new ranking objective function by enforcing the annotated relationships to have higher relevance scores. Unlike previous works, our proposed method can both facilitate the co-occurrence of relationships and mitigate the incompleteness problem. Experimental results show that our proposed method outperforms the state-of-the-art on the two widely used datasets. We also demonstrate its superiority in detecting zero-shot relationships.
Kongming Liang, Yuhong Guo, Hong Chang 0001, Xilin Chen 0001
AAAI3
2018 Style Transfer with Adversarial Learning for Cross-Dataset Person Re-identification
Furong Xu, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
ACCV (6)3
2018 Continuity-Discrimination Convolutional Neural Network for Visual Object Tracking
abstract
This paper proposes a novel model, named Continuity-Discrimination Convolutional Neural Network (CD-CNN), for visual object tracking. Existing state-of-the-art tracking methods do not deal with temporal relationship in video sequences, which leads to imperfect feature representations. To address this problem, CD-CNN models temporal appearance continuity based on the idea of temporal slowness. Mathematically, we prove that, by introducing temporal appearance continuity into tracking, the upper bound of target appearance representation error can be sufficiently small with high probability. Further, in order to alleviate inaccurate target localization and drifting, we propose a novel notion, object-centroid, to characterize not only objectness but also the relative position of the target within a given patch. Both temporal appearance continuity and object-centroid are jointly learned during offline training and then transferred for online tracking. We evaluate our tracker through extensive experiments on two challenging benchmarks and show its competitive tracking performance compared with state-of-the-art trackers.
Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
ICME3
2018 Parametric local multiview hamming distance metric learning
Deming Zhai, Xianming Liu 0005, Hong Chang 0001, Yi Zhen, Xilin Chen 0001, Maozu Guo 0001, Wen Gao 0001
Pattern Recognit.3
2017 Siamese recurrent architecture for visual tracking
abstract
Treating visual tracking as a matching problem, siamese architecture has drawn increasing interest recently. In this paper, we propose a novel siamese recurrent architecture that can enhance the similarity matching by leveraging contextual information. Specifically, the multi-directional Recurrent Neural Network (RNN) is employed to memorize the long-range contextual dependencies of object parts and learn the self-structure information of the object. We test the proposed method on a challenging benchmark, and it gain promising results compared with the existing tracking algorithms.
Xiaqing Xu, Bingpeng Ma, Hong Chang 0001, Xilin Chen 0001
ICIP3
2017 Incomplete Attribute Learning with auxiliary labels
abstract
Visual attribute learning is a fundamental and challenging problem for image understanding. Considering the huge semantic space of attributes, it is economically impossible to annotate all their presence or absence for a natural image via crowd-sourcing. In this paper, we tackle the incompleteness nature of visual attributes by introducing auxiliary labels into a novel transductive learning framework. By jointly predicting the attributes from the input images and modeling the relationship of attributes and auxiliary labels, the missing attributes can be recovered effectively. In addition, the proposed model can be solved efficiently in an alternative way by optimizing quadratic programming problems and updating parameters in closed-form solutions. Moreover, we propose and investigate different methods for acquiring auxiliary labels. We conduct experiments on three widely used attribute prediction datasets. The experimental results show that our proposed method can achieve the state-of-the-art performance with access to partially observed attribute annotations.
Kongming Liang, Yuhong Guo, Hong Chang 0001, Xilin Chen 0001
IJCAI3
2017 Special issue on selected and extended papers from the 2015 International Conference on Intelligence Science and Big Data Engineering (IScIDE 2015)
Shiguang Shan, Deng Cai 0001, Cheng Deng 0002, Hong Chang 0001
Neurocomputing4
2016 Deep Second-Order Siamese Network for Pedestrian Re-identification
Xuesong Deng, Bingpeng Ma, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
ACCV (2)3
2016 Attribute Conjunction Learning with Recurrent Neural Network
Kongming Liang, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
ECML/PKDD (1)2
2015 Online Visual Tracking via Coupled Object-Context Dictionary
abstract
(a) (b) Figure 1: Illustration of constructing the coupled dictionaries. The red rectangles in (a) and (b) represent the target bounding boxes. The green squares in (a), which are generated by sliding windows outside the target bounding box, correspond to basis patches involved in the background dictionary N. The blue squares in (a) are generated inside the target bounding box in a similar way, which constitute the noisy target dictionary P.
Mingquan Ye, Hong Chang 0001, Xilin Chen 0001
BMVC2
2015 A Unified Multiplicative Framework for Attribute Learning
abstract
Attributes are mid-level semantic properties of objects. Recent research has shown that visual attributes can benefit many traditional learning problems in computer vision community. However, attribute learning is still a challenging problem as the attributes may not always be predictable directly from input images and the variation of visual attributes is sometimes large across categories. In this paper, we propose a unified multiplicative framework for attribute learning, which tackles the key problems. Specifically, images and category information are jointly projected into a shared feature space, where the latent factors are disentangled and multiplied for attribute prediction. The resulting attribute classifier is category-specific instead of being shared by all categories. Moreover, our method can leverage auxiliary data to enhance the predictive ability of attribute classifiers, reducing the effort of instance-level attribute annotation to some extent. Experimental results show that our method achieves superior performance on both instance-level and category-level attribute prediction. For zero-shot learning based on attributes, our method significantly improves the state-of-the-art performance on AwA dataset and achieves comparable performance on CUB dataset.
Kongming Liang, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
ICCV2
2015 Instance-specific canonical correlation analysis
Deming Zhai, Yu Zhang 0006, Dit-Yan Yeung, Hong Chang 0001, Xilin Chen 0001, Wen Gao 0001
Neurocomputing4
2014 Representation Learning with Smooth Autoencoder
Kongming Liang, Hong Chang 0001, Zhen Cui 0001, Shiguang Shan, Xilin Chen 0001
ACCV (2)2
2014 Stacked Progressive Auto-Encoders (SPAE) for Face Recognition Across Poses
abstract
Identifying subjects with variations caused by poses is one of the most challenging tasks in face recognition, since the difference in appearances caused by poses may be even larger than the difference due to identity. Inspired by the observation that pose variations change non-linearly but smoothly, we propose to learn pose-robust features by modeling the complex non-linear transform from the non-frontal face images to frontal ones through a deep network in a progressive way, termed as stacked progressive auto-encoders (SPAE). Specifically, each shallow progressive auto-encoder of the stacked network is designed to map the face images at large poses to a virtual view at smaller ones, and meanwhile keep those images already at smaller poses unchanged. Then, stacking multiple these shallow auto-encoders can convert non-frontal face images to frontal ones progressively, which means the pose variations are narrowed down to zero step by step. As a result, the outputs of the topmost hidden layers of the stacked network contain very small pose variations, which can be used as the pose-robust features for face recognition. An additional attractiveness of the proposed method is that no pose estimation is needed for the test images. The proposed method is evaluated on two datasets with pose variations, i.e., MultiPIE and FERET datasets, and the experimental results demonstrate the superiority of our method to the existing works, especially to those 2D ones.
Meina Kan, Shiguang Shan, Hong Chang 0001, Xilin Chen 0001
CVPR3
2014 Deep Network Cascade for Image Super-resolution
Zhen Cui 0001, Hong Chang 0001, Shiguang Shan, Bineng Zhong 0001, Xilin Chen 0001
ECCV (5)2
2014 Modeling Video Dynamics with Deep Dynencoder
Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
ECCV (4)2
2014 Generalized Unsupervised Manifold Alignment
Zhen Cui 0001, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
NIPS2
2014 Joint sparse representation for video-based face recognition
Zhen Cui 0001, Hong Chang 0001, Shiguang Shan, Bingpeng Ma, Xilin Chen 0001
Neurocomputing2
2013 Progressive Image Restoration through Hybrid Graph Laplacian Regularization
abstract
In this paper, we propose a unified framework to perform progressive image restoration based on hybrid graph Laplacian regularized regression. We first construct a multi-scale representation of the target image by Laplacian pyramid, then progressively recover the degraded image in the scale space from coarse to fine so that the sharp edges and texture can be eventually recovered. On one hand, within each scale, a graph Laplacian regularization model represented by implicit kernel is learned which simultaneously minimizes the least square error on the measured samples and preserves the geometrical structure of the image data space by exploring non-local self-similarity. In this procedure, the intrinsic manifold structure is considered by using both measured and unmeasured samples. On the other hand, between two scales, the proposed model is extended to the parametric manner through explicit kernel mapping to model the inter-scale correlation, in which the local structure regularity is learned and propagated from coarser to finer scales. Experimental results on benchmark test images demonstrate that the proposed method achieves better performance than state-of-the-art image restoration algorithms.
Deming Zhai, Xianming Liu 0005, Debin Zhao, Hong Chang 0001, Wen Gao 0001
DCC4
2013 Temporally multiple dynamic textures synthesis using piecewise linear dynamic systems
abstract
Real-world nonlinear dynamic textures (DTs) usually consist of temporally multiple linear DTs which cannot be correctly modeled by previous works. In this paper, we propose piecewise linear dynamic systems (PLDS) to model temporally multiple DTs. PLDS simultaneously decides the temporal segmentation, models each DT segment with an LDS and the whole DT by switching between the LDS'. Experimental results verify that PLDS can capture the stochastic and dynamic nature of temporally multiple DTs and it synthesizes nonlinear DTs without decay or divergence. An EM-like algorithm iterating between sequence division and LDS' fitting is adopted to learn the model parameters.
Hong Chang 0001, Xilin Chen 0001
ICIP2
2013 Instance-specific canonical correlation analysis for pose alignment
abstract
Canonical correlation analysis (CCA) based methods achieve great success for pose alignment. However, CCA has limitations as a linear and global algorithm. Although some variants have been proposed to overcome the limitations, neither of them achieves locality and nonlinearity at the same time. In this paper, we propose a novel algorithm called Instance-Specific Canonical Correlation Analysis (ISCCA), which approximates the nonlinear data by computing the instance specific projections along the smooth curve of the manifold. Based on the framework of least squares regression, CCA is extended to the instance-specific case which obtains a set of locally-linear smooth but globally-nonlinear transformations. The optimization problem is proved to be convex and could be solved efficiently by alternating optimization. And the globally optimal solutions could be achieved with theoretical guarantee. Experimental results for pose alignment demonstrate the effectiveness of our proposed method.
Deming Zhai, Hong Chang 0001, Xilin Chen 0001, Wen Gao 0001
ICIP2
2013 Parametric Local Multimodal Hashing for Cross-View Similarity Search
Deming Zhai, Hong Chang 0001, Yi Zhen, Xianming Liu 0005, Xilin Chen 0001, Wen Gao 0001
IJCAI2
2013 Strip Features for Fast Object Detection
abstract
This paper presents a set of effective and efficient features, namely strip features, for detecting objects in real-scene images. Although shapes of a specific class usually have large intraclass variance, some basic local shape elements are relatively stable. Based on this observation, we propose a set of strip features to describe the appearances of those shape elements. Strip features capture object shapes with edgelike and ridgelike strip patterns, which significantly enrich the efficient features such as Haar-like and edgelet features. The proposed features can be efficiently calculated via two kinds of approaches. Moreover, the proposed features can be extended to a perturbed version (namely, perturbed strip features) to alleviate the misalignment caused by deformations. We utilize strip features for object detection under an improved boosting framework, which adopts a complexity-aware criterion to balance the discriminability and efficiency for feature selection. We evaluate the proposed approach for object detection on the public data sets, and the experimental results show the effectiveness and efficiency of the proposed approach.
Hong Chang 0001, Luhong Liang, Shiguang Shan, Xilin Chen 0001
IEEE Trans. Cybern.2
2012 Multi-layer Spectral Clustering for Video Segmentation
Xiaofei Di, Hong Chang 0001, Xilin Chen 0001
ACCV (2)2
2012 Grouping Active Contour Fragments for Object Recognition
Songlin Song, Hong Chang 0001, Xilin Chen 0001
ACCV (1)3
2012 Boosted translation-tolerable classifiers for fast object detection
Luhong Liang, Hong Chang 0001, Cherkeng Heng, Shiguang Shan, Xilin Chen 0001
Image Vis. Comput.3
2012 Multiview Metric Learning with Global Consistency and Local Smoothness
abstract
In many real-world applications, the same object may have different observations (or descriptions) from multiview observation spaces, which are highly related but sometimes look different from each other. Conventional metric-learning methods achieve satisfactory performance on distance metric computation of data in a single-view observation space, but fail to handle well data sampled from multiview observation spaces, especially those with highly nonlinear structure. To tackle this problem, we propose a new method calledMultiview Metric Learning with Global consistency and Local smoothness(MVML-GL) under a semisupervised learning setting, which jointly considers global consistency and local smoothness. The basic idea is to reveal the shared latent feature space of the multiview observations by embodying global consistency constraints and preserving local geometric structures. Specifically, this framework is composed of two main steps. In the first step, we seek a global consistent shared latent feature space, which not only preserves the local geometric structure in each space but also makes those labeled corresponding instances as close as possible. In the second step, the explicit mapping functions between the input spaces and the shared latent space are learned via regularized locally linear regression. Furthermore, these two steps both can be solved by convex optimizations in closed form. Experimental results with application to manifold alignment on real-world datasets of pose and facial expression demonstrate the effectiveness of the proposed method.
Deming Zhai, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001
ACM Trans. Intell. Syst. Technol.2
2011 A unified framework for locating and recognizing human actions
abstract
In this paper, we present a pose based approach for locating and recognizing human actions in videos. In our method, human poses are detected and represented based on deformable part model. To our knowledge, this is the first work on exploring the effectiveness of deformable part models in combining human detection and pose estimation into action recognition. Comparing with previous methods, ours have three main advantages. First, our method does not rely on any assumption on video preprocessing quality, such as satisfactory foreground segmentation or reliable tracking; Second, we propose a novel compact representation for human pose which works together with human detection and can well represent the spatial and temporal structures inside an action; Third, with human detection taken into consideration in our framework, our method has the ability to locate and recognize multiple actions in the same scene. Experiments on benchmark datasets and recorded cluttered videos verified the efficacy of our method.
Yuelei Xie, Hong Chang 0001, Zhe Li 0008, Luhong Liang, Xilin Chen 0001, Debin Zhao
CVPR2
2011 A novel coarse-to-fine hair segmentation method
abstract
Segmenting hair regions from human images facilitates many tasks like hair synthesis and hair style trends forecast. However, hair segmentation is quite challenging due to hair/background confusion and large hair pattern diversity. To address these problems to some extent, this paper proposes a novel coarse-to-fine hair segmentation method. In our approach, firstly, the recently proposed “Active Segmentation with Fixation” (ASF) is used to coarsely define an enclosed candidate region with high-recall (but possibly low-precision) of hair pixels and exclude considerable part of the backgrounds which are easily confused with hair. Then Graph Cuts (GC) method is applied to the candidate regions to remove additional false positives by incorporating hair-specific information. Specifically, Bayesian method is employed to select some reliable hair and background regions (seeds) among the ones over-segmented by Mean Shift. SVM classifier is then learnt online from these seeds and explored to predict hair/background likelihood probability, which is subsequently fed into GC algorithm. The novelty of the proposed approach lies in three folds: 1) an elaborate design of hair segmentation framework, which utilizes ASF to reduce the candidate hair regions and adopts GC to achieve more accurate hair region contours; 2) the region-based strategy for seed selection; 3) the exploration of the discriminative method, SVM, to predict the probability of each pixel belonging to hair and background regions. Extensive experimental results demonstrate the approach outperforms recently proposed methods.
Xiujuan Chai, Hongming Zhang 0011, Hong Chang 0001, Wei Zeng 0006, Shiguang Shan
FG4
2010 Manifold Alignment via Corresponding Projections
abstract
In this paper, we propose a novel manifold alignment method by learning the underlying common manifold with supervision of corresponding data pairs from different observation sets. Different from the previous algorithms of semi-supervised manifold alignment, our method learns the explicit corresponding projections from each original observation space to the common embedding space everywhere. Benefiting from this property, our method could process new test data directly rather than re-alignment. Furthermore, our approach doesn’t have any assumption on the data structures, thus it could handle more complex cases and get better results compared with previous work. In the proposed algorithm, manifold alignment is formulated as a minimization problem with proper constraints, which could be solved in an analytical manner with closed-form solution. Experimental results on pose manifold alignment of different objects and faces demonstrate the effectiveness of our proposed method.
Deming Zhai, Bo Li 0086, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001
BMVC3
2010 Boosted Sigma Set for Pedestrian Detection
abstract
This paper presents a new method to detect pedestrian in still image using Sigma sets as image region descriptors in the boosting framework. Sigma set encodes second order statistics of an image region implicitly in the form of a point set. Compared with the covariance matrix, the traditional second order statistics based region descriptor, which requires computationally demanding operations based on Riemannian manifold, Sigma set preserves similar robustness and discriminative power more efficiently because the classification on Sigma sets can be directly performed in vector space. Experimental results on the INRIA and the Daimler Chrysler pedestrian datasets show the effectiveness and efficiency of the proposed method.
Xiaopeng Hong, Hong Chang 0001, Xilin Chen 0001, Wen Gao 0001
ICPR2
2010 Low-Resolution Face Recognition via Coupled Locality Preserving Mappings
abstract
Practical face recognition systems are sometimes confronted with low-resolution face images. Traditional two-step methods solve this problem through employing super-resolution (SR). However, these methods usually have limited performance because the target of SR is not absolutely consistent with that of face recognition. Moreover, time-consuming sophisticated SR algorithms are not suitable for real-time applications. To avoid these limitations, we propose a novel approach for LR face recognition without any SR preprocessing. Our method based on coupled mappings (CMs), projects the face images with different resolutions into a unified feature space which favors the task of classification. These CMs are learned through optimizing the objective function to minimize the difference between the correspondences (i.e., low-resolution image and its high-resolution counterpart). Inspired by locality preserving methods for dimensionality reduction, we introduce a penalty weighting matrix into our objective function. Our method significantly improves the recognition performance. Finally, we conduct experiments on publicly available databases to verify the efficacy of our algorithm.
Bo Li 0086, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
IEEE Signal Process. Lett.2
2010 Sigma Set Based Implicit Online Learning for Object Tracking
abstract
This letter presents a novel object tracking approach within the Bayesian inference framework through implicit online learning. In our approach, the target is represented by multiple patches, each of which is encoded by a powerful and efficient region descriptor called Sigma set. To model each target patch, we propose to utilize the online one-class support vector machine algorithm, named Implicit online Learning with Kernels Model (ILKM). ILKM is simple, efficient, and capable of learning a robust online target predictor in the presence of appearance changes. Responses of ILKMs related to multiple target patches are fused by an arbitrator with an inference of possible partial occlusions, to make the decision and trigger the model update. Experimental results demonstrate that the proposed tracking approach is effective and efficient in ever-changing and cluttered scenes.
Xiaopeng Hong, Hong Chang 0001, Shiguang Shan, Bineng Zhong 0001, Xilin Chen 0001, Wen Gao 0001
IEEE Signal Process. Lett.2
2009 Coupled Metric Learning for Face Recognition with Degraded Images
Bo Li 0086, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
ACML2
2009 Semi-Supervised Discriminant Analysis via Spectral Transduction
abstract
Linear Discriminant Analysis (LDA) is a popular method for dimensionality reduction and classification. In real-world applications when there is no sufficient labeled data, LDA suffers from serious performance drop or even fails to work. In this paper, we propose a novel method called Spectral Transduction Semi-Supervised Discriminant Analysis (STSDA), which can alleviate such problem by utilizing both labeled and unlabeled data. Our method takes into consideration both label augmenting and local structure preserving. First, we formulate label transduction with labeled and unlabeled data as a constrained convex optimization problem and solve it efficiently with a closed-form solution by using orthogonal projector matrices. Then, unlabeled data with reliable class estimations are selected with a balanced strategy to augment the original labeled data set. At last, LDA with manifold regularization is performed. Experimental results on face recognition demonstrate the effectiveness of our proposed method.
Deming Zhai, Hong Chang 0001, Bo Li 0086, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001
BMVC2
2009 Sigma Set: A small second order statistical region descriptor
abstract
Given an image region of pixels, second order statistics can be used to construct a descriptor for object representation. One example is the covariance matrix descriptor, which shows high discriminative power and good robustness in many computer vision applications. However, operations for the covariance matrix on Riemannian manifolds are usually computationally demanding. This paper proposes a novel second order statistics based region descriptor, named “Sigma Set”, in the form of a small set of vectors, which can be uniquely constructed through Cholesky decomposition on the covariance matrix. Sigma Set is of low dimension, powerful and robust. Moreover, compared with the covariance matrix, Sigma Set is not only more efficient in distance evaluation and average calculation, but also easier to be enriched with first order statistics. Experimental results in texture classification and object tracking verify the effectiveness and efficiency of this novel object descriptor.
Xiaopeng Hong, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001
CVPR2
2009 Locality preserving constraints for super-resolution with neighbor embedding
abstract
In this paper, we revisit the manifold assumption which has been widely adopted in the learning-based image super-resolution. The assumption states that point-pairs from the high-resolution manifold share the local geometry with the corresponding low-resolution manifold. However, the assumption does not hold always, since the one-to-multiple mapping from LR to HR makes neighbor reconstruction ambiguous and results in blurring and artifacts. To minimize the ambiguous, we utilize Locality Preserving Constraints (LPC) to avoid confusions through emphasizing the consistency of localities on both manifolds explicitly. The LPC are combined with a MAP framework, and realized by building a set of cell-pairs on the coupled manifolds. Finally, we propose an energy minimization algorithm for the MAP with LPC which can reconstruct high quality images compared with previous methods. Experimental results show the effectiveness of our method.
Bo Li 0086, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
ICIP2
2009 Aligning Coupled Manifolds for Face Hallucination
abstract
Many learning-based super-resolution methods are based on the manifold assumption, which claims that point-pairs from the low-resolution representation manifold (LRM) and the corresponding high-resolution representation manifold (HRM) possess similar local geometry. However, the manifold assumption does not hold well on the original coupled manifolds (i.e., LRM and HRM) due to the nonisometric one-to-multiple mappings from low-resolution (LR) image patches to high-resolution (HR) ones. To overcome this limitation, we propose a solution from the perspective of manifold alignment. In this context, we perform alignment by learning two explicit mappings which project the point-pairs from the original coupled manifolds into the embeddings of the common manifold (CM). For the task of SR reconstruction, we treat HRM as target manifold and employ the manifold regularization to guarantee that the local geometry of CM is more consistent with that of HRM than LRM is. After alignment, we carry out the SR reconstruction based on neighbor embedding between the new couple of the CM and the target HRM. Besides, we extend our method by aligning the multiple coupled subsets instead of the whole coupled manifolds to address the issue of the global nonlinearity. Experimental results on face image super-resolution verify the effectiveness of our method.
Bo Li 0086, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001
IEEE Signal Process. Lett.2
2008 Parzen Discriminant Analysis
abstract
In this paper, we propose a non-parametric Discriminant Analysis method (no assumption on the distributions of classes), called Parzen Discriminant Analysis (PDA). Through a deep investigation on the non-parametric density estimation, we find that minimizing/maximizing the distances between each data sample and its nearby similar/dissimilar samples is equivalent to minimizing an upper bound of the Bayesian error rate. Based on this theoretical analysis, we define our criterion as maximizing the average local dissimilarity scatter with respect to a fixed average local similarity scatter. All local scatters are calculated in fixed size local regions, resembling the idea of Parzen estimation. Experiments in UCI machine learning database show that our method impressively outperforms other related neighbor based non-parametric methods.
Youhan Fang, Shiguang Shan, Hong Chang 0001, Xilin Chen 0001, Wen Gao 0001
ICPR3
2008 Hallucinating facial images and features
abstract
In facial image analysis, image resolution is an important factor which has great influence on the performance of face recognition systems. As for low-resolution face recognition problem, traditional methods usually carry out super-resolution firstly before passing the super-resolved image to a face recognition system. In this paper, we propose a new method which predicts high-resolution images and the corresponding features simultaneously. More specifically, we propose “feature hallucination” to project facial images with low-resolution into an expected feature space. As a result, the proposed method does not require super-resolution as an explicit preprocessing step. In addition, we explore a constrained hallucination that considers the local consistency in the image grid. In our method, we use the index of local visual primitives [5] as features and a block-based histogram distance to measure the similarity for the face recognition. Experimental results on FERET face database verify that the proposed method can improve both visual quality and recognition rate for low-resolution facial images.
Bo Li 0086, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001
ICPR2
2008 A Scalable Kernel-Based Semisupervised Metric Learning Algorithm with Out-of-Sample Generalization Ability
abstract
In recent years, metric learning in the semisupervised setting has aroused a lot of research interest. One type of semisupervised metric learning utilizes supervisory information in the form of pairwise similarity or dissimilarity constraints. However, most methods proposed so far are either limited to linear metric learning or unable to scale well with the data set size. In this letter, we propose a nonlinear metric learning method based on the kernel approach. By applying low-rank approximation to the kernel matrix, our method can handle significantly larger data sets. Moreover, our low-rank approximation scheme can naturally lead to out-of-sample generalization. Experiments performed on both artificial and real-world data show very promising results.
Dit-Yan Yeung, Hong Chang 0001, Guang Dai
Neural Comput.2
2008 Robust path-based spectral clustering
Hong Chang 0001, Dit-Yan Yeung
Pattern Recognit.1
2007 Locally Smooth Metric Learning with Application to Image Retrieval
abstract
In this paper, we propose a novel metric learning method based on regularized moving least squares. Unlike most previous metric learning methods which learn a global Mahalanobis distance, we define locally smooth metrics using local affine transformations which are more flexible. The data set after metric learning can preserve the original topological structures. Moreover, our method is fairly efficient and may be used as a preprocessing step for various subsequent learning tasks, including classification, clustering, and nonlinear dimensionality reduction. In particular, we demonstrate that our method can boost the performance of content-based image retrieval (CBIR) tasks. Experimental results provide empirical evidence for the effectiveness of our approach.
Dit-Yan Yeung, Hong Chang 0001
ICCV2
2007 A Scalable Kernel-Based Algorithm for Semi-Supervised Metric Learning
Dit-Yan Yeung, Hong Chang 0001, Guang Dai
IJCAI2
2007 Kernel-based distance metric learning for content-based image retrieval
Hong Chang 0001, Dit-Yan Yeung
Image Vis. Comput.1
2007 Learning the kernel matrix by maximizing a KFD-based class separability criterion
Dit-Yan Yeung, Hong Chang 0001, Guang Dai
Pattern Recognit.2
2007 A Kernel Approach for Semisupervised Metric Learning
abstract
While distance function learning for supervised learning tasks has a long history, extending it to learning tasks with weaker supervisory information has only been studied recently. In particular, some methods have been proposed for semisupervised metric learning based on pairwise similarity or dissimilarity information. In this paper, we propose a kernel approach for semisupervised metric learning and present in detail two special cases of this kernel approach. The metric learning problem is thus formulated as an optimization problem for kernel learning. An attractive property of the optimization problem is that it is convex and, hence, has no local optima. While a closed-form solution exists for the first special case, the second case is solved using an iterative majorization procedure to estimate the optimal solution asymptotically. Experimental results based on both synthetic and real-world data show that this new kernel approach is promising for nonlinear metric learning.
Dit-Yan Yeung, Hong Chang 0001
IEEE Trans. Neural Networks2
2006 A Manifold Regularization Approach to Calibration Reduction for Sensor-Network Based Tracking
Jeffrey Junfeng Pan, Qiang Yang 0001, Hong Chang 0001, Dit-Yan Yeung
AAAI3
2006 Graph Laplacian Kernels for Object Classification from a Single Example
abstract
Classification with only one labeled example per class is a challenging problem in machine learning and pattern recognition. While there have been some attempts to address this problem in the context of specific applications, very little work has been done so far on the problem under more general object classification settings. In this paper, we propose a graph-based approach to the problem. Based on a robust path-based similarity measure proposed recently, we construct a weighted graph using the robust path-based similarities as edge weights. A kernel matrix, called graph Laplacian kernel, is then defined based on the graph Laplacian. With the kernel matrix, in principle any kernel-based classifier can be used for classification. In particular, we demonstrate the use of a kernel nearest neighbor classifier on some synthetic data and real-world image sets, showing that our method can successfully solve some difficult classification tasks with only very few labeled examples.
Hong Chang 0001, Dit-Yan Yeung
CVPR (2)1
2006 Extending Kernel Fisher Discriminant Analysis with the Weighted Pairwise Chernoff Criterion
Guang Dai, Dit-Yan Yeung, Hong Chang 0001
ECCV (4)3
2006 Robust locally linear embedding
Hong Chang 0001, Dit-Yan Yeung
Pattern Recognit.1
2006 Locally linear metric adaptation with application to semi-supervised clustering and image retrieval
Hong Chang 0001, Dit-Yan Yeung
Pattern Recognit.1
2006 Relaxational metric adaptation and its application to semi-supervised clustering and content-based image retrieval
Hong Chang 0001, Dit-Yan Yeung, William Kwok-Wai Cheung
Pattern Recognit.1
2006 Extending the relevant component analysis algorithm for metric learning using both positive and negative equivalence constraints
Dit-Yan Yeung, Hong Chang 0001
Pattern Recognit.2
2005 Stepwise Metric Adaptation Based on Semi-Supervised Learning for Boosting Image Retrieval Performance
abstract
For a specific set of features chosen for representing images, the performance of a content-based image retrieval (CBIR) system depends critically on the similarity measure used. Based on a recently proposed semisupervised metric learning method called locally linear metric adaptation (LLMA), we propose in this paper a stepwise LLMA algorithm for boosting the retrieval performance of CBIR systems by incorporating relevance feedback from users collected over multiple query sessions. Unlike most existing metric learning methods which learn a global Mahalanobis metric, the transformation performed by LLMA is more general in that it is linear locally but nonlinear globally. Moreover, the efficiency problem is well addressed by the stepwise LLMA algorithm. We also report experimental results performed on a real-world color image database to demonstrate the effectiveness of our method. 1
Hong Chang 0001, Dit-Yan Yeung
BMVC1
2005 Robust Path-Based Spectral Clustering with Application to Image Segmentation
abstract
Spectral clustering and path-based clustering are two recently developed clustering approaches that have delivered impressive results in a number of challenging clustering tasks. However, they are not robust enough against noise and outliers in the data. In this paper, based on M-estimation from robust statistics, we develop a robust path-based spectral clustering method by defining a robust path-based similarity measure for spectral clustering. Our method is significantly more robust than spectral clustering and path-based clustering. We have performed experiments based on both synthetic and real-world data, comparing our method with some other methods. In particular, color images from the Berkeley segmentation dataset and benchmark are used in the image segmentation experiments. Experimental results show that our method consistently outperforms other methods due to its higher robustness.
Hong Chang 0001, Dit-Yan Yeung
ICCV1
2004 Super-Resolution through Neighbor Embedding
Hong Chang 0001, Dit-Yan Yeung, Yimin Xiong
CVPR (1)1
2004 Locally linear metric adaptation for semi-supervised clustering
abstract
Many supervised and unsupervised learning algorithms are very sensitive to the choice of an appropriate distance metric. While classification tasks can make use of class label information for metric learning, such information is generally unavailable in conventional clustering tasks. Some recent research sought to address a variant of the conventional clustering problem called semi-supervised clustering, which performs clustering in the presence of some background knowledge or supervisory information expressed as pairwise similarity or dissimilarity constraints. However, existing metric learning methods for semi-supervised clustering mostly perform global metric learning through a linear transformation. In this paper, we propose a new metric learning method which performs nonlinear transformation globally but linear transformation locally. In particular, we formulate the learning problem as an optimization problem and present two methods for solving it. Through some toy data sets, we show empirically that our locally linear metric adaptation (LLMA) method can handle some difficult cases that cannot be handled satisfactorily by previous methods. We also demonstrate the effectiveness of our method on some real data sets.
Hong Chang 0001, Dit-Yan Yeung
ICML1