EDBT 2026 Demo / reviewers in the wild / expert
Guosen Xie
dblp:150/4121 · also Guo-Sen Xie
· DBLP profile ↗
92ranked-venue papers
16as first author
65since 2021 · last 2026
0000-0002-5487-9845ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 9 first-author · 42 since 2021Artificial intelligence and machine learning · 52 · 13 first-author · 35 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReAlign: Text-to-Motion Generation via Step-Aware Reward-Guided AlignmentabstractText-to-motion generation, which synthesizes 3D human motions from text inputs, holds immense potential for applications in gaming, film, and robotics. Recently, diffusion-based methods have been shown to generate more diversity and realistic motion. However, there exists a misalignment between text and motion distributions in diffusion models, which leads to semantically inconsistent or low-quality motions. To address this limitation, we propose Reward-guided sampling Alignment (ReAlign), comprising a step-aware reward model to assess alignment quality during the denoising sampling and a reward-guided strategy that directs the diffusion process toward an optimally aligned distribution. This reward model integrates step-aware tokens and combines a text-aligned module for semantic consistency and a motion-aligned module for realism, refining noisy motions at each timestep to balance probability density and alignment. Extensive experiments of both motion generation and retrieval tasks demonstrate that our approach significantly improves text-motion alignment and motion quality compared to existing state-of-the-art methods. Wanjiang Weng, Xiaofeng Tan 0001, Junbo Wang 0003, Guosen Xie, Pan Zhou 0002, Hongsong Wang 0001 |
AAAI | 4 |
| 2026 | Mutually Causal Semantic Distillation Network for Zero-Shot Learning
Shiming Chen 0002, Shuhuang Chen, Guosen Xie, Xinge You |
Int. J. Comput. Vis. | 3 |
| 2026 | Foundation Model for Skeleton-Based Human Action UnderstandingabstractHuman action understanding serves as a foundational pillar in the field of intelligent motion perception.Skeletons serve as a modality- and device-agnostic representation for human modeling, and skeleton-based action understanding has potential applications in humanoid robot control and interaction. However, existing works often lack the scalability and generalization required to handle diverse action understanding tasks. There is no skeleton foundation model that can be adapted to a wide range of action understanding tasks. This paper presents a Unified Skeleton-based Dense Representation Learning (USDRL) framework, which serves as a foundational model for skeleton-based human action understanding. USDRL consists of a Transformer-based Dense Spatio-Temporal Encoder (DSTE), Multi-Grained Feature Decorrelation (MG-FD), and Multi-Perspective Consistency Training (MPCT). The DSTE module adopts two parallel streams to learn temporal dynamic and spatial structure features. The MG-FD module collaboratively performs feature decorrelation across temporal, spatial, and instance domains to reduce dimensional redundancy and enhance information extraction. The MPCT module employs both multi-view and multi-modal self-supervised consistency training. The former enhances the learning of high-level semantics and mitigates the impact of low-level discrepancies, while the latter effectively facilitates the learning of informative multimodal features. We perform extensive experiments on 25 benchmarks across across 9 skeleton-based action understanding tasks, covering coarse prediction, dense prediction, and transferred prediction. Our approach significantly outperforms the current state-of-the-art methods. We hope that this work would broaden the scope of research in skeleton-based action understanding and encourage more attention to dense prediction tasks. Hongsong Wang 0001, Wanjiang Weng, Junbo Wang 0003, Fang Zhao 0006, Guosen Xie, Xin Geng 0001, Liang Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Hierarchical Vision-Language Interaction for Facial Action Unit DetectionabstractFacial Action Unit (AU) detection seeks to recognize subtle facial muscle activations as defined by the Facial Action Coding System (FACS). A primary challenge w.r.t AU detection is the effective learning of discriminative and generalizable AU representations under conditions of limited annotated data. To address this, we propose a Hierarchical Vision-language Inter action for AU Understanding (HiVA) method, which leverages textual AU descriptions as semantic priors to guide and enhance AU detection. Specifically, HiVA employs a large language model to generate diverse and contextually rich AU descriptions to strengthen language-based representation learning. To capture both fine-grained and holistic vision-language associations, HiVA introduces an AU-aware dynamic graph module that facilitates the learning of AU-specific visual representations. These features are further integrated within a hierarchical cross-modal atten tion architecture comprising two complementary mechanisms: Disentangled Dual Cross-Attention (DDCA), which establishes fine-grained, AU-specific interactions between visual and textual features, and Contextual Dual Cross-Attention (CDCA), which models global inter-AU dependencies. This collaborative, cross modal learning paradigm enables HiVA to leverage multi-grained vision-based AU features in conjunction with refined language based AU details, culminating in robust and semantically en riched AU detection capabilities. Extensive experiments show that HiVA consistently surpasses state-of-the-art approaches. Besides, qualitative analyses reveal that HiVA produces semantically meaningful activation patterns, highlighting its efficacy in learning robust and interpretable cross-modal correspondences for comprehensive facial behavior analysis. Yong Li 0032, Yizhe Zhang 0001, Tianyi Zhang 0013, Muyun Jiang, Guosen Xie, Cuntai Guan |
IEEE Trans. Affect. Comput. | 7 |
| 2026 | Dual Adversarial Perturbations for Zero-Shot LearningabstractIn Zero-Shot Learning (ZSL), embedding-based methods learn a visual–semantic mapping that leverages the attribute knowledge of seen classes to predict the attributes of unseen classes, enabling knowledge transfer from seen to unseen classes. However, distributional discrepancies between seen and unseen classes introduce an inherent domain shift, and inter-class variations cause the same attribute to be expressed differently across categories. As a result, models trained on seen classes often struggle to accurately recognize attributes in unseen classes, limiting their generalization ability. To address these challenges, we propose DAPZSL, a dual adversarial perturbation framework that enhances the robustness of visual–semantic mappings through Feature-Level Adversarial Perturbation (FAP) and improves the model’s generalization ability via Weight-Level Adversarial Perturbation (WAP). Specifically, FAP generates semantically perturbed adversarial samples at the feature level, and incorporating these samples during training encourages the model to learn more robust visual–semantic mappings that are resilient to semantic variations, which improves attribute recognition on unseen classes. Meanwhile, WAP introduces adversarial perturbations into the model’s weight space, promoting a flatter loss landscape that alleviates overfitting to seen classes and enhances generalization. Extensive experiments on multiple benchmark datasets—including AWA2, SUN, and CUB—demonstrate that DAPZSL significantly outperforms existing ZSL models. Shiming Chen 0002, Guosen Xie, Chaojian Yu, Xinhua You, Qinmu Peng, Xinge You |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Through the Looking Glass: A Dual Perspective on Weakly Supervised Few-Shot SegmentationabstractMeta-learning aims to uniformly sample homologous support-query pairs, characterized by the same categories and similar attributes, and extract useful inductive biases through identical network architectures. However, this identical network design results in over-semantic homogenization. To address this, we propose a novel homologous but heterogeneous network. By treating support-query pairs as dual perspectives, we introduce heterogeneous visual aggregation (HA) modules to enhance complementarity while preserving semantic commonality. To further reduce semantic noise and amplify the uniqueness of heterogeneous semantics, we design a heterogeneous transport (HT) module. Finally, we propose heterogeneous CLIP (HC) textual information to enhance the generalization capability of multimodal models. In the weakly-supervised few-shot semantic segmentation (WFSS) task, with only 1/24 of the parameters of existing state-of-the-art models, TLG achieves a 13.2% improvement on Pascal- $5{^{\text {i}}}$ and a 7.9% improvement on COCO- $20{^{\text {i}}}$ . To the best of our knowledge, TLG is also the first weakly-supervised (image-level) model that outperforms fully supervised (pixel-level) models under the same backbone architectures. The code is available at https://github.com/jarch-ma/TLG. Jiaqi Ma 0006, Guosen Xie, Fang Zhao 0006, Zechao Li |
IEEE Trans. Image Process. | 2 |
| 2026 | Attack-Augmented Mixing-Contrastive Skeletal Representation LearningabstractContrastive learning facilitates the acquisition of informative skeleton representations for unsupervised action recognition by leveraging effective positive and negative sample pairs. However, most existing methods construct these pairs through weak or strong data augmentations, which typically rely on random appearance alterations of skeletons. While such augmentations are somewhat effective, they introduce semantic variations only indirectly and face two inherent limitations. First, simply modifying the appearance of skeletons often fails to reflect meaningful semantic variations. Second, random perturbations can unintentionally blur the boundary between positive and negative pairs, weakening the contrastive objective. To address these challenges, we propose an attack-driven augmentation framework that explicitly introduces semantic-level perturbations. This approach facilitates the generation of hard positives while guiding the model to mine more informative hard negatives. Building on this idea, we present Attack-Augmented Mixing-Contrastive Skeletal Representation Learning (A2MC), a novel framework that focuses on contrasting hard positive and hard negative samples for more robust representation learning. Within A2MC, we design an Attack-Augmentation (Att-Aug) module that integrates both targeted (attack-based) and untargeted (augmentation-based) perturbations to generate informative hard positive samples. In parallel, we propose the Positive-Negative Mixer (PNM), which blends hard positive and negative features to synthesize challenging hard negatives. These are then used to update a mixed memory bank for more effective contrastive learning. Comprehensive evaluations across three public benchmarks demonstrate that our approach, termed A2MC, achieves performance on par with or exceeding existing state-of-the-art methods. Binqian Xu, Xiangbo Shu, Jiachao Zhang, Rui Yan 0010, Guosen Xie |
IEEE Trans. Image Process. | 5 |
| 2026 | Location Matters: Frequency-Spatial Dual-Space Adaptation for Cross-Domain Few-Shot SegmentationabstractCurrent cross-domain few-shot semantic segmentation (CD-FSS) methods tend to overlook a fundamental yet domain-agnostic prior: the spatial correspondence between support and query images driven by the task itself. Unlike semantic similarity, this spatial correlation arises from the consistent structural layout of foreground objects across domains. To exploit this structural prior, we propose a novel frequency-spatial dual space adaptation (FDSA) framework, to learn domain-invariant structures and task-specific priors by jointly suppressing domain-specific redundancy in frequency domain and reinforcing geometric priors in spatial domain. Specifically, FDSA consists of two sequential modules, i.e., the frequency structural adapter (FSA) and the spatial geometry adapter (SGA). FSA performs image modulation in the frequency domain by emphasizing low-frequency foreground semantics and attenuating high-frequency noise, thus maintaining structural integrity of these input images. By contrast, SGA leverages handcrafted local descriptors to extract keypoints from both support and query images, generating Gaussian-based geometric priors that highlight desirable aligned regions. Additionally, we introduce spatial-guided SAM refinement (SSR) to extend our spatial geometric prior into the Segment Anything Model (SAM). SSR generates a soft Gaussian point prompt centered on the coarse mask, enabling SAM to refine segmentation masks without manual intervention. This integration effectively bridges task-specific localization with high-quality segmentation. Extensive experiments on four standard CD-FSS benchmarks demonstrate that our method achieves new state-of-the-art performance. Code is available at https://github.com/CVL-hub/FDSA.git. Guolei Sun, Yong Li 0032, Hongsong Wang 0001, Xiangbo Shu, Guosen Xie |
IEEE Trans. Image Process. | 6 |
| 2025 | SAM-Aware Graph Prompt Reasoning Network for Cross-Domain Few-Shot SegmentationabstractThe primary challenge of cross-domain few-shot segmentation (CD-FSS) is the domain disparity between the training and inference phases, which can exist in either the input data or the target classes. Previous models struggle to learn feature representations that generalize to various unknown domains from limited training domain samples. In contrast, the large-scale visual model SAM, pre-trained on tens of millions of images from various domains and classes, possesses excellent generalizability. In this work, we propose a SAM-aware graph prompt reasoning network (GPRN) that fully leverages SAM to guide CD-FSS feature representation learning and improve prediction accuracy. Specifically, we propose a SAM-aware prompt initialization module (SPI) to transform the masks generated by SAM into visual prompts enriched with high-level semantic information. Since SAM tends to divide an object into many sub-regions, this may lead to visual prompts representing the same semantic object having inconsistent or fragmented features. We further propose a graph prompt reasoning (GPR) module that constructs a graph among visual prompts to reason about their interrelationships and enable each visual prompt to aggregate information from similar prompts, thus achieving global semantic consistency. Subsequently, each visual prompt embeds its semantic information into the corresponding mask region to assist in feature representation learning. To refine the segmentation mask during testing, we also design a non-parameter adaptive point selection module (APS) to select representative point prompts from query predictions and feed them back to SAM to refine inaccurate segmentation results. Experiments on four standard CD-FSS datasets demonstrate that our method establishes new state-of-the-art results. Shi-Feng Peng, Guolei Sun, Yong Li 0032, Hongsong Wang 0001, Guosen Xie |
AAAI | 5 |
| 2025 | 3D-aware Select, Expand, and Squeeze Token for Aerial Action RecognitionabstractAerial Action Recognition (AAR) in videos captured by Unmanned Aerial Vehicles (UAVs) plays a vital role in numerous applications. However, current methods related to traditional action recognition primarily cater to fixed or near cameras, and rarely consider the movement disturbance of UAVs, including their varying attitudes and positions. Those characteristics of aerial videos bring moving objects in small regions compared to broad backgrounds and relative movement to the motion of objects, which reflect more sparse and disturbed semantic information for AAR. To address these issues, we present a novel framework, dubbed 3D-Tok, to Select, Expand, and Squeeze original visual tokens for obtaining compact yet diverse semantic-enhanced tokens. In particular, we present a 3D-token selector (3TS) to select complex yet diverse tokens in three channels, capturing the semantic awareness of moving objects in comparatively small regions. Additionally, to get rid of disturbed semantic information caused by the UAV flight, we present an Expand-Squeeze Converter (ESC) to adaptively expand and squeeze the 3D-selected tokens constrained by contrastive loss, thereby suppressing the semantic-irrelevant information and reinforce semantic-relevant information via the interpolation converting. By involving the token selecting, expanding, and squeezing into an all-in-one framework, 3D-Tok shows significant improvements on the UAV-Human dataset(↑9.5%), RoCoG-v2 dataset (↑23.5%), and Drone-Action dataset (↑5.7%). Luying Peng, Xiangbo Shu, Yazhou Yao, Guosen Xie |
AAAI | 4 |
| 2025 | Kernel-Aware Graph Prompt Learning for Few-Shot Anomaly DetectionabstractFew-shot anomaly detection (FSAD) aims to detect unseen anomaly regions with the guidance of very few normal support images from the same class. Existing FSAD methods usually find anomalies by directly designing complex text prompts to align them with visual features under the prevailing large vision-language model paradigm. However, these methods, almost always, neglect intrinsic contextual information in visual features, e.g., the interaction relationships between different vision layers, which is an important clue for detecting anomalies comprehensively. To this end, we propose a kernel-aware graph prompt learning framework, termed as KAG-prompt, by reasoning the cross-layer relations among visual features for FSAD. Specifically, a kernel-aware hierarchical graph is built by taking the different layer features focusing on anomalous regions of different sizes as nodes, meanwhile, the relationships between arbitrary pairs of nodes stand for the edges of the graph. By message passing over this graph, KAG-prompt can capture cross-layer contextual information, thus leading to more accurate anomaly prediction. Moreover, to integrate the information of multiple important anomaly signals in the prediction map, we propose a novel image-level scoring method based on multi-level information fusion. Extensive experiments on MVTecAD and VisA datasets show that KAG-prompt achieves state-of-the-art FSAD results for image-level/pixel-level anomaly detection. Fenfang Tao, Guosen Xie, Fang Zhao 0006, Xiangbo Shu |
AAAI | 2 |
| 2025 | USDRL: Unified Skeleton-Based Dense Representation Learning with Multi-Grained Feature DecorrelationabstractContrastive learning has achieved great success in skeleton-based representation learning recently. However, the prevailing methods are predominantly negative-based, necessitating additional momentum encoder and memory bank to get negative samples, which increases the difficulty of model training. Furthermore, these methods primarily concentrate on learning a global representation for recognition and retrieval tasks, while overlooking the rich and detailed local representations that are crucial for dense prediction tasks. To alleviate these issues, we introduce a Unified Skeleton-based Dense Representation Learning framework based on feature decorrelation, called USDRL, which employs feature decorrelation across temporal, spatial, and instance domains in a multi-grained manner to reduce redundancy among dimensions of the representations to maximize information extraction from features. Additionally, we design a Dense Spatio-Temporal Encoder (DSTE) to capture fine-grained action representations effectively, thereby enhancing the performance of dense prediction tasks. Comprehensive experiments, conducted on the benchmarks NTU-60, NTU-120, PKU-MMD I, and PKU-MMD II, across diverse downstream tasks including action recognition, action retrieval, and action detection, conclusively demonstrate that our approach significantly outperforms the current state-of-the-art (SOTA) approaches. Wanjiang Weng, Hongsong Wang 0001, Junbo Wang 0003, Guosen Xie |
AAAI | 5 |
| 2025 | Graph Interaction Prompt Network for Few-Shot Medical Image Anomaly DetectionabstractFew-shot medical image anomaly detection aims to detect and locate anomalies with limited data, playing a crucial role in clinical practice. In recent years, the large pre-trained vision-language model CLIP has demonstrated impressive performance across a variety of few-shot and zero-shot downstream tasks. However, CLIP mainly focuses on aligning text and images, emphasizing the semantics of global foreground objects rather than distinguishing local subtle normal or abnormal areas in the images. To address this challenge, we propose the Graph Interaction Prompt Network (GIPN), a framework that leverages graph interaction and text-prompt learning for precise anomaly detection in medical images. Specifically, we introduce a graph interaction prompt module that enables cross-layer interactions among visual features in a latent graph space, guiding the model to focus on challenging anomalous regions and enhancing feature representations. Additionally, we develop a dual-stream fusion strategy, which merges the hierarchical graph interaction features with the original vision-language features, better capturing critical cues for anomaly prediction under the guidance of text prompts. Extensive experiments on three medical anomaly datasets demonstrate that GIPN outperforms the current state-of-the-art few-shot medical anomaly detection approaches. Our code is available at https://github.com/CVL-hub/GIPN. Fenfang Tao, Tian-Zhu Xiang, Fang Zhao 0006, Guosen Xie |
BIBM | 5 |
| 2025 | Region-Aware Compositional Context Prompting for Zero-Shot Anomaly DetectionabstractZero-shot anomaly detection (ZSAD) aims to identify anomalies of unseen classes without requiring samples from those classes. Existing methods typically rely on pre-trained visual language models, such as CLIP, to detect anomalies by designing or learning generic text prompts and computing similarities with image features, which often fail to address the complexity and novelty of anomaly patterns, especially when the target domain exhibits significant differences from the source domain. To address the problems, we propose Region-aware Compositional Context Prompting (ReCo-CoP) for ZSAD, which dynamically generates contextual prompts by integrating both global and local visual information. Specifically, we introduce a Compositional Context Prompting (CCP) module that incorporates global visual features into the context through a set of basis vectors shared among images, and a Regional Context Prompting (RCP) module that optimizes the context based on image patch features, thereby enhancing the model’s ability to perceive local abnormal regions. Additionally, we combine dynamically generated prompts with static generic prompts to prevent the model from losing the essential general knowledge. Extensive experiments on 12 datasets from industrial and medical domains demonstrate the superior zero-shot detection performance of our model. The code is available at https://github.com/WenDongyp/ReCoCoP Guanglei Chu, Guosen Xie, Caifeng Shan, Fang Zhao 0006 |
ECAI | 4 |
| 2025 | Part in Part Embedding Network for Zero-Shot LearningabstractZero-shot learning (ZSL) seeks to utilize semantic information from seen classes encountered during training to effectively recognize unseen classes during testing. When dealing with fine-grained images, capturing local features heavily influences the accuracy of semantic descriptions. Meanwhile, local features are represented at various scales across different layers of a neural network, making it hard to capture their local details fully. To address these challenges, we propose a novel part in part embedding network, termed PPEN. Specifically, PPEN consists of two key modules: the cross-layer aggregation (CLA) module and the part in part attention (PIPA) module. The CLA module is designed to fuse and preserve features from multiple layers of the network, thereby maintaining the richness of information across different scales. Further, the PIPA module focuses on identifying local features that are most pertinent to the class semantic vectors, enhancing the alignment between visual features and semantic descriptions. We evaluate our approach on three ZSL benchmarks, i.e., CUB, SUN, and AWA2, and demonstrate the superiority and competitiveness of our proposed approach. Code is available at https://github.com/zhou834177226/PPENet. Zhexian Zhou, Liang Xiao 0001, Guosen Xie |
ICASSP | 3 |
| 2025 | Tensor-Aggregated LoRA in Federated Fine-Tuning
Binqian Xu, Xiangbo Shu, Jiachao Zhang, Yazhou Yao, Guosen Xie, Jinhui Tang 0001 |
ICCV | 6 |
| 2025 | A Conditional Probability Framework for Compositional Zero-Shot LearningabstractCompositional Zero-Shot Learning (CZSL) aims to recognize unseen combinations of known objects and attributes by leveraging knowledge from previously seen compositions. Traditional approaches primarily focus on disentangling attributes and objects, treating them as independent entities during learning. However, this assumption overlooks the semantic constraints and contextual dependencies inside a composition. For example, certain attributes naturally pair with specific objects (e.g., "striped" applies to "zebra" or "shirts" but not "sky" or "water"), while the same attribute can manifest differently depending on context (e.g., "young" in "young tree" vs. "young dog"). Thus, capturing attribute-object interdependence remains a fundamental yet long-ignored challenge in CZSL. In this paper, we adopt a Conditional Probability Framework (CPF) to explicitly model attribute-object dependencies. We decompose the probability of a composition into two components: the likelihood of an object and the conditional likelihood of its attribute. To enhance object feature learning, we incorporate textual descriptors to highlight semantically relevant image regions. These enhanced object features then guide attribute learning through a cross-attention mechanism, ensuring better contextual alignment. By jointly optimizing object likelihood and conditional attribute likelihood, our method effectively captures compositional dependencies and generalizes well to unseen compositions. Extensive experiments on multiple CZSL benchmarks demonstrate the superiority of our approach. Code is available at here. Peng Wu 0014, Qiuxia Lai, Hao Fang 0010, Guosen Xie, Yilong Yin, Xiankai Lu, Wenguan Wang |
ICCV | 4 |
| 2025 | Reliable and Diverse Hierarchical Adapter for Zero-shot Video ClassificationabstractAdapting pre-trained vision-language models to downstream tasks has emerged as a novel paradigm for zero-shot learning. Existing test-time adaptation (TTA) methods such as TPT attempt to fine-tune visual or textual representations to accommodate downstream tasks but still require expensive optimization costs. To this end, Training-free Dynamic Adapter (TDA) maintains a cache containing visual features for each category in a parameter-free manner and measures sample confidence based on prediction entropy of test samples. Inspired by TDA, this work aims to develop the first training-free adapter for zero-shot video classification. Capturing the intrinsic temporal relationships within video data to construct and maintain the video cache is key to extending TDA to the video domain. In this work, we propose a reliable and diverse Hierarchical Adapter for zero-shot video classification, which consists of Frame-level Cache Refiner and Video-level Cache Updater. Before each video sample enters the corresponding cache, it needs to be refined at frame level based on prediction entropy and temporal probability difference. Due to the limited capacity of the cache, we update the cache during inference based on the principle of diversity. Experiments on four popular video classification benchmarks demonstrate the effectiveness of Hierarchical Adapter. The code is available at https://github.com/Gwxer/Hierarchical-Adapter. Wenxuan Ge, Rui Yan 0010, Hongyu Qu, Guosen Xie, Xiangbo Shu |
IJCAI | 5 |
| 2025 | AnomalyControl: Highly-Aligned Anomalous Image Generation with Controlled Diffusion ModelabstractIn industrial scenarios, diverse anomalous images are difficult to acquire, significantly limiting the performance of industrial anomaly detection methods. Automatically generating anomalous images for anomaly detection has the potential to solve the above problem. However, existing anomaly generation models are still not satisfactory regarding the authenticity and controllability of anomaly generation. In this paper, we propose a controlled anomaly generation model named AnomalyControl to generate realistic anomalous images aligned highly with both text prompts and anomaly masks. First, we introduce a CLIP-guided anomaly prompt generator that leverages a CLIP text encoder to find anomaly text prompts most aligned with real anomalous images. Secondly, we propose an anomaly appearance and shape decoupling mechanism, which designs an embedding similarity loss to enforce the alignment between the anomaly text prompt and anomalies generated with different shapes at the same location, making the appearance of generated anomalies better maintain semantic consistency when the anomaly shape changes. Then, a training-free local control enhancement strategy is employed to provide stronger control intensity to anomaly regions during inference for finer alignment with anomaly masks. Finally, a hard sample generation module is proposed to create anomalous samples with subtle shapes and imperceptible anomaly appearances, enabling the downstream anomaly detection model to focus on learning low-saliency anomaly features. Extensive experiments demonstrate that anomalous images generated by our model outperform the state-of-the-art anomaly generation methods in terms of authenticity and consistency, and can significantly improve the performance of downstream anomaly detection tasks, especially anomaly localization. Yuanyi Duan, Qinlong Wu, Guosen Xie, Fang Zhao 0006, Caifeng Shan |
ACM Multimedia | 4 |
| 2025 | You Only Communicate Once: One-shot Federated Low-Rank Adaptation of MLLMabstractMultimodal Large Language Models (MLLMs) with Federated Learning (FL) can quickly adapt to privacy-sensitive tasks, but face significant challenges such as high communication costs and increased attack risks, due to their reliance on multi-round communication. To address this, One-shot FL (OFL) has emerged, aiming to complete adaptation in a single client-server communication. However, existing adaptive ensemble OFL methods still need more than one round of communication, because correcting heterogeneity-induced local bias relies on aggregated global supervision, meaning they still do not achieve true one-shot communication. In this work, we make the first attempt to achieve true one-shot communication for MLLMs under OFL, by investigating whether implicit (i.e., initial rather than aggregated) global supervision alone can effectively correct local training bias. Our key finding from the empirical study is that imposing directional supervision on local training substantially mitigates client conflicts and local bias. Building on this insight, we propose YOCO, in which directional supervision with sign-regularized LoRA B enforces global consistency, while sparsely regularized LoRA A preserves client-specific adaptability. Experiments demonstrate that YOCO cuts communication to $\sim$0.03\% of multi-round FL while surpassing those methods in several multimodal scenarios and consistently outperforming all one-shot competitors. Binqian Xu, Haiyang Mei, Zechen Bai, Jinjin Gong, Rui Yan 0010, Guosen Xie, Yazhou Yao, Basura Fernando, Xiangbo Shu |
NeurIPS | 6 |
| 2025 | STPM: Spatial-Temporal Token Pruning and Merging for Complex Activity RecognitionabstractLightweight video representation techniques have advanced significantly for simple activity recognition, but they still encounter several issues when applied to complex activity recognition: 1) The presence of numerous individuals and varying spatial positions makes it difficult for traditional token pruning methods to maintain accuracy. 2) Simply discarding entire frames may result in the loss of crucial clues. 3) To maintain parallel computing, applying the same pruning rate to every frame leads to significant redundancy in frames with low information content. To this end, we propose a lightweight and novel Spatial-Temporal Token Pruning and Merging (STPM) framework, specifically designed for complex action videos where human actors occupy a small spatial resolution within video frames. Our framework considers two critical factors: semantic importance and spatial-temporal redundancy, to further reduce overhead. For semantic importance, STPM captures class-specific attention scores by learning multiple class tokens within the transformer to guide token pruning. For spatial-temporal redundancy, STPM employs an anchor graph and temporal attention to perform spatial and temporal token merging, preserving appearance and temporal cues while eliminating semantic duplication and redundancy. We conduct extensive experiments on JRDB-PAR primarily using recently introduced video transformer backbones, e.g., MViT and ViT. Our framework achieves similar results while requiring 40% less computation. Yumeng Su, Jiachao Zhang, Rui Yan 0010, Pengpeng Li 0001, Guosen Xie, Xiangbo Shu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Multi-Granularity Aggregation Network for Remote Sensing Few-Shot SegmentationabstractFew-shot semantic segmentation (FSS) aims to segment a query image using a limited number of densely annotated support images from the same category. Most existing conventional FSS methods are tailored for coping with images from natural scenarios. Unlike natural images, remote sensing images usually have a similar background context among the support and query image pairs, and more severe intraclass inconsistency exists due to overhead shooting views. However, facing such realistic and challenging remote sensing FSS tasks, the existing methods seldom consider these intrinsic characteristics from a unified viewpoint, thus leading to inferior results. To solve the above dilemma, we propose a multi-granularity aggregation network (MGANet) to progressively capture multi-granularity discriminative information, for tackling the remote sensing FSS task. Specifically, MGANet consists of a multi-granularity similarity (MGS) module and an adaptive multiprototype aggregation (AMPA) module. To fully utilize background context, MGS extracts multi-granularity support and query feature maps from the backbone network to calculate a holistic correlation by incorporating the background information. Next, to alleviate the intraclass inconsistency of remote sensing images, AMPA decomposes the support foreground region into mainstay and auxiliary subregions by the guidance of reverse prediction on support features, thus generating three types of prototypes by masked average pooling (MAP) on these paired features and masks. Furthermore, these multiprototypes are collaboratively interacted with the query features to pursue reinforced discriminative features, relying on prototype-aware slot attention (PASA). Extensive experiments on iSAID-$5^{i}$and LoveDA-$2^{i}$demonstrate well the superiority of the proposed MGANet. The source code is available athttps://github.com/CVL-hub/MGANet/. Shi-Feng Peng, Guosen Xie, Fang Zhao 0006, Xiangbo Shu, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Uncertainty-Aware Transformer for Referring Camouflaged Object DetectionabstractReferring camouflaged object detection (Ref-COD) is a recently proposed task, aiming to segment specified camouflaged objects by leveraging visual reference, i.e., a small set of referring images with salient target objects. Ref-COD poses a considerable challenge due to the difficulty of discerning camouflaged objects from their highly similar backgrounds, as well as the significant feature differences between the camouflaged objects and the provided visual reference. To tackle the above dilemma, we propose a novel uncertainty-aware transformer for the Ref-COD task, termed UAT. UAT first utilizes a cross-attention mechanism to align and integrate visual reference to guide camouflaged feature learning, and then models dependencies between patches in a probabilistic manner to learn predictive uncertainty and excavate discriminative camouflaged features. Specifically, we first design a referring feature aggregation (RFA) module to align and incorporate referring features with camouflaged features, guiding targeted specific feature learning within the feature space of camouflaged images. Then, to enhance multi-level feature extraction, we develop a cross-attention encoder (CAE) to integrate global information and multi-scale semantics between adjacent layers to excavate critical camouflage cues. More importantly, we propose a transformer probabilistic decoder (TPD) to model the dependencies between patches as Gaussian random variables to capture uncertainty-aware camouflaged features. Extensive experiments on the golden Ref-COD benchmark demonstrate the superiority of UAT over existing state-of-the-art competitors. The proposed UAT also achieves competitive performance on several conventional COD datasets, further demonstrating its scalability. The source code is available at https://github.com/CVL-hub/UAT. Ranwan Wu, Tian-Zhu Xiang, Guosen Xie, Rongrong Gao, Xiangbo Shu, Fang Zhao 0006, Ling Shao 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Hierarchical Motion-Enhanced Matching Framework for Few-Shot Action RecognitionabstractFew-Shot Action Recognition (FSAR) aims to recognize novel class action with limited annotated training data from the same class. Most FSAR methods subconsciously follow the few-shot image classification solutions by solely focusing on appearance-level matching between support and query videos, such as part-level matching, frame-level matching, and segment-level matching. However, these methods, almost always, have two main limitations: 1) generally ignore the relationship among these part-, frame- and segment-level features and 2) may mismatch the same class actions under fast-term and slow-term dynamics. To this end, we present a novel Hierarchical Motion-enhanced Matching (HM${^{2}}$) framework to hierarchically learn the relation-aware multi-modal features, and jointly promote the multi-modal matching, including appearance-level matching on segments, frames, and parts, as well as the motion-level matching on dynamics. Specifically, we first propose a new Hierarchical Tokenizer (HT) to learn multi-modal features, namely utilizing a hierarchical Transformer to learn appearance-level features, along with a Slow-Fast Aware Motion (SFAM) strategy to learn motion-level features covering fast- and slow-term dynamics. Next, we propose a new Relation-aware Matcher (RM) to match the multi-modal features, by leveraging a Hierarchical Relational Graph Convolutional Network (H-RGCN) to capture the relationship among these appearance-level features. Further, a Dual Sample-to-Class Matching (DSCM) strategy is proposed to measure the bidirectional similarities among appearance- and motion-modal features by sample-to-class matching and class-to-sample matching. Extensive experiments on four golden FSAR datasets demonstrate significant performance improvements of HM${^{2}}$compared with the state-of-the-art methods. Hailiang Gao, Guosen Xie, Rui Yan 0010, Qiongjie Cui, Hongyu Qu, Xiangbo Shu |
IEEE Trans. Multim. | 2 |
| 2025 | AFANet: Adaptive Frequency-Aware Network for Weakly-Supervised Few-Shot Semantic SegmentationabstractFew-shot learning aims to recognize novel concepts by leveraging prior knowledge learned from a few samples. However, for visually intensive tasks such as few-shot semantic segmentation, pixel-level annotations are time-consuming and costly. Therefore, in this paper, we utilize the more challenging image-level annotations and propose an adaptive frequency-aware network (AFANet) for weakly-supervised few-shot semantic segmentation (WFSS). Specifically, we first propose a cross-granularity frequency-aware module (CFM) that decouples RGB images into high-frequency and low-frequency distributions and further optimizes semantic structural information by realigning them. Unlike most existing WFSS methods using the textual information from the multi-modal language-vision model, e.g., CLIP, in an offline learning manner, we further propose a CLIP-guided spatial-adapter module (CSM), which performs spatial domain adaptive transformation on textual information through online learning, thus providing enriched cross-modal semantic information for CFM. Extensive experiments on the Pascal-5iand COCO-20idatasets demonstrate that AFANet has achieved state-of-the-art performance. Jiaqi Ma 0006, Guosen Xie, Fang Zhao 0006, Zechao Li |
IEEE Trans. Multim. | 2 |
| 2025 | MVP-Shot: Multi-Velocity Progressive-Alignment Framework for Few-Shot Action RecognitionabstractRecent few-shot action recognition (FSAR) methods typically perform semantic matching on learned discriminative features to achieve promising performance. However, most FSAR methods focus on single-scale (e.g., frame-level, segment-level,etc.) feature alignment, which ignores that human actions with the same semantic may appear at different velocities. To this end, we develop a novel Multi-Velocity Progressive-alignment (MVP-Shot) framework to progressively learn and align semantic-related action features at multi-velocity levels. Concretely, a Multi-Velocity Feature Alignment (MVFA) module is designed to measure the similarity between features from support and query videos with different velocity scales and then merge all similarity scores in a residual fashion. To avoid the multiple velocity features deviating from the underlying motion semantic, our proposed Progressive Semantic-Tailored Interaction (PSTI) module injects velocity-tailored text information into the video feature via feature interaction on channel and temporal domains at different velocities. The above two modules compensate for each other to make more accurate query sample predictions under the few-shot settings. Experimental results show our method outperforms current state-of-the-art methods on multiple standard few-shot benchmarks (i.e., HMDB51, UCF101, Kinetics, SSv2-full, and SSv2-small). Hongyu Qu, Rui Yan 0010, Xiangbo Shu, Hailiang Gao, Guosen Xie |
IEEE Trans. Multim. | 6 |
| 2025 | Visual-Semantic Graph Matching Net for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to leverage additional semantic information to recognize unseen classes. To transfer knowledge from seen to unseen classes, most ZSL methods often learn a shared embedding space by simply aligning visual embeddings with semantic prototypes. However, methods trained under this paradigm often struggle to learn robust embedding space because they align the two modalities in an isolated manner among classes, which ignore the crucial class relationship during the alignment process. To address the aforementioned challenges, this article proposes a visual-semantic graph matching net (VSGMN), which leverages semantic relationships among classes to aid in visual-semantic embedding. VSGMN uses a graph build net (GBN) and a graph matching net (GMN) to achieve two-stage visual-semantic alignment. Specifically, GBN first uses an embedding-based approach to build visual and semantic graphs in the semantic space and align the embedding with its prototype for first-stage alignment. In addition, to supplement unseen class relationships in these graphs, GBN also builds the unseen class nodes based on semantic relationships. In the second stage, GMN continuously integrates neighbor and cross-graph information into the constructed graph nodes and aligns the node relationships between the two graphs under the class relationship constraint. Extensive experiments on three benchmark datasets demonstrate that VSGMN achieves superior performance in both conventional and generalized ZSL (GZSL) scenarios. The implementation of our VSGMN and experimental results are available at github: https://github.com/dbwfd/VSGMN. Bowen Duan 0001, Shiming Chen 0002, Yufei Guo 0001, Guosen Xie, Weiping Ding 0001, Yisong Wang 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Coarse-Fine Nested Network for Weakly Supervised Group Activity RecognitionabstractWeakly supervised group activity recognition (WSGAR) aims at identifying the overall behavior of multiple persons without any fine-grained supervision information (including individual position and action label). Traditional methods usually adopt a person-to-whole way: detect persons via off-the-shelf detectors, obtain person-level features, and integrate into the group-level features for training the classifier. However, these methods are unflexible due to serious reliance on the quality of detectors. To get rid of the detector, recent works learn several prototype tokens from noisy grid features with learnable weights directly, which treat all the local visual information equally and bring in redundant and ambiguous information to some extent. To this end, we propose a novel coarse-fine nested network (CFNN) to coarsely localize the key visual patches of activity and further finely learn the local features, as well as the global features. Specifically, we design a nested interactor (NI) to progressively model the spatiotemporal interactions of the learnable global token. According to the cue of spatial interaction in NI, we localize several key visual patches via a new coarse-grained spatial localizer (CSL). Then, we finally encode these localized visual patches with the help of global spatiotemporal dependency via a new fine-grained spatiotemporal selector (FSS). Extensive experiments on Volleyball and NBA datasets demonstrate the effectiveness of the proposed CFNN compared with the existing competitive methods. Code is available at: https://github.com/gexiaojingshelby/CFNN. Xiaojing Ge, Rui Yan 0010, Xiangbo Shu, Keke Chen, Guosen Xie |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Attribute Prompt Alignment Network for Zero-Shot LearningabstractIn the vanilla zero-shot learning (ZSL) paradigm, category attributes is the key for knowledge generalizable transfer from seen to unseen classes. By contrast, the current contrastive language-image pretraining (CLIP) model relies on the category names to achieve a more general ZSL-like prediction. When vanilla ZSL meets general CLIP, however, most existing methods on both sides struggle to benefit from each other. In this brief, we resort to attribute prompt tuning (APT) for improving the knowledge transferability from the pretrained CLIP model to the downstream ZSL framework for pursuing desirable feature representations. Our approach, termed as attribute prompt alignment network (APAN), leverages APT for cross-network feature alignment (CFA). In this way, we can investigate the effects of CLIP to vanilla ZSL task in the era of large model by the two branch APAN architecture. Specifically, APT takes as an input the templates of class attribute descriptions to produce attribute prompts, which are further used to both guide the localizations of visual regions across two frozen feature extraction networks, through a visual-semantic interaction attention. This enables APAN to progressively refine and align these cross-network features, thus resulting in generalizable feature representations that can capture fine-grained attribute information. For CFA, we simply introduce prediction alignment loss that constrains the predictions from these two cross-network visual features. Experimental results on three benchmark datasets well demonstrate that APAN outperforms the state-of-the-art methods by absorbing generalizable knowledge from CLIP models. Guosen Xie, Ting Guo 0004, Xiangbo Shu, Fang Zhao 0006, Zheng Zhang 0006, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Label-Efficient Few-Shot Semantic Segmentation with Unsupervised Meta-TrainingabstractThe goal of this paper is to alleviate the training cost for few-shot semantic segmentation (FSS) models. Despite that FSS in nature improves model generalization to new concepts using only a handful of test exemplars, it relies on strong supervision from a considerable amount of labeled training data for base classes. However, collecting pixel-level annotations is notoriously expensive and time-consuming, and small-scale training datasets convey low information density that limits test-time generalization. To resolve the issue, we take a pioneering step towards label-efficient training of FSS models from fully unlabeled training data, or additionally a few labeled samples to enhance the performance. This motivates an approach based on a novel unsupervised meta-training paradigm. In particular, the approach first distills pre-trained unsupervised pixel embedding into compact semantic clusters from which a massive number of pseudo meta-tasks is constructed. To mitigate the noise in the pseudo meta-tasks, we further advocate a robust Transformer-based FSS model with a novel prototype-based cross-attention design. Extensive experiments have been conducted on two standard benchmarks, i.e., PASCAL-5i and COCO-20i, and the results show that our method produces impressive performance without any annotations, and is comparable to fully supervised competitors even using only 20% of the annotations. Our code is available at: https://github.com/SSSKYue/UMTFSS. Jianwu Li, Kaiyue Shi, Guosen Xie, Xiaofeng Liu 0006, Jian Zhang 0002, Tianfei Zhou |
AAAI | 3 |
| 2024 | AdaFPP: Adapt-Focused Bi-Propagating Prototype Learning for Panoramic Activity RecognitionabstractPanoramic Activity Recognition (PAR) aims to identify multi-granul-arity behaviors performed by multiple persons in panoramic scenes, including individual activities, group activities, and global activities. Previous methods 1) heavily rely on manually annotated detection boxes in training and inference, hindering further practical deployment; or 2) directly employ normal detectors to detect multiple persons with varying size and spatial occlusion in panoramic scenes, blocking the performance gain of PAR. To this end, we consider learning a detector adapting varying-size occluded persons, which is optimized along with the recognition module in the all-in-one framework. Therefore, we propose a novel Adapt-Focused bi-Propagating Prototype learning (AdaFPP) framework to jointly recognize individual, group, and global activities in panoramic activity scenes by learning an adapt-focused detector and multi-granularity prototypes as the pretext tasks in an end-to-end way. Specifically, to accommodate the varying sizes and spatial occlusion of multiple persons in crowed panoramic scenes, we introduce a panoramic adapt-focuser, achieving the size-adapting detection of individuals by comprehensively selecting and performing fine-grained detections on object-dense sub-regions identified through original detections. In addition, to mitigate information loss due to inaccurate individual localizations, we introduce a bi-propagation prototyper that promotes closed-loop interaction and informative consistency across different granularities by facilitating bidirectional information propagation among the individual, group, and global levels. Extensive experiments demonstrate the significant performance of AdaFPP and emphasize its powerful applicability for PAR. Meiqi Cao, Rui Yan 0010, Xiangbo Shu, Guangzhao Dai, Yazhou Yao, Guosen Xie |
ACM Multimedia | 6 |
| 2024 | Rethinking attribute localization for zero-shot learning
Shuhuang Chen, Shiming Chen 0002, Guosen Xie, Xiangbo Shu, Xinge You, Xuelong Li 0001 |
Sci. China Inf. Sci. | 3 |
| 2024 | On the Number of Linear Regions of Convolutional Neural Networks With Piecewise Linear ActivationsabstractOne fundamental problem in deep learning is understanding the excellent performance of deep Neural Networks (NNs) in practice. An explanation for the superiority of NNs is that they can realize a large family of complicated functions, i.e., they have powerful expressivity. The expressivity of a Neural Network with Piecewise Linear activations (PLNN) can be quantified by the maximal number of linear regions it can separate its input space into. In this paper, we provide several mathematical results needed for studying the linear regions of Convolutional Neural Networks with Piecewise Linear activations (PLCNNs), and use them to derive the maximal and average numbers of linear regions for one-layer PLCNNs. Furthermore, we obtain upper and lower bounds for the number of linear regions of multi-layer PLCNNs. Our results suggest that deeper PLCNNs have more powerful expressivity than shallow PLCNNs, while PLCNNs have more expressivity than fully-connected PLNNs per parameter, in terms of the number of linear regions. Huan Xiong, Lei Huang 0015, Wenston J. T. Zang, Xiantong Zhen, Guosen Xie, Bin Gu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Holistic Prototype Attention Network for Few-Shot Video Object SegmentationabstractFew-shot video object segmentation (FSVOS) aims to segment dynamic objects of unseen classes by resorting to a small set of support images that contain pixel-level object annotations. Existing methods have demonstrated that the domain agent-based attention mechanism is effective in FSVOS by learning the correlation between support images and query frames. However, the agent frame contains redundant pixel information and background noise, resulting in inferior segmentation performance. Moreover, existing methods tend to ignore inter-frame correlations in query videos. To alleviate the above dilemma, we propose a holistic prototype attention network (HPAN) for advancing FSVOS. Specifically, HPAN introduces a prototype graph attention module (PGAM) and a bidirectional prototype attention module (BPAM), transferring informative knowledge from seen to unseen classes. PGAM generates local prototypes from all foreground features and then utilizes their internal correlations to enhance the representation of the holistic prototypes. BPAM exploits the holistic information from support images and video frames by fusing co-attention and self-attention to achieve support-query semantic consistency and inner-frame temporal consistency. Extensive experiments on YouTube-FSVOS have been provided to demonstrate the effectiveness and superiority of our proposed HPAN method. Our source code and models are available anonymously at https://github.com/NUST-Machine-Intelligence-Laboratory/HPAN. Tao Chen 0012, Xiruo Jiang, Yazhou Yao, Guosen Xie, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | GNDAN: Graph Navigated Dual Attention Network for Zero-Shot LearningabstractZero-shot learning (ZSL) tackles the unseen class recognition problem by transferring semantic knowledge from seen classes to unseen ones. Typically, to guarantee desirable knowledge transfer, a direct embedding is adopted for associating the visual and semantic domains in ZSL. However, most existing ZSL methods focus on learning the embedding from implicit global features or image regions to the semantic space. Thus, they fail to: 1) exploit the appearance relationship priors between various local regions in a single image, which corresponds to the semantic information and 2) learn cooperative global and local features jointly for discriminative feature representations. In this article, we propose the novel graph navigated dual attention network (GNDAN) for ZSL to address these drawbacks. GNDAN employs a region-guided attention network (RAN) and a region-guided graph attention network (RGAT) to jointly learn a discriminative local embedding and incorporate global context for exploiting explicit global embeddings under the guidance of a graph. Specifically, RAN uses soft spatial attention to discover discriminative regions for generating local embeddings. Meanwhile, RGAT employs an attribute-based attention to obtain attribute-based region features, where each attribute focuses on the most relevant image regions. Motivated by the graph neural network (GNN), which is beneficial for structural relationship representations, RGAT further leverages a graph attention network to exploit the relationships between the attribute-based region features for explicit global embedding representations. Based on the self-calibration mechanism, the joint visual embedding learned is matched with the semantic embedding to form the final prediction. Extensive experiments on three benchmark datasets demonstrate that the proposed GNDAN achieves superior performances to the state-of-the-art methods. Our code and trained models are available at https://github.com/shiming-chen/GNDAN. Shiming Chen 0002, Ziming Hong, Guosen Xie, Qinmu Peng, Xinge You, Weiping Ding 0001, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Deep Metric Learning Based on Meta-Mining Strategy With Semiglobal InformationabstractRecently, deep metric learning (DML) has achieved great success. Some existing DML methods propose adaptive sample mining strategies, which learn to weight the samples, leading to interesting performance. However, these methods suffer from a small memory (e.g., one training batch), limiting their efficacy. In this work, we introduce a data-driven method, meta-mining strategy with semiglobal information (MMSI), to apply meta-learning to learn to weight samples during the whole training, leading to an adaptive mining strategy. To introduce richer information than one training batch only, we elaborately take advantage of the validation set of meta-learning by implicitly adding additional validation sample information to training. Furthermore, motivated by the latest self-supervised learning, we introduce a dictionary (memory) that maintains very large and diverse information. Together with the validation set, this dictionary presents much richer information to the training, leading to promising performance. In addition, we propose a new theoretical framework that can formulate pairwise and tripletwise metric learning loss functions in a unified framework. This framework brings new insights to society and facilitates us to generalize our MMSI to many existing DML methods. We conduct extensive experiments on three public datasets, CUB200-2011, Cars-196, and Stanford Online Products (SOP). Results show that our method can achieve the state of the art or very competitive performance. Our source codes have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/MMSI. Xi Jiang 0001, Sheng Liu 0009, Xili Dai, Guosheng Hu, Xingguo Huang, Yazhou Yao, Guosen Xie, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Learning Anchor Transformations for 3D Garment AnimationabstractThis paper proposes an anchor-based deformation model, namely AnchorDEF, to predict 3D garment animation from a body motion sequence. It deforms a garment mesh template by a mixture of rigid transformations with extra nonlinear displacements. A set of anchors around the mesh surface is introduced to guide the learning of rigid transformation matrices. Once the anchor transformations are found, per-vertex nonlinear displacements of the garment template can be regressed in a canonical space, which reduces the complexity of deformation space learning. By explicitly constraining the transformed anchors to satisfy the consistencies of position, normal and direction, the physical meaning of learned anchor transformations in space is guaranteed for better generalization. Furthermore, an adaptive anchor updating is proposed to optimize the anchor position by being aware of local mesh topology for learning representative anchor transformations. Qualitative and quantitative experiments on different types of garments demonstrate that AnchorDEF achieves the state-of-the-art performance on 3D garment deformation prediction in motion, especially for loose-fitting garments. Fang Zhao 0006, Zekun Li 0002, Shaoli Huang, Junwu Weng, Tianfei Zhou, Guosen Xie, Jue Wang 0001, Ying Shan |
CVPR | 6 |
| 2023 | Swap-Reconstruction Autoencoder for Compositional Zero-Shot LearningabstractCompositional zero-shot learning (CZSL) aims to distinguish images from unseen compositional classes, which consist of state and object concepts that individually appear in some seen compositional images. The key challenge of CZSL is how to effectively mitigate the contextuality issue for achieving a desirable compositional transfer from seen classes to unseen ones. In CZSL, the visual appearances of the same state are inconsistent when combined with different objects. To address the above dilemma, we propose a swap-reconstruction autoencoder (SRA) to capture the intrinsic context of the ambiguous states. Specifically, SRA learns a consistent embedding space for multi-modal data. A swap-reconstruction mechanism is designed to disentangle the visual embedding of states and objects. The loss including a superclass-oriented state swap-reconstruction loss and object swap-reconstruction loss model the contextual relationship between states and objects. Extensive experiments demonstrate that SRA outperforms current state-of-the-art methods on the three benchmark datasets. Ting Guo 0004, Jiye Liang, Guosen Xie |
ICME | 3 |
| 2023 | MUP: Multi-granularity Unified Perception for Panoramic Activity RecognitionabstractPanoramic activity recognition is required to jointly identify multi-granularity human behaviors including individual actions, group activities, and global activities in multi-person videos. Previous methods encode these behaviors hierarchically through multiple stages, which disturb the inherent co-occurrence across multi-granularity behaviors in the same scene. To this end, we propose a novel Multi-granularity Unified Perception (MUP) framework that perceives different granularity behaviors universally to explore the co-occurrence motion pattern via the same parameters in an end-to-end fashion. To be specific, the proposed framework stacks three Unified Motion Encoding (UME) blocks for modeling multiple granularity behaviors with shared parameters. UME block mines intra-relevant and cross-relevant semantics synchronously from input feature sequences via Intra-granularity Motion Embedding (IME) and Cross-granularity Motion Prototyping (CMP). In particular, IME aims to model the interactions among visual features within each granularity based on the attention mechanism. CMP aims to aggregate features across different granularities (i.e., person to group) via several learnable prototypes. Extensive experiments demonstrate that MUP outperforms the state-of-the-art methods on JRDB-PAR and has satisfactory interpretability. Meiqi Cao, Rui Yan 0010, Xiangbo Shu, Jiachao Zhang, Jinpeng Wang 0001, Guosen Xie |
ACM Multimedia | 6 |
| 2023 | Foreground/Background-Masked Interaction Learning for Spatio-temporal Action DetectionabstractSpatio-temporal Action Detection (SAD) aims to recognize the multi-class actions, and meanwhile locate their spatio-temporal occurrence in untrimmed videos. Besides relying on the inherent inter-actor interactions, most previous SAD approaches model actor interactions between multi-actors and the whole frames or special parts (e.g., objects/hands). However, such approaches are relatively graceless by 1) roughly treating all various actors to equivalently interact with frames/parts or by 2) sumptuously borrowing multiple costly detectors to acquire the special parts. To solve the above dilemma, we propose a novel Foreground/Background-masked Interaction Learning (dubbed as FBI Learning) framework to learn the multi-actor features by attentively interacting with the hands-down foreground and background frames. Specifically, we first design a new Mask-guided Cross Attention (MCA) mechanism that calculates the masked cross-attentions to capture the compact relations between the actors and foreground/background regions. Next, we present a new Actor-guided Feature Aggregation (AFA) scheme that integrates foreground- and background-interacted actor features with the learnable actor-based weights. Finally, we construct a long-term feature bank that associates temporal context information to facilitate action classification. Extensive experiments are conducted on commonly available UCF101-24, MultiSports, and AVA v2.1/v2.2 datasets, which illustrate the competitive performance of FBI Learning against the state-of-the-art methods. Keke Chen, Xiangbo Shu, Guosen Xie, Rui Yan 0010, Jinhui Tang 0001 |
ACM Multimedia | 3 |
| 2023 | Group-wise interactive region learning for zero-shot recognition
Ting Guo 0004, Jiye Liang, Guosen Xie |
Inf. Sci. | 3 |
| 2023 | TransZero++: Cross Attribute-Guided Transformer for Zero-Shot LearningabstractZero-shot learning (ZSL) tackles the novel class recognition problem by transferring semantic knowledge from seen classes to unseen ones. Semantic knowledge is typically represented by attribute descriptions shared between different classes, which act as strong priors for localizing object attributes that represent discriminative region features, enabling significant and sufficient visual-semantic interaction for advancing ZSL. Existing attention-based models have struggled to learn inferior region features in a single image by solely using unidirectional attention, which ignore the transferable and discriminative attribute localization of visual features for representing the key semantic knowledge for effective knowledge transfer in ZSL. In this paper, we propose a cross attribute-guided Transformer network, termed TransZero++, to refine visual features and learn accurate attribute localization for key semantic knowledge representations in ZSL. Specifically, TransZero++ employs an attribute → visual Transformer sub-net (AVT) and a visual → attribute Transformer sub-net (VAT) to learn attribute-based visual features and visual-based attribute features, respectively. By further introducing feature-level and prediction-level semantical collaborative losses, the two attribute-guided transformers teach each other to learn semantic-augmented visual embeddings for key semantic knowledge representations via semantical collaborative learning. Finally, the semantic-augmented visual embeddings learned by AVT and VAT are fused to conduct desirable visual-semantic interaction cooperated with class semantic vectors for ZSL classification. Extensive experiments show that TransZero++ achieves the new state-of-the-art results on three golden ZSL benchmarks and on the large-scale ImageNet dataset. The project website is available at: https://shiming-chen.github.io/TransZero-pp/TransZero-pp.html. Shiming Chen 0002, Ziming Hong, Wenjin Hou, Guosen Xie, Yibing Song, Jian Zhao 0006, Xinge You, Shuicheng Yan, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | A Survey on Learning to RejectabstractLearning to reject is a special kind of self-awareness (the ability to know what you do not know), which is an essential factor for humans to become smarter. Although machine intelligence has become very accurate nowadays, it lacks such kind of self-awareness and usually acts as omniscient, resulting in overconfident errors. This article presents a comprehensive overview of this topic from three perspectives: confidence, calibration, and discrimination. Confidence is an important measurement for the reliability of model predictions. Rejection can be realized by setting thresholds on confidence. However, most models, especially modern deep neural networks, are usually overconfident. Therefore, calibration is a process to ensure confidence matching the actual likelihood of correctness, including two approaches: post-calibration and self-calibration. Calibration reflects the global characteristic of confidence, and the local distinguishing property of confidence is also important. In light of this, discrimination focuses on the performance of accepting positive samples while rejecting negative samples. As a binary classification problem, the challenge of discrimination comes from the missing and nonrepresentativeness of the negative data. Three discrimination tasks are comprehensively analyzed and discussed: failure rejection, unknown rejection, and fake rejection. By rejecting failures, the risk could be controlled especially for mission-critical applications. By rejecting unknowns, the awareness of the knowledge blind zone would be enhanced. By rejecting fakes, security and privacy could be protected. We provide a general taxonomy, organization, and discussion of the methods for solving these problems, which are studied separately in the literature. The connections between different approaches and future directions that are worth further investigation are also presented. With a discriminative and calibrated confidence, learning to reject will let the decision-making process be more practical, reliable, and secure. Xu-Yao Zhang, Guosen Xie, Xiuli Li, Tao Mei 0001, Cheng-Lin Liu 0001 |
Proc. IEEE | 2 |
| 2023 | Robust learning from noisy web data for fine-Grained recognition
Zhenhuang Cai, Guosen Xie, Xingguo Huang, Yazhou Yao, Zhenmin Tang |
Pattern Recognit. | 2 |
| 2023 | Towards Zero-Shot Learning: A Brief Review and an Attention-Based Embedding NetworkabstractZero-shot learning (ZSL), an emerging topic in recent years, targets at distinguishing unseen class images by taking images from seen classes for training the classifier. Existing works often build embeddings between global feature space and attribute space, which, however, neglect the treasure in image parts. Discrimination information is usually contained in the image parts, e.g., black and white striped area of a zebra is the key difference from a horse. As such, image parts can facilitate the transfer of knowledge among the seen and unseen categories. In this paper, we first conduct a brief review on ZSL with detailed descriptions of these methods. Next, to discover meaningful parts, we propose an end-to-end attention-based embedding network for ZSL, which contains two sub-streams: the attention part embedding (APE) stream, and the attention second-order embedding (ASE) stream. APE is used to discover multiple image parts based on attention. ASE is introduced for ensuring knowledge transfer stably by second-order collaboration. Furthermore, an adaptive thresholding strategy is proposed to suppress noise and redundant parts. Finally, a global branch is incorporated for the full use of global information. Experiments on four benchmarks demonstrate that our models achieve superior results under both ZSL and GZSL settings. Guosen Xie, Zheng Zhang 0006, Huan Xiong, Ling Shao 0001, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Saliency Guided Inter- and Intra-Class Relation Constraints for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation with only image-level labels aims to reduce annotation costs for the segmentation task. Existing approaches generally leverage class activation maps (CAMs) to locate the object regions for pseudo label generation. However, CAMs can only discover the most discriminative parts of objects, thus leading to inferior pixel-level pseudo labels. To address this issue, we propose a saliency guidedInter- andIntra-ClassRelationConstrained (I$^{2}$CRC) framework to assist the expansion of the activated object regions in CAMs. Specifically, we propose a saliency guided class-agnostic distance module to pull the intra-category features closer by aligning features to their class prototypes. Further, we propose a class-specific distance module to push the inter-class features apart and encourage the object region to have a higher activation than the background. Besides strengthening the capability of the classification network to activate more integral object regions in CAMs, we also introduce an object guided label refinement module to take a full use of both the segmentation prediction and the initial labels for obtaining superior pseudo-labels. Extensive experiments on PASCAL VOC 2012 and COCO datasets demonstrate well the effectiveness of I$^{2}$CRC over other state-of-the-art counterparts. Tao Chen 0012, Yazhou Yao, Lei Zhang 0054, Qiong Wang 0003, Guosen Xie, Fumin Shen |
IEEE Trans. Multim. | 5 |
| 2023 | Co-Communication Graph Convolutional Network for Multi-View Crowd CountingabstractWe study and address the multi-view crowd counting (MVCC) problem which poses more realistic challenges than single-view crowd counting for better facilitating crowd management/public safety systems. Its major challenge lies in how to fully distill and aggregate useful, complementary information among multiple camera views to create powerful ground-plane representations for wide-area crowd analysis. In this paper, we present a graph-based, multi-view learning model called Co-Communication Graph Convolutional Network (CoCo-GCN) to jointly investigate intra-view contextual dependencies and inter-view complementary relations. More specifically, CoCo-GCN builds a view-agnostic graph interaction space for each camera view to conduct efficient contextual reasoning, and extends the intra-view reasoning by using a novel Graph Communication Layer (GCL) to also take between-graph (cross-view), complementary information into account. Moreover, CoCo-GCN uses a new Co-Memory Layer (CoML) to jointly coarsen the graphs and close the ‘representational gap’ among them for further exploiting the compositional nature of graphs and learning more consistent representations. Finally, these jointly learned features of multiple views can be easily fused to create ground-plane representations for wide-area crowd counting. Experiments show that the proposed CoCo-GCN achieves state-of-the-art results on three MVCC datasets, i.e., PETS2009, DukeMTMC, and City Street, significantly improving the scene-level accuracy over previous models. Qiang Zhai, Fan Yang 0054, Xin Li 0079, Guosen Xie, Hong Cheng 0002, Zicheng Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Leveraging Balanced Semantic Embedding for Generative Zero-Shot LearningabstractGenerative (generalized) zero-shot learning [(G)ZSL] models aim to synthesize unseen class features by using only seen class feature and attribute pairs as training data. However, the generated fake unseen features tend to be dominated by the seen class features and thus classified as seen classes, which can lead to inferior performances under zero-shot learning (ZSL), and unbalanced results under generalized ZSL (GZSL). To address this challenge, we tailor a novel balanced semantic embedding generative network (BSeGN), which incorporates balanced semantic embedding learning into generative learning scenarios in the pursuit of unbiased GZSL. Specifically, we first design a feature-to-semantic embedding module (FEM) to distinguish real seen and fake unseen features collaboratively with the generator in an online manner. We introduce the bidirectional contrastive and balance losses for the FEM learning, which can guarantee a balanced prediction for the interdomain features. In turn, the updated FEM can boost the learning of the generator. Next, we propose a multilevel feature integration module (mFIM) from the cycle-consistency branch of BSeGN, which can mitigate the domain bias through feature enhancement. To the best of our knowledge, this is the first work to explore embedding and generative learning jointly within the field of ZSL. Extensive evaluations on four benchmarks demonstrate the superiority of BSeGN over its state-of-the-art counterparts. Guosen Xie, Xu-Yao Zhang, Tian-Zhu Xiang, Fang Zhao 0006, Zheng Zhang 0006, Ling Shao 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Kernelized Similarity Learning and Embedding for Dynamic Texture SynthesisabstractDynamic texture (DT) exhibits statistical stationarity in the spatial domain and stochastic repetitiveness in the temporal dimension, indicating that different frames of DT possess a high similarity correlation that is critical prior knowledge. However, existing methods cannot effectively learn a synthesis model for high-dimensional DT from a small number of training samples. In this article, we propose a novel DT synthesis method, which makes full use of similarity as prior knowledge to address this issue. Our method is based on the proposed kernel similarity embedding, which can not only mitigate the high dimensionality and small sample issues, but also has the advantage of modeling nonlinear feature relationships. Specifically, we first put forward two hypotheses that are essential for the DT model to generate new frames using similarity correlations. Then, we integrate kernel learning and the extreme learning machine into a unified synthesis model to learn kernel similarity embeddings for representing DTs. Extensive experiments on DT videos collected from the Internet and two benchmark datasets, i.e., Gatech Graphcut Textures and Dyntex, demonstrate that the learned kernel similarity embeddings can provide discriminative representations for DTs. Further, our method can preserve the long-term temporal continuity of the synthesized DT sequences with excellent sustainability and generalization. Meanwhile, it effectively generates realistic DT videos with higher speed and lower computation than the current state-of-the-art methods. The code and more synthesis videos are available at our project pagehttps://shiming-chen.github.io/Similarity-page/Similarit.html. Shiming Chen 0002, Peng Zhang 0040, Guosen Xie, Qinmu Peng, Zehong Cao, Wei Yuan 0001, Xinge You |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | TransZero: Attribute-Guided Transformer for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to recognize novel classes by transferring semantic knowledge from seen classes to unseen ones. Semantic knowledge is learned from attribute descriptions shared between different classes, which are strong prior for localization of object attribute for representing discriminative region features enabling significant visual-semantic interaction. Although few attention-based models have attempted to learn such region features in a single image, the transferability and discriminative attribute localization of visual features are typically neglected. In this paper, we propose an attribute-guided Transformer network to learn the attribute localization for discriminative visual-semantic embedding representations in ZSL, termed TransZero. Specifically, TransZero takes a feature augmentation encoder to alleviate the cross-dataset bias between ImageNet and ZSL benchmarks and improve the transferability of visual features by reducing the entangled relative geometry relationships among region features. To learn locality-augmented visual features, TransZero employs a visual-semantic decoder to localize the most relevant image regions to each attributes from a given image under the guidance of attribute semantic information. Then, the locality-augmented visual features and semantic vectors are used for conducting effective visual-semantic interaction in a visual-semantic embedding network. Extensive experiments show that TransZero achieves a new state-of-the-art on three ZSL benchmarks. The codes are available at: https://github.com/shiming-chen/TransZero. Shiming Chen 0002, Ziming Hong, Yang Liu 0069, Guosen Xie, Baigui Sun, Hao Li 0030, Qinmu Peng, Ke Lu 0002, Xinge You |
AAAI | 4 |
| 2022 | MSDN: Mutually Semantic Distillation Network for Zero-Shot LearningabstractThe key challenge of zero-shot learning (ZSL) is how to infer the latent semantic knowledge between visual and attribute features on seen classes, and thus achieving a desirable knowledge transfer to unseen classes. Prior works either simply align the global features of an image with its associated class semantic vector or utilize unidirectional attention to learn the limited latent semantic representations, which could not effectively discover the intrinsic semantic knowledge (e.g., attribute semantics) between visual and attribute features. To solve the above dilemma, we propose a Mutually Semantic Distillation Network (MSDN), which progressively distills the intrinsic semantic representations between visual and attribute features for ZSL. MSDN incorporates an attribute→visual attention sub-net that learns attribute-based visual features, and a visual→attribute attention sub-net that learns visual-based attribute features. By further introducing a semantic distillation loss, the two mutual attention sub-nets are capable of learning collaboratively and teaching each other throughout the training process. The proposed MSDN yields significant improvements over the strong baselines, leading to new state-of-the-art performances on three popular challenging benchmarks. Our codes have been available at: https://github.com/shiming-chen/MSDN. Shiming Chen 0002, Ziming Hong, Guosen Xie, Wenhan Yang, Qinmu Peng, Kai Wang 0036, Jian Zhao 0006, Xinge You |
CVPR | 3 |
| 2022 | Dynamic Prototype Convolution Network for Few-Shot Semantic SegmentationabstractThe key challenge for few-shot semantic segmentation (FSS) is how to tailor a desirable interaction among sup-port and query features and/or their prototypes, under the episodic training scenario. Most existing FSS methods im-plement such support/query interactions by solely leveraging plain operations - e.g., cosine similarity and feature concatenation - for segmenting the query objects. How-ever, these interaction approaches usually cannot well capture the intrinsic object details in the query images that are widely encountered in FSS, e.g., if the query object to be segmented has holes and slots, inaccurate segmentation al-most always happens. To this end, we propose a dynamic prototype convolution network (DPCN) to fully capture the aforementioned intrinsic details for accurate FSS. Specifi-cally, in DPCN, a dynamic convolution module (DCM) is firstly proposed to generate dynamic kernels from support foreground, then information interaction is achieved by con-volution operations over query features using these kernels. Moreover, we equip DPCN with a support activation mod-ule (SAM) and a feature filtering module (FFM) to generate pseudo mask and filter out background information for the query images, respectively. SAM and FFM together can mine enriched context information from the query features. Our DPCN is also flexible and efficient under the k-shot FSS setting. Extensive experiments on PASCAL-5iand COCO 20ishow that DPCN yields superior performances under both 1-shot and 5-shot settings. Jie Liu 0043, Yanqi Bao, Guosen Xie, Huan Xiong, Jan-Jakob Sonke, Efstratios Gavves |
CVPR | 3 |
| 2022 | Hierarchical Feature Alignment Network for Unsupervised Video Object Segmentation
Gensheng Pei, Fumin Shen, Yazhou Yao, Guosen Xie, Zhenmin Tang, Jinhui Tang 0001 |
ECCV (34) | 4 |
| 2022 | Semantic Compression Embedding for Generative Zero-Shot LearningabstractGenerative methods have been successfully applied in zero-shot learning (ZSL) by learning an implicit mapping to alleviate the visual-semantic domain gaps and synthesizing unseen samples to handle the data imbalance between seen and unseen classes. However, existing generative methods simply use visual features extracted by the pre-trained CNN backbone. These visual features lack attribute-level semantic information. Consequently, seen classes are indistinguishable, and the knowledge transfer from seen to unseen classes is limited. To tackle this issue, we propose a novel Semantic Compression Embedding Guided Generation (SC-EGG) model, which cascades a semantic compression embedding network (SCEN) and an embedding guided generative network (EGGN). The SCEN extracts a group of attribute-level local features for each sample and further compresses them into the new low-dimension visual feature. Thus, a dense-semantic visual space is obtained. The EGGN learns a mapping from the class-level semantic space to the dense-semantic visual space, thus improving the discriminability of the synthesized dense-semantic unseen visual features. Extensive experiments on three benchmark datasets, i.e., CUB, SUN and AWA2, demonstrate the significant performance gains of SC-EGG over current state-of-the-art methods and its baselines. Ziming Hong, Shiming Chen 0002, Guosen Xie, Wenhan Yang, Jian Zhao 0006, Yuanjie Shao, Qinmu Peng, Xinge You |
IJCAI | 3 |
| 2022 | Cross-modal propagation network for generalized zero-shot learning
Ting Guo 0004, Jianqing Liang, Jiye Liang, Guosen Xie |
Pattern Recognit. Lett. | 4 |
| 2022 | Self-Supervised Multi-Modal Hybrid Fusion Network for Brain Tumor SegmentationabstractAccurate medical image segmentation of brain tumors is necessary for the diagnosing, monitoring, and treating disease. In recent years, with the gradual emergence of multi-sequence magnetic resonance imaging (MRI), multi-modal MRI diagnosis has played an increasingly important role in the early diagnosis of brain tumors by providing complementary information for a given lesion. Different MRI modalities vary significantly in context, as well as in coarse and fine information. As the manual identification of brain tumors is very complicated, it usually requires the lengthy consultation of multiple experts. The automatic segmentation of brain tumors from MRI images can thus greatly reduce the workload of doctors and buy more time for treating patients. In this paper, we propose a multi-modal brain tumor segmentation framework that adopts the hybrid fusion of modality-specific features using a self-supervised learning strategy. The algorithm is based on a fully convolutional neural network. Firstly, we propose a multi-input architecture that learns independent features from multi-modal data, and can be adapted to different numbers of multi-modal inputs. Compared with single-modal multi-channel networks, our model provides a better feature extractor for segmentation tasks, which learns cross-modal information from multi-modal data. Secondly, we propose a new feature fusion scheme, named hybrid attentional fusion. This scheme enables the network to learn the hybrid representation of multiple features and capture the correlation information between them through an attention mechanism. Unlike popular methods, such as feature map concatenation, this scheme focuses on the complementarity between multi-modal data, which can significantly improve the segmentation results of specific regions. Thirdly, we propose a self-supervised learning strategy for brain tumor segmentation tasks. Our experimental results demonstrate the effectiveness of the proposed model against other state-of-the-art multi-modal medical segmentation methods. Feiyi Fang, Yazhou Yao, Tao Zhou 0002, Guosen Xie, Jianfeng Lu 0003 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Enhanced Feature Alignment for Unsupervised Domain Adaptation of Semantic SegmentationabstractUnsupervised domain adaptation for semantic segmentation aims to transfer knowledge from a labeled source domain to another unlabeled target domain. However, due to the label noise and domain mismatch, learning directly from source domain data tends to have poor performance. Though adversarial learning methods strive to reduce domain discrepancies by aligning feature distributions, traditional methods suffer from the training imbalance and feature distortion problems. Besides, due to the absence of target domain labels, the classifier is blind to features from the target domain during training. Consequently, the final classifier overfits the source domain features and usually fails to predict the structured outputs of the target domain. To alleviate these problems, we focus on enhancing the adversarial learning based feature alignment from three perspectives. First, a classification constrained discriminator is proposed to balance the adversarial training and alleviate the feature distortion problem. Next, to alleviate the classifier overfitting problem, self-training is collaboratively used to learn a domain robust classifier with target domain pseudo labels. Moreover, an efficient class centroid calculation module is proposed and the domain discrepancy is further reduced by aligning the feature centroids of the same class from different domains. Experimental evaluations on GTA5$\rightarrow$Cityscapes and SYNTHIA$\rightarrow$Cityscapes demonstrate state-of-the-art results compared to other counterpart methods. The source code and models have been made available at.11[Online]. Available:https://github.com/NUST-Machine-Intelligence-Laboratory/EFA. Tao Chen 0012, Shuihua Wang, Qiong Wang 0003, Zheng Zhang 0006, Guosen Xie, Zhenmin Tang |
IEEE Trans. Multim. | 5 |
| 2022 | Semantically Meaningful Class Prototype Learning for One-Shot Image SegmentationabstractOne-shot semantic image segmentation aims to segment the object regions for the novel class with only one annotated image. Recent works adopt the episodic training strategy to mimic the expected situation at testing time. However, these existing approaches simulate the test conditions too strictly during the training process, and thus cannot make full use of the given label information. Besides, these approaches mainly focus on the foreground-background target class segmentation setting. They only utilize binary mask labels for training. In this paper, we propose to leverage the multi-class label information during the episodic training. It will encourage the network to generate more semantically meaningful features for each category. After integrating the target class cues into the query features, we then propose a pyramid feature fusion module to mine the fused features for the final classifier. Furthermore, to take more advantage of the support image-mask pair, we propose a self-prototype guidance branch to support image segmentation. It can constrain the network for generating more compact features and a robust prototype for each semantic class. For inference, we propose a fused prototype guidance branch for the segmentation of the query image. Specifically, we leverage the prediction of the query image to extract the pseudo-prototype and combine it with the initial prototype. Then we utilize the fused prototype to guide the final segmentation of the query image. Extensive experiments demonstrate the superiority of our proposed approach. The source codes and models have been made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/SMCP. Tao Chen 0012, Guosen Xie, Yazhou Yao, Qiong Wang 0003, Fumin Shen, Zhenmin Tang, Jian Zhang 0002 |
IEEE Trans. Multim. | 2 |
| 2022 | Robust Learning From Noisy Web Images Via Data Purification for Fine-Grained RecognitionabstractManually labeling fine-grained datasetsis laborious and typically requires domain-specific expert knowledge. Conversely, a vast amount of web data is relatively easy to obtain with nearly no human effort. Therefore, learning from noisy web data for fine-grained tasks is attracting increasing attention in recent years. However, the presence of noise in web images is a huge obstacle for training robust fine-grained recognition models. To this end, we propose a novel approach to identify noisy images as well as specifically distinguish in- and out-of-distribution samples. It can purify the noisy web training set by discarding out-of-distribution noise and relabeling in-distribution noisy samples. Then we can train the model on the purified dataset to alleviate the harmful effects of noise and make the most of web images to achieve better performance. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is far superior to current state-of-the-art web-supervised methods. The data and source code of this work have been made publicly available at:https://github.com/NUST-Machine-Intelligence-Laboratory/Dataset-Purification. Chuanyi Zhang, Qiong Wang 0003, Guosen Xie, Qi Wu 0001, Fumin Shen, Zhenmin Tang |
IEEE Trans. Multim. | 3 |
| 2022 | Generalized Zero-Shot Learning With Multiple Graph Adaptive Generative NetworksabstractGenerative adversarial networks (GANs) for (generalized) zero-shot learning (ZSL) aim to generate unseen image features when conditioned on unseen class embeddings, each of which corresponds to one unique category. Most existing works on GANs for ZSL generate features by merely feeding the seen image feature/class embedding (combined with random Gaussian noise) pairs into the generator/discriminator for a two-player minimax game. However, the structure consistency of the distributions among the real/fake image features, which may shift the generated features away from their real distribution to some extent, is seldom considered. In this paper, to align the weights of the generator for better structure consistency between real/fake features, we propose a novel multigraph adaptive GAN (MGA-GAN). Specifically, a Wasserstein GAN equipped with a classification loss is trained to generate discriminative features with structure consistency. MGA-GAN leverages the multigraph similarity structures between sliced seen real/fake feature samples to assist in updating the generator weights in the local feature manifold. Moreover, correlation graphs for the whole real/fake features are adopted to guarantee structure correlation in the global feature manifold. Extensive evaluations on four benchmarks demonstrate well the superiority of MGA-GAN over its state-of-the-art counterparts. Guosen Xie, Zheng Zhang 0006, Guoshuai Liu, Fan Zhu 0001, Li Liu 0004, Ling Shao 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Scale-Aware Graph Neural Network for Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation (FSS) aims to segment unseen class objects given very few densely-annotated support images from the same class. Existing FSS methods find the query object by using support prototypes or by directly relying on heuristic multi-scale feature fusion. However, they fail to fully leverage the high-order appearance relationships between multi-scale features among the support-query image pairs, thus leading to an inaccurate localization of the query objects. To tackle the above challenge, we propose an end-to-end scale-aware graph neural network (SAGNN) by reasoning the cross-scale relations among the support-query images for FSS. Specifically, a scale-aware graph is first built by taking support-induced multi-scale query features as nodes and, meanwhile, each edge is modeled as the pairwise interaction of its connected nodes. By progressive message passing over this graph, SAGNN is capable of capturing cross-scale relations and overcoming object variations (e.g., appearance, scale and location), and can thus learn more precise node embeddings. This in turn enables it to predict more accurate foreground objects. Moreover, to make full use of the location relations across scales for the query image, a novel self-node collaboration mechanism is proposed to enrich the current node, which endows SAGNN the ability of perceiving different resolutions of the same objects. Extensive experiments on PASCAL-5iand COCO-20ishow that SAGNN achieves state-of-the-art results. Guosen Xie, Jie Liu 0043, Huan Xiong, Ling Shao 0001 |
CVPR | 1 |
| 2021 | Non-Salient Region Object Mining for Weakly Supervised Semantic SegmentationabstractSemantic segmentation aims to classify every pixel of an input image. Considering the difficulty of acquiring dense labels, researchers have recently been resorting to weak labels to alleviate the annotation burden of segmentation. However, existing works mainly concentrate on expanding the seed of pseudo labels within the image’s salient region. In this work, we propose a non-salient region object mining approach for weakly supervised semantic segmentation. We introduce a graph-based global reasoning unit to strengthen the classification network’s ability to capture global relations among disjoint and distant regions. This helps the network activate the object features outside the salient area. To further mine the non-salient region objects, we propose to exert the segmentation network’s self-correction ability. Specifically, a potential object mining module is proposed to reduce the false-negative rate in pseudo labels. Moreover, we propose a non-salient region masking module for complex images to generate masked pseudo labels. Our non-salient region masking module helps further discover the objects in the non-salient region. Extensive experiments on the PASCAL VOC dataset demonstrate state-of-the-art results compared to current methods. The source codes are available at https://github.com/NUST-Machine-Intelligence-Laboratory/nsrom. Yazhou Yao, Tao Chen 0012, Guosen Xie, Chuanyi Zhang, Fumin Shen, Qi Wu 0001, Zhenmin Tang, Jian Zhang 0002 |
CVPR | 3 |
| 2021 | Few-Shot Semantic Segmentation with Cyclic Memory NetworkabstractFew-shot semantic segmentation (FSS) is an important task for novel (unseen) object segmentation under the data-scarcity scenario. However, most FSS methods rely on unidirectional feature aggregation, e.g., from support prototypes to get the query prediction, and from high-resolution features to guide the low-resolution ones. This usually fails to fully capture the cross-resolution feature relationships and thus leads to inaccurate estimates of the query objects. To resolve the above dilemma, we propose a cyclic memory network (CMN) to directly learn to read abundant support information from all resolution features in a cyclic manner. Specifically, we first generate N pairs (key and value) of multi-resolution query features guided by the support feature and its mask. Next, we circularly take one pair of these features as the query to be segmented, and the rest N-1 pairs are written into an external memory accordingly, i.e., this leave-one-out process is conducted for N times. In each cycle, the query feature is updated by collaboratively matching its key and value with the memory, which can elegantly cover all the spatial locations from different resolutions. Furthermore, we incorporate the query feature re-adding and the query feature recursive updating mechanisms into the memory reading operation. CMN, equipped with these merits, can thus capture cross-resolution relationships and better handle the object appearance and scale variations in FSS. Experiments on PASCAL-5iand COCO-20iwell validate the effectiveness of our model for FSS. Guosen Xie, Huan Xiong, Jie Liu 0043, Yazhou Yao, Ling Shao 0001 |
ICCV | 1 |
| 2021 | HSVA: Hierarchical Semantic-Visual Adaptation for Zero-Shot LearningabstractZero-shot learning (ZSL) tackles the unseen class recognition problem, transferring semantic knowledge from seen classes to unseen ones. Typically, to guarantee desirable knowledge transfer, a common (latent) space is adopted for associating the visual and semantic domains in ZSL. However, existing common space learning methods align the semantic and visual domains by merely mitigating distribution disagreement through one-step adaptation. This strategy is usually ineffective due to the heterogeneous nature of the feature representations in the two domains, which intrinsically contain both distribution and structure variations. To address this and advance ZSL, we propose a novel hierarchical semantic-visual adaptation (HSVA) framework. Specifically, HSVA aligns the semantic and visual domains by adopting a hierarchical two-step adaptation, i.e., structure adaptation and distribution adaptation. In the structure adaptation step, we take two task-specific encoders to encode the source data (visual domain) and the target data (semantic domain) into a structure-aligned common space. To this end, a supervised adversarial discrepancy (SAD) module is proposed to adversarially minimize the discrepancy between the predictions of two task-specific classifiers, thus making the visual and semantic feature manifolds more closely aligned. In the distribution adaptation step, we directly minimize the Wasserstein distance between the latent multivariate Gaussian distributions to align the visual and semantic distributions using a common encoder. Finally, the structure and distribution adaptation are derived in a unified framework under two partially-aligned variational autoencoders. Extensive experiments on four benchmark datasets demonstrate that HSVA achieves superior performance on both conventional and generalized ZSL. The code is available at \url{https://github.com/shiming-chen/HSVA}. Shiming Chen 0002, Guosen Xie, Yang Liu 0069, Qinmu Peng, Baigui Sun, Hao Li 0030, Xinge You, Ling Shao 0001 |
NeurIPS | 2 |
| 2021 | VMAN: A Virtual Mainstay Alignment Network for Transductive Zero-Shot LearningabstractTransductive zero-shot learning (TZSL) extends conventional ZSL by leveraging (unlabeled) unseen images for model training. A typical method for ZSL involves learning embedding weights from the feature space to the semantic space. However, the learned weights in most existing methods are dominated by seen images, and can thus not be adapted to unseen images very well. In this paper, to align the (embedding) weights for better knowledge transfer between seen/unseen classes, we propose the virtual mainstay alignment network (VMAN), which is tailored for the transductive ZSL task. Specifically, VMAN is casted as a tied encoder-decoder net, thus only one linear mapping weights need to be learned. To explicitly learn the weights in VMAN, for the first time in ZSL, we propose to generate virtual mainstay (VM) samples for each seen class, which serve as new training data and can prevent the weights from being shifted to seen images, to some extent. Moreover, a weighted reconstruction scheme is proposed and incorporated into the model training phase, in both the semantic/feature spaces. In this way, the manifold relationships of the VM samples are well preserved. To further align the weights to adapt to more unseen images, a novel instance-category matching regularization is proposed for model re-training. VMAN is thus modeled as a nested minimization problem and is solved by a Taylor approximate optimization paradigm. In comprehensive evaluations on four benchmark datasets, VMAN achieves superior performances under the (Generalized) TZSL setting. Guosen Xie, Xu-Yao Zhang, Yazhou Yao, Zheng Zhang 0006, Fang Zhao 0006, Ling Shao 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Web-Supervised Network with Softly Update-Drop Training for Fine-Grained Visual ClassificationabstractLabeling objects at the subordinate level typically requires expert knowledge, which is not always available from a random annotator. Accordingly, learning directly from web images for fine-grained visual classification (FGVC) has attracted broad attention. However, the existence of noise in web images is a huge obstacle for training robust deep neural networks. In this paper, we propose a novel approach to remove irrelevant samples from the real-world web images during training, and only utilize useful images for updating the networks. Thus, our network can alleviate the harmful effects caused by irrelevant noisy web images to achieve better performance. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is much superior to state-of-the-art webly supervised methods. The data and source code of this work have been made anonymously available at: https://github.com/z337-408/WSNFGVC. Chuanyi Zhang, Yazhou Yao, Huafeng Liu 0004, Guosen Xie, Xiangbo Shu, Tianfei Zhou, Zheng Zhang 0006, Fumin Shen, Zhenmin Tang |
AAAI | 4 |
| 2020 | Region Graph Embedding Network for Zero-Shot Learning
Guosen Xie, Li Liu 0004, Fan Zhu 0001, Fang Zhao 0006, Zheng Zhang 0006, Yazhou Yao, Jie Qin 0004, Ling Shao 0001 |
ECCV (4) | 1 |
| 2020 | Unsupervised Domain Adaptation with Noise Resistible Mutual-Training for Person Re-identification
Fang Zhao 0006, Shengcai Liao, Guosen Xie, Jian Zhao 0006, Kaihao Zhang, Ling Shao 0001 |
ECCV (11) | 3 |
| 2020 | Classification Constrained Discriminator For Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptation for semantic segmentation aims to transfer knowledge from label-rich synthetic datasets to real-world images without any annotation. The traditional adversarial learning methods for domain adaptation learn to extract domain-invariant feature representations by aligning the feature distributions of both domains. However, these methods suffer from an imbalance in adversarial training and feature distortion. In this work, we propose a classification constrained discriminator to alleviate these problems. Specifically, we first propose to balance the adversarial training by eliminating any pooling layers or strided convolutions in the discriminator. Then, we propose to constrain the discriminator with an auxiliary classification loss to help the feature generator extract the domain-invariant features that are useful for segmentation rather than just ambiguous features to fool the domain discriminator. Extensive experiments demonstrate the superiority of our proposed approach. The source code and models have been made available at https://github.com/NUSTMachine-Intelligence-Laboratory/ccd. Tao Chen 0012, Jian Zhang 0002, Guosen Xie, Yazhou Yao, Xiaoshui Huang, Zhenmin Tang |
ICME | 3 |
| 2020 | CDIMC-net: Cognitive Deep Incomplete Multi-view Clustering NetworkabstractIn recent years, incomplete multi-view clustering, which studies the challenging multi-view clustering problem on missing views, has received growing research interests. Although a series of methods have been proposed to address this issue, the following problems still exist: 1) Almost all of the existing methods are based on shallow models, which is difficult to obtain discriminative common representations. 2) These methods are generally sensitive to noise or outliers since the negative samples are treated equally as the important samples. In this paper, we propose a novel incomplete multi-view clustering network, called Cognitive Deep Incomplete Multi-view Clustering Network (CDIMC-net), to address these issues. Specifically, it captures the high-level features and local structure of each view by incorporating the view-specific deep encoders and graph embedding strategy into a framework. Moreover, based on the human cognition, \emph{i.e.}, learning from easy to hard, it introduces a self-paced strategy to select the most confident samples for model training, which can reduce the negative influence of outliers. Experimental results on several incomplete datasets show that CDIMC-net outperforms the state-of-the-art incomplete multi-view clustering methods. Jie Wen 0001, Zheng Zhang 0006, Yong Xu 0001, Bob Zhang 0001, Lunke Fei, Guosen Xie |
IJCAI | 6 |
| 2020 | Discriminative margin-sensitive autoencoder for collective multi-view disease analysis
Zheng Zhang 0006, Qi Zhu 0001, Guosen Xie, Yi Chen 0023, Shuihua Wang |
Neural Networks | 3 |
| 2020 | SRSC: Selective, Robust, and Supervised Constrained Feature Representation for Image ClassificationabstractFeature representation learning, an emerging topic in recent years, has achieved great progress. Powerful learned features can lead to excellent classification accuracy. In this article, a selective and robust feature representation framework with a supervised constraint (SRSC) is presented. SRSC seeks a selective, robust, and discriminative subspace by transforming the original feature space into the category space. Particularly, we add a selective constraint to the transformation matrix (or classifier parameter) that can select discriminative dimensions of the input samples. Moreover, a supervised regularization is tailored to further enhance the discriminability of the subspace. To relax the hard zero-one label matrix in the category space, an additional error term is also incorporated into the framework, which can lead to a more robust transformation matrix. SRSC is formulated as a constrained least square learning (feature transforming) problem. For the SRSC problem, an inexact augmented Lagrange multiplier method (ALM) is utilized to solve it. Extensive experiments on several benchmark data sets adequately demonstrate the effectiveness and superiority of the proposed method. The proposed SRSC approach has achieved better performances than the compared counterpart methods. Guosen Xie, Zheng Zhang 0006, Li Liu 0004, Fan Zhu 0001, Xu-Yao Zhang, Ling Shao 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Exploiting Web Images for Multi-Output Classification: From Category to SubcategoriesabstractStudies present that dividing categories into subcategories contributes to better image classification. Existing image subcategorization works relying on expert knowledge and labeled images are both time-consuming and labor-intensive. In this article, we propose to select and subsequently classify images into categories and subcategories. Specifically, we first obtain a list of candidate subcategory labels from untagged corpora. Then, we purify these subcategory labels through calculating the relevance to the target category. To suppress the search error and noisy subcategory label-induced outlier images, we formulate outlier images removing and the optimal classification models learning as a unified problem to jointly learn multiple classifiers, where the classifier for a category is obtained by combining multiple subcategory classifiers. Compared with the existing subcategorization works, our approach eliminates the dependence on expert knowledge and labeled images. Extensive experiments on image categorization and subcategorization demonstrate the superiority of our proposed approach. Yazhou Yao, Fumin Shen, Guosen Xie, Li Liu 0004, Fan Zhu 0001, Jian Zhang 0002, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | SADIH: Semantic-Aware DIscrete HashingabstractDue to its low storage cost and fast query speed, hashing has been recognized to accomplish similarity search in largescale multimedia retrieval applications. Particularly, supervised hashing has recently received considerable research attention by leveraging the label information to preserve the pairwise similarities of data points in the Hamming space. However, there still remain two crucial bottlenecks: 1) the learning process of the full pairwise similarity preservation is computationally unaffordable and unscalable to deal with big data; 2) the available category information of data are not well-explored to learn discriminative hash functions. To overcome these challenges, we propose a unified Semantic-Aware DIscrete Hashing (SADIH) framework, which aims to directly embed the transformed semantic information into the asymmetric similarity approximation and discriminative hashing function learning. Specifically, a semantic-aware latent embedding is introduced to asymmetrically preserve the full pairwise similarities while skillfully handle the cumbersome n×n pairwise similarity matrix. Meanwhile, a semantic-aware autoencoder is developed to jointly preserve the data structures in the discriminative latent semantic space and perform data reconstruction. Moreover, an efficient alternating optimization algorithm is proposed to solve the resulting discrete optimization problem. Extensive experimental results on multiple large-scale datasets demonstrate that our SADIH can clearly outperform the state-of-the-art baselines with the additional benefit of lower computational costs. Zheng Zhang 0006, Guosen Xie, Yang Li 0140, Sheng Li 0001, Zi Huang |
AAAI | 2 |
| 2019 | Attentive Region Embedding Network for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to classify images from unseen categories, by merely utilizing seen class images as the training data. Existing works on ZSL mainly leverage the global features or learn the global regions, from which, to construct the embeddings to the semantic space. However, few of them study the discrimination power implied in local image regions (parts), which, in some sense, correspond to semantic attributes, have stronger discrimination than attributes, and can thus assist the semantic transfer between seen/unseen classes. In this paper, to discover (semantic) regions, we propose the attentive region embedding network (AREN), which is tailored to advance the ZSL task. Specifically, AREN is end-to-end trainable and consists of two network branches, i.e., the attentive region embedding (ARE) stream, and the attentive compressed second-order embedding (ACSE) stream. ARE is capable of discovering multiple part regions under the guidance of the attention and the compatibility loss. Moreover, a novel adaptive thresholding mechanism is proposed for suppressing redundant (such as background) attention regions. To further guarantee more stable semantic transfer from the perspective of second-order collaboration, ACSE is incorporated into the AREN. In the comprehensive evaluations on four benchmarks, our models achieve state-of-the-art performances under ZSL setting, and compelling results under generalized ZSL setting. Guosen Xie, Li Liu 0004, Xiao-Bo Jin, Fan Zhu 0001, Zheng Zhang 0006, Jie Qin 0004, Yazhou Yao, Ling Shao 0001 |
CVPR | 1 |
| 2019 | Fusion linear representation-based classification
Guosen Xie, Jiexin Pu |
Soft Comput. | 2 |
| 2019 | Scalable Supervised Asymmetric Hashing With Semantic and Latent Factor EmbeddingabstractCompact hash code learning has been widely applied to fast similarity search owing to its significantly reduced storage and highly efficient query speed. However, it is still a challenging task to learn discriminative binary codes for perfectly preserving the full pairwise similarities embedded in the high-dimensional real-valued features, such that the promising performance can be guaranteed. To overcome this difficulty, in this paper, we propose a novel scalable supervised asymmetric hashing (SSAH) method, which can skillfully approximate the full-pairwise similarity matrix based on maximum asymmetric inner product of two different non-binary embeddings. In particular, to comprehensively explore the semantic information of data, the supervised label information and the refined latent feature embedding are simultaneously considered to construct the high-quality hashing function and boost the discriminant of the learned binary codes. Specifically, SSAH learns two distinctive hashing functions in conjunction of minimizing the regression loss on the semantic label alignment and the encoding loss on the refined latent features. More importantly, instead of using only part of similarity correlations of data, the full-pairwise similarity matrix is directly utilized to avoid information loss and performance degeneration, and its cumbersome computation complexity on n ×n matrix can be dexterously manipulated during the optimization phase. Furthermore, an efficient alternating optimization scheme with guaranteed convergence is designed to address the resulting discrete optimization problem. The encouraging experimental results on diverse benchmark datasets demonstrate the superiority of the proposed SSAH method in comparison with many recently proposed hashing algorithms. Zheng Zhang 0006, Zhihui Lai 0001, Zi Huang, Wai Keung Wong, Guosen Xie, Li Liu 0004, Ling Shao 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | Scene classification-oriented saliency detection via the modularized prescription
Chunlei Yang, Jiexin Pu, Yongsheng Dong 0004, Guosen Xie, Yanna Si |
Vis. Comput. | 4 |
| 2018 | Video super-resolution based on spatial-temporal recurrent residual networks
Wenhan Yang, Jiashi Feng, Guosen Xie, Jiaying Liu 0001, Zongming Guo, Shuicheng Yan |
Comput. Vis. Image Underst. | 3 |
| 2018 | Approximately optimizing NDCG using pair-wise loss
Xiao-Bo Jin, Guanggang Geng, Guosen Xie, Kaizhu Huang |
Inf. Sci. | 3 |
| 2018 | Hybrid of extended locality-constrained linear coding and manifold ranking for salient object detection
Chunlei Yang, Xiangluo Wang, Jiexin Pu, Guosen Xie, Yongsheng Dong 0004, Lingfei Liang |
J. Vis. Commun. Image Represent. | 4 |
| 2017 | SDE: A Novel Selective, Discriminative and Equalizing Feature Representation for Visual Recognition
Guosen Xie, Xu-Yao Zhang, Shuicheng Yan, Cheng-Lin Liu 0001 |
Int. J. Comput. Vis. | 1 |
| 2017 | LG-CNN: From local parts to global discrimination for fine-grained recognition
Guosen Xie, Xu-Yao Zhang, Wenhan Yang, Mingliang Xu 0001, Shuicheng Yan, Cheng-Lin Liu 0001 |
Pattern Recognit. | 1 |
| 2017 | Extended Locality-Constrained Linear Self-Coding for Saliency DetectionabstractIn complex scenes, foreground saliency can hardly be detected completely, which may further result in the ambiguous cues of objects for other computer vision tasks. In this letter, an extended locality-constrained linear self-coding (eLLsC) scheme is proposed to assist to solve the saliency detection problem under the complex scenes. The locality of both spatial relation and feature distance is preserved in eLLsC, thus making the transformed code involved in the manifold ranking to prompt the generation of the saliency map with more complete foreground and clearer boundary. Experimental results on three saliency detection benchmarks demonstrate the effectiveness of the proposed hybrid method. Chunlei Yang, Jiexin Pu, Guosen Xie, Yongsheng Dong 0004 |
IEEE Signal Process. Lett. | 3 |
| 2017 | Hybrid CNN and Dictionary-Based Models for Scene Recognition and Domain AdaptationabstractConvolutional neural network (CNN) has achieved the state-of-the-art performance in many different visual tasks. Learned from a large-scale training data set, CNN features are much more discriminative and accurate than the handcrafted features. Moreover, CNN features are also transferable among different domains. On the other hand, traditional dictionary-based features (such as BoW and spatial pyramid matching) contain much more local discriminative and structural information, which is implicitly embedded in the images. To further improve the performance, in this paper, we propose to combine CNN with dictionary-based models for scene recognition and visual domain adaptation (DA). Specifically, based on the well-tuned CNN models (e.g., AlexNet and VGG Net), two dictionary-based representations are further constructed, namely, mid-level local representation (MLR) and convolutional Fisher vector (CFV) representation. In MLR, an efficient two-stage clustering method, i.e., weighted spatial and feature space spectral clustering on the parts of a single image followed by clustering all representative parts of all images, is used to generate a class-mixture or a class-specific part dictionary. After that, the part dictionary is used to operate with the multiscale image inputs for generating mid-level representation. In CFV, a multiscale and scale-proportional Gaussian mixture model training strategy is utilized to generate Fisher vectors based on the last convolutional layer of CNN. By integrating the complementary information of MLR, CFV, and the CNN features of the fully connected layer, the state-of-the-art performance can be achieved on scene recognition and DA problems. An interested finding is that our proposed hybrid representation (from VGG net trained on ImageNet) is also complementary to GoogLeNet and/or VGG-11 (trained on Place205) greatly. Guosen Xie, Xu-Yao Zhang, Shuicheng Yan, Cheng-Lin Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | End-to-End Online Writer Identification With Recurrent Neural NetworkabstractWriter identification is an important topic for pattern recognition and artificial intelligence. Traditional methods rely heavily on sophisticated hand-crafted features to represent the characteristics of different writers. In this paper, we propose an end-to-end framework for online text-independent writer identification by using a recurrent neural network (RNN). Specifically, the handwriting data of a particular writer are represented by a set of random hybrid strokes (RHSs). Each RHS is a randomly sampled short sequence representing pen tip movements ($xy$-coordinates) and pen-down or pen-up states. RHS is independent of the content and language involved in handwriting; therefore, writer identification at the RHS level is more general and convenient than the character level or the word level, which also requires character/word segmentation. The RNN model with bidirectional long short-term memory is used to encode each RHS into a fixed-length vector for final classification. All the RHSs of a writer are classified independently, and then, the posterior probabilities are averaged to make the final decision. The proposed framework is end-to-end and does not require any domain knowledge for handwriting data analysis. Experiments on both English (133 writers) and Chinese (186 writers) databases verify the advantages of our method compared with other state-of-the-art approaches. Xu-Yao Zhang, Guosen Xie, Cheng-Lin Liu 0001, Yoshua Bengio |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2017 | Discriminative Elastic-Net Regularized Linear RegressionabstractIn this paper, we aim at learning compact and discriminative linear regression models. Linear regression has been widely used in different problems. However, most of the existing linear regression methods exploit the conventional zero-one matrix as the regression targets, which greatly narrows the flexibility of the regression model. Another major limitation of these methods is that the learned projection matrix fails to precisely project the image features to the target space due to their weak discriminative capability. To this end, we present an elastic-net regularized linear regression (ENLR) framework, and develop two robust linear regression models which possess the following special characteristics. First, our methods exploit two particular strategies to enlarge the margins of different classes by relaxing the strict binary targets into a more feasible variable matrix. Second, a robust elastic-net regularization of singular values is introduced to enhance the compactness and effectiveness of the learned projection matrix. Third, the resulting optimization problem of ENLR has a closed-form solution in each iteration, which can be solved efficiently. Finally, rather than directly exploiting the projection matrix for recognition, our methods employ the transformed features as the new discriminate representations to make final image classification. Compared with the traditional linear regression model and some of its variants, our method is much more accurate in image classification. Extensive experiments conducted on publicly available data sets well demonstrate that the proposed framework can outperform the state-of-the-art methods. The MATLAB codes of our methods can be available at http://www.yongxu.org/lunwen.html. Zheng Zhang 0006, Zhihui Lai 0001, Yong Xu 0001, Ling Shao 0001, Guosen Xie |
IEEE Trans. Image Process. | 6 |
| 2016 | Computational Face Reader
Xiangbo Shu, Liyan Zhang 0001, Jinhui Tang 0001, Guosen Xie, Shuicheng Yan |
MMM (1) | 4 |
| 2016 | Age progression: Current technologies and applications
Xiangbo Shu, Guosen Xie, Zechao Li, Jinhui Tang 0001 |
Neurocomputing | 2 |
| 2015 | Task-Driven Feature Pooling for Image ClassificationabstractFeature pooling is an important strategy to achieve high performance in image classification. However, most pooling methods are unsupervised and heuristic. In this paper, we propose a novel task-driven pooling (TDP) model to directly learn the pooled representation from data in a discriminative manner. Different from the traditional methods (e.g., average and max pooling), TDP is an implicit pooling method which elegantly integrates the learning of representations into the given classification task. The optimization of TDP can equalize the similarities between the descriptors and the learned representation, and maximize the classification accuracy. TDP can be combined with the traditional BoW models (coding vectors) or the recent state-of-the-art CNN models (feature maps) to achieve a much better pooled representation. Furthermore, a self-training mechanism is used to generate the TDP representation for a new test image. A multi-task extension of TDP is also proposed to further improve the performance. Experiments on three databases (Flower-17, Indoor-67 and Caltech-101) well validate the effectiveness of our models. Guosen Xie, Xu-Yao Zhang, Xiangbo Shu, Shuicheng Yan, Cheng-Lin Liu 0001 |
ICCV | 1 |
| 2014 | Efficient Feature Coding Based on Auto-encoder Network for Image Classification
Guosen Xie, Xu-Yao Zhang, Cheng-Lin Liu 0001 |
ACCV (1) | 1 |
| 2014 | Integrating supervised subspace criteria with restricted Boltzmann Machine for feature extractionabstractRestricted Boltzmann Machine (RBM) is a widely used building-block in deep neural networks. However, RBM is an unsupervised model which can not exploit the rich supervised information of data. Therefore, we consider combining the descriptive (generative) ability of RBM with the discriminative ability of supervised subspace models, i.e., Fisher linear discriminant analysis (FDA), marginal Fisher analysis (MFA), and heat kernel MFA (hkMFA). Specifically, the hidden layer of RBM is regularized by the supervised subspace criteria, and the joint learning model can then be efficiently optimized by gradient descent and graph construction (used to define the scatter matrix in the subspace models) on mini-batch data. Compared with the traditional subspace models (FDA, MFA, hkMFA), the proposed hybrid models are essentially nonlinear and can be optimized by gradient descent instead of eigenvalue decomposition. More importantly, traditional subspace models can only reduce the dimensionality (because of linear transformation), while the proposed models can also increase the dimensionality for better class discrimination. Experiments on three databases demonstrate that the proposed hybrid models outperform both RBM and their counterpart subspace models (FDA, MFA, hkMFA) consistently. Guosen Xie, Xu-Yao Zhang, Yan-Ming Zhang 0001, Cheng-Lin Liu 0001 |
IJCNN | 1 |