Yazhou Yao

dblp:173/6775 · DBLP profile ↗
← Back
126ranked-venue papers
17as first author
91since 2021 · last 2026
0000-0002-0337-9410ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 94 · 12 first-author · 68 since 2021Artificial intelligence and machine learning · 60 · 9 first-author · 43 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Foundation-Adaptive Integrated Refinement for Generalized Category Discovery
abstract
The potential of Generalized Category Discovery (GCD) lies in its ability to identify previously undiscovered patterns in both labeled and unlabeled data by leveraging insights from partially labeled training samples. However, interference can arise due to the model's dual focus on discovering both novel and known categories, often leading to conflicts that obscure true patterns in the dataset. This paper presents a divide-and-conquer framework, Foundation-Adaptive Integrated Refinement (FAIR), which fine-tunes pretrained foundational weights for various purposes, divided into Foundation (pretrained weights), Adaptive (weights fine-tuned with a variance-preserving loss), and Integrated (weights adjusted for both labeled and unlabeled data). The Adaptive utilizes a newly proposed adaptive contrastive loss that introduces variances within classes to preserve the individuality of representations. The Integrated addresses inherent estimation errors while dynamically estimating the number of categories, incorporating a cosine-based perturbation mechanism as a relaxed margin to accommodate potential ground-truth deviations, rather than relying on biased estimates. Extensive experiments on six benchmark datasets demonstrate our method's effectiveness, outperforming state-of-the-art algorithms, especially on fine-grained datasets.
Yuwei Bian, Yazhou Yao, Haofeng Zhang 0001
AAAI3
2026 AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs
abstract
Visual abductive reasoning (VAR) is a challenging task that requires AI systems to infer the most likely explanation for incomplete visual observations. While recent MLLMs develop strong general-purpose multimodal reasoning capabilities, they remain fall short in abductive inference, as compared to human beings. To bridge this gap, we draw inspiration from the interplay between verbal and pictorial abduction in human cognition, and propose to strengthen abduction of MLLMs by mimicking such dual-mode behavior. Concretely, we introduce AbductiveMLLM comprising of two synergistic components: REASONER and IMAGINER. The REASONER operates in the verbal domain. It first explores a broad space of possible explanations using a blind LLM and then prunes visually incongruent hypotheses based on cross-modal causal alignment. The remaining hypotheses are introduced into the MLLM as targeted priors, steering its reasoning toward causally coherent explanations. The IMAGINER, on the other hand, further guides MLLMs by emulating human-like pictorial thinking. It conditions a text-to-image diffusion model on both the input video and the REASONER’s output embeddings to “imagine” plausible visual scenes that correspond to verbal explanation, thereby enriching MLLMs' contextual grounding. The two components are trained jointly in an end-to-end manner. Experiments on standard VAR benchmarks show that AbductiveMLLM achieves state-of-the-art performance, consistently outperforming traditional solutions and advanced MLLMs.
Boyu Chang, Qi Wang 0009, Zhixiong Nan, Yazhou Yao, Tianfei Zhou
AAAI5
2026 Beyond Quadratic: Linear-Time Change Detection with RWKV
abstract
Existing paradigms for remote sensing change detection are caught in a trade-off: CNNs excel at efficiency but lack global context, while Transformers capture long-range dependencies at a prohibitive computational cost. This paper introduces ChangeRWKV, a new architecture that reconciles this conflict. By building upon the Receptance Weighted Key Value (RWKV) framework, our ChangeRWKV uniquely combines the parallelizable training of Transformers with the linear-time inference of RNNs. Our approach core features two key innovations: a hierarchical RWKV encoder that builds multi-resolution feature representation, and a novel Spatial-Temporal Fusion Module (STFM) engineered to resolve spatial misalignments across scales while distilling fine-grained temporal discrepancies. ChangeRWKV not only achieves state-of-the-art performance on the LEVIR-CD benchmark, with an 85.46% IoU and 92.16% F1 score, but does so while drastically reducing parameters and FLOPs compared to previous leading methods. This work demonstrates a new, efficient, and powerful paradigm for operational-scale change detection.
Gensheng Pei, Tao Chen 0012, Xia Yuan, Haofeng Zhang 0001, Xiangbo Shu, Yazhou Yao
AAAI7
2026 Jo-SNC: Combating Noisy Labels Through Fostering Self- and Neighbor-Consistency
abstract
Label noise is pervasive in various real-world scenarios, posing challenges in supervised deep learning. Deep networks are vulnerable to such label-corrupted samples due to the memorization effect. One major stream of previous methods concentrates on identifying clean data for training. However, these methods often neglect imbalances in label noise across different mini-batches and devote insufficient attention to out-of-distribution noisy data. To this end, we propose a noise-robust method named Jo-SNC (Joint sample selection and model regularization based on Self- and Neighbor-Consistency). Specifically, we propose to employ the Jensen-Shannon divergence to measure the "likelihood" of a sample being clean or out-of-distribution. This process factors in the nearest neighbors of each sample to reinforce the reliability of clean sample identification. We design a self-adaptive, data-driven thresholding scheme to adjust per-class selection thresholds. While clean samples undergo conventional training, detected in-distribution and out-of-distribution noisy samples are trained following partial label learning and negative learning, respectively. Finally, we advance the model performance further by proposing a triplet consistency regularization that promotes self-prediction consistency, neighbor-prediction consistency, and feature consistency. Extensive experiments on various benchmark datasets and comprehensive ablation studies demonstrate the effectiveness and superiority of our approach over existing state-of-the-art methods.
Zeren Sun, Yazhou Yao, Tongliang Liu, Zechao Li, Fumin Shen, Jinhui Tang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 CAFCL: Class-aware flow-based contrastive learning for out-of-distribution detection
Shiyu Peng, Jingzhu Li, Mingtao Wei, Gensheng Pei, Yazhou Yao
Pattern Recognit.6
2026 Discriminative response pruning for robust and efficient deep networks under label noise
Shuwen Jin, Junzhu Mao, Zeren Sun, Yazhou Yao
Pattern Recognit. Lett.4
2026 DepMatch: Boosting Semi-Supervised Semantic Segmentation by Exploring Depth Difference Knowledge
abstract
Existing semi-supervised semantic segmentation (SSS) methods fail to explore the potential of depth information in unlabeled data, as they suffer from 1) inter-class depth similarity, and 2) intra-class depth discrepancy. To address these challenges, this paper proposes DepMatch, a simple yet effective approach that leverages depth difference knowledge to guide consistency learning. Specifically, a Class-wise Depth Disparity Perception (CDDP) module is designed to exploit depth difference information, driven by class prediction priors, facilitating robust feature learning. Depth-feature discrepancy set is first constructed and then reliable pixel pairs are selected for inter-class depth disparity knowledge distillation. Simultaneously, exponential normalization is applied to intra-category depth disparity for suppressing large outlier variations, and an entropy-based adaptive weight is derived to prioritize feature learning of high entropy areas. Moreover, we propose the Uncertain Logit Disparity Regulation (ULDR) module, which leverages the depth variations at class boundaries to promote the mutual regulation of uncertain pixel logit information, enhancing the model's spatial understanding. Experiments on five public benchmarks show that DepMatch can be seamlessly incorporated as a plug-and-play plugin into popular SSS frameworks, achieving significant performance improvements across various visual encoders. The source code and models are made available at https://github.com/NUST-Machine-Intelligence-Laboratory/DepMatch.
Jianjian Yin, Xiruo Jiang, Tao Chen 0012, Gensheng Pei, Yazhou Yao, Fumin Shen, Heng Tao Shen
IEEE Trans. Image Process.5
2026 Coarse Labels Matter: Revisiting the Role of Coarse-Grained Supervision in Fine-Grained Learning
abstract
The prohibitive cost of acquiring high-quality fine-grained annotations has spurred significant interest in leveraging readily available coarse labels for fine-grained learning. However, prevailing approaches tend to rely on increasingly sophisticated unsupervised methods to define fine-grained proxy tasks, with coarse labels often playing an auxiliary role. In this paper, we propose CSer, a framework designed to maximize the utility of coarse label information for Coarse-to-Fine learning. Specifically, to reconcile the conflict between preserving fine-grained feature diversity and maintaining strong coarse-grained supervision, our coarse-grained self-distillation strategy fortifies the backbone's discriminative power by distilling knowledge from the final classifier to intermediate layers. Concurrently, we introduce dense supervision on common component features within each coarse class, which are decoupled using Non-negative Matrix Factorization. This enhances responses to distinct components, thereby mitigating the simplicity bias in embeddings that can arise under coarse supervision. Moreover, we leverage relationships among intra-class samples to dynamically adjust the negative sampling strategy in contrastive learning, thereby constructing distinct fine-grained class relationships tailored to different coarse classes. Extensive experiments conducted on multiple benchmark datasets demonstrate the effectiveness of our method, yielding state-of-the-art results surpassing competing methods.
Xin-Yang Zhao, Pengyuan Zhang, Qiyuan Zhuang, Yazhou Yao, Xiu-Shen Wei
IEEE Trans. Image Process.4
2026 Fusion of Infrared and Visible Images Based on Iterative Dual-Branch Attention and Modality Discrepancy Guidance
abstract
The fusion of infrared and visible images aims to generate images that provide a more comprehensive description of the scene. Convolutional neural network and transformer are two commonly used methods in image fusion. The former focuses on extracting local features but lacks a global receptive field, and the latter can extract global information but ignores the discrepancy information between modalities. In view of this, we propose an efficient fusion network based on iterative dual-branch attention and modality discrepancy guidance (IDMDNet), which consists of three components. In the first component, we use shallow and deep feature extraction modules to extract features. In the second component, we first design an iterative dual-branch attention module to capture important features of each modality, thereby preserving important information. Secondly, to model the global information of the source image, we introduce a transformer module to achieve the fusion of shallow global information. Further, we design a modality discrepancy guided fusion module to promote the fusion of modality discrepancy information and ultimately achieve modality information complementation. In the third component, we introduce an invertible neural network to reduce the loss of feature information, thus achieving high quality image reconstruction. Finally, we construct a loss function for IDMDNet that includes intensity, gradient, and multiscale structure to motivate the network to preserve texture and target details. Experiments on infrared and visible image fusion benchmark datasets show that the proposed IDMDNet has competitive fusion performance. Remarkably, it can be effectively applied to other infrared and visible datasets without the need for fine-tuning, highlighting its good generalization ability. The source code of IDMDNet has been released athttps://github.com/AHUT-MILAGroup/IDMDNet.
Shuting Zhu, Yazhou Yao, Ping Zhong 0001
IEEE Trans. Multim.3
2025 3D-aware Select, Expand, and Squeeze Token for Aerial Action Recognition
abstract
Aerial Action Recognition (AAR) in videos captured by Unmanned Aerial Vehicles (UAVs) plays a vital role in numerous applications. However, current methods related to traditional action recognition primarily cater to fixed or near cameras, and rarely consider the movement disturbance of UAVs, including their varying attitudes and positions. Those characteristics of aerial videos bring moving objects in small regions compared to broad backgrounds and relative movement to the motion of objects, which reflect more sparse and disturbed semantic information for AAR. To address these issues, we present a novel framework, dubbed 3D-Tok, to Select, Expand, and Squeeze original visual tokens for obtaining compact yet diverse semantic-enhanced tokens. In particular, we present a 3D-token selector (3TS) to select complex yet diverse tokens in three channels, capturing the semantic awareness of moving objects in comparatively small regions. Additionally, to get rid of disturbed semantic information caused by the UAV flight, we present an Expand-Squeeze Converter (ESC) to adaptively expand and squeeze the 3D-selected tokens constrained by contrastive loss, thereby suppressing the semantic-irrelevant information and reinforce semantic-relevant information via the interpolation converting. By involving the token selecting, expanding, and squeezing into an all-in-one framework, 3D-Tok shows significant improvements on the UAV-Human dataset(↑9.5%), RoCoG-v2 dataset (↑23.5%), and Drone-Action dataset (↑5.7%).
Luying Peng, Xiangbo Shu, Yazhou Yao, Guosen Xie
AAAI3
2025 Twofold Debiasing Enhances Fine-Grained Learning with Coarse Labels
abstract
The Coarse-to-Fine Few-Shot (C2FS) task is designed to train models using only coarse labels, then leverages a limited number of subclass samples to achieve fine-grained recognition capabilities. This task presents two main challenges: coarse-grained supervised pre-training suppresses the extraction of critical fine-grained features for subcategory discrimination, and models suffer from overfitting due to biased distributions caused by limited fine-grained samples. In this paper, we propose the Twofold Debiasing (TFB) method, which addresses these challenges through detailed feature enhancement and distribution calibration. Specifically, we introduce a multi-layer feature fusion reconstruction module and an intermediate layer feature alignment module to combat the model's tendency to focus on simple predictive features directly related to coarse-grained supervision, while neglecting complex fine-grained level details. Furthermore, we mitigate the biased distributions learned by the fine-grained classifier using readily available coarse-grained sample embeddings enriched with fine-grained information. Extensive experiments conducted on five benchmark datasets demonstrate the efficacy of our approach, achieving state-of-the-art results that surpass competitive methods.
Xin-yang Zhao, Yazhou Yao
AAAI4
2025 Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection
abstract
The CLIP model has demonstrated significant advancements in aligning visual and language modalities through large-scale pre-training on image-text pairs, enabling strong zero-shot classification and retrieval capabilities on various domains. However, CLIP’s training remains computationally intensive, with high demands on both data processing and memory. To address these challenges, recent masking strategies have emerged, focusing on the selective removal of image patches to improve training efficiency. Although effective, these methods often compromise key semantic information, resulting in suboptimal alignment between visual features and text descriptions. In this work, we present a concise yet effective approach called Patch Generation-to-Selection (CLIP-PGS) to enhance CLIP’s training efficiency while preserving critical semantic con tent. Our method introduces a gradual masking process in which a small set of candidate patches is first pre-selected as potential mask regions. Then, we apply Sobel edge detection across the entire image to generate an edge mask that prioritizes the retention of the primary object areas. Finally, similarity scores between the candidate mask patches and their neighboring patches are computed, with optimal transport normalization refining the selection process to ensure a balanced similarity matrix. Our approach, CLIP-PGS, sets new state-of-the-art results in zero-shot classification and retrieval tasks, achieving superior performance in robustness evaluation and language compositionality benchmarks.
Gensheng Pei, Tao Chen 0012, Xinhao Cai, Xiangbo Shu, Tianfei Zhou, Yazhou Yao
CVPR7
2025 UNIALIGN: Scaling Multimodal Alignment within One Unified Model
abstract
We present UniAlign, a unified model to align an arbitrary number of modalities (e.g., image, text, audio, 3D point cloud, etc.) through one encoder and a single training phase. Existing solutions typically employ distinct encoders for each modality, resulting in increased parameters as the number of modalities grows. In contrast, UniAlign proposes a modality-aware adaptation of the powerful mixture- of-experts (MoE) schema and further integrates it with Low- Rank Adaptation (LoRA), efficiently scaling the encoder to accommodate inputs in diverse modalities while maintaining a fixed computational overhead. Moreover, prior work often requires separate training for each extended modality. This leads to task-specific models and further hinders the communication between modalities. To address this, we propose a soft modality binding strategy that aligns all modalities using unpaired data samples across datasets. Two additional training objectives are introduced to distill knowledge from well-aligned anchor modalities and prior multimodal models, elevating UniAlign into a high performance multimodal foundation model. Experiments on 11 benchmarks across 6 different modalities demonstrate that UniAlign could achieve comparable performance to SOTA approaches, while using merely 7.8M trainable parameters and maintaining an identical model with the same weight across all tasks.
Liulei Li, Huafeng Liu 0004, Yazhou Yao, Wenguan Wang
CVPR5
2025 Cycle-Consistent Learning for Joint Layout-to-Image Generation and Object Detection
Xinhao Cai, Qiuxia Lai, Gensheng Pei, Xiangbo Shu, Yazhou Yao, Wenguan Wang
ICCV5
2025 Exploiting Frequency Dynamics for Enhanced Multimodal Event-Based Action Recognition
Meiqi Cao, Xiangbo Shu, Xin Jiang 0010, Rui Yan 0010, Yazhou Yao, Jinhui Tang 0001
ICCV5
2025 Tensor-Aggregated LoRA in Federated Fine-Tuning
Binqian Xu, Xiangbo Shu, Jiachao Zhang, Yazhou Yao, Guosen Xie, Jinhui Tang 0001
ICCV5
2025 CA2C: A Prior-Knowledge-Free Approach for Robust Label Noise Learning via Asymmetric Co-Learning and Co-Training
Mengmeng Sheng, Zeren Sun, Tianfei Zhou, Xiangbo Shu, Jinshan Pan, Yazhou Yao
ICCV6
2025 Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization
abstract
Computer Vision (CV) has yet to fully achieve the zero-shot task generalization observed in Natural Language Processing (NLP), despite following many of the milestones established in NLP, such as large transformer models, extensive pre-training, and the auto-regression paradigm, among others. In this paper, we rethink the reality that CV adopts discrete and terminological task definitions (e.g., "image segmentation"), and conjecture it is a key barrier that hampers zero-shot task generalization. Our hypothesis is that without truly understanding previously-seen tasks—due to these terminological definitions—deep models struggle to generalize to novel tasks. To verify this, we introduce Explanatory Instructions, which provide an intuitive way to define CV task objectives through detailed linguistic transformations from input images to outputs. We create a large-scale dataset comprising 12 million "image input $\to$ explanatory instruction $\to$ output" triplets, and train an auto-regressive-based vision-language model (AR-based VLM) that takes both images and explanatory instructions as input. By learning to follow these instructions, the AR-based VLM achieves instruction-level zero-shot capabilities for previously-seen tasks and demonstrates strong zero-shot generalization for unseen CV tasks. Code and dataset will be open-sourced.
Yang Shen 0006, Xiu-Shen Wei, Yifan Sun 0003, YuXin Song 0001, Heyang Xu, Yazhou Yao, Errui Ding
ICML8
2025 OmniGaze: Reward-inspired Generalizable Gaze Estimation in the Wild
abstract
Current 3D gaze estimation methods struggle to generalize across diverse data domains, primarily due to $\textbf{i)}$ $\textit{the scarcity of annotated datasets}$, and $\textbf{ii)}$ $\textit{the insufficient diversity of labeled data}$. In this work, we present OmniGaze, a semi-supervised framework for 3D gaze estimation, which utilizes large-scale unlabeled data collected from diverse and unconstrained real-world environments to mitigate domain bias and generalize gaze estimation in the wild. First, we build a diverse collection of unlabeled facial images, varying in facial appearances, background environments, illumination conditions, head poses, and eye occlusions. In order to leverage unlabeled data spanning a broader distribution, OmniGaze adopts a standard pseudo-labeling strategy and devises a reward model to assess the reliability of pseudo labels. Beyond pseudo labels as 3D direction vectors, the reward model also incorporates visual embeddings extracted by an off-the-shelf visual encoder and semantic cues from gaze perspective generated by prompting a Multimodal Large Language Model to compute confidence scores. Then, these scores are utilized to select high-quality pseudo labels and weight them for loss computation. Extensive experiments demonstrate that OmniGaze achieves state-of-the-art performance on five datasets under both in-domain and cross-domain settings. Furthermore, we also evaluate the efficacy of OmniGaze as a scalable data engine for gaze estimation, which exhibits robust zero-shot generalization on four unseen datasets.
Hongyu Qu, Jianan Wei, Xiangbo Shu, Yazhou Yao, Wenguan Wang, Jinhui Tang 0001
NeurIPS4
2025 You Only Communicate Once: One-shot Federated Low-Rank Adaptation of MLLM
abstract
Multimodal Large Language Models (MLLMs) with Federated Learning (FL) can quickly adapt to privacy-sensitive tasks, but face significant challenges such as high communication costs and increased attack risks, due to their reliance on multi-round communication. To address this, One-shot FL (OFL) has emerged, aiming to complete adaptation in a single client-server communication. However, existing adaptive ensemble OFL methods still need more than one round of communication, because correcting heterogeneity-induced local bias relies on aggregated global supervision, meaning they still do not achieve true one-shot communication. In this work, we make the first attempt to achieve true one-shot communication for MLLMs under OFL, by investigating whether implicit (i.e., initial rather than aggregated) global supervision alone can effectively correct local training bias. Our key finding from the empirical study is that imposing directional supervision on local training substantially mitigates client conflicts and local bias. Building on this insight, we propose YOCO, in which directional supervision with sign-regularized LoRA B enforces global consistency, while sparsely regularized LoRA A preserves client-specific adaptability. Experiments demonstrate that YOCO cuts communication to $\sim$0.03\% of multi-round FL while surpassing those methods in several multimodal scenarios and consistently outperforming all one-shot competitors.
Binqian Xu, Haiyang Mei, Zechen Bai, Jinjin Gong, Rui Yan 0010, Guosen Xie, Yazhou Yao, Basura Fernando, Xiangbo Shu
NeurIPS7
2025 Controllable text-to-3D multi-object generation via integrating layout and multiview patterns
Shaorong Sun, Shuchao Pang, Yazhou Yao, Xiaoshui Huang
Comput. Graph.3
2025 An Empirical Study on Training Paradigms for Deep Supervised Hashing
Yang Shen 0006, Peng Wang 0023, Xiu-Shen Wei, Yazhou Yao
Int. J. Comput. Vis.4
2025 ChangeTitans: Toward Remote Sensing Change Detection With Neural Memory
abstract
Remote sensing change detection is essential for environmental monitoring, urban planning, and related applications. However, current methods often struggle to capture long-range dependencies while maintaining computational efficiency. Although Transformers can effectively model global context, their quadratic complexity poses scalability challenges, and existing linear attention approaches frequently fail to capture intricate spatiotemporal relationships. Drawing inspiration from the recent success of Titans in language tasks, we present ChangeTitans, the Titans-based framework for remote sensing change detection. Specifically, we propose VTitans, the first Titans-based vision backbone that integrates neural memory with segmented local attention, thereby capturing long-range dependencies while mitigating computational overhead. Next, we present a hierarchical VTitans-Adapter to refine multi-scale features across different network layers. Finally, we introduce TS-CBAM, a two-stream fusion module leveraging cross-temporal attention to suppress pseudo-changes and enhance detection accuracy. Experimental evaluations on four benchmark datasets (LEVIR-CD, WHU-CD, LEVIR-CD+, and SYSU-CD) demonstrate that ChangeTitans achieves state-of-the-art results, attaining 84.36% IoU and 91.52% F1-score on LEVIR-CD, while remaining computationally competitive. Our code and model are available at https://github.com/ChangeTitans/ChangeTitans.
Gensheng Pei, Yazhou Yao, Tianfei Zhou, Lizhong Ding 0003, Fumin Shen
IEEE Trans. Geosci. Remote. Sens.3
2025 NiCI-Pruning: Enhancing Diffusion Model Pruning via Noise in Clean Image Guidance
abstract
The substantial successes achieved by diffusion probabilistic models have prompted the study of their employment in resource-limited scenarios. Pruning methods have been proven effective in compressing discriminative models relying on the correlation between training losses and model performances. However, diffusion models employ an iterative process for generating high-quality images, leading to a breakdown of such connections. To address this challenge, we propose a simple yet effective method, named NiCI-Pruning (Noise in Clean Image Pruning), for the compression of diffusion models. NiCI-Pruning capitalizes the noise predicted by the model based on clean image inputs, favoring it as a feature for establishing reconstruction losses. Accordingly, Taylor expansion is employed for the proposed reconstruction loss to evaluate the parameter importance effectively. Moreover, we propose an interval sampling strategy that incorporates a timestep-weighted schema, alleviating the risk of misleading information obtained at later timesteps. We provide comprehensive experimental results to affirm the superiority of our proposed approach. Notably, our method achieves a remarkable average reduction of 30.4% in FID score increase across five different datasets compared to the state-of-the-art diffusion pruning method at equivalent pruning rates. Our code and models have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/NiCI-Pruning.
Junzhu Mao, Zeren Sun, Yazhou Yao, Tianfei Zhou, Liqiang Nie, Xian-Sheng Hua 0001
IEEE Trans. Image Process.3
2025 Prune and Merge: Efficient Token Compression for Vision Transformer With Spatial Information Preserved
abstract
Token compression is essential for reducing the computational and memory requirements of transformer models, enabling their deployment in resource-constrained environments. In this work, we propose an efficient and hardware-compatible token compression method called Prune and Merge. Our approach integrates token pruning and merging operations within transformer models to achieve layer-wise token compression. By introducing trainable merge and reconstruct matrices and utilizing shortcut connections, we efficiently merge tokens while preserving important information and enabling the restoration of pruned tokens. Additionally, we introduce a novel gradient-weighted attention scoring mechanism that computes token importance scores during the training phase, eliminating the need for separate computations during inference and enhancing compression efficiency. We also leverage gradient information to capture the global impact of tokens and automatically identify optimal compression structures. Extensive experiments on the ImageNet-1 k and ADE20 K datasets validate the effectiveness of our approach, achieving significant speed-ups with minimal accuracy degradation compared to state-of-the-art methods. For instance, on DeiT-Small, we achieve a 1.64× speed-up with only a 0.2% drop in accuracy on ImageNet-1k. Moreover, by compressing segmenter models and comparing with existing methods, we demonstrate the superior performance of our approach in terms of efficiency and effectiveness.
Junzhu Mao, Yang Shen 0006, Jinyang Guo 0002, Yazhou Yao, Xian-Sheng Hua 0001, Heng Tao Shen
IEEE Trans. Multim.4
2025 Semi-Supervised Semantic Segmentation With Multi-Constraint Consistency Learning
abstract
Consistency regularization has prevailed in semi-supervised semantic segmentation and achieved promising performance. However, existing methods typically concentrate on enhancing the Image-augmentation based Prediction consistency and optimizing the segmentation network as a whole, resulting in insufficient utilization of potential supervisory information. In this paper, we propose a Multi-Constraint Consistency Learning (MCCL) approach to facilitate the staged enhancement of the encoder and decoder. Specifically, we first design a feature knowledge alignment (FKA) strategy to promote the feature consistency learning of the encoder from image-augmentation. Our FKA encourages the encoder to derive consistent features for strongly and weakly augmented views from the perspectives of point-to-point alignment and prototype-based intra-class compactness. Moreover, we propose a self-adaptive intervention (SAI) module to increase the discrepancy of aligned intermediate feature representations, promoting Feature-perturbation based Prediction consistency learning. Self-adaptive feature masking and noise injection are designed in an instance-specific manner to perturb the features for robust learning of the decoder. Experimental results on Pascal VOC2012 and Cityscapes datasets demonstrate that our proposed MCCL achieves new state-of-the-art performance. The source code and models are made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/MCCL.
Jianjian Yin, Tao Chen 0012, Gensheng Pei, Huafeng Liu 0004, Yazhou Yao, Liqiang Nie, Xian-Sheng Hua 0001
IEEE Trans. Multim.5
2024 Adaptive Integration of Partial Label Learning and Negative Learning for Enhanced Noisy Label Learning
abstract
There has been significant attention devoted to the effectiveness of various domains, such as semi-supervised learning, contrastive learning, and meta-learning, in enhancing the performance of methods for noisy label learning (NLL) tasks. However, most existing methods still depend on prior assumptions regarding clean samples amidst different sources of noise (e.g., a pre-defined drop rate or a small subset of clean samples). In this paper, we propose a simple yet powerful idea called NPN, which revolutionizes Noisy label learning by integrating Partial label learning (PLL) and Negative learning (NL). Toward this goal, we initially decompose the given label space adaptively into the candidate and complementary labels, thereby establishing the conditions for PLL and NL. We propose two adaptive data-driven paradigms of label disambiguation for PLL: hard disambiguation and soft disambiguation. Furthermore, we generate reliable complementary labels using all non-candidate labels for NL to enhance model robustness through indirect supervision. To maintain label reliability during the later stage of model training, we introduce a consistency regularization term that encourages agreement between the outputs of multiple augmentations. Experiments conducted on both synthetically corrupted and real-world noisy datasets demonstrate the superiority of NPN compared to other state-of-the-art (SOTA) methods. The source code has been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/NPN.
Mengmeng Sheng, Zeren Sun, Zhenhuang Cai, Tao Chen 0012, Yazhou Yao
AAAI6
2024 Poly Kernel Inception Network for Remote Sensing Detection
abstract
Object detection in remote sensing images (RSIs) often suffers from several increasing challenges, including the large variation in object scales and the diverse-ranging context. Prior methods tried to address these challenges by expanding the spatial receptive field of the backbone, either through large-kernel convolution or dilated convolution. However, the former typically introduces considerable background noise, while the latter risks generating overly sparse feature representations. In this paper, we introduce the Poly Kernel Inception Network (PKINet) to handle the above challenges. PKINet employs multi-scale convolution kernels without dilation to extract object features of varying scales and capture local context. In addition, a Context Anchor Attention (CAA) module is introduced in parallel to capture long-range contextual information. These two components work jointly to advance the performance of PKINet on four challenging remote sensing detection benchmarks, namely DOTA-v1.0, DOTA-v1.5, HRSC2016, and DIOR-R.
Xinhao Cai, Qiuxia Lai, Wenguan Wang, Zeren Sun, Yazhou Yao
CVPR6
2024 VideoMAC: Video Masked Autoencoders Meet ConvNets
abstract
Recently, the advancement of self-supervised learning techniques, like masked autoencoders (MAE), has greatly influenced visual representation learning for images and videos. Nevertheless, it is worth noting that the predomi-nant approaches in existing masked image / video modeling rely excessively on resource-intensive vision transformers (ViTs) as the feature encoder. In this paper, we propose a new approach termed as VideoMAC, which combines video masked autoencoders with resource-friendly Con-vNets. Specifically, VideoMAC employs symmetric masking on randomly sampled pairs of video frames. To prevent the issue of mask pattern dissipation, we utilize ConvNets which are implemented with sparse convolutional operators as en-coders. Simultaneously, we present a simple yet effective masked video modeling (MVM) approach, a dual encoder architecture comprising an online encoder and an exponential moving average target encoder, aimed to facilitate inter-frame reconstruction consistency in videos. Additionally, we demonstrate that VideoMAC, empowering classical (ResNet) / modern (ConvNeXt) convolutional encoders to harness the benefits of MVM, outperforms ViT-based approaches on downstream tasks, including video object segmentation (+5.2% /6.4% J&F), body part propagation (+6.3% /3.1% mIoU), and human pose tracking (+10.2% / 11.1% [email protected]).
Gensheng Pei, Tao Chen 0012, Xiruo Jiang, Huafeng Liu 0004, Zeren Sun, Yazhou Yao
CVPR6
2024 SMP-Track: SAM in Multi-Pedestrian Tracking
abstract
Multiple Object Tracking (MOT) plays a crucial role in security data analysis as a fundamental problem for video surveillance. Our goal is to design a robust tracker for data damaged by attacks, while also emphasizing privacy protection. The mainstream paradigm for MOT is tracking-by-detection (TBD), which involves object detection followed by target association. In the association stage, most models rely on Intersection over Union (IoU) similarity of bounding boxes for short-range matching and cosine similarity of appearance features for long-range matching. However, both of these similarities contain a lot of redundant background regions except the target. To this end, we propose a new tracker named SMP-Track that integrates Segment Anything Model (SAM) into Multi-Pedestrian tracking method. Firstly we extract the pedestrian masks based on box prompt, focusing solely on the foreground information. Then we introduce a new similarity metric that combines the advantages of motion and foreground information (i.e., box-mask similarity). Extensive experiments demonstrate that SMP-Track increases main metrics on the MOT17 validation set, and achieves comparable performance to other state-of-the-art methods on the MOT17 and MOT20 test sets. Furthermore, by incorporating pedestrian masks, we reduce reliance on raw pedestrian images or features, making the model robust to corrupted data and mitigating the risk of privacy leakage.
Shiyin Wang, Huafeng Liu 0004, Qiong Wang 0003, Yazhou Yao
DSAA5
2024 Knowledge Transfer with Simulated Inter-image Erasing for Weakly Supervised Semantic Segmentation
Tao Chen 0012, Xiruo Jiang, Gensheng Pei, Zeren Sun, Yucheng Wang 0013, Yazhou Yao
ECCV (42)6
2024 Veil Privacy on Visual Data: Concealing Privacy for Humans, Unveiling for DNNs
Shuchao Pang, Ruhao Ma, Bing Li 0002, Yongbin Zhou, Yazhou Yao
ECCV (83)5
2024 Foster Adaptivity and Balance in Learning with Noisy Labels
Mengmeng Sheng, Zeren Sun, Tao Chen 0012, Shuchao Pang, Yucheng Wang 0013, Yazhou Yao
ECCV (27)6
2024 Relating CNN-Transformer Fusion Network for Remote Sensing Change Detection
abstract
While deep learning, particularly convolutional neural networks (CNNs), has revolutionized remote sensing (RS) change detection (CD), existing approaches often miss crucial features due to neglecting global context and incomplete change learning. Additionally, transformer networks struggle with low-level details. RCTNet addresses these limitations by introducing (1) an early fusion backbone to exploit both spatial and temporal features early on, (2) a Cross-Stage Aggregation (CSA) module for enhanced temporal representation, (3) a Multi-Scale Feature Fusion (MSF) module for enriched feature extraction in the decoder, and (4) an Efficient Self-deciphering Attention (ESA) module utilizing transformers to capture global information and fine-grained details for accurate change detection. Extensive experiments demonstrate RCTNet’s clear superiority over traditional RS image CD methods, showing significant improvement and an optimal balance between accuracy and computational cost. Our source codes and pre-trained models are available at: https://github.com/NUST-Machine-Intelligence-Laboratory/RCTNet.
Yuhao Gao, Gensheng Pei, Mengmeng Sheng, Zeren Sun, Tao Chen 0012, Yazhou Yao
ICME6
2024 Universal Organizer of Segment Anything Model for Unsupervised Semantic Segmentation
abstract
Unsupervised semantic segmentation (USS) aims to achieve high-quality segmentation without manual pixel-level annotations. Existing USS models provide coarse category classifi-cation for regions, but the results often have blurry and imprecise edges. Recently, a robust framework called the segment anything model (SAM) has been proven to deliver precise boundary object masks. Therefore, this paper proposes a universal organizer based on SAM, termed as UO-SAM, to enhance the mask quality of USS models. Specifically, using only the original image and the masks generated by the USS model, we extract visual features to obtain positional prompts for target objects. Then, we activate a local region optimizer that performs segmentation using SAM on a per-object basis. Finally, we employ a global region optimizer to incorporate global image information and refine the masks to obtain the final fine-grained masks. Compared to existing methods, our UO-SAM achieves state-of-the-art performance. Our codes are available at https://github.com/NUST-Machine-Intelligence-Laboratory/UO-SAM.
Gensheng Pei, Xinhao Cai, Qiong Wang 0003, Huafeng Liu 0004, Yazhou Yao
ICME6
2024 Progressively Robust Loss for Deep Learning with Noisy Labels
abstract
Learning with noisy labels (LNL) plays a pivotal role in arming deep neural networks (DNNs) to combat label noise. Early noise-robust functions tend to promote robustness against noisy labels at the cost of sacrificing data-fitting ability. Recent robust loss methods typically try to balance noise-robustness and learning capability. However, most of them generally descend to partially robust losses, which are still exposed to the risk of overfitting noisy labels. To this end, we propose a novel paradigm named progressively robust loss framework to dynamically guide existing noise-robust losses from fast convergence to noise-tolerant, which is in accord with the deep models’ memorization effect. Furthermore, our theoretical analysis of the upper bounds of empirical risk errors illustrates the increasing noise-robustness of our approach. Experimental results on two synthetic benchmarks (CIFAR-100N and CIFAR-80N) and two real-world noisy datasets (WebFG-496 and Webvision) demonstrate the superiority of our approach over state-of-the-art robust loss methods in dealing with noisy labels. The code is available at https://github.com/ptcepgce/ptcepgce.
Zhenhuang Cai, Yuanbo Chen, Chuanyi Zhang, Zeren Sun, Yazhou Yao
IJCNN6
2024 AdaFPP: Adapt-Focused Bi-Propagating Prototype Learning for Panoramic Activity Recognition
abstract
Panoramic Activity Recognition (PAR) aims to identify multi-granul-arity behaviors performed by multiple persons in panoramic scenes, including individual activities, group activities, and global activities. Previous methods 1) heavily rely on manually annotated detection boxes in training and inference, hindering further practical deployment; or 2) directly employ normal detectors to detect multiple persons with varying size and spatial occlusion in panoramic scenes, blocking the performance gain of PAR. To this end, we consider learning a detector adapting varying-size occluded persons, which is optimized along with the recognition module in the all-in-one framework. Therefore, we propose a novel Adapt-Focused bi-Propagating Prototype learning (AdaFPP) framework to jointly recognize individual, group, and global activities in panoramic activity scenes by learning an adapt-focused detector and multi-granularity prototypes as the pretext tasks in an end-to-end way. Specifically, to accommodate the varying sizes and spatial occlusion of multiple persons in crowed panoramic scenes, we introduce a panoramic adapt-focuser, achieving the size-adapting detection of individuals by comprehensively selecting and performing fine-grained detections on object-dense sub-regions identified through original detections. In addition, to mitigate information loss due to inaccurate individual localizations, we introduce a bi-propagation prototyper that promotes closed-loop interaction and informative consistency across different granularities by facilitating bidirectional information propagation among the individual, group, and global levels. Extensive experiments demonstrate the significant performance of AdaFPP and emphasize its powerful applicability for PAR.
Meiqi Cao, Rui Yan 0010, Xiangbo Shu, Guangzhao Dai, Yazhou Yao, Guosen Xie
ACM Multimedia5
2024 Enhancing Robustness in Learning with Noisy Labels: An Asymmetric Co-Training Approach
abstract
Label noise, an inevitable issue in various real-world datasets, tends to impair the performance of deep neural networks. A large body of literature focuses on symmetric co-training, aiming to enhance model robustness by exploiting interactions between models with distinct capabilities. However, the symmetric training processes employed in existing methods often culminate in model consensus, diminishing their efficacy in handling noisy labels. To this end, we propose an Asymmetric Co-Training (ACT) method to mitigate the detrimental effects of label noise. Specifically, we introduce an asymmetric training framework in which one model (i.e., RTM) is robustly trained with a selected subset of clean samples while the other (i.e., NTM) is conventionally trained using the entire training set. We propose two novel criteria based on agreement and discrepancy between models, establishing asymmetric sample selection and mining. Moreover, a metric, derived from the divergence between models, is devised to quantify label memorization, guiding our method in determining the optimal stopping point for sample mining. Finally, we propose to dynamically re-weight identified clean samples according to their reliability inferred from historical information. We additionally employ consistency regularization to achieve further performance improvement. Extensive experimental results on synthetic and real-world datasets demonstrate the effectiveness and superiority of our method.
Mengmeng Sheng, Zeren Sun, Gensheng Pei, Tao Chen 0012, Haonan Luo 0002, Yazhou Yao
ACM Multimedia6
2024 Delving Deeper Into Clean Samples for Combating Noisy Labels
Yiyou Gao, Zeren Sun, Yazhou Yao, Xiruo Jiang, Zhenmin Tang
PRCV (9)3
2024 Class Probability Space Regularization for semi-supervised semantic segmentation
Jianjian Yin, Tao Chen 0012, Yi Chen 0023, Yazhou Yao
Comput. Vis. Image Underst.5
2024 Two-stage fine-grained image classification model based on multi-granularity feature fusion
Yang Xu 0006, Biqi Wang, Zebin Wu 0001, Yazhou Yao, Zhihui Wei
Pattern Recognit.6
2024 Holistic Prototype Attention Network for Few-Shot Video Object Segmentation
abstract
Few-shot video object segmentation (FSVOS) aims to segment dynamic objects of unseen classes by resorting to a small set of support images that contain pixel-level object annotations. Existing methods have demonstrated that the domain agent-based attention mechanism is effective in FSVOS by learning the correlation between support images and query frames. However, the agent frame contains redundant pixel information and background noise, resulting in inferior segmentation performance. Moreover, existing methods tend to ignore inter-frame correlations in query videos. To alleviate the above dilemma, we propose a holistic prototype attention network (HPAN) for advancing FSVOS. Specifically, HPAN introduces a prototype graph attention module (PGAM) and a bidirectional prototype attention module (BPAM), transferring informative knowledge from seen to unseen classes. PGAM generates local prototypes from all foreground features and then utilizes their internal correlations to enhance the representation of the holistic prototypes. BPAM exploits the holistic information from support images and video frames by fusing co-attention and self-attention to achieve support-query semantic consistency and inner-frame temporal consistency. Extensive experiments on YouTube-FSVOS have been provided to demonstrate the effectiveness and superiority of our proposed HPAN method. Our source code and models are available anonymously at https://github.com/NUST-Machine-Intelligence-Laboratory/HPAN.
Tao Chen 0012, Xiruo Jiang, Yazhou Yao, Guosen Xie, Heng Tao Shen
IEEE Trans. Circuits Syst. Video Technol.4
2024 Spatial Structure Constraints for Weakly Supervised Semantic Segmentation
abstract
The image-level label has prevailed in weakly supervised semantic segmentation tasks due to its easy availability. Since image-level labels can only indicate the existence or absence of specific categories of objects, visualization-based techniques have been widely adopted to provide object location clues. Considering class activation maps (CAMs) can only locate the most discriminative part of objects, recent approaches usually adopt an expansion strategy to enlarge the activation area for more integral object localization. However, without proper constraints, the expanded activation will easily intrude into the background region. In this paper, we propose spatial structure constraints (SSC) for weakly supervised semantic segmentation to alleviate the unwanted object over-activation of attention expansion. Specifically, we propose a CAM-driven reconstruction module to directly reconstruct the input image from deep CAM features, which constrains the diffusion of last-layer object attention by preserving the coarse spatial structure of the image content. Moreover, we propose an activation self-modulation module to refine CAMs with finer spatial structure details by enhancing regional consistency. Without external saliency models to provide background clues, our approach achieves 72.7% and 47.0% mIoU on the PASCAL VOC 2012 and COCO datasets, respectively, demonstrating the superiority of our proposed approach. The source codes and models have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/SSC.
Tao Chen 0012, Yazhou Yao, Xingguo Huang, Zechao Li, Liqiang Nie, Jinhui Tang 0001
IEEE Trans. Image Process.2
2024 Anti-Collapse Loss for Deep Metric Learning
abstract
Deep metric learning (DML) aims to learn a discriminative high-dimensional embedding space for downstream tasks like classification, clustering, and retrieval. Prior literature predominantly focuses on pair-based and proxy-based methods to maximize inter-class discrepancy and minimize intra-class diversity. However, these methods tend to suffer from the collapse of the embedding space due to their over-reliance on label information. This leads to sub-optimal feature representation and inferior model performance. To maintain the structure of embedding space and avoid feature collapse, we propose a novel loss function called Anti-Collapse Loss. Specifically, our proposed loss primarily draws inspiration from the principle of Maximal Coding Rate Reduction. It promotes the sparseness of feature clusters in the embedding space to prevent collapse by maximizing the average coding rate of sample features or class proxies. Moreover, we integrate our proposed loss with pair-based and proxy-based methods, resulting in notable performance improvement. Comprehensive experiments on benchmark datasets demonstrate that our proposed method outperforms existing state-of-the-art methods. Extensive ablation studies verify the effectiveness of our method in preventing embedding space collapse and promoting generalization performance.
Xiruo Jiang, Yazhou Yao, Xili Dai, Fumin Shen, Liqiang Nie, Heng Tao Shen
IEEE Trans. Multim.2
2024 Learning With Imbalanced Noisy Data by Preventing Bias in Sample Selection
abstract
Learning with noisy labels has gained increasing attention because the inevitable imperfect labels in real-world scenarios can substantially hurt the deep model performance. Recent studies tend to regard low-loss samples as clean ones and discard high-loss ones to alleviate the negative impact of noisy labels. However, real-world datasets contain not only noisy labels but also class imbalance. The imbalance issue is prone to causing failure in the loss-based sample selection since the under-learning of tail classes also leans to produce high losses. To this end, we propose a simple yet effective method to address noisy labels in imbalanced datasets. Specifically, we proposeClass-Balance-based sampleSelection (CBS) to prevent the tail class samples from being neglected during training. We proposeConfidence-basedSampleAugmentation (CSA) for the chosen clean samples to enhance their reliability in the training process. To exploit selected noisy samples, we resort to prediction history to rectify labels of noisy samples. Moreover, we introduce theAverageConfidenceMargin (ACM) metric to measure the quality of corrected labels by leveraging the model's evolving training dynamics, thereby ensuring that low-quality corrected noisy samples are appropriately masked out. Lastly, consistency regularization is imposed on filtered label-corrected noisy samples to boost model performance. Comprehensive experimental results on synthetic and real-world datasets demonstrate the effectiveness and superiority of our proposed method, especially in imbalanced scenarios. The source code has been made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/CBS.
Huafeng Liu 0004, Mengmeng Sheng, Zeren Sun, Yazhou Yao, Xian-Sheng Hua 0001, Heng Tao Shen
IEEE Trans. Multim.4
2024 Deep Metric Learning Based on Meta-Mining Strategy With Semiglobal Information
abstract
Recently, deep metric learning (DML) has achieved great success. Some existing DML methods propose adaptive sample mining strategies, which learn to weight the samples, leading to interesting performance. However, these methods suffer from a small memory (e.g., one training batch), limiting their efficacy. In this work, we introduce a data-driven method, meta-mining strategy with semiglobal information (MMSI), to apply meta-learning to learn to weight samples during the whole training, leading to an adaptive mining strategy. To introduce richer information than one training batch only, we elaborately take advantage of the validation set of meta-learning by implicitly adding additional validation sample information to training. Furthermore, motivated by the latest self-supervised learning, we introduce a dictionary (memory) that maintains very large and diverse information. Together with the validation set, this dictionary presents much richer information to the training, leading to promising performance. In addition, we propose a new theoretical framework that can formulate pairwise and tripletwise metric learning loss functions in a unified framework. This framework brings new insights to society and facilitates us to generalize our MMSI to many existing DML methods. We conduct extensive experiments on three public datasets, CUB200-2011, Cars-196, and Stanford Online Products (SOP). Results show that our method can achieve the state of the art or very competitive performance. Our source codes have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/MMSI.
Xi Jiang 0001, Sheng Liu 0009, Xili Dai, Guosheng Hu, Xingguo Huang, Yazhou Yao, Guosen Xie, Ling Shao 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Robust-EQA: Robust Learning for Embodied Question Answering With Noisy Labels
abstract
Embodied question answering (EQA) is a recently emerged research field in which an agent is asked to answer the user's questions by exploring the environment and collecting visual information. Plenty of researchers turn their attention to the EQA field due to its broad potential application areas, such as in-home robots, self-driven mobile, and personal assistants. High-level visual tasks, such as EQA, are susceptible to noisy inputs, because they have complex reasoning processes. Before the profits of the EQA field can be applied to practical applications, good robustness against label noise needs to be equipped. To tackle this problem, we propose a novel label noise-robust learning algorithm for the EQA task. First, a joint training co-regularization noise-robust learning method is proposed for noisy filtering of the visual question answering (VQA) module, which trains two parallel network branches by one loss function. Then, a two-stage hierarchical robust learning algorithm is proposed to filter out noisy navigation labels in both trajectory level and action level. Finally, by taking purified labels as inputs, a joint robust learning mechanism is given to coordinate the work of the whole EQA system. Empirical results demonstrate that, under extremely noisy environments (45% of noisy labels) and low-level noisy environments (20% of noisy labels), the robustness of deep learning models trained by our algorithm is superior to the existing EQA models in noisy environments.
Haonan Luo 0002, Guosheng Lin, Fumin Shen, Xingguo Huang, Yazhou Yao, Heng Tao Shen
IEEE Trans. Neural Networks Learn. Syst.5
2024 Dual Dynamic Threshold Adjustment Strategy
abstract
Loss functions and sample mining strategies are essential components in deep metric learning algorithms. However, the existing loss function or mining strategy often necessitates the incorporation of additional hyperparameters, notably the threshold, which defines whether the sample pair is informative. The threshold provides a stable numerical standard for determining whether to retain the pairs. It is a vital parameter to reduce the redundant sample pairs participating in training. Nonetheless, finding the optimal threshold can be a time-consuming endeavor, often requiring extensive grid searches. Because the threshold cannot be dynamically adjusted in the training stage, we should conduct plenty of repeated experiments to determine the threshold. Therefore, we introduce a novel approach for adjusting the thresholds associated with both the loss function and the sample mining strategy. We design a static Asymmetric Sample Mining Strategy (ASMS) and its dynamic version, the Adaptive Tolerance ASMS (AT-ASMS), tailored for sample mining methods. ASMS utilizes differentiated thresholds to address the problems (too few positive pairs and too many redundant negative pairs) caused by only applying a single threshold to filter samples. The AT-ASMS can adaptively regulate the ratio of positive and negative pairs during training according to the ratio of the currently mined positive and negative pairs. This meta-learning-based threshold generation algorithm utilizes a single-step gradient descent to obtain new thresholds. We combine these two threshold adjustment algorithms to form the Dual Dynamic Threshold Adjustment Strategy (DDTAS). Experimental results show that our algorithm achieves competitive performance on the CUB200, Cars196, and SOP datasets. Our codes are available at https://github.com/NUST-Machine-Intelligence-Laboratory/DDTAS .
Xiruo Jiang, Yazhou Yao, Sheng Liu 0009, Fumin Shen, Liqiang Nie, Xian-Sheng Hua 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Semi-Supervised Semantic Segmentation With Region Relevance
abstract
Semi-supervised semantic segmentation aims to learn from a small amount of labeled data and plenty of unlabeled ones for the segmentation task. The most common approach is to generate pseudo-labels for unlabeled images to augment the training data. However, the noisy pseudo-labels will lead to cumulative classification errors and aggravate the local inconsistency in prediction. This paper proposes a Region Relevance Network (RRN) to alleviate the problem mentioned above. Specifically, we first introduce a local pseudo-label filtering module that leverages discriminator networks to assess the accuracy of the pseudo-label at the region level. A local selection loss is proposed to mitigate the negative impact of wrong pseudo-labels in consistency regularization training. In addition, we propose a dynamic region-loss correction module, which takes the merit of network diversity to further rate the reliability of pseudo-labels and correct the convergence direction of the segmentation network with a dynamic region loss. Extensive experiments are conducted on PASCAL VOC 2012 and Cityscapes datasets with varying amounts of labeled data, demonstrating that our proposed approach achieves state-of-the-art performance compared to current counterparts. Our code is available at: https://github.com/NUST-Machine-Intelligence-Laboratory/TorchSemiSeg2.
Tao Chen 0012, Qiong Wang 0003, Yazhou Yao
ICME4
2023 Guest Editorial: Learning from limited annotations for computer vision tasks
abstract
The past decade has witnessed remarkable achievements in computer vision, owing to the fast development of deep learning. With the advancement of computing power and deep learning algorithms, we can process and apply millions or even hundreds of millions of large-scale data to train robust and advanced deep learning models. In spite of the impressive success, current deep learning methods tend to rely on massive annotated training data and lack the capability of learning from limited exemplars. However, constructing a million-scale annotated dataset like ImageNet is time-consuming, labour-intensive and even infeasible in many applications. In certain fields, very limited annotated examples can be gathered due to various reasons such as privacy or ethical issues. Consequently, one of the pressing challenges in computer vision is to develop approaches that are capable of learning from limited annotated data. The purpose of this Special Issue is to collect high-quality articles on learning from limited annotations for computer vision tasks (e.g. image classification, object detection, semantic segmentation, instance segmentation and many others), publish new ideas, theories, solutions and insights on this topic and showcase their applications. In this Special Issue we received 29 papers, all of which underwent peer review. Of the 29 originally submitted papers, 9 have been accepted. The nine accepted papers can be clustered into two main categories: theoretical and applications. The papers that fall into the first category are by Liu et al., Li et al. and He et al. The second category of papers offers a direct solution to various computer vision tasks. These papers are by Ma et al., Wu et al., Rao et al., Sun et al., Hou et al. and Gong et al. A brief presentation of each of the papers in this Special Issue follows. Liu et al. present a Gaussianisation prototypical classifier (GPC) for few-shot classification which mainly focuses on solving the issue of prototype bias. GPC consists of handling the features with the Gaussianisation operation and estimating a reliable prototype using the maximum a posteriori method using base class features as prior information. The proposed method is simple yet effective, which does not use any extra labelled data or knowledge. Moreover, it's also a one-step prototype rectification method, which does not resort any complex continuous optimisation. The ablation study shows that GPC can benefit from features pretrained only with CE loss or jointly trained with self-supervised loss. The results demonstrate that the proposed method outperforms related work and other state-of-the-art methods. Li et al. present a novel hyperspectral unmixing method named ‘Global centralised and Structured discriminative Nonnegative Matrix Factorisation (GSNMF)’. The proposed GSNMF offers several distinct advantages over the traditional unmixing techniques. Constructed on the foundation of the manifold regularisation techniques, GSNMF captures the intrinsic structural information by using the local affinity and distant repulsion constraints concurrently. With the structured discriminative information, local affinity constraint ensures that similar elements share similar estimated abundances, while the distant repulsion constraint ensures that dissimilar elements have different abundances. All experiments and analyses have demonstrated that the proposed GSNMF exhibits a remarkable performance compared to the other methods. He et al. present a taxonomy of existing algorithms in the task of makeup transfer. Evaluation methods are proposed, existing methods are analysed and existing datasets are reviewed. Finally the current problems in the field of makeup transfer are discussed, and the trend of future research is analysed. Ma et al. present a dense transformer framework for person re-identification tasks. This paper introduces densely connected class tokens to connect any two layers implicitly. The framework, Denseformer, outperforms other vision transformer models on four widely used benchmarks, namely Market-1501, DukeMTMC-reID, MSMT17 and Occluded-Duke datasets with only a small amount of extra calculation cost. According to the visualisation results, the proposed Denseformer pays more attention to the main parts of human bodies, obtaining discriminative global features. The Denseformer is a general improvement on ViT and works well on other tasks that use ViT as a backbone according to the promising results. Wu et al. present a new homology-continuous-based makeup transformation method, which can be roughly divided into two network branches: the age compensation branch and the makeup transformation branch. Specifically, in the age compensation branch, based on the same source continuity the authors designed a new encoding module which can map the face vector into the corresponding high-dimensional vector space and realise the compensation for age by adjusting the vector direction. In the makeup transformation branch, this work designed a multi-style encoder to handle different types of makeup, such as Chinese Japanese Korean makeup etc. In addition, the proposed network structure is a two-pass encoder-decoder architecture which has good parallelism and can achieve better results with training and inference on GPU. Rao et al. present a novel end-to-end architecture for point completion by using a stack-style folding network called the SSFN. Due to the fact that the output shape code cannot completely represent semantically, they propose a Stack-Style Folding module that transforms the bottleneck output into the style code analogous to StyleGAN. Experiments on ShapeNet and KITTI datasets indicate that the proposed SSFN architecture achieves a decent visual quality and metric performance. Sun et al. present a method for a zero-shot temporal event localisation (ZSTEL) that leverage large-scale video and language models, for example, CLIP. They solve the two key problems for ZSTEL: (1) how to find the relevant region where the event is likely to occur, (2) how to determine event duration after the relevant region is found. They propose the query-guided optimisation for local frame relevance. Relying on the query-to-frame relationship, this method can find the most relevant local frame region where the event is most likely to occur, guided by a constructed objective. The experimental results on the two standard benchmark datasets, Charades-STA and ActivityCaptions have shown the effectiveness of the proposed approach. Hou et al. present a cutting-edge few-shot detection method for logo images. To avoid the misclassification between the base and novel classes, they add an extra classification head. They also apply the convolutional layer into regression heads to improve the accuracy of location by using the limited training data. Considering the characteristics of logo images, they add balanced feature pyramid with Deformable RoI Pooling and unfreeze region proposal network in the fine-tuning stage. The extensive comparative experimentation and ablation studies illustrate the advantage of the proposed method and the effectiveness of every component in the model. Gong et al. present a method for object detection with a long-tail distribution that includes a dual-balanced network and balanced classification loss. This work investigates how the long-tailed distribution impacts the sub-networks in the general two-stage object detection framework Faster-RCNN and finds that unbalanced proposal sampling and unbalanced classification logic deteriorate the performance of the model in terms of AP. They propose the balanced region proposal network and balanced the classification network to address the above issues. Experiments on the LVIS-v0.5 dataset demonstrate that the framework improves the performance of AP without sacrificing too much from the performance of head categories in long-tail distribution. All of the papers selected for this Special Issue show that the field of learning from limited annotations for computer vision tasks is steadily moving forward. The possibility of a weakly supervised learning paradigm will remain a source of inspiration for new techniques in the years to come. Firstly, we wish to express our thanks to Ph.D. students at Nanjing University of Science and Technology for their continuous assistance throughout this process. Also, we wish to express our gratitude to all the contributors who submitted novel scientific results in this special issue and to the anonymous reviewers, whose expert work allowed the realisation of this endeavor. We aspire that this effort should contribute to the further development of DL and increase the concern of the scientific and technological community in the respective area. Last, we should not omit to express our appreciation to the journal's Editors-in-Chief and the Editorial Office for their support throughout this venture. Data sharing is not applicable to this article as no new data were created or analysed in this study. Yazhou Yao is a professor at the School of Computer Science and Engineering and Nanjing University of Science and Technology. With the support of the China Scholarship Council, he received his Ph.D. degree in Computer Science, University of Technology Sydney, Australia at 2018. From July 2018 to July 2019, he worked as a Research Scientist at the Inception Institute of Artificial Intelligence, Abu Dhabi, UAE. His research interests include multimedia processing and machine learning. Wenguan Wang is currently a ZJU100 Young Professor at Zhejiang University. He received his Ph.D. degree from Beijing Institute of Technology in 2018. From 2016 to 2018, he was a joint Ph.D. candidate at the University of California, Los Angeles. From 2018 to 2019, he was a senior scientist at the Inception Institute of Artificial Intelligence, UAE. From 2020 to 2022, he worked as a postdoc researcher at ETH Zurich, Switzerland. After that, he worked as a lecturer and ARC DECRA Fellow at the University of Technology Sydney. His current research interests include computer vision, image processing and deep learning. Qiang Wu received the BEng and MEng degrees in electronic engineering from the Harbin Institute of Technology, Harbin, China, in 1996 and 1998, respectively, and the Ph.D. degree in computing science from the University of Technology Sydney, Sydney, Australia, in 2004. He is currently an Associate Professor and a Core Member of the Global Big Data Technologies Centre, University of Technology Sydney. He has published more than 70 refereed papers, including those published in prestigious journals and top international conferences. His major research interests include computer vision, image processing, pattern recognition, machine learning and multimedia processing. He has served as the chair and/or a Programme Committee Member for a number of international conferences. Dongfang Liu is an Assistant Professor in the Department of Computer Engineering at the Rochester Institute of Technology (RIT). He earned his Ph.D. degree from Purdue University. Dr. Dongfang Liu's research focus on embodied AI and creates general AI solutions to address significant societal challenges. His ongoing work consists of: (1) developing attention-guided perception models that behave like a human's perpetual capacity; and (2) developing structured and human-centred recognition systems that comprehend the surrounding visual world. His publication portfolio includes papers from major conferences in the artificial intelligence and robotics fields, such as CVPR, ECCV, ICCV, ICLR, NIPS, ICML, AAAI, IJCAI, ACL, EMNLP, WWW, WACV, IROS etc. He currently serves on the senior programme committee for AAAI and IJCAI and as an associate editor for IEEE Transactions on Circuits and Systems for Video Technology (TCSVT). Jin Zheng received the BS and MS degrees from Liaoning Technical University, in 2001 and 2004, respectively, and the Ph.D. degree from the School of Computer Science and Engineering, Beihang University, in 2009. She joined the School of Computer Science and Engineering, Beihang University, in 2009. In 2014, she visited Harvard University, MA, USA, as a Visiting Scholar for 1 year. Her current research interests include object detection, tracking and recognition, among other similar interests.
Yazhou Yao, Wenguan Wang, Qiang Wu 0001, Dongfang Liu
IET Comput. Vis.1
2023 Information bottleneck and selective noise supervision for zero-shot learning
Lei Zhou 0008, Yang Liu 0357, Pengcheng Zhang 0003, Xiao Bai 0001, Lin Gu 0003, Jun Zhou 0001, Yazhou Yao, Tatsuya Harada, Edwin R. Hancock
Mach. Learn.7
2023 Depth and Video Segmentation Based Visual Attention for Embodied Question Answering
abstract
Embodied Question Answering (EQA) is a newly defined research area where an agent is required to answer the user's questions by exploring the real-world environment. It has attracted increasing research interests due to its broad applications in personal assistants and in-home robots. Most of the existing methods perform poorly in terms of answering and navigation accuracy due to the absence of fine-level semantic information, stability to the ambiguity, and 3D spatial information of the virtual environment. To tackle these problems, we propose a depth and segmentation based visual attention mechanism for Embodied Question Answering. First, we extract local semantic features by introducing a novel high-speed video segmentation framework. Then guided by the extracted semantic features, a depth and segmentation based visual attention mechanism is proposed for the Visual Question Answering (VQA) sub-task. Further, a feature fusion strategy is designed to guide the navigator's training process without much additional computational cost. The ablation experiments show that our method effectively boosts the performance of the VQA module and navigation module, leading to 4.9 % and 5.6 % overall improvement in EQA accuracy on House3D and Matterport3D datasets respectively.
Haonan Luo 0002, Guosheng Lin, Yazhou Yao, Fayao Liu, Zichuan Liu, Zhenmin Tang
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Robust learning from noisy web data for fine-Grained recognition
Zhenhuang Cai, Guosen Xie, Xingguo Huang, Yazhou Yao, Zhenmin Tang
Pattern Recognit.5
2023 Co-mining: Mining informative samples with noisy labels
Zhenhuang Cai, Huafeng Liu 0004, Yazhou Yao, Zhenmin Tang
Signal Process.4
2023 Motion Stimulation for Compositional Action Recognition
abstract
Recognizing the unseen combinations of action and different objects, namely (zero-shot) compositional action recognition, is extremely challenging for conventional action recognition algorithms in real-world applications. Previous methods focus on enhancing the dynamic clues of objects that appear in the scene by building region features or tracklet embedding from ground-truths or detected bounding boxes. These methods rely heavily on manual annotation or the quality of detectors, which are inflexible for practical applications. In this work, we aim to mining the temporal clues from moving objects or hands without explicit supervision. Thus, we propose a novel Motion Stimulation (MS) block, which is specifically designed to mine dynamic clues of the local regions autonomously from adjacent frames. Furthermore, MS consists of the following three steps: motion feature extraction, motion feature recalibration, and action-centric excitation. The proposed MS block can be directly and conveniently integrated into existing video backbones to enhance the ability of compositional generalization for action recognition algorithms. Extensive experimental results on three action recognition datasets, the Something-Else, IKEA-Assembly and EPIC-KITCHENS datasets, indicate the effectiveness and interpretability of our MS block.
Yuhui Zheng, Zhao Zhang 0001, Yazhou Yao, Xijian Fan, Qiaolin Ye
IEEE Trans. Circuits Syst. Video Technol.4
2023 Multi-Granularity Denoising and Bidirectional Alignment for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation (WSSS) models relying on class activation maps (CAMs) have achieved desirable performance comparing to the non-CAMs-based counterparts. However, to guarantee WSSS task feasible, we need to generate pseudo labels by expanding the seeds from CAMs which is complex and time-consuming, thus hindering the design of efficient end-to-end (single-stage) WSSS approaches. To tackle the above dilemma, we resort to the off-the-shelf and readily accessible saliency maps for directly obtaining pseudo labels given the image-level class labels. Nevertheless, the salient regions may contain noisy labels and cannot seamlessly fit the target objects, and saliency maps can only be approximated as pseudo labels for simple images containing single-class objects. As such, the achieved segmentation model with these simple images cannot generalize well to the complex images containing multi-class objects. To this end, we propose an end-to-end multi-granularity denoising and bidirectional alignment (MDBA) model, to alleviate the noisy label and multi-class generalization issues. Specifically, we propose the online noise filtering and progressive noise detection modules to tackle image-level and pixel-level noise, respectively. Moreover, a bidirectional alignment mechanism is proposed to reduce the data distribution gap at both input and output space with simple-to-complex image synthesis and complex-to-simple adversarial learning. MDBA can reach the mIoU of 69.5% and 70.2% on validation and test sets for the PASCAL VOC 2012 dataset. The source codes and models have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/MDBA.
Tao Chen 0012, Yazhou Yao, Jinhui Tang 0001
IEEE Trans. Image Process.2
2023 Hierarchical Graph Pattern Understanding for Zero-Shot Video Object Segmentation
abstract
The optical flow guidance strategy is ideal for obtaining motion information of objects in the video. It is widely utilized in video segmentation tasks. However, existing optical flow-based methods have a significant dependency on optical flow, which results in poor performance when the optical flow estimation fails for a particular scene. The temporal consistency provided by the optical flow could be effectively supplemented by modeling in a structural form. This paper proposes a new hierarchical graph neural network (GNN) architecture, dubbed hierarchical graph pattern understanding (HGPU), for zero-shot video object segmentation (ZS-VOS). Inspired by the strong ability of GNNs in capturing structural relations, HGPU innovatively leverages motion cues (i.e., optical flow) to enhance the high-order representations from the neighbors of target frames. Specifically, a hierarchical graph pattern encoder with message aggregation is introduced to acquire different levels of motion and appearance features in a sequential manner. Furthermore, a decoder is designed for hierarchically parsing and understanding the transformed multi-modal contexts to achieve more accurate and robust results. HGPU achieves state-of-the-art performance on four publicly available benchmarks (DAVIS-16, YouTube-Objects, Long-Videos and DAVIS-17). Code and pre-trained model can be found at https://github.com/NUST-Machine-Intelligence-Laboratory/HGPU.
Gensheng Pei, Fumin Shen, Yazhou Yao, Tao Chen 0012, Xian-Sheng Hua 0001, Heng Tao Shen
IEEE Trans. Image Process.3
2023 Hierarchical Co-Attention Propagation Network for Zero-Shot Video Object Segmentation
abstract
Zero-shot video object segmentation (ZS-VOS) aims to segment foreground objects in a video sequence without prior knowledge of these objects. However, existing ZS-VOS methods often struggle to distinguish between foreground and background or to keep track of the foreground in complex scenarios. The common practice of introducing motion information, such as optical flow, can lead to overreliance on optical flow estimation. To address these challenges, we propose an encoder-decoder-based hierarchical co-attention propagation network (HCPN) capable of tracking and segmenting objects. Specifically, our model is built upon multiple collaborative evolutions of the parallel co-attention module (PCM) and the cross co-attention module (CCM). PCM captures common foreground regions among adjacent appearance and motion features, while CCM further exploits and fuses cross-modal motion features returned by PCM. Our method is progressively trained to achieve hierarchical spatio-temporal feature propagation across the entire video. Experimental results demonstrate that our HCPN outperforms all previous methods on public benchmarks, showcasing its effectiveness for ZS-VOS. Code and pre-trained model can be found at https://github.com/NUST-Machine-Intelligence-Laboratory/HCPN.
Gensheng Pei, Yazhou Yao, Fumin Shen, Xingguo Huang, Heng Tao Shen
IEEE Trans. Image Process.2
2023 Saliency Guided Inter- and Intra-Class Relation Constraints for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation with only image-level labels aims to reduce annotation costs for the segmentation task. Existing approaches generally leverage class activation maps (CAMs) to locate the object regions for pseudo label generation. However, CAMs can only discover the most discriminative parts of objects, thus leading to inferior pixel-level pseudo labels. To address this issue, we propose a saliency guidedInter- andIntra-ClassRelationConstrained (I$^{2}$CRC) framework to assist the expansion of the activated object regions in CAMs. Specifically, we propose a saliency guided class-agnostic distance module to pull the intra-category features closer by aligning features to their class prototypes. Further, we propose a class-specific distance module to push the inter-class features apart and encourage the object region to have a higher activation than the background. Besides strengthening the capability of the classification network to activate more integral object regions in CAMs, we also introduce an object guided label refinement module to take a full use of both the segmentation prediction and the initial labels for obtaining superior pseudo-labels. Extensive experiments on PASCAL VOC 2012 and COCO datasets demonstrate well the effectiveness of I$^{2}$CRC over other state-of-the-art counterparts.
Tao Chen 0012, Yazhou Yao, Lei Zhang 0054, Qiong Wang 0003, Guosen Xie, Fumin Shen
IEEE Trans. Multim.2
2023 FECANet: Boosting Few-Shot Semantic Segmentation With Feature-Enhanced Context-Aware Network
abstract
Few-shot semantic segmentation is the task of learning to locate each pixel of the novel class in the query image with only a few annotated support images. The current correlation-based methods construct pair-wise feature correlations to establish the many-to-many matching because the typical prototype-based approaches cannot learn fine-grained correspondence relations. However, the existing methods still suffer from the noise contained in naive correlations and the lack of context semantic information in correlations. To alleviate these problems mentioned above, we propose a Feature-Enhanced Context-Aware Network (FECANet). Specifically, a feature enhancement module is proposed to suppress the matching noise caused by inter-class local similarity and enhance the intra-class relevance in the naive correlation. In addition, we propose a novel correlation reconstruction module that encodes extra correspondence relations between foreground and background and multi-scale context semantic features, significantly boosting the encoder to capture a reliable matching pattern. Experiments on PASCAL-$5^{i}$and COCO-$20^{i}$datasets demonstrate that our proposed FECANet leads to remarkable improvement compared to previous state-of-the-arts, demonstrating its effectiveness. The source codes and models have been made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/FECANET.
Huafeng Liu 0004, Tao Chen 0012, Qiong Wang 0003, Yazhou Yao, Xian-Sheng Hua 0001
IEEE Trans. Multim.5
2023 Attention Map Guided Transformer Pruning for Occluded Person Re-Identification on Edge Device
abstract
Due to its significant capability of modeling long-range dependencies, vision transformer (ViT) has achieved promising success in both holistic and occluded person re-identification (Re-ID) tasks. However, the inherent problems of transformers such as the huge computational cost and memory footprint are still two unsolved issues that will block the deployment of ViT based person Re-ID models on resource-limited edge devices. Our goal is to reduce both the inference complexity and model size without sacrificing the comparable accuracy on person Re-ID, especially for tasks with occlusion. To this end, we propose a novel attention map guided (AMG) transformer pruning method, which removes both redundant tokens and heads with the guidance of the attention map in a hardware-friendly way. We first calculate the entropy in the key dimension and sum it up for the whole map, and the corresponding head parameters of maps with high entropy will be removed for model size reduction. Then we combine the similarity and first-order gradients of key tokens along the query dimension for token importance estimation and remove redundant key and value tokens to further reduce the inference complexity. Comprehensive experiments on Occluded DukeMTMC and Market-1501 demonstrate the effectiveness of our proposals. For example, our proposed pruning strategy on ViT-Base enjoys29.4%FLOPssavings with0.2%drop on Rank-1 and0.4%improvement on mAP, respectively. Code and models have been made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/AMG.
Junzhu Mao, Yazhou Yao, Zeren Sun, Xingguo Huang, Fumin Shen, Heng Tao Shen
IEEE Trans. Multim.2
2023 Boosting Robust Learning Via Leveraging Reusable Samples in Noisy Web Data
abstract
Webly-supervised fine-grained visual classification (FGVC) has attracted increasing attention in recent years because of the unaffordable cost of obtaining correctly-labeled large-scale fine-grained datasets. However, due to the existence of label noise in web images and the high memorization capacity of deep neural networks, training deep fine-grained (FG) models directly through web images tends to have an inferior recognition ability. In the literature, to alleviate this issue, loss correction methods try to estimate the noise transition matrix, but the inevitable false correction would cause accumulated errors. Sample selection methods identify clean (“easy”) samples based on the fact that small losses can alleviate the accumulated errors. However, “hard” and mislabeled examples that can both boost the robustness of FG models are also dropped. To this end, we propose a certainty-based reusable sample selection and correction approach, termed as CRSSC, for coping with label noise in training deep FG models with web images. Our key idea is to additionally identify and correct reusable samples, and then leverage them together with clean examples to update the network. Furthermore, in order to endow our model with the capability to capture richer and more discriminative feature representations, we propose a cross-layer attention-based feature refinement (CLAR) block. We demonstrate the superiority of the proposed approach from both theoretical and experimental perspectives.
Zeren Sun, Yazhou Yao, Xiu-Shen Wei, Fumin Shen, Jian Zhang 0002, Xian-Sheng Hua 0001
IEEE Trans. Multim.2
2023 Guided by Meta-Set: A Data-Driven Method for Fine-Grained Visual Recognition
abstract
The lack of sufficient training data has been one obstacle to fine-grained visual classification research because labeling subcategories generally requires specialist knowledge. As one optional approach to alleviating the data-hunger problem, leveraging web images as training data is drawing increasing attention. Nevertheless, web images potentially have false labels, which can misguide the training process. Although several works have been proposed to deal with label noise, it still can be difficult for the network to tackle complex real-world noisy labels without any prior knowledge. In the literature, we propose to leverage a small and clean meta-set to provide reliable prior knowledge for tackling noisy web images. Specifically, our method trains a network with two peer predicting heads, which learn from noisy web images (web head) and meta ones (meta head), respectively. The meta head produces pseudo soft labels for web images to revise their training loss, which can overcome the high noise ratio problem. Furthermore, a selection net is trained in a meta-learning strategy to identify in- and out-of-distribution noisy images. Then in-distribution ones are reused for training with pseudo soft labels produced by the meta head as supervision, while out-of-distribution ones are discarded. In this manner, the misguidance caused by label noise is remarkably alleviated and in-distribution noisy samples are properly exploited to boost model performance. The superiority of our proposed approach is demonstrated by mathematical theory with great interpretability as well as extensive experimental results on the real-world dataset WebFG-496.
Chuanyi Zhang, Guosheng Lin, Qiong Wang 0003, Fumin Shen, Yazhou Yao, Zhenmin Tang
IEEE Trans. Multim.5
2022 PNP: Robust Learning from Noisy Labels by Probabilistic Noise Prediction
abstract
Label noise has been a practical challenge in deep learning due to the strong capability of deep neural networks in fitting all training data. Prior literature primarily resorts to sample selection methods for combating noisy labels. However, these approaches focus on dividing samples by order sorting or threshold selection, inevitably introducing hyperparameters (e.g., selection ratio / threshold) that are hard-to-tune and dataset-dependent. To this end, we propose a simple yet effective approach named PNP (Probabilistic Noise Prediction) to explicitly model label noise. Specifically, we simultaneously train two networks, in which one predicts the category label and the other predicts the noise type. By predicting label noise probabilistically, we identify noisy samples and adopt dedicated optimization objectives accordingly. Finally, we establish a joint loss for network update by unifying the classification loss, the auxiliary constraint loss, and the in-distribution consistency loss. Comprehensive experimental results on synthetic and realworld datasets demonstrate the superiority of our proposed method. The source code and models have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/PNP.
Zeren Sun, Fumin Shen, Qiong Wang 0003, Xiangbo Shu, Yazhou Yao, Jinhui Tang 0001
CVPR6
2022 Hierarchical Feature Alignment Network for Unsupervised Video Object Segmentation
Gensheng Pei, Fumin Shen, Yazhou Yao, Guosen Xie, Zhenmin Tang, Jinhui Tang 0001
ECCV (34)3
2022 Exploring Linear Feature Disentanglement for Neural Networks
abstract
Non-linear activation functions, e.g., Sigmoid, ReLU, and Tanh, have achieved great success in neural networks (NNs). Due to the complex non-linear characteristic of samples, the objective of those activation functions is to project samples from their original feature space to a linear separable feature space. This phenomenon ignites our interest in exploring whether all features need to be transformed by all nonlinear functions in current typical NNs, i.e., whether there exists a part of features arriving at the linear separable feature space in the intermediate layers, that does not require further non-linear variation but an affine transformation instead. To validate the above hypothesis, we explore the problem of linear feature disentanglement for neural networks in this paper. Specifically, we devise a learnable mask module to distinguish between linear and non-linear features. Through our designed experiments we found that some features reach the linearly separable space earlier than the others and can be detached partly from the NNs. The explored method also provides a readily feasible pruning strategy which barely affects the performance of the original model. We conduct our experiments on four datasets and present promising results.
Tiantian He 0004, Zhibin Li 0002, Yongshun Gong, Yazhou Yao, Xiushan Nie, Yilong Yin
ICME4
2022 Feature Difference Enhancement Fusion for Remote Sensing Image Change Detection
Gensheng Pei, Tao Chen 0012, Yazhou Yao
PRCV (3)5
2022 Unsupervised Pre-training for 3D Object Detection with Transformer
Maosheng Sun, Xiaoshui Huang, Zeren Sun, Qiong Wang 0003, Yazhou Yao
PRCV (3)5
2022 Few-Shot Object Detection via Understanding Convolution and Attention
Jiaxing Tong, Tao Chen 0012, Qiong Wang 0003, Yazhou Yao
PRCV (1)4
2022 DBFC-Net: a uniform framework for fine-grained cross-media retrieval
Qiong Wang 0003, Youdong Guo, Yazhou Yao
Multim. Syst.3
2022 Dense Semantics-Assisted Networks for Video Action Recognition
abstract
Most existing action recognition approaches directly leverage the video-level features to recognize human actions from videos. Although these methods have made remarkable progress, the accuracy is still unsatisfied. When the test video involves complex backgrounds and activities, existing methods usually suffer from a significant drop in accuracy. Human action is inherently a high-level concept. Merely applying a video classification model without a detailed semantic understanding of the video content, e.g., objects, scene context, object motions, object interactions, is inadequate to tackle the challenges for action recognition. Fine-level semantic understanding of videos generates elementary semantic concepts from the raw video data, such as the semantics of objects and background regions. It can be employed to bridge the gap between the raw video data and the high-level concept of human actions. In this work, we leverage dense semantic segmentation masks, which encode rich semantic details, provide extra information for the network training, and improve the performance of action recognition. We propose a novel deep architecture which is named as Dense Semantics-Assisted Convolutional Neural Networks (DSA-CNNs) to effectively utilize dense semantic information of video by a bottom-up attention way in the spatial stream, while by the way of branch fusion in the temporal stream. To verify the effectiveness of our approach, we conduct extensive experiments on publicly available datasets – UCF101, HMDB51, and Kinetics. The experimental results demonstrate that our approach substantially improves existing methods and achieves very competitive performance. It also shows that our approach is superior to other related methods that utilize extra information for action recognition.
Haonan Luo 0002, Guosheng Lin, Yazhou Yao, Zhenmin Tang, Qingyao Wu, Xian-Sheng Hua 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Self-Supervised Multi-Modal Hybrid Fusion Network for Brain Tumor Segmentation
abstract
Accurate medical image segmentation of brain tumors is necessary for the diagnosing, monitoring, and treating disease. In recent years, with the gradual emergence of multi-sequence magnetic resonance imaging (MRI), multi-modal MRI diagnosis has played an increasingly important role in the early diagnosis of brain tumors by providing complementary information for a given lesion. Different MRI modalities vary significantly in context, as well as in coarse and fine information. As the manual identification of brain tumors is very complicated, it usually requires the lengthy consultation of multiple experts. The automatic segmentation of brain tumors from MRI images can thus greatly reduce the workload of doctors and buy more time for treating patients. In this paper, we propose a multi-modal brain tumor segmentation framework that adopts the hybrid fusion of modality-specific features using a self-supervised learning strategy. The algorithm is based on a fully convolutional neural network. Firstly, we propose a multi-input architecture that learns independent features from multi-modal data, and can be adapted to different numbers of multi-modal inputs. Compared with single-modal multi-channel networks, our model provides a better feature extractor for segmentation tasks, which learns cross-modal information from multi-modal data. Secondly, we propose a new feature fusion scheme, named hybrid attentional fusion. This scheme enables the network to learn the hybrid representation of multiple features and capture the correlation information between them through an attention mechanism. Unlike popular methods, such as feature map concatenation, this scheme focuses on the complementarity between multi-modal data, which can significantly improve the segmentation results of specific regions. Thirdly, we propose a self-supervised learning strategy for brain tumor segmentation tasks. Our experimental results demonstrate the effectiveness of the proposed model against other state-of-the-art multi-modal medical segmentation methods.
Feiyi Fang, Yazhou Yao, Tao Zhou 0002, Guosen Xie, Jianfeng Lu 0003
IEEE J. Biomed. Health Informatics2
2022 Self-Supervised Depth Completion From Direct Visual-LiDAR Odometry in Autonomous Driving
abstract
In this work, a simple yet effective deep neural network is proposed to generate the dense depth map of the scene by exploiting both LiDAR sparse point cloud and the monocular camera image. Specifically, a feature pyramid network is firstly employed to extract feature maps from images across time. Then the relative pose is calculated by minimizing the feature distance between aligned pixels from inter-frame feature maps. Finally, the feature maps and the relative pose are further applied to compute the feature-metric loss for training the depth completion network. The key novelty of this work lies in that a self-supervised mechanism is presented to train the depth completion network by directly using visual-LiDAR odometry between consecutive frames. Comprehensive experiments and ablation studies on benchmark dataset KITTI demonstrate the superior performance over other state-of-the-art methods in terms of pose estimation and depth completion. The detailed performance of the proposed approach (referred to asSelfCompDVLO) can be found on the KITTI depth completion benchmark. The source code, models, and data have been made available at GitHub.
Zhenbo Song, Jianfeng Lu 0003, Yazhou Yao, Jian Zhang 0002
IEEE Trans. Intell. Transp. Syst.3
2022 Semantically Meaningful Class Prototype Learning for One-Shot Image Segmentation
abstract
One-shot semantic image segmentation aims to segment the object regions for the novel class with only one annotated image. Recent works adopt the episodic training strategy to mimic the expected situation at testing time. However, these existing approaches simulate the test conditions too strictly during the training process, and thus cannot make full use of the given label information. Besides, these approaches mainly focus on the foreground-background target class segmentation setting. They only utilize binary mask labels for training. In this paper, we propose to leverage the multi-class label information during the episodic training. It will encourage the network to generate more semantically meaningful features for each category. After integrating the target class cues into the query features, we then propose a pyramid feature fusion module to mine the fused features for the final classifier. Furthermore, to take more advantage of the support image-mask pair, we propose a self-prototype guidance branch to support image segmentation. It can constrain the network for generating more compact features and a robust prototype for each semantic class. For inference, we propose a fused prototype guidance branch for the segmentation of the query image. Specifically, we leverage the prediction of the query image to extract the pseudo-prototype and combine it with the initial prototype. Then we utilize the fused prototype to guide the final segmentation of the query image. Extensive experiments demonstrate the superiority of our proposed approach. The source codes and models have been made available athttps://github.com/NUST-Machine-Intelligence-Laboratory/SMCP.
Tao Chen 0012, Guosen Xie, Yazhou Yao, Qiong Wang 0003, Fumin Shen, Zhenmin Tang, Jian Zhang 0002
IEEE Trans. Multim.3
2022 Exploiting Web Images for Fine-Grained Visual Recognition by Eliminating Open-Set Noise and Utilizing Hard Examples
abstract
Labeling objects at a subordinate level typically requires expert knowledge, which is not always available when using random annotators. As such, learning directly from web images for fine-grained recognition has attracted broad attention. However, the presence of label noise and hard examples in web images are two obstacles for training robust fine-grained recognition models. Therefore, in this paper, we propose a novel approach for removing irrelevant samples from real-world web images during training, while employing useful hard examples to update the network. Thus, our approach can alleviate the harmful effects of irrelevant noisy web images and hard examples to achieve better performance. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is far superior to current state-of-the-art web-supervised methods. The data and source code of this work have been made publicly available at:https://github.com/NUST-Machine-Intelligence-Laboratory/Advanced-Softly-Update-Drop.
Huafeng Liu 0004, Chuanyi Zhang, Yazhou Yao, Xiu-Shen Wei, Fumin Shen, Zhenmin Tang, Jian Zhang 0002
IEEE Trans. Multim.3
2022 Guest Editorial: Learning From Noisy Multimedia Data
abstract
This special issue provides a premier forum for researchers in multimedia big data to share challenges and recent advancements in learning from noisy multimedia data. The multimedia age and its proliferation of devices and platforms is fueling exponential data growth. As computational power and deep learning algorithms rapidly evolve, the web has become a rich source of potential training data for robust machine learning, with search engines such as Google and Bing, Twitter, TikTok, Instagram, and short video sharing platforms offering large-scale data points in the hundreds of millions. The concurrent shift in the Internet to richer web data modalities such as text, audio, image, and video reveal further opportunities to leverage large-scale data for the automatic construction of a variety of datasets for model training and testing. However, the ubiquity of multimedia data means noise is a fundamental challenge, with ‘label noise’ and ‘domain mismatch’ the most critical issues in automatically collected datasets. Learning from noisy multimedia data tends towards poor performance, making it increasingly essential to address these challenges.
Jian Zhang 0002, Alan Hanjalic, Ramesh Jain 0001, Xian-Sheng Hua 0001, Shin'ichi Satoh 0001, Yazhou Yao, Dan Zeng 0001
IEEE Trans. Multim.6
2022 Multimodal Marketing Intent Analysis for Effective Targeted Advertising
abstract
People’s daily information sharing and acquisition through the Internet has become more and more popular. The comprehensive multimodal marketing advertorial generated by ‘We Media’ accounts besides the normal social news is gaining its importance on social media platforms. In order to achieve effective advertising, the marketing intent understanding is a key step towards generating targeted advertising strategies (push advertorials to specific people at a specific time). However, advertorials in real are usually designed to pretend as normal social news with a wide range of contents. This poses big challenges to the platforms on accurately recognizing and analyzing the marketing intents behind the advertorials. As a pioneering study, we address this new problem of multimodal-based marketing intent analysis and answer three core questions: (1) does a piece of social news contain marketing intent? (2) what is the topic of marketing intent? (3) what is the extent of marketing intent? Towards this end, we propose a novel Multimodal-based Marketing Intent Analysis scheme (MMIA) to estimate the marketing intent embedded in the multimodal contents. Specifically, a novel supervised neural autoregressive model (SmiDocNADE) is proposed to enhance the discriminative capacity of the learned hidden features so that a single system is capable of solving the three questions. In order to effectively model inter-correlations between images and text in advertorials, we fuse multimodal data and extract features by Graph Convolution Networks as an enhancement to SmiDocNADE. The extensive evaluations demonstrate the advantages of our proposed system in multimodal-based marketing intent analysis from multiple aspects.
Lu Zhang 0062, Jialie Shen 0001, Jian Zhang 0002, Jingsong Xu, Zhibin Li 0002, Yazhou Yao, Litao Yu
IEEE Trans. Multim.6
2021 Non-Salient Region Object Mining for Weakly Supervised Semantic Segmentation
abstract
Semantic segmentation aims to classify every pixel of an input image. Considering the difficulty of acquiring dense labels, researchers have recently been resorting to weak labels to alleviate the annotation burden of segmentation. However, existing works mainly concentrate on expanding the seed of pseudo labels within the image’s salient region. In this work, we propose a non-salient region object mining approach for weakly supervised semantic segmentation. We introduce a graph-based global reasoning unit to strengthen the classification network’s ability to capture global relations among disjoint and distant regions. This helps the network activate the object features outside the salient area. To further mine the non-salient region objects, we propose to exert the segmentation network’s self-correction ability. Specifically, a potential object mining module is proposed to reduce the false-negative rate in pseudo labels. Moreover, we propose a non-salient region masking module for complex images to generate masked pseudo labels. Our non-salient region masking module helps further discover the objects in the non-salient region. Extensive experiments on the PASCAL VOC dataset demonstrate state-of-the-art results compared to current methods. The source codes are available at https://github.com/NUST-Machine-Intelligence-Laboratory/nsrom.
Yazhou Yao, Tao Chen 0012, Guosen Xie, Chuanyi Zhang, Fumin Shen, Qi Wu 0001, Zhenmin Tang, Jian Zhang 0002
CVPR1
2021 Jo-SRC: A Contrastive Approach for Combating Noisy Labels
abstract
Due to the memorization effect in Deep Neural Networks (DNNs), training with noisy labels usually results in inferior model performance. Existing state-of-the-art methods primarily adopt a sample selection strategy, which selects small-loss samples for subsequent training. However, prior literature tends to perform sample selection within each mini-batch, neglecting the imbalance of noise ratios in different mini-batches. Moreover, valuable knowledge within high-loss samples is wasted. To this end, we propose a noise-robust approach named Jo-SRC (Joint Sample Selection and Model Regularization based on Consistency). Specifically, we train the network in a contrastive learning manner. Predictions from two different views of each sample are used to estimate its "likelihood" of being clean or out-of-distribution. Furthermore, we propose a joint loss to advance the model generalization performance by introducing consistency regularization. Extensive experiments have validated the superiority of our approach over existing state-of-the-art methods. The source code and models have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/Jo-SRC.
Yazhou Yao, Zeren Sun, Chuanyi Zhang, Fumin Shen, Qi Wu 0001, Jian Zhang 0002, Zhenmin Tang
CVPR1
2021 Webly Supervised Fine-Grained Recognition: Benchmark Datasets and An Approach
abstract
Learning from the web can ease the extreme dependence of deep learning on large-scale manually labeled datasets. Especially for fine-grained recognition, which targets at distinguishing subordinate categories, it will significantly reduce the labeling costs by leveraging free web data. Despite its significant practical and research value, the webly supervised fine-grained recognition problem is not extensively studied in the computer vision community, largely due to the lack of high-quality datasets. To fill this gap, in this paper we construct two new benchmark webly supervised fine-grained datasets, termed WebFG-496 and WebiNat-5089, respectively. In concretely, WebFG-496 consists of three sub-datasets containing a total of 53,339 web training images with 200 species of birds (Web-bird), 100 types of aircrafts (Web-aircraft), and 196 models of cars (Web-car). For WebiNat-5089, it contains 5089 sub-categories and more than 1.1 million web training images, which is the largest webly supervised fine-grained dataset ever. As a minor contribution, we also propose a novel webly supervised method (termed "Peer-learning") for benchmarking these datasets. Comprehensive experimental results and analyses on two new benchmark datasets demonstrate that the proposed method achieves superior performance over the competing baseline models and states-of-the-art. Our benchmark datasets and the source codes of Peer-learning have been made available at https://github.com/NUST-Machine-Intelligence-Laboratory/weblyFG-dataset.
Zeren Sun, Yazhou Yao, Xiu-Shen Wei, Fumin Shen, Jianxin Wu 0001, Jian Zhang 0002, Heng Tao Shen
ICCV2
2021 Few-Shot Semantic Segmentation with Cyclic Memory Network
abstract
Few-shot semantic segmentation (FSS) is an important task for novel (unseen) object segmentation under the data-scarcity scenario. However, most FSS methods rely on unidirectional feature aggregation, e.g., from support prototypes to get the query prediction, and from high-resolution features to guide the low-resolution ones. This usually fails to fully capture the cross-resolution feature relationships and thus leads to inaccurate estimates of the query objects. To resolve the above dilemma, we propose a cyclic memory network (CMN) to directly learn to read abundant support information from all resolution features in a cyclic manner. Specifically, we first generate N pairs (key and value) of multi-resolution query features guided by the support feature and its mask. Next, we circularly take one pair of these features as the query to be segmented, and the rest N-1 pairs are written into an external memory accordingly, i.e., this leave-one-out process is conducted for N times. In each cycle, the query feature is updated by collaboratively matching its key and value with the memory, which can elegantly cover all the spatial locations from different resolutions. Furthermore, we incorporate the query feature re-adding and the query feature recursive updating mechanisms into the memory reading operation. CMN, equipped with these merits, can thus capture cross-resolution relationships and better handle the object appearance and scale variations in FSS. Experiments on PASCAL-5iand COCO-20iwell validate the effectiveness of our model for FSS.
Guosen Xie, Huan Xiong, Jie Liu 0043, Yazhou Yao, Ling Shao 0001
ICCV4
2021 Link Prediction with Multiple Structural Attentions in Multiplex Networks
abstract
Many real networks can be viewed as multiplex networks with more than one layers. As different layers are usually not independent from each other, they can provide complementary information in the task of link prediction. In this paper, with the help of attention mechanism, we dig the structural correlations among different layers of the multiplex network as well as the network structural information of the target layer to make more precise link predictions. Specifically, we introduce three different attentions, namely the intra-layer distance/degree attention, the intra-layer neighbourhood attention, and the interlayer structural attention, to calculate both the influence among nodes in the same layer and the link correlations in different layers. Compared with other state-of-the art methods which usually require the information of node attributes or edge types, we only utilize the topological information of the network and thus provide a more general link prediction solution for multiplex network. We conduct comprehensive experiments on several real-world datatsets of different scales. By comparing with the state-of-the-art link prediction algorithms, we show the advantages of our algorithm, and the effectiveness of different attentions. Also, through visual case studies we uncover some intuitions about the relationship between the graph structure and the existence of a link. We make our source code anonymously available at: (will be released after review)
Shangrong Huang, Quanyu Ma, Chao Yang 0015, Yazhou Yao
IJCNN4
2021 CAA: Candidate-Aware Aggregation for Temporal Action Detection
abstract
Temporal action detection aims to locate specific segments of action instances in an untrimmed video. Most existing approaches commonly extract the features of all candidate video segments and then classify them separately. However, they may neglect the underlying relationship among candidates unconsciously. In this paper, we propose a novel model termed Candidate-Aware Aggregation (CAA) to tackle this problem. In CAA, we design the Global Awareness (GA) module to exploit long-range relations among all candidates from a global perspective, which enhances the features of action instances. The GA module is then embedded into a multi-level hierarchical network named FENet, to aggregate local features in adjacent candidates to suppress background noise. As a result, the relationship among candidates is explicitly captured from both local and global perspectives, which ensures more accurate prediction results for the candidates. Extensive experiments conducted on two popular benchmarks ActivityNet-1.3 and THUMOS-14 demonstrate the superiority of CAA comparing to the state-of-the-art methods.
Yifan Ren, Xing Xu 0001, Fumin Shen, Yazhou Yao, Huimin Lu 0001
ACM Multimedia4
2021 Video Representation Learning with Graph Contrastive Augmentation
abstract
Contrastive-based self-supervised learning for image representations has significantly closed the gap with supervised learning. A natural extension of image-based contrastive learning methods to the video domain is to fully exploit the temporal structure presented in videos. We propose a novel contrastive self-supervised video representation learning framework, termed Graph Contrastive Augmentation (GCA), by constructing a video temporal graph and devising a graph augmentation that is designed to enhance the correlation across frames of videos and developing a new view for exploring temporal structure in videos. Specifically, we construct the temporal graph in the video by leveraging the relational knowledge behind the correlated sequence video features. Afterwards, we apply the proposed graph augmentation to generate another graph view by cooperating random corruption of the original graph to enhance the diversity of the intrinsic structure of the temporal graph. To this end, we provide two different kinds of contrastive learning methods to train our framework using temporal relationships concealed in videos as self-supervised signals. We perform empirical experiments on downstream tasks, action recognition and video retrieval, using the learned video representation, and the results demonstrate that with the graph view of temporal structure, our proposed GCA remarkably improves performance against or on par with the recent methods.
Jingran Zhang, Xing Xu 0001, Fumin Shen, Yazhou Yao, Jie Shao 0001, Xiaofeng Zhu 0001
ACM Multimedia4
2021 Curriculum-Based Meta-learning
abstract
Meta-learning offers an effective solution to learn new concepts with scarce supervision through an episodic training scheme: a series of target-like tasks sampled from base classes are sequentially fed into a meta-learner to extract common knowledge across tasks, which can facilitate the quick acquisition of task-specific knowledge of the target task with few samples. Despite its noticeable improvements, the episodic training strategy samples tasks randomly and uniformly, without considering their hardness and quality, which may not progressively improve the meta-leaner's generalization ability. In this paper, we present a Curriculum-Based Meta-learning (CubMeta) method to train the meta-learner using tasks from easy to hard. Specifically, the framework of CubMeta is in a progressive way, and in each step, we design a module named BrotherNet to establish harder tasks and an effective learning scheme for obtaining an ensemble of stronger meta-learners. In this way, the meta-learner's generalization ability can be progressively improved, and better performance can be obtained even with fewer training tasks. We evaluate our method for few-shot classification on two benchmarks - mini-ImageNet and tiered-ImageNet, where it achieves consistent performance improvements on various meta-learning paradigms.
Ji Zhang 0012, Jingkuan Song, Yazhou Yao, Lianli Gao
ACM Multimedia3
2021 Extracting Useful Knowledge from Noisy Web Images via Data Purification for Fine-Grained Recognition
abstract
Fine-grained visual recognition tasks typically require training data with reliable acquisition and annotation processes. Acquiring such datasets with precise fine-grained annotations is very expensive and time-consuming. Conversely, a vast amount of web data is relatively easy to obtain with nearly no human effort. Nevertheless, the presence of label noise in web images becomes a huge obstacle for training robust fine-grained recognition models. In this work, we investigate the noisy label problem and propose a method that can specifically distinguish in- and out-of-distribution noisy samples. It can purify the web training data by discarding out-of-distribution noisy images and relabeling in-distribution ones. After purification, we can train the model on a less noisy web training set to achieve better robustness and performance. Extensive experiments on three real-world web datasets for fine-grained visual recognition demonstrate the superiority of our approach.
Chuanyi Zhang, Yazhou Yao, Xing Xu 0001, Jie Shao 0001, Jingkuan Song, Zechao Li, Zhenmin Tang
ACM Multimedia2
2021 Local Self-Attention on Fine-grained Cross-media Retrieval
abstract
Due to the heterogeneity gap, the data representation of different media is inconsistent and belongs to different feature spaces. Therefore, it is challenging to measure the fine-grained gap between them. To this end, we propose an attention space training method to learn common representations of different media data. Specifically, we utilize local self-attention layers to learn the common attention space between different media data. We propose a similarity concatenation method to understand the content relationship between features. To further improve the robustness of the model, we also train a local position encoding to capture the spatial relationships between features. In this way, our proposed method can effectively reduce the gap between different feature distributions on cross-media retrieval tasks. It also improves the fine-grained recognition performance by attaching attention to high-level semantic information. Extensive experiments and ablation studies demonstrate that our proposed method achieves state-of-the-art performance. At the same time, our approach provides a new pipeline for fine-grained cross-media retrieval. The source code and models are publicly available at: https://github.com/NUST-Machine-Intelligence-Laboratory/SAFGCMHN.
Yazhou Yao, Qiong Wang 0003, Zhenmin Tang
MMAsia2
2021 Knowledge memorization and generation for action recognition in still images
Wankou Yang, Yazhou Yao, Fatih Porikli
Pattern Recognit.3
2021 Exploiting textual queries for dynamically visual disambiguation
abstract
Due to the high cost of manual annotation, learning directly from the web has attracted broad attention. One issue that limits the performance of current webly supervised models is the problem of visual polysemy. In this work, we present a novel framework that resolves visual polysemy by dynamically matching candidate text queries with retrieved images. Specifically, our proposed framework includes three major steps: we first discover and then dynamically select the text queries according to the keyword-based image search results, we employ the proposed saliency-guided deep multi-instance learning (MIL) network to remove outliers and learn classification models for visual disambiguation. Compared to existing methods, our proposed approach can figure out the right visual senses, adapt to dynamic changes in the search results, remove outliers, and jointly learn the classification models . Extensive experiments and ablation studies on CMU-Poly-30 and MIT-ISD datasets demonstrate the effectiveness of our proposed approach.
Zeren Sun, Yazhou Yao, Jimin Xiao, Lei Zhang 0054, Jian Zhang 0002, Zhenmin Tang
Pattern Recognit.2
2021 VMAN: A Virtual Mainstay Alignment Network for Transductive Zero-Shot Learning
abstract
Transductive zero-shot learning (TZSL) extends conventional ZSL by leveraging (unlabeled) unseen images for model training. A typical method for ZSL involves learning embedding weights from the feature space to the semantic space. However, the learned weights in most existing methods are dominated by seen images, and can thus not be adapted to unseen images very well. In this paper, to align the (embedding) weights for better knowledge transfer between seen/unseen classes, we propose the virtual mainstay alignment network (VMAN), which is tailored for the transductive ZSL task. Specifically, VMAN is casted as a tied encoder-decoder net, thus only one linear mapping weights need to be learned. To explicitly learn the weights in VMAN, for the first time in ZSL, we propose to generate virtual mainstay (VM) samples for each seen class, which serve as new training data and can prevent the weights from being shifted to seen images, to some extent. Moreover, a weighted reconstruction scheme is proposed and incorporated into the model training phase, in both the semantic/feature spaces. In this way, the manifold relationships of the VM samples are well preserved. To further align the weights to adapt to more unseen images, a novel instance-category matching regularization is proposed for model re-training. VMAN is thus modeled as a nested minimization problem and is solved by a Taylor approximate optimization paradigm. In comprehensive evaluations on four benchmark datasets, VMAN achieves superior performances under the (Generalized) TZSL setting.
Guosen Xie, Xu-Yao Zhang, Yazhou Yao, Zheng Zhang 0006, Fang Zhao 0006, Ling Shao 0001
IEEE Trans. Image Process.3
2021 Deep Unsupervised Self-Evolutionary Hashing for Image Retrieval
abstract
Hashing methods have proven to be effective in the field of large-scale image retrieval. In recent years, the performance of hashing algorithms based on deep learning has greatly exceeded that of non-deep methods. However, most of the outstanding hashing methods are supervised models that heavily rely on annotated labels. In order to circumvent the huge overhead of labeling large-scale datasets, some unsupervised hashing algorithms have been proposed, such as pseudo labels and pseudo pairs. Since the image labels are strictly unavailable, some hyper-parameters in these methods are difficult to be selected, e.g., the final result is very sensitive to the picked number of categories or the chosen threshold of similarity for pairs. In addition, the calculation of pseudo-labels in high-dimensional space is not only computationally complex, but also has low precision. Therefore, in order to alleviate these issues in this paper, we propose a simple but effective Deep Unsupervised Self-evolutionary Hashing (DUSH) algorithm, which utilizes a curriculum learning strategy to iteratively select pseudo pairs from easy to hard in low dimensional Hamming space. Extensive experiments are conducted on four popular datasets, including two single-label datasets and two multi-label datasets, and the results show that our method can significantly outperform the state-of-the-art methods.
Haofeng Zhang 0001, Yazhou Yao, Zheng Zhang 0006, Li Liu 0004, Jian Zhang 0002, Ling Shao 0001
IEEE Trans. Multim.3
2020 Web-Supervised Network with Softly Update-Drop Training for Fine-Grained Visual Classification
abstract
Labeling objects at the subordinate level typically requires expert knowledge, which is not always available from a random annotator. Accordingly, learning directly from web images for fine-grained visual classification (FGVC) has attracted broad attention. However, the existence of noise in web images is a huge obstacle for training robust deep neural networks. In this paper, we propose a novel approach to remove irrelevant samples from the real-world web images during training, and only utilize useful images for updating the networks. Thus, our network can alleviate the harmful effects caused by irrelevant noisy web images to achieve better performance. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is much superior to state-of-the-art webly supervised methods. The data and source code of this work have been made anonymously available at: https://github.com/z337-408/WSNFGVC.
Chuanyi Zhang, Yazhou Yao, Huafeng Liu 0004, Guosen Xie, Xiangbo Shu, Tianfei Zhou, Zheng Zhang 0006, Fumin Shen, Zhenmin Tang
AAAI2
2020 Motion-Attentive Transition for Zero-Shot Video Object Segmentation
abstract
In this paper, we present a novel Motion-Attentive Transition Network (MATNet) for zero-shot video object segmentation, which provides a new way of leveraging motion information to reinforce spatio-temporal object representation. An asymmetric attention block, called Motion-Attentive Transition (MAT), is designed within a two-stream encoder, which transforms appearance features into motion-attentive representations at each convolutional stage. In this way, the encoder becomes deeply interleaved, allowing for closely hierarchical interactions between object motion and appearance. This is superior to the typical two-stream architecture, which treats motion and appearance separately in each stream and often suffers from overfitting to appearance information. Additionally, a bridge network is proposed to obtain a compact, discriminative and scale-sensitive representation for multi-level encoder features, which is further fed into a decoder to achieve segmentation results. Extensive experiments on three challenging public benchmarks (i.e., DAVIS-16, FBMS and Youtube-Objects) show that our model achieves compelling performance against the state-of-the-arts. Code is available at: https://github.com/tfzhou/MATNet.
Tianfei Zhou, Shunzhou Wang, Yi Zhou 0007, Yazhou Yao, Jianwu Li, Ling Shao 0001
AAAI4
2020 Region Graph Embedding Network for Zero-Shot Learning
Guosen Xie, Li Liu 0004, Fan Zhu 0001, Fang Zhao 0006, Zheng Zhang 0006, Yazhou Yao, Jie Qin 0004, Ling Shao 0001
ECCV (4)6
2020 Classification Constrained Discriminator For Domain Adaptive Semantic Segmentation
abstract
Unsupervised domain adaptation for semantic segmentation aims to transfer knowledge from label-rich synthetic datasets to real-world images without any annotation. The traditional adversarial learning methods for domain adaptation learn to extract domain-invariant feature representations by aligning the feature distributions of both domains. However, these methods suffer from an imbalance in adversarial training and feature distortion. In this work, we propose a classification constrained discriminator to alleviate these problems. Specifically, we first propose to balance the adversarial training by eliminating any pooling layers or strided convolutions in the discriminator. Then, we propose to constrain the discriminator with an auxiliary classification loss to help the feature generator extract the domain-invariant features that are useful for segmentation rather than just ambiguous features to fool the domain discriminator. Extensive experiments demonstrate the superiority of our proposed approach. The source code and models have been made available at https://github.com/NUSTMachine-Intelligence-Laboratory/ccd.
Tao Chen 0012, Jian Zhang 0002, Guosen Xie, Yazhou Yao, Xiaoshui Huang, Zhenmin Tang
ICME4
2020 Hsi Road: A Hyper Spectral Image Dataset For Road Segmentation
abstract
Road segmentation is a challenging task in the field of self-driving research. This paper present a road dataset built by hyper spectral imaging (HSI) cameras instead of the widely-used RGB cameras. HSI image is informative in spectrums and full of potential for natural environment perception. In this article, a first-of-its-kind HSI road segmentation dataset is built with careful annotation in both urban and rural scenes. It contains 3799 scenes with RGB and NIR bands as well as their respective masks. Unlike many existing datasets that provide urban scenes in RGB images only, our dataset expands the sensing spectrum to 28 bands and includes various kinds of road surfaces, such as asphalt, cement, dirt and sand, under rural and natural scenes. We also provide benchmark performances based on the recently popular segmentation algorithms on this dataset. The dataset is released at github‡.‡https://github.com/NUST-Machine-Intelligence-Laboratory/hsi_road
Jiarou Lu, Huafeng Liu 0004, Yazhou Yao, Shuyin Tao, Zhenmin Tang, Jianfeng Lu 0003
ICME3
2020 Web-Supervised Network for Fine-Grained Visual Classification
abstract
Fine-grained visual classification (FGVC) is a tough task due to its high annotation cost of the fine-grained subcategories. To build a large-scale dataset at low manual cost, straightforwardly learning from web images for FGVC has attracted broad attention. However, there exist two characteristics in the need of concerning for the web dataset: 1) Noisy images; 2) A large proportion of hard examples. In this paper, we propose a simple yet effective approach to deal with noisy images and hard examples during training. Our method is a pure web-supervised method for FGVC. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is much superior to the state-of-the-art web-supervised methods. The data and source code of this work have been posted available at: https://github.com/NUST-Machine-Intelligence-Laboratory/WSNFG.
Chuanyi Zhang, Yazhou Yao, Jiachao Zhang, Jian Zhang 0002, Zhenmin Tang
ICME2
2020 Set and Rebase: Determining the Semantic Graph Connectivity for Unsupervised Cross-Modal Hashing
abstract
The label-free nature of unsupervised cross-modal hashing hinders models from exploiting the exact semantic data similarity. Existing research typically simulates the semantics by a heuristic geometric prior in the original feature space. However, this introduces heavy bias into the model as the original features are not fully representing the underlying multi-view data relations. To address the problem above, in this paper, we propose a novel unsupervised hashing method called Semantic-Rebased Cross-modal Hashing (SRCH). A novel ‘Set-and-Rebase’ process is defined to initialize and update the cross-modal similarity graph of training data. In particular, we set the graph according to the intra-modal feature geometric basis and then alternately rebase it to update the edges within according to the hashing results. We develop an alternating optimization routine to rebase the graph and train the hashing auto-encoders with closed-form solutions so that the overall framework is efficiently trained. Our experimental results on benchmarked datasets demonstrate the superiority of our model against state-of-the-art algorithms.
Yuming Shen, Haofeng Zhang 0001, Yazhou Yao, Li Liu 0004
IJCAI4
2020 PyRetri: A PyTorch-based Library for Unsupervised Image Retrieval by Deep Convolutional Neural Networks
abstract
Despite significant progress of applying deep learning methods to the field of content-based image retrieval, there has not been a software library that covers these methods in a unified manner. In order to fill this gap, we introduce PyRetri, an open source library for deep learning based unsupervised image retrieval. The library encapsulates the retrieval process in several stages and provides functionality that covers various prominent methods for each stage. The idea underlying its design is to provide a unified platform for deep learning based image retrieval research, with high usability and extensibility. The project source code, with usage examples, sample data and pre-trained models are available at https://github.com/PyRetri/.
Benyi Hu, Renjie Song, Xiu-Shen Wei, Yazhou Yao, Xian-Sheng Hua 0001, Yuehu Liu
ACM Multimedia4
2020 CRSSC: Salvage Reusable Samples from Noisy Data for Robust Learning
abstract
Due to the existence of label noise in web images and the high memorization capacity of deep neural networks, training deep fine-grained (FG) models directly through web images tends to have an inferior recognition ability. In the literature, to alleviate this issue, loss correction methods try to estimate the noise transition matrix, but the inevitable false correction would cause severe accumulated errors. Sample selection methods identify clean ("easy") samples based on the fact that small losses can alleviate the accumulated errors. However, "hard" and mislabeled examples that can both boost the robustness of FG models are also dropped. To this end, we propose a certainty-based reusable sample selection and correction approach, termed as CRSSC, for coping with label noise in training deep FG models with web images. Our key idea is to additionally identify and correct reusable samples, and then leverage them together with clean examples to update the networks. We demonstrate the superiority of the proposed approach from both theoretical and experimental perspectives.
Zeren Sun, Xian-Sheng Hua 0001, Yazhou Yao, Xiu-Shen Wei, Guosheng Hu, Jian Zhang 0002
ACM Multimedia3
2020 Bridging the Web Data and Fine-Grained Visual Recognition via Alleviating Label Noise and Domain Mismatch
abstract
To distinguish the subtle differences among fine-grained categories, a large amount of well-labeled images are typically required. However, manual annotations for fine-grained categories is an extremely difficult task as it usually has a high demand for professional knowledge. To this end, we propose to directly leverage web images for fine-grained visual recognition. Our work mainly focuses on two critical issues including "label noise" and "domain mismatch" in the web images. Specifically, we propose an end-to-end deep denoising network (DDN) model to jointly solve these problems in the process of web images selection. To verify the effectiveness of our proposed approach, we first collect web images by using the labels in fine-grained datasets. Then we apply the proposed deep denoising network model for noise removal and domain mismatch alleviation. We leverage the selected web images as the training set for fine-grained categorization models learning. Extensive experiments and ablation studies demonstrate state-of-the-art performance gained by our proposed approach, which, at the same time, delivers a new pipeline for fine-grained visual categorization that is to be highly effective for real-world applications.
Yazhou Yao, Xian-Sheng Hua 0001, Guanyu Gao, Zeren Sun, Zhibin Li 0002, Jian Zhang 0002
ACM Multimedia1
2020 Data-driven Meta-set Based Fine-Grained Visual Recognition
abstract
Constructing fine-grained image datasets typically requires domain-specific expert knowledge, which is not always available for crowd-sourcing platform annotators. Accordingly, learning directly from web images becomes an alternative method for fine-grained visual recognition. However, label noise in the web training set can severely degrade the model performance. To this end, we propose a data-driven meta-set based approach to deal with noisy web images for fine-grained recognition. Specifically, guided by a small amount of clean meta-set, we train a selection net in a meta-learning manner to distinguish in- and out-of-distribution noisy images. To further boost the robustness of the model, we also learn a labeling net to correct the labels of in-distribution noisy data. In this way, our proposed method can alleviate the harmful effects caused by out-of-distribution noise and properly exploit the in-distribution noisy samples for training. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is much superior to state-of-the-art noise-robust methods.
Chuanyi Zhang, Yazhou Yao, Xiangbo Shu, Zechao Li, Zhenmin Tang, Qi Wu 0001
ACM Multimedia2
2020 Field-wise Learning for Multi-field Categorical Data
abstract
We propose a new method for learning with multi-field categorical data. Multi-field categorical data are usually collected over many heterogeneous groups. These groups can reflect in the categories under a field. The existing methods try to learn a universal model that fits all data, which is challenging and inevitably results in learning a complex model. In contrast, we propose a field-wise learning method leveraging the natural structure of data to learn simple yet efficient one-to-one field-focused models with appropriate constraints. In doing this, the models can be fitted to each category and thus can better capture the underlying differences in data. We present a model that utilizes linear models with variance and low-rank constraints, to help it generalize better and reduce the number of parameters. The model is also interpretable in a field-wise manner. As the dimensionality of multi-field categorical data can be very high, the models applied to such data are mostly over-parameterized. Our theoretical analysis can potentially explain the effect of over-parametrization on the generalization of our model. It also supports the variance constraints in the learning objective. The experiment results on two large-scale datasets show the superior performance of our model, the trend of the generalization error bound, and the interpretability of learning outcomes. Our code is available at https://github.com/lzb5600/Field-wise-Learning.
Zhibin Li 0002, Jian Zhang 0002, Yongshun Gong, Yazhou Yao, Qiang Wu 0001
NeurIPS4
2020 Multi-model Network for Fine-Grained Cross-Media Retrieval
Jiemi Bai, Yazhou Yao, Qiong Wang 0003, Wankou Yang, Fumin Shen
PRCV (2)2
2020 A Novel CNN Architecture for Real-Time Point Cloud Recognition in Road Environment
Duyao Fan, Yazhou Yao, Yunfei Cai, Xiangbo Shu, Wankou Yang
PRCV (1)2
2020 Road segmentation with image-LiDAR data fusion in deep neural network
Huafeng Liu 0004, Yazhou Yao, Zeren Sun, Xiangrui Li, Ke Jia, Zhenmin Tang
Multim. Tools Appl.2
2020 CAN-GAN: Conditioned-attention normalized GAN for face age synthesis
Chenglong Shi, Jiachao Zhang, Yazhou Yao, Yunlian Sun, Huaming Rao, Xiangbo Shu
Pattern Recognit. Lett.3
2020 Pseudo distribution on unseen classes for generalized zero shot learning
Haofeng Zhang 0001, Jingren Liu, Yazhou Yao, Yang Long 0001
Pattern Recognit. Lett.3
2020 Towards Automatic Construction of Diverse, High-Quality Image Datasets
abstract
The availability of labeled image datasets has been shown critical for high-level image understanding, which continuously drives the progress of feature designing and models developing. However, constructing labeled image datasets is laborious and monotonous. To eliminate manual annotation, in this work, we propose a novel image dataset construction framework by employing multiple textual queries. We aim at collecting diverse and accurate images for given queries from the Web. Specifically, we formulate noisy textual queries removing and noisy images filtering as a multi-view and multi-instance learning problem separately. Our proposed approach not only improves the accuracy but also enhances the diversity of the selected images. To verify the effectiveness of our proposed approach, we construct an image dataset with 100 categories. The experiments show significant performance gains by using the generated data of our approach on several tasks, such as image classification, cross-dataset generalization, and object detection. The proposed method also consistently outperforms existing weakly supervised and web-supervised approaches.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Li Liu 0004, Fan Zhu 0001, Dongxiang Zhang, Heng Tao Shen
IEEE Trans. Knowl. Data Eng.1
2020 Approximate Kernel Selection via Matrix Approximation
abstract
Kernel selection is of fundamental importance for the generalization of kernel methods. This article proposes an approximate approach for kernel selection by exploiting the approximability of kernel selection and the computational virtue of kernel matrix approximation. We define approximate consistency to measure the approximability of the kernel selection problem. Based on the analysis of approximate consistency, we solve the theoretical problem of whether, under what conditions, and at what speed, the approximate criterion is close to the accurate one, establishing the foundations of approximate kernel selection. We introduce two selection criteria based on error estimation and prove the approximate consistency of the multilevel circulant matrix (MCM) approximation and Nyström approximation under these criteria. Under the theoretical guarantees of the approximate consistency, we design approximate algorithms for kernel selection, which exploits the computational advantages of the MCM and Nyström approximations to conduct kernel selection in a linear or quasi-linear complexity. We experimentally validate the theoretical results for the approximate consistency and evaluate the effectiveness of the proposed kernel selection algorithms.
Lizhong Ding 0001, Shizhong Liao, Yong Liu 0018, Li Liu 0004, Fan Zhu 0001, Yazhou Yao, Ling Shao 0001, Xin Gao 0001
IEEE Trans. Neural Networks Learn. Syst.6
2020 Exploiting Web Images for Multi-Output Classification: From Category to Subcategories
abstract
Studies present that dividing categories into subcategories contributes to better image classification. Existing image subcategorization works relying on expert knowledge and labeled images are both time-consuming and labor-intensive. In this article, we propose to select and subsequently classify images into categories and subcategories. Specifically, we first obtain a list of candidate subcategory labels from untagged corpora. Then, we purify these subcategory labels through calculating the relevance to the target category. To suppress the search error and noisy subcategory label-induced outlier images, we formulate outlier images removing and the optimal classification models learning as a unified problem to jointly learn multiple classifiers, where the classifier for a category is obtained by combining multiple subcategory classifiers. Compared with the existing subcategorization works, our approach eliminates the dependence on expert knowledge and labeled images. Extensive experiments on image categorization and subcategorization demonstrate the superiority of our proposed approach.
Yazhou Yao, Fumin Shen, Guosen Xie, Li Liu 0004, Fan Zhu 0001, Jian Zhang 0002, Heng Tao Shen
IEEE Trans. Neural Networks Learn. Syst.1
2019 Attentive Region Embedding Network for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to classify images from unseen categories, by merely utilizing seen class images as the training data. Existing works on ZSL mainly leverage the global features or learn the global regions, from which, to construct the embeddings to the semantic space. However, few of them study the discrimination power implied in local image regions (parts), which, in some sense, correspond to semantic attributes, have stronger discrimination than attributes, and can thus assist the semantic transfer between seen/unseen classes. In this paper, to discover (semantic) regions, we propose the attentive region embedding network (AREN), which is tailored to advance the ZSL task. Specifically, AREN is end-to-end trainable and consists of two network branches, i.e., the attentive region embedding (ARE) stream, and the attentive compressed second-order embedding (ACSE) stream. ARE is capable of discovering multiple part regions under the guidance of the attention and the compatibility loss. Moreover, a novel adaptive thresholding mechanism is proposed for suppressing redundant (such as background) attention regions. To further guarantee more stable semantic transfer from the perspective of second-order collaboration, ACSE is incorporated into the AREN. In the comprehensive evaluations on four benchmarks, our models achieve state-of-the-art performances under ZSL setting, and compelling results under generalized ZSL setting.
Guosen Xie, Li Liu 0004, Xiao-Bo Jin, Fan Zhu 0001, Zheng Zhang 0006, Jie Qin 0004, Yazhou Yao, Ling Shao 0001
CVPR7
2019 SegEQA: Video Segmentation Based Visual Attention for Embodied Question Answering
abstract
Embodied Question Answering (EQA) is a newly defined research area where an agent is required to answer the user's questions by exploring the real world environment. It has attracted increasing research interests due to its broad applications in automatic driving system, in-home robots, and personal assistants. Most of the existing methods perform poorly in terms of answering and navigation accuracy due to the absence of local details and vulnerability to the ambiguity caused by complicated vision conditions. To tackle these problems, we propose a segmentation based visual attention mechanism for Embodied Question Answering. Firstly, We extract the local semantic features by introducing a novel high-speed video segmentation framework. Then by the guide of extracted semantic features, a bottom-up visual attention mechanism is proposed for the Visual Question Answering (VQA) sub-task. Further, a feature fusion strategy is proposed to guide the training of the navigator without much additional computational cost. The ablation experiments show that our method boosts the performance of VQA module by 4.2% (68.99% vs 64.73%) and leads to 3.6% (48.59% vs 44.98%) overall improvement in EQA accuracy.
Haonan Luo 0002, Guosheng Lin, Zichuan Liu, Fayao Liu, Zhenmin Tang, Yazhou Yao
ICCV6
2019 Dynamically Visual Disambiguation of Keyword-based Image Search
abstract
Due to the high cost of manual annotation, learning directly from the web has attracted broad attention. One issue that limits their performance is the problem of visual polysemy. To address this issue, we present an adaptive multi-model framework that resolves polysemy by visual disambiguation. Compared to existing methods, the primary advantage of our approach lies in that our approach can adapt to the dynamic changes in the search results. Our proposed framework consists of two major steps: we first discover and dynamically select the text queries according to the image search results, then we employ the proposed saliency-guided deep multi-instance learning network to remove outliers and learn classification models for visual disambiguation. Extensive experiments demonstrate the superiority of our proposed approach.
Yazhou Yao, Zeren Sun, Fumin Shen, Li Liu 0004, Limin Wang 0002, Fan Zhu 0001, Lizhong Ding 0001, Gangshan Wu, Ling Shao 0001
IJCAI1
2019 Clustering-driven unsupervised deep hashing for image retrieval
Haofeng Zhang 0001, Yazhou Yao, Wankou Yang, Li Liu 0004
Neurocomputing4
2019 Deep representation learning for road detection using Siamese network
Huafeng Liu 0004, Xiangrui Li, Yazhou Yao, Zhenmin Tang
Multim. Tools Appl.4
2019 Exploiting textual and visual features for image categorization
Yazhou Yao, Wankou Yang, Qiong Wang 0003, Yunfei Cai, Zhenmin Tang
Pattern Recognit. Lett.1
2019 Extracting Privileged Information for Enhancing Classifier Learning
abstract
The accuracy of data-driven learning approaches is often unsatisfactory when the training data is inadequate either in quantity or quality. Manually labeled privileged information (PI), e.g., attributes, tags or properties, is usually incorporated to improve classifier learning. However, the process of manually labeling is time-consuming and labor-intensive. Moreover, due to the limitations of personal knowledge, manually labeled PI may not be rich enough. To address these issues, we propose to enhance classifier learning by exploring PI from untagged corpora, which can effectively eliminate the dependency on manually labeled data and obtain much richer PI. In detail, we treat each selected PI as a subcategory and learn one classifier for each subcategory independently. The classifiers for all subcategories are integrated together to form a more powerful category classifier. Particularly, we propose a novel instancelevel multi-instance learning (MIL) model to simultaneously select a subset of training images from each subcategory and learn the optimal SVM classifiers based on the selected images. Extensive experiments on four benchmark datasets demonstrate the superiority of our proposed approach.
Yazhou Yao, Fumin Shen, Jian Zhang 0002, Li Liu 0004, Zhenmin Tang, Ling Shao 0001
IEEE Trans. Image Process.1
2019 Extracting Multiple Visual Senses for Web Learning
abstract
Labeled image datasets have played a critical role in high-level image understanding. However, the process of manual labeling is both time consuming and labor intensive. To reduce the dependence on manually labeled data, there have been increasing research efforts on learning visual classifiers by directly exploiting web images. One issue that limits their performance is the problem of polysemy. Existing unsupervised approaches attempt to reduce the influence of visual polysemy by filtering out irrelevant images, but do not directly address polysemy. To this end, in this paper, we present a multimodal framework that solves the problem of polysemy by allowing sense-specific diversity in search results. Specifically, we first discover a list of possible semantic senses from untagged corpora to retrieve sense-specific images. Then, we merge visual similar semantic senses and prune noise by using the retrieved images. Finally, we train one visual classifier for each selected semantic sense and use the learned sense-specific classifiers to distinguish multiple visual senses. Extensive experiments on classifying images into sense-specific categories and reranking search results demonstrate the superiority of our proposed approach.
Yazhou Yao, Fumin Shen, Jian Zhang 0002, Li Liu 0004, Zhenmin Tang, Ling Shao 0001
IEEE Trans. Multim.1
2018 Discovering and Distinguishing Multiple Visual Senses for Polysemous Words
abstract
To reduce the dependence on labeled data, there have been increasing research efforts on learning visual classifiers by exploiting web images. One issue that limits their performance is the problem of polysemy. To solve this problem, in this work, we present a novel framework that solves the problem of polysemy by allowing sense-specific diversity in search results. Specifically, we first discover a list of possible semantic senses to retrieve sense-specific images. Then we merge visual similar semantic senses and prune noises by using the retrieved images. Finally, we train a visual classifier for each selected semantic sense and use the learned sense-specific classifiers to distinguish multiple visual senses. Extensive experiments on classifying images into sense-specific categories and re-ranking search results demonstrate the superiority of our proposed approach.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Wankou Yang, Zhenmin Tang
AAAI1
2018 Extracting Privileged Information from Untagged Corpora for Classifier Learning
abstract
The performance of data-driven learning approaches is often unsatisfactory when the training data is inadequate either in quantity or quality. Manually labeled privileged information (PI), \eg attributes, tags or properties, is usually incorporated to improve classifier learning. However, the process of manually labeling is time-consuming and labor-intensive. To address this issue, we propose to enhance classifier learning by extracting PI from untagged corpora, which can effectively eliminate the dependency on manually labeled data. In detail, we treat each selected PI as a subcategory and learn one classifier for per subcategory independently. The classifiers for all subcategories are then integrated together to form a more powerful category classifier. Particularly, we propose a new instance-level multi-instance learning (MIL) model to simultaneously select a subset of training images from each subcategory and learn the optimal classifiers based on the selected images. Extensive experiments demonstrate the superiority of our approach.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Wankou Yang, Xian-Sheng Hua 0001, Zhenmin Tang
IJCAI1
2017 A new web-supervised method for image dataset constructions
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Xian-Sheng Hua 0001, Jingsong Xu, Zhenmin Tang
Neurocomputing1
2017 Exploiting Web Images for Dataset Construction: A Domain Robust Approach
abstract
Labeled image datasets have played a critical role in high-level image understanding. However, the process of manual labeling is both time-consuming and labor intensive. To reduce the cost of manual labeling, there has been increased research interest in automatically constructing image datasets by exploiting web images. Datasets constructed by existing methods tend to have a weak domain adaptation ability, which is known as the “dataset bias problem.” To address this issue, we present a novel image dataset construction framework that can be generalized well to unseen target domains. Specifically, the given queries are first expanded by searching the Google Books Ngrams Corpus to obtain a rich semantic description, from which the visually nonsalient and less relevant expansions are filtered out. By treating each selected expansion as a “bag” and the retrieved images as “instances,” image selection can be formulated as a multi-instance learning problem with constrained positive bags. We propose to solve the employed problems by the cutting-plane and concave-convex procedure algorithm. By using this approach, images from different distributions can be kept while noisy images are filtered out. To verify the effectiveness of our proposed approach, we build an image dataset with 20 categories. Extensive experiments on image classification, cross-dataset generalization, diversity comparison, and object detection demonstrate the domain robustness of our dataset.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Xian-Sheng Hua 0001, Jingsong Xu, Zhenmin Tang
IEEE Trans. Multim.1
2016 Automatic image dataset construction with multiple textual metadata
abstract
The goal of this work is to automatically collect a large number of highly relevant images from the Internet for given queries. A novel image dataset construction framework is proposed by employing multiple textual metadata. In specific, the given queries are first expanded by searching in the Google Books Ngrams Corpora to obtain a richer semantic description, from which the visually non-salient and less relevant expansions are then filtered. After retrieving images from the Internet with filtered expansions, we further filter noisy images by clustering and progressively Convolutional Neural Networks (CNN). To verify the effectiveness of our proposed method, we construct a dataset with 10 categories, which is not only much larger than but also have comparable cross-dataset generalization ability with manually labeled dataset STL-10 and CIFAR-10.
Yazhou Yao, Jian Zhang 0002, Fumin Shen, Xian-Sheng Hua 0001, Jingsong Xu, Zhenmin Tang
ICME1
2016 A Domain Robust Approach For Image Dataset Construction
abstract
There have been increasing research interests in automatically constructing image dataset by collecting images from the Internet. However, existing methods tend to have a weak domain adaptation ability, known as the "dataset bias problem". To address this issue, in this work, we propose a novel image dataset construction framework which can generalize well to unseen target domains. In specific, the given queries are first expanded by searching in the Google Books Ngrams Corpora (GBNC) to obtain a richer semantic description, from which the noisy query expansions are then filtered out. By treating each expansion as a "bag" and the retrieved images therein as "instances", we formulate image filtering as a multi-instance learning (MIL) problem with constrained positive bags. By this approach, images from different data distributions will be kept while with noisy images filtered out. Comprehensive experiments on two challenging tasks demonstrate the effectiveness of our proposed approach.
Yazhou Yao, Xian-Sheng Hua 0001, Fumin Shen, Jian Zhang 0002, Zhenmin Tang
ACM Multimedia1
2016 Extracting Visual Knowledge from the Internet: Making Sense of Image Data
Yazhou Yao, Jian Zhang 0002, Xian-Sheng Hua 0001, Fumin Shen, Zhenmin Tang
MMM (1)1