VLDB 2026 Research / reviewers in the wild / expert
Huaxi Huang
dblp:184/0802
· DBLP profile ↗
11ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0002-6837-6747ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Establishing Nuanced Multimodal Attention for Weakly Supervised Semantic Segmentation of Remote Sensing ScenesabstractWeakly Supervised Semantic Segmentation (WSSS) with image-level labels reduces reliance on pixel-level annotations for remote sensing (RS) imagery. However, in natural scenes, WSSS frequently faces challenges such as imprecise localization, extraneous activations, and class ambiguity. These challenges are particularly pronounced in RS images, characterized by complex backgrounds, substantial scale variations, and dense small-object distributions, complicating the distinction between intra-class variations and inter-class similarities. To tackle these challenges, we introduce a class-constrained multi-modal attention framework aimed at enhancing the localization accuracy of class activation maps (CAMs). Specifically, we design class-specific tokens to capture the visual characteristics of each target class. As these tokens initially lack explicit constraints, we integrate the textual branch of the RemoteCLIP model to leverage class-related linguistic priors, which collaborate with visual features to encode the specific semantics of diverse objects. Furthermore, the multi-modal collaborative optimization module dynamically establishes tailored attention mechanisms for both global and regional features, thereby improving class discriminability among targets to mitigate challenges like inter-class similarity and dense small-object distributions. By refining class-specific attention, textual semantic attention, and patch-level pairwise affinity weights, the quality of generated pseudo-masks is markedly enhanced. Concurrently, to ensure domain-invariant feature learning, we align the backbone features with the CLIP visual embedding by minimizing the distribution disparity between the two in the latent space, semantic consistency is therefore preserved. The experimental results validate the effectiveness and robustness of our proposed method, achieving significant performance improvements on two representative RS WSSS datasets. Junjie Zhang 0002, Huaxi Huang, Fangyu Wu 0001, Hongwen Yu |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | Hierarchical Multi-Prototype Discrimination: Boosting Support-Query Matching for Few-Shot SegmentationabstractFew-shot segmentation (FSS) aims at training a model on base classes with sufficient annotations and then tasking the model with predicting a binary mask to identify novel class pixels with limited labeled images. Mainstream FSS methods adopt a support-query matching paradigm that activates target regions of the query image according to their similarity with a single support class prototype. However, this prototype vector is inclined to overfit the support images, leading to potential under-matching in latent query object regions and incorrect mismatches with base class features in the query image. To address these issues, this study reformulates conventional single foreground prototype matching to a multi-prototype matching paradigm. In this paradigm, query features exhibiting high confidence with non-target prototypes will be categorized as background. Specifically, the target query features are drawn closer to the novel class prototype through a Masked Cross-Image Encoding (MCE) module and a Semantic Multi-prototype Matching (SMM) module is incorporated to collaboratively filter unexpected base class regions on multi-scale features. Furthermore, we devise an adaptive class activation map, termed target-aware class activation map (TCAM) to preserve semantically coherent regions that might be inadvertently suppressed under pixel-wise matching guidance. Experimental results on PASCAL-5$^{i}$and COCO-20$^{i}$datasets demonstrate the advantage of the proposed novel modules, with the holistic approach outperforming compared state-of-the-art methods. Wenbo Xu 0004, Huaxi Huang, Yongshun Gong, Litao Yu, Qiang Wu 0001, Jian Zhang 0002 |
IEEE Trans. Multim. | 2 |
| 2024 | Few-shot classification guided by generalization error bound
Fan Liu 0003, Sai Yang, Delong Chen, Huaxi Huang, Jun Zhou 0001 |
Pattern Recognit. | 4 |
| 2024 | Few-Shot Classification Model Compression via School LearningabstractFew-shot classification (FSC) is a challenging task due to limitation in accessing training data. Recent methods often employ highly complex networks to obtain high-quality features, but this may not be suitable for resource-limited applications. To tackle this challenge, we introduce Few-Shot Classification Model Compression (FSC-MC), a new task aimed at enhancing the FSC performance of lightweight and low-capacity models by learning from more complex models. We also propose a novel two-level learning strategy called School Learning to accomplish the FSC-MC task by mimicking the real learning process in the social school life. In this new learning paradigm, the first level performs preview learning, in which each student is equipped with a preparer to perform self-learning on the base set. The second level is the team learning, consisting of a complex teacher network and several lightweight student networks organized into a team. One student network is randomly chosen as the leader network, while the remaining student networks serve as member networks. The leader network simultaneously learns knowledge from the teacher network and all member networks. Conversely, each member network receives knowledge from both the teacher network and the leader network. Ultimately, the leader network is deployed for FSC evaluation, resulting in effective model compression. Extensive experiments in the FSC-MC setting demonstrate that School Learning outperforms 17 state-of-the-art knowledge distillation methods including both offline methods and online methods, enabling lightweight models to achieve outstanding FSC performance. Sai Yang, Fan Liu 0003, Delong Chen, Huaxi Huang, Jun Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | PADDLES: Phase-Amplitude Spectrum Disentangled Early Stopping for Learning with Noisy LabelsabstractConvolutional Neural Networks (CNNs) are powerful in learning patterns of different vision tasks, but they are sensitive to label noise and may overfit to noisy labels during training. The early stopping strategy averts updating CNNs during the early training phase and is widely employed in the presence of noisy labels. Motivated by biological findings that the amplitude spectrum (AS) and phase spectrum (PS) in the frequency domain play different roles in the animal’s vision system, we observe that PS, which captures more semantic information, can increase the robustness of CNNs to label noise, more so than AS can. We thus propose early stops at different times for AS and PS by disentangling the features of some layer(s) into AS and PS using Discrete Fourier Transform (DFT) during training. Our proposed Phase-AmplituDe DisentangLed Early Stopping (PADDLES) method is shown to be effective on both synthetic and real-world label-noise datasets. PADDLES out-performs other early stopping methods and obtains state-of-the-art performance. Huaxi Huang, Olivier Salvado, Thierry Rakotoarivelo, Dadong Wang, Tongliang Liu |
ICCV | 1 |
| 2023 | Masked Cross-image Encoding for Few-shot SegmentationabstractFew-shot segmentation (FSS) is a dense prediction task that aims to infer the pixel-wise labels of unseen classes using only a limited number of annotated images. The key challenge in FSS is to classify the labels of query pixels using class prototypes learned from the few labeled support exemplars. Prior approaches to FSS have typically focused on learning class-wise descriptors independently from support images, thereby ignoring the rich contextual information and mutual dependencies among support-query features. To address this limitation, we propose a joint learning method termed Masked Cross-Image Encoding (MCE), which is designed to capture common visual properties that describe object details and to learn bidirectional inter-image dependencies that enhance feature interaction. MCE is more than a visual representation enrichment module; it also considers cross-image mutual dependencies and implicit guidance. Experiments on FSS benchmarks PASCAL-5iand COCO-20idemonstrate the advanced meta-learning ability of the proposed method. Wenbo Xu 0004, Huaxi Huang, Litao Yu, Qiang Wu 0001, Jian Zhang 0002 |
ICME | 2 |
| 2022 | TOAN: Target-Oriented Alignment Network for Fine-Grained Image Categorization With Few Labeled SamplesabstractIn this paper, we study the fine-grained categorization problem under the few-shot setting, i.e., each fine-grained class only contains a few labeled examples, termed Fine-Grained Few-Shot classification (FGFS). The core predicament in FGFS is the high intra-class variance yet low inter-class fluctuations in the dataset. In traditional fine-grained classification, the high intra-class variance can be somewhat relieved by conducting the supervised training on the abundant labeled samples. However, with few labeled examples, it is hard for the FGFS model to learn a robust class representation with the significantly higher intra-class variance. Moreover, the inter- and intra-class variance are closely related. The significant intra-class variance in FGFS often aggravates the low inter-class variance issue. To address the above challenges, we propose a Target-Oriented Alignment Network (TOAN) to tackle the FGFS problem from both intra- and inter-class perspective. To reduce the intra-class variance, we propose a target-oriented matching mechanism to reformulate the spatial features of each support image to match the query ones in the embedding space. To enhance the inter-class discrimination, we devise discriminative fine-grained features by integrating local compositional concept representations with the global second-order pooling. We conducted extensive experiments on four public datasets for fine-grained categorization, and the results show the proposed TOAN obtains the state-of-the-art. Huaxi Huang, Junjie Zhang 0002, Litao Yu, Jian Zhang 0002, Qiang Wu 0001, Chang Xu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | PTN: A Poisson Transfer Network for Semi-supervised Few-shot LearningabstractThe predicament in semi-supervised few-shot learning (SSFSL) is to maximize the value of the extra unlabeled data to boost the few-shot learner. In this paper, we propose a Poisson Transfer Network (PTN) to mine the unlabeled information for SSFSL from two aspects. First, the Poisson Merriman–Bence–Osher (MBO) model builds a bridge for the communications between labeled and unlabeled examples. This model serves as a more stable and informative classifier than traditional graph-based SSFSL methods in the message-passing process of the labels. Second, the extra unlabeled samples are employed to transfer the knowledge from base classes to novel classes through contrastive learning. Specifically, we force the augmented positive pairs close while push the negative ones distant. Our contrastive transfer scheme implicitly learns the novel-class embeddings to alleviate the over-fitting problem on the few labeled data. Thus, we can mitigate the degeneration of embedding generality in novel classes. Extensive experiments indicate that PTN outperforms the state-of-the-art few-shot and SSFSL models on miniImageNet and tieredImageNet benchmark datasets. Huaxi Huang, Junjie Zhang 0002, Jian Zhang 0002, Qiang Wu 0001, Chang Xu 0002 |
AAAI | 1 |
| 2021 | Low-Rank Pairwise Alignment Bilinear Network For Few-Shot Fine-Grained Image ClassificationabstractDeep neural networks have demonstrated advanced abilities on various visual classification tasks, which heavily rely on the large-scale training samples with annotated ground-truth. However, it is unrealistic always to require such annotation in real-world applications. Recently, Few-Shot learning (FS), as an attempt to address the shortage of training samples, has made significant progress in generic classification tasks. Nonetheless, it is still challenging for current FS models to distinguish the subtle differences between fine-grained categories given limited training data. To filling the classification gap, in this paper, we address the Few-Shot Fine-Grained (FSFG) classification problem, which focuses on tackling the fine-grained classification under the challenging few-shot learning setting. A novel low-rank pairwise bilinear pooling operation is proposed to capture the nuanced differences between the support and query images for learning an effective distance metric. Moreover, a feature alignment layer is designed to match the support image features with query ones before the comparison. We name the proposed model Low-Rank Pairwise Alignment Bilinear Network (LRPABN), which is trained in an end-to-end fashion. Comprehensive experimental results on four widely used fine-grained classification data sets demonstrate that our LRPABN model achieves the superior performances compared to state-of-the-art methods. Huaxi Huang, Junjie Zhang 0002, Jian Zhang 0002, Jingsong Xu, Qiang Wu 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | Compare More Nuanced: Pairwise Alignment Bilinear Network for Few-Shot Fine-Grained LearningabstractThe recognition ability of human beings is developed in a progressive way. Usually, children learn to discriminate various objects from coarse to fine-grained with limited supervision. Inspired by this learning process, we propose a simple yet effective model for the Few-Shot Fine-Grained (FSFG) recognition, which tries to tackle the challenging fine-grained recognition task using meta-learning. The proposed method, named Pairwise Alignment Bilinear Network (PABN), is an end-to-end deep neural network. Unlike traditional deep bilinear networks for fine-grained classification, which adopt the self-bilinear pooling to capture the subtle features of images, the proposed model uses a novel pairwise bilinear pooling to compare the nuanced differences between base images and query images for learning a deep distance metric. In order to match base image features with query image features, we design feature alignment losses before the proposed pairwise bilinear pooling. Experiment results on four fine-grained classification datasets and one generic few-shot dataset demonstrate that the proposed model outperforms both the state-of-the-art few-shot fine-grained and general few-shot methods. Huaxi Huang, Junjie Zhang 0002, Jian Zhang 0002, Qiang Wu 0001, Jingsong Xu |
ICME | 1 |
| 2016 | Multi-view Representative and Informative Induced Active Learning
Huaxi Huang, Changqing Zhang 0002, Qinghua Hu, Pengfei Zhu 0001 |
PRICAI | 1 |