Gaoxia Jiang

dblp:193/6994 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0002-2343-1132ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 8 first-author · 14 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 GCIB: Causal Intervention Guided Graph Information Bottleneck Framework
abstract
Graph neural networks (GNNs) have demonstrated impressive performance in a broad spectrum of fields, but always suffer from the generalization problem when confronted with out-of-distribution (OOD) scenarios. Information bottleneck (IB) principle, which endeavors to learn the minimally sufficient representations for downstream tasks, has been shown to be a promising strategy in dealing with this problem. However, the IB-based methods do not inherently distinguish between causal and non-causal parts in the graph, leading to underperforming OOD generalization ability. In this paper, we develop the Graph Causal Information Bottleneck (GCIB) framework, a causal extension of the IB for graph data, which is capable of jointly compressing abundant information and capturing causal dependency from the input graph. Specifically, we endow graph IB with the ability of maintaining causal control by incorporating the underlying causal structure and introducing intervention operation. On this basis, we formulate the learning objective for GCIB and present its specific implementation. Graph representations learned by GCIB can effectively preserve causal information that fundamentally determines graph properties, resulting in outstanding OOD generalization ability. Extensive experiments on both synthetic and real-world datasets demonstrate the superiority of GCIB over state-of-the-art baselines.
Hangyuan Du, Lixin Cui, Gaoxia Jiang, Liang Bai 0001, Wenjian Wang 0001
AAAI4
2026 CLIP-VC: Augmented prompt integrating visual components for weakly supervised semantic segmentation
Gaoxia Jiang
Pattern Recognit.3
2026 A multi-model dynamic filtering of label noise for regression
Gaoxia Jiang, Wenjian Wang 0001
Pattern Recognit.1
2026 A Dual Correction Guarantee Mechanism for Numerical Label Noise
Yaqing Guo, Lingxuan Cui, Gaoxia Jiang, Senyu Hou, Hang Xu 0009, Wenjian Wang 0001
IEEE Trans. Knowl. Data Eng.3
2026 Class-Aware Multi-Granularity Co-Diffusion Models for Learning With Noisy Labels on Imbalanced Datasets
abstract
Data quality is essential for the performance of deep neural networks in various fields. However, label noise and class imbalance are common data issues, which cause deep learning models to overfit in real-world scenarios. Recent research solves the learning with noisy labels (LNL) problem by employing label correction or loss adjustment methods, which often rely on uncertainty estimation. Unfortunately, these methods usually do not work well with imbalanced datasets. To address both noisy and imbalanced data biases, we analyze the limitations of current discriminative models in uncertainty estimation, and propose Class-aware Multi-granularity Co-Diffusion models (CaMCoD), which leverage a generative uncertainty and inconsistency loss adjustment method to generate labels more robustly. Specifically, we reframe the LNL problem as a robust diffusion-generative process, i.e., labels are generated by gradually refining an initial random guess. First, we use coarse-grained uncertainty from the diffusion model to achieve more accurate confidence estimates. This will guide the model to generate correct labels on a broader level. Then, we leverage the fine-grained inconsistency of co-diffusion models during reverse denoising to determine the learnable weight for each sample, which can mitigate the risk of the model overfitting to noisy samples. Finally, we apply class-aware loss adjustments to reduce data bias caused by class imbalance. Experiments on both synthetic and real-world datasets demonstrate that our method perform well in imbalanced and noisy scenarios. We provide our code on GitHub:https://github.com/SenyuHou/CaMCoD.
Senyu Hou, Gaoxia Jiang, Yaqing Guo, Wenjian Wang 0001
IEEE Trans. Knowl. Data Eng.2
2025 CGFNet: Frequency-Domain Causal Discovery and Dual-Path Spectral Filtering for Wildfire Prediction
Hangyuan Du, Dengke Su, Liang Bai 0001, Gaoxia Jiang, Lu Bai 0001, Wenjian Wang 0001
IEEE Big Data4
2025 Contrastive Anomalous User Detection in Recommender Systems via Multi-Semantic Paths
Hangyuan Du, Liang Bai 0001, Gaoxia Jiang, Lu Bai 0001, Wenjian Wang 0001
IEEE Big Data4
2025 Directional Label Diffusion Model for Learning from Noisy Labels
abstract
In image classification, the label quality of training data critically influences model generalization, especially for deep neural networks (DNNs). Traditionally, learning from noisy labels (LNL) can improve the generalization of DNNs through complex architectures or a series of robust techniques, but its performance improvement is limited by the discriminative paradigm. Unlike traditional ways, we resolve the LNL problems from the perspective of robust label generation, based on diffusion models within the generative paradigm. To expand the diffusion model into a robust classifier that explicitly accommodates more noise knowledge, we propose a Directional Label Diffusion (DLD) model. It disentangles the diffusion process into two paths, i.e., directional diffusion and random diffusion. Specifically, directional diffusion simulates the corruption of true labels into a directed noise distribution, prioritizing the removal of likely noise, whereas random diffusion introduces inherent randomness to support label recovery. This architecture enable DLD to gradually infer labels from an initial random state, interpretably diverging from the specified noise distribution. To adapt the model to diverse noisy environments, we design a low-cost label pre-correction method that automatically supplies more accurate label information to the diffusion model, without requiring manual intervention or additional iterations. Our approach outperforms state-of-the-art methods on both simulated and real-world noisy datasets. Code is available at https://github.com/SenyuHou/DLD.
Senyu Hou, Gaoxia Jiang, Jia Zhang 0023, Shangrong Yang, Husheng Guo, Yaqing Guo, Wenjian Wang 0001
CVPR2
2025 Noisy Multi-Label Learning through Co-Occurrence-Aware Diffusion
abstract
Noisy labels often compel models to overfit, especially in multi-label classification tasks. Existing methods for noisy multi-label learning (NML) primarily follow a discriminative paradigm, which relies on noise transition matrix estimation or small-loss strategies to correct noisy labels. However, they remain substantial optimization difficulties compared to noisy single-label learning. In this paper, we propose a Co-Occurrence-Aware Diffusion (CAD) model, which reformulates NML from a generative perspective. We treat features as conditions and multi-labels as diffusion targets, optimizing the diffusion model for multi-label learning with theoretical guarantees. Benefiting from the diffusion model's strength in capturing multi-object semantics and structured label matrix representation, we can effectively learn the posterior mapping from features to true multi-labels. To mitigate the interference of noisy labels in the forward process, we guide generation using pseudo-clean labels reconstructed from the latent neighborhood space, replacing original point-wise estimates with neighborhood-based proxies. In the reverse process, we further incorporate label co-occurrence constraints to enhance the model's awareness of incorrect generation directions, thereby promoting robust optimization. Extensive experiments on both synthetic (Pascal-VOC, MS-COCO) and real-world (NUS-WIDE) noisy datasets demonstrate that our approach outperforms state-of-the-art methods.
Senyu Hou, Yuru Ren, Gaoxia Jiang, Wenjian Wang 0001
NeurIPS3
2025 An interpretable sample selection framework against numerical label noise
Gaoxia Jiang, Wenjian Wang 0001
Mach. Learn.1
2025 Outlier-trimmed dual-interval smoothing loss for sample selection in learning with noisy labels
Senyu Hou, Maolong Xu, Gaoxia Jiang, Yaqing Guo, Wenjian Wang 0001
Neural Networks3
2025 Rethinking Oversampling With Class Alliance Constraints From Data Complexity Perspective
abstract
Class overlap is a major factor of data complexity that hampers classifier performance, particularly in imbalanced learning scenarios. Most existing oversampling methods rely on conservative seed sample selection and decoupled synthesis strategies, which limit sample diversity and fail to effectively control overlap risk. This paper proposes a novel oversampling framework called TMACO (Class Alliance-Constrained Oversampling), which integrates data complexity considerations into both seed selection and sample generation. First, TMACO selects seed sample units using a class alliance constraint that jointly considers spatial geometry and class distribution to enhance diversity and representativeness. Second, it generates synthetic samples based on three-point units to ensure regional stability. Third, a region-level filtering mechanism is applied to prevent synthetic samples from intruding into majority class areas. Extensive experiments on benchmark and real-world datasets demonstrate that TMACO consistently improves minority class performance and overall classification accuracy compared to state-of-the-art oversampling techniques. The proposed method also offers interpretable parameter control and adapts well to varying task objectives.
Mingming Han, Husheng Guo, Gaoxia Jiang, Wenjian Wang 0001
IEEE Trans. Knowl. Data Eng.3
2024 Which Is More Effective in Label Noise Cleaning, Correction or Filtering?
abstract
Most noise cleaning methods adopt one of the correction and filtering modes to build robust models. However, their effectiveness, applicability, and hyper-parameter insensitivity have not been carefully studied. We compare the two cleaning modes via a rebuilt error bound in noisy environments. At the dataset level, Theorem 5 implies that correction is more effective than filtering when the cleaned datasets have close noise rates. At the sample level, Theorem 6 indicates that confident label noises (large noise probabilities) are more suitable to be corrected, and unconfident noises (medium noise probabilities) should be filtered. Besides, an imperfect hyper-parameter may have fewer negative impacts on filtering than correction. Unlike existing methods with a single cleaning mode, the proposed Fusion cleaning framework of Correction and Filtering (FCF) combines the advantages of different modes to deal with diverse suspicious labels. Experimental results demonstrate that our FCF method can achieve state-of-the-art performance on benchmark datasets.
Gaoxia Jiang, Jia Zhang 0023, Wenjian Wang 0001, Deyu Meng
AAAI1
2024 Maximum a posteriori estimation and filtering algorithm for numerical label noise
Gaoxia Jiang, Zhengying Li, Wenjian Wang 0001
Appl. Intell.1
2024 Noise cleaning for nonuniform ordinal labels based on inter-class distance
Gaoxia Jiang, Wenjian Wang 0001
Appl. Intell.1
2024 A general elevating framework for label noise filters
Qingqiang Chen, Gaoxia Jiang, Fuyuan Cao, Changqian Men, Wenjian Wang 0001
Pattern Recognit.2
2021 A Unified Sample Selection Framework for Output Noise Filtering: An Error-Bound Perspective
abstract
The existence of output noise will bring difficulties to supervised learning. Noise filtering, aiming to detect and remove polluted samples, is one of the main ways to deal with the noise on outputs. However, most of the filters are heuristic and could not explain the filtering influence on the generalization error (GE) bound. The hyper-parameters in various filters are specified manually or empirically, and they are usually unable to adapt to the data environment. The filter with an improper hyper-parameter may overclean, leading to a weak generalization ability. This paper proposes a unified framework of optimal sample selection (OSS) for the output noise filtering from the perspective of error bound. The covering distance filter (CDF) under the framework is presented to deal with noisy outputs in regression and ordinal classification problems. Firstly, two necessary and sufficient conditions for a fixed goodness of fit in regression are deduced from the perspective of GE bound. They provide the unified theoretical framework for determining the filtering effectiveness and optimizing the size of removed samples. The optimal sample size has the adaptability to the environmental changes in the sample size, the noise ratio, and noise variance. It offers a choice of tuning the hyper-parameter and could prevent filters from overcleansing. Meanwhile, the OSS framework can be integrated with any noise estimator and produces a new filter. Then the covering interval is proposed to separate low-noise and high-noise samples, and the effectiveness is proved in regression. The covering distance is introduced as an unbiased estimator of high noises. Further, the CDF algorithm is designed by integrating the cover distance with the OSS framework. Finally, it is verified that the CDF not only recognizes noise labels correctly but also brings down the prediction errors on real apparent age data set. Experimental results on benchmark regression and ordinal classification data sets demonstrate that the CDF outperforms the state-of-the-art filters in terms of prediction ability, noise recognition, and efficiency.
Gaoxia Jiang, Wenjian Wang 0001, Jiye Liang
J. Mach. Learn. Res.1
2019 A novel distance measure for time series: Maximum shifting correlation distance
Gaoxia Jiang, Wenjian Wang 0001
Pattern Recognit. Lett.1
2017 Markov cross-validation for time series model evaluations
Gaoxia Jiang, Wenjian Wang 0001
Inf. Sci.1
2017 Error estimation based on variance analysis of k-fold cross-validation
Gaoxia Jiang, Wenjian Wang 0001
Pattern Recognit.1