EDBT 2026 Demo / reviewers in the wild / expert
Xiuwen Gong
dblp:160/9976
· DBLP profile ↗
39ranked-venue papers
11as first author
36since 2021 · last 2026
0000-0002-1078-1571ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 10 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 12 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SWIM-Merging: Reweighting singular values of weight matrices for efficient test-time model merging
Boqi Li 0002, Xiuwen Gong |
Neurocomputing | 4 |
| 2025 | CUGF: A Reliable and Fair Recommendation FrameworkabstractRecommendation systems (RS) play a crucial role in assisting decision-making but often suffer from either a lack of credibility or unfairness problems. A few recommendation models have endeavored to address the problem from only one aspect, and approaches to solving both problems remain to be explored. This paper aims to construct a generalized fairness-based recommendation framework that can also provide the credibility of recommendation models. Generally, we propose a reliable and fair recommendation framework called Conformalized User Group Fairness (CUGF) based on the inspiration of conformal prediction. Specifically, we construct dynamic prediction sets that are guaranteed to cover the true item with a user pre-specified probability to ensure credibility while designing novel fairness metrics based on empirical risks to guarantee the fairness of users across different groups. Furthermore, we design a novel CUGF Algorithm to optimize the parameter γ that dominates the prediction sets and also the fairness. Besides, we conduct extensive experiments by applying CUGF on top of various recommendation models and representative datasets to validate its effectiveness with respect to recommendation performance (in terms of average set size) and fairness (in terms of the two defined fairness metrics), the results of which demonstrate the validity of the proposed framework. Nitin Bisht, Xiuwen Gong, Guandong Xu |
AAAI | 2 |
| 2025 | Conformal Prediction for Partial Label LearningabstractPartial label learning (PLL) allows each instance to be annotated with a set of candidate labels, but only one is the ground-truth label. Although the state-of-the-art (SOTA) PLL models have shown competitive performance, they cannot get rid of the negative influence from the noisy false-positive labels during the training process. This leads to a large extent of uncertainty of PLL models’ prediction, and it becomes unreliable to trust a PLL model’s performance only by its prediction accuracy. To bridge this gap, we develop a new framework to quantify the uncertainty for PLL models with valid confidence guarantee, which is named as Conformal Prediction for Partial Label Learning (CP-PLL). This framework can be implemented on top of any PLL method to quantify their predictive confidence in terms of average prediction set size with a use-specified error rate or coverage/confidence level (i.e., probability). We prove that the coverage guarantee in PLL still holds, that is, the ground-truth label can be covered in the constructed prediction set with the user pre-defined error rate α when we use the noisy calibration data to carlibrate the PLL models, which yields to a probability interval of [1- α, 1- α + 1/n+1 + ε]. Extensive experiments are conducted on SOTA PLL methods and benchmark datasets to verify the effectiveness of the proposed framework. Xiuwen Gong, Nitin Bisht, Guandong Xu |
AAAI | 1 |
| 2025 | ENSUR: Equitable and Statistically Unbiased RecommendationabstractAlthough Recommender Systems (RS) have been well-developed for various fields of applications, they often suffer from a crisis of platform credibility with respect to RS confidence and fairness, which may drive users away, threatening the platform's long-term success. In recent years, some works have tried to solve these issues; however, they lack strong statistical guarantees. Therefore, there is an urgent need to solve both issues with a unifying framework with robust statistical guarantees. In this paper, we propose a novel and reliable framework called Equitable and Statistically Unbiased Recommendation (ENSUR)) to dynamically generate prediction sets for users across various groups, which are guaranteed 1) to include ground-truth items with user-predefined high confidence/probability (e.g., 90\%); 2) to ensure user fairness across different groups; 3) to have minimum efficient average prediction set sizes.
We further design an efficient algorithm named Guaranteed User Fairness Algorithm (GUFA) to optimize the proposed method and derive upper bounds of risk and fairness metrics to speed up the optimization process.
Moreover, we provide rigorous theoretical analysis concerning risk and fairness control and minimum set size. Extensive experiments validate the effectiveness of the proposed framework, which aligns with our theoretical analysis. Nitin Bisht, Xiuwen Gong, Guandong Xu |
ICML | 2 |
| 2025 | Unified Knowledge-Guided Molecular Graph Encoder with multimodal fusion and multi-task learning
Mukun Chen, Xiuwen Gong, Shirui Pan, Jia Wu 0001, Bo Du 0001, Wenbin Hu 0001 |
Neural Networks | 2 |
| 2025 | Knowledge-aware contrastive heterogeneous molecular graph learningabstractMolecular representation learning is pivotal in predicting molecular properties and advancing drug design. Traditional methodologies, which predominantly rely on homogeneous graph encoding, are limited by their inability to integrate external knowledge and represent molecular structures across different levels of granularity. To address these limitations, we propose a paradigm shift by encoding molecular graphs into heterogeneous structures, introducing a novel framework: Knowledge-aware Contrastive Heterogeneous Molecular Graph Learning. This approach leverages contrastive learning to enrich molecular representations with embedded external knowledge. KCHML conceptualizes molecules through three distinct graph views-molecular, elemental, and pharmacological-enhanced by heterogeneous molecular graphs and a dual message-passing mechanism. This design offers a comprehensive representation for property prediction, as well as for downstream tasks such as drug-drug interaction prediction. Extensive benchmarking demonstrates KCHML's superiority over state-of-the-art molecular property prediction models, underscoring its ability to capture intricate molecular features. Mukun Chen, Jia Wu 0001, Shirui Pan, Bo Du 0001, Xiuwen Gong, Wenbin Hu 0001 |
PLoS Comput. Biol. | 6 |
| 2025 | DSMT: Dual-Stage Multiscale Transformer for Hyperspectral Snapshot Compressive ImagingabstractSnapshot compressive imaging (SCI) compresses a 3D hyperspectral image (HSI) into a 2D measurement, significantly improving imaging efficiency while preserving the spatial and spectral information inherent in HSI. However, reconstructing high-quality HSIs from compressed measurements remains a core challenge due to the complexity of the inverse problem. Transformer-based methods have recently shown promising performance in HSI reconstruction. Nonetheless, effectively capturing local information, long-range dependencies, and multi-scale features within a reasonable computational cost remains a significant challenge. In this paper, we propose a dual-stage multiscale Transformer (DSMT) tailored for HSI reconstruction, which adopts a coarse-to-fine framework to enhance reconstruction accuracy and network generalization. Specifically, we design a novel U-Net architecture with a dual-branch encoder, where two separate branches process distinct features and are fused to achieve more refined reconstruction results. Full-scale skip connections are introduced to strengthen feature fusion across different stages. To further improve performance, we develop a novel self-attention mechanism called dual-window multiscale multi-head self-attention (DWM-MSA). By utilizing two differently sized windows, DWM-MSA captures long-range dependencies and local information at multiple scales, significantly boosting reconstruction quality. Additionally, we introduce a hybrid positional embedding method, conditional/relative positional embedding (CRPE), which dynamically models both spatial and spectral dependencies, effectively enhancing the Transformer's capacity for HSI reconstruction. Extensive quantitative and qualitative experiments on both the simulated and the real data are conducted to demonstrate the superior performance, stability, and generalization ability of our DSMT. Code of this project is at https://github.com/chenx2000/DSMT. Fulin Luo, Xi Chen 0087, Tan Guo, Xiuwen Gong, Lefei Zhang, Ce Zhu |
IEEE Trans. Image Process. | 4 |
| 2025 | GITANet: Group Interactive Threshold-Based Attention Network for Hyperspectral Image ClassificationabstractGroup convolution networks have shown great potential in hyperspectral image (HSI) classification because of their ability to divide total spectral bands into multiple groups and focus on fine discrimination within different spectral ranges. Most group convolution networks process parallel spectral groups independently; however, they neglect the important relevance of nearby spectral ranges. Moreover, the feature maps from different spectral groups are not considered for recalibration in the existing attention. To address these issues, we propose a novel group interactive threshold attention network (GITANet). In the network, a stratified-split-concatenation strategy, which not only splits all bands into multiple groups for intragroup convolution but also propagates the intergroup information via the stratified concatenation operation between different groups, is designed for bandwise group convolution. Relying on the high dependencies among nearby spectra, the cross-group interactive attention block is designed to encourage significant spectral features. Subsequently, from different spectral ranges, a learnable threshold generation block is built to estimate the information validity of each pixel. On the basis of this threshold, soft threshold spatial attention is developed in the bandwise encoder-decoder architecture, which emphasizes high-value spatial areas during the fusion of group convolutional features. Therefore, complementary and discriminative spectral-spatial features are obtained to improve the performance of HSI classification. The experimental results on three HSI datasets illustrate that GITANet is superior to several state-of-the-art networks. Maixia Fu, Xiuwen Gong, Yingying Niu, Fulin Luo |
IEEE Trans. Multim. | 4 |
| 2024 | Dual-Window Multiscale Transformer for Hyperspectral Snapshot Compressive ImagingabstractCoded aperture snapshot spectral imaging (CASSI) system is an effective manner for hyperspectral snapshot compressive imaging. The core issue of CASSI is to solve the inverse problem for the reconstruction of hyperspectral image (HSI). In recent years, Transformer-based methods achieve promising performance in HSI reconstruction. However, capturing both long-range dependencies and local information while ensuring reasonable computational costs remains a challenging problem. In this paper, we propose a Transformer-based HSI reconstruction method called dual-window multiscale Transformer (DWMT), which is a coarse-to-fine process, reconstructing the global properties of HSI with the long-range dependencies. In our method, we propose a novel U-Net architecture using a dual-branch encoder to refine pixel information and full-scale skip connections to fuse different features, enhancing the extraction of fine-grained features. Meanwhile, we design a novel self-attention mechanism called dual-window multiscale multi-head self-attention (DWM-MSA), which utilizes two different-sized windows to compute self-attention, which can capture the long-range dependencies in a local region at different scales to improve the reconstruction performance. We also propose a novel position embedding method for Transformer, named con-abs position embedding (CAPE), which effectively enhances positional information of the HSIs. Extensive experiments on both the simulated and the real data are conducted to demonstrate the superior performance, stability, and generalization ability of our DWMT. Code of this project is at https://github.com/chenx2000/DWMT. Fulin Luo, Xi Chen 0087, Xiuwen Gong, Weiwen Wu, Tan Guo |
AAAI | 3 |
| 2024 | Does Label Smoothing Help Deep Partial Label Learning?abstractAlthough deep partial label learning (deep PLL) classifiers have shown their competitive performance, they are heavily influenced by the noisy false-positive labels leading to poorer performance as the training progresses. Meanwhile, existing deep PLL research lacks theoretical guarantee on the analysis of correlation between label noise (or ambiguity degree) and classification performance. This paper addresses the above limitations with label smoothing (LS) from both theoretical and empirical aspects. In theory, we prove lower and upper bounds of the expected risk to show that label smoothing can help deep PLL. We further derive the optimal smoothing rate to investigate the conditions, i.e., when label smoothing benefits deep PLL. In practice, we design a benchmark solution and a novel optimization algorithm called Label Smoothing-based Partial Label Learning (LS-PLL). Extensive experimental results on benchmark PLL datasets and various deep architectures validate that label smoothing does help deep PLL in improving classification performance and learning distinguishable representations, and the best results can be achieved when the empirical smoothing rate approximately approaches the optimal smoothing rate in theoretical findings. Code is publicly available at https://github.com/kalpiree/LS-PLL. Xiuwen Gong, Nitin Bisht, Guandong Xu |
ICML | 1 |
| 2024 | Contrastive Transformer Masked Image Hashing for Degraded Image Retrieval
Xiaobo Shen 0001, Haoyu Cai, Xiuwen Gong, Yuhui Zheng |
IJCAI | 3 |
| 2024 | Unsupervised Deep Graph Structure and Embedding Learning
Xiaobo Shen 0001, Xiuwen Gong, Shirui Pan |
IJCAI | 3 |
| 2024 | Contrastive Learning Drug Response Models from Natural Language Supervision
Kun Li 0009, Xiuwen Gong, Jia Wu 0001, Wenbin Hu 0001 |
IJCAI | 2 |
| 2024 | EMVCC: Enhanced Multi-View Contrastive Clustering for Hyperspectral ImagesabstractCross-view consensus representation plays a critical role in hyperspectral images (HSIs) clustering. Recent multi-view contrastive cluster methods utilize contrastive loss to extract contextual consensus representation. However, these methods have a fatal flaw: contrastive learning may treat similar heterogeneous views as positive sample pairs and dissimilar homogeneous views as negative sample pairs. At the same time, the data representation via self-supervised contrastive loss is not specifically designed for clustering. Thus, to tackle this challenge, we propose a novel multi-view clustering method, i.e., Enhanced Multi-View Contrastive Clustering (EMVCC). First, the spatial multi-view is designed to learn the diverse features for contrastive clustering, and the globally relevant information of spectrum-view is extracted by Transformer, enhancing the spatial multi-view differences between neighboring samples. Then, a joint self-supervised loss is designed to constrain the consensus representation from different perspectives to efficiently avoid false negative pairs. Specifically, to preserve the diversity of multi-view information, the features are enhanced by using probabilistic contrastive loss, and the data is projected into a semantic representation space, ensuring that the similar samples in this space are closer in distance. Finally, we design a novel clustering loss that aligns the view feature representation with high confidence pseudo-labels for promoting the network to learn cluster-friendly features. In the training process, the joint self-supervised loss is used to optimize the cross-view features.Abundant experiment studies on numerous benchmarks verify the superiority of EMVCC in comparison to some state-of-the-art clustering methods. The codes are available at https://github.com/YiLiu1999/EMVCC. Fulin Luo, Yi Liu 0038, Xiuwen Gong, Zhixiong Nan, Tan Guo |
ACM Multimedia | 3 |
| 2024 | Dimensionality Reduction via Multiple Neighborhood-Aware Nonlinear Collaborative Analysis for Hyperspectral Image ClassificationabstractLocal collaborative representation (CR) has drawn much attention in exploring data relationships due to considering local knowledge in the global linear combination, subsequently, local CR-based graph embedding methods have been applied to dimensionality reduction of hyperspectral image (HSI). However, HSI data with nonlinear distribution cannot be handled with pure linear combination accurately. Furthermore, the existing local knowledge in terms of binary relations between pairwise neighbors makes it hard to learn the accurate local structure among neighborhood sets through local CR-based graph embedding. To this end, this paper proposes a novel multiple neighborhood-aware nonlinear collaborative analysis (MNNCA) method. Relying on the primary and secondary neighborhoods, a dual-level neighborhood reconstruction is designed to search for optimal neighbors and mine the common attributes within the neighborhood. With the reconstruction information, a nonlinear extend multiple neighborhood-aware collaborative representation (NE-MNACR) model is built on nonlinear geodesic constraint and multi-neighborhood-aware items. It can explore the collaborative relationship among multiple neighborhood sets in the nonlinear space of HSI data. By preserving the multivariate local structure instead of pairwise local relations, a pair of collaborative structure preservation graphs are constructed to realize the final embedding of HSI data. Experimental results on serval HSI data sets demonstrate the superior performance of the proposed MNNCA method and NE-MNACR model in comparison with some state-of-the-art DR methods and local CR models. Maixia Fu, Xiuwen Gong, Fulin Luo |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Learnable Background Endmember With Subspace Representation for Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) aims to label each hyperspectral image (HSI) pixel as background or anomaly, in a totally unsupervised manner. Thus, a fine background representation is vital to obtain good HAD performance. This article introduces background endmember representation and proposes a novel HAD method termed learnable background endmember with subspace representation (LEBSR). First, the HSI is unmixed to simultaneously obtain the background endmembers and their abundances. The three constraints of$p$-norm, sum-to-one, and nonnegativity work together to promote a more meaningful and accurate background endmember representation. In addition, a mapping matrix with orthogonality is jointly optimized to transform the priori backgrounds into the background endmember subspace, and then the mapped priori backgrounds are approximated to the low-rank representation (LRR) with the background endmembers. With the methodology, the backgrounds can be well reconstructed under the guidance of the priori background information to accurately detect anomalous pixels with the reconstructed residuals. The experimental results on several HSI datasets verify the superior performance of LEBSR than the state-of-the-art methods.https://github.com/HalongL/HAD-LEBSR Tan Guo, Fulin Luo, Xiuwen Gong, Lei Zhang 0038, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | DCENet: Diff-Feature Contrast Enhancement Network for Semi-Supervised Hyperspectral Change DetectionabstractMulti-temporal Hyperspectral images (HSIs) have wide applications in change detection (CD) of different land-covers for their rich spectral features and image details. Traditional supervised learning-based HSI CD algorithms often rely on a substantial number of labeled samples. However, it requires a significant cost in sample annotation. In this paper, we propose a diff-feature contrast enhancement network (DCENet) for semi-supervised HSI CD, which leverages a limited number of labeled samples to guide the training process and a large number of unlabeled samples to improve the confidence of change detection. To achieve this, a differential fusion attention (DFA) sub-network is constructed to extract temporal features from the initial input HSI patches. The dual-branch siamese enhancement module (SEM) is utilized to enhance the generalization of differential features in the feature maps. Herein, multi-scale Kullback-Leibler divergence and feature-enhanced probabilistic contrast loss are designed to constrain the SEM. The proposed method excels at detecting subtle changes in bi-temporal HSIs simultaneously improving the generalization performance of networks. The visual and quantitative experimental results on four HSI datasets show that the proposed DCENet outperforms the compared state-of-the-art methods for HSI CD. Codes: https://github.com/Zhoutya/ChangeDetection-DCENet. Fulin Luo, Tianyuan Zhou, Tan Guo, Xiuwen Gong, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | CrCD: Multidirection-MLP-Based Cross-Contrastive Disambiguation for Hyperspectral Image Partial Label Learning
Xiaoyu Tian, Fulin Luo, Xiuwen Gong, Tan Guo, Bo Du 0001, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | HSCN: Semi-Supervised ALS Point Cloud Semantic Segmentation via Hybrid Structure Constraint NetworkabstractSemi-supervised learning (SSL) plays a crucial role in airborne laser scanning (ALS) point cloud semantic segmentation to reduce the cost of sample labeling. However, the prevailing semi-supervised approaches mainly focus on local and global feature aggregation, which neglects the multiscale and neighborhood structure properties of ALS point clouds. To address this issue, we propose an innovative approach called the Hybrid Structured Constraint Network (HSCN) for semi-supervised semantic segmentation of ALS point clouds. HSCN makes full use of a large number of unlabeled samples to guide the model training under limited labeled samples. To process the unlabeled samples, we construct a global awareness loss (GAL) to constrain the global distribution of point clouds. Then, we also design a multiscale geometric structure similarity metric loss (MLS) to align the neighborhood structure for point clouds at different scales. In addition, we utilize the multiscale features to develop a multilevel fused pseudo-label generator for obtaining high-value pseudo-labels, and then a pseudo-label loss (PLS) is constructed to reduce the class mean probability discrepancies. The proposed HSCN fully utilizes the multiscale and neighborhood structure properties of unlabeled samples to achieve a robust model under limited labeled samples. Extensive experimental analysis using three benchmark datasets (i.e., ISPRS, LASDU, and DFC2019) reveals that our proposed method achieves comparable advantages to some existing advanced fully supervised approaches, even only 0.1% labeled samples for model training. The code is available athttps://github.com/SC-shendazt/HSCN. Fulin Luo, Tan Guo, Xiuwen Gong, Wenqiang Shu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Relation Preference Oriented High-order Sampling for RecommendationabstractThe introduction of knowledge graphs (KG) into recommendation systems (RS) has been proven to be effective because KG introduces a variety of relations between items. In fact, users have different relation preferences depending on the relationship in KG. Existing GNN-based models largely adopt random neighbor sampling strategies to process propagation; however, these models cannot aggregate biased relation preference local information for a specific user, and thus cannot effectively reveal the internal relationship between users' preferences. This will reduce the accuracy of recommendations, while also limiting the interpretability of the results. Mukun Chen, Xiuwen Gong, YH Jin, Wenbin Hu 0001 |
WSDM | 2 |
| 2023 | CLNode: Curriculum Learning for Node ClassificationabstractNode classification is a fundamental graph-based task that aims to predict the classes of unlabeled nodes, for which Graph Neural Networks (GNNs) are the state-of-the-art methods. Current GNNs assume that nodes in the training set contribute equally during training. However, the quality of training nodes varies greatly, and the performance of GNNs could be harmed by two types of low-quality training nodes: (1) inter-class nodes situated near class boundaries that lack the typical characteristics of their corresponding classes. Because GNNs are data-driven approaches, training on these nodes could degrade the accuracy. (2) mislabeled nodes. In real-world graphs, nodes are often mislabeled, which can significantly degrade the robustness of GNNs. To mitigate the detrimental effect of the low-quality training nodes, we present CLNode, which employs a selective training strategy to train GNN based on the quality of nodes. Specifically, we first design a multi-perspective difficulty measurer to accurately measure the quality of training nodes. Then, based on the measured qualities, we employ a training scheduler that selects appropriate training nodes to train GNN in each epoch. To evaluate the effectiveness of CLNode, we conduct extensive experiments by incorporating it in six representative backbone GNNs. Experimental results on real-world networks demonstrate that CLNode is a general framework that can be combined with various GNNs to improve their accuracy and robustness. Xiaowen Wei, Xiuwen Gong, Yibing Zhan, Bo Du 0001, Yong Luo 0002, Wenbin Hu 0001 |
WSDM | 2 |
| 2023 | Multilevel Context Feature Fusion for Semantic Segmentation of ALS Point CloudabstractSemantic segmentation of airborne laser scanning (ALS) point clouds using deep learning is a hot research in remote sensing and photogrammetry. A current trend is to aggregate contextual features from different scales for boosting network generalization and diversity discrimination capabilities. One main challenge is how to achieve effective fusion with multiscale information. In this letter, we propose a muti-level context feature fusion network (MCFN) for semantic segmentation of ALS point cloud based on an encoder-decoder structure. More specifically, we design the squeeze-expansion shared MLP module (SE-MLP) following kernel point convolution (KPConv) in the encoding stage, which can extend the receptive field of KPConv. To aggregate low-level features and high-level representations, we establish channel self-attention between skip connections. In the decoding stage, we develop a cross-layer attention fusion module (CAF) to generate additional discriminative channel features by fusing multi-scale features at different upsampling layers. Experiments on the ISPRS and LASDU datasets demonstrate the superiority of the proposed method. Code: https://github.com/SC-shendazt/MCFN. Fulin Luo, Tan Guo, Xiuwen Gong, Jingyun Xue, Hanshan Li |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | A Unifying Probabilistic Framework for Partially Labeled Data LearningabstractPartially labeled data learning (PLDL), including partial label learning (PLL) and partial multi-label learning (PML), has been widely used in nowadays data science. Researchers attempt to construct different specific models to deal with the different classification tasks for PLL and PML scenarios respectively. The main challenge in training classifiers for PLL and PML is how to deal with ambiguities caused by the noisy false-positive labels in the candidate label set. The state-of-the-art strategy for both scenarios is to perform disambiguation by identifying the ground-truth label(s) directly from the candidate label set, which can be summarized into two categories: 'the identifying method' and 'the embedding method'. However, both kinds of methods are constructed by hand-designed heuristic modeling under considerations like feature/label correlations with no theoretical interpretation. Instead of adopting heuristic or specific modeling, we propose a novel unifying framework called A Unifying Probabilistic Framework for Partially Labeled Data Learning (UPF-PLDL), which is derived from a clear probabilistic formulation, and brings existing research on PLL and PML under one theoretical interpretation with respect to information theory. Furthermore, the proposed UPF-PLDL also unifies 'the identifying method' and 'the embedding method' into one integrated framework, which naturally incorporates the feature and label correlation considerations. Comprehensive experiments on synthetic and real-world datasets for both PLL and PML scenarios clearly demonstrate the superiorities of the derived framework. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001, Fulin Luo |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Classification via Structure-Preserved Hypergraph Convolution Network for Hyperspectral ImageabstractGraph convolutional network (GCN) as a combination of deep learning and graph learning has gained increasing attention in hyperspectral image (HSI) classification. However, most GCN methods consider the simple point-to-point structure between two pixels rather than the high-order structure of multiple pixels, which is contradict with the real feature distribution of ground object. And the nonlinear property of HSI also brings challenge for precise structural representation in GCN. To tackle these problems, this work proposes a structure preserved hypergraph convolution network (SPHGCN). It first builds a multiple neighborhood reconstruction (MNR) model to reveal the essential resemblance of multiple pixels in nonlinear spectral feature space. With the high-order structure, SPHGCN designs the hypergraph convolution operation for irregular feature aggregation among similar pixels from different regions, which achieves more discriminative features from multiple pixel nodes. Meanwhile, a structure preservation layer is built to optimize the distribution of convolutional features under the guidance of high-order structure. Moreover, SPHGCN integrates local regular convolution and irregular hypergraph convolution to learn the structured semantic feature of HSI. This strategy breaks the boundary restriction in traditional convolution and aggregates semantic feature across different image patches. Experiments on three HSI data sets indicate that SPHGCN outperforms a few state-of-the-art methods for HSI classification. Fulin Luo, Maixia Fu, Yingying Niu, Xiuwen Gong |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Anomaly Detection of Hyperspectral Image With Hierarchical Antinoise Mutual-Incoherence- Induced Low-Rank RepresentationabstractHyperspectral image (HSI) anomaly detection (AD) generally considers background pixels as low-rank distribution and anomaly pixels as sparse distribution. However, it is usually difficult to construct an accurate background dictionary for the background pixels composed of different land-covers, and completely separate sparse anomaly targets from various complicated background pixels with complex mixed noise interference. To address these challenges, we propose an anti-noise hierarchical mutual-incoherence-induced discriminative learning (AHMID) method for AD of HSI. A structural incoherence constraint is designed to constrain the inherent dissimilarity and incoherence between background and anomalies for improving their separability. Then, a first-order statistic constraint is conducted on targets to enhance the anomaly representation, and a decentralization constraint is used on background to suppress the background representation. Meanwhile, a mixed noise model is constructed by ℓ1,1-norm and Frobenius norm to improve the anti-noise performance. Finally, a hierarchical alternating strategy is developed to gradually optimize the background and anomalies. Experiments on six HSI AD datasets show that the proposed method outperforms a few state-of-the-art AD algorithms. Code: https://github.com/HalongL/HAD-AHMID. Tan Guo, Fulin Luo, Xiuwen Gong, Lei Zhang 0038 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Dual-View Spectral and Global Spatial Feature Fusion Network for Hyperspectral Image ClassificationabstractFor hyperspectral image (HSI) classification, two branch networks generally use the convolution neural networks (CNNs) to extract the spatial features and the long short-term memory (LSTM) to learn the spectral features. However, CNN with a local kernel neglects the global properties of the whole HSI. LSTM doesn’t consider the macroscopic and detailed information of spectra. In this paper, we propose a dual-view spectral and global spatial feature fusion network (DSGSF) to extract the spatial-spectral features for HSI classification, including a spatial subnetwork and a spectral subnetwork. In the spatial subnetwork, we propose a global spatial feature representation model based on the encoder-decoder structure with channel attention and spatial attention to learn the global spatial features. In the spectral subnetwork, we design a dual-view spectral feature aggregation model with view attention to learn the diversity of spectral features. By fusing the two subnetworks, we construct DSGSF to extract the spatial-spectral features of HSI with strong discriminating performance. Experimental results on three public datasets illustrate that the proposed method can achieve competitive results compared with the state-of-the-art methods. Code: https://github.com/RZWang-WH/DSGSF. Tan Guo, Fulin Luo, Xiuwen Gong, Lei Zhang 0038, Xinbo Gao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Multiscale Diff-Changed Feature Fusion Network for Hyperspectral Image Change DetectionabstractFor hyperspectral image (HSI) change detection (CD), multiscale features are usually used to construct the detection models. However, the existing studies only consider the multiscale features containing changed and unchanged components, which is difficult to represent the subtle changes between bitemporal HSIs in each scale. To address this problem, we propose a multiscale diff-changed feature fusion network (MSDFFN) for HSI CD, which improves the ability of feature representation by learning the refined change components between bitemporal HSIs under different scales. In this network, a temporal feature encoder–decoder subnetwork, which combines a reduced inception (RI) module and a cross-layer attention module to highlight the significant features, is designed to extract the temporal features of HSIs. A bidirectional diff-changed feature representation (BDFR) module is proposed to learn the fine changed features of bitemporal HSIs at various scales to enhance the discriminative performance of the subtle change. A multiscale attention fusion (MSAF) module is developed to adaptively fuse the changed features of various scales. The proposed method can not only discover the subtle change in bitemporal HSIs but also improve the discriminating power for HSI CD. Experimental results on three HSI datasets show that MSDFFN outperforms a few state-of-the-art methods. Fulin Luo, Tianyuan Zhou, Tan Guo, Xiuwen Gong, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Recurrent Residual Dual Attention Network for Airborne Laser Scanning Point Cloud Semantic SegmentationabstractKernel point convolution (KPConv) can effectively represent the point features of point cloud data. However, KPConv-based methods just consider the local information of each point, which is very difficult to characterize the intrinsic properties of ALS point clouds for complex laser scanning conditions. Therefore, we rethink KPConv and propose a recurrent residual dual attention network (RRDAN) based on the encoder-decoder structure for the semantic segmentation of ALS point cloud data. In the encoder stage, we design an attention kernel point convolution (AKPConv) block by using a scaling factor of batch normalization to highlight the significant channel information. Then, we use the AKPConv block to develop a recurrent residual kernel attention (RRKA) module to iteratively aggregate the local neighborhood features. In the decoder stage, we design a global and local channel attention (GLCA) module with global connection and local 1D convolution to interact the global and local information after fusing the upsampled high-level representations and the skip-connected low-level features. In addition, to reduce the influence of the long-tail distribution of reflection intensity, we apply gamma transformation to correct the data as normal distribution. The proposed RRDAN can achieve diversified feature aggregation to implement the refined semantic segmentation of ALS point clouds. We evaluate our method on two ALS datasets (i.e., ISPRS and DCF2019) to demonstrate its performance compared to a few advanced methods. Code: https://github.com/SC-shendazt/RRDAN. Fulin Luo, Tan Guo, Xiuwen Gong, Jingyun Xue, Hanshan Li |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Attention Multihop Graph and Multiscale Convolutional Fusion Network for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) for hyperspectral image (HSI) classification have generated good progress. Meanwhile, graph convolutional networks (GCNs) have also attracted considerable attention by using unlabeled data, broadly and explicitly exploiting correlations between adjacent parcels. However, the CNN with a fixed square convolution kernel is not flexible enough to deal with irregular patterns, while the GCN using the superpixel to reduce the number of nodes will lose the pixel-level features, and the features from the two networks are always partial. In this paper, to make good use of the advantages of CNN and GCN, we propose a novel multiple feature fusion model termed attention multi-hop graph and multi-scale convolutional fusion network (AMGCFN), which includes two sub-networks of multi-scale fully CNN and multi-hop GCN to extract the multi-level information of HSI. Specifically, the multi-scale fully CNN aims to comprehensively capture pixel-level features with different kernel sizes, and a multi-head attention fusion module is used to fuse the multi-scale pixel-level features. The multi-hop GCN systematically aggregates the multi-hop contextual information by applying multi-hop graphs on different layers to transform the relationships between nodes, and a multi-head attention fusion module is adopted to combine the multi-hop features. Finally, we design a cross attention fusion module to adaptively fuse the features of two sub-networks. AMGCFN makes full use of multi-scale convolution and multi-hop graph features, which is conducive to the learning of multi-level contextual semantic features. Experimental results on three benchmark HSI datasets show that AMGCFN has better performance than a few state-of-the-art methods. Fulin Luo, Huiping Zhuang, Zhenyu Weng, Xiuwen Gong, Zhiping Lin 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Discriminative Metric Learning for Partial Label LearningabstractOne simple strategy to deal with ambiguity in partial label learning (PLL) is to regard all candidate labels equally as the ground-truth label, and then solve the PLL problem using existing multiclass classification algorithms. However, due to the noisy false-positive labels in the candidate set, these approaches are readily mislead and do not generalize well in testing. Consequently, the method of identifying the ground-truth label straight from the candidate label set has grown popular and effective. When the labeling information in PLL is ambiguous, we ought to take advantage of the data's underlying structure, such as label and feature interdependencies, to conduct disambiguation. Furthermore, while metric learning is an excellent method for supervised learning classification that takes feature and label interdependencies into account, it cannot be used to solve the weekly supervised learning PLL problem directly due to the ambiguity of labeling information in the candidate label set. In this article, we propose an effective PLL paradigm called discriminative metric learning for partial label learning (DML-PLL), which aims to learn a Mahanalobis distance metric discriminatively while identifying the ground-truth label iteratively for PLL. We also design an efficient algorithm to alternatively optimize the metric parameter and the latent ground-truth label in an iterative way. Besides, we prove the convergence of the designed algorithms by two proposed lemmas. We additionally study the computational complexity of the proposed DML-PLL in terms of training and testing time for each iteration. Extensive experiments on both controlled UCI datasets and real-world PLL datasets from diverse domains demonstrate that the proposed DML-PLL regularly outperforms the compared approaches in terms of prediction accuracy. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Partial Multi-Label Learning via Large Margin Nearest Neighbour EmbeddingsabstractTo deal with ambiguities in partial multi-label learning (PML), existing popular PML research attempts to perform disambiguation by direct ground-truth label identification. However, these approaches can be easily misled by noisy false-positive labels in the iteration of updating the model parameter and the latent ground-truth label variables. When labeling information is ambiguous, we should depend more on underlying structure of data, such as label and feature correlations, to perform disambiguation for partially labeled data. Moreover, large margin nearest neighbour (LMNN) is a popular strategy that considers data structure in classification. However, due to the ambiguity of labeling information in PML, traditional LMNN cannot be used to solve the PML problem directly. In addition, embedding is an effective technology to decrease the noise information of data. Inspried by LMNN and embedding technology, we propose a novel PML paradigm called Partial Multi-label Learning via Large Margin Nearest Neighbour Embeddings (PML-LMNNE), which aims to conduct disambiguation by projecting labels and features into a lower-dimension embedding space and reorganize the underlying structure by LMNN in the embedding space simultaneously. An efficient algorithm is designed to implement the proposed method and the convergence rate of the algorithm is analyzed. Moreover, we present a theoretical analysis of the generalization error bound for the proposed PML-LMNNE, which shows that the generalization error converges to the sum of two times the Bayes error over the labels when the number of instances goes to infinity. Comprehensive experiments on artificial and real-world datasets demonstrate the superiorities of the proposed PML-LMNNE. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001 |
AAAI | 1 |
| 2022 | Partial Label Learning via Label Influence FunctionabstractTo deal with ambiguities in partial label learning (PLL), state-of-the-art strategies implement disambiguations by identifying the ground-truth label directly from the candidate label set. However, these approaches usually take the label that incurs a minimal loss as the ground-truth label or use the weight to represent which label has a high likelihood to be the ground-truth label. Little work has been done to investigate from the perspective of how a candidate label changing a predictive model. In this paper, inspired by influence function, we develop a novel PLL framework called Partial Label Learning via Label Influence Function (PLL-IF). Moreover, we implement the framework with two specific representative models, an SVM model and a neural network model, which are called PLL-IF+SVM and PLL-IF+NN method respectively. Extensive experiments conducted on various datasets demonstrate the superiorities of the proposed methods in terms of prediction accuracy, which in turn validates the effectiveness of the proposed PLL-IF framework. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001 |
ICML | 1 |
| 2022 | Generalized Large Margin $k$NN for Partial Label LearningabstractTo deal with noises in partial label learning (PLL), existing approaches try to perform disambiguation either by identifying the ground-truth label or by averaging the candidate labels. However, these methods can be easily misled by the false-positive noisy labels in the candidate set, and fail to generalize well in testing. When labeling information is ambiguous, learning paradigms should depend more on underlying data structure. Large margin nearest neighbour (LMNN) is a popular strategy to consider instance and class correlations in supervised learning, but can not be directly used in weakly-supervised PLL due to the ambiguity of labeling information. In this paper, we first define similarly and differently labeled pairs as well as the similarity weight to evaluate the similarties between any two instances. We then propose a novel PLL method called Generalized Large Margin$k$NN for Partial Label Learning (GLMNN-PLL), which adapts the framework of LMNN to PLL by modifying the constraint from ‘the same class’ to ‘similarly-labeled’. GLMNN-PLL aims to learn a new metric and perform disambiguation by reorganizing the underlying data structure, that is, making similarly labeled instances closer to each other while making differently labeled instances seperated by a large margin. As two close instances with shared labels do not necessarily belong to the same class, we put a weight on each instance pair. An efficient algorithm is designed to optimize the proposed method and the convergence is analyzed in this paper. Moreover, we present a theoretical analysis of the generalization error bound for GLMNN-PLL. Comprehensive experiments on controlled UCI datasets as well as real-world partial label datasets from various domains demonstrate the superiorities of the proposed method. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Top- Partial Label MachineabstractTo deal with ambiguities in partial label learning (PLL), the existing PLL methods implement disambiguations, by either identifying the ground-truth label or averaging the candidate labels. However, these methods can be easily misled by the false-positive labels in the candidate label set. We find that these ambiguities often originate from the noise caused by highly correlated or overlapping candidate labels, which leads to the difficulty in identifying the ground-truth label on the first attempt. To give the trained models more tolerance, we first propose the top-k partial loss and convex top-k partial hinge loss. Based on the losses, we present a novel top-k partial label machine (TPLM) for partial label classification. An efficient optimization algorithm is proposed based on accelerated proximal stochastic dual coordinate ascent (Prox-SDCA) and linear programming (LP). Moreover, we present a theoretical analysis of the generalization error for TPLM. Comprehensive experiments on both controlled UCI datasets and real-world partial label datasets demonstrate that the proposed method is superior to the state-of-the-art approaches. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Fast Multi-label LearningabstractEmbedding approaches have become one of the most pervasive techniques for multi-label classification. However, the training process of embedding methods usually involves a complex quadratic or semidefinite programming problem, or the model may even involve an NP-hard problem. Thus, such methods are prohibitive on large-scale applications. More importantly, much of the literature has already shown that the binary relevance (BR) method is usually good enough for some applications. Unfortunately, BR runs slowly due to its linear dependence on the size of the input data. The goal of this paper is to provide a simple method, yet with provable guarantees, which can achieve competitive performance without a complex training process. To achieve our goal, we provide a simple stochastic sketch strategy for multi-label classification and present theoretical results from both algorithmic and statistical learning perspectives. Our comprehensive empirical studies corroborate our theoretical findings and demonstrate the superiority of the proposed methods. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001 |
IJCAI | 1 |
| 2021 | Understanding Partial Multi-Label Learning via Mutual InformationabstractTo deal with ambiguities in partial multilabel learning (PML), state-of-the-art methods perform disambiguation by identifying ground-truth labels directly. However, there is an essential question:“Can the ground-truth labels be identified precisely?". If yes, “How can the ground-truth labels be found?". This paper provides affirmative answers to these questions. Instead of adopting hand-made heuristic strategy, we propose a novel Mutual Information Label Identification for Partial Multilabel Learning (MILI-PML), which is derived from a clear probabilistic formulation and could be easily interpreted theoretically from the mutual information perspective, as well as naturally incorporates the feature/label relevancy considerations. Extensive experiments on synthetic and real-world datasets clearly demonstrate the superiorities of the proposed MILI-PML. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001 |
NeurIPS | 1 |
| 2020 | Online Metric Learning for Multi-Label ClassificationabstractExisting research into online multi-label classification, such as online sequential multi-label extreme learning machine (OSML-ELM) and stochastic gradient descent (SGD), has achieved promising performance. However, these works lack an analysis of loss function and do not consider label dependency. Accordingly, to fill the current research gap, we propose a novel online metric learning paradigm for multi-label classification. More specifically, we first project instances and labels into a lower dimension for comparison, then leverage the large margin principle to learn a metric with an efficient optimization algorithm. Moreover, we provide theoretical analysis on the upper bound of the cumulative loss for our method. Comprehensive experiments on a number of benchmark multi-label datasets validate our theoretical approach and illustrate that our proposed online metric learning (OML) algorithm outperforms state-of-the-art methods. Xiuwen Gong, Dong Yuan 0001, Wei Bao 0001 |
AAAI | 1 |
| 2015 | Effectively Predicting Whether and When a Topic Will Become Prevalent in a Social NetworkabstractEffective forecasting of future prevalent topics plays animportant role in social network business development.It involves two challenging aspects: predicting whethera topic will become prevalent, and when. This cannotbe directly handled by the existing algorithms in topicmodeling, item recommendation and action forecasting.The classic forecasting framework based on time seriesmodels may be able to predict a hot topic when a seriesof periodical changes to user-addressed frequency in asystematic way. However, the frequency of topics discussedby users often changes irregularly in social networks.In this paper, a generic probabilistic frameworkis proposed for hot topic prediction, and machine learningmethods are explored to predict hot topic patterns.Two effective models, PreWHether and PreWHen, areintroduced to predict whether and when a topic will becomeprevalent. In the PreWHether model, we simulatethe constructed features of previously observed frequencychanges for better prediction. In the PreWHen model,distributions of time intervals associated with the emergenceto prevalence of a topic are modeled. Extensiveexperiments on real datasets demonstrate that ourmethod outperforms the baselines and generates moreeffective predictions. Weiwei Liu 0003, Zhi-Hong Deng 0001, Xiuwen Gong, Frank Jiang 0001, Ivor W. Tsang |
AAAI | 3 |
| 2015 | Mining Top K Spread Sources for a Specific Topic and a Given NodeabstractIn social networks, nodes (or users) interested in specific topics are often influenced by others. The influence is usually associated with a set of nodes rather than a single one. An interesting but challenging task for any given topic and node is to find the set of nodes that represents the source or trigger for the topic and thus identify those nodes that have the greatest influence on the given node as the topic spreads. We find that it is an NP-hard problem. This paper proposes an effective framework to deal with this problem. First, the topic propagation is represented as the Bayesian network. We then construct the propagation model by a variant of the voter model. The probability transition matrix (PTM) algorithm is presented to conduct the probability inference with the complexity O(θ(3)log2θ), while θ is the number nodes in the given graph. To evaluate the PTM algorithm, we conduct extensive experiments on real datasets. The experimental results show that the PTM algorithm is both effective and efficient. Weiwei Liu 0003, Zhi-Hong Deng 0001, Longbing Cao, Xiuwen Gong |
IEEE Trans. Cybern. | 6 |