Xinyu Xiong

dblp:193/1435 · DBLP profile ↗
← Back
13ranked-venue papers
2as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 FGTBT: Frequency-guided task-balancing transformer for unified facial landmark detection
Jun Wan 0005, Xinyu Xiong, Zhihui Lai 0001, Jie Zhou 0009, Wenwen Min
Inf. Sci.2
2025 Towards Realistic Semi-supervised Medical Image Classification
abstract
Existing semi-supervised learning (SSL) approaches follow the idealized closed-world assumption, neglecting the challenges present in realistic medical scenarios, such as open-set distribution and imbalanced class distribution. Although some methods in natural domains attempt to address the open-set problem, they are insufficient for medical domains, where intertwined challenges like class imbalance and small inter-class lesion discrepancies persist. Thus, this paper presents a novel self-recalibrated semantic training framework, which is tailored for SSL in medical imaging by ingeniously harvesting realistic unlabeled samples. Inspired by the observation that certain open-set samples share some similar disease-related representations with in-distribution samples, we first propose an informative sample selection strategy that identifies high-value samples to serve as augmentations, thereby effectively enriching the semantics of known categories. Furthermore, we adopt a compact semantic clustering strategy to address the semantic confusion raised by the above newly introduced open-set semantics. Moreover, to mitigate the interference of class imbalance in open-set SSL, we introduce a less biased dual-balanced classifier with similarity pseudo-label regularization and category-customized regularization. Extensive experiments on a variety of medical image datasets demonstrate the superior performance of our proposed method over state-of-the-art Closed-set and Open-set SSL methods.
Wenxue Li 0003, Lie Ju, Peng Xia 0005, Xinyu Xiong, Lei Zhu 0002, ZongYuan Ge
AAAI5
2025 GlassWizard: Harvesting Diffusion Priors for Glass Surface Detection
Wenxue Li 0003, Tian Ye 0001, Xinyu Xiong, Jinbin Bai, Wenxuan Song, Zhaohu Xing, Lie Ju, Guanbin Li, Lei Zhu 0003
ICCV3
2025 PDC-Net: Pattern Divide-and-Conquer Network for Pelvic Radiation Injury Segmentation
Xinyu Xiong, Wuteng Cao, Zihuang Wu, Guanbin Li, Qiyuan Qin
MICCAI (4)1
2025 HFS-SAM2: Segment Anything Model 2 With High-Frequency Feature Supplementation for Camouflaged Object Detection
abstract
Camouflaged Object Detection (COD) aims to identify objects seamlessly blended with their backgrounds. While effective solutions exist for camouflaged animals, detecting camouflaged plants presents unique challenges and remains an open problem. This work introduces a novel Plant Camouflage Detection (PCD) method leveraging the Segment Anything Model 2 (SAM2). Our approach enhances the vanilla SAM2 decoder with specialized frequency-aware modules to improve the performance on PCD. Specifically, we employ a laplacian pyramid to extract high-frequency image components and introduce a High-Frequency Supplementation (HFS) module to enhance crucial spatial details for identifying camouflaged plants. The Multi-Scale Extraction (MSE) module is leveraged to capture rich multi-scale information, after which the features from the last three encoder layers are fused through a Cross-Layer Aggregation (CLA) module to obtain the aggregated high-level semantic features. A Semantic Gap Reduction (SGR) module is further proposed to bridge the semantic gap between high-level and shallow features during fusion. Finally, a Reverse Feature Mining (RFM) module is designed to highlight complementary regions and fine details. Extensive experiments on five datasets, encompassing both plant and animal camouflage detection, demonstrate the superior performance of our method compared to state-of-the-art approaches.
Zihuang Wu, Xinyu Xiong, Guangwei Gao, Hongwei Li 0017
IEEE Signal Process. Lett.2
2025 Free Meal: Boosting Semi-Supervised Polyp Segmentation by Harvesting Negative Samples
abstract
Existing semi-supervised polyp segmentation methods assume that unlabeled images are positive, containing lesions to be annotated, while neglecting negative samples that are widely available in practice. This letter reveals that harvesting lesion-free negative samples can effectively boost polyp segmentation performance. Directly extending the labeled set with negative samples is sub-optimal since it introduces potential class imbalance. To overcome this challenge, we first introduce a data augmentation strategy named TypeMix. By fusing unlabeled samples with negative samples, the network can better benefit from diverse features provided by negatives while alleviating the potential side effects. Furthermore, it is observed that the number of negative samples significantly exceeds that of lesion samples. To reduce redundancy and improve training efficiency, we propose a dynamic informativeness-aware sampling strategy, prioritizing the active selection of high-valuable negative samples. Extensive experiments on public datasets demonstrate that our simple but effective strategies are enough to consistently outperform other state-of-the-art methods, offering new possibilities for future work from a data collection perspective.
Xinyu Xiong, Wenxue Li 0003, Duojun Huang
IEEE Signal Process. Lett.1
2024 BA-SAM: Boundary-Aware Adaptation of Segment Anything Model for Medical Image Segmentation
abstract
The Segment Anything Model (SAM) has demonstrated remarkable capabilities in its performance on natural images. However, it faces considerable challenges when applied to medical datasets. Specifically, the performance of vanilla SAM is degraded and lacks generalisability when processing medical images with large domain gaps. What’s worse, many medical segmentation tasks highly demand accurate boundary identification, while existing SAM variants struggle with this need. To overcome the above challenges, we propose BA-SAM, a Segment Anything Model variant that can achieve better performance on medical images. Specifically, based on the idea of parameter-efficient fine-tuning (PEFT), we first add a parallel tuneable CNN encoder to better extract local details using convolutional operations, while most parts of the original ViT encoder in SAM are set frozen. Moreover, we use a Boundary-Aware Attention (BAA) module in the CNN-Branch to encourage the framework to better capture boundary-related features. Extensive experiments on three public datasets demonstrate that the proposed BA-SAM further improvements over existing state-of-the-art methods.
Xinyu Xiong, Huihui Fang, Yanwu Xu 0001
BIBM2
2024 AlignSAM: Aligning Segment Anything Model to Open Context via Reinforcement Learning
abstract
Powered by massive curated training data, Segment Any-thing Model (SAM) has demonstrated its impressive generalization capabilities in open-world scenarios with the guidance of prompts. However, the vanilla SAM is class-agnostic and heavily relies on user-provided prompts to segment objects of interest. Adapting this method to diverse tasks is crucial for accurate target identification and to avoid suboptimal segmentation results. In this paper, we propose a novel framework, termed AlignSAM, designed for automatic prompting for aligning SAM to an open context through reinforcement learning. Anchored by an agent, AlignSAM enables the generality of the SAM model across diverse downstream tasks while keeping its parameters frozen. Specifically, AlignSAM initiates a prompting agent to iteratively refine segmentation predictions by interacting with the foundational model. It integrates a reinforcement learning policy network to provide informative prompts to the foundational models. Additionally, a semantic recal-ibration module is introduced to provide fine-grained labels of prompts, enhancing the model's proficiency in handling tasks encompassing explicit and implicit semantics. Experiments conducted on various challenging segmentation tasks among existing foundation models demonstrate the superiority of the proposed AlignSAM over state-of-the-art approaches. Project page: https://github.com/Duojun-Huang/AIignSAM-CVPR2024.
Duojun Huang, Xinyu Xiong, Jichang Li, Zequn Jie, Lin Ma 0002, Guanbin Li
CVPR2
2024 TP-DRSeg: Improving Diabetic Retinopathy Lesion Segmentation with Explicit Text-Prompts Assisted SAM
Wenxue Li 0003, Xinyu Xiong, Peng Xia 0005, Lie Ju, ZongYuan Ge
MICCAI (8)2
2024 MoE-Polyp: Shifting More Attention to Small Polyp Segmentation via Mixture-of-Experts
Zihuang Wu, Xinyu Xiong, Ying Chen 0052
MMAsia2
2024 HybridVPS: Hybrid-Supervised Video Polyp Segmentation Under Low-Cost Labels
abstract
Deep polyp segmentation methods have shown remarkable potential in boosting diagnostic efficiency. Nevertheless, these methods rely on sufficient pixel-wise annotated data, which is time-consuming and labor-intensive to acquire in clinical practice. This challenge is further escalated under the polyp segmentation scenario due to the massive video frames. To alleviate annotating burden, in this letter, we propose a label-efficient polyp segmentation framework named HybridVPS, which drastically reduces the annotation cost while maintaining satisfactory performance. Our core insight is to take full advantage of the similar semantics between consecutive video frames. Specifically, only a few frames require pixel-wise annotations, while the cheap scribble annotations are enough for the remaining part. To fully leverage the coarse location information provided by scribble annotations, we introduce an adaptive label prompter, which utilizes pixel-wise annotation to provide reliable guidance for scribble-annotated neighboring frames, thus facilitating the overall accuracy of the segmentation. Extensive experiments on the large-scale video polyp dataset SUN-SEG demonstrate the superiority of our approach. HybridVPS achieves comparable performance to the fully supervised scheme while requiring only 2% of the pixel-level annotations.
Wenxue Li 0003, Xinyu Xiong, Fugui Fan
IEEE Signal Process. Lett.2
2023 Long-term Wind Power Forecasting with Hierarchical Spatial-Temporal Transformer
abstract
Wind power is attracting increasing attention around the world due to its renewable, pollution-free, and other advantages. However, safely and stably integrating the high permeability intermittent power energy into electric power systems remains challenging. Accurate wind power forecasting (WPF) can effectively reduce power fluctuations in power system operations. Existing methods are mainly designed for short-term predictions and lack effective spatial-temporal feature augmentation. In this work, we propose a novel end-to-end wind power forecasting model named Hierarchical Spatial-Temporal Transformer Network (HSTTN) to address the long-term WPF problems. Specifically, we construct an hourglass-shaped encoder-decoder framework with skip-connections to jointly model representations aggregated in hierarchical temporal scales, which benefits long-term forecasting. Based on this framework, we capture the inter-scale long-range temporal dependencies and global spatial correlations with two parallel Transformer skeletons and strengthen the intra-scale connections with downsampling and upsampling operations. Moreover, the complementary information from spatial and temporal features is fused and propagated in each other via Contextual Fusion Blocks (CFBs) to promote the prediction further. Extensive experimental results on two large-scale real-world datasets demonstrate the superior performance of our HSTTN over existing solutions.
Lingbo Liu, Xinyu Xiong, Guanbin Li, Liang Lin 0004
IJCAI3
2019 Channel Estimation for Millimeter Wave Wideband Massive MIMO Systems via Tensor Decomposition
abstract
The acquisition of channel state information is crucial in massive multiple-input multiple-output (MIMO) systems when large scale antenna array is used. For Millimeter Wave (mmWave) wideband massive MIMO-OFDM systems, both the number of antennas and subcarriers are very large, so it is very difficult for us to get accurate channel information. In this paper, we propose a tensor-based channel estimation scheme for mmWave wideband channels. Firstly, we transform the two- dimensional mmWave wideband channel matrix into a three-dimensional tensor due to the large number of antennas and subcarriers. Secondly, by utilizing the PARAFAC decomposition of the tensor, the estimated three factor matrices can be obtained. Then, channel parameters can be extracted from the estimated factor matrices. Simulation results show that our proposed method outperforms the conventional low complexity compressed sensing based method, such as orthogonal matching pursuit (OMP) method, while maintaining the same complexity order.
Long Cheng 0009, Guangrong Yue, Xinyu Xiong, Shaoqian Li
VTC Spring3