Xin Luo 0009

dblp:53/5106-9 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0003-1641-5713ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 A Cost-effective Solution for Remote Sensing Image Segmentation via Train/Test-Time Adaptation
abstract
Remote Sensing Image (RSI) segmentation has made significant strides, emerging as a leading solution for interpreting remote sensing data. However, due to the substantial domain gap between different remote sensors and limited computational resources, existing RSI segmentation methods often suffer from poor generalization performance. To address these challenges, we propose a COst-effective and Whole-process Domain Adaptation solution, namely COWDA, which adapts models at both the train and test time through three key phases: 1) Source-data Domain Alignment: We employ traditional image stylization techniques to translate source images into target-style alternatives, which avoids the computationally intensive need for auxiliary neural network models. 2) Target-data Train-time Fine-tuning: We propose a joint positive and negative learning (JPNL) algorithm that adds both positive and negative samples to effectively learn domain-invariant knowledge from noisy pseudo-labeled target data. 3) Test-time Adaptation: We propose an entropy-weighted test-time adaptation strategy to update the trained model with online test samples, further enhancing its performance. Extensive experiments on two widely-used domain adaptation benchmarks for remote sensing show that COWDA improves state-of-the-art counterparts by 1.4% and 2.6% F1 scores, respectively.
Wei Chen 0009, Xin Luo 0009, Yu-Lin He, Tianhang Guo, Yuhua Tang
ICASSP2
2024 Crots: Cross-Domain Teacher-Student Learning for Source-Free Domain Adaptive Semantic Segmentation
Xin Luo 0009, Wei Chen 0009, Zhengfa Liang, Chen Li 0034
Int. J. Comput. Vis.1
2023 Domain Generalized Fundus Image Segmentation via Dual-Level Mixing
abstract
Single domain generalization plays a vital role in signal processing tasks, which is capable of extracting domain-invariant knowledge from a single source domain such that the learned model can be generalizable to unseen domains. However, many existing methods require the help of auxiliary tasks, bringing extra computation cost. Aiming at this pitfall, this study proposes Dual-Level Mixing (DLM) to boost the diversity of the single source domain and enhance the generalization performance. Specifically, at the input level, we apply different image augmentations to get different variants of the same image. Patches from augmented images are spatially mixed to get a perturbed image as the input, which enlarges the scale of training data and boosts the diversity. Meanwhile, at the feature level, we characterize feature statistics as Gaussian distributions. Then, we resample the feature statistics to renormalize the latent features, which simulates the potential style variance of different domains. In summary, the proposed DLM method synergies input-level and feature-level mixing strategies, leading to enhanced generalization performance. Experimental results of cross-domain fundus image segmentation demonstrate that either input-level mixing or feature-level mixing can effectively promote the performance of domain generalization. Moreover, the collaboration of dual-level mixing strategies leads to superior or comparable performance to domain adaptation counterparts that rely on target data.
Xin Luo 0009, Wei Chen 0009, Chen Li 0034, Bin Zhou 0004, Yusong Tan
ICASSP1
2023 A knowledge-based learning framework for self-supervised pre-training towards enhanced recognition of biomedical microscopy images
abstract
Self-supervised pre-training has become the priory choice to establish reliable neural networks for automated recognition of massive biomedical microscopy images, which are routinely annotation-free, without semantics, and without guarantee of quality. Note that this paradigm is still at its infancy and limited by closely related open issues: (1) how to learn robust representations in an unsupervised manner from unlabeled biomedical microscopy images of low diversity in samples? and (2) how to obtain the most significant representations demanded by a high-quality segmentation? Aiming at these issues, this study proposes a knowledge-based learning framework (TOWER) towards enhanced recognition of biomedical microscopy images, which works in three phases by synergizing contrastive learning and generative learning methods: (1) Sample Space Diversification: Reconstructive proxy tasks have been enabled to embed a priori knowledge with context highlighted to diversify the expanded sample space; (2) Enhanced Representation Learning: Informative noise-contrastive estimation loss regularizes the encoder to enhance representation learning of annotation-free images; (3) Correlated Optimization: Optimization operations in pre-training the encoder and the decoder have been correlated via image restoration from proxy tasks, targeting the need for semantic segmentation. Experiments have been conducted on public datasets of biomedical microscopy images against the state-of-the-art counterparts (e.g., SimCLR and BYOL), and results demonstrate that: TOWER statistically excels in all self-supervised methods, achieving a Dice improvement of 1.38 percentage points over SimCLR. TOWER also has potential in multi-modality medical image analysis and enables label-efficient semi-supervised learning, e.g., reducing the annotation cost by up to 99% in pathological classification.
Wei Chen 0009, Chen Li 0034, Dan Chen 0001, Xin Luo 0009
Neural Networks4
2023 Adversarial style discrepancy minimization for unsupervised domain adaptation
abstract
Mainstream unsupervised domain adaptation (UDA) methods align feature distributions across different domains via adversarial learning. However, most of them focus on global distribution alignment, ignoring the fine-grained domain discrepancy. Besides, they generally require auxiliary models, bringing extra computation costs. To tackle these issues, this study proposes an UDA method that differentiates individual samples without the help of extra models. To this end, we introduce a novel discrepancy metric, termed style discrepancy, to distinguish different target samples. We also propose a paradigm for adversarial style discrepancy minimization (ASDM). Specifically, we fix the parameters of the feature extractor and maximize style discrepancy to update the classifier, which helps detect more hard samples. Adversely, we fix the parameters of the classifier and minimize the style discrepancy to update the feature extractor, pushing those hard samples near the support of the source distribution. Such adversary helps to progressively detect and adapt more hard samples, leading to fine-grained domain adaptation. Experiments on different UDA tasks validate the effectiveness of ASDM. Overall, without any extra models, ASDM reaches a 46.9% mIoU in the GTA5 to Cityscapes benchmark and an 84.7% accuracy in the VisDA-2017 benchmark, outperforming many existing adversarial-learning-based methods.
Xin Luo 0009, Wei Chen 0009, Zhengfa Liang, Chen Li 0034, Yusong Tan
Neural Networks1
2022 Adaptive Pseudo Labeling for Source-Free Domain Adaptation in Medical Image Segmentation
abstract
Domain adaptation is common but challenging in signal processing tasks due to the intrinsic discrepancy, especially in difficult-to-label medical image segmentation application scenarios. Pseudo labeling methods are widely utilized to compensate for the scarcity of annotation. However, most existing methods set the fixed thresholds to select highly-confident predictions as pseudo labels, inevitably generating false labels with noise. In this paper, we combine the dual-classifiers consistency and predictive category-aware confidence to form a novel regularization for pseudo-label denoising. The dual-classifiers consistency helps promote the robustness of pseudo labels. Meanwhile, category-aware confidence is utilized as adaptive pixel-wise weights, avoiding the need for handcrafted thresholds. The adapted model is refined by the rectified pseudo labels without source domain samples. The proposed method is model-independent and thus can be plug-and-play to improve existing UDA methods. We validated it on the cross-modality medical image segmentation and obtained more competitive results.
Chen Li 0034, Wei Chen 0009, Xin Luo 0009, Yu-Lin He, Yusong Tan
ICASSP3
2021 Tri-Directional Tasks Complementary Learning for Unsupervised Domain Adaptation of Cross-modality Medical Image Semantic Segmentation
abstract
Cross-modality adaptation is challenging due to the internal domain discrepancy in appearance and representation. When the trained model of source domain is transferred to the target domain, the domain shift will reduce accuracy. Meanwhile, unsupervised domain adaptation has the potential to recover this degradation among medical images of different modalities, so it is of clinical significance and meaningful in bioinformatics. However, previous related works usually try to align domains in a single direction or two directions, failing to take advantage of the complementary relationship between different directions and alignment tasks. In this paper, we propose the Tri-directional learning framework to solve domain shift in the task of medical image semantic segmentation. The proposed framework is able to synergize image style transformation, mask segmentation and edge segmentation. The above three tasks are mutually boosted through complementary training in each iteration. In this way, our method performs cross-modality medical images semantic segmentation from labeled source domain (MRI) to unlabeled target domain (CT). The experimental results demonstrate the effectiveness of the proposed method. For the task of cardiac structure segmentation from cross-modality medical images, our proposed framework achieves state-of-the-art performance. The code is available at https://github.com/lichen14/TriDL.
Chen Li 0034, Wei Chen 0009, Mingfei Wu, Xin Luo 0009, Yu-Lin He, Yusong Tan
BIBM4
2021 AttENT: Domain-Adaptive Medical Image Segmentation via Attention-Aware Translation and Adversarial Entropy Minimization
abstract
Due to the intrinsic domain shift among different modalities, it is nontrivial to directly apply a well-trained model into other cross-modality medical images. Unsupervised domain adaptation (UDA) has the potential to reduce such domain shift. However, existing UDA methods try to align domains in either image level or in feature level, failing to consider the unified relationship between cross-modality images and their corresponding features. In this paper, we propose a novel UDA framework for domain adaptive medical image segmentation. The proposed framework synergizes both pixel space and entropy space for domain alignment. Specifically, in the pixel space, we introduce the attention mechanism into CycleGAN, and enhance the semantic and geometric consistency of the target organs during the image style transformation. In entropy space, we utilize entropy minimization principle to force consistent image segmentation between well-annotated source domain and non-annotated target domain. The aligned ensemble of two representation spaces enables a well-trained segmentation model to effectively transfer from source domain to target domain. The experimental results demonstrate the effectiveness of the proposed method. For the task of multi-organs segmentation from cross-modality medical images, our proposed framework achieves state-of-the-art performance, with some specific metric even superior to those of supervised methods. The code is available at https://github.com/lichen14/AttENT.
Chen Li 0034, Xin Luo 0009, Wei Chen 0009, Yu-Lin He, Mingfei Wu, Yusong Tan
BIBM2
2021 Multi-Scale Cascade Disparity Refinement Stereo Network
abstract
Stereo matching has attracted much attention in recent years. Traditional methods can quickly generate a disparity result, but the accuracy is low. On the contrary, methods based on neural networks can achieve a high accuracy level, but they are difficult to reach the real-time level. Therefore, this paper presents MCDRNet, which combines traditional methods with neural networks to achieve real-time and accurate stereo matching results. Concretely, our network first generates a rough disparity map based on the traditional ADCensus algorithm. Then we design a novel Multi-Scale Cascade Network to refine the disparity map from coarse to fine. We evaluate our best-trained model on the KITTI official website. The results show that our network is much faster than most current top-performing methods(31×than CSPN, 56×than GANet, etc.). Meanwhile, it is more accurate than traditional stereo methods(SGM, SPS-St) and other fast 2D convolution networks(Fast DS-CS, DispNetC, etc.), demonstrating the rationalities and feasibilities of our method.
Xiaogang Jia, Wei Chen 0009, Zhengfa Liang, Xin Luo 0009, Mingfei Wu, Yusong Tan, Libo Huang 0002
ICASSP4
2021 Fast and Accurate Lane Detection via Frequency Domain Learning
abstract
It is desirable to maintain both high accuracy and runtime efficiency in lane detection. State-of-the-art methods mainly address the efficiency problem by direct compression of high-dimensional features. These methods usually suffer from information loss and cannot achieve satisfactory accuracy performance. To ensure the diversity of features and subsequently maintain information as much as possible, we introduce multi-frequency analysis into lane detection. Specifically, we propose a multi-spectral feature compressor (MSFC) based on two-dimensional (2D) discrete cosine transform (DCT) to compress features while preserving diversity information. We group features and associate each group with an individual frequency component, which incurs only 1/7 overhead of one-dimensional convolution operation but preserves more information. Moreover, to further enhance the discriminability of features, we design a multi-spectral lane feature aggregator (MSFA) based on one-dimensional (1D) DCT to aggregate features from each lane according to their corresponding frequency components. The proposed method outperforms the state-of-the-art methods (including LaneATT and UFLD) on TuSimple, CULane, and LLAMAS benchmarks. For example, our method achieves 76.32% F1 at 237 FPS and 76.98% F1 at 164 FPS on CULane, which is 1.23% and 0.30% higher than LaneATT. Our code and models are available at https://github.com/harrylin-hyl/MSLD.
Yu-Lin He, Wei Chen 0009, Zhengfa Liang, Dan Chen 0001, Yusong Tan, Xin Luo 0009, Chen Li 0034, Yulan Guo
ACM Multimedia6
2020 Attention Unet++: A Nested Attention-Aware U-Net for Liver CT Image Segmentation
abstract
Liver cancer is one of the cancers with the highest mortality. In order to help doctors diagnose and treat liver lesion, an automatic liver segmentation model is urgently needed due to manually segmentation is time-consuming and error-prone. In this paper, we propose a nested attention-aware segmentation network, named Attention UNet++. Our proposed method has a deep supervised encoder-decoder architecture and a redesigned dense skip connection. Attention UNet++ introduces attention mechanism between nested convolutional blocks so that the features extracted at different levels can be merged with a task-related selection. Besides, due to the introduction of deep supervision, the prediction speed of the pruned network is accelerated at the cost of modest performance degradation. We evaluated proposed model on MICCAI 2017 Liver Tumor Segmentation (LiTS) Challenge Dataset. Attention UNet++ achieved very competitive performance for liver segmentation.
Chen Li 0034, Yusong Tan, Wei Chen 0009, Xin Luo 0009, Yuanming Gao, Xiaogang Jia, Zhiying Wang 0003
ICIP4
2020 CenterRepp: Predict Central Representative Point Set's Distribution For Detection
abstract
Object detection has long been an important issue in the discipline of scene understanding. Existing researches mainly focus on the object itself, ignoring its surrounding environment. In fact, the surrounding environment provides abundant information to help detectors classify and locate objects. This paper proposes CRPDet, viz. CenterRepp Detector, a framework for object detection. The main function of CRPDet is accomplished by the CenterRepp module, which takes into account the surrounding environment by predicting the distribution of the central representative points. CenterRepp converts labeled object frames into the mean and standard variance of the sampling points' distribution. This helps increase the receptive field of objects, breaking the limitation of object frames. CenterRepp defines a position-fixed center point with significant weights, avoiding to sample all points in the surroundings. Experiments on the COCO test-dev detection benchmark demonstrates that our proposed CRPDet has comparable performance with state-of-the-art detectors, achieving 39.4 mAP with 51 FPS tested under single size input.
Yu-Lin He, Limeng Zhang, Wei Chen 0009, Xin Luo 0009, Xiaogang Jia, Chen Li 0034
ICPR4
2020 ANU-Net: Attention-based nested U-Net to exploit full resolution features for medical image segmentation
Chen Li 0034, Yusong Tan, Wei Chen 0009, Xin Luo 0009, Yu-Lin He, Yuanming Gao
Comput. Graph.4