VLDB 2026 Research / reviewers in the wild / expert
Xiao Luan
dblp:130/5073
· DBLP profile ↗
20ranked-venue papers
10as first author
13since 2021 · last 2026
0000-0003-0010-7361ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Text-Guided Modality Fusion for Brain Tumor Segmentation
Xiao Luan, Xiongfeng Huang, Linghui Liu, Weisheng Li 0001, Xinbo Gao 0001 |
Pattern Recognit. | 1 |
| 2026 | Mutually Guided Fusion Learning for Collaborative Camouflaged Object SegmentationabstractCollaborative camouflaged object segmentation (CoCOS) is a challenging task, focusing on identifying objects that blend closely with their backgrounds by jointly processing intraclass images. Existing methods fail to fully leverage the shared features (e.g., shape, texture, and contour) from these intraclass images, which leads to poor segmentation performance in relatively complex scenarios. To address this issue, we propose a novel mutually guided fusion refinement network (MFRNet), which improves the model performance by more effectively collaborating and optimizing the shared information. Specifically, it includes feature encoding, single-image branch feature enhancement, multiimage branch feature enhancement, and mutual guidance. After the feature encoding step, we design the graph convolution self-attention (GCS) and spatial context exploration (SCE) modules to enhance multilevel features of the single-image and multiimage branches, respectively. Moreover, we propose a mutual guidance fusion (MGF) module to utilize cross-scene image information for mutual guidance and progressive refinement, enhancing intraclass collaboration for improving target feature distinction. Extensive experimental results demonstrate that our MFRNet significantly outperforms existing CoCOS methods, achieving a mean E-measure score of 0.846 on the CoCOD8K dataset. Our code will be published at https://github.com/another-u/MFRNet. Chen Li 0048, Xiao Luan, Linghui Liu, Yanzhao Su, Yule Fu, Weisheng Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Beyond Probability Guided Categorization: A Subspace Projection-Aware Approach for Medical Image SegmentationabstractMost current medical image segmentation (MIS) methods focus on the performance of models while overlooking the interpretability of models. For existing segmentation models, the number of channels of the output feature maps is equal to the number of segmentation categories. The category of each pixel is determined by the channel with the highest probability, yet the reasons for this decision are unclear. Inspired by the principle that samples from the same category are spatially closer, we propose a novel subspace projection method to enhance both the performance and interpretability of MIS models. Our approach replaces the output layer of baseline models with a new convolutional layer that enhances channel-wise representational power, and concatenates the channel features of each pixel into feature vectors to integrate channel-level information. Considering the sparse structure of feature vectors, an overcomplete dictionary learning is applied to extract the most discriminative features. These feature vectors are projected into category subspaces, with pixel classification determined based on the smallest projection distance. To guarantee that feature vectors within the same category subspace are grouped closely and to promote nonoverlapping feature learning across subspaces, we introduce the distance loss and the subspace orthogonality loss. Our method is evaluated on ten popular MIS models using two publicly available datasets. Experimental results show that our approach outperforms baseline models in terms of both segmentation accuracy and the interpretability of model decision. Xiao Luan, Yule Fu, Linghui Liu, Chen Li 0048, Weisheng Li 0001 |
BIBM | 1 |
| 2025 | CSSNet: A 3D medical image segmentation network based on compressed sparse dual-branch structure
Xiao Luan, Yule Fu, Linghui Liu, Weisheng Li 0001 |
Pattern Recognit. Lett. | 1 |
| 2025 | Reversible Feature Learning for Brain Tumor Segmentation With Incomplete ModalitiesabstractAccurate brain tumor segmentation is vital for clinical diagnosis and treatment. Due to motion artifacts and image damage, it is challenging to obtain accurate segmentation results of brain tumors in the presence of incomplete MRI modalities. We propose a multimodal reversible feature learning method to tackle this problem. This method can fully explore the potential feature similarities and complementarities between MRI modalities. To compensate the information of missing modalities, we propose a reversible feature interaction module. It explores information similarity among existing modalities as priors, which are used to reconstruct features of missing modalities at the feature level. With enhanced discriminative information, the model suppresses the noise in the missing modalities. Furthermore, we propose a dual-scale attention module to enhance the detail restoration and reconstruction accuracy of images. Comparison results on the BRATS challenge datasets show the superiority of our method over current popular methods. The code used in this study is available athttps://github.com/fybgogogo/reverse. Yanbing Fan, Linghui Liu, Xiao Luan, Weisheng Li 0001 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Learning from Inside: Self-driven Intra-modality Siamese Knowledge Generation and Inter-modality Alignment for Chest X-rays Vision-Language Pre-trainingabstractSince pathology occupies only a small portion of an X-ray, which means that a large portion of the information may be irrelevant to the paired radiology report, the Chest X-rays Report Understanding (CRU) task focuses on how to utilize small regions of the case to improve the performance of medical VLP. However, existing studies have neglected the fine-grained false negative samples of medical visual representations, resulting in their poor performance in CRU scenarios, which we attribute this to the fine-grained feature collapse problem. To address this issue, we propose an intra-modality siamese knowledge generation and inter-modality alignment framework, termed Chest X-rays Report Understanding Framework(CRUF). CRUF leverages the siamese knowledge in image-text pairs as guiding signals to distinguish fine-grained false negative and negative samples within the modality, and further narrows the distance between false negative and positive samples between modalities, accurately aligning the case regions of each image with the corresponding medical terms. Experimental results on multiple downstream medical image datasets covering tasks such as image classification, object detection, and semantic segmentation demonstrate the stability and outstanding performance of our framework. Code is available at https://github.com/cl-red/CRUF. Lihong Qiao, Yucheng Shu, Xiao Luan, Bin Xiao 0002 |
BIBM | 4 |
| 2024 | Boosting Robust Multi-Focus Image Fusion With Frequency Mask and Hyperdimensional ComputingabstractMulti-focus image fusion (MFIF) creates an image from different source images with various sensors or optical settings as the devices can’t focus all objects at different distances. Most of the MFIF methods have several limitations in encoder enough features from the images and the result are not robust. To overcome the primary issue, we present a robust fusion algorithm based on the Frequency mask and the Hyperdimensional computing. We propose the Frequency Mask Filter (FMF) to get the narrow-band signals by encoding the frequency domain vector through the mask filter in the frequency domain. The Hyperdimensional encoder uses monogenic mapping, in which the multi-modulation features (MMF) such as the frequency, phase and amplitude are dynamically selected to obtain robust focus maps. Generated by multiscale monogenic representations of each image, the narrow-band image are mapped to hypervector encoding. Hyperdimensional encoder shows the energetic and structural information and leads to robust fusion results. Our proposed method is far superior to the existing MFIF method in terms of both objective evaluation metrics and visual effects on three publicly available datasets.Additionally, our proposed method requires only 0.88 seconds and has a parameter count of 0.13 million for multi-focus image fusion. Lihong Qiao, Shixin Wu, Bin Xiao 0002, Yucheng Shu, Xiao Luan, Sicheng Lu, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Multi-level dynamic error coding for face recognition with a contaminated single sample per person
Xiao Luan, Linghui Liu, Weisheng Li 0001 |
Pattern Recognit. Lett. | 1 |
| 2023 | A Symmetrical Siamese Network Framework With Contrastive Learning for Pose-Robust Face RecognitionabstractFace recognition has achieved remarkable success owing to the development of deep learning. However, most of existing face recognition models perform poorly against pose variations. We argue that, it is primarily caused by pose-based long-tailed data - imbalanced distribution of training samples between profile faces and near-frontal faces. Additionally, self-occlusion and nonlinear warping of facial textures caused by large pose variations also increase the difficulty in learning discriminative features of profile faces. In this study, we propose a novel framework called Symmetrical Siamese Network (SSN), which can simultaneously overcome the limitation of pose-based long-tailed data and pose-invariant features learning. Specifically, two sub-modules are proposed in the SSN, i.e., Feature-Consistence Learning sub-Net (FCLN) and Identity-Consistence Learning sub-Net (ICLN). For FCLN, the inputs are all face images on training dataset. Inspired by the contrastive learning, we simulate pose variations of faces and constrain the model to focus on the consistent areas between the original face image and its corresponding virtual pose face images. For ICLN, only profile images are used as inputs, and we propose to adopt Identity Consistence Loss to minimize the intra-class feature variation across different poses. The collaborative learning of two sub-modules guarantees that the parameters of network are updated in a relatively equal probability between near-frontal face images and profile images, so that the pose-based long-tailed problem can be effectively addressed. The proposed SSN shows comparable results over the state-of-the-art methods on several public datasets. In this study, LightCNN is selected as the backbone of SSN, and existing popular networks also can be used into our framework for pose-robust face recognition. Xiao Luan, Zibiao Ding, Linghui Liu, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | Partial-to-Partial Point Cloud Registration Based on Multi-Level Semantic-Structural Cognitionabstract3D point cloud registration attempt to establish spatial correspondences between the source point cloud and the target point cloud. It is a fundamental task in computer vision and multimedia applications. Recently, many learning-based methods have been proposed and achieved promising performance. However, in the Partial-to-Partial (PtP) registration problem, the existence of a large number of external points may greatly handicap the effectiveness of these methods. In this paper, we propose to address the PtP issue under a novel multi-task cognition framework. At the global semantic level, we introduce a multi-scale feature exchanging network to actively evaluate the matching credibility. For the local structural learning, an inner-inter attention fusion branch is applied to generate discriminative features. Moreover, we integrate a novel alternating correspondences searching mechanism with a flexible bi-directional dislocation loss to perform robust learning, and a simple yet effective SVD weighting scheme is introduced at the inference stage. Experiment results on two challenging PtP 3D point cloud registration data sets show that our proposed method outperforms all the SOTA methods with higher precision and robustness. Yucheng Shu, Zongzhuang Hou, Bin Xiao 0002, Xiuli Bi, Xiao Luan, Weisheng Li 0001 |
ICME | 5 |
| 2022 | Learning Unsupervised Face Normalization Through Frontal View ReconstructionabstractFace normalization from large pose is a challenging problem. Many Generative Adversarial Network (GAN) based models can infer frontal view of profile faces, while they require paired faces and pose label. Instead, we focus on frontal face synthesis with unpaired and unlabeled training data. We present a Frontal View Reconstruction based GAN (FVR-GAN) for large pose face normalization and recognition. The generator of FVR-GAN can be considered as a dual-input auto-encoder, where the identity encoder extracts identity features from an identity image and the template encoder extracts contour features from a frontal image. The decoder combines those two kinds of features and synthesizes a corresponding frontal face. To learn face normalization effectively, we incorporate the Frontal View Reconstruction (FVR) operation into training stage. The FVR operation includes self-reconstruction and frontalization mapping. A group of sub-discriminators which receive different facial parts are employed for discrimination. Considering that different face parts have different contributions to discrimination, we introduce a dynamic weighting mechanism to balance the output of sub-discriminators. FVR-GAN can recover high-quality frontal images under arbitrary poses. Experimental results on datasets of Multi-PIE, IJB-A, LFW and CFP demonstrate the efficacy of our model in terms of quality of synthesized images and recognition accuracy. Xiao Luan, Jiezhong Zheng, Weisheng Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Rubik-Net: Learning Spatial Information via Rotation-Driven Convolutions for Brain SegmentationabstractThe accurate segmentation of brain tissue in Magnetic Resonance Image (MRI) slices is essential for assessing neurological conditions and brain diseases. However, it is challenging to segment MRI slices because of the low contrast between different brain tissues and the partial volume effect. 2-Dimensional (2-D) convolutional networks cannot handle such volumetric image data well because they overlook spatial information between MRI slices. Although 3-Dimensional (3-D) convolutions capture volumetric spatial information, they have not been fully exploited to enhance representative ability of deep networks; moreover, they may lead to overfitting with insufficient training data. In this paper, we propose a novel convolutional mechanism, termed Rubik convolution, to capture multi-dimensional information between MRI slices. Rubik convolution rotates the axis of a set of consecutive slices, enabling 2-D convolution kernels to extract features of each axial plane simultaneously. Next, feature maps are rotated back to fuse multidimensional information by the Max-View-Maps. Furthermore, we propose an efficient 2-D convolutional network, namely Rubik-Net, where the residual connections and the bottleneck structure are used to enhance information transmission and reduce the number of network parameters. The Rubik-Net shows promising results on iSeg2017, iSeg2019, IBSR and BrainWeb datasets in terms of segmentation accuracy. In particular, we achieved the best results in 95th percentile Hausdorff distance and average surface distance in cerebrospinal fluid segmentation on the most challenging iSeg2019 dataset. The experiments indicate that Rubik-Net improves the accuracy and efficiency of medical image segmentation. Moreover, Rubik convolution can be easily embedded into existing 2-D convolutional networks. Xiao Luan, Xinyu zheng, Weisheng Li 0001, Linghui Liu, Yucheng Shu |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | Collaborative Learning With a Multi-Branch Framework for Feature EnhancementabstractFeature representation is highly important for many computer vision tasks. A broad range of prior studies have been proposed to strengthen representation ability of architectures via built-in blocks. However, during the forward propagation, the reduction in feature map scales still leads to the lack of representation ability. In this paper, we focus on boosting the representational power of a convolutional network by the multi-branch framework that we term the BranchNet. Each branch is directly supervised by label information to enrich the hierarchy features in BranchNet. Based on this framework, we further propose a collaborative learning loss and a soft target loss to transfer knowledge from deeper layers to shallow layers. BranchNet is an efficient training framework without extra parameters introduced in inference and can be integrated in existing networks, e.g., VGG, ResNet, and DenseNet. We evaluate BranchNet on all of these models and find that our method outperforms the baseline models on the widely-used CIFAR and ImageNet datasets. In particular, on the CIFAR-100 dataset, the classification error of ResNet-164 with BranchNet decreases by 4.51 percent. We also conduct experiments on the representative computer vision tasks of instance segmentation and class activation mapping, further verifying the superiority of BranchNet over the baseline models. Models and code are available athttps://github.com/zyyupup/BranchNet/. Xiao Luan, Weihua Ou, Linghui Liu, Weisheng Li 0001, Yucheng Shu, Hongmin Geng |
IEEE Trans. Multim. | 1 |
| 2020 | AFT-Net: Active Fusion-Transduction for Multi-stream Medical Image SegmentationabstractAs an important building block in automatic medical applications, image segmentation has made a great progress due to the data-driving mechanism of deep architecture. Recently, numerous methods have been proposed to boost the segmentation performance based on U-shape network. However, they often built feature encoders with only one data routine, which have limited the representation ability of the networks. Although some methods applied multiple learning paths to fix this problem, the deep supervision techniques are required to monitor the training status at individual path, which may bring extra burden to practical usage of the algorithm. Additionally, under these frameworks, the semantic gap between different paths may interfere with model's learning performance, and the potential transduction ability of skip connections still needs further investigation. To address these issues, we introduce a novel medical image segmentation framework, namely AFT-Net, in which an attention-based data fusion model is proposed to effectively cooperate with multi-stream encoder. By progressively accumulating the features from different paths, our method can establish meaningful connections between structural and semantic features, while keeping an integral and flexible layout without deeply customized supervisions. Extensive experiments on two medical image data sets demonstrate that our method is able to acquire image features with both diversity and quality, thereby outperforms current state-of-the-art segmentation methods. Yucheng Shu, Bin Xiao 0002, Xiao Luan, Linghui Liu, Chunlong Hu |
ICTAI | 4 |
| 2020 | Global-Local Mutual Guided Learning for Person Re-identification
Junheng Chen, Xiao Luan, Weisheng Li 0001 |
PRCV (2) | 2 |
| 2018 | Pose-robust face recognition with Huffman-LBP enhanced by Divide-and-Rule strategy
Lifang Zhou, Yue-Wei Du, Weisheng Li 0001, Jian-Xun Mi, Xiao Luan |
Pattern Recognit. | 5 |
| 2018 | Robust discriminative nonnegative dictionary learning for occluded face recognition
Weihua Ou, Xiao Luan, Jianping Gou, Quan Zhou 0004, Wenjun Xiao, Xiangguang Xiong, Wu Zeng |
Pattern Recognit. Lett. | 2 |
| 2014 | Extracting sparse error of robust PCA for face recognition in the presence of varying illumination and occlusion
Xiao Luan, Bin Fang 0001, Linghui Liu, Weibin Yang, Jiye Qian |
Pattern Recognit. | 1 |
| 2013 | Face recognition with contiguous occlusion using linear regression and level set method
Xiao Luan, Bin Fang 0001, Linghui Liu, Lifang Zhou |
Neurocomputing | 1 |
| 2013 | Exploiting local intensity information in Chan-Vese model for noisy image segmentation
Linghui Liu, Xiao Luan |
Signal Process. | 4 |