VLDB 2026 Research / reviewers in the wild / expert
Liuxin Bao
dblp:355/0767
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0003-3725-1586ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Text-Image Guided Retrieval Network for Triple-Modal Images Few-Shot Semantic SegmentationabstractThis letter presents a novel framework for triple-modal few-shot semantic segmentation involving visible, depth, and thermal modalities. While multi-modal approaches aim to provide complementary data, existing models often suffer from a heavy reliance on the visual appearance of support samples and lack effective semantic guidance during cross-modal fusion. To address these issues, we build a Text-Image Guided Retrieval Network (TIGRNet), where we leverage a Vision-Language Model (VLM) to inject semantic priors into the segmentation process. Specifically, we propose a Text Guided Fusion (TGF) module that dynamically regulates modality weights using global semantic cues, thereby suppressing noise and enhancing reliable features. Furthermore, we introduce a dual-branch retrieval strategy consisting of an Image-Guided Retrieval(IGR) module to capture visual appearance and a Text-Guided Retrieval(TGR) module to retrieve regions consistent with target semantics, reducing the burden of precise visual matching. Experiments on the VDT-2048-5$^{i}$dataset demonstrate that our TIGRNet outperforms existing state-of-the-art models in both 1-shot and 5-shot settings. The dataset and code are available athttps://github.com/FengshuoChen/TIGRNet. Fengshuo Chen, Liuxin Bao, Juting Miao, Xiaofei Zhou 0003 |
IEEE Signal Process. Lett. | 2 |
| 2026 | Reparameterization-Driven Depthwise Separable Large-Kernel Network for Lightweight Salient Object Detection of Strip Steel Surface DefectsabstractWith the rapid development of neural networks, strip steel surface defect detection, as an important task in computer vision, has achieved remarkable progress. However, state-of-the-art methods still face a tradeoff between accuracy and efficiency. High-performing models are usually large and computationally expensive, whereas lightweight models often suffer from limited detection accuracy. To address this issue, we first propose a spatial channel enhancement (SCE) module, which consists of a reparameterizable depthwise large-kernel convolution and a reparameterizable pointwise (RepPw) convolution. The proposed SCE module enlarges the receptive field and strengthens long-range spatial and channel interactions while preserving computational efficiency. Based on the SCE module, we propose a novel lightweight saliency model for strip steel surface defects, namely, reparameterization-driven depthwise separable large-kernel network (RepDSLKNet). RepDSLKNet employs SCE modules to build an encoder and a decoder, and utilizes cascaded channel attention (CCA) modules for the feature fusion. The lightweight architecture can effectively extract and fuse the semantic information and detailed features of strip steel surface defects, thereby improving the accuracy and speed of detection with a small model size. With an input size of $224 \times 224$ , our RepDSLKNet has only 0.47 M parameters and 0.42 G FLOPs during inference. Compared to the current state-of-the-art methods, our approach achieves a 19-fold improvement in throughput and a twofold reduction in latency. Experiments on two public strip steel defect datasets demonstrate that RepDSLKNet delivers competitive performance against state-of-the-art methods. Xiaofei Zhou 0003, Zhenkun Mo, Gongyang Li, Liuxin Bao, Xiaobin Xu 0002, Jiyong Zhang 0001 |
IEEE Trans. Cybern. | 4 |
| 2025 | Multi-modal feature integration network for Visible-Depth-Thermal salient object detection
Fengyv Cui, Xiaofei Zhou 0003, Liuxin Bao, Bin Wan, Jiyong Zhang 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Bidirectionally Guided Multi-Scale Feature Decoding Network for High-Resolution Salient Object DetectionabstractABSTRACT With the advancement of modern camera technology, the resolution and quality of images have been significantly improved. High‐resolution images can provide more detailed and clearer information, but they may also introduce more noise, which interferes with the detection and localization of salient objects. To address this issue, existing high‐resolution salient object detection methods either design complex network structures or adopt multi‐modal fusion. However, these approaches often consume significant computing and storage resources. This leads to redundancy of irrelevant features and loss of critical details. In this paper, we propose a network called bidirectionally guided multi‐scale feature decoding network for high‐resolution salient object detection. The model incorporates a bidirectional guidance method to explore the complementarity between encoding and decoding features, thereby achieving a comprehensive combination and enhancement of features. Additionally, in the decoder, multi‐scale encoding features are obtained and utilized sequentially to enhance feature learning and improve the accuracy of salient object detection. Specifically, our model consists of an encoder, a guided multi‐scale feature enhancement (GMFE) module, a guided feature fusion (GFF) module, and a multi‐scale feature decoder (MFD) module. First, multi‐scale encoding features are extracted through the encoder. These features are then fed into the GMFE module to enhance the multi‐scale encoding features under the guidance of saliency map derived from the decoding features of the previous layer. Subsequently, in the GFF module, the enhanced encoding features are fused with the decoding features from the previous layer. Finally, in the MFD module, the bidirectionally guided multi‐scale encoding features is integrated to generate an accurate saliency map. Experiments on two high‐resolution and two low‐resolution datasets demonstrate that our model outperforms on high‐resolution datasets while maintaining competitive performance on low‐resolution datasets, underscoring its effectiveness across varying image qualities. Jiangping Tang, Shuyao Guo, Xiaofei Zhou 0003, Liuxin Bao, Jiyong Zhang 0001 |
IET Image Process. | 4 |
| 2025 | GLNet: Global-Local Fusion Network for Strip Steel Surface Defects DetectionabstractSurface defect detection in strip steel is a critical task in industrial quality control. However, existing methods struggle with capturing both local details and global context effectively. In this paper, we propose the Global-Local Fusion Network (GLNet) for strip steel surface defect detection, which combines the advantages of VMamba's global feature extraction and CNN's local feature modeling. GLNet employs an encoder-decoder structure, where the encoder consists of two parallel branches: one based on VMamba for capturing global features and the other using ResNet50 for extracting local features. In the decoder, a Global-Local Fusion (GLF) module integrates these features using the Cross Prototype Objective Enhancement (CPOE) and Selective Spatial and Channel Attention (SSCA) modules. The CPOE module facilitates the interaction and fusion between global and local features, while the SSCA module digs the multi-scale information from the global feature through dynamic attention to guide the feature aggregation. Extensive experiments on the ESDIs dataset, demonstrate that GLNet achieves state-of-the-art performance in defect detection, surpassing 13 existing methods in both quantitative and qualitative metrics. Liuxin Bao, Xiaofei Zhou 0003, Xiaobin Xu 0002 |
IEEE Signal Process. Lett. | 2 |
| 2025 | IFENet: Interaction, Fusion, and Enhancement Network for V-D-T Salient Object DetectionabstractVisible-depth-thermal (VDT) salient object detection (SOD) aims to highlight the most visually attractive object by utilizing the triple-modal cues. However, existing models don't give sufficient exploration of the multi-modal correlations and differentiation, which leads to unsatisfactory detection performance. In this paper, we propose an interaction, fusion, and enhancement network (IFENet) to conduct the VDT SOD task, which contains three key steps including the multi-modal interaction, the multi-modal fusion, and the spatial enhancement. Specifically, embarking on the Transformer backbone, our IFENet can acquire multi-scale multi-modal features. Firstly, the inter-modal and intra-modal graph-based interaction (IIGI) module is deployed to explore inter-modal channel correlation and intra-modal long-term spatial dependency. Secondly, the gated attention-based fusion (GAF) module is employed to purify and aggregate the triple-modal features, where multi-modal features are filtered along spatial, channel, and modality dimensions, respectively. Lastly, the frequency split-based enhancement (FSE) module separates the fused feature into high-frequency and low-frequency components to enhance spatial information (i.e., boundary details and object location) of the salient object. Extensive experiments are performed on VDT-2048 dataset, and the results show that our saliency model consistently outperforms 13 state-of-the-art models. Our code and results are available at https://github.com/Lx-Bao/IFENet. Liuxin Bao, Xiaofei Zhou 0003, Bolun Zheng, Runmin Cong, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | AGFNet: Adaptive Gated Fusion Network for RGB-T Semantic SegmentationabstractRGB-T semantic segmentation can effectively pop-out objects from challenging scenarios (e.g., low illumination and low contrast environments) by combining RGB and thermal infrared images. However, the existing cutting-edge RGB-T semantic segmentation methods often present insufficient exploration of multi-modal feature fusion, where they overlook the differences between the two modalities. In this paper, we propose an adaptive gated fusion network (AGFNet) to conduct RGB-T semantic segmentation, where the multi-modal features are combined via the gating mechanisms and the spatial details are enhanced via the introduction of edge information. Specifically, the AGFNet employs a cross-modal adaptive gated-attention fusion (CAGF) module to aggregate the RGB and thermal features, where we give a sufficient exploration of the complementarity between the two-modal features via the gated attention unit (GAU). Particularly, in GAU, the gates can be used to purify the features, and the channel and spatial attention mechanisms are further employed to enhance the two-modal features interactively. Then, we design an edge detection (ED) module to learn the object-related edge cues, which simultaneously incorporates local detail information from low-level features and global location information from high-level features. After that, we deploy the edge guidance (EG) module to emphasize the spatial details of the fused features. Next, we deploy the contextual elevation (CE) module to enrich the contextual information of features by iteratively introducing the sine and cosine functions. Finally, considering that the quality of thermal images is usually lower than that of RGB images, we progressively integrate the multi-level RGB encoder features with multi-level decoder features, thereby focusing more on appearance information. Following this way, we can acquire the final high-quality segmentation result. Extensive experiments are performed on three public datasets including MFNet, PST900 and FMB datasets, and the experimental results show that our method achieves competitive performance when compared with the 22 state-of-the-art methods. Xiaofei Zhou 0003, Liuxin Bao, Haibing Yin, Qiuping Jiang, Jiyong Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | GINet:Graph interactive network with semantic-guided spatial refinement for salient object detection in optical remote sensing images
Chenwei Zhu, Xiaofei Zhou 0003, Liuxin Bao, Hongkui Wang, Shuai Wang 0003, Zunjie Zhu, Chenggang Yan 0001, Jiyong Zhang 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | Quality-Aware Selective Fusion Network for V-D-T Salient Object DetectionabstractDepth images and thermal images contain the spatial geometry information and surface temperature information, which can act as complementary information for the RGB modality. However, the quality of the depth and thermal images is often unreliable in some challenging scenarios, which will result in the performance degradation of the two-modal based salient object detection (SOD). Meanwhile, some researchers pay attention to the triple-modal SOD task, namely the visible-depth-thermal (VDT) SOD, where they attempt to explore the complementarity of the RGB image, the depth image, and the thermal image. However, existing triple-modal SOD methods fail to perceive the quality of depth maps and thermal images, which leads to performance degradation when dealing with scenes with low-quality depth and thermal images. Therefore, in this paper, we propose a quality-aware selective fusion network (QSF-Net) to conduct VDT salient object detection, which contains three subnets including the initial feature extraction subnet, the quality-aware region selection subnet, and the region-guided selective fusion subnet. Firstly, except for extracting features, the initial feature extraction subnet can generate a preliminary prediction map from each modality via a shrinkage pyramid architecture, which is equipped with the multi-scale fusion (MSF) module. Then, we design the weakly-supervised quality-aware region selection subnet to generate the quality-aware maps. Concretely, we first find the high-quality and low-quality regions by using the preliminary predictions, which further constitute the pseudo label that can be used to train this subnet. Finally, the region-guided selective fusion subnet purifies the initial features under the guidance of the quality-aware maps, and then fuses the triple-modal features and refines the edge details of prediction maps through the intra-modality and inter-modality attention (IIA) module and the edge refinement (ER) module, respectively. Extensive experiments are performed on VDT-2048 dataset, and the results show that our saliency model consistently outperforms 13 state-of-the-art methods with a large margin. Our code and results are available at https://github.com/Lx-Bao/QSFNet. Liuxin Bao, Xiaofei Zhou 0003, Xiankai Lu, Yaoqi Sun, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | 360$^{\circ }$ Omnidirectional Salient Object Detection with Multi-scale Interaction and Densely-Connected Prediction
Haowei Dai, Liuxin Bao, Kunye Shen, Xiaofei Zhou 0003, Jiyong Zhang 0001 |
ICIG (1) | 2 |
| 2023 | Aggregating transformers and CNNs for salient object detection in optical remote sensing images
Liuxin Bao, Xiaofei Zhou 0003, Bolun Zheng, Haibing Yin, Zunjie Zhu, Jiyong Zhang 0001, Chenggang Yan 0001 |
Neurocomputing | 1 |