VLDB 2026 Research / reviewers in the wild / expert
Xiaogang Song 0001
dblp:42/8957-1
· DBLP profile ↗
18ranked-venue papers
17as first author
18since 2021 · last 2026
0000-0001-9841-9624ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 9 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Camouflaged object detection via deep assistance sparsely weighted attention
Xiaogang Song 0001, Peirui Li, Haoyu Yuan, Xinhong Hei 0001 |
Comput. Vis. Image Underst. | 1 |
| 2026 | Depth-Guided Magnitude Spectrum and Local Reconstruction Network for Camouflaged Object Detection
Xiaogang Song 0001, Zixin Yue, Xinhong Hei 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | TGNet: Texture-enhanced guidance network for RGB-D salient object detection
Xiaogang Song 0001, Xinhong Hei 0001 |
Expert Syst. Appl. | 1 |
| 2026 | BIP-CENet: A Bilateral Prior-Collaborative Enhancement Network with dual-domain priors for low-light image enhancement
Xinhong Hei 0001, Xiaogang Song 0001, Zetian Zhang, Haiyan Tu, Yuping Tan, Xiujuan Zheng, Anlin Zhang |
Knowl. Based Syst. | 4 |
| 2026 | Depth correction and edge guidance network for RGB-D salient object detection
Xiaogang Song 0001, Bingxing Wei, Xinhong Hei 0001 |
Pattern Recognit. | 1 |
| 2026 | Multi-Clue Sliding Window Attention for Camouflaged Object DetectionabstractThe aim of camouflaged object detection (COD) is to discern concealed objects within the background. Due to issues such as high similarity to the surrounding environment, small size, occlusions, COD is considered a highly challenging task. In this paper, we propose a novel COD framework, named multi-clue sliding window attention network (MCSWA-Net), stressing in utilizing prior knowledge at different semantic levels to guide the detection of camouflaged objects via multi-scale sliding window attention (MSWA). To this end, we first devise the dynamic local detail capture (DLC) module and the global interactive decoder (GID) module to generate both local and global guidance clues. Particularly, each block of the DLC module produces local prior clue by processing corresponding image features at each stage from the encoder. And the GID module fuses all adjacent encoder features, generates global prior clue by combining fusion features of multi-semantic levels. Further, to make full use of prior clues guiding the detection of camouflaged objects at multi-semantic levels, we design the multi-scale guidance attention fusion (MAF) module and use two prior clues to refine the image features via the group fusion and the MSWA separately. Experiments conducted on four COD benchmark datasets, and results demonstrate that our MCSWA-Net is superior to state-of-the-art (SOTA) COD methods. In addition, we explore the detection capabilities of our MCSWA-Net for the downstream vision tasks related to COD, such as polyp segmentation, COVID-19 lung infection segmentation, and industrial defect detection. Experimental results show the proposed method has high degree of generality. Xiaogang Song 0001, Haoyu Yuan, Xinhong Hei 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | A method for absolute pose regression based on cascaded attention modules
Xiaogang Song 0001, Weixuan Guo, Xinhong Hei 0001 |
Comput. Vis. Image Underst. | 1 |
| 2025 | Three-dimensional human pose estimation based on multi-scale spatial-temporal transformer
Xiaogang Song 0001, Yongxin Cui, Jichen Chen, Xinhong Hei 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Spatial and channel enhanced self-attention network for efficient single image super-resolution
Xiaogang Song 0001, Yuping Tan, Xinchao Pang, Lei Zhang 0081, Xinhong Hei 0001 |
Neurocomputing | 1 |
| 2025 | GDVIFNet: A generated depth and visible image fusion network with edge feature guidance for salient object detection
Xiaogang Song 0001, Yuping Tan, Xiaochang Li, Xinhong Hei 0001 |
Neural Networks | 1 |
| 2025 | Self-Supervised Monocular Depth Estimation With Progressive Enhancement of Local-to-Global Visual PerceptionabstractSelf-supervised monocular depth estimation trains by utilizing the structure of the data itself without relying on ground-truth depth labels, gaining widespread attention in fields such as autonomous driving. However, many existing methods adopt the popular encoder-decoder structure, but this has deficiencies in refining both local and global visual clues, limiting its performance in detail recovery and spatial modeling. In this paper, we propose LGEDepth, a novel method to progressively enhance local to global visual perception. In LGEDepth, we design two crucial components to comprehensively refine local and global information after the encoder-decoder. The first is the patch-wise refinement tokenizer (PWRT), which fully refines the detailed information in local regions based on a local attention strategy with low performance overhead, effectively enhancing the local visual perception of the model. The second is the hierarchical interaction module (HIM), which better aggregates multi-scale contextual information through a cross-level interaction manner while determining the optimal depth bins that adapt to the depth distribution characteristics of the scene, effectively enhancing the global visual perception of the model. The experimental results on the KITTI, Cityscapes, Make3D and RUGD datasets demonstrate that LGEDepth achieves state-of-the-art performance and exhibits strong generalization ability, outperforming existing competitors. Xiaogang Song 0001, Bingxing Wei, Xinhong Hei 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | BTDGNet: A Dual-Guided Camouflaged Object Detection Network Leveraging Boundary and Texture InformationabstractCamouflaged object detection aims to identify objects that blend seamlessly with their background, posing a greater challenge compared to general object detection tasks. Due to its ability to recognize camouflaged objects, such detection models hold significant practical value across various fields. To accurately identify camouflaged targets in various complex environments, we designed a dual-guided camouflaged object detection network based on boundary and texture information(BTDGNet). The process consists of two main stages. The first stage is the localization stage, which leverages a convolutional neural network (CNN) to capture boundary and texture information of objects. These features are then fused to achieve coarse localization of the camouflaged objects. In the second stage, the recognition stage, we employ a Transformer to extract global information from the image, enhancing the differentiation between foreground and background. An interactive fusion module is designed to fully exploit and integrate both global and local features, producing precise prediction images. By leveraging boundary and texture information, the model's adaptability to different camouflaged objects is improved. The integration of local and global features enhances the model's detection accuracy from various perspectives, ultimately building a camouflaged object detection model suitable for a wide range of complex scenarios. The proposed method was extensively compared with other state-of-the-art methods across four public datasets, and the results demonstrated superior performance. Furthermore, benefiting from our dual-guidance strategy that leverages both texture and boundary information, our model demonstrates robust performance. We conducted tests on detection tasks across four different domains, and the results confirm that our model can accurately segment camouflaged objects in complex scenes. Xiaogang Song 0001, Xiaochang Li, Xinhong Hei 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | TransBoNet: Learning camera localization with Transformer Bottleneck and Attention
Xiaogang Song 0001, Hongjuan Li, Li Liang 0008, Weiwei Shi 0003, Guo Xie, Xinhong Hei 0001 |
Pattern Recognit. | 1 |
| 2024 | Salient Object Detection With Dual-Branch Stepwise Feature Fusion and Edge RefinementabstractIn recent years, Transformers have been gradually applied in salient object detection tasks with good results. However, the Transformer’s global modeling capabilities can lead to the loss of local details that are important in salient object detection tasks. A feature extraction backbone based on a convolutional neural network (CNN) is good at extracting local detail features due to the gradual expansion of the receptive field but is limited by the size of the receptive field, resulting in an insufficient ability to extract global semantic features. Therefore, this paper combines the Transformer with a CNN and presents a dual-branch encoder to ensure that the features extracted contain rich global semantic information as well as local detail features. In addition, due to the different features extracted by the Transformer and CNN, noise may be introduced in the fusion of the two features, so different features need to be processed correspondingly during fusion. The fusion enhancement module (FEM) we propose fuses the features of the two branches step by step. A hybrid attention mechanism is used to carry out weighted fusion of different features. This progressive approach minimizes the differences between the features of the two branches so that the merged features retain the semantic and detail features extracted by the two branches to the greatest extent. Considering the loss of detailed information caused by repeated downsampling, we propose an edge refinement module (ERM) to address the need for accurate outline prediction. This module leverages salient features to obtain edge features and gradually refines the prediction results by incorporating these edge features. It makes full use of the connection between salient features and edge features and does not introduce additional edges to extract branches. Extensive experimental evaluations conducted on five benchmark tests demonstrate the superior performance of our method compared to other existing approaches. Code can be found athttps://github.com/gfq1605694825/DSRNet-main. Xiaogang Song 0001, Fuqiang Guo, Lei Zhang 0081, Xinhong Hei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | A Universal Multi-View Guided Network for Salient Object and Camouflaged Object DetectionabstractSalient object detection and camouflaged object detection have attracted increasing attention due to their significant practical applications. While these two domains share similarities in recognition methods and object characteristics, they also exhibit distinctions. In this paper, we propose a novel multi-view guided network for camouflaged and salient object detection, utilizing the Transformer as the backbone network for feature extraction. Capitalizing on shared characteristics, we introduce a CNN-based multi-view encoder and a multi-view fusion module, enhancing the acquisition of multi-perspective information while minimizing the increase in computational cost. Moreover, recognizing domain differences, we incorporate an attention exploration module, seamlessly integrating multi-view features with globally extracted features from the backbone network. This integration involves simultaneous exploration from both positional and color perspectives, unearthing valuable information to identify salient and camouflaged objects. Our approach maximizes shared characteristics between the two tasks while effectively addressing their differences, leading to precise object identification—be it for camouflaged or salient objects. Extensive experiments on nine challenging benchmark datasets demonstrate the superior performance of our method across four widely used evaluation metrics, outperforming 34 state-of-the-art methods. Furthermore, we applied our method to other visually-related tasks, such as polyp segmentation and defect detection. The results further demonstrate the versatility of our model. The source code and results of our method are available athttps://github.com/1900zpf/MVGNet. Xiaogang Song 0001, Xinhong Hei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Unsupervised Monocular Estimation of Depth and Visual Odometry Using Attention and Depth-Pose Consistency LossabstractRecent studies have shown that joint depth and pose estimation using convolutional neural networks (CNNs) can learn unlabelled monocular frames. However, three problems remain: 1) CNNs can only extract local features due to the limited receptive field, 2) scale ambiguity is inherent in the monocular task, and 3) illness regions violate the photometric consistency assumption and produce large errors. We propose a novel framework, ADPDepth, with corresponding effective strategies to ameliorate the above problems. First, a PCAtt module is designed to capture the correlation between channels and efficiently extract multiscale spatial information using a multibranch parallel strategy. Second, depth-pose consistency loss is proposed based on the geometric consistency in depth and pose to constrain the scale between samples, eliminate scale ambiguity and obtain a globally consistent scale. To further improve performance, a cover mask is derived from depth-pose consistency for filtering dynamic objects and outliers to reduce the adverse effects of these illness regions. Extensive experiments are conducted on the KITTI, NYU-Depth and Make3D datasets. Based on public benchmarks, the experimental results confirm that the proposed ADPDepth framework achieves state-of-the-art performance. The effectiveness of each strategy is also verified in subsequent ablation experiments. Xiaogang Song 0001, Haoyue Hu, Li Liang 0008, Weiwei Shi 0003, Guo Xie, Xinhong Hei 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Local motion feature extraction and spatiotemporal attention mechanism for action recognition
Xiaogang Song 0001, Li Liang 0008, Xinhong Hei 0001 |
Vis. Comput. | 1 |
| 2023 | Image super-resolution with multi-scale fractal residual attention network
Xiaogang Song 0001, Wanbo Liu, Li Liang 0008, Weiwei Shi 0003, Guo Xie, Xinhong Hei 0001 |
Comput. Graph. | 1 |