VLDB 2026 Research / reviewers in the wild / expert
Baoyang Mu
dblp:349/9531
· DBLP profile ↗
14ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0003-2898-0461ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BigCounter: A Bidirectional-Guided Network With Scene-Semantics-Driven Fusion for RGB-Thermal Crowd CountingabstractAccurate crowd counting has become increasingly essential for public safety management and Internet of Video Things (IoVT) applications, driven by rapid population growth and urbanization. However, RGB-Thermal (RGB-T) crowd counting remains challenging due to poor recognition of small targets and degraded performance under extreme conditions such as low-light environments. To address these issues, we propose a bidirectional-guided network with scene-semantics driven fusion for RGB-T crowd counting (BigCounter) that enhances robustness and generalization in complex scenes. BigCounter comprises of three parallel branches: a primary branch, a dynamic illumination auxiliary enhancement branch (DIAEB), and a high-resolution auxiliary enhancement branch (HAEB), which respectively improve robustness under illumination variations and accuracy for small target detection. Moreover, a cross-layer scene-driven fusion module (CLSFM) and a cross-modal semantic-driven fusion module (CMSFM) are designed to strengthen structural consistency and explore semantic complementarity between modalities. Through multi-branch collaboration and semantic-aware fusion, BigCounter significantly enhances feature representation. Extensive experiments on two benchmark RGB-T datasets demonstrate that BigCounter achieves superior accuracy and generalization compared with state-of-the-art methods. Xiaomin Fan, Feng Shao 0001, Baoyang Mu, Xiongli Chai, Zhongjie Zhu, Zhiyi Mo |
IEEE Internet Things J. | 3 |
| 2026 | Edge-Embedded Bidirectional Interactive Network for Lightweight Surface Defect DetectionabstractConvolutional neural networks (CNN)-based surface defect detection methods have achieved remarkable results. However, most existing approaches face two key challenges: first, they typically require substantial computational resources to learn rich features, making it difficult to balance performance and computational cost; second, when dealing with complex backgrounds and various defect shapes, they often lose edge details. To address these challenges, we propose an Edge-embedded Bidirectional Interactive Network (EBINet) for lightweight surface defect detection. Specifically, we introduce an edge generation module (EGM), which enhances feature details by interacting with low-level and high-level features, providing high-quality edge details for defect regions. In addition, we design a bidirectional interactive decoder (BID) that consists of a self-recognition module (SRM), cross-attention fusion module (CAFM), and multiscale edge embedded module (MEEM). This decoder gradually integrates features from different stages and thus can effectively capture inter-layer feature relationships to generate high-quality saliency maps. Extensive experiments conducted on four public defect datasets demonstrate that the lightweight EBINet offers strong competitiveness and superior performance, requiring only 3.57M parameters and 2.30G FLOPs for a 256×256 input image. Feng Shao 0001, Dongze Jin, Baoyang Mu, Hangwei Chen |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2026 | USformer: A U-Shaped Structure Transformer for RGB-Thermal Semantic Segmentation and Traffic Scene UnderstandingabstractRecent advancements in multimodal approaches, particularly RGB-thermal (RGB-T) segmentation, have significantly promote the development of Intelligent Transportation Systems (ITS). However, existing methods still encounter challenges related to modality discrepancy and the effective integration of multi-scale features. To address these issues, we propose the U-shaped Structure Transformer (USformer) for RGB-T semantic segmentation. We improve the feature flow of existing methods by designing a novel U-shaped encoding network that integrates inter-layer fusion and cross-modal fusion. Specifically, our method introduces an inter-layer interaction mechanism that facilitates the iterative fusion of high-level semantic and low-level detail features. For each layer, our fusion process is divided into two stages: the Cross-Modal and -Scale Auxiliary (CMSA) module enforces distribution alignment across modalities and scales, while the Cross-Attention Feature Merger (CAFM) allows each modality to refine its own feature selection by employing a multi-head cross-attention mechanism. These modules effectively adapt and integrate well-established attention designs into our U-shaped encoding architecture, thereby achieving efficient multi-modal feature alignment and fusion. Finally, we utilize the Mask2Former decoder to aggregate the fused features from multiple layers and improve the segmentation across various object sizes and complex scenes. Extensive experiments on four RGB-T datasets demonstrate that our proposed USformer achieves state-of-the-art performance. Feng Shao 0001, Baoyang Mu, Xiongli Chai, Qiuping Jiang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Rethinking Lightweight RGB-Thermal Salient Object Detection With Local and Global Perception NetworkabstractRGB–thermal salient object detection (RGB-T SOD) aims to segment the most intriguing parts through RGB and thermal images. However, the high computational costs and large model sizes of existing methods prevent its deployment in edge computing platforms. To cope with it, a new lightweight RGB-T SOD method is proposed, namely, local and global perception network (LGPNet), which relieves the pressure of data transmission and processing and improves robustness in complicated surroundings. Specifically, high-level lightweight fusion (HLF) blocks and low-level lightweight fusion (LLF) blocks are designed to reduce the difference between two modalities and perform multimodal feature fusion. Unlike existing convolutional neural network-based lightweight SOD methods, HLF and LLF blocks combine the spatial inductive bias of convolutional neural networks with the global perception of Transformer, which is able to perform feature extraction and fusion with fewer parameters and larger receptive field. Experimental results show that our method is able to compete with the state-of-the-art RGB-T SOD methods while having only 7.35 learnable parameters[M], 6.40G floating-point operations and real-time speed (33 frames per second (FPS) on PyTorch framework and 224 FPS on TensorRT framework). Dongze Jin, Feng Shao 0001, Zhengxuan Xie, Baoyang Mu, Hangwei Chen |
IEEE Internet Things J. | 4 |
| 2025 | RGBT-Booster: Detail-Boosted Fusion Network for RGB-Thermal Crowd Counting With Local Contrastive LearningabstractWith the swift development of the Internet of Video Things (IOVT), crowd counting has demerged as an indispensable technology in the domains of intelligent transportation and video surveillance. However, due to the insufficient extraction of detail head information and the limited ability to reduce the multimodality differences, the existing methods still have large errors in accurate RGB-thermal (RGB-T) crowd counting. To this end, we propose a novel RGB-T crowd counting network, i.e., RGBT-Booster, to effectively deal with the aforementioned challenges. In RGBT-Booster, by introducing additional detail auxiliary branches for RGB and thermal infrared images and the proposed enhanced detail fusion module (EDFM), we can obtain richer low-level head detail features. In addition, we also propose a local contrastive learning (LCL) to further reduce the multimodality differences for accurate crowd counting. Experimental results on two public RGB-T crowd counting datasets (i.e., RGBT crowd counting (RGBT-CC) and DroneRGBT) and one RGB-Depth (RGB-D) crowd counting dataset (i.e., ShanghaiTechRGBD) show that the proposed RGBT-Booster achieves effective and superior counting performance, compared with previous methods. The source code and datasets used in the experiments will be released athttps://github.com/QSBAOYANGMU/RGBT-Booster. Baoyang Mu, Feng Shao 0001, Zhengxuan Xie, Long Xu 0001, Qiuping Jiang |
IEEE Internet Things J. | 1 |
| 2025 | Art Comes From Life: Artistic Image Aesthetics Assessment via Attribute Knowledge AmalgamationabstractAssessing the aesthetic quality and visual appeal of artworks has become one of the hotspots in current research. The existing artistic image aesthetics assessment (AIAA) methods directly learn aesthetics from images, while ignoring the impact of variations in visual attributes on human aesthetic perception, which hampers the further development of AIAA. To address this issue, this paper presents a new AIAA method based on attribute knowledge amalgamation, named AKA-Net. Specifically, we initially learn common attribute aesthetic rules (e.g., composition and color) through pre-training on natural aesthetic images. Then, we devise a multi-model amalgamation strategy based on contrastive learning to transfer different types of prior attribute knowledge into a single target model, enabling flexible and efficient aesthetic prediction. Finally, an attribute-aware feature enhancement module (AFEM) is introduced to better establish the relationship between aesthetic quality and attribute knowledge. Experimental results on three public benchmark AIAA databases demonstrate that the proposed AKA-Net outperforms the state-of-the-art AIAA metrics. Hangwei Chen, Feng Shao 0001, Xiongli Chai, Baoyang Mu, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | GADFNet: Geometric Priors Assisted Dual-Projection Fusion Network for Monocular Panoramic Depth EstimationabstractPanoramic depth estimation is crucial for acquiring comprehensive 3D environmental perception information, serving as a foundational basis for numerous panoramic vision tasks. The key challenge in panoramic depth estimation is how to address various distortions in 360° omnidirectional images. Most panoramic images are displayed as 2D equirectangular projections, which exhibit significant distortion, particularly with the severe fisheye effect near the equatorial regions. Traditional depth estimation methods for perspective images are unsuitable for such projections. On the other hand, cubemap projection consists of six distortion-free perspective images, allowing the use of existing depth estimation methods. However, the boundaries between faces of a cubemap projection introduce discontinuities, causing a loss of global information when using cube maps alone. In this work, we propose an innovative geometric priors assisted dual-projection fusion network (GADFNet) that leverages geometric priors of panoramic images and the strengths of both projection types to enhance the accuracy of panoramic depth estimation. Specifically, to better focus the network on key areas, we introduce a distortion perception module (DPM) and incorporate geometric information into the loss function. To more effectively extract global information from the equirectangular projection branch, we propose a scene understanding module (SUM), which captures features from different dimensions. Additionally, to achieve effective fusion of the two projections, we design a dual projection adaptive fusion module (DPAFM) to dynamically adjust the weights of the two branches during fusion. Extensive experiments conducted on four public datasets (including both virtual and real-world scenarios) demonstrate that our proposed GADFNet outperforms existing methods, achieving superior performance. Chengchao Huang, Feng Shao 0001, Hangwei Chen, Baoyang Mu, Long Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | A Mutual Head Knowledge Distillation Framework for Lightweight RGB-T Crowd CountingabstractAs an important technology in the fields of intelligent transportation and public safety, crowd counting that can obtain pedestrian flow information has attracted extensive attention from academic and industrial communities. However, existing RGB-T crowd counting methods cannot effectively balance the counting accuracy and computational complexity in practical applications. For this, we propose a Mutual Head Knowledge Distillation Framework (MHKDF) to obtain a lightweight RGB-T crowd counting network for efficient and accurate pedestrian number estimation. Specifically, to avoid the influence of parameter and structure differences between teacher and student networks on the distillation effect, we propose a Cooperative Mutual Knowledge Distillation (CMKD) strategy to comprehensively and dynamically transfer the crowd analysis ability of the complex teacher model (MHKDF-T) to the lightweight student model (MHKDF-S). In addition, the upper bound of the performance of the student network depends on the teacher model with high accuracy. Therefore, to take advantage of the complementary advantages of frequency domain and spatial domain feature fusion, we propose a Multi-Modal Spatial-Frequency Hybrid Fusion Module (MSFHFM) to futher improve counting accuracy of MHKDF-T. Comprehensive experiments on two RGB-T crowd counting datasets demonstrate that our MHKDF-S achieves competitive performance with only 5.68 FLOPs and 4.89M parameters. Our code will be released at https://github.com/BaoYangCC/MHKDF. Baoyang Mu, Feng Shao 0001, Hangwei Chen, Xuejin Wang, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | MISF-Net: Modality-Invariant and -Specific Fusion Network for RGB-T Crowd CountingabstractTo accurately perform crowd counting, utilizing the complementary relationship between RGB and thermal images to analyze the crowd has become the focus of current research. Due to different imaging principles, multi-modal images often contain different contents, which are their modality-specific information. For example, RGB images contain more texture and color details, while thermal images contain thermal radiation information. Meanwhile, they also describe the same target content, e.g., crowds, which are modality-invariant. However, existing methods only design different modules to directly fuse RGB and thermal image features, which did not fully consider the above facts. In this paper, by analyzing the similarities and differences between multi-modal images, we propose a Modality-Invariant and -Specific Fusion Network (MISF-Net) for RGB-T Crowd Counting. Specifically, we design a modality decomposition and fusion module (MDFM), which decomposes RGB and thermal image features into modality-invariant and -specific features by using the similarity and difference supervision between multi-modal features. Besides, reconstruction supervision is also used to prevent network learning from generating bias. After that, different fusion strategies are applied to the invariant and specific features, respectively. In addition, to adapt to the variations in size of different pedestrians, we design a modality-invariant fusion module (MIFM). Finally, after the fusion decoder, MISF-Net can obtain a more accurate crowd density map. Comprehensive experiments on the RGB-T crowd counting dataset show that our MISF-Net can achieve competitive performance. Baoyang Mu, Feng Shao 0001, Zhengxuan Xie, Hangwei Chen, Zhongjie Zhu, Qiuping Jiang |
IEEE Trans. Multim. | 1 |
| 2024 | CAFCNet: Cross-modality asymmetric feature complement network for RGB-T salient object detection
Dongze Jin, Feng Shao 0001, Zhengxuan Xie, Baoyang Mu, Hangwei Chen, Qiuping Jiang |
Expert Syst. Appl. | 4 |
| 2024 | Hallucinated-PQA: No reference point cloud quality assessment via injecting pseudo-reference features
Baoyang Mu, Feng Shao 0001, Hangwei Chen, Qiuping Jiang, Long Xu 0001, Yo-Sung Ho |
Expert Syst. Appl. | 1 |
| 2024 | Visual Prompt Multibranch Fusion Network for RGB-Thermal Crowd CountingabstractAs population growth and urbanization continue, accurate crowd counting is increasingly important for public safety management and the Internet of Video Things (IOVT). However, RGB and thermal infrared (RGB-T) crowd counting still faces challenges in improving feature extraction capability for RGB streams and reducing multimodality differences. For this, we propose a visual prompt multibranch fusion network (VPMFNet) to tackle the above challenges. Specifically, to improve the ability of crowd analysis of the RGB stream in RGB-T crowd counting, through designing the prompt enhancement module, we take the prior features of head perception in the crowd as visual prompt cues to embed into the RGB stream. In terms of RGB and thermal image feature fusion, we fully reduce the modality differences from the perspectives of local fusion, global fusion, and multireceptive field fusion to accurately estimate the pedestrian number. Various experiments on two RGB-T crowd counting data sets demonstrate that our VPMFNet achieves a smaller estimation error in the number of pedestrians. Besides, our VPMFNet outperforms existing methods (i.e., multicolumn convolutional neural network, BL, SANet, UCNet, HDFNet, BBSNet, BL+IDAM, BL+CSCA, dual-branch enhanced feature fusion network, and GETANet) on the RGB-D data set. Our code will be released athttps://github.com/QSBAOYANGMU/VPMFNet. Baoyang Mu, Feng Shao 0001, Zhengxuan Xie, Hangwei Chen, Qiuping Jiang, Yo-Sung Ho |
IEEE Internet Things J. | 1 |
| 2024 | Plain-PCQA: No-Reference Point Cloud Quality Assessment by Analysis of Plain Visual and Geometrical ComponentsabstractIn reviewing the research progress in Point Cloud Quality Assessment (PCQA), two main pathways have emerged, i.e., 2D projections and 3D point descriptors. The former primarily focuses on visual information, while the latter concentrates on crucial geometrical information in three-dimensional space. However, the current studies lack a thorough investigation of the impact of visual components and seldom pay special attention to plane-point fusion strategies. To comprehensively represent features and effectively tackle various types of impairments, we propose an end-to-end learning paradigm, only considering plain visual and geometrical factors called Plain-PCQA, for quantitatively evaluating objective metrics of 3D dense point clouds associated with human perception. Firstly, we explore a sophisticated preprocessing technique. The entire point clouds are packaged into six projections by moving virtual cameras, which can conveniently increase the visual samples during the training stage. Given the high resolution of the projected image, we have opted for a relatively lightweight network, namely ResNet-18, as the backbone to enable higher resolution input data. Five cropped patches from the projected image are collectively fed into this network. In light of the presence of some invalid information in the projections, a mask weight is devised to calculate the significance of each patch based on its effective informational content. Secondly, dual neural networks, comprising of a No-Reference (NR) branch and a Degraded-Reference (DR) branch, are designed with fundamental visual components to provide quantitative quality metrics. Specifically, the NR branch utilizes the feature output of each block in the Vision Transformer (ViT) model to obtain long-range low-level and high-level visual NR quality. The DR branch employs KLT (Karhunen-Loève Transform) to acquire the principal component information of an image as the macro-structural image, and then feeds the difference between input images and macro-structural images into a network for DR quality extraction. Thirdly, a Plane-Point Interaction Transformer (P2IT) is presented by incorporating texture and semantic features in 2D projections and geometrical features in 3D spaces to characterize the complete features with a connected 2D-3D feature representation. With these elaborately designed deep features, the proposed model can achieve competitive performances relying solely on plain visual and geometrical components. The experimental results demonstrate the potential of the proposed approach in multiple representative databases, which surpasses existing state-of-the-art methods significantly. Xiongli Chai, Feng Shao 0001, Baoyang Mu, Hangwei Chen, Qiuping Jiang, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Dynamic Weighted Fusion and Progressive Refinement Network for Visible-Depth-Thermal Salient Object DetectionabstractThe introduction of depth/thermal modality has significantly enhanced the performance of dual-modal salient object detection (SOD) methods. However, depth maps and thermal images are prone to environmental interference, making them insufficient for providing salient information. To address this challenge, triple-modal SOD methods have been proposed. However, these methods often overlook the detrimental effects of defective modalities during fusion, leading to subpar performance. To tackle this issue, we present a novel dynamic weighted fusion and progressive refinement network (DWFPRNet) for Visible-Depth-Thermal (V-D-T) SOD. Specifically, we first use the dual-modal fusion module (DFM) to fuse dual modalities, thereby obtaining fused features. Subsequently, the modality selective fusion module (MSFM) mines complementary information between fused features, considering both fusion features and the quality of feature maps, to achieve weighted fusion. Finally, we design a progressive refinement decoder (PRD) to realize interaction and multi-scale learning among different scale features and generate high-quality saliency maps. Extensive experiments conducted on the VDT-2048 public dataset demonstrate that our method outperforms existing state-of-the-art multi-modal methods. Feng Shao 0001, Baoyang Mu, Hangwei Chen, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |