EDBT 2026 Demo / reviewers in the wild / expert
Wujie Zhou
dblp:117/5608
· DBLP profile ↗
138ranked-venue papers
67as first author
121since 2021 · last 2027
0000-0002-3055-2493ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 64 · 31 first-author · 53 since 2021Applied, interdisciplinary, general and emerging computing · 36 · 21 first-author · 36 since 2021Artificial intelligence and machine learning · 29 · 8 first-author · 25 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | CDV-PCQA: Content-distortion-guided dynamic viewpoint quality assessment for 3D point clouds
Qihao Liang, Li Li 0014, Ting Luo 0001, Gangyi Jiang, Wujie Zhou, Linwei Zhu, Zhouyan He |
Expert Syst. Appl. | 5 |
| 2026 | MLANet: Multilevel aggregation network for binocular eye-fixation prediction
Wujie Zhou, Jiabao Ma, Yulai Zhang, Lu Yu 0003, Weijia Gao, Ting Luo 0001 |
Signal Process. Image Commun. | 1 |
| 2026 | Multi-stream interaction network with cross-modal contrast distillation for co-salient object detection
Wujie Zhou, Bingying Wang, Xiena Dong, Caie Xu, Fangfang Qiang |
Signal Process. Image Commun. | 1 |
| 2026 | Hierarchical Multi-Modal Enhancement for Robust Transmission Line Detection
Shengdong Zhang, Xiaoqin Zhang 0002, Shaohua Wan 0001, Yujing M. Jiang, Wujie Zhou, LinLin Shen, Wenqi Ren |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2026 | Lightweight Scope Integration Network for Rail Surface Defect DetectionabstractRail-surface defect detection (RSDD) is a key technology for ensuring the safety and efficiency of railroad transportation. Existing models enhance the robustness in complex scenarios by using complementary information from visible light (RGB) image and depth data. However, incorporating depth data increases the computational complexity, which consumes more resources, increases the risk of overfitting, and decreases the model efficiency. Optimizing and streamlining multimodal data processing remains challenging. To address this issue and enable fast and accurate RSDD, this study introduces a novel lightweight scope integration network (LSINet). This method efficiently and precisely fuses the features using a lightweight double-context-aware module and refines the image quality using an adaptive Markov field smoothing module in a hierarchical decoding process. A lightweight two-dimensional scanning method that captures long-range dependencies and improves computational efficiency. We integrate this approach with local range extremes to enhance the multimodal feature fusion. Furthermore, to refine the defect edges and improve the detection accuracy, we integrate the Markov random field module into the defect segmentation and optimization using its ability to model interpixel spatial correlation, particularly for small and ambiguous defects. In designing the model, we focused on optimizing the number of parameters (e.g., computational resources). Experimental evaluations using various rail surface images demonstrated that the proposed method enhanced the defect detection accuracy and recall while reaching fast detection speeds. According to extensive experiments conducted using the industrial RGB-D dataset (i.e., NEU RSDDS-AUG), LSINet outperformed 15 state-of-the-art methods with only 27.6 million parameters. Code and results are publicly available at https://github.com/Wuyue15/LSINet. Wujie Zhou, Fangfang Qiang, Weiqing Yan |
IEEE Trans. Big Data | 1 |
| 2026 | RCNet: Dual-Network Resonance Collaboration via Mutual Learning for RGB-D Road Defect Detection
Wujie Zhou, Zijun Ju, Runmin Cong, Weiqing Yan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Decouple-Then-Synergize: A Self-Paced Collaborative Learning Network for RGB-T Snowy Urban Scene ParsingabstractFusing RGB and thermal infrared images is essential for advancing urban scene analysis. However, both modalities exhibit severe performance degradation under snowy conditions. Although independent enhancement modules can partially mitigate this issue, stacking multiple modules with different functions increases model complexity and may cause intermodular interference. To address these limitations, we propose a "decouple-then-synergize" framework that decouples the task into frequency-oriented enhancement and spatial semantic fusion, implemented by FRENet (frequency restoration enhancement network) and SIFNet (spatial interactive fusion network), respectively. FRENet uses an asymmetric enhancement strategy that selectively sharpens RGB color gradients while amplifying faint thermal targets. It incorporates a precise spectral refinement module to restore high-frequency details. SIFNet introduces a Mamba zipper fusion module to achieve robust interaction of high-level semantics and performs a reconstruction task to implicitly integrate thermal features into the RGB stream. To ensure effective collaboration between the two networks, we design a self-paced curriculum that manages bidirectional knowledge exchange at both the sample and pixel levels. This approach enables the networks to evolve into their enhanced versions, namely FRENet-collaborative learning (CL) and SIFNet-CL. Extensive experiments on the SUS and PST900 datasets demonstrate that our framework outperforms state-of-the-art scene parsing methods. The code and associated results are available at https://github.com/Lyb-2001/SPCL. Wujie Zhou, Yiben Li, Qiuping Jiang, Runmin Cong, Weisi Lin |
IEEE Trans. Image Process. | 1 |
| 2026 | Turbidity-Similarity Decoupling: Feature-Consistent Mutual Learning for Underwater Salient Object DetectionabstractUnderwater salient object detection (USOD) faces two major challenges that hinder accurate detection: substantial image noise owing to water turbidity and low foreground-background contrast caused by high visual similarity. In this study, a dual-model architecture based on mutual learning is proposed to address these issues. First, DenoisedNet, which focuses on addressing water turbidity issues, is developed. Using a separation-denoising-enhancement processing framework, it suppresses noise while maintaining target feature integrity through domain separation and cleaning enhancement modules. Second, SearchNet is designed to address the foreground-background similarity issue. It achieves precise localization through pseudo-label generation and layer-by-layer search mechanisms. To enable both networks to address these challenges collaboratively, a feature-consistent mutual-learning strategy is proposed, which aligns encoded features and prediction results, via evaluation and cross modes, respectively. This strategy enables their respective strengths to be complemented and the challenges of USOD to be solved more comprehensively. Our DenoisedNet and SearchNet outperform the best existing methods on the USOD10K and USOD benchmarks, achieving MAE improvements of 4.52%/5.52% and 1.61%/8.94%, respectively. The source code is available at https://github.com/BeibeiIsFreshman/DSNet_CL. Wujie Zhou, Beibei Tang, Runmin Cong, Qiuping Jiang |
IEEE Trans. Image Process. | 1 |
| 2025 | Cross-modal State Space Modeling for Real-time RGB-thermal Wild Scene Semantic SegmentationabstractThe integration of RGB and thermal data can significantly improve semantic segmentation performance in wild environments for field robots. Nevertheless, multi-source data processing (e.g. Transformer-based approaches) imposes significant computational overhead, presenting challenges for resource-constrained systems. To resolve this critical limitation, we introduced CM-SSM, an efficient RGB-thermal semantic segmentation architecture leveraging a cross-modal state space modeling (SSM) approach. Our framework comprises two key components. First, we introduced a cross-modal 2D-selective-scan (CM-SS2D) module to establish SSM between RGB and thermal modalities, which constructs cross-modal visual sequences and derives hidden state representations of one modality from the other. Second, we developed a cross-modal state space association (CM-SSA) module that effectively integrates global associations from CM-SS2D with local spatial features extracted through convolutional operations. In contrast with Transformer-based approaches, CM-SSM achieves linear computational complexity with respect to image resolution. Experimental results show that CM-SSM achieves state-of-the-art performance on the CART dataset with fewer parameters and lower computational cost. Further experiments on the PST900 dataset demonstrate its generalizability. Codes are available at https://github.com/xiaodonguo/CMSSM. Zi'ang Lin, Luwen Hu, Tong Liu 0009, Wujie Zhou |
IROS | 6 |
| 2025 | Cross-attention fusion and edge-guided fully supervised contrastive learning network for rail surface defect detection
Jinxin Yang, Wujie Zhou |
Appl. Intell. | 2 |
| 2025 | Multilevel attention imitation knowledge distillation for RGB-thermal transmission line detection
Wujie Zhou, Tong Liu 0009 |
Expert Syst. Appl. | 2 |
| 2025 | IIBNet: Inter- and Intra-Information balance network with Self-Knowledge distillation and region attention for RGB-D indoor scene parsing
Fangfang Qiang, Wujie Zhou, Weiqing Yan, Lv Ye |
Expert Syst. Appl. | 3 |
| 2025 | Feature Contrast Difference and Enhanced Network for RGB-D Indoor Scene Classification in Internet of ThingsabstractThe era of smart connectivity spawned by the Internet of Things (IoT) has made the need to achieve environmental perception and understanding of different scenarios increasingly urgent. Among the many scenarios, indoor scene classification has attracted much attention because of its relevance to the daily lives of people, ranging from comfort regulation in living spaces to the optimal allocation of resources in offices, and a variety of approaches for this task have emerged. However, increasing accuracy remains a crucial objective due to the complexity and disorder of indoor scenes. Therefore, we propose a feature contrast difference and enhanced network for RGB-D indoor scene classification, FCDENet. First, the red, green, and blue and Depth images express different information. Therefore, we built a feature contrast difference module for the first two low-level features to extend the receptive fields of the different features, utilizing differential contrast to complement each other. Second, the high-level feature semantic information is abstract. Therefore, we introduced information cluster blocks, which are used to aggregate feature points with similar attributes into compact clusters after being parsed by an initial frequency transform, enabling instantiated representations of the semantic information. Finally, to further enhance the integrated features, we introduced a wavelet transform block in the cross-layer decoding process. In contrast to conventional decoding methods, we employed a wavelet transform for initial denoising cross-layer features and used multiple pooling structures to supplement local information, gradually weighting to achieve higher prediction accuracy. Extensive experiments on two typical indoor datasets, NYUDv2 and SUN RGB-D, show that our results exhibit excellent performance. In addition, to better demonstrate the reliability of the method, we conducted generalizability experiments on other datasets, and the proposed method provides a robust solution to the challenges of multiple scenarios in the era of IoT smart connectivity. The code is available athttps://github.com/XUEXIKUAIL/FCDENet. Wujie Zhou, Bitao Jian, Yuanyuan Liu 0004 |
IEEE Internet Things J. | 1 |
| 2025 | Hybrid Knowledge Distillation for RGB-T Crowd Density Estimation in Smart Surveillance SystemsabstractCrowd density estimation is a practical application task in which speed efficiency is as crucial as the accuracy of the results. Hence, we propose the hybrid knowledge distillation network (HKDNet) for RGB-thermal (RGB-T) crowd density estimation to address the limitations of computational cost and training time from the perspective of ensuring accuracy. We efficiently combine the advantages of traditional convolution and self-attention through a multimodal interactive transform. Subsequently, cross-graph convolution is extended to an interactive space for multimodal relationship reasoning. Finally, from the perspectives of channel and space, the crowd density map is obtained using channel synergy supplementation and spatial detail filling. In contrast to the computing resources required when using a heavyweight teacher network, the proposed HKDNet uses only approximately 6% of the parameters of the teacher network by abstracting the teacher network into a network hierarchy to generate a lightweight and efficient student network. Extensive experiments show that the proposed HKDNet performed well on two RGB-T crowd density estimation datasets. The code and models are available athttps://github.com/WBangG/HKDNet. Wujie Zhou, Weiqing Yan, Qiuping Jiang |
IEEE Internet Things J. | 1 |
| 2025 | Distilled wave-aware network: Efficient RGB-D co-salient object detection via wave learning and three-stage distillation
Zhangping Tu, Xiaohong Qian, Wujie Zhou |
Inf. Sci. | 3 |
| 2025 | FDFNet-S*: frequency domain fusion networks for RGB-D mirror segmentation by contrastive knowledge refinement
Caie Xu, Wujie Zhou |
Multim. Syst. | 3 |
| 2025 | Few-shot semantic segmentation in complex industrial components
Caie Xu, Jin Gan, Yu Wang 0246, Minglei Tu, Wujie Zhou |
Multim. Tools Appl. | 7 |
| 2025 | AESeg: Affinity-enhanced segmenter using feature class mapping knowledge distillation for efficient RGB-D semantic segmentation of indoor scenes
Wujie Zhou, Yuxiang Xiao, Fangfang Qiang, Xiena Dong, Caie Xu, Lu Yu 0003 |
Neural Networks | 1 |
| 2025 | Beyond boundaries: Hierarchical-contrast unsupervised temporal action localization with high-coupling feature learning
Yuanyuan Liu 0004, Leyuan Liu 0001, Wujie Zhou, Chang Tang |
Pattern Recognit. | 6 |
| 2025 | Multi-branch Space Sharing Feature Aggregation for contrastive multi-view clustering
Yuanyang Zhang, Weiqing Yan, Chang Tang, Wujie Zhou |
Pattern Recognit. | 4 |
| 2025 | SACNet: Saliency-Aided Aggregation Consensus Network for RGB-D Co-Salient Object DetectionabstractRed, green, and blue-depth (RGB-D) co-salient object detection identifies and segments co-appearing salient targets in various related images and corresponding depth maps. Existing methods directly blend depth features with RGB features. However, they often overlook that, in depth maps, the low contrast of neighboring objects affects salient regions that do not correspond to the saliency regions in the RGB image, leading to unsatisfactory results. We design a novel network called saliency-aided aggregation consensus network (SACNet) for RGB-D co-salient object detection. SACNet incorporates two key modules: the multimodal weight-sharing fusion and feature consensus aggregation module. The multimodal weight-sharing fusion effectively integrates RGB and depth modal information by calibrating depth features and sharing weights, which helps extract feature consensus. The feature consensus aggregation module clusters information in individual image saliency for inference, distinguishing between co-salient and non-co-salient objects, and aggregates consensus features across the image group. To enhance the discriminative power of the consensus representation, individual image saliency information is stored and updated in a queue. SACNet achieved competitive results on two challenging benchmark datasets. Zhangping Tu, Xiaohong Qian, Wujie Zhou |
IEEE Signal Process. Lett. | 3 |
| 2025 | Transformer-Prompted Network: Efficient Audio-Visual Segmentation via Transformer and Prompt LearningabstractAudio–visual segmentation (AVS) is a challenging task that focuses on segmenting sound-producing objects within video frames by leveraging audio signals. Existing convolutional neural networks (CNNs) and Transformer-based methods extract features separately from modality-specific encoders and then use fusion modules to integrate the visual and auditory features. We propose an effective Transformer-prompted network, TPNet, which utilizes prompt learning with a Transformer to guide the CNN in addressing AVS tasks. Specifically, during feature encoding, we incorporate a frequency-based prompt-supplement module to fine-tune and enhance the encoded features through frequency-domain methods. Furthermore, during audio–visual fusion, we integrate a self-supplementing cross-fusion module that uses self-attention, two-dimensional selective scanning, and cross-attention mechanisms to merge and enhance audio–visual features effectively. The prompt features undergo the same processing in cross-modal fusion, further refining the fused features to achieve more accurate segmentation results. Finally, we apply self-knowledge distillation to the network, further enhancing the model performance. Extensive experiments on the AVSBench dataset validate the effectiveness of TPNet. Xiaohong Qian, Wujie Zhou |
IEEE Signal Process. Lett. | 3 |
| 2025 | PFCNet: Enhancing Rail Surface Defect Detection With Pixel-Aware Frequency Conversion NetworksabstractApplying computer vision techniques to rail surface defect detection (RSDD) is crucial for preventing catastrophic accidents. However, challenges such as complex backgrounds and irregular defect shapes persist. Previous methods have focused on extracting salient object information from a pixel perspective, thereby neglecting valuable high- and low-frequency image information, which can better capture global structural information. In this study, we design a pixel-aware frequency conversion network (PFCNet) to explore RSDD from a frequency domain perspective. We use different attention mechanisms and frequency enhancement for high-level and shallow features to explore local details and global structures comprehensively. In addition, we design a dual-control reorganization module to refine the features across levels. We conducted extensive experiments on an industrial RGB-D dataset (NEU RSDDS-AUG), and PFCNet achieved superior performance. The code and results are publicly available at https://github.com/Wuyue15/PFCNet. Fangfang Qiang, Wujie Zhou, Weiqing Yan |
IEEE Signal Process. Lett. | 3 |
| 2025 | Efficient RGB-D Co-Salient Object Detection via Modality-Aware PromptingabstractRGB-D co-salient object detection (Co-SOD) aims to identify and segment co-occurring salient objects in a set of correlated images and depth maps. Most existing RGB-D Co-SOD methods fully fine-tune the dual-stream encoder-decoder architecture and fuse the RGB and depth features using a complex feature fusion strategy, which is expensive to train owing to the large number of parameters that need to be updated during the feature extraction and fusion process. In addition, current methods do not pay sufficient attention to differentiate co- salient information from non-co- salient information effectively. This interfering information affects the localization of co-salient targets. Therefore, this study proposes a simple and effective modality-aware prompting network (MAPNet) for efficient RGB-D Co-SOD. MAPNet mainly performs RGB-D Co-SOD through two approaches, namely modal fusion and consensus feature extraction, using a multimodal prompt generator (MPG) and consensus feature extraction module (CFEM), respectively. Specifically, the MPG module guides the depth features in the fine-tuned backbone network from the RGB features obtained in the frozen backbone network for fusion in hyperbolic spaces to generate multilevel modal cues that are subsequently injected into the fine-tuned backbone network for efficient modal fusion. The CFEM uses RGB features to generate an image salient prior, combines the salient prior with the highest level of fusion features to obtain the central point, and uses the salient features closer to the central point as the consensus features of the image group. In addition, contrast loss is introduced to separate the synergistic and non-synergistic salient features to obtain pure co-salient features. The trained MAPNet delivered state-of-the-art performance on three benchmark datasets (RGB-D CoSal1k, RGB-D CoSal150, and RGB-D CoSeg183), with the structure-measure improved by 2.1% on the RGB-D CoSeg183 dataset. The codes are available athttps://github.com/trumpetor/MAPNet. Note to Practitioners—This study presents a straightforward and effective modality-aware prompting network (MAPNet) designed for efficient RGB-D Co-SOD. Initially, the MPG module of MAPNet enables the effective integration of RGB and depth modalities through prompt learning. Subsequently, the CFEM employs the pixel group centroid proxy and top-k selection mechanism to extract high-level integrated features and salient prior consensus features, which serve as coordinated saliency features for image groups. Finally, the coordinated obtained salient and integrated features are input into the decoder to generate predictions. Zhangping Tu, Xiaohong Qian, Wujie Zhou |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Transmission Line Detection Through Auxiliary Feature Registration With Knowledge DistillationabstractThe periodic inspection of transmission lines ensures a stable power supply across various regions. Deep learning methods, particularly those based on multimodal approaches, have made significant advancements in this field. Current multimodal fusion algorithms do not prioritize the integration of supplementary information from other modal sources. Moreover, these methods frequently depend on substantial parameterization to achieve optimal performance, severely hampering deployment on mobile devices. To enhance the supportive role of the auxiliary modality and minimize model parameters, we proposed the knowledge distillation-based auxiliary feature registration network (AFRNet-S*) for RGB-T transmission line detection (TLD). The method incorporates an auxiliary feature registration module during fusion, utilizing unique complementarity between primary and auxiliary feature modalities for precise feature registration. For knowledge distillation, we devised response-mimicking distillation, semantic supplementary distillation, and cross-view feature polymerization distillation tailored for the student network (AFRNet-S, without knowledge distillation) to achieve model compression. AFRNet-S can learn comprehensive final spatial, precise semantic, and complementary feature integration information from AFRNet-T through these distillation methods. Extensive experiments conducted on an industrial TLD dataset demonstrate that the proposed AFRNet-S* (AFRNet-S with knowledge distillation) outperforms the existing state-of-the-art methods, thereby demonstrating its generality and effectiveness. Our AFRNet-S*, compared to AFRNet-T, has a reduced parameter count from 44.92M to 16.86M and a decrease in FLOPs from 21.13G to 6.41G. This further optimizes the deployment of our network on edge devices, such as unmanned aerial vehicles (drones), in the future. The code and results of our approach are available athttps://github.com/WangYuSenn/AFRNet. Note to Practitioners—This study introduces the AFRNet for TLD, utilizing knowledge distillation (KD) techniques. We achieved a better balance between model performance and parameter count by integrating the KD task into the TLD task, facilitating easier deployment on mobile devices. The proposed AFRNet-S* effectively balances model performance and parameter count, ensuring efficient performance of the TLD task. Addressing RGB-T TLD-specific challenges remains a crucial area for future research. For instance, in scenarios where a target occupies a small area within an image, removing excessive background noise is essential to enhance the efficiency of target localization. Furthermore, we aim to utilize AFRNet-S* on drones for transmission line inspection in the future to address practical application challenges. Wujie Zhou, Xiaohong Qian |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Depth Enhancement Mask Mapping Network With Multi-Teacher Distillation for RGB-D Scene ParsingabstractThe proposed method in the study, called DEMMNet-KD, aims to address the challenges of scene parsing in RGB-D indoor scenes. The study recognizes the importance of lightweight models for pixel-intensive tasks like scene parsing, and proposes a depth enhancement mask mapping network (DEMMNet) with multi-teacher knowledge distillation approach to achieve both reduced model size and high accuracy. DEMMNet introduces a new segmentation paradigm called class-position separate segmentation to reduce the difficulty of delineating and classifying the segmentation region. It also includes a depth spatial enhancement module based on the attention mechanism to effectively utilize depth information in a single-stream network design. To improve performance without increasing complexity, DEMMNet-KD utilizes Mix Transformer encoder (MiT) as a backbone and employs multi-teacher knowledge distillation. This allows semantic knowledge to be transferred from large teacher networks with different performances to the student network. The experiments conducted on NYUDv2 and SUN RGB-D datasets demonstrate that DEMMNet-KD achieves state-of-the-art performance with fewer additional learnable parameters and minimal computational effort. Utilizing MiT-B2 as the backbone network, the DEMMNet-KD method attained mean Intersection over Union (mIoU) scores of 54.71% and 51.03% on the two datasets, respectively, thereby achieving an optimal balance between accuracy and efficiency. Furthermore, the scalability of this method has been demonstrated across various domains and modalities. Overall, the proposed DEMMNet-KD approach provides an efficient solution for scene parsing in RGB-D indoor scenes, enhancing the capabilities of bionic binocular robots in perceiving their environments and serving human society effectively. The code is available at https://github.com/SHARKALAKALA/DEMMNet. Yuxiang Xiao, Jiajun Meng, Fangfang Qiang, Xiena Dong, Wujie Zhou |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | MVoxTi-DNeRF: Explicit Multi-Scale Voxel Interpolation and Temporal Encoding Network for Efficient Dynamic Neural Radiance FieldabstractNeural radiance fields have revolutionized the field of novel view synthesis, achieving remarkable results. However, traditional approaches based on implicit representations, particularly those built upon NeRF, suffer from slow rendering speeds due to the need for numerous MLP evaluations. Recently, there has been a promising shift towards explicit representations using voxel grids, which has significantly improved reconstruction times for static scenes. Nonetheless, extending these methods from static scenes to dynamic scenes is a non-trivial task as it requires accounting for the changing geometry and appearance of the scene over time. In this paper, we propose an efficient dynamic neural radiance field with multi-scale explicit voxel interpolation and temporal encoding. We leverage an explicit voxel structure to store the 3D dynamic features, while employing a lightweight MLP to estimate the displacement, thereby significantly enhancing the reconstruction speed. In our canonical module, we incorporated temporal information encoding in density estimation and color estimation to rectify the error estimation of the displacement in deformation module. In addition, a multi-scale voxel interpolation is designed to accommodate large-scale motions while meticulously capturing intricate details in small-scale motions in the density estimation module. In experiment, we evaluate our MVoxTi-DNeRF method on both synthetic and real scenes, where it achieves superior or comparable rendering quality compared to state of the art methods, while remaining computationally efficient (more than$60\times $faster than the original DNeRF). More experiment results and test code are available athttps://github.com/CHenYYff/MVoxTi-DNeRFNote to Practitioners—Neural Radiance Fields(NeRF) can create highly detailed and realistic 3D models from 2D images via a neural network. It is particularly popular for its capacity to generate novel views of a scene, enabling it to synthesize images from previously unobserved viewpoints. NeRF has found applications in a wide range of fields, including virtual reality, augmented reality, gaming, and so on. In this study, we introduces an efficient dynamic NeRF method. First, our network leverages an optimized explicit voxel grid to store 3D dynamic features and employs a lightweight MLP to decode these deformation features, significantly accelerating the training process. Second, to correct the error estimation related to deformation displacement, we introduce encoding of temporal information into density and color estimation in our canonical module, which fortifies the canonical field’s perception of temporal information. Third, we utilize multi-scale voxel interpolation to capture different-scale motion in the density estimation module, where minor motions are modeled using nearby voxels, while motion within a broader range is captured through more distant voxels. The consideration of multi-scale voxel features diminishes the detrimental impact caused by inaccurate displacement estimation. Weiqing Yan, Yanshun Chen, Wujie Zhou, Runmin Cong |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | MSNet: Multiple Strategy Network With Bidirectional Fusion for Detecting Salient Objects in RGB-D ImagesabstractVarious salient object detection (SOD) approaches have been developed to identify visually attractive objects in scenes captured in RGB-D (RGB and depth) images. High-level features often provide abstract semantics, and low-level features include more details such as textures and spatial structures. Hence, effectively fusing multimodal information from different levels has become a major area of development. We propose a multiple-strategy network (MSNet) with bidirectional fusion for RGB-D SOD that incorporates multilevel feature fusion and cross-modal aggregation into a multisupervised framework. We first use a multiple-strategy fusion module to transmit high-level semantic features along a top-down progressive pathway to generate a series of appearance features. Thereafter, a self-refinement module further refines and optimizes the saliency map. Furthermore, a depth optimization module strengthens depth information extraction, especially from low-quality depth maps. Extensive experimental results on seven benchmark datasets reveal the superiority and efficacy of the proposed MSNet, compared with state-of-the-art RGB-D SOD approaches.Note to Practitioners—This study introduces a RGB-D SOD network known as Multiple-Strategy Network (MSNet) with bidirectional fusion. Initially, we employ a multiple-strategy fusion module to transmit high-level semantic features in a top-down progressive manner, generating a sequence of appearance features. Subsequently, we apply a self-refinement module to further enhance and optimize the saliency map. Additionally, we incorporate a depth optimization module to improve the extraction of depth information, particularly from low-quality depth maps. Wujie Zhou, Weiwei Qiu |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | RDNet-KD: Recursive Encoder, Bimodal Screening Fusion, and Knowledge Distillation Network for Rail Defect DetectionabstractRail defect detection (RDD) plays a crucial role in ensuring rail transportation safety. Recently, bimodal algorithms have become mainstream; however, the asymmetry in the information of RGB and depth makes it difficult to find a suitable bimodal information fusion algorithm. In addition, it is difficult to deploy most of the existing methods on mobile devices. To solve these problems, we propose a recursive encoder and bimodal information screening fusion with a knowledge distillation network (RDNet-KD) for RDD. First, we propose the recursive encoder-based depth information augmentation (REDA) algorithm. It recursively learns to expand the channel depth information to alleviate the quality problem of depth information. Second, we propose a similarity-driven bimodal information screening fusion (SICF) module. This evaluates the complementarity of information from two modalities by computing the similarity of their hierarchical feature maps to screen useful information for fusion. Third, we introduce the global location and interrelation-based dual contextual knowledge distillation method to enhance the performance of the compact model. Therefore, it is possible to deploy the network on mobile devices. Based on the extensive experiments performed on the RGB-D rail defect dataset NEU RSDDS-AUG, we validate the competitiveness of our RDNet-KD, considering the prediction quality and operational efficiency relative to 12 state-of-the-art methods. The RDNet-KD code and results are available at https://github.com/legendfantasy/RDNet-KD.Note to Practitioners—This study introduces a recursive encoder and bimodal information screening fusion with a knowledge distillation network (RDNet-KD) for RDD in RGB-D images. Our method enhances depth information quality and effectively selects valuable information from both modalities using the similarity as a coefficient to evaluate the complementary capabilities of the modal information. Furthermore, to compress the model, we introduce knowledge distillation (KD) to balance the number of parameters and detection results and propose a novel KD method that transfers knowledge from the teacher network to the student network. Wujie Zhou, Jinxin Yang, Weiqing Yan, Meixin Fang |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Enhancing RGB-D Mirror Segmentation With a Neighborhood-Matching and Demand-Modal Adaptive Network Using Knowledge DistillationabstractRecent breakthroughs in computer vision have led to remarkable progress in the areas of autonomous vehicles and robotics. However, ordinary objects such as mirrors pose unique challenges to computer vision systems owing to occlusion, reflection, and distortion. Moreover, existing deep learning models suffer from issues such as excessive parameters and high computational complexity, making it challenging to implement numerous studies offline. To address these issues, we propose an innovative solution: a neighborhood-matching and demand-modal adaptive network using knowledge distillation (KD), called NDANet-S$^{\ast }$, specifically designed for red-green-blue depth mirror segmentation. NDANet-S$^{\ast }$operates by iteratively matching detailed and semantic difference between neighborhood features during the encoding phase. It then complements information across different modalities through demand-modal adaptation, enhancing heteromodal cross-complementation during the KD stage. In the decoding phase, semantic enhancement features and iterative encoding features are deeply integrated, forming a strong foundation for multistage progressive knowledge transfer in the KD process. Furthermore, we introduce a multistage teacher-assisted KD scheme, guided by sample complexity, to work synergistically with the mirror segmentation model. This innovative scheme includes a sample complexity rater, heterogeneous cross-complementarity, and hierarchical progressive knowledge transfer. Experimental evaluations on publicly available datasets indicate that NDANet-S$^{\ast }$significantly enhances segmentation accuracy while preserving a consistent number of parameters. Additionally, it achieves state-of-the-art performance in mirror segmentation. The source code for our model is publicly available and can be accessed at:https://github.com/2021nihao/NMDANet. Note to Practitioners—This study presents a neighborhood-matching and demand-modal adaptive network for RGB-D mirror segmentation, incorporating knowledge distillation (KD). The combination of the KD framework and segmentation network enables sample complexity discrimination, cross-modal distillation, and multilevel distillation (guided by the former) to achieve targeted distillation. The newly proposed NDANet-S$^{\ast }$surpasses current state-of-the-art methods, with a reduction of 93.2% in parameters and 87.65% in floating-point operations (FLOPs) compared to NDANet-T. Wujie Zhou, Yuanyuan Liu 0004, Ting Luo 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2025 | Asymmetric Dual-Encoder Network With Clustering and Mutual Contrast Loss for the Semantic Segmentation of Remote-Sensing ImagesabstractIn recent years, the semantic segmentation of multimodal remote-sensing images using convolutional methods has received significant attention. Owing to the localized nature of convolutional operations, existing methods use the attention method to obtain global relations. However, effectively complementing and eliminating the global and local relations between two modalities has become a major challenge. In addition, category imbalance often affects the model performance as a result of the high resolution of remote-sensing images. In this study, we propose an asymmetric dual encoder network with clustering mutual contrast loss. Specifically, we use a convolutional neural network and Transformer as the backbone to extract two modal features in parallel to learn the local and global information, respectively. Next, our proposed multimodal hierarchical interaction module and dynamic weight inspired block efficiently fuse the multimodal features to complement the local and global information. The fused features are fed into our proposed local context extraction module and global context extraction module. Furthermore, to address the challenge presented by feature class imbalances, we apply a clustering algorithm to classify each pixel, which is subsequently inter-supervised using the inter-contrast loss. Extensive experiments on benchmark datasets show that the proposed model is extremely effective in the semantic segmentation of remotely sensed images and that it outperforms current state-of-the-art networks both quantitatively and qualitatively. Our code and results can be found at https://github.com/LYZ00918/AMCNet. Wujie Zhou, Yangzhen Li, Yuanyuan Liu 0004 |
IEEE Trans. Big Data | 1 |
| 2025 | Differential Modal Multistage Adaptive Fusion Networks via Knowledge Distillation for RGB-D Mirror SegmentationabstractMirrors play a significant role in our daily lives and are ubiquitous. However, deep learning computer vision models find them challenging owing to the negative impact of reflected information on scene understanding. This study addresses two key challenges faced by multimodal models. First, the cross-modal variability of features at different stages is generally overlooked by contemporary backbone networks. Second, good performance has only been achieved at an unacceptable computational expense, owing to the numerous parameters used. To address the first challenge, we propose a differential-mode multistage adaptive fusion network (differential mode refers to images generated by different sensors that are differentiated to complement each other) that incorporates two-step fusion in the coding stage to account for the degrees of difference among the cross-modal features. In the first stage, wherein considerable differences in modal features exist, multi-angle fusion is performed. In the second stage, wherein the differences are smaller, a hierarchical adaptive fusion strategy is employed. Regarding the second challenge, we introduce a companion training framework for mirror segmentation that combines knowledge distillation and contrastive learning. Our proposed scheme achieves state-of-the-art performance on an available mirror segmentation dataset without requiring numerous parameters. Wujie Zhou, Weiwei Qiu |
IEEE Trans. Big Data | 1 |
| 2025 | Knowledge Distillation and Contrastive Learning for Detecting Visible-Infrared Transmission Lines Using Separated Stagger Registration NetworkabstractMultimodal transmission-line detection (TLD) and other vision-related tasks in smart grids have garnered increasing attention due to advances in deep-learning technologies and escalating need for reliable power supplies. However, current TLD methodologies encounter several limitations. First, complex weather conditions often introduce substantial background noise, resulting in inaccurate object detection. Second, high parameter counts of extant models impede their deployment in real-world applications. Third, insufficient data samples results in overfitting and instability. To address these challenges, we proposed a separated stagger registration network (SSRNet-S$^\ast)$, augmented with knowledge distillation (KD) and contrastive learning, specifically designed for RGB-T TLD. This method integrates a separated stagger registration mechanism into the fusion module to investigate relationships between cross-modal features. This approach enhances feature representation and effectively reduces background noise. Additionally, we devised a joint training framework incorporating KD and contrastive learning and proposed a hierarchical distillation strategy to compress the model while mitigating the impact of limited data samples. Complementary features were captured at various stages of SSRNet-S$^\ast$by employing three levels of distillation. Extensive experiments on a TLD dataset demonstrated that both SSRNet-T and SSRNet-S$^\ast$(with KD) outperform state-of-the-art methods. When using P2T-Large and P2T-Tiny as backbone networks in SSRNet-T and SSRNet-S$^\ast$, respectively, the number of parameters decreased from 68.37M to 15.06M, and the computational floating-point operations decreased from 26.99G to 3.01G. Our code and results are available at https://github.com/WangYuSenn/SSRNet-KD. Wujie Zhou, Xiaohong Qian |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | Deep Incomplete Multi-View Clustering via Dynamic Imputation and Triple Alignment With Dual OptimizationabstractIn recent years, Incomplete Multi-View Clustering (IMVC) has become an important and challenging task. Although several methods have been proposed to address IMVC, they still have the following drawbacks: i) Due to the presence of missing samples in the views, clustering prototypes obtained from different views may have positional deviations, leading to inaccurate positioning of cluster centers, thus affecting the accuracy of clustering results. ii) Repair strategies based on cross-view prediction and adversarial generation have high computational costs and heavily rely on model performance. Neighbor-based repair strategies may result in inaccurate neighbor selection due to the presence of noise. iii) Models learned solely from complete data often perform better than models learned from both complete and incomplete data, especially when there are semantic differences between the repaired data and the missing data. To address the aforementioned issues, this paper proposes a Dynamic Imputation and Triple Alignment with Dual-Optimization for Deep Incomplete Multi-View Clustering (DITA-IMVC). Specifically, by accurately representing advanced features, we propose a cross-view dynamic structure learning strategy for missing view repair, where the dynamic structural relationships between high-semantic features within each view are calculated to obtain highly related samples from different views. To address positional deviations from different views, we propose a triple cross-view alignment with prototype, feature, and clustering assignment, which preserves the consistency among different views. Finally, we design a dual-optimization process for both complete view features and repaired features via alternating iterations to fully utilize the incomplete view data, thereby improving clustering performance. To demonstrate the effectiveness of our DITA-IMVC, extensive experiments conducted on different standard datasets show that it yields superior clustering results compared to existing methods. Weiqing Yan, Kanglong Liu, Wujie Zhou, Chang Tang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | MDNet: Mamba-Effective Diffusion-Distillation Network for RGB-Thermal Urban Dense PredictionabstractIn recent years, significant progress has been achieved in urban dense prediction tasks, particularly with advancements in deep learning models and novel architectures that enhance segmentation accuracy and computational efficiency. However, the following challenges persist: i) Existing modal fusion methods typically adopt convolutional neural networks (CNNs) or transformer (Trans)-based methods, which lead to inadequate global modeling or excessive computation owing to the introduction of quadratic complexity modeling; and ii) existing dense prediction networks typically utilize discriminative networks (codecs), which result in networks with insufficient discriminative properties. To address these issues, we propose the Mamba-effective diffusion-distillation network (MDNet) for RGB-thermal urban dense prediction. First, a new Mamba-effective fusion module is proposed, which efficiently models long-range pixel-level features using Mamba and generates pixel-level adaptive weights to fully utilize complementary modal information. Second, inspired by human self-reflection, a new diffusion self-distillation (DSD) strategy is proposed. The DSD generates coarse-grained binary semantic information via conditional multimodal image diffusion, which serves as self-distillation labels to improve the discriminative properties of the network. Experimental results demonstrate that the proposed MDNet achieves state-of-the-art performance on the MFNet dataset with fewer parameters and reduced computational effort. Extended experiments on the PST900 dataset further illustrate the effectiveness and generalizability of MDNet. The source code and results are available athttps://github.com/Tortoisewhp/MDNet. Wujie Zhou, Hongping Wu, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | MASDG: Multiview Augmented Single-Source Domain Generalization Method for Robust Remote Sensing Building ExtractionabstractDespite advances in deep learning for remote sensing building extraction (RSBE), Multi-target Domain RSBE (MD-RSBE) remains challenging, as it requires transferring knowledge from a labeled source domain to multiple unlabeled target domains, with domain shifts in texture, style, and semantics. Existing domain adaptation (DA) and generalization (DG) methods face significant limitations: DA requires target-domain training, while DG needs multi-source training, leading to high training costs and low generalization in practical MD-RSBE scenarios. To address this, we propose a Multi-view Augmented Single-source Domain Generalization (MASDG) method, which effectively mitigates domain shifts across RS source and target domains for robust MD-RSBE performance by enriching the diversity of the source domain through multi-view augmentation and enforcing semantic consistency. Specifically, MASDG consists of three key components: Texture-level Domain Augmentation (TDA) module, Style-level Domain Augmentation (SDA) module and Semantic-invariant Representation Learning (SRL). To mitigate texture-level domain shift, TDA first introduces parameter-optimized multi-layer random convolution to modify the texture of source image, generating texture-augmented image pairs for simulating real-world texture diversity across various RS domains. Then, with each image pair from TDA, SDA employs two paralleled encoders, namely the general feature encoder and the batch-guided style encoder, to formulate multi-view building features, further mitigating style-level domain shift. Finally, SRL ensures semantic-invariant representation learning via a dual mechanism, including multi-view segmentation loss and semantic consistency loss. The former generates predictions from diverse feature views (original, texture-augmented, style-augmented, etc.), while the latter performs semantic alignment by minimizing distribution discrepancies among predictions, bridging semantic inconsistency to enable robust segmentation. Extensive experiments across three different MD-RSBE settings with 7 different target domains demonstrate that our MASDG outperforms existing state-of-the-art methods by a significant margin. Yunjiao Liu, Yuanyuan Liu 0004, Kejun Liu, Chang Tang, Wujie Zhou, Zhe Chen 0013, Wei Xiang 0001, Hongyan Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2025 | PMTSeg: Prompt-Driven Multimodal Transformer for Task-Adapted Remote Sensing Image SegmentationabstractMultimodal remote sensing image segmentation (MRSIS) is important for intelligent remote sensing image (RS) interpretation, which encompasses three distinct tasks: semantic segmentation, instance segmentation, and panoptic segmentation. Existing methods typically address individual tasks with specialized models, limiting generalization and real-world applicability. Multi-task learning approaches have introduced separated task heads to unify tasks, yet we identify two key challenges when directly applying them to MRSIS: (1) the modality gap, arising from semantic discrepancies and granularity discrepancies across RS modalities, and (2) the task gap, due to varying preferences in learning different segmentation tasks. To overcome these challenges, we propose PMTSeg—a novel Prompt-driven Multimodal Transformer for task-adapted MRSIS. PMTSeg integrates three key components: (1) Task-common Multimodal Affinity Approximation (TMAA), (2) Task-common Multi-scale Semantic Fusion (TMSF), and (3) a unified Prompt-driven Segmentation Head (PSH). First, TMAA addresses the modality gap by approximating inter-modal affinity matrices, extracting task-common features across modalities and aligning semantic information. Then, TMSF further integrates these features using the scale-matched fusion at multiple scales to produce enriched, multi-scale task-common features. Moreover, to address the task gap, the PSH leverages task-adapted text prompts and task-adapted contrastive loss to model relationships across tasks, enabling adaptive optimization for robust and universal MRSIS performance. Extensive experiments on three MRSIS datasets—VALID, SEMCITY TOULOUSE, and UBCV2—demonstrate that PMTSeg significantly surpasses state-of-the-art methods in all three segmentation tasks, offering a unified and accurate solution to MRSIS. Kejun Liu, Xuesong Yan 0001, Yuanyuan Liu 0004, Chang Tang, Yibing Zhan, Wujie Zhou, Hongyan Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2025 | HLMamba: Hybrid Lightweight Mamba-Based Fusion Network for Dense Prediction of Remote Sensing Images
Wujie Zhou, Penghan Yang, Yuanyuan Liu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | DA-Net: A Double Alignment Multimodal Learning Network for Point Cloud Quality AssessmentabstractExisting multimodal point cloud quality assessment (PCQA) methods usually integrate 3D and 2D information to simulate human visual perception of distortions. However, due to the lack of consideration of spatial correspondence, they have difficulty to learn consistent distortion representations from different modalities in the same region of the PC. In addition, they also ignore the heterogeneity of modalities and rely on complex fusion mechanisms (e.g., attention) to integrate multimodal features. Both lead to limited performance and increased computational complexity. To address these limitations, we propose a novel double alignment multimodal learning network (DA-Net), which introduces two key alignment strategies. Specifically, the first is spatial pre-alignment strategy, which generates informative 2D patch for each 3D patch via an adaptive patch projection module (APPM), ensuring accurate spatial correspondence of different modalities prior to feature extraction. The second is a uniform feature alignment strategy, which includes feature disentanglement module (FDM) and feature mapping module (FMM) to relieve heterogeneity of modalities and guide the optimization of 2D and 3D encoder. Finally, multimodal features are simply integrated and regressed to obtain the quality score. Experimental results demonstrate that the DA-Net exhibits outstanding performance and generalization ability. It also achieves lower computational complexity compared with other multimodal PCQA methods. The source codes of DA-Net will be available at https://github.com/Rphone/DA-Net. Xinqiang Wu, Zhouyan He, Ting Luo 0001, Gangyi Jiang, Wujie Zhou, Linwei Zhu, Weisi Lin |
IEEE Trans. Image Process. | 5 |
| 2025 | Transferring Prior Thermal Knowledge for Snowy Urban Scene Semantic SegmentationabstractRGB-thermal (RGB-T) semantic segmentation enables intelligent vehicles to understand environments while operating in urban scenes. However, the research encounters two main challenges: 1) scarcity of training samples under snowy conditions and 2) challenge in applying the model in practice. To address the first challenge, we proposed a publicly accessible RGB-T semantic segmentation dataset in snowy urban scenes (SUS dataset). The SUS dataset comprises 1035 pairs of precisely registered RGB-T images, and provides pixel-level semantic annotations for five categories for all images. To tackle the second challenge, we introduced MCNet-S*, a novel semantic segmentation model that leverages knowledge distillation (KD). The KD structure consists of an RGB-T teacher model, named MCNet-T, and an RGB student model, named MCNet-S. Within MCNet-T, we proposed a cross-modal dual association (CDA) module to enhance utilization of RGB-T information in snowy urban scenes. Within MCNet-S, a depth-wise separable pyramid (DSP) module was proposed to improve the efficiency of RGB information utilization and align the feature dimensions with those of MCNet-T. Between MCNet-S and MCNet-T, memory-based contrastive learning distillation (MCLD) was proposed to transfer the prior thermal knowledge, improving the segmentation accuracy of MCNet-S and obtaining optimized MCNet-S*. Extensive experiments on the SUS and MFNet datasets show that the proposed models outperform state-of-the-art models. The SUS dataset and codes are available at https://github.com/xiaodonguo/SUS_dataset. Tong Liu 0009, Yefeng Mou, Bohan Ren, Wujie Zhou |
IEEE Trans. Intell. Transp. Syst. | 9 |
| 2025 | Local and Global Structure-Guided No-Reference Point Cloud Quality AssessmentabstractAs a crucial representation of 3D data, a point cloud (PC) can accurately capture the geometry, structure, and color information of objects. However, various quality problems arise owing to device noise, data acquisition errors, and compression algorithms, limiting the application of PCs. Therefore, assessing PC quality to determine its suitability for applications is a challenging task. In this work, a local and global structure-guided feature extraction and attention network (LGS-Net) is introduced for no-reference PC quality assessment (PCQA). This approach incorporates cluster construction (CC), local structure-guided cluster feature extraction (LSFE), and global structure-guided attention (GSA) modules. First, owing to the heightened sensitivity of the human visual system (HVS) to structural information, a graph filter is employed to identify high-frequency clusters. Within the LSFE module, a multiscale strategy is employed to ensure that structural information effectively influences both the geometry and color information. Simultaneously, the multiscale features within the cluster are dynamically fine-tuned using feature channel weight reassignment. To account for the impact of interclusters on overall quality, a GSA module is introduced to establish global dependencies between local clusters. This approach enables the extraction of final geometry, color, and structure information, which are ultimately used for accurate quality assessment. Extensive experimental results show that the proposed method outperforms the existing state-of-the-art PCQA methods using two publicly available subjective datasets. Zhouyan He, Qihao Liang, Gangyi Jiang, Mei Yu 0001, Yeyao Chen, Ting Luo 0001, Wujie Zhou |
IEEE Trans. Multim. | 7 |
| 2025 | Emotion-Oriented Cross-Modal Prompting and Alignment for Human-Centric Emotional Video CaptioningabstractHuman-centric Emotional Video Captioning (H-EVC) aims to generate fine-grained, emotion-related sentences for human-based videos, enhancing the understanding of human emotions and facilitating human-computer emotional interaction. However, existing video captioning methods often overlook subtle emotional clues and interactions in videos. As a result, the generated captions frequently lack emotional information. To address this, we proposeEmotion-orientedCross-modalPrompting andAlignment (ECPA), which improves HEVC accuracy by modeling fine-grained visual-textual emotion clues. Using large foundation models, ECPA introduces two learnable prompting strategies: visual emotion prompting (VEP) and textual emotion prompting (TEP), along with an emotion-oriented cross-modal alignment (ECA) module. VEP uses two levels of visual prompts,i.e., emotion recognition (ER) and action unit (AU), to focus on both coarse and fine visual emotional features. TEP devise two-level learnable textual prompts,i.e., sentence-level emotional tokens and word-level masked tokens to capture global and local textual emotion representations. ECA introduces another two levels of emotion-oriented prompt alignment learning mechanisms: the ER-sentence level and the AU-word level alignment losses. Both enhance the model's ability to capture and integrate both global and local cross-modal emotion semantics, thereby enabling the generation of fine-grained emotional linguistic descriptions in video captioning. Experiments show ECPA significantly outperforms state-of-the-art methods on various H-EVC datasets (relative improvements of 9.98%, 5.72%, 4.46%, 24.52% on MAFW, and 12.82%, 20.27%, 4.22%, 5.01% on EmVidCap across four evaluation metrics) and supports zero-shot tasks on MSVD and MSRVTT, demonstrating strong applicability and generalization. Yu Wang 0246, Yuanyuan Liu 0004, Shunping Zhou, Chang Tang, Wujie Zhou, Zhe Chen 0013 |
IEEE Trans. Multim. | 6 |
| 2025 | Multiview Representation Learning via Information-Theoretic OptimizationabstractMultiview data, characterized by rich features, are crucial in many machine learning applications. However, effectively extracting intraview features and integrating interview information present significant challenges in multiview learning (MVL). Traditional deep network-based approaches often involve learning multiple layers to derive latent. In these methods, the features of different classes are typically implicitly embedded rather than systematically organized. This lack of structure makes it challenging to explicitly map classes to independent principal subspaces in the feature space, potentially causing class overlap and confusion. Consequently, the capability of these representations to accurately capture the intrinsic structure of the data remains uncertain. In this article, we introduce an innovative multiview representation learning (MVRL) by maximizing two information-theoretic metrics: intraview coding rate reduction and interview mutual information. Specifically, in the intraview representation learning, we aim to optimize feature representations by maximizing the coding rate difference between the entire dataset and individual classes. This process expands the feature representation space while compressing the representations within each class, resulting in more compact feature representations within each viewpoint. Subsequently, we align and fuse these view-specific features through space transformation and cross-sample fusion to achieve consistent representation across multiple views. Finally, we maximize information transmission to maintain consistency and correlation among data representations across views. By maximizing mutual information between the consensus representations and view-specific representations, our method ensures that the learned representations capture more concise intrinsic features and correlations among different views, thereby enhancing the performance and generalization ability of MVL. Experiments show that the proposed methods have achieved excellent performance. Weiqing Yan, Shuochen Yao, Chang Tang, Wujie Zhou |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Anchor-Sharing and Cluster-Wise Contrastive Network for Multiview Representation LearningabstractMultiview clustering (MVC) has gained significant attention as it enables the partitioning of samples into their respective categories through unsupervised learning. However, there are a few issues as follows: 1) many existing deep clustering methods use the same latent features to achieve the conflict objectives, namely, reconstruction and view consistency. The reconstruction objective aims to preserve view-specific features for each individual view, while the view-consistency objective strives to obtain common features across all views; 2) some deep embedded clustering (DEC) approaches adopt view-wise fusion to obtain consensus feature representation. However, these approaches overlook the correlation between samples, making it challenging to derive discriminative consensus representations; and 3) many methods use contrastive learning (CL) to align the view's representations; however, they do not take into account cluster information during the construction of sample pairs, which can lead to the presence of false negative pairs. To address these issues, we propose a novel multiview representation learning network, called anchor-sharing and clusterwise CL (CwCL) network for multiview representation learning. Specifically, we separate view-specific learning and view-common learning into different network branches, which addresses the conflict between reconstruction and consistency. Second, we design an anchor-sharing feature aggregation (ASFA) module, which learns the sharing anchors from different batch data samples, establishes the bipartite relationship between anchors and samples, and further leverages it to improve the samples' representations. This module enhances the discriminative power of the common representation from different samples. Third, we design CwCL module, which incorporates the learned transition probability into CL, allowing us to focus on minimizing the similarity between representations from negative pairs with a low transition probability. It alleviates the conflict in previous sample-level contrastive alignment. Experimental results demonstrate that our method outperforms the state-of-the-art performance. Weiqing Yan, Yuanyang Zhang, Chang Tang, Wujie Zhou, Weisi Lin |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | IRFR-Net: Interactive Recursive Feature-Reshaping Network for Detecting Salient Objects in RGB-D ImagesabstractUsing attention mechanisms in saliency detection networks enables effective feature extraction, and using linear methods can promote proper feature fusion, as verified in numerous existing models. Current networks usually combine depth maps with red-green-blue (RGB) images for salient object detection (SOD). However, fully leveraging depth information complementary to RGB information by accurately highlighting salient objects deserves further study. We combine a gated attention mechanism and a linear fusion method to construct a dual-stream interactive recursive feature-reshaping network (IRFR-Net). The streams for RGB and depth data communicate through a backbone encoder to thoroughly extract complementary information. First, we design a context extraction module (CEM) to obtain low-level depth foreground information. Subsequently, the gated attention fusion module (GAFM) is applied to the RGB depth (RGB-D) information to obtain advantageous structural and spatial fusion features. Then, adjacent depth information is globally integrated to obtain complementary context features. We also introduce a weighted atrous spatial pyramid pooling (WASPP) module to extract the multiscale local information of depth features. Finally, global and local features are fused in a bottom-up scheme to effectively highlight salient objects. Comprehensive experiments on eight representative datasets demonstrate that the proposed IRFR-Net outperforms 11 state-of-the-art (SOTA) RGB-D approaches in various evaluation indicators. Wujie Zhou, Qinling Guo, Jingsheng Lei, Lu Yu 0003, Jenq-Neng Hwang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Multiattentive Perception and Multilayer Transfer Network Using Knowledge Distillation for RGB-D Indoor Scene ParsingabstractScene parsing has gained wide attention in the field of computer vision, with emerging methods and techniques providing superior solutions. Although some methods have improved performance, they tend to neglect the number of model parameters and computational size, which makes achieving real-time operation in practical applications challenging. To address these limitations, we propose a multiattentive perception and multilayer transfer network that employs knowledge distillation (MPMTNet-KD), which is generated by a student network (MPMTNet-S) under the guidance of a teacher network (MPMTNet-T) with the aid of our proposed multilayer transfer knowledge distillation (KD) methods. To capture complete information from different modalities, a multiattentive perception module (MAPM) is introduced to mine features from various perspectives, and hetero-oriented sensing (HOS) convolution is utilized to integrate cross-layer features in a single and holistic manner. Importantly, we introduce multilayer transfer KD to explore the different knowledge types between layers, as well as intraclass and interclass correlations. In addition, we use the discrete cosine transform (DCT) approach combined with filtering during the KD process to mitigate noise that may be induced by the depth map, thereby improving the depth information and further enhancing the knowledge transfer effect. We conducted comprehensive experiments on two challenging indoor benchmark datasets, namely NYUDv2 and SUN RGB-D. Compared with existing methods, the proposed MPMTNet-KD reduces the number of parameters from 125.8 M in MPMTNet-T to 28.3 M in MPMTNet-S, achieving a mean intersection over union (mIoU) of 54.9% in the indoor scene parsing task. MPMTNet-KD was also evaluated on two additional public datasets, namely MFNet and PST900, to demonstrate its generalization capacity. The source code is available at https://github.com/XUEXIKUAIL/MPMTNet. Wujie Zhou, Bitao Jian, Yuanyuan Liu 0004, Qiuping Jiang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Asymmetrical Contrastive Learning Network via Knowledge Distillation for No-Service Rail Surface Defect DetectionabstractOwing to extensive research on deep learning, significant progress has recently been made in trackless surface defect detection (SDD). Nevertheless, existing algorithms face two main challenges. First, while depth features contain rich spatial structure features, most models only accept red-green-blue (RGB) features as input, which severely constrains performance. Thus, this study proposes a dual-stream teacher model termed the asymmetrical contrastive learning network (ACLNet-T), which extracts both RGB and depth features to achieve high performance. Second, the introduction of the dual-stream model facilitates an exponential increase in the number of parameters. As a solution, we designed a single-stream student model (ACLNet-S) that extracted RGB features. We leveraged a contrastive distillation loss via knowledge distillation (KD) techniques to transfer rich multimodal features from the ACLNet-T to the ACLNet-S pixel by pixel and channel by channel. Furthermore, to compensate for the lack of contrastive distillation loss that focuses exclusively on local features, we employed multiscale graph mapping to establish long-range dependencies and transfer global features to the ACLNet-S through multiscale graph mapping distillation loss. Finally, an attentional distillation loss based on the adaptive attention decoder (AAD) was designed to further improve the performance of the ACLNet-S. Consequently, we obtained the ACLNet-S*, which achieved performance similar to that of ACLNet-T, despite having a nearly eightfold parameter count gap. Through comprehensive experimentation using the industrial RGB-D dataset NEU RSDDS-AUG, the ACLNet-S* (ACLNet-S with KD) was confirmed to outperform 16 state-of-the-art methods. Moreover, to showcase the generalization capacity of ACLNet-S*, the proposed network was evaluated on three additional public datasets, and ACLNet-S* achieved comparable results. The code is available at https://github.com/Yuride0404127/ACLNet-KD. Wujie Zhou, Xiaohong Qian, Meixin Fang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Hybrid Knowledge Distillation Network for RGB-D Co-Salient Object DetectionabstractThe aim of RGB-D Co-salient object detection (RGB-D Co-SOD) is to locate the most prominent objects within a provided collection of correlated RGB and depth images. The development of the Transformer has resulted in significant advancements in RGB-D Co-SOD. However, existing methods overlook the considerable computational and parametric costs associated with using the Transformer. Although compact models are computationally efficient, they suffer from performance degradation, which limits their practical applicability. This is because the reduction of model parameters weakens their feature representation capability. To bridge the performance gap between compact and complex models, we propose a hybrid knowledge distillation (KD) network, HKDNet-S*, to perform the RGB-D Co-SOD task. This method incorporates positive-negative logits approximation KD to guide the student network (HKDNet-S) in effectively learning the interrelationships among samples with multiple attributes by considering both positive and negative logits. HKDNet-S* primarily consists of the group cosaliency semantic exploration module and the positive and negative logits approximation KD method. Specifically, we employ a trained RGB-D Co-SOD model as a teacher model (HKDNet-T) to train the HKDNet-S with a limited number of participants using KD. Through extensive experiments on three challenging benchmark datasets (RGBD CoSal1k, RGBD CoSal150, and RGBD CoSeg183), we demonstrate that HKDNet-S* achieves superior accuracy while utilizing fewer parameters in comparison to the existing state-of-the-art methods. Zhangping Tu, Wujie Zhou, Xiaohong Qian, Weiqing Yan |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2025 | Knowledge Distillation SegFormer-Based Network for RGB-T Semantic SegmentationabstractDeep-learning-based semantic segmentation has received increasing research attention in recent years. However, owing to complex architectures, existing approaches have failed to achieve high accuracies in real-time applications. In this article, a novel knowledge distillation (KD) SegFormer-based network, called KDSNet-S*, is proposed to explore the tradeoff between accuracy and efficiency. Specifically, a structured KD scheme is designed to transfer the rich advanced features of a teacher network (KDSNet-T) to a student network (KDSNet-S). Thereafter, the KDSNet-S network learns the precise segmentation ability of the KDSNet-T network. Additionally, a multifield perceptual fusion model is proposed to learn more integrated features for a single modality and obtain discriminative and comprehensive feature representations. Furthermore, a high-level feature integration module is introduced to refine multimodality high-level features. Finally, multilevel features are fused, and a label-decoupling-based three-stream decoder that decomposes the original semantic segmentation map into center and contour diffusion maps for different supervision tasks is introduced. Experimental results on two public red-green–blue-thermal semantic segmentation datasets indicate the superiority of KDSNet-S* over compared state-of-the-art methods. The KDSNet-S* reduces parameters and floating-point operations per second by 91.1% and 81.9%, respectively, compared with the KDSNet-T. The source codes and results will be available athttps://github.com/purple-ting/KDSNet. Wujie Zhou, Tingting Gong, Weiqing Yan |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2024 | Progressive Adjacent-Layer coordination symmetric cascade network for semantic segmentation of Multimodal remote sensing images
Xiaomin Fan, Wujie Zhou, Xiaohong Qian, Weiqing Yan |
Expert Syst. Appl. | 2 |
| 2024 | CAGNet: Coordinated attention guidance network for RGB-T crowd counting
Wujie Zhou, Weiqing Yan, Xiaohong Qian |
Expert Syst. Appl. | 2 |
| 2024 | MJPNet-S*: Multistyle Joint-Perception Network With Knowledge Distillation for Drone RGB-Thermal Crowd Density Estimation in Smart CitiesabstractCrowd density estimation has gained significant research interest owing to its potential in various industries and social applications. Therefore, this paper proposes a multistyle joint-perception network based on a knowledge distillation-trained student network (MJPNet-S*) for drone-based red–green–blue, thermal/depth (RGB-T/D) crowd density estimation tasks. To provide superior accuracy and efficiency, a novel trimodal working module effectively combines the modalities to facilitate comprehensive extraction and utilization. A two-step strategy comprising high-and low-level fusion is employed in which the high-level features capture relational reasoning and a one-dimensional projection relationship module captures multisensory field information with high-quality semantics. A shallow injection fusion module leverages the multiscale and channel relationships at the low level to combine full-text information interactively. Finally, to reduce resource consumption, a neighboring collaborative distillation method enables the lightweight student network to achieve superior performance by increasing the speed by 92 reducing the number of parameters by 83 of the teacher. Extensive experiments demonstrate that the proposed MJPNet-S* performs remarkably well on two RGB-T datasets. The code will be made public at https://github.com/WBangG/MJPNet. Wujie Zhou, Xiena Dong, Meixin Fang, Weiqing Yan, Ting Luo 0001 |
IEEE Internet Things J. | 1 |
| 2024 | Contrastive learning-based knowledge distillation for RGB-thermal urban scene semantic segmentation
Wujie Zhou, Tong Liu 0009 |
Knowl. Based Syst. | 2 |
| 2024 | Graph Enhancement and Transformer Aggregation Network for RGB-Thermal Crowd CountingabstractCrowd counting has received significant attention in recent years due to its practical applications. In order to address the specific characteristics of RGB and thermal images, we have developed the graph enhancement and transformer aggregation network (GETANet) for generating representative density maps. Our approach incorporates several innovative modules to enhance accuracy. Firstly, we introduced a position-adaptive module that effectively counts individuals’ positions and integrates features extracted from the main framework. Furthermore, we leveraged the advantages of graph convolutional networks (GCNs), which integrate spatial information and exploit relationships between nodes. Specifically, we designed a dual GCN module that further improves the model’s performance by considering the spatial context and relationships among individuals in the crowd. To capture global image information and improve overall performance, we integrated a vision transformer into our model architecture. The vision transformer effectively captures global dependencies and enhances the model’s ability to understand complex crowd scenes. Additionally, we designed a transformer information aggregation module that integrates information from multiple levels, resulting in a highly precise prediction map. Through comprehensive experiments on benchmark datasets such as RGBT-CC and DroneRGBT, our GETANet demonstrated its effectiveness in RGB-thermal crowd counting tasks. Moreover, GETANet showcased remarkable generalization results on the ShanghaiTech-RGBD dataset. Our code has been made publicly available on GitHub at https://github.com/panyi95/GETANet. Wujie Zhou, Meixin Fang, Fangfang Qiang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2024 | DMFTNet: dense multimodal fusion transfer network for free-space detection
Jiabao Ma, Wujie Zhou, Meixin Fang, Ting Luo 0001 |
Multim. Syst. | 2 |
| 2024 | PGGNet: Pyramid gradual-guidance network for RGB-D indoor scene semantic segmentation
Wujie Zhou, Gao Xu, Meixin Fang, Shanshan Mao, Rongwang Yang, Lu Yu 0003 |
Signal Process. Image Commun. | 1 |
| 2024 | Semantic Progressive Guidance Network for RGB-D Mirror SegmentationabstractExisting salient target detection methods tend to use a single-mirror segmentation strategy, which ignores feature hierarchy information in the frequency domain and lacks fine-grained correspondence. To address these challenges, we propose a new semantic progressive guidance network (SPGNet). To mine sufficient effective information, we propose the wavelet bidirectional focusing (WBF) module to aggregate sub-band features through a bidirectional wavelet transform and fuse them with low-level features to deepen the detail mining. We also introduce the Gaussian fusion complementary (GFC) module, which adopts Gaussian filtering technology to optimize the feature space and then efficiently extracts the contour information through enhanced feature processing. In addition, we propose a global correlation bootstrapping (GCB) module that constructs region-to-pixel correlations from a global perspective to achieve fine-grained correspondence. The proposed model achieves competitive results on a benchmark dataset. Wujie Zhou, Weiqing Yan |
IEEE Signal Process. Lett. | 2 |
| 2024 | PiSFANet: Pillar Scale-Aware Feature Aggregation Network for Real-Time 3D Pedestrian DetectionabstractDetecting 3D pedestrian from point cloud data in real-time while accounting for scale is crucial in various robotic and autonomous driving applications. Currently, the most successful methods for 3D object detection rely on voxel-based techniques, but these tend to be computationally inefficient for deployment in aerial scenarios. Conversely, the pillar-based approach exclusively employs 2D convolution, requiring fewer computational resources, albeit potentially sacrificing detection accuracy compared to voxel-based methods. Previous pillar-based approaches suffered from inadequate pillar feature encoding. In this letter, we introduce a real-time and scale-aware 3D Pedestrian Detection, which incorporates a robust encoder network designed for effective pillar feature extraction. The Proposed TriFocus Attention module (TriFA), which integrates external attention and similar attention strategies based on Squeeze and Exception. By comprehensively supervising the point-wise, channel-wise, and pillar-wise of pillar features, it enhances the encoding ability of pillars, suppresses noise in pillar features, and enhances the expression ability of pillar features. The proposed Bidirectional Scale-Aware Feature Pyramid module (BiSAFP) integrates a scale-aware module into the multi-scale pyramid structure. This addition enhances its ability to perceive pedestrian within low-level features. Moreover, it ensures that the significance of feature maps across various feature levels is fully taken into account. BiSAFP represents a lightweight multi-scale pyramid network that minimally impacts inference time while substantially boosting network performance. Our approach achieves real-time detection, processing up to 30 frames per second (FPS). Weiqing Yan, Shile Liu, Chang Tang, Wujie Zhou |
IEEE Signal Process. Lett. | 4 |
| 2024 | Self-Knowledge Distillation-Based Staged Extraction and Multiview Collection Network for RGB-D Mirror SegmentationabstractIntegrating RGB and depth information could improve mirror segmentation performance. Therefore, it is important to extract and use both types of information. Previous algorithms have often used existing backbone networks for feature extraction, frequently ignoring the differences between different feature layers. To address this issue, we propose an algorithm for staged feature extraction using a multitype backbone network that combines the features of convolutional neural networks and transformers. In addition, we developed a multiview collector to extract cross-modal fusion features from various perspectives. Furthermore, we applied the self-knowledge distillation technique to the proposed algorithm, thereby improving the model's performance. The proposed model achieved competitive results on a benchmark dataset. Xiaoxiao Ran, Wujie Zhou |
IEEE Signal Process. Lett. | 3 |
| 2024 | Lightweight Dual Stream Network With Knowledge Distillation for RGB-D Scene ParsingabstractSignificant progress has been made in the field of indoor scene parsing. The increasing demand for lightweight networks is due to the limited hardware capacity of mobile devices. However, there has been a lack of research on the design of lightweight networks for indoor scene parsing. Therefore, we propose lightweight dual stream network (LDSNet) with knowledge distillation (KD) for RGB-D indoor scene parsing. Initially, we developed a two-stream network with three versions (LDSNet-tiny*, LDSNet-small*, and LDSNet-base, where * represents the model after KD) for different scenarios. In the main stream, we designed an integrated joint enhancement module that captures valuable information from both RGB and depth features. This information is then processed by the cascading integration module to generate the final map. To improve the performance of the model, we included an auxiliary extraction module in the auxiliary stream to specifically extract feature information for KD. During the training process, we used hierarchical context loss to distill features and obtain LDSNet-tiny* and LDSNet-small*. We conducted experiments on the NYUDv2 and SUN RGB-D datasets, which demonstrated that our LDSNet-base achieves superior results, while LDSNet-tiny* and LDSNet-small* also exhibit satisfactory performance. Wujie Zhou, Xiaoxiao Ran, Meixin Fang |
IEEE Signal Process. Lett. | 2 |
| 2024 | CMPFFNet: Cross-Modal and Progressive Feature Fusion Network for RGB-D Indoor Scene Semantic SegmentationabstractDepth information can contribute to the semantic segmentation of scenes from red–green–blue (RGB) images. Therefore, the amount of information that can be obtained from RGB and RGB-depth (RGB-D) images is significantly greater for this task. However, RGB and RGB-D modalities are different in terms of object representation. Features that are extracted from these modalities and fused effectively are key to scene semantic segmentation. In addition, complete segmentation requires the fusion of multiscale features to unify global information. However, existing approaches primarily use multiscale features for sequential integration. This study introduces a cross-modal and progressive feature fusion network (CMPFFNet) for semantic segmentation of indoor scenes in RGB-D images. First, a multimodal adaptive alignment fusion (MAAF) module based on an attention mechanism is introduced. This module aligns the two modal channels by additive attention and then computes the spatial similarity between the two modalities based on the dot product to incorporate the complementary information of the depth modality into the RGB modality. In addition, a reverse attention augmentation (RAA) module is introduced to augment the more abstract high-level features for two adjacent multilevel features using the concrete semantic information of the lower-level features in them. After augmenting the extracted multilevel features, a multilevel feature progressive fusion (MFPF) module is deployed; this module sequentially fuses the neighboring two features progressively with emphasis on the spatial semantics. The network uses the Segformer network with high performance as a backbone in multiple computer vision tasks to enhance the segmentation capability. Experimental results obtained from two publicly available datasets of indoor scenes reveal that the proposed CMPFFNet outperforms existing models in semantic segmentation of indoor scenes of RGB-D images.Note to Practitioners—This study introduces a cross-modal and progressive feature fusion network (CMPFFNet) for indoor scene semantic segmentation in RGB-D images. The complementary information of the depth modality is incorporated into the RGB modality in both channel and spatial forms to form a discriminative representation for easy segmentation. A multilevel feature aggregation decoder is proposed to predict the results of semantic segmentation of scenes. The network uses the Segformer network with high performance as a backbone in multiple computer vision tasks to enhance the segmentation capability. Wujie Zhou, Yuxiang Xiao, Weiqing Yan, Lu Yu 0003 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | Modal Evaluation Network via Knowledge Distillation for No-Service Rail Surface Defect DetectionabstractDeep learning techniques have largely solved the problem of rail surface defect detection (SDD), however, two aspects have yet to be addressed. In most existing approaches, two red–green–blue and depth (RGB-D) streams are indiscriminately fused across modalities, ignoring the fact that RGB and depth images produce different feature qualities in different scenes. Additionally, in their focus on performance, previous studies have overlooked the fact that models produce several parameters, resulting in unrealistic practical applications. To address these challenges, we designed a modal evaluation network (MENet) via knowledge distillation (KD) (MENet-S*) for a no-service rail SDD to adaptively manage information in each scenario and achieve model compression. First, to dynamically adjust the feature distribution and quality, dynamic and static feature coding ideas are introduced. Second, modal evaluation distillation is introduced, which allows a compact model (MENet-S) to learn the feature evaluation process of a complex model (MENet-T). Third, to enable MENet-S to learn the dynamic encoding process of MENet-T and to improve the feature representation of MENet-S, we propose accessible knowledge distillation. Furthermore, multitiered KD is introduced to facilitate the learning of MENet-S. Based on extensive experiments using the industrial RGB-D dataset NEU RSDDS-AUG, we observed that MENet-S* (MENet-S with KD) outperformed 16 state-of-the-art methods. In addition, to demonstrate the generalization capability of MENet-S*, we evaluated the proposed network on three additional public datasets, and MENet-S* achieved competitive results. The source codes and results are available at https://github.com/hjklearn/MENet-KD. Wujie Zhou, Jiankang Hong, Weiqing Yan, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | DGPINet-KD: Deep Guided and Progressive Integration Network With Knowledge Distillation for RGB-D Indoor Scene AnalysisabstractSignificant advancements in RGB-D semantic segmentation have been made owing to the increasing availability of robust depth information. Most researchers have combined depth with RGB data to capture complementary information in images. Although this approach improves segmentation performance, it requires excessive model parameters. To address this problem, we propose DGPINet-KD, a deep-guided and progressive integration network with knowledge distillation (KD) for RGB-D indoor scene analysis. First, we used branching attention and depth guidance to capture coordinated, precise location information and extract more complete spatial information from the depth map to complement the semantic information for the encoded features. Second, we trained the student network (DGPINet-S) with a well-trained teacher network (DGPINet-T) using a multilevel KD. Third, an integration unit was developed to explore the contextual dependencies of the decoding features and to enhance relational KD. Comprehensive experiments on two challenging indoor benchmark datasets, NYUDv2 and SUN RGB-D, demonstrated that DGPINet-KD achieved improved performance in indoor scene analysis tasks compared with existing methods. Notably, on the NYUDv2 dataset, DGPINet-KD (DGPINet-S with KD) achieves a pixel accuracy gain of 1.7% and a class accuracy gain of 2.3% compared with DGPINet-S. In addition, compared with DGPINet-T, the proposed DGPINet-KD (DGPINet-S with KD) utilizes significantly fewer parameters (29.3M) while maintaining accuracy. The source code is available at https://github.com/XUEXIKUAIL/DGPINet. Wujie Zhou, Bitao Jian, Meixin Fang, Xiena Dong, Yuanyuan Liu 0004, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Lightweight and Efficient Multimodal Prompt Injection Network for Scene Parsing of Remote Sensing Scene ImagesabstractScene parsing of high-resolution remote sensing images with complex backgrounds has received extensive attention in recent years. As unimodal networks are significantly affected by weather conditions, reflecting complex ground conditions fully and accurately is difficult; therefore, multimodal scene analysis is particularly important. Current multimodal scene-parsing networks often employ a dual-coding architecture to achieve high-performance segmentation. Because prompt learning allows models to understand and capture contextual information more effectively, the proposed prompt injection module (PIM) extracts relevant information from frozen normalized digital surface model (nDSM) features and integrates it into the infrared, red, and green (IRRG) branches through a modal embedding block. To extract the contextual semantic relationships between the local and global features in the image efficiently, we also design a dynamic filter block for feature enhancement. This design facilitates the mutual complementarity and guidance of information between the two modalities and optimizes fusion. The experimental results demonstrate that lightweight and effective multimodal prompt injection network (LENet) outperforms most current state-of-the-art lightweight methods on two public datasets, achieving comparable accuracy to that of traditional methods. It has only 10.81 M parameters, with 2.72 GFLOPS. Our code and results are available athttps://github.com/LYZ00918/LENet. Yangzhen Li, Wujie Zhou, Jiajun Meng, Weiqing Yan |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Multitarget Domain Adaptation Building Instance Extraction of Remote Sensing Imagery With Domain-Common Approximation LearningabstractDeep learning-based building instance extraction on remote sensing imagery (RSI) has achieved tremendous success under the large-scale labeled training data. However, multi-target domain adaptation building instance extraction (MD-BIE) is still a challenge task that involves transferring knowledge from a source domain to multiple unlabeled target domains, which poses various semantic gaps between and within multiple domains,e.g., style, illumination, resolution, density, scale, etc. Most current methods for single-target domain adaptation are not applicable to the more realistic MD-BIE task. To this end, we propose a novel Domain-common Approximation Learning (DAL) for both modelling intra-domain and inter-domain adaptation, thus obtaining robust MD-BIE. DAL contains three main modules: multi-domain style transfer (MST), multi-domain feature approximation (MFA), and multi-domain cascaded instance extraction (MCIE). To alleviate the semantic gaps between multiple domains for inter-domain adaptation, we first employ the MST to learn multiple target-domain-like features that preserve both the styles of target domains and the content of the source domain, and then use the MFA to approximate these features towards a central domain-common space, thus producing domain-common semantic representations. Moreover, we develop the MCIE with hierarchical extraction losses for intra-domain adaptation to extract precise building instance contours from the domain-common semantic representations, further eliminating the potential gaps within multiple domains. By co-learning these three modules in an end-to-end manner, the DAL bridges the semantic gaps between and within multiple domains. Extensive experiments on different popular MD-BIS tasks (SAB → Crowd & WHU, Crowd → SAB & WHU, SAB → Crowd & SAB & WHU and SAB → WHU) show that our DAL outperforms the current methods by a significant margin. Fayong Zhang, Kejun Liu, Yuanyuan Liu 0004, Wujie Zhou, Hongyan Zhang 0001, Lizhe Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | MSTNet-KD: Multilevel Transfer Networks Using Knowledge Distillation for the Dense Prediction of Remote-Sensing ImagesabstractRecently, methods based on convolutional neural networks have achieved good results in the dense prediction of remote-sensing images, particularly when employing normalized digital surface models. However, most existing methods use multiscale convolution and attention methods to mine multimodal feature information without considering the differences and complementarities between the two features. Moreover, previous studies have prioritized model segmentation performance and ignored parametric issues, which makes it difficult to deploy the model in practical applications. To address this challenge, we designed a multilevel semantic transfer network (MSTNet) for the dense prediction of remote-sensing images using a knowledge-distillation approach to adaptively select useful semantic information for the transfer network. We designed a multilevel semantic knowledge alignment distillation framework (MSKA) to enable a compact student model to learn the semantic information extracted from a complex model. The MSKA framework comprises three main components: cross-layer semantic alignment, dynamic semantic aggregation, and softening learning for semantic information transfer and predictive label softening. Experiments on the Vaihingen and Potsdam datasets showed that the student network employing the MSKA framework achieved excellent segmentation performance with only 8.88M parameters and 2.09 gigaFLOPs in terms of computational costs compared with current state-of-the-art methods. The source code and results are available at https://github.com/LYZ00918/MSKANet. Wujie Zhou, Yangzhen Li, Juan Huan, Yuanyuan Liu 0004, Qiuping Jiang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Remote Sensing Image Scene Classification via Graph Template Enhancement and Supplementation Network With Dual-Teacher Knowledge DistillationabstractImage-processing techniques used for remote sensing images (RSIs) have addressed the issues associated with scene classification. However, two challenges remain; first, most bi-modal methods integrate the extracted features indiscriminately while neglecting certain more relevant features, and bi-modal feature qualities may vary in different scenarios. Second, previous studies have improved the feature extraction performance at the expense of computational speed and complexity, hindering the wide application of improved versions. To address these issues, we constructed a graph template enhancement and supplementation network (GTESNet) with dual-teacher knowledge distillation (KD), GTESNet-S$^{\ast }$, to emphasize and manage the extracted features adaptively and compress the model. First, we developed a graph template enhancement (GTE) module that combined long-range contextual information with clustered template features. Second, an exchange-correlation fusion (ECF) module was introduced to allocate feature distribution dynamically and achieve integration. Third, we constructed feedback complementary decoders (FCDs) consisting of two subdecoders with a cascade connection that used a feedback mechanism to supplement the final output. A dual-teacher distillation method that utilized simple confidence perceptions to assign weights to both teachers was implemented during KD. Furthermore, we introduced clustered templates and graph-relation distillations (GRDs), allowing the student GTESNet-S to learn the clustering ability and semantic association process of the teacher GTESNet-T. Finally, we introduced simple progressive decoder distillation (PDD) to facilitate the learning of GTESNet-S. Extensive experiments on two publicly available datasets demonstrated that the GTESNet outperformed similar state-of-the-art (SOTA) methods. The code is available at:https://github.com/MAXHAN22/GTESNet. Wujie Zhou, Penghan Yang, Yuanyuan Liu 0004, Runmin Cong, Qiuping Jiang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | STONet-S*: A Knowledge-Distilled Approach for Semantic Segmentation in Remote Sensing ImagesabstractSemantic segmentation of remote sensing images is a critical research domain. The integration of cross-modal features enhances stability in intricate environments. Despite the impressive performance of existing methods, their complexity and parameter demands remain significant. Our proposed STONet-S$^{\ast }$, a stepped transmission optimization network (STONet) with knowledge distillation (KD), extracts insights from a pretrained extensive teacher network and transfers them to an untrained compact student network. Initially, a group enhancement and interaction unit (GEIU) correct for background noise influence and seamlessly integrates cross-modal features. Additionally, we introduce a stepped transmission decoder (STD) comprising a stepped capture module (SCM) and a self-reverse revision module (SRRM) to capture multiscale information from the ground up. Furthermore, leveraging the frequency domain, we employ frequency-awareness KD using a discrete cosine transform (DCT) and octave convolution to separate high and low-frequency maps, which are subsequently transferred to the student network. Last, detail-delivery and stepped-response KD (SRKD) mechanisms enhance the learning capacity of the student network. Through extensive experimentation on two datasets, STONet-S$^{\ast }$demonstrates superior segmentation accuracy by achieving remarkable results with only 7.19 M parameters. The corresponding code repository can be accessed at:https://github.com/MAXHAN22/STONet. Wujie Zhou, Penghan Yang, Weiwei Qiu, Fangfang Qiang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Transmission Line Detection Through Bi-Directional Guided Registration With Knowledge DistillationabstractTransmission line (TL) inspection plays a crucial role in maintaining a reliable electricity supply to all regions. Computer vision methods, especially those utilizing infrared images, have achieved significant advancements in this field. However, many existing multimodal fusion methods utilize conventional attention mechanisms or simple meta-additions to combine different modalities without proper alignment. Moreover, these methods often rely on a large number of parameters to achieve better performance. To enhance the fusion of disparate modalities and minimize model parameters, we propose bidirectional guided registration via knowledge distillation (BGRNet-S*) for RGB-T transmission line detection (TLD). This approach incorporates a bidirectional registration mechanism within the fusion module and achieves parameter reduction through our knowledge distillation (KD) method. Based on the bidirectional guidance of the non-local position encoding module, accurate feature registration between modes can be achieved. Additionally, we designed response distillation and spatial semantic distillation for our student network (BGRNet-S). Extensive experiments on TLD datasets demonstrate that both our BGRNet-T and BGRNet-S*(BGRNet-S with KD) achieve excellent performance using state-of-the-art methods. When using Shunted-B and Shunted-T as the backbones of the BGRNet-T and BGRNet-S*networks, respectively, the number of parameters was reduced from 52.34M to 13.09M, and the calculated floating-point numbers decreased from 25.23G to 8.93G. Wujie Zhou, Chuanming Ji, Meixin Fang |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | EGFNet: Edge-Aware Guidance Fusion Network for RGB-Thermal Urban Scene ParsingabstractUrban scene parsing is the core of the intelligent transportation system, and RGB–thermal urban scene parsing has recently attracted increasing research interest in the field of computer vision. However, most existing approaches fail to perform good boundary extraction for prediction maps and cannot fully use high-level features. In addition, these methods simply fuse the features from RGB and thermal modalities but are unable to obtain comprehensive fused features. To address these problems, an edge-aware guidance fusion network (EGFNet) was developed in this study for RGB–thermal urban scene parsing. First, a prior edge map generated using the RGB and thermal images were introduced to capture detailed information in the prediction map and then embed the prior edge cues into the feature maps. To fuse the RGB and thermal information effectively, a multimodal fusion module was designed that guarantees adequate cross-modal fusion. Considering the importance of high-level semantic information, global and semantic information modules were proposed to extract rich semantic information from the high-level features. For decoding, simple elementwise addition was utilized for cascaded feature fusion. Finally, to improve the parsing accuracy, multitask deep supervision was applied to the semantic and boundary maps. Extensive experiments were performed on benchmark datasets to demonstrate the effectiveness of the proposed EGFNet and its superior performance compared with the state-of-the-art methods. Shaohua Dong, Wujie Zhou, Caie Xu, Weiqing Yan |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Normalized Cyclic Loop Network for Rail Surface Defect Detection Using Knowledge DistillationabstractIn recent years, the application of computer vision for detecting rail defects has shown promising results. However, as the accuracy of the models improves, they become more complex with a large number of parameters, making it challenging to use them in practical scenarios. Although some lightweight models have been proposed to reduce the number of parameters, maintaining satisfactory performance remains difficult. Therefore, we propose a normalized cyclic loop network (NCLNet) using knowledge distillation (KD), called NCLNet-S*, for rail surface defect detection (RSDD). This model aims to be as lightweight as possible while preserving excellent accuracy. It achieves this by constructing an inter-modal feature cyclic enhancement structure to maximize the use of complementary features between modalities, incorporating a normalized fusion filter module for adaptive weighting of useful knowledge, and transferring knowledge from the teacher network (NCLNet-T) to the student network (NCLNet-S) using our proposed KD framework. Additionally, to enhance the performance and robustness of the NCLNet-S* model, we introduce intermodal cross-layer self-distillation. Extensive experimental results demonstrate that our proposed NCLNet-S* (NCLNet-S with KD) achieves high accuracy while remaining lightweight compared to state-of-the-art models. Furthermore, we conduct experiments on the publicly available RGBD-SOD datasets and achieve satisfactory performance, demonstrating the generality of our NCLNet-S*. Wujie Zhou, Xiaohong Qian |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Morphology-Guided Network via Knowledge Distillation for RGB-D Mirror SegmentationabstractMirror segmentation is an emerging computer vision task that is extensively applied in various fields. However, it presents significant challenges to existing segmentation methods when irregular shapes are involved. Most methods are designed for deployment on heavy-duty host machines that demand substantial computational resources and storage capacity, which limits their feasibility for deployment on mobile devices, where efficient and resource-friendly solutions are required. Therefore, we propose a morphology-guided network (MGNet) with knowledge distillation, called MGNet-S*, to achieve the efficiency required for deployment in mobile devices. In this network, we introduce an erosion dilation fusion module that leverages morphological knowledge to extract texture details from intrinsic features. This module incorporates different optimization strategies for multimodal features. Furthermore, it provides a knowledge-distillation framework specifically tailored to the proposed MGNet-S*. The MGNet-S* includes three effective distillation modules: a semi-soft label, misaligned features, and adaptive aggregation types. These modules facilitate the efficient transfer of knowledge from the MGNet teacher to MGNet student, allowing the lightweight network, MGNet-S*, to achieve remarkable performance. Numerous experiments proved that our proposed MGNet-S* outperformed state-of-the-art methods, achieving an 88.6% reduction in parameter count and 82.5% reduction in floating-point operations compared to those of the MGNet teacher network. Wujie Zhou, Yuqi Cai, Fangfang Qiang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | DSANet-KD: Dual Semantic Approximation Network via Knowledge Distillation for Rail Surface Defect DetectionabstractOwing to the development of convolutional neural networks (CNNs), the detection of defects on rail surfaces has significantly improved. Although existing methods achieve good results, they incur huge computational and parameter costs associated with CNNs. The usual approach to this problem is to design lightweight models that meet the needs of real-world applications; however, the performance is often compromised. To address the aforementioned problems, we designed a dual semantic approximation network via knowledge distillation (DSANet-KD, a student model with knowledge distillation) for rail surface defect detection; it focuses on both foreground and background knowledge and obtains more accurate prediction results. This model comprises an adaptive 3D spatial integration module, feature-optimization decoding module, and dual semantic approximation knowledge-distillation framework. Specifically, we employed a thoroughly trained teacher defect detection network equipped with dual semantic approximation information as an experienced teacher to guide the training of a student defect detection network. Experimental results showed that the proposed DSANet-KD achieved better accuracy with a smaller number of parameters than the state-of-the-art methods. To demonstrate the generalizability of DSANet-KD, experiments were conducted on a publicly available RGBD-SOD dataset, whose source code is available at: https://github.com/hjklearn/DSANet-KD. Wujie Zhou, Jiankang Hong, Xiaoxiao Ran, Weiqing Yan, Qiuping Jiang |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | MC3Net: Multimodality Cross-Guided Compensation Coordination Network for RGB-T Crowd CountingabstractOwing to the expansion in processing of industrial information through advances in machine learning, the demand for accurate crowd counting in various applications is increasing. We propose a multimodality cross-guided compensation coordination network (MC$^{3}$Net) for accurate red–green–blue and thermal (RGB-T) crowd counting. The network includes modules of intricate interactive fusion, feature difference compensation, and complementary attention enhancement. We use ConvNext as the backbone and process the three streams from RGB, thermal, and spliced RGB-T inputs. The multimodality data are sequentially guided and fused hierarchically, fully combining features extracted from the RGB and thermal images. Thereafter, difference compensation is applied to compress fusion and splicing features. Redundant information is removed. Then, feature mismatch is mitigated to enhance complementary information, reduce the loss of details, and finally obtain crowd statistics. Results from extensive experiments on the RGBT-CC dataset indicate the robustness and effectiveness of MC$^{3}$Net, which also achieves high performance on the DroneRGBT dataset and ShanghaiTechRGBD dataset, outperforming existing crowd counting methods. The code and models are available at: https://github.com/WBangG/MC3Net. Wujie Zhou, Jingsheng Lei, Weiqing Yan, Lu Yu 0003 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | UTLNet: Uncertainty-Aware Transformer Localization Network for RGB-Depth Mirror SegmentationabstractMirror segmentation, an emerging discipline in the field of computer vision, involves the identification and marking of mirrors in an image. Current mirror segmentation methods rely on fixed mirror elements as features for object segmentation. However, these methods do not account for the varied quality of feature images obtained under complex real-world conditions, leading to inaccurate segmentation results. To address these limitations, we propose a novel uncertainty-aware transformer localization network (UTLNet) for RGB-D mirror segmentation. Our approach draws inspiration from biomimicry, specifically the behavior pattern of human observation. We aim to explore features from different angles and focus on complex features that are challenging to determine during the coding stage. Additionally, we employ graph convolution to construct complementary dual-modal fusion features. Furthermore, we design a multiscale interaction transformer module using the shifted-window self-attention mechanism to acquire precise position information. In our experiments, the proposed UTLNet surpasses the current state-of-the-art mirror segmentation method as well as alternative task-specific methods. It achieves superior performance across various evaluation scenarios. Wujie Zhou, Yuqi Cai, Weiqing Yan, Lu Yu 0003 |
IEEE Trans. Multim. | 1 |
| 2024 | DHFNet: dual-decoding hierarchical fusion network for RGB-thermal semantic segmentation
Yuqi Cai, Wujie Zhou, Lu Yu 0003, Ting Luo 0001 |
Vis. Comput. | 2 |
| 2023 | Global contextually guided lightweight network for RGB-thermal urban scene understanding
Tingting Gong, Wujie Zhou, Xiaohong Qian, Jingsheng Lei, Lu Yu 0003 |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | CGINet: Cross-modality grade interaction network for RGB-T crowd counting
Wujie Zhou, Xiaohong Qian, Shanshan Mao, Rongwang Yang, Lu Yu 0003 |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | DRNet: Dual-stage refinement network with boundary inference for RGB-D semantic segmentation of indoor scenes
Enquan Yang, Wujie Zhou, Xiaohong Qian, Jingsheng Lei, Lu Yu 0003 |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | MENet: Lightweight multimodality enhancement network for detecting salient objects in RGB-thermal images
Wujie Zhou, Xiaohong Qian, Jingsheng Lei, Lu Yu 0003, Ting Luo 0001 |
Neurocomputing | 2 |
| 2023 | CCFNet: Cross-Complementary fusion network for RGB-D scene parsing of clothing images
Gao Xu, Wujie Zhou, Xiaohong Qian, Lv Ye, Jingsheng Lei, Lu Yu 0003 |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | AMCFNet: Asymmetric multiscale and crossmodal fusion network for RGB-D semantic segmentation in indoor service robots
Wujie Zhou, Yuchun Yue, Meixin Fang, Shanshan Mao, Rongwang Yang, Lu Yu 0003 |
J. Vis. Commun. Image Represent. | 1 |
| 2023 | Edge Detection Guide Network for Semantic Segmentation of Remote-Sensing ImagesabstractThe acquisition of high-resolution satellite and airborne remote sensing images has been significantly simplified due to the rapid development of sensor technology. Several practical applications of high-resolution remote sensing images (HRRSIs) are based on semantic segmentation. However, single-modal HRRSIs are difficult to classify accurately in the complex situation of some scene objects; therefore, the semantic segmentation of multi-source information fusion is gaining popularity. The inherent difference between multimodal features and the semantic gap between multi-level features typically affect the performance of existing multi-mode fusion methods. We propose a multimodal fusion network based on edge detection to address these issues. This method aids multimodal information fusion by utilizing spatial information contained in the boundary. An edge detection guide module is included in the feature extraction stage to realize the boundary information through the fusion of details and semantics between high-level and low-level features. The boundary information is extended into the well-designed multimodal adaptive fusion block (MAFB) to obtain the multimodal fusion features. Furthermore, a residual adaptive fusion block (RAFB) and a spatial position module (SPM) in the feature decoding stage have been designed to fuse multi-level features from the standpoint of local and global dependence. We compared our method to several state-of-the-art (SOTA) methods using the International Society for Photogrammetry and Remote Sensing’s (ISPRS) Vaihingen and Potsdam datasets. The final results demonstrate that our method achieves excellent performance. Jianhui Jin, Wujie Zhou, Rongwang Yang, Lv Ye, Lu Yu 0003 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Adjacent Bi-Hierarchical Network for Scene Parsing of Remote Sensing ImagesabstractDriven by the rapid development and application of earth observation sensors, the scene parsing of remote sensing images (RSIs) has attracted extensive research attention in recent years. Restricted by the limited local receptive field of successive convolution layers, traditional models of scene parsing cannot effectively and interactively utilize the local-global information and digital surface model (DSM) of RSIs. Comparatively, accurate scene parsing faces more challenges because of unbalanced categories, small targets, and more complex scenes. To address these challenges, herein, we propose a novel adjacent bi-hierarchical network (ABHNet). Specifically, we introduce a DSM-enhanced (DSE) module to excavate characteristic DSM information from DSM images and enhance the red, green, and blue (RGB) features by exploiting informative cues between RGB and DSM modalities. In addition, an adjacent context exploration (ACE) module is proposed, which contains current and adjacent branches. The branches first exploit multiscale complementary characteristics of multilevel features and then integrate these features by applying adjacent exploration. Our model includes five ACE modules—three are deployed to activate detailed features and two obtain deep-guided features. The mutual collaboration of deep and detailed features is more beneficial to the segmentation of small objects. Extensive experiments on two remote sensing benchmark datasets (ISPRS Potsdam and Vaihingen) showed that the proposed ABHNet qualitatively and quantitatively outperformed other methods. Jiabao Ma, Wujie Zhou, Jingsheng Lei, Lu Yu 0003 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Deep Video Stabilization via Robust Homography EstimationabstractVideo stabilization can improve the visual quality of videos that have been captured on mobile devices or other handheld cameras, which are more prone to shaking and motion artifacts. Most of the existing deep video stabilization methods adopts optical flow-based, which produce artifacts and distortions caused by pixel-level warping and enquire expensive computation time. In this paper, we present a novel unsupervised deep video stabilization approach that addresses the influence of moving objects on video stabilization through robust homography estimation. Specifically, we design a foreground mask estimation module as a preprocessing step using a pre-trained semantic segmentation guided method to distinguish the foreground and background regions, enabling us to estimate camera motion via analyzing the background motion. Additionally, we design a low-level confidence feature extraction module to improve motion alignment loss and ensure robust motion estimation. By integrating the learned low-level confidence features with the foreground mask, we can then design a motion estimation module that captures the consistent spatial correspondence between frames through local and global feature extraction. At last, the learnt robust homography is leveraged to stabilize videos. Our method outperforms related state-of-the-art approaches in both quality and quantity on three public benchmarks while remaining computationally efficient. Weiqing Yan, Yiqiu Sun 0001, Wujie Zhou, Zhaowei Liu 0001, Runmin Cong |
IEEE Signal Process. Lett. | 3 |
| 2023 | CSANet: Contour and Semantic Feature Alignment Fusion Network for Rail Surface Defect DetectionabstractRail surface defect detection for traffic safety has received considerable attention. With the development of deep learning, numerous methods for combining RGB and depth information have been proposed. However, these methods directly fuse raw features extracted from the backbone, which can lead to ineffective use of the complementary information of the two modalities. In this study, we developed a contour and semantic feature alignment fusion network (CSANet) with bidirectional feature alignment to explore the internal consistency of cross-modal features from both contour and semantic perspectives. First, an adjacency contour feature extraction module was designed to capture high-quality contour information from adjacent low-level features. Second, an attention-aware graph convolution embedded semantic feature extraction module was designed to explore long-range dependencies and extract semantic information. Third, a bidirectional alignment mechanism was designed to explore the internal consistency of contours and semantics between bimodal features. Experimental results on the industrial RGB-D dataset (NEU RSDDS-AUG) revealed that the proposed CSANet outperformed 12 state-of-the-art algorithms in four evaluation metrics. Jinxin Yang, Wujie Zhou, Ruiming Wu, Meixin Fang |
IEEE Signal Process. Lett. | 2 |
| 2023 | MMSMCNet: Modal Memory Sharing and Morphological Complementary Networks for RGB-T Urban Scene Semantic SegmentationabstractCombining color (RGB) images with thermal images can facilitate semantic segmentation of poorly lit urban scenes. However, for RGB-thermal (RGB-T) semantic segmentation, most existing models address cross-modal feature fusion by focusing only on exploring the samples while neglecting the connections between different samples. Additionally, although the importance of boundary, binary, and semantic information is considered in the decoding process, the differences and complementarities between different morphological features are usually neglected. In this paper, we propose a novel RGB-T semantic segmentation network, called MMSMCNet, based on modal memory fusion and morphological multiscale assistance to address the aforementioned problems. For this network, in the encoding part, we used SegFormer for feature extraction of bimodal inputs. Next, our modal memory sharing module implements staged learning and memory sharing of sample information across modal multiscales. Furthermore, we constructed a decoding union unit comprising three decoding units in a layer-by-layer progression that can extract two different morphological features according to the information category and realize the complementary utilization of multiscale cross-modal fusion information. Each unit contains a contour positioning module based on detail information, a skeleton positioning module with deep features as the primary input, and a morphological complementary module for mutual reinforcement of the first two types of information and construction of semantic information. Based on this, we constructed a new supervision strategy, that is, a multi-unit-based complementary supervision strategy. Extensive experiments using two standard datasets showed that MMSMCNet outperformed related state-of-the-art methods. The code is available at:https://github.com/2021nihao/MMSMCNet. Wujie Zhou, Weiqing Yan, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Graph Attention Guidance Network With Knowledge Distillation for Semantic Segmentation of Remote Sensing ImagesabstractDeep learning has become a popular method for studying the semantic segmentation of high-resolution remote sensing images (HRRSIs). Existing methods have adopted convolutional neural networks to achieve better segmentation accuracy of HRRSIs, and the success of these models often depends on the model complexity and parameter quantity. However, the deployment of these models on equipment with limited resources is a significant challenge. To solve this problem, a lightweight student network framework—a graph attention guidance network (GAGNet) with knowledge distillation, called GAGNet-S*—is proposed in this study, which distills knowledge from pretrained large teacher network (GAGNet-T) and builds reliable weak labels to optimize untrained student network (GAGNet-S). Inspired by the graph convolution network, this study designs a graph convolution module called the attention-graph decoder, which combines attention mechanisms with graph convolution to optimize image features and improve segmentation accuracy in the semantic segmentation task of HRRSIs. In addition, a dense cross-decoder was designed for multiscale dense fusion, which utilizes rich semantic information in the high-level features to guide and refine the low-level features from the bottom up. Extensive experiments showed that GAGNet-S* (GAGNet-S with knowledge distillation) achieved excellent segmentation performance on two widely used datasets: Potsdam and Vaihingen. The code and models are available at https://github.com/F8AoMn/GAGNet-KD. Wujie Zhou, Xiaomin Fan, Weiqing Yan, Shengdao Shan, Qiuping Jiang, Jenq-Neng Hwang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | GSGNet-S*: Graph Semantic Guidance Network via Knowledge Distillation for Optical Remote Sensing Image Scene AnalysisabstractIn recent years, optical remote sensing image (ORSI) scene analysis has attracted increasing interest. However, existing networks show a trend of bifurcation. Lightweight networks have very high inference speed but poor inference of contextual information in highly complex backgrounds. In contrast, networks with high-performance contextual information reasoning capability require many parameters and are computationally expensive. Since the knowledge distillation method can greatly lighten the model, we propose a graph semantic guided network (GSGNet) that utilizes knowledge refinement for ORSI scenario analysis, which has a high inference speed while maintaining practical contextual inference capability. Rich semantic and detailed information facilitates semantic segmentation of optical remote sensing images. We design adjacent dynamic capture and local-global map inference modules that can effectively extract low-level spatial details and high-level contextual semantics. To improve the attention map relearning performance of the distillation method, we designed semantically guided fusion modules to locate spatial information and refine edge information. We also employed a structural relationship transfer distillation method in which the structural relationship knowledge of the teacher model (GSGNet-T) was used to guide the student model (GSGNet-S). We compared the performances of GSGNet-T and the GSGNet-S with knowledge distillation (GSGNet-S*) with those of several state-of-the-art methods on the Vaihingen and Potsdam datasets. Extensive experiments showed that GSGNet-S* outperformed most advanced methods with only 19.61M parameters and a computation cost of 2.9G FLOPs. The experimental results and code of our network can be accessed at the following URL: https://github.com/LYZ00918/GSGNet-KD. Wujie Zhou, Yangzhen Li, Weiqing Yan, Meixin Fang, Qiuping Jiang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | WaveNet: Wavelet Network With Knowledge Distillation for RGB-T Salient Object DetectionabstractIn recent years, various neural network architectures for computer vision have been devised, such as the visual transformer and multilayer perceptron (MLP). A transformer based on an attention mechanism can outperform a traditional convolutional neural network. Compared with the convolutional neural network and transformer, the MLP introduces less inductive bias and achieves stronger generalization. In addition, a transformer shows an exponential increase in the inference, training, and debugging times. Considering a wave function representation, we propose the WaveNet architecture that adopts a novel vision task-oriented wavelet-based MLP for feature extraction to perform salient object detection in RGB (red-green-blue)-thermal infrared images. In addition, we apply knowledge distillation to a transformer as an advanced teacher network to acquire rich semantic and geometric information and guide WaveNet learning with this information. Following the shortest-path concept, we adopt the Kullback-Leibler distance as a regularization term for the RGB features to be as similar to the thermal infrared features as possible. The discrete wavelet transform allows for the examination of frequency-domain features in a local time domain and time-domain features in a local frequency domain. We apply this representation ability to perform cross-modality feature fusion. Specifically, we introduce a progressively cascaded sine-cosine module for cross-layer feature fusion and use low-level features to obtain clear boundaries of salient objects through the MLP. Results from extensive experiments indicate that the proposed WaveNet achieves impressive performance on benchmark RGB-thermal infrared datasets. The results and code are publicly available at https://github.com/nowander/WaveNet. Wujie Zhou, Qiuping Jiang, Runmin Cong, Jenq-Neng Hwang |
IEEE Trans. Image Process. | 1 |
| 2023 | LSNet: Lightweight Spatial Boosting Network for Detecting Salient Objects in RGB-Thermal ImagesabstractMost recent methods for RGB (red-green-blue)-thermal salient object detection (SOD) involve several floating-point operations and have numerous parameters, resulting in slow inference, especially on common processors, and impeding their deployment on mobile devices for practical applications. To address these problems, we propose a lightweight spatial boosting network (LSNet) for efficient RGB-thermal SOD with a lightweight MobileNetV2 backbone to replace a conventional backbone (e.g., VGG, ResNet). To improve feature extraction using a lightweight backbone, we propose a boundary boosting algorithm that optimizes the predicted saliency maps and reduces information collapse in low-dimensional features. The algorithm generates boundary maps based on predicted saliency maps without incurring additional calculations or complexity. As multimodality processing is essential for high-performance SOD, we adopt attentive feature distillation and selection and propose semantic and geometric transfer learning to enhance the backbone without increasing the complexity during testing. Experimental results demonstrate that the proposed LSNet achieves state-of-the-art performance compared with 14 RGB-thermal SOD methods on three datasets while improving the numbers of floating-point operations (1.025G) and parameters (5.39M), model size (22.1 MB), and inference speed (9.95 fps for PyTorch, batch size of 1, and Intel i5-7500 processor; 93.53 fps for PyTorch, batch size of 1, and NVIDIA TITAN V graphics processor; 936.68 fps for PyTorch, batch size of 20, and graphics processor; 538.01 fps for TensorRT and batch size of 1; and 903.01 fps for TensorRT/FP16 and batch size of 1). The code and results can be found from the link of https://github.com/zyrant/LSNet. Wujie Zhou, Yun Zhu 0011, Jingsheng Lei, Rongwang Yang, Lu Yu 0003 |
IEEE Trans. Image Process. | 1 |
| 2023 | Embedded Control Gate Fusion and Attention Residual Learning for RGB-Thermal Urban Scene ParsingabstractThe semantic segmentation of road scenes is an important task in autonomous driving. Deep learning has enabled the development of a variety of semantic segmentation networks using RGB and depth data. However, poor lighting conditions and long-distance sensing limit the applicability of RGB and depth cameras. Nevertheless, many existing methods still rely on precise depth maps for scene segmentation. Unlike depth information, thermal imaging provides a visual heat representation that remains accurate under a variety of lighting conditions and over longer distances. For robust and accurate segmentation of scenes collected during autonomous driving, we used the advanced MobileNetV2 network for feature extraction and a fusion strategy with an embedded control gate. In addition, we adopted an encoder–decoder scheme for semantic segmentation and developed an attention residual learning strategy to restore the resolution of the feature map. Finally, semantic and boundary supervision is introduced to optimize parameters of the proposed network. Experimental results show that the proposed network outperforms existing networks on segmentation of urban scenes, and our network can be generalized to depth data. Wujie Zhou, Jingsheng Lei, Lu Yu 0003 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | PGDENet: Progressive Guided Fusion and Depth Enhancement Network for RGB-D Indoor Scene ParsingabstractScene parsing is a fundamental task in computer vision. Various RGB-D (color and depth) scene parsing methods based on fully convolutional networks have achieved excellent performance. However, color and depth information are different in nature and existing methods cannot optimize the cooperation of high-level and low-level information when aggregating modal information, which introduces noise or loss of key information in the aggregated features and generates inaccurate segmentation maps. The features extracted from the depth branch are weak because of the low quality of the depth map, which results in unsatisfactory feature representation. To address these drawbacks, we propose a progressive guided fusion and depth enhancement network (PGDENet) for RGB-D indoor scene parsing. First, high-quality RGB images are used to improve depth data through a depth enhancement module, in which the depth maps are strengthened in terms of channel and spatial correlations. Then, we integrate information from the RGB and enhance depth modalities using a progressive complementary fusion module, in which we start with high-level semantic information and move down layerwise to guide the fusion of adjacent layers while reducing hierarchy-based differences. Extensive experiments are conducted on two public indoor scene datasets, and the results show that the proposed PGDENet outperforms state-of-the-art methods in RGB-D scene parsing. Wujie Zhou, Enquan Yang, Jingsheng Lei, Jian Wan 0001, Lu Yu 0003 |
IEEE Trans. Multim. | 1 |
| 2023 | DBCNet: Dynamic Bilateral Cross-Fusion Network for RGB-T Urban Scene Understanding in Intelligent VehiclesabstractUnderstanding urban scenes is a fundamental capability required of intelligent vehicles. Depth cues provide useful geometric information for semantic segmentation, thus complementing RGB (color) data. Although single-modal RGB images are improved by depth information, semantic segmentation may be degraded in poor-visibility conditions. Thermal imaging can address some limitations of depth data. Therefore, we leverage the multimodal information in RGB-and-thermal (RGB-T) images by introducing a dynamic bilateral cross-fusion network (DBCNet) for RGB-T urban scene understanding. First, RGB-T features extracted by a given backbone are regrouped as high- or low-level features. Second, multimodal high-level features are sent to a dynamic bilateral cross-fusion module for further refinement. Third, a bounded high-level semantic-feature integration module is added to provide feature guidance, and a multitask supervision mechanism is used for fine-tuning. Extensive experiments on two RGB-T urban scene-understanding datasets indicate that DBCNet aggregates multilevel deep features effectively and outperforms state-of-the-art deep-learning scene-understanding methods. Wujie Zhou, Tingting Gong, Jingsheng Lei, Lu Yu 0003 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2022 | Edge-Aware Guidance Fusion Network for RGB-Thermal Scene ParsingabstractRGB–thermal scene parsing has recently attracted increasing research interest in the field of computer vision. However, most existing methods fail to perform good boundary extraction for prediction maps and cannot fully use high-level features. In addition, these methods simply fuse the features from RGB and thermal modalities but are unable to obtain comprehensive fused features. To address these problems, we propose an edge-aware guidance fusion network (EGFNet) for RGB–thermal scene parsing. First, we introduce a prior edge map generated using the RGB and thermal images to capture detailed information in the prediction map and then embed the prior edge information in the feature maps. To effectively fuse the RGB and thermal information, we propose a multimodal fusion module that guarantees adequate cross-modal fusion. Considering the importance of high-level semantic information, we propose a global information module and a semantic information module to extract rich semantic information from the high-level features. For decoding, we use simple elementwise addition for cascaded feature fusion. Finally, to improve the parsing accuracy, we apply multitask deep supervision to the semantic and boundary maps. Extensive experiments were performed on benchmark datasets to demonstrate the effectiveness of the proposed EGFNet and its superior performance compared with state-of-the-art methods. The code and results can be found at https://github.com/ShaohuaDong2021/EGFNet. Wujie Zhou, Shaohua Dong, Caie Xu, Yaguan Qian |
AAAI | 1 |
| 2022 | Filter Pruning via Feature Discrimination in Deep Neural Networks
Yaguan Qian, Bin Wang 0062, Xiaohui Guan, Zhaoquan Gu, Xiang Ling 0001, Shaoning Zeng, Haijiang Wang 0003, Wujie Zhou |
ECCV (21) | 10 |
| 2022 | Robust Network Architecture Search via Feature Distortion Restraining
Yaguan Qian, Shenghui Huang, Bin Wang 0062, Xiang Ling 0001, Xiaohui Guan, Zhaoquan Gu, Shaoning Zeng, Wujie Zhou, Haijiang Wang 0003 |
ECCV (5) | 8 |
| 2022 | RLLNet: a lightweight remaking learning network for saliency redetection on RGB-D images
Wujie Zhou, Jingsheng Lei, Lu Yu 0003 |
Sci. China Inf. Sci. | 1 |
| 2022 | GCNet: Grid-like context-aware network for RGB-thermal semantic segmentation
Wujie Zhou, Yueli Cui, Lu Yu 0003, Ting Luo 0001 |
Neurocomputing | 2 |
| 2022 | HFNet: Hierarchical feedback network with multilevel atrous spatial pyramid pooling for RGB-D saliency detection
Wujie Zhou, Jingsheng Lei, Lu Yu 0003, Ting Luo 0001 |
Neurocomputing | 1 |
| 2022 | GEBNet: Graph-Enhancement Branch Network for RGB-T Scene ParsingabstractRGB-T (red–green–blue and thermal) scene parsing has recently drawn considerable research attention. Although existing methods efficiently conduct RGB-T scene parsing, their performance remains limited by a small receptive field. Unlike methods that capture the global context by fusing multiscale features or using an attention mechanism, we propose a graph-enhancement branch network (GEBNet), which uses long-range dependencies obtained from the branch to refine a coarse semantic map generated by the decoder. Semantic and detail modules embedded in the graph-enhancement branch fuse high- and low-level features. Furthermore, inspired by the ability of graph neural networks to capture the global context, we integrate a novel graph-enhancement module into the network branch to obtain global information from both high-level semantic information and low-level details. Results from extensive experiments on the MFNet and PST900 datasets demonstrate the high performance of the proposed GEBNet and the contributions of its main components to the parsing performance. Shaohua Dong, Wujie Zhou, Xiaohong Qian, Lu Yu 0003 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Hierarchical Decoding Network Based on Swin Transformer for Detecting Salient Objects in RGB-T ImagesabstractAlthough conventional deep convolutional neural networks are effective for contextual semantic segmentation of objects, recent vision transformers can capture global information of an image and are better at capturing semantic associations over longer ranges. In addition, some existing saliency detection methods disregard the guidance of high-level semantic information for low-level features during decoding, and only use layer-by-layer transmission for encoding. Therefore, we propose a hierarchical decoding network based on a swin transformer to perform red–green–blue and thermal (RGB-T) salient object detection (SOD). First, a sine–cosine fusion module performs multimodality intersections and exploits complementarity. As a second fusion stage, an advanced semantic information guidance module adjusts high-level semantic information and low-level detailed characteristics. Finally, a global saliency perception module fuses cross-layer information in a top-down path. Comprehensive experiments demonstrate that the proposed network outperforms 12 state-of-the-art methods on three RGB-T SOD datasets. Wujie Zhou, Lv Ye, Lu Yu 0003 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Depth Repeated-Enhancement RGB Network for Rail Surface Defect InspectionabstractSurface defect inspection of railways is important to ensure safe transportation. However, challenging conditions, such as uneven illumination and similar foreground and background, hinder defect inspection. With the development of deep learning and the wide application of the computer vision, defect inspection has made great progress. Accordingly, we propose a depth repeated-enhancement RGB (red–green–blue) network (DRERNet) for rail surface defect inspection. DRERNet fully uses depth and RGB information to better inspect defects on rail surfaces using an encoder–decoder architecture. In the encoder, a novel cross modality enhancement fusion module uses details from RGB maps and location information from depth maps to perform cross-modality fusion. In the decoder, the details and location information in a multimodality complementation module are repeatedly used to progressively refine the DRERNet prediction. We performed extensive experiments, and compared the proposed DRERNet with 10 state-of-the-art methods on the industrial NEU RSDDS-AUG RGB-depth dataset. The comparison results demonstrate that DRERNet consistently performs better than other methods in the all evaluation measures. Wujie Zhou, Weiwei Qiu, Lu Yu 0003 |
IEEE Signal Process. Lett. | 2 |
| 2022 | MGCNet: Multilevel Gated Collaborative Network for RGB-D Semantic Segmentation of Indoor SceneabstractRGB-D semantic segmentation of indoor scenes has long been an enduring research topic. However, because of the intrinsic differences in modal information and large gaps in multi-level feature cues, adopting the traditional U-Net framework provides suboptimal indoor scene segmentation. In this paper, we consider an effective feature exploration approach to achieve accurate segmentation. Specifically, it consists of three steps. First, in the encoder, we design a difference-exploration fusion module, which extracts the difference weights of the two modalities to guide them for fusion, so as to achieve intrinsically consistent feature fusion. The gated decoder module relates to the remaining two steps. Second, we use a gating unit for each level of fusion information to reduce the difference between layers, which also increases the unique distinction of a specific layer while avoiding the exclusion between layers of information. Finally, we use a serial-parallel alternation strategy to increase the ability to capture contextual knowledge. Considering the above three steps, we construct the multilevel gated collaborative network (MGCNet). Extensive experiments indicate the performance of the proposed MGCNet can compete favorably against state-of-the-art models under three standard metrics. Enquan Yang, Wujie Zhou, Xionghong Qian, Lu Yu 0003 |
IEEE Signal Process. Lett. | 2 |
| 2022 | RTLNet: Recursive Triple-Path Learning Network for Scene Parsing of RGB-D ImagesabstractScene parsing approaches have attracted extensive attention in recent years; although several methods have been developed for scene parsing, most include complex modules for both cross-modality fusion between RGB and depth images in the encoder and image scale level recovery in the decoder under label supervision for high inference accuracy. Cross-modality information in the encoder may be diluted when processed through the decoder, and the supervision results may not be reused effectively, which adversely affects scene parsing. To address these problems, we propose a recursive triple-path learning network (RTLNet) for cross-modality interactions in the decoder using global context and cross-modality fusion modules. The proposed modules fully use cross-modality information to reduce information loss. To enhance the robustness of RTLNet, we add a path to reuse the initial predictions from the decoder and introduce a ladder-shaped feature consistency module to further leverage multiscale features. Experiments are conducted with the proposed RTLNet and nine recent RGB-D indoor scene parsing methods on the NYUv2 and SUN-RGBD indoor scene datasets; the results show that the RTLNet outperforms the other methods. Yuchun Yue, Wujie Zhou, Jingsheng Lei, Lu Yu 0003 |
IEEE Signal Process. Lett. | 2 |
| 2022 | ECFFNet: Effective and Consistent Feature Fusion Network for RGB-T Salient Object DetectionabstractUnder ideal environmental conditions, RGB-based deep convolutional neural networks can achieve high performance for salient object detection (SOD). In scenes with cluttered backgrounds and many objects, depth maps have been combined with RGB images to better distinguish spatial positions and structures during SOD, achieving high accuracy. However, under low-light and uneven lighting conditions, RGB and depth information may be insufficient for detection. Thermal images are insensitive to lighting and weather conditions, being able to capture important objects even during nighttime. By combining thermal images and RGB images, we propose an effective and consistent feature fusion network (ECFFNet) for RGB-T SOD. In ECFFNet, an effective cross-modality fusion module fully fuses features of corresponding sizes from the RGB and thermal modalities. Then, a bilateral reversal fusion module performs bilateral fusion of foreground and background information, enabling the full extraction of salient object boundaries. Finally, a multilevel consistent fusion module combines features across different levels to obtain complementary information. Comprehensive experiments on three RGB-T SOD datasets show that the proposed ECFFNet outperforms 12 state-of-the-art methods under different evaluation indicators. Wujie Zhou, Qinling Guo, Jingsheng Lei, Lu Yu 0003, Jenq-Neng Hwang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | CEGFNet: Common Extraction and Gate Fusion Network for Scene Parsing of Remote Sensing ImagesabstractScene parsing of high spatial resolution (HSR) remote sensing images has achieved notable progress in recent years by the adoption of convolutional neural networks. However, for scene parsing of multimodal remote sensing images, effectively integrating complementary information remains challenging. For instance, the decrease in feature map resolution through a neural network causes loss of spatial information, likely leading to blurred object boundaries and misclassification of small objects. In addition, object scales on a remote sensing image vary substantially, undermining the parsing performance. To solve these problems, we propose an end-to-end common extraction and gate fusion network (CEGFNet) to capture both high-level semantic features and low-level spatial details for scene parsing of remote sensing images. Specifically, we introduce a gate fusion module to extract complementary features from spectral data and digital surface model data. A gate mechanism removes redundant features in the data stream and extracts complementary features that improve multimodal feature fusion. In addition, a global context module and a multilayer aggregation decoder handle scale variations between objects and the loss of spatial details due to downsampling, respectively. The proposed CEGFNet was quantitatively evaluated on benchmark scene parsing datasets containing HSR remote sensing images, and it achieved state-of-the-art performance. Wujie Zhou, Jianhui Jin, Jingsheng Lei, Jenq-Neng Hwang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | DEFNet: Dual-Branch Enhanced Feature Fusion Network for RGB-T Crowd CountingabstractMost existing crowd counting approaches use limited information of RGB (red–green–blue) images and fail to suitably extract potential pedestrians in unconstrained scenarios. Moreover, complementary depth maps do not provide information of locations where people are more likely to be present. However, by incorporating optical and thermal information, the recognition of pedestrians may be enhanced considerably. In fact, thermal imaging information is robust to weather and lighting scenarios, and information from targets can be extracted even at nighttime. By combining RGB and thermal imaging information, we propose a dual-branch enhanced feature fusion network (DEFNet) for RGB-T (RGB and thermal) crowd counting. In DEFNet, an intensive data-enhancement module fuses complementary features of the same sizes from the RGB and thermal modalities, thus combining various rich receptive fields and generating powerful fused RGB-T features. These features describe both spatial structures and appearance details, highlighting information of crowd location. Then, an efficient dilation fusion module applies convolutions to the RGB -T features to obtain flexible and specific features, effectively eliminating the influence of background on the crowd information for density map prediction. Finally, high- and low-level features are used to efficiently obtain density maps through a fusion decoding module. Experimental results on an RGB-T crowd counting dataset indicate that the proposed DEFNet outperforms existing approaches. Furthermore, DEFNet can be generalized to handle RGB and depth data. Wujie Zhou, Jingsheng Lei, Lv Ye, Lu Yu 0003 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | MFFENet: Multiscale Feature Fusion and Enhancement Network For RGB-Thermal Urban Road Scene ParsingabstractCompared with traditional handcrafted features, deep learning has greatly improved the performance of scene parsing. However, it remains challenging under various environmental conditions caused by imaging limitations. Thermal imaging cameras have several advantages over cameras for the visible spectrum, such as operation in total darkness, robustness to shadow effects, insensitivity to illumination variations, and strong ability to penetrate smog and haze. These advantages of thermal imaging cameras make them ideal for the scene parsing of semantic objects in daytime and nighttime. In this paper, we propose a novel multiscale feature fusion and enhancement network (MFFENet) for accurate parsing of RGB–thermal urban road scenes even when the quality of the available RGB data is compromised. The proposed MFFENet consists of two encoders, a feature fusion layer, and a multi-label supervision layer. We concatenate the multi-scale features with the features that contain global semantic information. Furthermore, we explore the cross-modal fusion of RGB and thermal features at multiple stages, rather than fusing them once at the low or high stage. Then, we propose a spatial attention mechanism module that provides a higher weight to (focuses more on) the foreground area, allowing MFFENet to emphasize foreground objects. Finally, multi-label supervision is introduced to optimize parameters of the proposed MFFENet. Experimental results confirm that the proposed MFFENet outperforms similar high-performing methods. Wujie Zhou, Xinyang Lin, Jingsheng Lei, Lu Yu 0003, Jenq-Neng Hwang |
IEEE Trans. Multim. | 1 |
| 2022 | CCAFNet: Crossflow and Cross-Scale Adaptive Fusion Network for Detecting Salient Objects in RGB-D ImagesabstractOwing to the widespread adoption of depth sensors, salient object detection (SOD) supported by depth maps for reliable complementary information is being increasingly investigated. Existing SOD models mainly exploit the relation between an RGB image and its corresponding depth information across three fusion domains: input RGB-D images, extracted feature maps, and output salient object. However, these models do not leverage the crossflows between high- and low-level information well. Moreover, the decoder in these models uses conventional convolution that involves several calculations. To further improve RGB-D SOD, we propose a crossflow and cross-scale adaptive fusion network (CCAFNet) to detect salient objects in RGB-D images. First, a channel fusion module allows for effective fusing depth and high-level RGB features. This module extracts accurate semantic information features from high-level RGB features. Meanwhile, a spatial fusion module combines low-level RGB and depth features with accurate boundaries and subsequently extracts detailed spatial information from low-level depth features. Finally, a purification loss is proposed to precisely learn the boundaries of salient objects and obtain additional details of the objects. The results of comprehensive experiments on seven common RGB-D SOD datasets indicate that the performance of the proposed CCAFNet is comparable to those of state-of-the-art RGB-D SOD models. Wujie Zhou, Yun Zhu 0011, Jingsheng Lei, Jian Wan 0001, Lu Yu 0003 |
IEEE Trans. Multim. | 1 |
| 2021 | SEV-Net: Residual network embedded with attention mechanism for plant disease severity detectionabstractSummary Early and accurate assessment of plant disease severity is key to preventing disease attack. Traditional detection methods rely on manual vision to distinguish between types of disease infection, but this is time consuming, laborious and inaccurate. To address this problem, this paper proposes a deep learning‐based attentional network model (SEV‐Net) for plant disease severity identification and classification. The network embeds the improved channel and spatial attention module into the residual block of ResNet. The proposed attention module reduces the redundancy of information between channels and focuses on the most information‐rich regions of the feature map. In this experiment, SEV‐Net achieved an accuracy of 97.59% and 95.37% for multiple and single plant (Tomato) disease severity classification, which was better than existing attentional networks (SE‐Net and CBAM). Moreover, the combination of visualization techniques showed that SEV‐Net was adept at distinguishing small variations between plant diseases, proving the feasibility and effectiveness of the network. Furthermore, we have also designed and developed an Android application for real‐time classification of plant disease severity. The system deploys the SEV‐Net network model, which has higher classification accuracy and faster recognition speed. Yun Zhao 0002, Jiagui Chen, Xing Xu 0006, Jingsheng Lei, Wujie Zhou |
Concurr. Comput. Pract. Exp. | 5 |
| 2021 | Attention-based contextual interaction asymmetric network for RGB-D saliency prediction
Xinyue Zhang 0010, Wujie Zhou, Jingsheng Lei |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Multiscale multilevel context and multimodal fusion for RGB-D salient object detection
Junwei Wu 0001, Wujie Zhou, Ting Luo 0001, Lu Yu 0003, Jingsheng Lei |
Signal Process. | 2 |
| 2021 | Multi-layer fusion network for blind stereoscopic 3D visual quality prediction
Wujie Zhou, Xinyang Lin, Jingsheng Lei, Lu Yu 0003, Ting Luo 0001 |
Signal Process. Image Commun. | 1 |
| 2021 | TSFNet: Two-Stage Fusion Network for RGB-T Salient Object DetectionabstractSalient object detection (SOD) based on convolutional neural networks has achieved remarkable success. However, further improving the detection performance on challenging scenes (e.g., low-light scenes) requires additional investigation. Thermal infrared imaging captures thermal radiation from the surface of objects. Thus, it is insensitive to lighting conditions and can provide uniform imaging of objects. Accordingly, we propose a two-stage fusion network (TSFNet) integrating RGB and thermal information for RGB-T SOD. For the first fusion stage, we propose a feature-wise fusion module that captures and aggregates united information and intersecting information in each local region of the RGB and thermal images, and then independent decoding is applied to the RGB and thermal features. For the second fusion stage, we propose a bilateral auxiliary fusion module that extracts auxiliary spatial features from the foreground and background of the thermal and RGB modalities. Finally, we use multiple supervision to further improve the SOD performance. Comprehensive experiments demonstrate that TSFNet outperforms 11 state-of-the-art models under various indicators on three RGB-T SOD datasets. Qinling Guo, Wujie Zhou, Jingsheng Lei, Lu Yu 0003 |
IEEE Signal Process. Lett. | 2 |
| 2021 | Two-Stage Cascaded Decoder for Semantic Segmentation of RGB-D ImagesabstractExploiting RGB and depth information can boost the performance of semantic segmentation. However, owing to the differences between RGB images and the corresponding depth maps, such multimodal information should be effectively used and combined. Most existing methods use the same fusion strategy to explore multilevel complementary information at various levels, likely ignoring different feature contributions at various levels for segmentation. To address this problem, we propose a network using a two-stage cascaded decoder (TCD), embedding a detail polishing module, to effectively integrate high- and low-level features and suppress noise from low-level details. Additionally, we introduce a depth filter and fusion module to extract informative regions from depth cues with the guidance of RGB images. The proposed TCD network achieves comparable performance to state-of-the-art RGB-D semantic segmentation methods on the benchmark NYUDv2 and SUN RGB-D datasets. Yuchun Yue, Wujie Zhou, Jingsheng Lei, Lu Yu 0003 |
IEEE Signal Process. Lett. | 2 |
| 2021 | MRINet: Multilevel Reverse-Context Interactive-Fusion Network for Detecting Salient Objects in RGB-D ImagesabstractThe use of RGB-D information for salient object detection (SOD) is being increasingly explored. Traditional multilevel models handle both low- and high-level features similarly, as they use the same number of features for blending. Unlike these models, in this paper, we propose multilevel reverse-context interactive-fusion (MRI) network (MRINet) for RGB-D SOD. Specifically, first, we extract and reuse different numbers of features depending on their level; the deeper the information, the more times do we perform the extraction. Deeper information contains more semantic cues, which are important for locating salient regions. Thereafter, we use an RGB MRI block (MRIB) to merge RGB information at different levels; furthermore, we use depth features as auxiliary information and an RGB-D MRIB for full merging with RGB information. RGB and RGB-D MRIBs can reconstruct the high-level feature map in high resolution and integrate the low-level feature map to enhance boundary details. Extensive experiments demonstrate the effectiveness of the proposed MRINet and its state-of-the-art performance in RGB-D SOD. Wujie Zhou, Sijia Pan, Jingsheng Lei, Lu Yu 0003 |
IEEE Signal Process. Lett. | 1 |
| 2021 | Parallax-Estimation-Enhanced Network With Interweave Consistency Feature Fusion for Binocular Salient Object DetectionabstractSalient object detection (SOD) has received extensive attention in recent years, and many models have been developed. However, most SOD models only consider monocular images and not binocular images, which resemble the human vision and can better reflect human perception for distinguishing salient objects. To leverage the information in binocular images, we propose herein a first-of-its-kind parallax-estimation-enhanced network (PEENet) for binocular SOD. More specifically, we use a weighted binocular fusion module and a parallax correlation fusion module to explore the complementary and different information in binocular images. In addition, a parallax enhancing module and interweave consistency fusion use complementary saliency information and parallax information to enhance saliency and parallax representations. Finally, a transformation module avoids global and local information loss during decoding. Experiments were performed to validate the effectiveness and robustness of the proposed PEENet, which outperforms 10-RGB/RGB-D SOD methods on two binocular SOD datasets. Yun Zhu 0011, Wujie Zhou, Lu Yu 0003 |
IEEE Signal Process. Lett. | 2 |
| 2021 | GMNet: Graded-Feature Multilabel-Learning Network for RGB-Thermal Urban Scene Semantic SegmentationabstractSemantic segmentation is a fundamental task in computer vision, and it has various applications in fields such as robotic sensing, video surveillance, and autonomous driving. A major research topic in urban road semantic segmentation is the proper integration and use of cross-modal information for fusion. Here, we attempt to leverage inherent multimodal information and acquire graded features to develop a novel multilabel-learning network for RGB-thermal urban scene semantic segmentation. Specifically, we propose a strategy for graded-feature extraction to split multilevel features into junior, intermediate, and senior levels. Then, we integrate RGB and thermal modalities with two distinct fusion modules, namely a shallow feature fusion module and deep feature fusion module for junior and senior features. Finally, we use multilabel supervision to optimize the network in terms of semantic, binary, and boundary characteristics. Experimental results confirm that the proposed architecture, the graded-feature multilabel-learning network, outperforms state-of-the-art methods for urban scene semantic segmentation, and it can be generalized to depth data. Wujie Zhou, Jingsheng Lei, Lu Yu 0003, Jenq-Neng Hwang |
IEEE Trans. Image Process. | 1 |
| 2021 | Salient Object Detection in Stereoscopic 3D Images Using a Deep Convolutional Residual AutoencoderabstractIn recent years, the detection of distinctive objects in stereoscopic 3D images has drawn increasing attention. Unlike 2D salient object detection, salient object detection in stereoscopic 3D images is highly challenging. Hence, we propose a novel Deep Convolutional Residual Autoencoder (DCRA) for end-to-end salient object detection in stereoscopic 3D images. The core trainable architecture of the salient object detection model employs raw stereoscopic 3D images as the inputs and their corresponding ground truth saliency masks as the labels. A convolutional residual module is applied to both the encoder and the decoder as a basic building block in the DCRA, and long-range skip connections are employed to bypass the equal-sized feature maps between the encoder and the decoder. To explore the complex relationships and exploit the complementarity between RGB (photometric) and depth (geometric) information, multiple feature map fusion modules are constructed. These modules integrate texture and structure information between the RGB and depth branches of the encoder and fuse their features over several multiscale layers. Finally, to efficiently optimize DCRA parameters, a supervision pyramid based on boundary loss and background prior loss is adopted, which employs supervised learning over the multiscale layers in the decoder to prevent vanishing gradients and accelerate the training at the fusion stage. We compare the proposed DCRA with state-of-the-art methods on two challenging benchmark datasets. The results of these experiments demonstrate that our proposed DCRA performs favorably against the comparison models. Wujie Zhou, Junwei Wu 0001, Jingsheng Lei, Jenq-Neng Hwang, Lu Yu 0003 |
IEEE Trans. Multim. | 1 |
| 2021 | Global and Local-Contrast Guides Content-Aware Fusion for RGB-D Saliency PredictionabstractMany RGB-D visual attention models have been proposed with diverse fusion models; thus, the main challenge lies in the differences in the results between the different models. To address this challenge, we propose a local-global fusion model for fixation prediction on an RGB-D image; this method combines global and local information through a content-aware fusion module (CAFM) structure. First, it comprises a channel-based upsampling block for exploiting global contextual information and scaling up this information to the same resolution as the input. Second, our Deconv block contains a contrast feature module to utilize multilevel local features stage-by-stage for superior local feature representation. The experimental results demonstrate that the proposed model exhibits competitive performance on two databases. Wujie Zhou, Jingsheng Lei, Lu Yu 0003 |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2020 | Spot evasion attacks: Adversarial examples for license plate recognition systems with convolutional neural networks
Yaguan Qian, Dan-feng Ma, Bin Wang 0062, Jun Pan 0004, Jiamin Wang 0003, Zhaoquan Gu, Jian-Hai Chen, Wujie Zhou, Jing-Sheng Lei |
Comput. Secur. | 8 |
| 2020 | Asymmetric Deeply Fused Network for Detecting Salient Objects in RGB-D ImagesabstractMost RGB-D salient object detection (SOD) models use the same network to process RGB images and their corresponding depth maps. Subsequently, these models perform direct concatenation and summation at deep or shallow layers. However, these models ignore the complementarity of multi-level features extracted from RGB images and depth maps. This paper presents an asymmetric deeply fused network (ADFNet) for RGB-D SOD. Two different backbone networks, i.e., ResNet-50 and VGG-16, are utilized to process RGB images and related depth maps. We use an aggregation decoder and adaptive attention transformer module (AATM) to avoid information loss in the decoding process. Additionally, we use an attention early fusion module (AEFM) and deep fusion module (DFM) to deal with the deep features in various complex situations. Experiments validate the effectiveness of the proposed ADFNet, which outperforms thirteen recent RGB-D SOD models in the analysis of five public RGB-D SOD datasets. Wujie Zhou, Jingsheng Lei |
IEEE Signal Process. Lett. | 2 |
| 2020 | GFNet: Gate Fusion Network With Res2Net for Detecting Salient Objects in RGB-D ImagesabstractThe performance of recent RGB-D salient object detectors has significantly improved owing to their integration of convolutional neural networks (CNNs). However, most existing salient object detection (SOD) methods represent features using a VGGNet backbone, which lacks the ability to retain complete RGB and depth modals and must compensate by applying several skip connections. In this letter, we propose a gate fusion network (GFNet) with Res2Net architecture to solve this problem. GFNet consists of two interacting Res2Net block encoder streams and four gate fusion block (GFB) decoders to interconnect the streams and fuse features. Res2Net blocks have a robust feature retention mechanism to ensure that the decoders can learn complete information, while the GFB formulates the interdependences of the encoders and eliminates noise via a gate mechanism. We evaluated GFNet using two popular RGB-D salient detection benchmark datasets (NJU2000 and NLPR) and achieved state-of-the art performance. Wujie Zhou, Lu Yu 0003 |
IEEE Signal Process. Lett. | 1 |
| 2019 | Deep blind quality evaluator for multiply distorted images based on monogenic binary coding
Wujie Zhou, Lu Yu 0003, Yaguan Qian, Weiwei Qiu, Yang Zhou 0011, Ting Luo 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | A novel robust color image watermarking method using RGB correlations
Fangyan Zhang, Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Wujie Zhou |
Multim. Tools Appl. | 6 |
| 2019 | Deep Road Scene UnderstandingabstractRoad scene understanding is a difficult task in autonomous driving. In this letter, we propose a novel deep encoder-decoder architecture for road scene understanding in an end-to-end manner. This core trainable understanding engine includes an encoder network, a decoder network with two streams, and a pixel-level fusion network with classification layer. The encoder network is composed of the front-end model of the classical convolution neural network, VGGNet. The decoder network with two streams includes multi-scale skip connection modules to reduce the down-scaling effect. Finally, a fusion network fuses the two-level information from the two streams of the decoder network for precise pixel-level classification. Additionally, the convolution layer is added to each skip connection module to increase the depth of the architecture. Our architecture achieves outstanding performance on the publicly available CamVid dataset and significantly outperforms previous architectures. This deep architecture is ideal for road scene understanding. Wujie Zhou, Sijia Lv, Qiuping Jiang, Lu Yu 0003 |
IEEE Signal Process. Lett. | 1 |
| 2018 | Local and Global Feature Learning for Blind Quality Evaluation of Screen Content and Natural Scene ImagesabstractThe blind quality evaluation of screen content images (SCIs) and natural scene images (NSIs) has become an important, yet very challenging issue. In this paper, we present an effective blind quality evaluation technique for SCIs and NSIs based on a dictionary of learned local and global quality features. First, a local dictionary is constructed using local normalized image patches and conventional -means clustering. With this local dictionary, the learned local quality features can be obtained using a locality-constrained linear coding with max pooling. To extract the learned global quality features, the histogram representations of binary patterns are concatenated to form a global dictionary. The collaborative representation algorithm is used to efficiently code the learned global quality features of the distorted images using this dictionary. Finally, kernel-based support vector regression is used to integrate these features into an overall quality score. Extensive experiments involving the proposed evaluation technique demonstrate that in comparison with most relevant metrics, the proposed blind metric yields significantly higher consistency in line with subjective fidelity ratings. Wujie Zhou, Lu Yu 0003, Yang Zhou 0011, Weiwei Qiu, Mingwei Wu 0001, Ting Luo 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Blind 3D image quality assessment based on self-similarity of binocular features
Wujie Zhou, Shuangshuang Zhang, Lu Yu 0003, Weiwei Qiu, Yang Zhou 0011, Ting Luo 0001 |
Neurocomputing | 1 |
| 2017 | Local gradient patterns (LGP): An effective local-statistical-feature extraction scheme for no-reference image quality assessment
Wujie Zhou, Lu Yu 0003, Weiwei Qiu, Yang Zhou 0011, Mingwei Wu 0001 |
Inf. Sci. | 1 |
| 2017 | Blind quality estimator for 3D images based on binocular combination and extreme learning machine
Wujie Zhou, Lu Yu 0003, Yang Zhou 0011, Weiwei Qiu, Mingwei Wu 0001, Ting Luo 0001 |
Pattern Recognit. | 1 |
| 2016 | Utilizing binocular vision to facilitate completely blind 3D image quality measurement
Wujie Zhou, Lu Yu 0003, Weiwei Qiu, Ting Luo 0001, Zhongpeng Wang, Mingwei Wu 0001 |
Signal Process. | 1 |
| 2016 | Binocular Responses for No-Reference 3D Image Quality AssessmentabstractPerceptual quality assessment of distorted three-dimensional (3D) images has become a fundamental yet challenging issue in the field of 3D imaging. In this paper, we propose a general-purpose blind/no-reference (NR) 3D image quality assessment (IQA) metric that utilizes the complementary local patterns (the local magnitude pattern and the proposed generalized local directional pattern) of binocular energy response (BER) and binocular rivalry response (BRR). The main technical contribution of this research is that binocular visual perception and local structural distribution are considered for NR 3D-IQA. More specifically, the metric simulates the binocular visual perception using BER and BRR. Subsequently, the local patterns of the binocular responses' encoding maps are used to form various binocular quality-predictive features, which will change in the presence of distortions. After feature extraction, we use k-nearest neighbors-based machine learning to drive the overall quality score. We tested our proposed metric against two publicly available 3D databases; these tests confirm that the proposed metric's results consistently align with human subjective judgments. Wujie Zhou, Lu Yu 0003 |
IEEE Trans. Multim. | 1 |
| 2014 | Reduced-reference stereoscopic image quality assessment based on view and disparity zero-watermarks
Wujie Zhou, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
Signal Process. Image Commun. | 1 |
| 2014 | PMFS: A Perceptual Modulated Feature Similarity Metric for Stereoscopic Image Quality AssessmentabstractStereoscopic image quality assessment (SIQA) is an important and challenging issue in three dimensional applications. In this letter, a perceptual modulated feature similarity (PMFS) metric for SIQA is proposed by considering the monocular and binocular perception properties. Specifically, stereoscopic image is first classified into monocular occlusion and binocular rivalry regions. Then, feature similarities between the original and distorted stereoscopic images are defined and measured for the monocular occlusion and binocular rivalry regions as the local monocular and binocular quality maps, respectively. Monocular and binocular just noticeable difference visual saliency models are presented to construct a modulation function to derive monocular and binocular quality scores. Finally, those scores are integrated into an overall quality score by support vector regression. Extensive experiments performed on LIVE phase II and MICT asymmetric databases demonstrate that the proposed PMFS metric can achieve much higher consistency with the subjective quality scores than some state-of-the-art SIQA metrics. Wujie Zhou, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
IEEE Signal Process. Lett. | 1 |
| 2013 | Generating Partial Covering Array for Locating Faulty Interactions in Combinatorial Testing
Ziyuan Wang 0001, Wujie Zhou, Weifeng Zhang 0001, Baowen Xu |
SEKE | 3 |
| 2012 | A New Approach to Evaluate Path Feasibility and Coverage Ratio of EFSM Based on Multi-objective Optimization
Zhenyu Chen 0001, Baowen Xu, Zhiyi Zhang 0004, Wujie Zhou |
SEKE | 5 |
| 2008 | Infrared and Visible Image Fusion via Multiscale Receptive Field Amplification Fusion NetworkabstractInfrared and visible image fusion, which highlights radiometric and detailed texture information and completely and accurately describes objects, is a long-standing and well-studied task in computer vision. Existing convolutional neural network-based approaches that leverage end-to-end networks to fuse infrared and visible images have made significant progress. However, most approaches typically extract the features in the encoder segment and use a coarse fusion strategy. Unlike these algorithms, this study proposes a multiscale receptive field amplification fusion network (MRANet) to effectively extract the local and global features from images. Particularly, we extract long-range information in the encoder segment using a convolutional residual structure as the main backbone and a simplified uniformer as an auxiliary backbone, both of which are ResNet-inspired. Additionally, we propose an effective multiscale fusion strategy based on an attention mechanism to integrate the two modalities. Extensive experiments demonstrate that MRANet performs efficiently on image fusion datasets. Chuanming Ji, Wujie Zhou, Jingsheng Lei, Lv Ye |
IEEE Signal Process. Lett. | 2 |