Weiqing Yan

dblp:168/3958 · DBLP profile ↗
← Back
65ranked-venue papers
15as first author
58since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 10 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 2 first-author · 18 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 ATGFB-MFF: Adaptive Text-Guided Fiber Bundle Feature Fusion with LLMs for Multimodal Sentiment Analysis and Emotion Recognition in Conversations
abstract
Multimodal Sentiment Analysis (MSA) and Emotion Recognition in Conversations (ERC) have rapidly developed into pivotal tasks in artificial intelligence. Large Language Models (LLMs) offer powerful semantic reasoning and computational capabilities, showing great potential for understanding emotional content. However, when applied to multimodal sentiment data, LLMs face significant challenges, including the inability to directly process heterogeneous data, difficulties in coping with feature misalignment and suboptimal cross-modal fusion. To address these challenges, we propose a novel multimodal sentiment inference framework named ATGFB-MFF which grounded in fiber bundle theory. This method decomposes multimodal features into an adaptive text-guided shared semantic space and fiber offset spaces to achieve structured alignment and fusion. Then the fused features are converted into structured pseudo-token sequences for effective inference via frozen LLMs. We also introduce two loss functions respectively called shared space consistency loss and fiber offset regularization loss which are used to improve representation stability. Extensive experiments on four benchmark datasets demonstrate that ATGFB-MFF consistently outperforms state-of-the-art baselines. These results highlight the efficacy of geometric structural modeling in unlocking the potential of LLMs for multimodal sentiment inference.
Zhaowei Liu 0001, Weiqing Yan, Peng Song 0002, Yongchao Song, Rufei Gao
WWW3
2026 Collaborative Subgraph Learning based Spectrum Sensing under Partial Observations
abstract
Data-driven spectrum sensing is a key technology for addressing complex challenges in Cognitive Radio Networks (CRNs). Traditional methods are typically designed for simple single-band scenarios and perform poorly in practical wideband applications. In real-world systems, a single Secondary User (SU) is often restricted by energy, time, and hardware capabilities during real-time sensing. Consequently, only local and fragmented frequency information can be obtained. This partial sensing leads to a severe lack of training data. Additionally, the lack of historical records for emerging frequency bands, combined with data incompleteness due to resource constraints, creates training bottlenecks for data-driven models and limits the reliability of sensing. To address these challenges, this paper proposes a novel framework based on Collaborative Subgraph Learning and Hyperbolic Graph Neural Networks (GNNs). This approach enables Secondary Users to perform collaborative sensing through distributed subgraph learning. By utilizing GNNs to extract features and model multi-band correlations, a new distributed GNNs architecture is designed to efficiently detect wideband spectrum occupancy, even with partial observations. Within this framework, all frequency bands in the wideband spectrum pool are treated as a unified graph, while the bands observed by each SU form a subgraph. Subsequently, the complete spectrum graph is constructed through the joint training and aggregation of these subgraphs. By integrating hyperbolic geometry into GNNs, this method better captures the hierarchical structure of spectrum patterns, providing a more accurate and efficient sensing model. Experimental results demonstrate that, compared to the second-best HCNNs model, the proposed framework improves sensing accuracy by 3.8% on average across various test environments, while reducing key resource consumption by 18.4% on average.
Zhaowei Liu 0001, Weiqing Yan, Yongchao Song, Anzuo Jiang
WWW4
2026 GLF-Net: Global-local fusion network for radar signal modulation recognition
Xingnong Liu, Xiaolin Du, Xiaolong Chen 0001, Guolong Cui, Jibin Zheng, Wenming Ma, Jinglei Liu, Zhaowei Liu 0001, Weiqing Yan
Expert Syst. Appl.9
2026 Hybrid explicit and implicit encoding for multi-view representation learning
Shuochen Yao, Yusheng Zhang, Weiqing Yan, Chang Tang, Guanghui Yue 0001, Kaile Su
Pattern Recognit.3
2026 HGS-3DSeg: Identity-Encoding Half-Gaussian Splatting for Memory-Efficient 3D Reconstruction and Segmentation
abstract
Recent advancements in 3D Gaussian Splatting have achieved high-quality and real-time novel view synthesis for 3D scenes. However, this method primarily focuses on appearance and geometric modeling, lacks the ability to comprehend scenes with fine granularity at the object level. Some techniques enhance Gaussian Splatting, empowering it with the capability to perform unified 3D reconstruction and segmentation on real-world 3D scenes. Despite its improvements, these methods still face challenges in segmentation accuracy and reconstruction quality using 3D Gaussian representation. Additionally, the use of 16-bit identity codes to distinguish different Gaussian groups significantly increases memory overhead. To address these issues, we propose Identity-Encoding Half-Gaussian (ID-HGS) kernels. Our approach introduces a plane to split each Gaussian into two parts with distinct opacity values, enabling precise reconstruction of details and object boundaries. We replace the adaptive density control (ADC) used in Gaussian Grouping with localized half—gaussian point management (LHPM), which performs finer densification in under-reconstructed regions. LHPM resets pathological Gaussians and optimizes Gaussian density, reducing their impact on segmentation accuracy. Furthermore, we assign a global contribution score to each Gaussian and prune low-contribution Gaussians during training, saving memory and accelerating training. Compared to Gaussian Grouping, our method improves both reconstruction quality and segmentation accuracy while effectively controlling memory usage. Extensive experiments demonstrate that HGS-3D outperforms prior Gaussian Grouping on both reconstruction and segmentation: it achieves higher mask accuracy on the LERF-Localization benchmark and reduces peak memory usage while improving render quality on the mipnerf360 dataset.
Weiqing Yan, Kaile Su, Chang Tang
IEEE Trans Autom. Sci. Eng.1
2026 Lightweight Scope Integration Network for Rail Surface Defect Detection
abstract
Rail-surface defect detection (RSDD) is a key technology for ensuring the safety and efficiency of railroad transportation. Existing models enhance the robustness in complex scenarios by using complementary information from visible light (RGB) image and depth data. However, incorporating depth data increases the computational complexity, which consumes more resources, increases the risk of overfitting, and decreases the model efficiency. Optimizing and streamlining multimodal data processing remains challenging. To address this issue and enable fast and accurate RSDD, this study introduces a novel lightweight scope integration network (LSINet). This method efficiently and precisely fuses the features using a lightweight double-context-aware module and refines the image quality using an adaptive Markov field smoothing module in a hierarchical decoding process. A lightweight two-dimensional scanning method that captures long-range dependencies and improves computational efficiency. We integrate this approach with local range extremes to enhance the multimodal feature fusion. Furthermore, to refine the defect edges and improve the detection accuracy, we integrate the Markov random field module into the defect segmentation and optimization using its ability to model interpixel spatial correlation, particularly for small and ambiguous defects. In designing the model, we focused on optimizing the number of parameters (e.g., computational resources). Experimental evaluations using various rail surface images demonstrated that the proposed method enhanced the defect detection accuracy and recall while reaching fast detection speeds. According to extensive experiments conducted using the industrial RGB-D dataset (i.e., NEU RSDDS-AUG), LSINet outperformed 15 state-of-the-art methods with only 27.6 million parameters. Code and results are publicly available at https://github.com/Wuyue15/LSINet.
Wujie Zhou, Fangfang Qiang, Weiqing Yan
IEEE Trans. Big Data4
2026 Mutual Calibration Network for Multi-View Clustering
abstract
Multi-view clustering, which uses information from multiple views to partition data into distinct clusters, has garnered significant attention. Existing MLP-based and GCN-based algorithms primarily focus on enhancing performance by extracting node attribute and graph structure features, and then integrating them for clustering. However, features extracted by MLP and GCN may differ in quality due to the high sparsity and noise in multi-view data. Directly integrating these features can lead to feature contamination. To address this issue, we propose a novel mutual calibration network for multi-view clustering (McMVC). Specifically, the attribute and structural features from multi-view data are integrated separately using the attention fusion module. A classifier is then employed to obtain high-confidence pseudo-labels for these fused features. We design a Cluster-level Mutual Calibration (CLMC) module that uses the pseudo-labels to mutually calibrate view features while maintaining compact cluster structures. During training, the attribute and structural features are not directly integrated to avoid feature contamination but are jointly optimized by reliable class information. Concurrently, we construct a Centroid-level Contrastive Calibration (CLCC) module to map view features into their inner centroid space and learn more discriminative centroid representations. Our approach outperforms state-of-the-art methods in multi-view data clustering, as demonstrated by extensive experiments on six real-world benchmark datasets. The source code is available at https://github.com/YuangXiao/ McMVC.
Yuang Xiao, Chang Tang, Weiqing Yan, Yuanyuan Liu 0004, Xinwang Liu 0002
IEEE Trans. Circuits Syst. Video Technol.4
2026 Wavelet Spectral-Spatial Mamba Network for Hyperspectral Image Classification
abstract
Spectral–spatial feature modeling plays a crucial role in hyperspectral image (HSI) classification. However, existing models based on convolutional neural networks (CNNs) and Transformers still face a trade-off between feature modeling capability and computational efficiency. Although recent wavelet-based HSI classification methods have demonstrated the advantages of frequency-domain analysis, they typically rely on a single wavelet basis, which limits their ability to capture diverse spectral–spatial patterns across different frequency bands. To address these issues, we propose Wavelet Spectral-Spatial Mamba (WSSMamba) by combining wavelet transform with state space modeling for HSI classification. WSSMamba introduces an Adaptive Wavelet Fusion Module (AWFM) to perform multi-scale frequency domain decomposition using multiple wavelet bases. This allows the model to extract both low-frequency global structure and high-frequency local details. A Wavelet Feature Enhancement (WFE) module is also designed to improve feature discriminability by applying channel and spatial attention mechanisms. Furthermore, we propose a Spectral-Spatial Cross-Fusion Strategy (SSCFS), which uses multi-directional state modeling to dynamically integrate high-frequency information. Extensive experiments on benchmark datasets demonstrate that WSSMamba outperforms state-of-the-art methods in classification performance.
Yongchao Song, Zhaowei Liu 0001, Weiqing Yan, Zengmao Wang, Xuan Wang 0021
IEEE Trans. Circuits Syst. Video Technol.4
2026 RCNet: Dual-Network Resonance Collaboration via Mutual Learning for RGB-D Road Defect Detection
Wujie Zhou, Zijun Ju, Runmin Cong, Weiqing Yan
IEEE Trans. Circuits Syst. Video Technol.4
2026 FD-SCU: Frequency Decomposition-Based Spectrum Collaborative Upsampling for Point Cloud Color Attribute
abstract
Existing point cloud color upsampling methods typically treat color upsampling as an interpolation problem within a local color or implicit feature domain. This largely overlooks the ability of the frequency domain to capture color correlations in local point sets. To address this limitation, we propose a spectrum collaborative strategy that uses frequency decomposition on voxel blocks (VBs) to enhance point cloud color reconstruction. We first voxelize the low-resolution (LR) color point cloud to generate multiple VBs and introduce a virtual filling strategy that adaptively assigns colors to empty voxels in each VB, ensuring that the irregularly distributed color information fully occupies the VB. We then apply the discrete cosine transform, known for its strong frequency-domain representation of locally smooth signals, to each color-filled VB to obtain frequency coefficients. These frequency coefficients are separated into high-frequency (HF) and low-frequency (LF) components. The LF coefficients, together with the LR color point cloud, are fed into a multi-scale cross-domain feature extraction module to capture deep features. Next, a Gaussian perturbation-based feature expansion generates upsampled color features, which are used to regress a coarse upsampled color point cloud. Finally, a high-frequency-guided residual refinement module uses the HF coefficients to refine the coarse upsampled result and produce a high-fidelity color point cloud. Extensive experiments demonstrate that our method achieves superior performance compared to state-of-the-art methods. Our code will be publicly available at https://github.com/wangwenchaoxx/FD-SCU.
Hao Liu 0044, Hui Yuan 0001, Raouf Hamzaoui, Weiqing Yan, Junhui Hou
IEEE Trans. Image Process.5
2026 Structure-Aware Conditional Diffusion Generation for Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering (IMVC) has attracted increasing attention in recent years, owing to the prevalence of missing data in real-world multi-view scenarios. Existing imputation-based IMVC methods partially mitigate the impact of missing information but still face three key limitations: (i) overlooking latent structural relationships among samples, which leads to imputed representations deviating from the true distribution; (ii) decoupling imputation from clustering, which reduces the discriminability of the recovered representations; and (iii) exhibiting low efficiency, which makes it difficult to balance recovery quality and inference speed under complex missing scenarios. To address these issues, we propose a Structure-Aware Conditional Diffusion Generation (SACDG) framework. During training, SACDG first models local structural relationships via adaptive neighborhood graphs and injects them as conditional priors into the diffusion model, where a cross-attention mechanism integrates these priors into the noise prediction process to learn structure-aware generative capability. Meanwhile, a semantic distribution alignment module is introduced to leverage pseudo-labels for enforcing cross-view consistency, thereby enhancing semantic discriminability. During inference, SACDG integrates cross-view structural information through cross-view adjacency fusion to guide the reverse denoising trajectory, and employs deterministic DDIM sampling to efficiently and stably recover the representations of missing views. Extensive comparative experiments and ablation studies on multiple benchmark datasets demonstrate that SACDG achieves superior clustering performance and improved efficiency over state-of-the-art methods. Our code is available athttps://github.com/zhangyuanyang21/SACDG.
Yuanyang Zhang, Yijie Lin 0001, Xinhang Wan, Jie Xu 0044, Li Yao 0003, Weiqing Yan, Chang Tang
IEEE Trans. Knowl. Data Eng.6
2026 Fetatrack: Foreground-aware dynamic template co-evolution for visual object tracking
Yibing Zhang, Xuan Wang 0021, Yongchao Song, Weiqing Yan, Aoran Wang
Vis. Comput.4
2025 Incomplete Multi-view Clustering via Diffusion Contrastive Generation
abstract
Incomplete multi-view clustering (IMVC) has garnered increasing attention in recent years due to the common issue of missing data in multi-view datasets. The primary approach to address this challenge involves recovering the missing views before applying conventional multi-view clustering methods. Although imputation-based IMVC methods have achieved significant improvements, they still encounter notable limitations: 1) heavy reliance on paired data for training the data recovery module, which is impractical in real scenarios with high missing data rates; 2) the generated data often lacks diversity and discriminability, resulting in suboptimal clustering results. To address these shortcomings, we propose a novel IMVC method called Diffusion Contrastive Generation (DCG). Motivated by the consistency between the diffusion and clustering processes, DCG learns the distribution characteristics to enhance clustering by applying forward diffusion and reverse denoising processes to intra-view data. By performing contrastive learning on a limited set of paired multi-view samples, DCG can align the generated views with the real views, facilitating accurate recovery of views across arbitrary missing view scenarios. Additionally, DCG integrates instance-level and category-level interactive learning to exploit the consistent and complementary information available in multi-view data, achieving robust and end-to-end clustering. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches.
Yuanyang Zhang, Yijie Lin 0001, Weiqing Yan, Li Yao 0003, Xinhang Wan, Guanzhou Ke, Jie Xu 0044
AAAI3
2025 IIBNet: Inter- and Intra-Information balance network with Self-Knowledge distillation and region attention for RGB-D indoor scene parsing
Fangfang Qiang, Wujie Zhou, Weiqing Yan, Lv Ye
Expert Syst. Appl.4
2025 Hybrid Knowledge Distillation for RGB-T Crowd Density Estimation in Smart Surveillance Systems
abstract
Crowd density estimation is a practical application task in which speed efficiency is as crucial as the accuracy of the results. Hence, we propose the hybrid knowledge distillation network (HKDNet) for RGB-thermal (RGB-T) crowd density estimation to address the limitations of computational cost and training time from the perspective of ensuring accuracy. We efficiently combine the advantages of traditional convolution and self-attention through a multimodal interactive transform. Subsequently, cross-graph convolution is extended to an interactive space for multimodal relationship reasoning. Finally, from the perspectives of channel and space, the crowd density map is obtained using channel synergy supplementation and spatial detail filling. In contrast to the computing resources required when using a heavyweight teacher network, the proposed HKDNet uses only approximately 6% of the parameters of the teacher network by abstracting the teacher network into a network hierarchy to generate a lightweight and efficient student network. Extensive experiments show that the proposed HKDNet performed well on two RGB-T crowd density estimation datasets. The code and models are available athttps://github.com/WBangG/HKDNet.
Wujie Zhou, Weiqing Yan, Qiuping Jiang
IEEE Internet Things J.3
2025 Multi-branch Space Sharing Feature Aggregation for contrastive multi-view clustering
Yuanyang Zhang, Weiqing Yan, Chang Tang, Wujie Zhou
Pattern Recognit.2
2025 PFCNet: Enhancing Rail Surface Defect Detection With Pixel-Aware Frequency Conversion Networks
abstract
Applying computer vision techniques to rail surface defect detection (RSDD) is crucial for preventing catastrophic accidents. However, challenges such as complex backgrounds and irregular defect shapes persist. Previous methods have focused on extracting salient object information from a pixel perspective, thereby neglecting valuable high- and low-frequency image information, which can better capture global structural information. In this study, we design a pixel-aware frequency conversion network (PFCNet) to explore RSDD from a frequency domain perspective. We use different attention mechanisms and frequency enhancement for high-level and shallow features to explore local details and global structures comprehensively. In addition, we design a dual-control reorganization module to refine the features across levels. We conducted extensive experiments on an industrial RGB-D dataset (NEU RSDDS-AUG), and PFCNet achieved superior performance. The code and results are publicly available at https://github.com/Wuyue15/PFCNet.
Fangfang Qiang, Wujie Zhou, Weiqing Yan
IEEE Signal Process. Lett.4
2025 MVoxTi-DNeRF: Explicit Multi-Scale Voxel Interpolation and Temporal Encoding Network for Efficient Dynamic Neural Radiance Field
abstract
Neural radiance fields have revolutionized the field of novel view synthesis, achieving remarkable results. However, traditional approaches based on implicit representations, particularly those built upon NeRF, suffer from slow rendering speeds due to the need for numerous MLP evaluations. Recently, there has been a promising shift towards explicit representations using voxel grids, which has significantly improved reconstruction times for static scenes. Nonetheless, extending these methods from static scenes to dynamic scenes is a non-trivial task as it requires accounting for the changing geometry and appearance of the scene over time. In this paper, we propose an efficient dynamic neural radiance field with multi-scale explicit voxel interpolation and temporal encoding. We leverage an explicit voxel structure to store the 3D dynamic features, while employing a lightweight MLP to estimate the displacement, thereby significantly enhancing the reconstruction speed. In our canonical module, we incorporated temporal information encoding in density estimation and color estimation to rectify the error estimation of the displacement in deformation module. In addition, a multi-scale voxel interpolation is designed to accommodate large-scale motions while meticulously capturing intricate details in small-scale motions in the density estimation module. In experiment, we evaluate our MVoxTi-DNeRF method on both synthetic and real scenes, where it achieves superior or comparable rendering quality compared to state of the art methods, while remaining computationally efficient (more than$60\times $faster than the original DNeRF). More experiment results and test code are available athttps://github.com/CHenYYff/MVoxTi-DNeRFNote to Practitioners—Neural Radiance Fields(NeRF) can create highly detailed and realistic 3D models from 2D images via a neural network. It is particularly popular for its capacity to generate novel views of a scene, enabling it to synthesize images from previously unobserved viewpoints. NeRF has found applications in a wide range of fields, including virtual reality, augmented reality, gaming, and so on. In this study, we introduces an efficient dynamic NeRF method. First, our network leverages an optimized explicit voxel grid to store 3D dynamic features and employs a lightweight MLP to decode these deformation features, significantly accelerating the training process. Second, to correct the error estimation related to deformation displacement, we introduce encoding of temporal information into density and color estimation in our canonical module, which fortifies the canonical field’s perception of temporal information. Third, we utilize multi-scale voxel interpolation to capture different-scale motion in the density estimation module, where minor motions are modeled using nearby voxels, while motion within a broader range is captured through more distant voxels. The consideration of multi-scale voxel features diminishes the detrimental impact caused by inaccurate displacement estimation.
Weiqing Yan, Yanshun Chen, Wujie Zhou, Runmin Cong
IEEE Trans Autom. Sci. Eng.1
2025 Progressive Feature Enhancement Network for Automated Colorectal Polyp Segmentation
abstract
In recent years, colorectal polyp segmentation has attracted increasing attention in academia and industry. Although most existing methods can achieve commendable outcomes, they often confront difficulty when localizing challenging polyps with complex background, variable shape/size, and ambiguous boundary, because of the limitations in modeling global context and in cross-layer feature interaction. To cope with these challenges, this paper proposes a novel Progressive Feature Enhancement Network (PFENet) for polyp segmentation. Specifically, PFENet follows an encoder-decoder structure and utilizes the pyramid vision transformer as the encoder to capture multi-scale long-term dependencies at different stages. A cross-stage feature enhancement (CFE) module is embedded in each stage. The CFE module enhances the feature representation ability from interaction among adjacent stages, which helps integrate scale information for recognizing polyps with complex background and variable shape/size. In addition, a foreground boundary co-enhancement (FBC) module is used at each decoder to simultaneously enhance the foreground and boundary information by incorporating the output of the adjacent high stage and the coarse segmentation map, which is generated by fusing features of all four stages via a coarse map generation module. Through top-down connections of FBC modules, PFENet can progressively refine the prediction in a coarse-to-fine manner. Extensive experiments show the effectiveness of our PFENet in the polyp segmentation task, with the mIoU and mDic values over 0.886 and 0.931 tested on two in-domain datasets and over 0.735 and 0.809 tested on three out-of-domain datasets.Note to Practitioners—Automated and accurate polyp segmentation in colonoscopy images is a critical prerequisite for subsequent detection, removal, and diagnosis of polyps in clinical practice. This paper proposes a novel deep neural network for polyp segmentation, termed PFENet, with a CFE module to enhance the feature representation ability for better capturing polyps with complex background and variable shape/size, and a FBC module to simultaneously enhance the foreground and boundary information on the feature representation provided by the CFE module. Qualitative and quantitative results on five public datasets show that our PFENet yields accurate predictions and is superior to 9 state-of-the-art polyp segmentation methods. The proposed PFENet will facilitate potential computer-aided diagnosis systems in clinical practice, in which it can better promote medical decision-making than competing methods in polyp detection and removal.
Guanghui Yue 0001, Houlu Xiao, Tianwei Zhou, Songbai Tan, Yun Liu 0009, Weiqing Yan
IEEE Trans Autom. Sci. Eng.6
2025 RDNet-KD: Recursive Encoder, Bimodal Screening Fusion, and Knowledge Distillation Network for Rail Defect Detection
abstract
Rail defect detection (RDD) plays a crucial role in ensuring rail transportation safety. Recently, bimodal algorithms have become mainstream; however, the asymmetry in the information of RGB and depth makes it difficult to find a suitable bimodal information fusion algorithm. In addition, it is difficult to deploy most of the existing methods on mobile devices. To solve these problems, we propose a recursive encoder and bimodal information screening fusion with a knowledge distillation network (RDNet-KD) for RDD. First, we propose the recursive encoder-based depth information augmentation (REDA) algorithm. It recursively learns to expand the channel depth information to alleviate the quality problem of depth information. Second, we propose a similarity-driven bimodal information screening fusion (SICF) module. This evaluates the complementarity of information from two modalities by computing the similarity of their hierarchical feature maps to screen useful information for fusion. Third, we introduce the global location and interrelation-based dual contextual knowledge distillation method to enhance the performance of the compact model. Therefore, it is possible to deploy the network on mobile devices. Based on the extensive experiments performed on the RGB-D rail defect dataset NEU RSDDS-AUG, we validate the competitiveness of our RDNet-KD, considering the prediction quality and operational efficiency relative to 12 state-of-the-art methods. The RDNet-KD code and results are available at https://github.com/legendfantasy/RDNet-KD.Note to Practitioners—This study introduces a recursive encoder and bimodal information screening fusion with a knowledge distillation network (RDNet-KD) for RDD in RGB-D images. Our method enhances depth information quality and effectively selects valuable information from both modalities using the similarity as a coefficient to evaluate the complementary capabilities of the modal information. Furthermore, to compress the model, we introduce knowledge distillation (KD) to balance the number of parameters and detection results and propose a novel KD method that transfers knowledge from the teacher network to the student network.
Wujie Zhou, Jinxin Yang, Weiqing Yan, Meixin Fang
IEEE Trans Autom. Sci. Eng.3
2025 PU-GSM: A Latent Geometry-Guided Self-Similarity Model for Point Cloud Upsampling
abstract
Existing point cloud upsampling methods typically treat upsampling as a local interpolation problem, neglecting the importance of global correlations within point sets, which can limit their performance. To address this limitation, we exploit the inherent self-similarity of point clouds from a global perspective and propose PU-GSM, a latent geometry-guided self-similarity model for upsampling. We first generate a lower-resolution sparse sub-point cloud (SPC) by downsampling the input point cloud (IPC). Then, we introduce a latent geometry-guided self-similarity model (LGSM) that learns a point distribution on the underlying surface of SPC by exploiting the inherent self-similarity of IPC. Next, we reuse the LGSM for the remaining points (i.e., the points left after removing SPC from IPC). Afterward, we introduce a gradient-aware dual domain refiner to generate and calibrate the upsampled point cloud from the learned point distribution. Finally, we propose an inference-free latent vector matching approach to regularize the upsampled point cloud by enhancing the feature similarity between the upsampled point cloud and the ground truth in latent space. Extensive experiments show that PU-GSM achieves better upsampling results compared to state-of-the-art methods. Our code will be available at: https://github.com/liuhaoyun/PU-GSM.
Hao Liu 0044, Hui Yuan 0001, Raouf Hamzaoui, Weiqing Yan
IEEE Trans. Circuits Syst. Video Technol.4
2025 Deep Incomplete Multi-View Clustering via Dynamic Imputation and Triple Alignment With Dual Optimization
abstract
In recent years, Incomplete Multi-View Clustering (IMVC) has become an important and challenging task. Although several methods have been proposed to address IMVC, they still have the following drawbacks: i) Due to the presence of missing samples in the views, clustering prototypes obtained from different views may have positional deviations, leading to inaccurate positioning of cluster centers, thus affecting the accuracy of clustering results. ii) Repair strategies based on cross-view prediction and adversarial generation have high computational costs and heavily rely on model performance. Neighbor-based repair strategies may result in inaccurate neighbor selection due to the presence of noise. iii) Models learned solely from complete data often perform better than models learned from both complete and incomplete data, especially when there are semantic differences between the repaired data and the missing data. To address the aforementioned issues, this paper proposes a Dynamic Imputation and Triple Alignment with Dual-Optimization for Deep Incomplete Multi-View Clustering (DITA-IMVC). Specifically, by accurately representing advanced features, we propose a cross-view dynamic structure learning strategy for missing view repair, where the dynamic structural relationships between high-semantic features within each view are calculated to obtain highly related samples from different views. To address positional deviations from different views, we propose a triple cross-view alignment with prototype, feature, and clustering assignment, which preserves the consistency among different views. Finally, we design a dual-optimization process for both complete view features and repaired features via alternating iterations to fully utilize the incomplete view data, thereby improving clustering performance. To demonstrate the effectiveness of our DITA-IMVC, extensive experiments conducted on different standard datasets show that it yields superior clustering results compared to existing methods.
Weiqing Yan, Kanglong Liu, Wujie Zhou, Chang Tang
IEEE Trans. Circuits Syst. Video Technol.1
2025 Lane Detection for Autonomous Driving: Comprehensive Reviews, Current Challenges, and Future Predictions
abstract
Lane detection is crucial for autonomous driving systems (ADS), utilizing sensors like cameras and LiDAR to identify lanes and understand vehicle position, direction, and lane shape. It provides data support for the control system to make informed driving decisions. In this survey, we review recent advancements in lane detection, focusing on both 2D techniques and emerging 3D methods. We begin with an overview of the significance of lane detection in ADS, followed by an analysis of the evolution of 2D techniques over the past decade, covering traditional and deep learning approaches. We also examine recent advancements in 3D lane detection. Additionally, we summarize evaluation metrics and popular datasets in the field. Finally, we discuss current challenges and future directions in lane detection, aiming to provide valuable insights for researchers and developers in this technology.
Jiping Bi, Yongchao Song, Yahong Jiang, Xuan Wang 0021, Zhaowei Liu 0001, Siwen Quan, Weiqing Yan
IEEE Trans. Intell. Transp. Syst.10
2025 Self-Supervised Semantic Soft Label Learning Network for Deep Multi-View Clustering
abstract
Multi-view clustering, which identifies shared semantics from different perspectives and classifies data samples into distinct categories using unsupervised methods, is gaining increasing interest. This task primarily focuses on learning consistent multi-view feature representations and clustering labels. Current approaches for achieving consistent multi-view feature representations often use techniques such as cascading, weight fusion, and attention mechanism fusion. These methods reconstruct features based on original low-level features via encoder-decoder, which often contain visual private information, leading to misleading feature representations. Furthermore, in the clustering label learning process, many methods use a two-stage approach: first, they achieve consistent feature representations, and then they apply hard labeling methods like K-means or spectral clustering to obtain clustering labels. Single-stage methods typically derive consistent labels through a linear coding layer based on consistent representation learning. These methods do not fully utilize the multi-view view semantic information, and consistent representation learning may be impaired when some low-quality views are present, leading to the generation of inaccurate semantic labels. To address these issues, we propose a Self-supervised Semantic Soft Label Learning Network for Deep Multi-view Clustering. Specifically, we introduce a consensus high-level feature learning module that uses a shared MLP layer to transform low-level features into a high-level feature space. To enhance the consistency between high-level features from different views, we maximize mutual information between these features and introduce the U-Projection module, which improves the expressive power of the consensus feature via resampling the features and concatenating the fused features before and after sampling operations. Additionally, we propose a self-supervised semantic label learning module that employs a dual-branch approach to independently learn consistent view-specific semantic labels through contrastive learning, while deriving view-consensus semantic labels from shared high-level features extracted from multiple views. Finally, KL divergence is used to align the view-consensus labels with the view-specific labels. A series of extensive experiments have shown that our approach yields superior clustering results compared to existing techniques.
Weiqing Yan, Tingyu Yang, Chang Tang
IEEE Trans. Multim.1
2025 Multiview Representation Learning via Information-Theoretic Optimization
abstract
Multiview data, characterized by rich features, are crucial in many machine learning applications. However, effectively extracting intraview features and integrating interview information present significant challenges in multiview learning (MVL). Traditional deep network-based approaches often involve learning multiple layers to derive latent. In these methods, the features of different classes are typically implicitly embedded rather than systematically organized. This lack of structure makes it challenging to explicitly map classes to independent principal subspaces in the feature space, potentially causing class overlap and confusion. Consequently, the capability of these representations to accurately capture the intrinsic structure of the data remains uncertain. In this article, we introduce an innovative multiview representation learning (MVRL) by maximizing two information-theoretic metrics: intraview coding rate reduction and interview mutual information. Specifically, in the intraview representation learning, we aim to optimize feature representations by maximizing the coding rate difference between the entire dataset and individual classes. This process expands the feature representation space while compressing the representations within each class, resulting in more compact feature representations within each viewpoint. Subsequently, we align and fuse these view-specific features through space transformation and cross-sample fusion to achieve consistent representation across multiple views. Finally, we maximize information transmission to maintain consistency and correlation among data representations across views. By maximizing mutual information between the consensus representations and view-specific representations, our method ensures that the learned representations capture more concise intrinsic features and correlations among different views, thereby enhancing the performance and generalization ability of MVL. Experiments show that the proposed methods have achieved excellent performance.
Weiqing Yan, Shuochen Yao, Chang Tang, Wujie Zhou
IEEE Trans. Neural Networks Learn. Syst.1
2025 Anchor-Sharing and Cluster-Wise Contrastive Network for Multiview Representation Learning
abstract
Multiview clustering (MVC) has gained significant attention as it enables the partitioning of samples into their respective categories through unsupervised learning. However, there are a few issues as follows: 1) many existing deep clustering methods use the same latent features to achieve the conflict objectives, namely, reconstruction and view consistency. The reconstruction objective aims to preserve view-specific features for each individual view, while the view-consistency objective strives to obtain common features across all views; 2) some deep embedded clustering (DEC) approaches adopt view-wise fusion to obtain consensus feature representation. However, these approaches overlook the correlation between samples, making it challenging to derive discriminative consensus representations; and 3) many methods use contrastive learning (CL) to align the view's representations; however, they do not take into account cluster information during the construction of sample pairs, which can lead to the presence of false negative pairs. To address these issues, we propose a novel multiview representation learning network, called anchor-sharing and clusterwise CL (CwCL) network for multiview representation learning. Specifically, we separate view-specific learning and view-common learning into different network branches, which addresses the conflict between reconstruction and consistency. Second, we design an anchor-sharing feature aggregation (ASFA) module, which learns the sharing anchors from different batch data samples, establishes the bipartite relationship between anchors and samples, and further leverages it to improve the samples' representations. This module enhances the discriminative power of the common representation from different samples. Third, we design CwCL module, which incorporates the learned transition probability into CL, allowing us to focus on minimizing the similarity between representations from negative pairs with a low transition probability. It alleviates the conflict in previous sample-level contrastive alignment. Experimental results demonstrate that our method outperforms the state-of-the-art performance.
Weiqing Yan, Yuanyang Zhang, Chang Tang, Wujie Zhou, Weisi Lin
IEEE Trans. Neural Networks Learn. Syst.1
2025 Hybrid Knowledge Distillation Network for RGB-D Co-Salient Object Detection
abstract
The aim of RGB-D Co-salient object detection (RGB-D Co-SOD) is to locate the most prominent objects within a provided collection of correlated RGB and depth images. The development of the Transformer has resulted in significant advancements in RGB-D Co-SOD. However, existing methods overlook the considerable computational and parametric costs associated with using the Transformer. Although compact models are computationally efficient, they suffer from performance degradation, which limits their practical applicability. This is because the reduction of model parameters weakens their feature representation capability. To bridge the performance gap between compact and complex models, we propose a hybrid knowledge distillation (KD) network, HKDNet-S*, to perform the RGB-D Co-SOD task. This method incorporates positive-negative logits approximation KD to guide the student network (HKDNet-S) in effectively learning the interrelationships among samples with multiple attributes by considering both positive and negative logits. HKDNet-S* primarily consists of the group cosaliency semantic exploration module and the positive and negative logits approximation KD method. Specifically, we employ a trained RGB-D Co-SOD model as a teacher model (HKDNet-T) to train the HKDNet-S with a limited number of participants using KD. Through extensive experiments on three challenging benchmark datasets (RGBD CoSal1k, RGBD CoSal150, and RGBD CoSeg183), we demonstrate that HKDNet-S* achieves superior accuracy while utilizing fewer parameters in comparison to the existing state-of-the-art methods.
Zhangping Tu, Wujie Zhou, Xiaohong Qian, Weiqing Yan
IEEE Trans. Syst. Man Cybern. Syst.4
2025 Knowledge Distillation SegFormer-Based Network for RGB-T Semantic Segmentation
abstract
Deep-learning-based semantic segmentation has received increasing research attention in recent years. However, owing to complex architectures, existing approaches have failed to achieve high accuracies in real-time applications. In this article, a novel knowledge distillation (KD) SegFormer-based network, called KDSNet-S*, is proposed to explore the tradeoff between accuracy and efficiency. Specifically, a structured KD scheme is designed to transfer the rich advanced features of a teacher network (KDSNet-T) to a student network (KDSNet-S). Thereafter, the KDSNet-S network learns the precise segmentation ability of the KDSNet-T network. Additionally, a multifield perceptual fusion model is proposed to learn more integrated features for a single modality and obtain discriminative and comprehensive feature representations. Furthermore, a high-level feature integration module is introduced to refine multimodality high-level features. Finally, multilevel features are fused, and a label-decoupling-based three-stream decoder that decomposes the original semantic segmentation map into center and contour diffusion maps for different supervision tasks is introduced. Experimental results on two public red-green–blue-thermal semantic segmentation datasets indicate the superiority of KDSNet-S* over compared state-of-the-art methods. The KDSNet-S* reduces parameters and floating-point operations per second by 91.1% and 81.9%, respectively, compared with the KDSNet-T. The source codes and results will be available athttps://github.com/purple-ting/KDSNet.
Wujie Zhou, Tingting Gong, Weiqing Yan
IEEE Trans. Syst. Man Cybern. Syst.3
2024 Progressive Adjacent-Layer coordination symmetric cascade network for semantic segmentation of Multimodal remote sensing images
Xiaomin Fan, Wujie Zhou, Xiaohong Qian, Weiqing Yan
Expert Syst. Appl.4
2024 CAGNet: Coordinated attention guidance network for RGB-T crowd counting
Wujie Zhou, Weiqing Yan, Xiaohong Qian
Expert Syst. Appl.3
2024 Adaptive multi-channel Bayesian Graph Neural Network
Zhaowei Liu 0001, Yingjie Wang 0002, Weiqing Yan
Neurocomputing5
2024 MJPNet-S*: Multistyle Joint-Perception Network With Knowledge Distillation for Drone RGB-Thermal Crowd Density Estimation in Smart Cities
abstract
Crowd density estimation has gained significant research interest owing to its potential in various industries and social applications. Therefore, this paper proposes a multistyle joint-perception network based on a knowledge distillation-trained student network (MJPNet-S*) for drone-based red–green–blue, thermal/depth (RGB-T/D) crowd density estimation tasks. To provide superior accuracy and efficiency, a novel trimodal working module effectively combines the modalities to facilitate comprehensive extraction and utilization. A two-step strategy comprising high-and low-level fusion is employed in which the high-level features capture relational reasoning and a one-dimensional projection relationship module captures multisensory field information with high-quality semantics. A shallow injection fusion module leverages the multiscale and channel relationships at the low level to combine full-text information interactively. Finally, to reduce resource consumption, a neighboring collaborative distillation method enables the lightweight student network to achieve superior performance by increasing the speed by 92 reducing the number of parameters by 83 of the teacher. Extensive experiments demonstrate that the proposed MJPNet-S* performs remarkably well on two RGB-T datasets. The code will be made public at https://github.com/WBangG/MJPNet.
Wujie Zhou, Xiena Dong, Meixin Fang, Weiqing Yan, Ting Luo 0001
IEEE Internet Things J.5
2024 Colorectal endoscopic image enhancement via unsupervised deep learning
Guanghui Yue 0001, Lvyin Duan, Jingfeng Du, Weiqing Yan, Shuigen Wang, Tianfu Wang 0001
Multim. Tools Appl.5
2024 Boundary uncertainty aware network for automated polyp segmentation
Guanghui Yue 0001, Guibin Zhuo, Weiqing Yan, Tianwei Zhou, Chang Tang, Peng Yang 0011, Tianfu Wang 0001
Neural Networks3
2024 Semantic Progressive Guidance Network for RGB-D Mirror Segmentation
abstract
Existing salient target detection methods tend to use a single-mirror segmentation strategy, which ignores feature hierarchy information in the frequency domain and lacks fine-grained correspondence. To address these challenges, we propose a new semantic progressive guidance network (SPGNet). To mine sufficient effective information, we propose the wavelet bidirectional focusing (WBF) module to aggregate sub-band features through a bidirectional wavelet transform and fuse them with low-level features to deepen the detail mining. We also introduce the Gaussian fusion complementary (GFC) module, which adopts Gaussian filtering technology to optimize the feature space and then efficiently extracts the contour information through enhanced feature processing. In addition, we propose a global correlation bootstrapping (GCB) module that constructs region-to-pixel correlations from a global perspective to achieve fine-grained correspondence. The proposed model achieves competitive results on a benchmark dataset.
Wujie Zhou, Weiqing Yan
IEEE Signal Process. Lett.4
2024 PiSFANet: Pillar Scale-Aware Feature Aggregation Network for Real-Time 3D Pedestrian Detection
abstract
Detecting 3D pedestrian from point cloud data in real-time while accounting for scale is crucial in various robotic and autonomous driving applications. Currently, the most successful methods for 3D object detection rely on voxel-based techniques, but these tend to be computationally inefficient for deployment in aerial scenarios. Conversely, the pillar-based approach exclusively employs 2D convolution, requiring fewer computational resources, albeit potentially sacrificing detection accuracy compared to voxel-based methods. Previous pillar-based approaches suffered from inadequate pillar feature encoding. In this letter, we introduce a real-time and scale-aware 3D Pedestrian Detection, which incorporates a robust encoder network designed for effective pillar feature extraction. The Proposed TriFocus Attention module (TriFA), which integrates external attention and similar attention strategies based on Squeeze and Exception. By comprehensively supervising the point-wise, channel-wise, and pillar-wise of pillar features, it enhances the encoding ability of pillars, suppresses noise in pillar features, and enhances the expression ability of pillar features. The proposed Bidirectional Scale-Aware Feature Pyramid module (BiSAFP) integrates a scale-aware module into the multi-scale pyramid structure. This addition enhances its ability to perceive pedestrian within low-level features. Moreover, it ensures that the significance of feature maps across various feature levels is fully taken into account. BiSAFP represents a lightweight multi-scale pyramid network that minimally impacts inference time while substantially boosting network performance. Our approach achieves real-time detection, processing up to 30 frames per second (FPS).
Weiqing Yan, Shile Liu, Chang Tang, Wujie Zhou
IEEE Signal Process. Lett.1
2024 CMPFFNet: Cross-Modal and Progressive Feature Fusion Network for RGB-D Indoor Scene Semantic Segmentation
abstract
Depth information can contribute to the semantic segmentation of scenes from red–green–blue (RGB) images. Therefore, the amount of information that can be obtained from RGB and RGB-depth (RGB-D) images is significantly greater for this task. However, RGB and RGB-D modalities are different in terms of object representation. Features that are extracted from these modalities and fused effectively are key to scene semantic segmentation. In addition, complete segmentation requires the fusion of multiscale features to unify global information. However, existing approaches primarily use multiscale features for sequential integration. This study introduces a cross-modal and progressive feature fusion network (CMPFFNet) for semantic segmentation of indoor scenes in RGB-D images. First, a multimodal adaptive alignment fusion (MAAF) module based on an attention mechanism is introduced. This module aligns the two modal channels by additive attention and then computes the spatial similarity between the two modalities based on the dot product to incorporate the complementary information of the depth modality into the RGB modality. In addition, a reverse attention augmentation (RAA) module is introduced to augment the more abstract high-level features for two adjacent multilevel features using the concrete semantic information of the lower-level features in them. After augmenting the extracted multilevel features, a multilevel feature progressive fusion (MFPF) module is deployed; this module sequentially fuses the neighboring two features progressively with emphasis on the spatial semantics. The network uses the Segformer network with high performance as a backbone in multiple computer vision tasks to enhance the segmentation capability. Experimental results obtained from two publicly available datasets of indoor scenes reveal that the proposed CMPFFNet outperforms existing models in semantic segmentation of indoor scenes of RGB-D images.Note to Practitioners—This study introduces a cross-modal and progressive feature fusion network (CMPFFNet) for indoor scene semantic segmentation in RGB-D images. The complementary information of the depth modality is incorporated into the RGB modality in both channel and spatial forms to form a discriminative representation for easy segmentation. A multilevel feature aggregation decoder is proposed to predict the results of semantic segmentation of scenes. The network uses the Segformer network with high performance as a backbone in multiple computer vision tasks to enhance the segmentation capability.
Wujie Zhou, Yuxiang Xiao, Weiqing Yan, Lu Yu 0003
IEEE Trans Autom. Sci. Eng.3
2024 Dual-Constraint Coarse-to-Fine Network for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) is an important yet challenging task, with great application values in industrial defect detection, medical care, etc. The challenges mainly come from the high intrinsic similarities between target objects and background. In this paper, inspired by the biological studies that object detection consists of two steps, i.e., search and identification, we propose a novel framework, named DCNet, for accurate COD. DCNet explores candidate objects and extra object-related edges through two constraints (object area and boundary) and detects camouflaged objects in a coarse-to-fine manner. Specifically, we first exploit an area-boundary decoder (ABD) to obtain initial region cues and boundary cues simultaneously by fusing multi-level features of the backbone. Then, an area search module (ASM) is embedded into each level of the backbone to adaptively search coarse regions of objects with the assistance of region cues from the ABD. After the ASM, an area refinement module (ARM) is utilized to identify fine regions of objects by fusing adjacent-level features with the guidance of boundary cues. Through the deep supervision strategy, DCNet can finally localize the camouflaged objects precisely. Extensive experiments on three benchmark COD datasets demonstrate that our DCNet is superior to 12 state-of-the-art COD methods. In addition, DCNet shows promising results on two COD-related tasks, i.e., industrial defect detection and polyp segmentation.
Guanghui Yue 0001, Houlu Xiao, Hai Xie, Tianwei Zhou, Wei Zhou 0021, Weiqing Yan, Baoquan Zhao, Tianfu Wang 0001, Qiuping Jiang
IEEE Trans. Circuits Syst. Video Technol.6
2024 Modal Evaluation Network via Knowledge Distillation for No-Service Rail Surface Defect Detection
abstract
Deep learning techniques have largely solved the problem of rail surface defect detection (SDD), however, two aspects have yet to be addressed. In most existing approaches, two red–green–blue and depth (RGB-D) streams are indiscriminately fused across modalities, ignoring the fact that RGB and depth images produce different feature qualities in different scenes. Additionally, in their focus on performance, previous studies have overlooked the fact that models produce several parameters, resulting in unrealistic practical applications. To address these challenges, we designed a modal evaluation network (MENet) via knowledge distillation (KD) (MENet-S*) for a no-service rail SDD to adaptively manage information in each scenario and achieve model compression. First, to dynamically adjust the feature distribution and quality, dynamic and static feature coding ideas are introduced. Second, modal evaluation distillation is introduced, which allows a compact model (MENet-S) to learn the feature evaluation process of a complex model (MENet-T). Third, to enable MENet-S to learn the dynamic encoding process of MENet-T and to improve the feature representation of MENet-S, we propose accessible knowledge distillation. Furthermore, multitiered KD is introduced to facilitate the learning of MENet-S. Based on extensive experiments using the industrial RGB-D dataset NEU RSDDS-AUG, we observed that MENet-S* (MENet-S with KD) outperformed 16 state-of-the-art methods. In addition, to demonstrate the generalization capability of MENet-S*, we evaluated the proposed network on three additional public datasets, and MENet-S* achieved competitive results. The source codes and results are available at https://github.com/hjklearn/MENet-KD.
Wujie Zhou, Jiankang Hong, Weiqing Yan, Qiuping Jiang
IEEE Trans. Circuits Syst. Video Technol.3
2024 Lightweight and Efficient Multimodal Prompt Injection Network for Scene Parsing of Remote Sensing Scene Images
abstract
Scene parsing of high-resolution remote sensing images with complex backgrounds has received extensive attention in recent years. As unimodal networks are significantly affected by weather conditions, reflecting complex ground conditions fully and accurately is difficult; therefore, multimodal scene analysis is particularly important. Current multimodal scene-parsing networks often employ a dual-coding architecture to achieve high-performance segmentation. Because prompt learning allows models to understand and capture contextual information more effectively, the proposed prompt injection module (PIM) extracts relevant information from frozen normalized digital surface model (nDSM) features and integrates it into the infrared, red, and green (IRRG) branches through a modal embedding block. To extract the contextual semantic relationships between the local and global features in the image efficiently, we also design a dynamic filter block for feature enhancement. This design facilitates the mutual complementarity and guidance of information between the two modalities and optimizes fusion. The experimental results demonstrate that lightweight and effective multimodal prompt injection network (LENet) outperforms most current state-of-the-art lightweight methods on two public datasets, achieving comparable accuracy to that of traditional methods. It has only 10.81 M parameters, with 2.72 GFLOPS. Our code and results are available athttps://github.com/LYZ00918/LENet.
Yangzhen Li, Wujie Zhou, Jiajun Meng, Weiqing Yan
IEEE Trans. Geosci. Remote. Sens.4
2024 EGFNet: Edge-Aware Guidance Fusion Network for RGB-Thermal Urban Scene Parsing
abstract
Urban scene parsing is the core of the intelligent transportation system, and RGB–thermal urban scene parsing has recently attracted increasing research interest in the field of computer vision. However, most existing approaches fail to perform good boundary extraction for prediction maps and cannot fully use high-level features. In addition, these methods simply fuse the features from RGB and thermal modalities but are unable to obtain comprehensive fused features. To address these problems, an edge-aware guidance fusion network (EGFNet) was developed in this study for RGB–thermal urban scene parsing. First, a prior edge map generated using the RGB and thermal images were introduced to capture detailed information in the prediction map and then embed the prior edge cues into the feature maps. To fuse the RGB and thermal information effectively, a multimodal fusion module was designed that guarantees adequate cross-modal fusion. Considering the importance of high-level semantic information, global and semantic information modules were proposed to extract rich semantic information from the high-level features. For decoding, simple elementwise addition was utilized for cascaded feature fusion. Finally, to improve the parsing accuracy, multitask deep supervision was applied to the semantic and boundary maps. Extensive experiments were performed on benchmark datasets to demonstrate the effectiveness of the proposed EGFNet and its superior performance compared with the state-of-the-art methods.
Shaohua Dong, Wujie Zhou, Caie Xu, Weiqing Yan
IEEE Trans. Intell. Transp. Syst.4
2024 DSANet-KD: Dual Semantic Approximation Network via Knowledge Distillation for Rail Surface Defect Detection
abstract
Owing to the development of convolutional neural networks (CNNs), the detection of defects on rail surfaces has significantly improved. Although existing methods achieve good results, they incur huge computational and parameter costs associated with CNNs. The usual approach to this problem is to design lightweight models that meet the needs of real-world applications; however, the performance is often compromised. To address the aforementioned problems, we designed a dual semantic approximation network via knowledge distillation (DSANet-KD, a student model with knowledge distillation) for rail surface defect detection; it focuses on both foreground and background knowledge and obtains more accurate prediction results. This model comprises an adaptive 3D spatial integration module, feature-optimization decoding module, and dual semantic approximation knowledge-distillation framework. Specifically, we employed a thoroughly trained teacher defect detection network equipped with dual semantic approximation information as an experienced teacher to guide the training of a student defect detection network. Experimental results showed that the proposed DSANet-KD achieved better accuracy with a smaller number of parameters than the state-of-the-art methods. To demonstrate the generalizability of DSANet-KD, experiments were conducted on a publicly available RGBD-SOD dataset, whose source code is available at: https://github.com/hjklearn/DSANet-KD.
Wujie Zhou, Jiankang Hong, Xiaoxiao Ran, Weiqing Yan, Qiuping Jiang
IEEE Trans. Intell. Transp. Syst.4
2024 MC3Net: Multimodality Cross-Guided Compensation Coordination Network for RGB-T Crowd Counting
abstract
Owing to the expansion in processing of industrial information through advances in machine learning, the demand for accurate crowd counting in various applications is increasing. We propose a multimodality cross-guided compensation coordination network (MC$^{3}$Net) for accurate red–green–blue and thermal (RGB-T) crowd counting. The network includes modules of intricate interactive fusion, feature difference compensation, and complementary attention enhancement. We use ConvNext as the backbone and process the three streams from RGB, thermal, and spliced RGB-T inputs. The multimodality data are sequentially guided and fused hierarchically, fully combining features extracted from the RGB and thermal images. Thereafter, difference compensation is applied to compress fusion and splicing features. Redundant information is removed. Then, feature mismatch is mitigated to enhance complementary information, reduce the loss of details, and finally obtain crowd statistics. Results from extensive experiments on the RGBT-CC dataset indicate the robustness and effectiveness of MC$^{3}$Net, which also achieves high performance on the DroneRGBT dataset and ShanghaiTechRGBD dataset, outperforming existing crowd counting methods. The code and models are available at: https://github.com/WBangG/MC3Net.
Wujie Zhou, Jingsheng Lei, Weiqing Yan, Lu Yu 0003
IEEE Trans. Intell. Transp. Syst.4
2024 Perceptual Quality Assessment of Retouched Face Images
abstract
Nowadays, it is a common practice to retouch face images before sharing them on websites, social media, and even identification cards. In response, increased criticisms have appeared about taking photo retouching to an extreme. This naturally leads to the necessity of designing perceptual quality assessment methods that can measure how much a retouched face image has strayed from reality. However, such an issue has seldom been considered. In this paper, we conduct both subjective and objective studies to advance this field. Firstly, we construct a benchmark database (termed SZU-RFD) via subjective experiments. SZU-RFD consists of 200 high-quality images with Asian faces and 1,600 retouched images generated by three popular photo-editing tools under different settings. Secondly, considering that retouching usually distorts the image texture, we propose a novel no-reference (NR) quality assessment method, named TANet, for retouched face images by taking the textural artifact into account. Specifically, a texture enhancement module is embedded into the shallow layer to help the network focus on textural information, and a multi-task learning strategy is applied to improve the performance of the main task with the assistance of an auxiliary task, i.e., texture recognition. Extensive experiments on the constructed SZU-RFD show that our proposed TANet correlates well with subjective perceptual judgments and is superior to 19 mainstream NR image quality assessment methods in evaluating retouched face images.
Guanghui Yue 0001, Honglv Wu, Qiuping Jiang, Tianwei Zhou, Weiqing Yan, Tianfu Wang 0001
IEEE Trans. Multim.5
2024 UTLNet: Uncertainty-Aware Transformer Localization Network for RGB-Depth Mirror Segmentation
abstract
Mirror segmentation, an emerging discipline in the field of computer vision, involves the identification and marking of mirrors in an image. Current mirror segmentation methods rely on fixed mirror elements as features for object segmentation. However, these methods do not account for the varied quality of feature images obtained under complex real-world conditions, leading to inaccurate segmentation results. To address these limitations, we propose a novel uncertainty-aware transformer localization network (UTLNet) for RGB-D mirror segmentation. Our approach draws inspiration from biomimicry, specifically the behavior pattern of human observation. We aim to explore features from different angles and focus on complex features that are challenging to determine during the coding stage. Additionally, we employ graph convolution to construct complementary dual-modal fusion features. Furthermore, we design a multiscale interaction transformer module using the shifted-window self-attention mechanism to acquire precise position information. In our experiments, the proposed UTLNet surpasses the current state-of-the-art mirror segmentation method as well as alternative task-specific methods. It achieves superior performance across various evaluation scenarios.
Wujie Zhou, Yuqi Cai, Weiqing Yan, Lu Yu 0003
IEEE Trans. Multim.4
2023 GCFAgg: Global and Cross-View Feature Aggregation for Multi-View Clustering
abstract
Multi-view clustering can partition data samples into their categories by learning a consensus representation in unsupervised way and has received more and more attention in recent years. However, most existing deep clustering methods learn consensus representation or view-specific representations from multiple views via view-wise aggregation way, where they ignore structure relationship of all samples. In this paper, we propose a novel multi-view clustering network to address these problems, called Global and Cross-view Feature Aggregation for Multi-View Clustering (GCFAggMVC). Specifically, the consensus data presentation from multiple views is obtained via cross-sample and cross-view feature aggregation, which fully explores the complementary of similar samples. Moreover, we align the consensus representation and the view-specific representation by the structure-guided contrastive learning module, which makes the view-specific representations from different samples with high structure relationship similar. The proposed module is a flexible multi-view data representation module, which can be also embedded to the incomplete multi-view data clustering task via plugging our module into other frameworks. Extensive experiments show that the proposed method achieves excellent performance in both complete multi-view data clustering tasks and incomplete multi-view data clustering tasks.
Weiqing Yan, Yuanyang Zhang, Chenlei Lv, Chang Tang, Guanghui Yue 0001, Weisi Lin
CVPR1
2023 Deep Video Stabilization via Robust Homography Estimation
abstract
Video stabilization can improve the visual quality of videos that have been captured on mobile devices or other handheld cameras, which are more prone to shaking and motion artifacts. Most of the existing deep video stabilization methods adopts optical flow-based, which produce artifacts and distortions caused by pixel-level warping and enquire expensive computation time. In this paper, we present a novel unsupervised deep video stabilization approach that addresses the influence of moving objects on video stabilization through robust homography estimation. Specifically, we design a foreground mask estimation module as a preprocessing step using a pre-trained semantic segmentation guided method to distinguish the foreground and background regions, enabling us to estimate camera motion via analyzing the background motion. Additionally, we design a low-level confidence feature extraction module to improve motion alignment loss and ensure robust motion estimation. By integrating the learned low-level confidence features with the foreground mask, we can then design a motion estimation module that captures the consistent spatial correspondence between frames through local and global feature extraction. At last, the learnt robust homography is leveraged to stabilize videos. Our method outperforms related state-of-the-art approaches in both quality and quantity on three public benchmarks while remaining computationally efficient.
Weiqing Yan, Yiqiu Sun 0001, Wujie Zhou, Zhaowei Liu 0001, Runmin Cong
IEEE Signal Process. Lett.1
2023 Heterogeneous Network Representation Learning Approach for Ethereum Identity Identification
abstract
Recently, network representation learning has been widely used to mine and analyze network characteristics, and it is also applied to blockchain, but most of the embedding methods in blockchain ignore the heterogeneity of network, so it is difficult to accurately describe the characteristics of the transaction. As smart society evolves, Ethereum makes smart contracts reality, while the mine of transaction characteristics appearing on the Ethereum platform is scarce; thus, there is an urgent need to mine Ethereum from contract and transfer. In this article, we propose a heterogeneous network representation learning method to mine implicit information inside Ethereum transactions. Specifically, we construct an Ethereum transaction network by collecting transaction data from normal and phishing Ethereum accounts. Then, we propose a walk strategy that combines timestamps and transaction amounts to represent the information that occurs at the time of a transaction. To mine the types of nodes and edges, we use a heterogeneous network representation learning method to map the transaction network to a low-dimensional space. Finally, we improve the accuracy of the embedding results in the node classification task, which has important implications for Ethereum mining as well as identity recognition.
Zhaowei Liu 0001, Weiqing Yan
IEEE Trans. Comput. Soc. Syst.4
2023 MMSMCNet: Modal Memory Sharing and Morphological Complementary Networks for RGB-T Urban Scene Semantic Segmentation
abstract
Combining color (RGB) images with thermal images can facilitate semantic segmentation of poorly lit urban scenes. However, for RGB-thermal (RGB-T) semantic segmentation, most existing models address cross-modal feature fusion by focusing only on exploring the samples while neglecting the connections between different samples. Additionally, although the importance of boundary, binary, and semantic information is considered in the decoding process, the differences and complementarities between different morphological features are usually neglected. In this paper, we propose a novel RGB-T semantic segmentation network, called MMSMCNet, based on modal memory fusion and morphological multiscale assistance to address the aforementioned problems. For this network, in the encoding part, we used SegFormer for feature extraction of bimodal inputs. Next, our modal memory sharing module implements staged learning and memory sharing of sample information across modal multiscales. Furthermore, we constructed a decoding union unit comprising three decoding units in a layer-by-layer progression that can extract two different morphological features according to the information category and realize the complementary utilization of multiscale cross-modal fusion information. Each unit contains a contour positioning module based on detail information, a skeleton positioning module with deep features as the primary input, and a morphological complementary module for mutual reinforcement of the first two types of information and construction of semantic information. Based on this, we constructed a new supervision strategy, that is, a multi-unit-based complementary supervision strategy. Extensive experiments using two standard datasets showed that MMSMCNet outperformed related state-of-the-art methods. The code is available at:https://github.com/2021nihao/MMSMCNet.
Wujie Zhou, Weiqing Yan, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.3
2023 Graph Attention Guidance Network With Knowledge Distillation for Semantic Segmentation of Remote Sensing Images
abstract
Deep learning has become a popular method for studying the semantic segmentation of high-resolution remote sensing images (HRRSIs). Existing methods have adopted convolutional neural networks to achieve better segmentation accuracy of HRRSIs, and the success of these models often depends on the model complexity and parameter quantity. However, the deployment of these models on equipment with limited resources is a significant challenge. To solve this problem, a lightweight student network framework—a graph attention guidance network (GAGNet) with knowledge distillation, called GAGNet-S*—is proposed in this study, which distills knowledge from pretrained large teacher network (GAGNet-T) and builds reliable weak labels to optimize untrained student network (GAGNet-S). Inspired by the graph convolution network, this study designs a graph convolution module called the attention-graph decoder, which combines attention mechanisms with graph convolution to optimize image features and improve segmentation accuracy in the semantic segmentation task of HRRSIs. In addition, a dense cross-decoder was designed for multiscale dense fusion, which utilizes rich semantic information in the high-level features to guide and refine the low-level features from the bottom up. Extensive experiments showed that GAGNet-S* (GAGNet-S with knowledge distillation) achieved excellent segmentation performance on two widely used datasets: Potsdam and Vaihingen. The code and models are available at https://github.com/F8AoMn/GAGNet-KD.
Wujie Zhou, Xiaomin Fan, Weiqing Yan, Shengdao Shan, Qiuping Jiang, Jenq-Neng Hwang
IEEE Trans. Geosci. Remote. Sens.3
2023 GSGNet-S*: Graph Semantic Guidance Network via Knowledge Distillation for Optical Remote Sensing Image Scene Analysis
abstract
In recent years, optical remote sensing image (ORSI) scene analysis has attracted increasing interest. However, existing networks show a trend of bifurcation. Lightweight networks have very high inference speed but poor inference of contextual information in highly complex backgrounds. In contrast, networks with high-performance contextual information reasoning capability require many parameters and are computationally expensive. Since the knowledge distillation method can greatly lighten the model, we propose a graph semantic guided network (GSGNet) that utilizes knowledge refinement for ORSI scenario analysis, which has a high inference speed while maintaining practical contextual inference capability. Rich semantic and detailed information facilitates semantic segmentation of optical remote sensing images. We design adjacent dynamic capture and local-global map inference modules that can effectively extract low-level spatial details and high-level contextual semantics. To improve the attention map relearning performance of the distillation method, we designed semantically guided fusion modules to locate spatial information and refine edge information. We also employed a structural relationship transfer distillation method in which the structural relationship knowledge of the teacher model (GSGNet-T) was used to guide the student model (GSGNet-S). We compared the performances of GSGNet-T and the GSGNet-S with knowledge distillation (GSGNet-S*) with those of several state-of-the-art methods on the Vaihingen and Potsdam datasets. Extensive experiments showed that GSGNet-S* outperformed most advanced methods with only 19.61M parameters and a computation cost of 2.9G FLOPs. The experimental results and code of our network can be accessed at the following URL: https://github.com/LYZ00918/GSGNet-KD.
Wujie Zhou, Yangzhen Li, Weiqing Yan, Meixin Fang, Qiuping Jiang
IEEE Trans. Geosci. Remote. Sens.4
2023 Benchmarking Polyp Segmentation Methods in Narrow-Band Imaging Colonoscopy Images
abstract
In recent years, there has been significant progress in polyp segmentation in white-light imaging (WLI) colonoscopy images, particularly with methods based on deep learning (DL). However, little attention has been paid to the reliability of these methods in narrow-band imaging (NBI) data. NBI improves visibility of blood vessels and helps physicians observe complex polyps more easily than WLI, but NBI images often include polyps with small/flat appearances, background interference, and camouflage properties, making polyp segmentation a challenging task. This paper proposes a new polyp segmentation dataset (PS-NBI2K) consisting of 2,000 NBI colonoscopy images with pixel-wise annotations, and presents benchmarking results and analyses for 24 recently reported DL-based polyp segmentation methods on PS-NBI2K. The results show that existing methods struggle to locate polyps with smaller sizes and stronger interference, and that extracting both local and global features improves performance. There is also a trade-off between effectiveness and efficiency, and most methods cannot achieve the best results in both areas simultaneously. This work highlights potential directions for designing DL-based polyp segmentation methods in NBI colonoscopy images, and the release of PS-NBI2K aims to drive further development in this field.
Guanghui Yue 0001, Guibin Zhuo, Tianwei Zhou, Jingfeng Du, Weiqing Yan, Jingwen Hou, Weide Liu, Tianfu Wang 0001
IEEE J. Biomed. Health Informatics6
2023 Toward Multicenter Skin Lesion Classification Using Deep Neural Network With Adaptively Weighted Balance Loss
abstract
Recently, deep neural network-based methods have shown promising advantages in accurately recognizing skin lesions from dermoscopic images. However, most existing works focus more on improving the network framework for better feature representation but ignore the data imbalance issue, limiting their flexibility and accuracy across multiple scenarios in multi-center clinics. Generally, different clinical centers have different data distributions, which presents challenging requirements for the network's flexibility and accuracy. In this paper, we divert the attention from framework improvement to the data imbalance issue and propose a new solution for multi-center skin lesion classification by introducing a novel adaptively weighted balance (AWB) loss to the conventional classification network. Benefiting from AWB, the proposed solution has the following advantages: 1) it is easy to satisfy different practical requirements by only changing the backbone; 2) it is user-friendly with no tuning on hyperparameters; and 3) it adaptively enables small intraclass compactness and pays more attention to the minority class. Extensive experiments demonstrate that, compared with solutions equipped with state-of-the-art loss functions, the proposed solution is more flexible and more competent for tackling the multi-center imbalanced skin lesion classification task with considerable performance on two benchmark datasets. In addition, the proposed solution is proved to be effective in handling the imbalanced gastrointestinal disease classification task and the imbalanced DR grading task. Code is available at https://github.com/Weipeishan2021.
Guanghui Yue 0001, Peishan Wei, Tianwei Zhou, Qiuping Jiang, Weiqing Yan, Tianfu Wang 0001
IEEE Trans. Medical Imaging5
2022 Efficient Multiple Kernel Clustering via Spectral Perturbation
abstract
Clustering is a fundamental task in the machine learning and data mining community. Among existing clustering methods, multiple kernel clustering (MKC) has been widely investigated due to its effectiveness to capture non-linear relationships among samples. However, most of the existing MKC methods bear intensive computational complexity in learning an optimal kernel and seeking the final clustering partition. In this paper, based on the spectral perturbation theory, we propose an efficient MKC method that reduces the computational complexity from O(n3) to O(nk2 + k3), with n and k denoting the number of data samples and the number of clusters, respectively. The proposed method recovers the optimal clustering partition from base partitions by maximizing the eigen gaps to approximate the perturbation errors. An equivalent optimization objective function is introduced to obtain base partitions. Furthermore, a kernel weighting scheme is embedded to capture the diversity among multiple kernels. Finally, the optimal partition, base partitions, and kernel weights are jointly learned in a unified framework. An efficient alternate iterative optimization algorithm is designed to solve the resultant optimization problem. Experimental results on various benchmark datasets demonstrate the superiority of the proposed method when compared to other state-of-the-art ones in terms of both clustering efficacy and efficiency.
Chang Tang, Zhenglai Li, Weiqing Yan, Guanghui Yue 0001, Wei Zhang 0049
ACM Multimedia3
2022 Bipartite Graph-based Discriminative Feature Learning for Multi-View Clustering
abstract
Multi-view clustering is an important technique in machine learning research. Existing methods have improved in clustering performance, most of them learn graph structure depending on all samples, which are high complexity. Bipartite graph-based multi-view clustering can obtain clustering result by establishing the relationship between the sample points and small anchor points, which improve the efficiency of clustering. Most bipartite graph-based clustering methods only focus on topological graph structure learning depending on sample nodes, ignore the influence of node features. In this paper, we propose bipartite graph-based discriminative feature learning for multi-view clustering, which combines bipartite graph learning and discriminative feature learning to a unified framework. Specifically, the bipartite graph learning is proposed via multi-view subspace representation with manifold regularization terms. Meanwhile, our feature learning utilizes data pseudo-labels obtained by fused bipartite graph to seek projection direction, which make the same label be closer and make data points with different labels be far away from each other. At last, the proposed manifold regularization terms establish the relationship between constructed bipartite graph and new data representation. By leveraging the interactions between structure learning and discriminative feature learning, we are able to select more informative features and capture more accurate structure of data for clustering. Extensive experimental results on different scale datasets demonstrate our method achieves better or comparable clustering performance than the results of state-of-the-art methods.
Weiqing Yan, Jinglei Liu, Guanghui Yue 0001, Chang Tang
ACM Multimedia1
2022 Sparse LiDAR and Binocular Stereo Fusion Network for 3D Object Detection
Weiqing Yan, Kaiqi Su, Jinlai Ren, Runmin Cong, Shuigen Wang
PRCV (3)1
2022 Single image dehazing using generative adversarial networks based on an attention mechanism
abstract
Abstract Most existing image dehazing methods rely on the solution of the atmospheric scattering model or supervised learning based on paired images. However, owing to incomplete prior knowledge and the lack of paired hazy and haze‐free images of the same scenes as training samples, their performances for single image dehazing are unsatisfactory. Here, the authors present an unpaired image learning method based on the attention mechanism for single image dehazing problems. The method uses the constraint transfer learning ability and circulatory structure of CycleGAN to carry out an unsupervised image dehazing task for unpaired data. Considering the complexity of the haze distribution in actual imaging and human visual characteristics, the improved channel attention and domain attention mechanisms are integrated into the network to process different features and different regions non‐uniformly. The experimental results show that the proposed method achieves good results on both synthetic datasets and real hazy images.
Yongli Ma, Fei Jia, Weiqing Yan, Zhaowei Liu 0001, Mengying Ni
IET Image Process.4
2022 Stereo VoVNet-CNN for 3D object detection
Kaiqi Su, Weiqing Yan, Meiqi Gu
Multim. Tools Appl.2
2020 Shape-optimizing mesh warping method for stereoscopic panorama stitching
Weiqing Yan, Guanghui Yue 0001, Yanwei Yu, Kai Wang 0014, Chang Tang, Xiangrong Tong
Inf. Sci.1
2020 Perceptual objective quality assessment of stereoscopic stitched images
Weiqing Yan, Guanghui Yue 0001, Yuming Fang 0001, Hua Chen 0004, Chang Tang, Gangyi Jiang
Signal Process.1
2020 Referenceless Quality Evaluation of Tone-Mapped HDR and Multiexposure Fused Images
abstract
Nowadays, the standard dynamic range (SDR) image acquired at a fixed exposure exposes weakness in portraying fine-grained details of real scenes. The high dynamic range (HDR) image and other types of SDR images generated by multiexposure fusion techniques provide us new choices for scene representation. To display on SDR screens, an HDR image must be tone-mapped to an SDR one. Since different tone-mapping/fusion algorithms produce images with varying visual quality levels, it naturally desires a quality evaluation model for comparison. This article proposes an effective model in the absence of the reference image. By analyzing the characteristics of tone-mapped HDR and multiexposure fused images, we first extract multiple quality-sensitive features from the following aspects: 1) colorfulness; 2) exposure; and 3) naturalness. Then, the model is built by bridging all extracted features and associated subjective ratings via support vector regression. Extensive experiments on publicly available databases prove the superiority of our model over the state-of-the-art referenceless quality evaluation ones.
Guanghui Yue 0001, Weiqing Yan, Tianwei Zhou
IEEE Trans. Ind. Informatics2
2019 Diversity and consistency learning guided spectral embedding for multi-view clustering
Zhenglai Li, Chang Tang, Jiajia Chen 0010, Weiqing Yan, Xinwang Liu 0002
Neurocomputing5
2017 Content-aware disparity adjustment for different stereo displays
Weiqing Yan, Chunping Hou, Baoliang Wang, Laihua Wang
Multim. Tools Appl.1
2017 Stereoscopic Image Stitching Based on a Hybrid Warping Model
abstract
Traditional image editing techniques cannot be directly used to process stereoscopic media, as extra constraints are required to ensure consistent changes between the left and right images. In this paper, we propose a hybrid warping model for stereoscopic image stitching by combining projective and content-preserving warping. First, a uniform homography algorithm is proposed to prewarp the left and right images, and thus ensure consistent changes. Second, a content-preserving warping is introduced to locally refine alignment and reduce vertical disparities. Finally, a seam-cutting-based algorithm is used to find a blending seam, and the multiband blending algorithm is used to produce the final stitched image. Experimental results show that the proposed method can effectively stitch stereoscopic images, which not only avoids local distortions, but also reduces vertical disparities reasonably.
Weiqing Yan, Chunping Hou, Jianjun Lei 0001, Yuming Fang 0001, Zhouye Gu, Nam Ling
IEEE Trans. Circuits Syst. Video Technol.1
2015 View generation with DIBR for 3D display system
Laihua Wang, Chunping Hou, Jianjun Lei 0001, Weiqing Yan
Multim. Tools Appl.4