VLDB 2026 Research / reviewers in the wild / expert
Meixin Fang
dblp:297/2701
· DBLP profile ↗
12ranked-venue papers
0as first author
12since 2021 · last 2025
0009-0006-1327-4416ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RDNet-KD: Recursive Encoder, Bimodal Screening Fusion, and Knowledge Distillation Network for Rail Defect DetectionabstractRail defect detection (RDD) plays a crucial role in ensuring rail transportation safety. Recently, bimodal algorithms have become mainstream; however, the asymmetry in the information of RGB and depth makes it difficult to find a suitable bimodal information fusion algorithm. In addition, it is difficult to deploy most of the existing methods on mobile devices. To solve these problems, we propose a recursive encoder and bimodal information screening fusion with a knowledge distillation network (RDNet-KD) for RDD. First, we propose the recursive encoder-based depth information augmentation (REDA) algorithm. It recursively learns to expand the channel depth information to alleviate the quality problem of depth information. Second, we propose a similarity-driven bimodal information screening fusion (SICF) module. This evaluates the complementarity of information from two modalities by computing the similarity of their hierarchical feature maps to screen useful information for fusion. Third, we introduce the global location and interrelation-based dual contextual knowledge distillation method to enhance the performance of the compact model. Therefore, it is possible to deploy the network on mobile devices. Based on the extensive experiments performed on the RGB-D rail defect dataset NEU RSDDS-AUG, we validate the competitiveness of our RDNet-KD, considering the prediction quality and operational efficiency relative to 12 state-of-the-art methods. The RDNet-KD code and results are available at https://github.com/legendfantasy/RDNet-KD.Note to Practitioners—This study introduces a recursive encoder and bimodal information screening fusion with a knowledge distillation network (RDNet-KD) for RDD in RGB-D images. Our method enhances depth information quality and effectively selects valuable information from both modalities using the similarity as a coefficient to evaluate the complementary capabilities of the modal information. Furthermore, to compress the model, we introduce knowledge distillation (KD) to balance the number of parameters and detection results and propose a novel KD method that transfers knowledge from the teacher network to the student network. Wujie Zhou, Jinxin Yang, Weiqing Yan, Meixin Fang |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Asymmetrical Contrastive Learning Network via Knowledge Distillation for No-Service Rail Surface Defect DetectionabstractOwing to extensive research on deep learning, significant progress has recently been made in trackless surface defect detection (SDD). Nevertheless, existing algorithms face two main challenges. First, while depth features contain rich spatial structure features, most models only accept red-green-blue (RGB) features as input, which severely constrains performance. Thus, this study proposes a dual-stream teacher model termed the asymmetrical contrastive learning network (ACLNet-T), which extracts both RGB and depth features to achieve high performance. Second, the introduction of the dual-stream model facilitates an exponential increase in the number of parameters. As a solution, we designed a single-stream student model (ACLNet-S) that extracted RGB features. We leveraged a contrastive distillation loss via knowledge distillation (KD) techniques to transfer rich multimodal features from the ACLNet-T to the ACLNet-S pixel by pixel and channel by channel. Furthermore, to compensate for the lack of contrastive distillation loss that focuses exclusively on local features, we employed multiscale graph mapping to establish long-range dependencies and transfer global features to the ACLNet-S through multiscale graph mapping distillation loss. Finally, an attentional distillation loss based on the adaptive attention decoder (AAD) was designed to further improve the performance of the ACLNet-S. Consequently, we obtained the ACLNet-S*, which achieved performance similar to that of ACLNet-T, despite having a nearly eightfold parameter count gap. Through comprehensive experimentation using the industrial RGB-D dataset NEU RSDDS-AUG, the ACLNet-S* (ACLNet-S with KD) was confirmed to outperform 16 state-of-the-art methods. Moreover, to showcase the generalization capacity of ACLNet-S*, the proposed network was evaluated on three additional public datasets, and ACLNet-S* achieved comparable results. The code is available at https://github.com/Yuride0404127/ACLNet-KD. Wujie Zhou, Xiaohong Qian, Meixin Fang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | MJPNet-S*: Multistyle Joint-Perception Network With Knowledge Distillation for Drone RGB-Thermal Crowd Density Estimation in Smart CitiesabstractCrowd density estimation has gained significant research interest owing to its potential in various industries and social applications. Therefore, this paper proposes a multistyle joint-perception network based on a knowledge distillation-trained student network (MJPNet-S*) for drone-based red–green–blue, thermal/depth (RGB-T/D) crowd density estimation tasks. To provide superior accuracy and efficiency, a novel trimodal working module effectively combines the modalities to facilitate comprehensive extraction and utilization. A two-step strategy comprising high-and low-level fusion is employed in which the high-level features capture relational reasoning and a one-dimensional projection relationship module captures multisensory field information with high-quality semantics. A shallow injection fusion module leverages the multiscale and channel relationships at the low level to combine full-text information interactively. Finally, to reduce resource consumption, a neighboring collaborative distillation method enables the lightweight student network to achieve superior performance by increasing the speed by 92 reducing the number of parameters by 83 of the teacher. Extensive experiments demonstrate that the proposed MJPNet-S* performs remarkably well on two RGB-T datasets. The code will be made public at https://github.com/WBangG/MJPNet. Wujie Zhou, Xiena Dong, Meixin Fang, Weiqing Yan, Ting Luo 0001 |
IEEE Internet Things J. | 4 |
| 2024 | Graph Enhancement and Transformer Aggregation Network for RGB-Thermal Crowd CountingabstractCrowd counting has received significant attention in recent years due to its practical applications. In order to address the specific characteristics of RGB and thermal images, we have developed the graph enhancement and transformer aggregation network (GETANet) for generating representative density maps. Our approach incorporates several innovative modules to enhance accuracy. Firstly, we introduced a position-adaptive module that effectively counts individuals’ positions and integrates features extracted from the main framework. Furthermore, we leveraged the advantages of graph convolutional networks (GCNs), which integrate spatial information and exploit relationships between nodes. Specifically, we designed a dual GCN module that further improves the model’s performance by considering the spatial context and relationships among individuals in the crowd. To capture global image information and improve overall performance, we integrated a vision transformer into our model architecture. The vision transformer effectively captures global dependencies and enhances the model’s ability to understand complex crowd scenes. Additionally, we designed a transformer information aggregation module that integrates information from multiple levels, resulting in a highly precise prediction map. Through comprehensive experiments on benchmark datasets such as RGBT-CC and DroneRGBT, our GETANet demonstrated its effectiveness in RGB-thermal crowd counting tasks. Moreover, GETANet showcased remarkable generalization results on the ShanghaiTech-RGBD dataset. Our code has been made publicly available on GitHub at https://github.com/panyi95/GETANet. Wujie Zhou, Meixin Fang, Fangfang Qiang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | DMFTNet: dense multimodal fusion transfer network for free-space detection
Jiabao Ma, Wujie Zhou, Meixin Fang, Ting Luo 0001 |
Multim. Syst. | 3 |
| 2024 | PGGNet: Pyramid gradual-guidance network for RGB-D indoor scene semantic segmentation
Wujie Zhou, Gao Xu, Meixin Fang, Shanshan Mao, Rongwang Yang, Lu Yu 0003 |
Signal Process. Image Commun. | 3 |
| 2024 | Lightweight Dual Stream Network With Knowledge Distillation for RGB-D Scene ParsingabstractSignificant progress has been made in the field of indoor scene parsing. The increasing demand for lightweight networks is due to the limited hardware capacity of mobile devices. However, there has been a lack of research on the design of lightweight networks for indoor scene parsing. Therefore, we propose lightweight dual stream network (LDSNet) with knowledge distillation (KD) for RGB-D indoor scene parsing. Initially, we developed a two-stream network with three versions (LDSNet-tiny*, LDSNet-small*, and LDSNet-base, where * represents the model after KD) for different scenarios. In the main stream, we designed an integrated joint enhancement module that captures valuable information from both RGB and depth features. This information is then processed by the cascading integration module to generate the final map. To improve the performance of the model, we included an auxiliary extraction module in the auxiliary stream to specifically extract feature information for KD. During the training process, we used hierarchical context loss to distill features and obtain LDSNet-tiny* and LDSNet-small*. We conducted experiments on the NYUDv2 and SUN RGB-D datasets, which demonstrated that our LDSNet-base achieves superior results, while LDSNet-tiny* and LDSNet-small* also exhibit satisfactory performance. Wujie Zhou, Xiaoxiao Ran, Meixin Fang |
IEEE Signal Process. Lett. | 4 |
| 2024 | DGPINet-KD: Deep Guided and Progressive Integration Network With Knowledge Distillation for RGB-D Indoor Scene AnalysisabstractSignificant advancements in RGB-D semantic segmentation have been made owing to the increasing availability of robust depth information. Most researchers have combined depth with RGB data to capture complementary information in images. Although this approach improves segmentation performance, it requires excessive model parameters. To address this problem, we propose DGPINet-KD, a deep-guided and progressive integration network with knowledge distillation (KD) for RGB-D indoor scene analysis. First, we used branching attention and depth guidance to capture coordinated, precise location information and extract more complete spatial information from the depth map to complement the semantic information for the encoded features. Second, we trained the student network (DGPINet-S) with a well-trained teacher network (DGPINet-T) using a multilevel KD. Third, an integration unit was developed to explore the contextual dependencies of the decoding features and to enhance relational KD. Comprehensive experiments on two challenging indoor benchmark datasets, NYUDv2 and SUN RGB-D, demonstrated that DGPINet-KD achieved improved performance in indoor scene analysis tasks compared with existing methods. Notably, on the NYUDv2 dataset, DGPINet-KD (DGPINet-S with KD) achieves a pixel accuracy gain of 1.7% and a class accuracy gain of 2.3% compared with DGPINet-S. In addition, compared with DGPINet-T, the proposed DGPINet-KD (DGPINet-S with KD) utilizes significantly fewer parameters (29.3M) while maintaining accuracy. The source code is available at https://github.com/XUEXIKUAIL/DGPINet. Wujie Zhou, Bitao Jian, Meixin Fang, Xiena Dong, Yuanyuan Liu 0004, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Transmission Line Detection Through Bi-Directional Guided Registration With Knowledge DistillationabstractTransmission line (TL) inspection plays a crucial role in maintaining a reliable electricity supply to all regions. Computer vision methods, especially those utilizing infrared images, have achieved significant advancements in this field. However, many existing multimodal fusion methods utilize conventional attention mechanisms or simple meta-additions to combine different modalities without proper alignment. Moreover, these methods often rely on a large number of parameters to achieve better performance. To enhance the fusion of disparate modalities and minimize model parameters, we propose bidirectional guided registration via knowledge distillation (BGRNet-S*) for RGB-T transmission line detection (TLD). This approach incorporates a bidirectional registration mechanism within the fusion module and achieves parameter reduction through our knowledge distillation (KD) method. Based on the bidirectional guidance of the non-local position encoding module, accurate feature registration between modes can be achieved. Additionally, we designed response distillation and spatial semantic distillation for our student network (BGRNet-S). Extensive experiments on TLD datasets demonstrate that both our BGRNet-T and BGRNet-S*(BGRNet-S with KD) achieve excellent performance using state-of-the-art methods. When using Shunted-B and Shunted-T as the backbones of the BGRNet-T and BGRNet-S*networks, respectively, the number of parameters was reduced from 52.34M to 13.09M, and the calculated floating-point numbers decreased from 25.23G to 8.93G. Wujie Zhou, Chuanming Ji, Meixin Fang |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | AMCFNet: Asymmetric multiscale and crossmodal fusion network for RGB-D semantic segmentation in indoor service robots
Wujie Zhou, Yuchun Yue, Meixin Fang, Shanshan Mao, Rongwang Yang, Lu Yu 0003 |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | CSANet: Contour and Semantic Feature Alignment Fusion Network for Rail Surface Defect DetectionabstractRail surface defect detection for traffic safety has received considerable attention. With the development of deep learning, numerous methods for combining RGB and depth information have been proposed. However, these methods directly fuse raw features extracted from the backbone, which can lead to ineffective use of the complementary information of the two modalities. In this study, we developed a contour and semantic feature alignment fusion network (CSANet) with bidirectional feature alignment to explore the internal consistency of cross-modal features from both contour and semantic perspectives. First, an adjacency contour feature extraction module was designed to capture high-quality contour information from adjacent low-level features. Second, an attention-aware graph convolution embedded semantic feature extraction module was designed to explore long-range dependencies and extract semantic information. Third, a bidirectional alignment mechanism was designed to explore the internal consistency of contours and semantics between bimodal features. Experimental results on the industrial RGB-D dataset (NEU RSDDS-AUG) revealed that the proposed CSANet outperformed 12 state-of-the-art algorithms in four evaluation metrics. Jinxin Yang, Wujie Zhou, Ruiming Wu, Meixin Fang |
IEEE Signal Process. Lett. | 4 |
| 2023 | GSGNet-S*: Graph Semantic Guidance Network via Knowledge Distillation for Optical Remote Sensing Image Scene AnalysisabstractIn recent years, optical remote sensing image (ORSI) scene analysis has attracted increasing interest. However, existing networks show a trend of bifurcation. Lightweight networks have very high inference speed but poor inference of contextual information in highly complex backgrounds. In contrast, networks with high-performance contextual information reasoning capability require many parameters and are computationally expensive. Since the knowledge distillation method can greatly lighten the model, we propose a graph semantic guided network (GSGNet) that utilizes knowledge refinement for ORSI scenario analysis, which has a high inference speed while maintaining practical contextual inference capability. Rich semantic and detailed information facilitates semantic segmentation of optical remote sensing images. We design adjacent dynamic capture and local-global map inference modules that can effectively extract low-level spatial details and high-level contextual semantics. To improve the attention map relearning performance of the distillation method, we designed semantically guided fusion modules to locate spatial information and refine edge information. We also employed a structural relationship transfer distillation method in which the structural relationship knowledge of the teacher model (GSGNet-T) was used to guide the student model (GSGNet-S). We compared the performances of GSGNet-T and the GSGNet-S with knowledge distillation (GSGNet-S*) with those of several state-of-the-art methods on the Vaihingen and Potsdam datasets. Extensive experiments showed that GSGNet-S* outperformed most advanced methods with only 19.61M parameters and a computation cost of 2.9G FLOPs. The experimental results and code of our network can be accessed at the following URL: https://github.com/LYZ00918/GSGNet-KD. Wujie Zhou, Yangzhen Li, Weiqing Yan, Meixin Fang, Qiuping Jiang |
IEEE Trans. Geosci. Remote. Sens. | 5 |