VLDB 2026 Research / reviewers in the wild / expert
Fangfang Qiang
dblp:196/2879
· DBLP profile ↗
9ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0003-2256-3851ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-stream interaction network with cross-modal contrast distillation for co-salient object detection
Wujie Zhou, Bingying Wang, Xiena Dong, Caie Xu, Fangfang Qiang |
Signal Process. Image Commun. | 5 |
| 2026 | Lightweight Scope Integration Network for Rail Surface Defect DetectionabstractRail-surface defect detection (RSDD) is a key technology for ensuring the safety and efficiency of railroad transportation. Existing models enhance the robustness in complex scenarios by using complementary information from visible light (RGB) image and depth data. However, incorporating depth data increases the computational complexity, which consumes more resources, increases the risk of overfitting, and decreases the model efficiency. Optimizing and streamlining multimodal data processing remains challenging. To address this issue and enable fast and accurate RSDD, this study introduces a novel lightweight scope integration network (LSINet). This method efficiently and precisely fuses the features using a lightweight double-context-aware module and refines the image quality using an adaptive Markov field smoothing module in a hierarchical decoding process. A lightweight two-dimensional scanning method that captures long-range dependencies and improves computational efficiency. We integrate this approach with local range extremes to enhance the multimodal feature fusion. Furthermore, to refine the defect edges and improve the detection accuracy, we integrate the Markov random field module into the defect segmentation and optimization using its ability to model interpixel spatial correlation, particularly for small and ambiguous defects. In designing the model, we focused on optimizing the number of parameters (e.g., computational resources). Experimental evaluations using various rail surface images demonstrated that the proposed method enhanced the defect detection accuracy and recall while reaching fast detection speeds. According to extensive experiments conducted using the industrial RGB-D dataset (i.e., NEU RSDDS-AUG), LSINet outperformed 15 state-of-the-art methods with only 27.6 million parameters. Code and results are publicly available at https://github.com/Wuyue15/LSINet. Wujie Zhou, Fangfang Qiang, Weiqing Yan |
IEEE Trans. Big Data | 3 |
| 2025 | IIBNet: Inter- and Intra-Information balance network with Self-Knowledge distillation and region attention for RGB-D indoor scene parsing
Fangfang Qiang, Wujie Zhou, Weiqing Yan, Lv Ye |
Expert Syst. Appl. | 2 |
| 2025 | AESeg: Affinity-enhanced segmenter using feature class mapping knowledge distillation for efficient RGB-D semantic segmentation of indoor scenes
Wujie Zhou, Yuxiang Xiao, Fangfang Qiang, Xiena Dong, Caie Xu, Lu Yu 0003 |
Neural Networks | 3 |
| 2025 | PFCNet: Enhancing Rail Surface Defect Detection With Pixel-Aware Frequency Conversion NetworksabstractApplying computer vision techniques to rail surface defect detection (RSDD) is crucial for preventing catastrophic accidents. However, challenges such as complex backgrounds and irregular defect shapes persist. Previous methods have focused on extracting salient object information from a pixel perspective, thereby neglecting valuable high- and low-frequency image information, which can better capture global structural information. In this study, we design a pixel-aware frequency conversion network (PFCNet) to explore RSDD from a frequency domain perspective. We use different attention mechanisms and frequency enhancement for high-level and shallow features to explore local details and global structures comprehensively. In addition, we design a dual-control reorganization module to refine the features across levels. We conducted extensive experiments on an industrial RGB-D dataset (NEU RSDDS-AUG), and PFCNet achieved superior performance. The code and results are publicly available at https://github.com/Wuyue15/PFCNet. Fangfang Qiang, Wujie Zhou, Weiqing Yan |
IEEE Signal Process. Lett. | 2 |
| 2025 | Depth Enhancement Mask Mapping Network With Multi-Teacher Distillation for RGB-D Scene ParsingabstractThe proposed method in the study, called DEMMNet-KD, aims to address the challenges of scene parsing in RGB-D indoor scenes. The study recognizes the importance of lightweight models for pixel-intensive tasks like scene parsing, and proposes a depth enhancement mask mapping network (DEMMNet) with multi-teacher knowledge distillation approach to achieve both reduced model size and high accuracy. DEMMNet introduces a new segmentation paradigm called class-position separate segmentation to reduce the difficulty of delineating and classifying the segmentation region. It also includes a depth spatial enhancement module based on the attention mechanism to effectively utilize depth information in a single-stream network design. To improve performance without increasing complexity, DEMMNet-KD utilizes Mix Transformer encoder (MiT) as a backbone and employs multi-teacher knowledge distillation. This allows semantic knowledge to be transferred from large teacher networks with different performances to the student network. The experiments conducted on NYUDv2 and SUN RGB-D datasets demonstrate that DEMMNet-KD achieves state-of-the-art performance with fewer additional learnable parameters and minimal computational effort. Utilizing MiT-B2 as the backbone network, the DEMMNet-KD method attained mean Intersection over Union (mIoU) scores of 54.71% and 51.03% on the two datasets, respectively, thereby achieving an optimal balance between accuracy and efficiency. Furthermore, the scalability of this method has been demonstrated across various domains and modalities. Overall, the proposed DEMMNet-KD approach provides an efficient solution for scene parsing in RGB-D indoor scenes, enhancing the capabilities of bionic binocular robots in perceiving their environments and serving human society effectively. The code is available at https://github.com/SHARKALAKALA/DEMMNet. Yuxiang Xiao, Jiajun Meng, Fangfang Qiang, Xiena Dong, Wujie Zhou |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | Graph Enhancement and Transformer Aggregation Network for RGB-Thermal Crowd CountingabstractCrowd counting has received significant attention in recent years due to its practical applications. In order to address the specific characteristics of RGB and thermal images, we have developed the graph enhancement and transformer aggregation network (GETANet) for generating representative density maps. Our approach incorporates several innovative modules to enhance accuracy. Firstly, we introduced a position-adaptive module that effectively counts individuals’ positions and integrates features extracted from the main framework. Furthermore, we leveraged the advantages of graph convolutional networks (GCNs), which integrate spatial information and exploit relationships between nodes. Specifically, we designed a dual GCN module that further improves the model’s performance by considering the spatial context and relationships among individuals in the crowd. To capture global image information and improve overall performance, we integrated a vision transformer into our model architecture. The vision transformer effectively captures global dependencies and enhances the model’s ability to understand complex crowd scenes. Additionally, we designed a transformer information aggregation module that integrates information from multiple levels, resulting in a highly precise prediction map. Through comprehensive experiments on benchmark datasets such as RGBT-CC and DroneRGBT, our GETANet demonstrated its effectiveness in RGB-thermal crowd counting tasks. Moreover, GETANet showcased remarkable generalization results on the ShanghaiTech-RGBD dataset. Our code has been made publicly available on GitHub at https://github.com/panyi95/GETANet. Wujie Zhou, Meixin Fang, Fangfang Qiang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | STONet-S*: A Knowledge-Distilled Approach for Semantic Segmentation in Remote Sensing ImagesabstractSemantic segmentation of remote sensing images is a critical research domain. The integration of cross-modal features enhances stability in intricate environments. Despite the impressive performance of existing methods, their complexity and parameter demands remain significant. Our proposed STONet-S$^{\ast }$, a stepped transmission optimization network (STONet) with knowledge distillation (KD), extracts insights from a pretrained extensive teacher network and transfers them to an untrained compact student network. Initially, a group enhancement and interaction unit (GEIU) correct for background noise influence and seamlessly integrates cross-modal features. Additionally, we introduce a stepped transmission decoder (STD) comprising a stepped capture module (SCM) and a self-reverse revision module (SRRM) to capture multiscale information from the ground up. Furthermore, leveraging the frequency domain, we employ frequency-awareness KD using a discrete cosine transform (DCT) and octave convolution to separate high and low-frequency maps, which are subsequently transferred to the student network. Last, detail-delivery and stepped-response KD (SRKD) mechanisms enhance the learning capacity of the student network. Through extensive experimentation on two datasets, STONet-S$^{\ast }$demonstrates superior segmentation accuracy by achieving remarkable results with only 7.19 M parameters. The corresponding code repository can be accessed at:https://github.com/MAXHAN22/STONet. Wujie Zhou, Penghan Yang, Weiwei Qiu, Fangfang Qiang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Morphology-Guided Network via Knowledge Distillation for RGB-D Mirror SegmentationabstractMirror segmentation is an emerging computer vision task that is extensively applied in various fields. However, it presents significant challenges to existing segmentation methods when irregular shapes are involved. Most methods are designed for deployment on heavy-duty host machines that demand substantial computational resources and storage capacity, which limits their feasibility for deployment on mobile devices, where efficient and resource-friendly solutions are required. Therefore, we propose a morphology-guided network (MGNet) with knowledge distillation, called MGNet-S*, to achieve the efficiency required for deployment in mobile devices. In this network, we introduce an erosion dilation fusion module that leverages morphological knowledge to extract texture details from intrinsic features. This module incorporates different optimization strategies for multimodal features. Furthermore, it provides a knowledge-distillation framework specifically tailored to the proposed MGNet-S*. The MGNet-S* includes three effective distillation modules: a semi-soft label, misaligned features, and adaptive aggregation types. These modules facilitate the efficient transfer of knowledge from the MGNet teacher to MGNet student, allowing the lightweight network, MGNet-S*, to achieve remarkable performance. Numerous experiments proved that our proposed MGNet-S* outperformed state-of-the-art methods, achieving an 88.6% reduction in parameter count and 82.5% reduction in floating-point operations compared to those of the MGNet teacher network. Wujie Zhou, Yuqi Cai, Fangfang Qiang |
IEEE Trans. Intell. Transp. Syst. | 3 |