VLDB 2026 Research / reviewers in the wild / expert
Wen-Liang Du 0002
dblp:214/9676 · also Wenliang Du 0002
· DBLP profile ↗
26ranked-venue papers
5as first author
24since 2021 · last 2026
0000-0002-9234-0912ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 4 first-author · 9 since 2021Computer networks · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLIPDet3D: Vision-Language Collaborative Distillation for 3D Object DetectionabstractMulti-view 3D object detection plays a vital role in autonomous driving systems due to its ability to perceive complex scenes accurately. However, real-world driving data often exhibits a long-tailed distribution, causing significant drops in detection accuracy for rare categories in existing methods. To mitigate this issue, we propose CLIPDet3D, a novel vision-language collaborative framework for multi-view 3D object detection. First, to tackle the difficulty of capturing the semantic information of rare categories, a Vision-Language Collaborative Learning strategy is proposed to incorporate class-level semantic priors from CLIP. Second, a Depth Feature Contrastive Distillation module is designed to overcome the large depth estimation error for rare categories by aligning depth features between a teacher and a student network. Furthermore, to alleviate the difficulty in focusing on regions of rare categories, a Dual-Stream Prompt Attention mechanism is devised to inject learnable prompts and compute attention along both horizontal and vertical BEV directions. Evaluations on the nuScenes dataset demonstrate that CLIPDet3D achieves state-of-the-art accuracy while maintaining efficient inference. Jiaqi Zhao 0001, Huanfeng Hu, Yong Zhou 0003, Wen-Liang Du 0002, Kunyang Sun, Rui Yao 0006, Qigong Sun |
AAAI | 4 |
| 2026 | Unified Representation Causal Prompt Distillation for Re-Inference-Free Lifelong Person Re-IdentificationabstractLifelong person re-identification (LReID) aims to retrieve the target person from sequentially collected data. Due to significant domain gaps between datasets and the continuous increase of training data from different scenarios, weak inter-domain generalization and catastrophic forgetting issues have remained major challenges for LReID. To tackle these issues, a novel LReID method called Unified Representation Causal Prompt Distillation (URCPD) is proposed. Specifically, to reduce domain gaps among different scene datasets and improve model inter-domain generalization capability, a Feature Decoupling Style Transfer module (FDST) is proposed to map new features into a unified feature space. Furthermore, to reduce the accumulated forgetting of old knowledge during the training stage, a Causal Prompt Distillation module (CPD) is introduced. This module eliminates the re-inference process for distillation and embeds memory prompts to combat catastrophic forgetting. Extensive experiments on five classic LReID seen datasets and seven unseen datasets demonstrate that our method significantly outperforms state-of-the-art methods. Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006 |
AAAI | 4 |
| 2026 | Causal Decoupling Domain Generalization for Remote Sensing Change DetectionabstractWhile current state-of-the-art Remote Sensing Change Detection (RSCD) methods can achieve impressive results on individual datasets, they become unreliable in unseen environments and imaging conditions, with performance metrics declining by as much as 60% to 80%. Simultaneously, variable environments and complex imaging conditions are the main characteristics of remote sensing data, calling for generalizable RSCD methods. To address this issue, we propose a novel RSCD method capable of domain generalization—CDDGNet. This method is based on causal decoupling theory, which progressively decouples invariant change features from variable domain features to extract generalizable characteristics. This enables a network trained on a single domain to accurately identify change regions in other domains. Specifically, firstly, the Causal Feature Adaptation Module is proposed to preliminarily decouple and simplify feature information during the encoding process by using wavelet transformation and feature energy spectralization methods. Secondly, the Causal Feature Fusion Module is presented to fully decouple features and aggregate significant change features during the decoding process through frequency domain processing and feature re-attention mechanisms. Thirdly, the Decoupling Effect Loss Function is proposed to optimize the process by evaluating the effectiveness of causal decoupling. Extensive experiments have shown that our model significantly outperforms existing methods across multiple groups of generalization tasks with varying levels of difficulty. Jiaqi Zhao 0001, Jianpeng Xie 0001, Yong Zhou 0003, Wen-Liang Du 0002, Hancheng Zhu, Rui Yao 0006 |
AAAI | 4 |
| 2026 | SpaceFormer: Spatial Position Contextual Semantics Embedding for Multi-View 3D Object Detectionabstract3D object detection aims to accurately localize and recognize objects in 3D space. It serves as a fundamental task for reliable perception in intelligent transportation systems, enabling the monitoring of diverse traffic participants such as vehicles, pedestrians, cyclists, and public transport. Recently, transformer-based methods have gained significant attention in multi-view 3D object detection due to their strong global reasoning capabilities. However, their limited capacity to model spatial positional information hinders accurate object localization, especially in complex and large-scale scenes. To address this limitation, SpaceFormer is proposed as a novel transformer-based multi-view 3D object detector. Specifically, a Contextual Visual Prompts Learning strategy is proposed to enhance the perception of small and sparse traffic participants by incorporating contextual priors. To further suppress background interference, a Semantics-guided Depth Estimation method is proposed to refine depth representations using high-level semantic information. Furthermore, a Spatial Position Embedding mechanism is proposed to improve the spatial localization capability of the transformer by integrating geometric position and polar spatial embedding. Extensive experiments on the nuScenes benchmark demonstrate that SpaceFormer achieves state-of-the-art performance with 55.5% mAP and 62.9% NDS. These improvements indicate not only methodological advances but also practical benefits for intelligent transportation systems, enhancing safety, reliability, and efficiency in real-world deployments. Jiaqi Zhao 0001, Huanfeng Hu, Wen-Liang Du 0002, Yong Zhou 0003, Kunyang Sun, Rui Yao 0006, Abdulmotaleb El Saddik |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | ReDiffDet: Rotation-equivariant Diffusion Model for Oriented Object DetectionabstractThe diffusion model has been successfully applied to various detection tasks. However, it still faces several challenges when used for oriented object detection: objects that are arbitrarily rotated require the diffusion model to encode their orientation information; uncontrollable random boxes inaccurately locate objects with dense arrangements and extreme aspect ratios; oriented boxes result in the misalignment between them and image features. To overcome these limitations, we propose ReDiffDet, a framework that formulates oriented object detection as a rotation-equivariant denoising diffusion process. First, we represent an oriented box as a 2D Gaussian distribution, forming the basis of the denoising paradigm. The reverse process can be proven to be rotation-equivariant within this representation and model framework. Second, we design a conditional encoder with conditional boxes to prevent boxes from being randomly placed across the entire image. Third, we propose an aligned decoder for alignment between oriented boxes and image features. The extensive experiments demonstrate ReDiffDet achieves promising performance and significantly outperforms the diffusion-based baseline detector. Codes are available at https://github.com/wokaikaixinxin/ReDiffDet. Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006 |
CVPR | 5 |
| 2025 | GSDet: Gaussian Splatting for Oriented Object DetectionabstractOriented object detection has advanced with the development of convolutional neural networks (CNNs) and transformers. However, modern detectors still rely on predefined object candidates, such as anchors in CNN-based methods or queries in transformer-based methods, which struggle to capture spatial information effectively. To address the limitations, we propose GSDet, a novel framework that formulates oriented object detection as Gaussian splatting. Specifically, our approach performs detection within a 3D feature space constructed from image features, where 3D Gaussians are employed to represent oriented objects. These 3D Gaussians are projected onto the image plane to form 2D Gaussians, which are then transformed into oriented boxes. Furthermore, we optimize the mean, anisotropic covariance, and confidence scores of these randomly initialized 3D Gaussians, using a decoder that incorporates 3D Gaussian sampling. Moreover, our method exhibits flexibility, enabling adaptive control and a dynamic number of Gaussians during inference. Experiments on 3 datasets indicate that GSDet achieves AP50 gains of 0.7% on DIOR-R, 0.3% on DOTA-v1.0, and 0.55% on DOTA-v1.5 when evaluated with adaptive control and outperforms mainstream detectors. Zeyu Ding 0010, Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Hancheng Zhu, Rui Yao 0006 |
IJCAI | 4 |
| 2025 | Beyond Individual and Point: Next POI Recommendation via Region-aware Dynamic Hypergraph with Dual-level ModelingabstractNext POI recommendation contributes to the prosperity of various intelligent location-based services. Existing studies focus on exploring sequential patterns and POI interactions using sequential and graph-based methods to enhance recommendation performance. However, they don't effectively exploit geographical information. In addition, methods that focus on modeling mobility patterns using individual limited data may suffer from data sparsity and the information cocoons problem. Moreover, most graph structures focus on adjacent nodes, failing to capture potential high-order associations among POIs. To address these challenges, we propose the Region-aware dynamic Hypergraph learning method with Dual-level interaction Modeling (ReHDM), which exploits users' dynamic mobility beyond individual and point. Specifically, ReHDM utilizes regional encoding to mine the potential spatial relationships among POIs with coarse-grained geographical information. By incorporating POI-level and trajectory-level associations within a hypergraph convolutional network, ReHDM comprehensively captures cross-user collaborative information. Furthermore, ReHDM captures not only dependencies among POIs within each trajectory for a single user, but also the high-order collaborative information across individual user trajectories and associated users' trajectories. Experimental results on three public datasets demonstrate the superiority of ReHDM to the state-of-the-art. Zhuo Gu, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Wen-Liang Du 0002 |
IJCAI | 7 |
| 2025 | Counterfactual Knowledge Maintenance for Unsupervised Domain AdaptationabstractTraditional unsupervised domain adaptation (UDA) struggles to extract rich semantics due to backbone limitations. Recent large-scale pre-trained visual-language models (VLMs) have shown strong zero-shot learning capabilities in UDA tasks. However, directly using VLMs results in a mixture of semantic and domain-specific information, complicating knowledge transfer. Complex scenes with subtle semantic differences are prone to misclassification, which in turn can result in the loss of features that are crucial for distinguishing between classes. To address these challenges, we propose a novel counterfactual knowledge maintenance UDA framework. Specifically, we employ counterfactual disentanglement to separate the representation of semantic information from domain features, thereby reducing domain bias. Furthermore, to clarify ambiguous visual information specific to classes, we maintain the discriminative knowledge of both visual and textual information. This approach synergistically leverages multimodal information to preserve modality-specific distinguishable features. We conducted extensive experimental evaluations on several public datasets to demonstrate the effectiveness of our method. The source code is available at https://github.com/LiYaolab/CMKUDA Yong Zhou 0003, Jiaqi Zhao 0001, Wen-Liang Du 0002, Rui Yao 0006, Bing Liu 0016 |
IJCAI | 4 |
| 2025 | RQFormer: Rotated Query Transformer for end-to-end oriented object detection
Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik |
Expert Syst. Appl. | 5 |
| 2025 | MoViM: A Hybrid CNN Vision Mamba Network for Lightweight Semantic Segmentation of Multimodal Remote Sensing ImagesabstractThe “Others” category in multimodal remote sensing images is characterized by high intra-class variability. Therefore, existing lightweight semantic segmentation models struggle with this category due to limitations in capturing both local details and global dependencies efficiently. We propose MoViM, a lightweight model that integrates a hybrid Vision Mamba (ViM) and CNN backbone to capture global contextual information and local details effectively. In addition, the MoViM also features an Inverted Stem for efficient multimodal fusion, a Global Semantics Extraction (GSE) module for enhanced global feature representation, and a Global-Local Feature Fusion (GLF) module for context-aware feature integration. Extensive experiments on WHU-OPT-SAR and Potsdam datasets demonstrate that MoViM achieves state-of-the-art performance, particularly in the “Others” category, while maintaining low computational complexity. Our codes are available at https://github.com/WenliangDu/MoViM. Wen-Liang Du 0002, Jiaqi Zhao 0001, Rui Yao 0006, Yong Zhou 0003 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | FA-MSVNet: multi-scale and multi-view feature aggregation methods for stereo 3D reconstruction
Yong Zhou 0003, Jiaqi Zhao 0001, Wen-Liang Du 0002, Rui Yao 0006 |
Multim. Tools Appl. | 4 |
| 2025 | Grid-distance-based selection for fine-grained object detection in aerial images
Jiaqi Zhao 0001, Qingfeng Ou, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006 |
Pattern Recognit. Lett. | 4 |
| 2025 | DDCI: Unsupervised Domain Adaptation for Remote Sensing Images Based on Diffusion Causal DistillationabstractThe distribution of remote sensing (RS) images can vary significantly due to seasonal changes and lighting conditions, making it difficult for deep learning models to generalize effectively across different RS datasets. This variation leads to a domain gap that hampers model performance when applied to new, unseen data. To tackle this challenge, we introduce DDCI, a novel unsupervised domain adaptation (UDA) framework designed to bridge the domain gap in RS image perception. Our framework consists of two key components, i.e., the adaptation diffusion distillation (ADD) module and the consistent causal intervention (CCI) module. The ADD module addresses the domain gap by aligning the source and target domains. It enhances the representation of the target domain by distilling semantic knowledge from the teacher model of the source domain. This process allows the target domain to benefit from the rich features of the source domain, leading to improved model generalization. The CCI module focuses on removing spurious correlations between domain-agnostic knowledge and domain-specific knowledge. By carefully considering the distinct characteristics of the target domain while preserving the specificity of the source domain, the CCI module ensures that only relevant, causal information is transferred between domains. This prevents overfitting to irrelevant domain-specific features and enhances model robustness. We demonstrate the effectiveness of the DDCI framework on RS scene classification tasks, utilizing four widely recognized RS datasets. Our results show significant performance improvements, underscoring the potential of this approach to boost the adaptability of deep learning models across diverse RS image datasets. Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | GLFRNet: Global-Local Feature Refusion Network for Remote Sensing Image Instance SegmentationabstractInstance segmentation is a significant way for remote sensing image (RSI) interpretation. The large number, sharp variation of sizes, and complex background of objects raise higher demands for instance segmentation models. The synergistic usage of global and local features has drawn great attention due to its superior performance but has not been fully explored in mainstream instance segmentation methods. In this work, a global-local feature refusion network (GLFRNet) with two fusion procedures is proposed to fully utilize coarse-grained and fine-grained features for RSI instance segmentation. In this model, the backbone integrates both convolutional neural network (CNN)-based and VMamba-based branches to extract local and global features, respectively. Three novel models are proposed to leverage the features adaptively, i.e., the cross-dim feature fusion (CDFF) module, the semantic complementary feature fusion (SCFF) module, and the guided feature refusion module (GFRM). The CDFF module is designed to aggregate features flexibly by fusing features from two backbones with different attention modules in the first fusion procedure. The GFRM and SCFF module are proposed in the refusion procedure to generate accurate segmentation results. Inspired by agent attention, the GFRM dynamically assembles detailed features for mask generation by refusing local and global features with the guidance of fusion results from CDFF. The SCFF module complements the significant features by enhancing and integrating global, local, and detailed features, and finally generates masks of instances. Extensive experiments demonstrate that GLFRNet outperforms the second-best model by 1.9, 1.3, and 0.3 in mask average precisions (APs) on NWPU VHR-10, WHU Building, and iSAID datasets. Jiaqi Zhao 0001, Yari Wang, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | ST-Mamba: Spatio-Temporal Synergistic Model for Remote Sensing Change DetectionabstractThe advancement of remote sensing and deep learning has spurred interest in high-resolution image change detection (CD). However, pseudo-changes in multi-temporal images, due to complex scenes and variable imaging conditions, often lead to significant misdetection in current methods. To address this problem, we propose a new CD framework: Spatio-Temporal Mamba (ST-Mamba), which consists of three key components. Firstly, a Mamba-based Feature Extraction Module (MFEM) is designed as the encoder to extract essential features from multi-temporal images by leveraging Mamba’s capability to capture inherent information in long data sequences. Secondly, a Spatio-Temporal Synergy Module (STSM) is developed to unify the background features of multi-temporal feature maps into a common domain by employing the state-space model for spatio-temporal modeling. Finally, a Spatio-Temporal Fusion Module (STFM) is created to guide the fusion of image features at different scales and across channels by utilizing a feature map of the unified background features. Experimental results on five widely used change detection datasets show significant improvements over current state-of-the-art methods. Jiaqi Zhao 0001, Jianpeng Xie 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Similarity Regulation and Calibration Alignment for Weakly Supervised Text-Based Person Re-IdentificationabstractTraditional text-based person re-identification relies on identity labels. However, it is impossible to annotate large datasets, since identity annotation is expensive and time-consuming. Weakly supervised text-based person re-identification, where only text–image pairs are available without annotation of identities, is very practical in real life. While dealing with the weakly supervised person re-identification, two issues should be strengthed, i.e., alignment caused by different modal, and cross-modal matching ambiguity caused by the lack of identity labels. In this article, we propose a similarity regulation and calibration alignment (SRCA) framework, which consists of two unimodal encoders for images and text, respectively, and a multi-modal encoder for the masked language modeling task. First, a similarity regulation (SR) strategy is proposed to relax the strict one-to-one constraints for the local similarities between different pairs by introducing a novel soft objective. The soft objective can adjust hard objectives to achieve soft cross-modal alignment by establishing a many-to-many relationship between two modalities. Second, the calibration alignment (CA) module is proposed to improve intra-class compactness by modeling pseudo-label assignment as optimal transport. The ambiguity of cross-modal matching can be reduced by aligning features and pseudo-labels of different modalities and gradually calibrating the distribution of pseudo-labels. Experimental results show that our method has achieved obvious advantages compared with existing methods and also demonstrated competitive performance compared with fully supervised methods. Ao Fu, Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | CCFL: Customized Client Federated Learning for Unsupervised Person Re-identificationabstractFederated learning-based person re-identification (Re-ID) aims to address the issue of data silos in surveillance systems caused by increasingly stringent regulations on sensitive data. However, due to differences in data collection locations, times, and scales, severe non-independent and identically distributed (non-IID) characteristics exist across different Re-ID datasets. Existing federated learning-based Re-ID methods often adopt a unified model structure, which prevents the model from adapting well to diverse data environments, thereby significantly degrading the overall Re-ID performance. To address the challenges of training neural networks on non-IID data across different datasets, we propose a customizable federated learning framework. First, customizable clients allow each organization to freely select suitable neural network training methods and model architectures based on local data scales and prior knowledge, thus improving training outcomes. Second, since traditional federated learning frameworks cannot achieve knowledge fusion through parameter exchange between models with different architectures, we introduce an independent model, referred to as the interaction model, specifically designed for knowledge exchange among clients. The interaction model learns parameters (knowledge) from local models on each client through distillation learning. Subsequently, the interaction model is uploaded to the server, where it undergoes parameter fusion (knowledge exchange) with interaction models from other clients. Finally, the interaction model, enriched with knowledge from other clients, guides local model training through knowledge distillation. It is worth noting that selecting a lightweight interaction model, while potentially impacting Re-ID performance, can significantly reduce communication costs between the server and clients. Yong Zhou 0003, Fayao Liu, Jiaqi Zhao 0001, Hancheng Zhu, Wen-Liang Du 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Fine-grained semantic oriented embedding set alignment for text-based person search
Jiaqi Zhao 0001, Ao Fu, Yong Zhou 0003, Wen-Liang Du 0002, Rui Yao 0006 |
Image Vis. Comput. | 4 |
| 2024 | A Mamba-Diffusion Framework for Multimodal Remote Sensing Image Semantic SegmentationabstractRecent advances in deep learning have made significant progress in multimodal remote sensing semantic segmentation. However, current methods face challenges in maintaining geometric consistency, particularly when dealing with large objects, resulting in fragmented segmentation masks. We propose a Mamba-diffusion framework to preserve geometric consistency in segmentation masks. This framework preserves geometric consistency by introducing a generative diffusion-based semantic segmentation pipeline and developing a Mamba-based multimodal fusion model. The fusion model fuses the multimodal images in multiple scales and scanning mechanisms by a double cross-fusion (DCF) module. Then, the cross-modal information is further integrated by a dual-splitting structured state-space (DS-S4) model. Finally, the diffusion-based segmentation pipeline predicts semantic masks by progressively refining random Gaussian noise, guided by fused multimodal features. Our experimental results, verified on WHU-OPT-SAR and Hunan datasets, demonstrate that the proposed framework surpasses state-of-the-art (SOTA) methods by a considerable margin. Our codes are available athttps://github.com/WenliangDu/MambaDiffusion. Wen-Liang Du 0002, Yang Gu 0005, Jiaqi Zhao 0001, Hancheng Zhu, Rui Yao 0006, Yong Zhou 0003 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | OrientedFormer: An End-to-End Transformer-Based Oriented Object Detector in Remote Sensing ImagesabstractOriented object detection in remote sensing images is a challenging task due to objects being distributed in multiorientation. Recently, end-to-end transformer-based methods have achieved success by eliminating the need for post-processing operators compared to traditional convolutional neural network (CNN)-based methods. However, directly extending transformers to oriented object detection presents three main issues: 1) objects rotate arbitrarily, necessitating the encoding of angles along with position and size; 2) the geometric relations of oriented objects are lacking in self-attention, due to the absence of interaction between content and positional queries; and 3) oriented objects cause misalignment, mainly between values and positional queries in cross-attention, making accurate classification and localization difficult. In this article, we propose an end-to-end transformer-based oriented object detector, consisting of three dedicated modules to address these issues. First, Gaussian positional encoding (PE) is proposed to encode the angle, position, and size of oriented boxes using Gaussian distributions. Second, Wasserstein self-attention is proposed to introduce geometric relations and facilitate interaction between content and positional queries by utilizing Gaussian Wasserstein distance scores. Third, oriented cross-attention is proposed to align values and positional queries by rotating sampling points around the positional query according to their angles. Experiments on six datasets DIOR-R, a series of DOTA, HRSC2016, and ICDAR2015 show the effectiveness of our approach. Compared with previous end-to-end detectors, the OrientedFormer gains 1.16 and 1.21 AP50 on DIOR-R and DOTA-v1.0, respectively, while reducing training epochs from$3\times $to$1\times $. The code is available athttps://github.com/wokaikaixinxin/OrientedFormer. Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Multi-Modal LiDAR Point Cloud Semantic Segmentation with Salience Refinement and Boundary PerceptionabstractPoint cloud segmentation is essential for scene understanding, which provides advanced information for many applications, such as autonomous driving, robots, and virtual reality. To improve the accuracy and robustness of point cloud segmentation, many researchers have attempted to fuze camera images to complement the color and texture information. The common fusion strategy is the combination of convolutional operations with concatenation, element-wise addition or element-wise multiplication. However, conventional convolutional operators tend to confine the fusion of modal features within their receptive fields, which can be incomplete and limited. In addition, the inability of encoder–decoder segmentation networks to explicitly perceive segmentation boundary information results in semantic ambiguity and classification errors at object edges. These errors are further amplified in point cloud segmentation tasks, significantly affecting the accuracy of point cloud segmentation. To address the above issues, we propose a novel self-attention multi-modal fusion semantic segmentation network for point cloud semantic segmentation. Firstly, to effectively fuze different modal features, we propose a self-cross fusion module (SCF), which models long-range modality dependencies and transfers complementary image information to the point cloud to fully leverage the modality-specific advantages. Secondly, we design the salience refinement module (SR), which calculates the importance of channels in the feature maps and global descriptors to enhance the representation capability of salient modal features. Finally, we propose the local-aware anisotropy loss measure the element-level importance in the data and explicitly provide boundary information for the model, which alleviates the inherent semantic ambiguity problem in segmentation networks. Extensive experiments on two benchmark datasets demonstrate that our proposed method surpasses current state-of-the-art methods. Yong Zhou 0003, Zeming Xie, Jiaqi Zhao 0001, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | A Semi-Supervised Image-to-Image Translation Framework for SAR-Optical Image MatchingabstractSynthetic Aperture Radar (SAR) and optical image matching aims to acquire correspondences from a certain pair of SAR and optical images. Recent advances in the image-to-image translation provided a way to simplify the SAR-optical image matching into the SAR-SAR or optical-optical image matchings. Existing image-to-image translations mainly focus on supervised or unsupervised learning. However, gathering sufficient amounts of aligned training data for supervised learning is challenging, while unsupervised learning cannot guarantee enough correct correspondences. In this work, we investigate the applicability of semi-supervised image-to-image translation for SAR-optical image matching such that both aligned and unaligned SAR-optical images could be used. To this end, we combine the benefits of both supervised and unsupervised well-known image-to-image translation methods, i.e., Pix2pix and CycleGAN, and propose a simple yet effective semi-supervised image-to-image translation framework. Through extensive experimental comparisons to baseline methods, we verify the effectiveness of the proposed framework in both semi-supervised and fully-supervised settings. Our codes are available at https://github.com/WenliangDu/Semi-I2I. Wen-Liang Du 0002, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Zhiwen Shao, Xiaolin Tian 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Few-Shot Object Detection via Context-Aware Aggregation for Remote Sensing ImagesabstractFew-shot object detection methods have made prodigious progress in recent years. However, these methods are designed for optical images at a single scale, which leads to significantly degraded detection performance due to object scale variation of remote sensing images. In this letter, we propose a few-shot object detection method for the problem of scale variation in remote sensing images. More specifically, our model contains two main components: a context-aware pixel aggregation (CPA) that allows the model to adapt to objects at different scales through different scale convolution and a context-aware feature aggregation (CFA) that enhances context awareness to obtain more semantic information through a graph convolution network (GCN). Experiments on the DIOR dataset demonstrate that our model can achieve a satisfying detection performance on remote sensing images, and our model performs significantly better than the state-of-the-art model. Yong Zhou 0003, Jiaqi Zhao 0001, Hancheng Zhu, Rui Yao 0006, Wen-Liang Du 0002 |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Facial action unit detection via hybrid relational reasoning
Zhiwen Shao, Yong Zhou 0003, Bing Liu 0016, Hancheng Zhu, Wen-Liang Du 0002, Jiaqi Zhao 0001 |
Vis. Comput. | 5 |
| 2018 | An Automatic Evaluation Platform for Feature Matching Algorithms Based on an Orbital Optical Pushbroom Stereo Imaging SystemabstractThanks to the high robustness of feature-based matching algorithms, they can be applied in remote sensing (RS) applications with complex image changes. However, evaluating feature matching algorithms on RS images is still challenging, because creating ground truth matching data sets of RS images costs a lot of both computing time and human operator time. Therefore, in this letter, we present an evaluation platform for simulating RS ground truth data sets of feature-points correspondences based on a customizable orbital optical pushbroom stereo imaging system. With the help of the proposed platform, evaluating feature matching algorithms could be fully automatic and customized. The performance of three state-of-the-art feature matching algorithms based on local transformation constraint is evaluated and discussed on the proposed platform with comprehensive experiments. The evaluation results of the platform are also compared with the manual evaluation results of 10 pairs of real RS stereo images. The evaluation results show that the proposed platform indeed offers an efficient way for evaluating RS feature matching algorithms. Wen-Liang Du 0002, Xiaolin Tian 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2017 | An automatic image registration evaluation model on dense feature points by pinhole camera simulationabstractEvaluating image registration methods in images with dense feature points is challenging, because it's hard to discriminate real inliers in thousands of resulting correspondences processed by image registration methods. Therefore, this paper presents a dense feature points simulation which could provide ground truth of correspondences for evaluating image registration methods automatically. Moreover, the dense feature points are created by simulating pinhole camera model, parallax of reference image and sensed image, radial distortion of camera lens and random outliers. The performance of five state-of-art image registration methods is evaluated and discussed on the dense correspondences simulated by proposed model. The evaluation results show that the proposed model indeed offers a practical way for evaluating image registration methods on dense feature points. Wen-Liang Du 0002, Xiaolin Tian 0001 |
ICIP | 1 |