VLDB 2026 Research / reviewers in the wild / expert
Yijing Wang 0004
dblp:56/3525-4
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-0862-6564ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ID-Splat: Propagating Object Identities for Segmenting 3D Aerial-view ScenesabstractHigh-resolution Earth Observation technologies present unprecedented opportunities for geospatial analysis, yet traditional 2D aerial-view semantic segmentation remains limited by its inability to model spatial relationships and handle object occlusions. While 3D Aerial-view Segmentation (3DAS) has emerged to address these limitations, existing methods predominantly rely on 2D discriminative models pre-trained on natural scenes. These models struggle to accurately recognize aerial-view imagery, resulting in suboptimal performance due to significant domain discrepancies. This paper introduces ID-Splat, a novel object-centric framework that directly leverages multi-view object identities without discriminative information to enhance 3D semantic understanding. ID-Splat implements a two-stage process: first, Mask-object Tracking combines SAM and Point Tracking to establish robust and consistent object identities across multi-view aerial images; second, Object Integration & Propagation assigns these identities to 3D Gaussian Splatting (3DGS) points, enabling complete 3D segmentation through semantic propagation. Experimental results on the 3D-AS dataset demonstrate that ID-Splat significantly outperforms existing methods, particularly under sparse supervision conditions. ID-Splat also achieves state-of-the-art performance while reducing the need for extensive labeled data by effectively leveraging the inherent 3D structure. Yijing Wang 0004, Xu Tang 0004, Xiangrong Zhang, Jingjing Ma 0001 |
AAAI | 1 |
| 2025 | Cross-Modal Remote Sensing Image-Text Retrieval via Context and Uncertainty-Aware PromptabstractThe cross-modal remote sensing image-text retrieval (CMRSITR) is a lively research topic in the remote sensing (RS) community. Benefiting from the large pretrained image-text models, many successful CMRSITR methods have been proposed in recent years. Although their performance is attractive, there are still some challenges. First, fine-tuning large pretrained models requires a significant amount of computational resources. Second, most large models are pretrained by natural images, which reduces their effectiveness in processing RS images. To tackle these challenges, we propose a new CMRSITR network named context and uncertainty-aware prompt (CUP). First, prompt tuning theory is introduced into CUP to eliminate the burden of optimization resources. By training the prompt tokens rather than all parameters, the large model's knowledge can be transferred to CMRSITR tasks with small trainable parameters. Second, considering the differences between natural-image-based prior clues and RS images, apart from adopting the free-prompt tokens, we develop a prompt generation module (PGM) to produce the RS-oriented prompt tokens. The specific prompt tokens are rich in object-level messages of RS images, which help CUP narrow the gaps between natural large models and RS images. Third, we further design an uncertainty estimation module (UEM) to whittle down the uncertainties caused by the model and data. This way, can not only the semantic misalignment and intraclass diversity imbalance problems be mitigated but also the RS clues can be deeply explored. Competitive experimental results counted on three public benchmark datasets demonstrate that our CUP can achieve competitive performance in the CMRSITR task compared with many existing methods. Our source codes are available at: https://github.com/TangXu-Group/Cross-modal-remote-sensing-image-and-text-retrieval-models/tree/main/CUP. Yijing Wang 0004, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Pseudo-Viewpoint Regularized 3D Gaussian Splatting For Remote Sensing Few-Shot Novel View SynthesisabstractIn remote sensing (RS), Few-Shot Novel View Synthesis (FS-NVS) focuses on creating images of unobserved viewpoints using limited training images. Recently, 3D Gaussian Splatting (3DGS) has drawn scholars’ attention by its increasing rendering speeds and providing an explicit neural representation for 3D scenes. However, 3DGS tends to overfit limited training data. To tackle this challenge, we propose a Pseudoview Regularized 3DGS (PR3DGS) FSNVS method for RS scenarios. Our PR3DGS method introduces a pseudo-views regularization module to discriminate synthetic RS images generated from training- or pseudo-viewpoints. Therefore, our PR3DGS method can effectively mitigate overfitting in seen views and enhance the model’s capability to generate more realistic RS images from novel viewpoints. Besides, the excellent experimental results on the LEVIR-NVS dataset demonstrate the effectiveness of our method in RS FSNVS. Yijing Wang 0004, Xu Tang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0001, Licheng Jiao |
IGARSS | 1 |
| 2023 | Exchange Data Augmentation for Change DetectionabstractChange detection is developed to automatically identify semantic changes between remote sensing (RS) images captured at different points of time in a specific geographic location. Due to the limited availability of annotated data and the high complexity of the change detection problem, the performance of change detection models can not meet our expectations. To address this issue, many methods are proposed to solve this issue. However, ignoring the characteristics of the change detection task limits their performance. Therefore, we introduce a novel exchange data enhancement method (EDEM) strategy to generate image pairs to help the neural network to capture the temporal consistency in the data. We evaluate the proposed approach on two publicly available datasets and compare it with several state-of-the-art methods. The experimental results demonstrate that our proposed approach can effectively improve the performance of change detection models, achieving state-of-the-art performance on both datasets. JunYi Duan, Yijing Wang 0004, Xu Tang 0004, Jingjing Ma 0001, Yuqun Yang |
IGARSS | 2 |
| 2023 | Multi-Scale Interaction Prototypical Network For Few-Shot Remote Sensing Scene ClassificationabstractFew-shot remote sensing scene classification (FSRSSC) aims to make the model quickly adapt to new scenes with a small amount of annotation data. The large intra-class variance and high inter-class similarity in remote sensing (RS) scenes make this task more challenging. To this end, we propose a multi-scale interaction prototypical network, which pays attention to capturing multi-scale information of images during model learning, and then generates a prototype representation of mixed query information through a feature interaction module, thereby enhancing the rapid learning ability of the model, so as to reducing of the within-class and between-class variance ratio in RS scenes. The positive experimental results on UC-Merced and NWPU datasets demonstrate the effectiveness of our model in FSRSSC. Shiji Pei, Yijing Wang 0004, Jingjing Ma 0001, Xu Tang 0004, Yuqun Yang |
IGARSS | 2 |
| 2023 | Interacting-Enhancing Feature Transformer for Cross-Modal Remote-Sensing Image and Text RetrievalabstractCross-modal remote sensing image-text retrieval (CMRSITR) is a challenging topic in the remote sensing (RS) community. It has gained growing attention because it can be flexibly used in many practical applications. In the current deep era, with the help of deep convolutional neural networks (DCNNs), many successful CMRSITR methods have been proposed. Most of them first learn valuable features from RS images and texts respectively. Then, the obtained visual and textual features are mapped into a common space for the final retrieval. The above operations are feasible, however, two difficulties are still to be solved. One is that the semantics within the visual and textual features are misaligned due to the independent learning manner. The other one is that the deep links between RS images and texts cannot be fully explored by simple common space mapping. To overcome the above challenges, we propose a new model named interacting-enhancing feature transformer (IEFT) for CMRSITR, which regards the RS images and texts as a whole. First, a simple feature embedding module (FEM) is developed to map images and texts into the visual and textual feature spaces. Second, an information interacting-enhancing module (IIEM) is designed to simultaneously model the inner relationships between RS images and texts and enhance the visual features. IIEM consists of three feature interacting-enhancing (FIE) blocks, each of which contains an inter-modality relationship interacting (IMRI) sub-block and a visual feature enhancing (VFE) sub-block. The duty of IMRI is to exploit the hidden relations between cross-modal data, while the responsibility of VFE is to improve the visual features. By combining them, semantic bias can be mitigated, and the complex contents of RS images can be studied. Finally, the retrieval module (RM) is constructed to generate the matching scores for deciding the search results. Extensive experiments are conducted on four public RS data sets. The positive results demonstrate that our IEFT can achieve superior retrieval performance compared with many existing methods. Our source codes are available at https://github.com/TangXu-Group/Cross-modal-remote-sensing-image-and-text-retrieval-models/tree/main/IEFT. Xu Tang 0004, Yijing Wang 0004, Jingjing Ma 0001, Xiangrong Zhang, Fang Liu 0034, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Multi-Scale Interactive Transformer for Remote Sensing Cross-Modal Image-Text RetrievalabstractCross-modal Remote sensing (RS) image-text retrieval (CMR-SITR) plays a crucial role in the RS community. A common way for CMRSITR is to extract texts and RS images' feature representations separately and then measure their similarities in the specific or common feature space. Recently, along with the booming of deep convolutional neural networks (DCNNs), these kinds of methods are vivid and achieve successes in their own applications. However, they neglect the inherent relationships between different features, and they are always heavy. To overcome the limitations mentioned above, we propose a new model for CMRSITR in this paper, named multi-scale interactive transformer (MSIT). MSIT first adopts simple feature learning models for texts and RS images which could ensure the whole model is not heavy. Then, MSIT introduces transformer encoders to enhance features' usefulness by considering the potential relations between different representations. Also, a lightweight multi-scale feature learning module is proposed to mine more plentiful contents from RS images. Finally, instead of outputting the features, MSIT produces matching scores for texts and RS images, which can be used to decide the retrieval results directly. The experimental results on two RS datasets indicate our modal is effective for CMRSITR. Yijing Wang 0004, Jingjing Ma 0001, Mingteng Li, Xu Tang 0004, Xiao Han 0012, Licheng Jiao |
IGARSS | 1 |