VLDB 2026 Research / reviewers in the wild / expert
Chubo Deng
dblp:278/6555
· DBLP profile ↗
9ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0003-3469-5624ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scene-Specific Multiprototype Network for Remote Sensing Scene Graph GenerationabstractRemote sensing scene graph generation aims to capture both objects and their semantic relationships, offering a comprehensive understanding of complex scenes. However, two major challenges hinder the performance of existing methods. First, remote sensing images often contain a large number of objects, many of which are unrelated. Performing global feature interactions across all objects introduces noise from irrelevant pairs, degrading feature quality. Second, relationship categories in remote sensing scenes exhibit significant intra-class variation across different contexts, and long-tailed distribution further complicates learning due to limited samples for tail classes. To address these issues, we propose the Scene-specific Multi-Prototype Network (SSMP). Our method performs contextual interactions selectively based on object and relationship categories, reducing interference from irrelevant features. Moreover, we introduce a scene-specific multi-prototype classification framework that better captures the diverse visual manifestations of each relationship class, while also improving discrimination under long-tailed distributions. Experimental results demonstrate that the proposed model achieves state-of-the-art (SOTA) performance, with a minimum improvement of 4.5% and a maximum improvement of 21.3% in mR@20 on the PredCls task over baseline models. Zhongyan Hou, Chubo Deng, Qiwei Yan, Tong Ling, Wanxuan Lu, Yingyan Hou, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Hypergraph-Guided Multimodal Prototype for Remote Sensing Scene UnderstandingabstractNoticeable achievements have been made in entity-level perception tasks (e.g., object detection) in remote sensing (RS) image interpretation. But for RS images carrying rich content, individual perception cannot well obtain the interaction patterns between entities. The recognition of relationships between entities is the key to deeply understanding RS scenes. In this article, we propose a hypergraph-guided multimodal prototype network (HMPNet), which performs relation recognition by matching relation representations with multimodal predicate prototypes. To overcome the imbalance of modal information in the matching process, a multimodal calibration strategy is devised, taking into account the image subprototype and text subprototype, which makes prediction results more reliable. Meanwhile, to align image and text subprototypes and explore relevant semantic patterns, the multimodal hypergraph is constructed to efficiently capture the associations between heterogeneous prototypes. Experimental results show that the performance of our model can reach the state-of-the-art (SOTA) level on the RS scene graph generation (SGG) task. Chubo Deng, Qiwei Yan, Liangyu Xu, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Ringmo-SenseV2: Remote Sensing Foundation Model for Spatiotemporal Prediction Based on Multisource Heterogeneous Time-Series DataabstractThe rapid development of Remote Sensing (RS) technology has generated a vast amount of heterogeneous time series data from various sources, including drone videos, satellite time-series images, and multi-object trajectories. Effectively processing and analyzing this multi-source heterogeneous data for accurate spatiotemporal prediction is crucial in fields such as environmental protection and disaster response. In this paper, we propose a universal predictive foundation model named Ringmo-SenseV2 to learn the general evolutionary patterns of RS elements from massive heterogeneous data. Ringmo- SenseV2 features a Mixture-of-Heterogeneous-Experts (MoHE) Transformer, which unifies the modeling of multi-source heterogeneous time-series data. Additionally, to better capture the complex dependencies across different spatiotemporal locations, we introduce a hypergraph translator, treating embeddings of different spatiotemporal locations as nodes and employing hypergraph convolution for information propagation. Furthermore, to enhance the model’s adaptability to different evolution speeds during pre-training, we implement the Adaptive tube Masking (AM) strategy, which controls prediction difficulty by adaptively setting mask proportions for sequences with varying evolution speeds. Extensive experiments demonstrate that Ringmo-SenseV2 exhibits outstanding performance across various RS prediction tasks. Further tests on scene graph generation for RS images showcase the model’s ability to extract image features, thereby enhancing image perception tasks. Liangyu Xu, Wanxuan Lu, Leiyi Hu, Heming Yang 0003, Chubo Deng, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2025 | ReCon1M: A Large-Scale Benchmark Dataset for Relation Comprehension in Remote Sensing ImageryabstractScene graph generation (SGG) is a high-level visual understanding and reasoning task aimed at extracting entities (such as objects) and their interrelationships from images. Significant progress has been made in the study of SGG in natural images in recent years, but its exploration in the domain of remote sensing images remains very limited. The complex characteristics of remote sensing images necessitate higher time and manual interpretation costs for annotation compared to natural images. The lack of a large-scale public SGG benchmark is a major impediment to the advancement of SGG-related research in aerial imagery. In this article, we introduce the first publicly available large-scale, million-level relation dataset in the field of remote sensing images, which is named ReCon1M. Specifically, our dataset is built upon FAIR1M and comprises 22 262 images. It includes annotations for 873 761 object bounding boxes across 60 categories and 1 052 223 relation triplets across 59 categories based on these bounding boxes. We provide a detailed description of the dataset’s characteristics and statistical information. In addition, an efficient global context-aware network (EGCAN) is proposed to improve inference efficiency in dense relation prediction through an object-pair pre-screening mechanism. By integrating visual, spatial, and semantic features, EGCAN captures fine-grained pairwise features and object-level contextual information to enhance its ability to discriminate relation. We conduct two object detection tasks and three subtasks within SGG on this dataset, assessing the performance of mainstream methods on these tasks. The experimental results show that the proposed EGCAN achieves state-of-the-art (SOTA) performance in 17 out of 24 accuracy metrics across three tasks and delivers the best performance in frames per second (FPS) for model inference. The ReCon1M dataset and related resources are available athttps://recon1m-dataset.github.io/ Qiwei Yan, Chubo Deng, Zhongyan Hou, Wanxuan Lu, Fanglong Yao, Lingxiang Hao, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | SoPerModel: Leveraging Social Perception for Multi-Agent Trajectory PredictionabstractTrajectory prediction is an essential task within various automation systems. Recent studies have highlighted that the social interactions among multiple agents are crucial for accurate predictions, relying on empirically derived human-imposed constraints to model these interactions. However, from a sociological perspective, agents’ interactions exhibit significant inherent randomness. Dependence on a priori knowledge may lead to biased estimations of data distributions across different scenarios, failing to account for this randomness. Consequently, such methodologies often do not comprehensively capture the full spectrum of social influences, thus limiting the models’ predictive efficacy. To address these issues, we propose a novel multi-agent trajectory prediction framework, SoPerModel, which incorporates a freeform social evolution module (FSEM) and a local perception attention mechanism (LPA). The FSEM enables SoPerModel to naturally capture representative social interactions among agents without the reliance on additional human-derived priors. Through LPA, the model integrates both local and global social interaction information and leverages them to enhance trajectory prediction performance. Our framework is empirically evaluated on real-world trajectory prediction datasets, and the results demonstrate that our approach achieves a highly competitive performance compared with state-of-the-art models. Heming Yang 0003, Changyuan Tian 0001, Wanxuan Lu, Chubo Deng, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Improper Normalization on Hyperspectral Data Cause Cheating Result on Retrieval TaskabstractWe address the problems inherent in studying wrongful result coming from analyzing the hyperspectral data in Baiyangdian Lake. Let’s consider the following two cases: (i) Normalize on the entire set, then break into training/test sets and make predictions on the test set. (ii) Break into training/test sets, then normalize training/test sets separately, and make predictions on the test set. The main purpose of this paper is to present the case (i) yields a systematic overestimation of prediction capability that is triggered by the normalization. We provide theoretical analysis and use our example to demonstrate in some circumstances, the R-square criteria can reach to one for any data sets. This perfect R-square rate neither implies the data sets are well collected, nor does it give a good assessment on the algorithms it used, it happens when we apply (i) and combine Z-score transformation with leave-one-out cross validation together. Normalization is a fundamental block that most scientists will do before applying their algorithm, the paper indicates that the results of data analysis may be overestimated on a large scale. Chubo Deng |
IGARSS | 1 |
| 2023 | RingMo-Sense: Remote Sensing Foundation Model for Spatiotemporal Prediction via Spatiotemporal Evolution DisentanglingabstractRemote sensing spatiotemporal prediction aims to infer future trends from historical spatiotemporal data, e.g., videos and time series images, has a broad application prospect in many fields. The foundation model is a promising research direction for spatiotemporal information mining because of its robust feature extraction capability, and has made rapid progress in natural scenes. Nevertheless, due to the spatially multi-scale and temporally multi-scale properties in remote sensing data, these methods still encounter bottlenecks when applied to remote sensing. Therefore, we propose a foundation model for remote sensing spatiotemporal prediction via spatiotemporal evolution decoupling, abbreviated as RingMo-Sense. Considering spatial affinity, temporal continuity, and spatiotemporal interaction, we construct spatial, temporal, and spatiotemporal triple-branch prediction networks. Specifically, we use parameter-sharing and progressive joint training strategies to achieve stable long-range prediction and parameter reduction simultaneously. In addition, we build a remote sensing spatiotemporal dataset by collecting various remote sensing videos and time series images. The experimental results on six downstream spatiotemporal tasks demonstrate that the proposed model yields competitive performance. Fanglong Yao, Wanxuan Lu, Heming Yang 0003, Liangyu Xu, Leiyi Hu, Nayu Liu, Chubo Deng, Deke Tang, Changshuo Chen, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2022 | An Instance-Based Multitask Graph Network for Complex Facility Recognition in Remote Sensing ImageryabstractWith the availability of very high-resolution remote sensing imagery, the fine-grained recognition of complex geospatial facilities has become possible. We can view these facilities as a combination of component objects with specific functions and distribution. However, the existing methods are insufficient in modeling spatial relations of component objects. In this article, we propose an instance-based multitask graph network (IBMG-Net) for complex facility recognition. Specifically, we perform pixel-level component objects prediction and facility recognition simultaneously and achieve performance improvement of both tasks by joint multitasking training. Given the component information, we build an instance-based graph neural network (IBGN) where components are defined as nodes and their spatial relations are encoded as edges. The IBGN module aims to flexibly model spatial relations of complex facility. To enhance the feature representation of component objects, we utilize the multiscale region of interest module (MS-ROI) to retain all scale-specific features and the sparse context information module (SCM) to aggregate long-range context information. In addition, we build a new multitask dataset for complex facility recognition in remote sensing (MCF dataset) to verify the effectiveness of our method and alleviate the lack of pixel-level labeled multitask datasets in remote sensing. Extensive experiments on MCF also indicate that the significant performance improvement of our approach to complex facility recognition. Jingquan Peng, Xian Sun 0001, Chubo Deng, Fanglong Yao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Exploring a Fine-Grained Multiscale Method for Cross-Modal Remote Sensing Image RetrievalabstractRemote sensing (RS) cross-modal text–image retrieval has attracted extensive attention for its advantages of flexible input and efficient query. However, traditional methods ignore the characteristics of multiscale and redundant targets in RS image, leading to the degradation of retrieval accuracy. To cope with the problem of multiscale scarcity and target redundancy in RS multimodal retrieval task, we come up with a novel asymmetric multimodal feature matching network (AMFMN). Our model adapts to multiscale feature inputs, favors multisource retrieval methods, and can dynamically filter redundant features. AMFMN employs the multiscale visual self-attention (MVSA) module to extract the salient features of RS image and utilizes visual features to guide the text representation. Furthermore, to alleviate the positive samples ambiguity caused by the strong intraclass similarity in RS image, we propose a triplet loss function with dynamic variable margin based on prior similarity of sample pairs. Finally, unlike the traditional RS image-text dataset with coarse text and higher intraclass similarity, we construct a fine-grained and more challenging Remote sensing Image-Text Match dataset (RSITMD), which supports RS image retrieval through keywords and sentence separately and jointly. Experiments on four RS text–image datasets demonstrate that the proposed model can achieve state-of-the-art performance in cross-modal RS text–image retrieval task. Wenkai Zhang 0002, Kun Fu 0001, Chubo Deng, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |