EDBT 2026 Demo / reviewers in the wild / expert
Tengfei Gong
dblp:261/9867
· DBLP profile ↗
10ranked-venue papers
4as first author
9since 2021 · last 2025
0000-0002-8465-0144ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploring Flexibility in Incremental Few-Shot Object DetectionabstractIncremental few-shot object detection (iFSD) is critical for real-world applications, enabling rapid adaptation to novel categories with minimal data while mitigating catastrophic forgetting. However, existing methods lack flexibility, particularly in feature representation. The pursuit of a flexible approach to iFSD presents a substantial challenge. To address this, we propose an Attention-Based Feature Aggregation (AFA) that dynamically refines feature representations guided by limited support samples, and a Conditional Classifier (CC) that dynamically refines the generated class prototypes based on the existing knowledge, while conditioning on the limited support images, enhancing flexibility and adaptability. We conducted comprehensive experiments on the MS COCO and LVIS datasets to validate the superiority of our approach. Dongdong Gong, Tengfei Gong, Yaxiong Chen, Jinglin Yuan, Shengwu Xiong 0001 |
ICME | 2 |
| 2025 | Temporal-Aware Spatial Interaction Transformer for Crop Yield Prediction Based on Multisensor Satellite Image Time SeriesabstractCrop yield prediction is crucial for agricultural decision making. Satellite Image Time Series (SITS) data, which provide continuous temporal observations of vegetation changes, have become a standard for accurate prediction. Recent studies have shown that the integration of multimodal data from different satellite sensors significantly enhances the performance of SITS-based crop yield predictions. However, existing methods often rely on simplistic combinations of multimodal data. Temporal inconsistency between different modals is not considered. In addition, the influence of spatial interaction on crop growth is ignored. To address this issue, we propose the TASI-Transformer (Temporal-Aware Spatial Interaction Transformer) for crop yield prediction using multisensor satellite image time series, which incorporates two innovative modules: a Temporal Enhanced Position Encoding (TEPE) module that incorporates crop growth dates to extract unique temporal information from different satellite time series. A Spatial Enhanced Multimodal Interaction (SEMI) module that learns the impact of spatial relationships between different regions and the interaction between multiple modals. Experimental results on the SICKLE and CROPNET datasets demonstrate that the proposed method achieves state-of-the-art performance, with a Mean Absolute Percentage Error (MAPE) as low as 26.99% in SICKLE using actual season data, and a Root Mean Squared Error (RMSE) of 6.85, Coefficient of Determination(R²) of 0.51, and Pearson Correlation (CORR) of 0.71 in CROPNET, outperforming existing methods. Tengfei Gong, Xinchao Zhu, Yaxiong Chen, Shengwu Xiong 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | Global-Local Fusion With Semantic Information Guidance for Accurate Small Object Detection in UAV Aerial ImagesabstractIn recent years, the rapid development of the unmanned aerial vehicle (UAV) technology has generated a large number of aerial photography images captured by UAV. Consequently, the object detection in UAV aerial images has emerged as a recent research focus. However, due to the flexible flight heights and diverse shooting angles of UAV, two significant challenges have arisen in UAV aerial images: extreme variation in target scale and the presence of numerous small targets. To address these challenges, this article introduces a semantic information-guided fusion module specifically tailored for small targets. This module utilizes high-level semantic information to guide and align the underlying texture information, thereby enhancing the semantic representation of small targets at the feature level and subsequently improving the model’s ability to detect them. In addition, this article introduces a novel global–local fusion detection strategy to strengthen the detection of small targets. We have redesigned the foreground region assembly method to address the drawbacks of previous methods that involved multiple inferences. Extensive experiments conducted on the VisDrone and UAVDT datasets demonstrate that our two self-designed modules can significantly enhance the detection capability of small targets compared with the YOLOX-M model. Our code is publicly available at:https://github.com/LearnYZZ/GLSDet. Yaxiong Chen, Zhengze Ye, Haokai Sun 0001, Tengfei Gong, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Discover the Unknown Ones in Fine-Grained Ship DetectionabstractRemote sensing image-based ship identification technology has great applications in areas such as national defense and fishery management. However, existing remote sensing ship studies mainly focus on a closed environment and overlook actual sea conditions, while new military ships will be encountered. These unknown categories of ships will be ignored or misclassified by existing models, dramatically affecting the accurate assessment of the maritime situation. Furthermore, existing unknown detection methods for natural images fail to tackle the remote sensing ship detection problem for the property of high similarity in overall appearance. To cope with this problem, this paper proposes a fine-grained unknown ship detection network. Firstly, we explore a class-balanced proposal sampler to avoid inefficient information learning. Secondly, we propose a finegrained memory bank-based contrastive learning strategy to separate different categories. Finally, to further separate unknown classes, we adopt an uncertainty-aware unknown learner with logit to reduce the uncertainty of fine-grained predictions. Experiments conducted in three public ship detection datasets ShipRSImageNet, DOSR, and HRSC2016 show that the method not only achieves good detection on unknown class ships, but also improves the detection accuracy on known classes. The code is available at https://github.com/FoRGEU/DUONet. Tengfei Gong, Yaxiong Chen, Shengwu Xiong 0001, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Domain Mapping Network for Remote Sensing Cross-Domain Few-Shot ClassificationabstractIt is a challenging task to recognize novel categories with only a few labeled remote sensing images. Currently, meta-learning solves the problem by learning prior knowledge from another dataset where the classes are disjoint. However, the existing methods assume the training dataset comes from the same domain as the test dataset. For remote sensing images, test dataset may come from different domains. It is impossible to collect a training dataset for each domain. Meta-learning and transfer learning are widely used to tackle the few-shot classification and the cross-domain classification, respectively. However, it is difficult to recognize novel categories from various domains with only a few images. In this paper, a Domain Mapping Network (DMN) is proposed to cope with the few-shot classification under domain shift. DMN trains an efficient few-shot classification model on the source domain and then adapts the model to the target domain. Specifically, dual autoencoders are exploited to fit the source and target domain distribution. First, DMN learns an autoencoder on the source domain to fit the source domain distribution. Then, a target autoencoder is initiated from the source domain autoencoder and further updated with a few target images. To ensure the distribution alignment, cycle-consistency losses are proposed to jointly train the source autoencoder and target autoencoder. Extensive experiments are conducted to validate the generalizable and superiority of the proposed method. Xiaoqiang Lu, Tengfei Gong, Xiangtao Zheng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Human action recognition by multiple spatial clues network
Xiangtao Zheng, Tengfei Gong, Xiaoqiang Lu, Xuelong Li 0001 |
Neurocomputing | 2 |
| 2022 | Meta Self-Supervised Learning for Distribution Shifted Few-Shot Scene ClassificationabstractFew-shot classification tries to recognize novel remote sensing image categories with a few shot samples. However, current methods assume that the test dataset shares the same domain with the labeled training dataset where prior knowledge is learned. It is infeasible to collect a training dataset for each domain, since remote sensing images may come from various domains. Exploiting the existing labeled dataset from another domain (source domain) to help the target dataset (target domain) classification would be efficient. In this paper, both meta-learning and self-supervised learning are jointly conducted for few-shot classification. Specifically, meta-learning is executed over a pre-trained network for few-shot classification. Furthermore, self-supervised learning is exploited to fit the target domain distribution by training on unlabeled target domain images. Experiments are conducted on NWPU, EuroSAT and Merced datasets to validate the effectiveness. Tengfei Gong, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Generalized Scene Classification From Small-Scale Datasets With Multitask LearningabstractRemote sensing images contain a wealth of spatial information. Efficient scene classification is a necessary precedent step for further application. Despite the great practical value, the mainstream methods using deep convolutional neural networks (CNNs) are generally pretrained on other large datasets (such as ImageNet) and thus fail to capture the specific visual characteristics of remote sensing images. For another, it lacks the generalization ability to new tasks when training a new CNN from scratch with an existing remote sensing dataset. This article addresses the dilemma and uses multiple small-scale datasets to learn a generalized model for efficient scene classification. Since the existing datasets are heterogeneous and cannot be directly combined to train a network, a multitask learning network (MTLN) is developed. The MTLN treats each small-scale dataset as an individual task and uses complementary information contained in multiple tasks to improve generalization. Concretely, the MTLN consists of a shared branch for all tasks and multiple task-specific branches with each for one task. The shared branch extracts shared features for all tasks to achieve information sharing among tasks. The task-specific branch distills the shared features into task-specific features toward the optimal estimation of each specific task. By jointly learning shared features and task-specific features, the MTLN maintains both generalization and discrimination abilities. Two types of MTL scenarios are explored to validate the effectiveness of the proposed method: one is to complete multiple scene classification tasks and the other is to jointly perform scene classification and semantic segmentation. Xiangtao Zheng, Tengfei Gong, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Cross-Domain Scene Classification by Integrating Multiple Incomplete SourcesabstractCross-domain scene classification identifies scene categories by learning knowledge from a labeled data set (source domain) to an unlabeled data set (target domain), where the source data and the target data are sampled from different distributions. A lot of domain adaptation methods are used to reduce the distribution shift across domains, and most existing methods assume that the source domain shares the same categories with the target domain. It is usually hard to find a source domain that covers all categories in the target domain. Some works exploit multiple incomplete source domains to cover the target domain. However, in such setting, the categories of each source domain are a subset of the target-domain categories, and the target domain contains “unknown” categories for each source domain. The existence of unknown categories results in the conventional domain adaptation unsuitable. Known and unknown categories should be treated separately. Therefore, a separation mechanism is proposed to separate the known and unknown categories in this article. First, multiple-source classifiers trained on the multiple source domains are used to coarsely separate the known/unknown categories in the target domain. The target images with high similarities to source images are selected as known categories, and the target images with low similarities are selected as unknown categories. Then, a binary classifier trained using the selected images is used to finely separate all target-domain images. Finally, only the known categories are implemented in the cross-domain alignment and classification. The target images get labels by integrating the hypotheses of multiple-source classifiers on the known categories. Experiments are conducted on three cross-domain data sets to demonstrate the effectiveness of the proposed method. Tengfei Gong, Xiangtao Zheng, Xiaoqiang Lu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Multisource Compensation Network for Remote Sensing Cross-Domain Scene ClassificationabstractCross-domain scene classification refers to the scene classification task in which the training set (termed source domain) and the test set (termed target domain) come from different distributions. Various domain adaptation methods have been developed to reduce the distribution discrepancy between different domains. However, current domain adaptation methods assume that the source domain and target domain share the same categories. In reality, it is hard to find a source domain that can completely cover all the categories of target domain. In this article, we propose to use multiple complementary source domains to form the categories of target domain. A multisource compensation network (MSCN) is proposed to tackle these challenges: distribution discrepancy and category incompleteness. First, a pretrained convolutional neural network (CNN) is exploited to learn the feature representation for each domain. Second, a cross-domain alignment module is developed to reduce the domain shift between source and target domains. Domain shift is reduced by mapping the two domain features into a common feature space. Finally, a classifier complement module is proposed to align categories in multiple sources and learn a target classifier. Two cross-domain classification data sets are constructed using four heterogeneous remote sensing scene classification data sets. Extensive experiments are conducted on these datasets to validate the effectiveness of the proposed method. The proposed method can achieve 81.23% and 81.97% average accuracies on two-source-complementary data set and three-source-complementary data set, respectively. Xiaoqiang Lu, Tengfei Gong, Xiangtao Zheng |
IEEE Trans. Geosci. Remote. Sens. | 2 |