EDBT 2026 Demo / reviewers in the wild / expert
Qibin He 0001
dblp:304/3890-1
· DBLP profile ↗
10ranked-venue papers
6as first author
10since 2021 · last 2024
0000-0003-2158-559XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Prompting Multi-Modal Image Segmentation with Semantic GroupingabstractMulti-modal image segmentation is one of the core issues in computer vision. The main challenge lies in integrating common information between modalities while retaining specific patterns for each modality. Existing methods typically perform full fine-tuning on RGB-based pre-trained parameters to inherit the powerful representation of the foundation model. Although effective, such paradigm is not optimal due to weak transferability and scarce downstream data. Inspired by the recent success of prompt learning in language models, we propose the Grouping Prompt Tuning Framework (GoPT), which introduces explicit semantic grouping to learn modal-related prompts, adapting the frozen pre-trained foundation model to various downstream multi-modal segmentation tasks. Specifically, a class-aware uni-modal prompter is designed to balance intra- and inter-modal semantic propagation by grouping modality-specific class tokens, thereby improving the adaptability of spatial information. Furthermore, an alignment-induced cross-modal prompter is introduced to aggregate class-aware representations and share prompt parameters among different modalities to assist in modeling common statistics. Extensive experiments show the superiority of our GoPT, which achieves SOTA performance on various downstream multi-modal image segmentation tasks by training only < 1% model parameters. Qibin He 0001 |
AAAI | 1 |
| 2024 | Orientation-Aware Multi-Modal Learning for Road Intersection Identification and MappingabstractAccurate identification of road intersections is the pivotal task for automatic construction of high-definition maps, particularly in unstructured scenes. Existing methods predominantly rely on single-modal data and thus show an obvious unimodal limitation, i.e., lack of contextual information. Moreover, these approaches overlook the benefits of leveraging multi-modal data fusion and representation learning that is crucial for generalizability. To this end, we propose a novel orientation-aware multi-modal learning paradigm, which formulates intersection identification as an oriented object detection task. Specifically, heterogeneous fusion is introduced to harmonize disparate data modalities, i.e., vector maps, point clouds, and vehicle trajectories, into a unified feature space. Concurrently, we present trigonometry-induced adaptive regression to elevate orientation estimation, while mitigating issues related to scale imbalance and boundary confusion through dual-objective matching with spatial adaptation. To evaluate our methodology, we assemble the first-of-its-kind multi-modal benchmark tailored for complex low-speed environments, complete with fine-grained semantic annotations for intersections. Comprehensive empirical analyses, including ablation studies, affirm both the superior performance of our proposed framework and the efficacy of its constituent modules. Qibin He 0001, Zhongyang Xiao, Ze Huang, Hongyuan Yuan, Li Sun 0005 |
ICRA | 1 |
| 2024 | DLC: Dynamic Loss Correction for Cross-Domain Remotely Sensed SegmentationabstractDue to the diversity of acquisition conditions and imaging mechanisms in remote sensing, the generalization of semantic segmentation models trained with labeled data in the source domain to other unlabeled target domains is hindered. Existing mainstream self-training-based methods provide pseudo-labels to target data as ground truth to utilize target domain evidence for unsupervised domain adaptation (UDA). However, the label shift and domain gap between different domains inevitably introduce noise into pseudo-labeled target data, that is, misclassified pixels. As a consequence, we present a dynamic loss correction (DLC) framework for cross-domain semantic segmentation, which mitigates domain discrepancy by formally modeling the noise distribution of pseudo-labels in the target domain with noise transition matrix (NTM). Specifically, to promote the model output to fit the true label distribution, we employ the high-order consistency information of neighbor representations to estimate NTM and correct the supervision signal without heuristically setting anchors. Furthermore, smooth geometric constraints are introduced to regularize the mutual improvement of NTM derivation and segmentation model optimization in a data-driven manner, thereby compensating for the lack of target domain knowledge. Extensive experimental results on four cross-domain remotely sensed segmentation tasks highlight the generalization capability and competitiveness of the presented method, including cross-scene, cross-band, and cross-modal transfer. Our results and code are available athttps://github.com/heqibin/dlc. Qibin He 0001, Wenhui Diao, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | A Self-Supervised Cross-Modal Remote Sensing Foundation Model with Multi-Domain Representation and Cross-Domain FusionabstractThe construction of a basic model to extract generalized features from a large number of multimodal data is a new challenge in the field of remote sensing. Compared with natural scene images, When faced with a complex application scenario of remote sensing of multi-sensor acquisition, models that are suitable for a specific task are difficult to generalize to new scenarios. In this paper, we propose a model architecture based on the concepts of multi-domain representation and cross-domain fusion. By extracting strong generalization features from massive multi-modal data, a single foundation model can accomplish generalization interpretation for multiple downstream tasks. Experimental results show that the proposed model performs well on multiple downstream tasks, which validates the feasibility of the remote sensing cross-modal foundation model in the interpretation task. Yingchao Feng, Peijin Wang, Wenhui Diao, Qibin He 0001, Huiyang Hu, Hanbo Bi, Xian Sun 0001, Kun Fu 0001 |
IGARSS | 4 |
| 2023 | RingMo: A Remote Sensing Foundation Model With Masked Image ModelingabstractDeep learning approaches have contributed to the rapid development of remote sensing (RS) image interpretation. The most widely used training paradigm is to use ImageNet pretrained models to process RS data for specified tasks. However, there are issues such as domain gap between natural and RS scenes and the poor generalization capacity of RS models. It makes sense to develop a foundation model with general RS feature representation. Since a large amount of unlabeled data is available, the self-supervised method has more development significance than the fully supervised method in RS. However, most of the current self-supervised methods use contrastive learning, whose performance is sensitive to data augmentation, additional information, and selection of positive and negative pairs. In this article, we leverage the benefits of generative self-supervised learning (SSL) for RS images and propose an RS foundationmodel framework called RingMo, which consists of two parts. First, a large-scale dataset is constructed by collecting two million RS images from satellite and aerial platforms, covering multiple scenes and objects around the world. Second, we propose an RS foundation model training method designed for dense and small objects in complicated RS scenes. We show that the foundation model trained on our dataset with RingMo method achieves state-of-the-art (SOTA) on eight datasets across four downstream tasks, demonstrating the effectiveness of the proposed framework. Through in-depth exploration, we believe it is time for RS researchers to embrace generative SSL and leverage its general representation capabilities to speed up the development of RS applications. Xian Sun 0001, Peijin Wang, Wanxuan Lu, Zicong Zhu, Qibin He 0001, Junxi Li, Xuee Rong, Zhujun Yang, Qinglin He, Ruiping Wang 0001, Jiwen Lu, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | CODet: Component Object Detector Extracting Structural Features Based on Target CharacteristicsabstractDeep learning technology has promoted the object detection task in the remote sensing (RS) field to move toward better performance and more demanding requirements. Except for rigid body objects, component objects (COs) with more complex characteristics remain a detection challenge. Its “partial rules and overall disorder” characteristic limits the model learning ability to the structural features. And the internal noise and relatively sparse arrangement are not conducive to optimizing the model by the existing sample assignment strategies. We propose CODet to detect COs in RS scenes. It consists of a cross-hierarchy feature fusion module (CFM) and a noise-sparse sample assignment (NSA) strategy. CFM learns the potential representation and relative position relationship of components by fusing different level features. NSA redefines the optimization process of sample assignment. It aims to alleviate the problems of classification–localization misalignment (CLM) and the positive–negative sample imbalance (PNI) caused by the object’s internal noise and sparse arrangement. The method is verified on the proposed COD dataset of six categories of COs, reaching an average mAP/mAP50of 54.3/86.0. To be closer to the task requirements of the practical RS scene, we also propose a RS large-scale images inference framework. It includes a dataset (APRoI, labeled with COs and rigid body objects), a large-scale image inference strategy, and a set of evaluation metrics. With CODet as the core, the framework can effectively reduce the inference time by three to four times on images with an average of more than 100 million pixels. Zicong Zhu, Xian Sun 0001, Wenhui Diao, Kaiqiang Chen, Qibin He 0001, Guangluan Xu, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Multimodal Remote Sensing Image Segmentation With Intuition-Inspired Hypergraph ModelingabstractMultimodal remote sensing (RS) image segmentation aims to comprehensively utilize multiple RS modalities to assign pixel-level semantics to the studied scenes, which can provide a new perspective for global city understanding. Multimodal segmentation inevitably encounters the challenge of modeling intra- and inter-modal relationships, $i.e$ ., object diversity and modal gaps. However, the previous methods are usually designed for a single RS modality, limited by the noisy collection environment and poor discrimination information. Neuropsychology and neuroanatomy confirm that the human brain performs the guiding perception and integrative cognition of multimodal semantics through intuitive reasoning. Therefore, establishing a semantic understanding framework inspired by intuition to realize multimodal RS segmentation becomes the main motivation of this work. Drived by the superiority of hypergraphs in modeling high-order relationships, we propose an intuition-inspired hypergraph network ( $I^{2}HN$ ) for multimodal RS segmentation. Specifically, we present a hypergraph parser to imitate guiding perception to learn intra-modal object-wise relationships. It parses the input modality into irregular hypergraphs to mine semantic clues and generate robust mono-modal representations. In addition, we also design a hypergraph matcher to dynamically update the hypergraph structure from the explicit correspondence of visual concepts, similar to integrative cognition, to improve cross-modal compatibility when fusing multimodal features. Extensive experiments on two multimodal RS datasets show that the proposed $I^{2}HN$ outperforms the state-of-the-art models, achieving F1/mIoU accuracy 91.4%/82.9% on the ISPRS Vaihingen dataset, and 92.1%/84.2% on the MSAW dataset. Qibin He 0001, Xian Sun 0001, Wenhui Diao, Fanglong Yao, Kun Fu 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | GAL: Graph-Induced Adaptive Learning for Weakly Supervised 3D Object DetectionabstractWeakly Supervised 3D Object Detection (WS3DOD) aims to perform 3D object detection with little reliance on 3D labels, which greatly reduces the cost of 3D annotations. In recent literature, the pseudo-label-based approach brings impressive performance, which generates 3D pseudo-labels from 2D bounding boxes. Despite their success, two key issues remain unresolved that reduce the quality of 3D pseudo-labels: 1) the existing local object locating algorithm can not capture complete clusters of points globally, and 2) the existing algorithm can not capture sparse points caused by the unevenly distributed points obtained by LiDAR cameras. Hence, we propose GAL, a Graph-induced Adaptive Learning algorithm, to generate 3D pseudo-labels. First, we propose the Cluster Locating algorithm based on the Minimum Spanning Tree (MST) to globally locate the objects, which can leverage the characteristic that points inside an object are compact while points between objects are discrete. Second, we propose a density-guided adaptive learning algorithm to optimise the Cluster Locating algorithm, named Cuboid Drift. Cuboid Drift considers the inhomogeneous distribution of reflected points on different reflective surfaces of LiDAR imaging. Finally, 3D pseudo-labels generated by GAL are leveraged to train 3D detectors. Extensive experiments on the challenging KITTI and DAIR-V2X-V dataset demonstrate that GAL without 3D labels can be comparable with strongly supervised approaches and outperforms the previous state-of-the-art WS3DOD methods. Moreover, our method saves 88% of the time spent on pseudo-label generation. Dongshuo Yin, Nayu Liu, Fanglong Yao, Qibin He 0001, Shiyao Yan, Xian Sun 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | DABNet: Deformable Contextual and Boundary-Weighted Network for Cloud Detection in Remote Sensing ImagesabstractIn recent years, deep convolutional neural networks (DCNNs) have made significant progress in cloud detection tasks, and the detection accuracy has been greatly improved. However, most existing CNN-based models have high computational complexity, which limits their practical application, especially for spaceborne optical remote sensing. In addition, most of the methods cannot make adaptive adjustments based on the structural information of the clouds, and blurred boundaries often occur in the detection results. In order to address these problems, this article proposes a lightweight network (DABNet) to achieve high-accuracy detection of complex clouds, not only a clearer boundary but also lower false-alarm rate. Specifically, a deformable context feature pyramid module is proposed to improve the adaptive modeling capability of multiscale features. Besides, a boundary-weighted loss function is designed to direct the network to focus on cloud boundary information and optimize the relevant detection results. The proposed method has been validated on two data sets: the public GF-1 WFV benchmark and our self-built GF-2 cloud detection data set with higher spatial resolution. The experimental results exhibit that DABNet achieves state-of-the-art performance while only using 4.12M parameters and 8.29G multiadds. Qibin He 0001, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Multi-Object Tracking in Satellite Videos With Graph-Based Multitask ModelingabstractRecently, satellite video has become an emerging means of earth observation, providing the possibility of tracking moving objects. However, the existing multi-object trackers are commonly designed for natural scenes without considering the characteristics of remotely sensed data. In addition, most trackers are composed of two independent stages of detection and reidentification (ReID), which means that they cannot be mutually promoted. To this end, we propose an end-to-end online framework, which is called TGraM, for multi-object tracking in satellite videos. It models multi-object tracking as a graph information reasoning procedure from the multitask learning perspective. Specifically, a graph-based spatiotemporal reasoning module is presented to mine the potential high-order correlations between video frames. Furthermore, considering the inconsistency of optimization objectives between detection and ReID, a multitask gradient adversarial learning strategy is designed to regularize each task-specific network. In addition, aiming at the data scarcity in this field, a large-scale and high-resolution Jilin-1 satellite video dataset for multi-object tracking (AIR-MOT) is built for the experiments. Compared with state-of-the-art multi-object trackers, TGraM achieves efficient collaborative learning between detection and ReID, improving the tracking accuracy by 1.2 multiple object tracking accuracy. The code and dataset will be available online (https://github.com/HeQibin/TGraM). Qibin He 0001, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |