VLDB 2026 Research / reviewers in the wild / expert
Shunping Ji
dblp:123/0960
· DBLP profile ↗
31ranked-venue papers
3as first author
25since 2021 · last 2026
0000-0002-3088-1481ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 12 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Opt3DGS: Optimizing 3D Gaussian Splatting with Adaptive Exploration and Curvature-Aware Exploitationabstract3D Gaussian Splatting (3DGS) has emerged as a leading framework for novel view synthesis, yet its core optimization challenges remain underexplored. We identify two key issues in 3DGS optimization: entrapment in suboptimal local optima and insufficient convergence quality. To address these, we propose Opt3DGS, a robust framework that enhances 3DGS through a two-stage optimization process of adaptive exploration and curvature-guided exploitation. In the exploration phase, an Adaptive Weighted Stochastic Gradient Langevin Dynamics (SGLD) method enhances global search to escape local optima. In the exploitation phase, a Local Quasi-Newton Direction-guided Adam optimizer leverages curvature information for precise and efficient convergence. Extensive experiments on diverse benchmark datasets demonstrate that Opt3DGS achieves state-of-the-art rendering quality by refining the 3DGS optimization process without modifying its underlying representation. Jiagang Chen, Shunping Ji |
AAAI | 4 |
| 2025 | Point Cloud Mamba: Point Cloud Learning via State Space ModelabstractRecently, state space models have exhibited strong global modeling capabilities and linear computational complexity in contrast to transformers. This research focuses on applying such architecture to more efficiently and effectively model point cloud data globally with linear computational complexity. In particular, for the first time, we demonstrate that Mamba-based point cloud methods can outperform previous methods based on transformer or multi-layer perceptrons (MLPs). To enable Mamba to process 3-D point cloud data more effectively, we propose a novel Consistent Traverse Serialization method to convert point clouds into 1-D point sequences while ensuring that neighboring points in the sequence are also spatially adjacent. Consistent Traverse Serialization yields six variants by permuting the order of x, y, and z coordinates, and the synergistic use of these variants aids Mamba in comprehensively observing point cloud data. Furthermore, to assist Mamba in handling point sequences with different orders more effectively, we introduce point prompts to inform Mamba of the sequence’s arrangement rules. Finally, we propose positional encoding based on spatial coordinate mapping to inject positional information into point cloud sequences more effectively. Point Cloud Mamba surpasses the state-of-the-art (SOTA) point-based method PointNeXt and achieves new SOTA performance on the ScanObjectNN, ModelNet40, ShapeNetPart, and S3DIS datasets. It is worth mentioning that when using a more powerful local feature extraction module, our PCM achieves 79.6 mIoU on S3DIS, significantly surpassing the previous SOTA models, DeLA and PTv3, by 5.5 mIoU and 4.9 mIoU, respectively. Tao Zhang 0042, Haobo Yuan, Lu Qi 0001, Jiangning Zhang, Qianyu Zhou 0001, Shunping Ji, Shuicheng Yan, Xiangtai Li |
AAAI | 6 |
| 2025 | Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMsabstractRecent advancements in multimodal large language models (MLLM) have shown a strong ability in visual perception, reasoning abilities, and vision-language understanding. However, the visual matching ability of MLLMs is rarely studied, despite finding the visual correspondence of objects is essential in computer vision. Our research reveals that the matching capabilities in recent MLLMs still exhibit systematic shortcomings, even with current strong MLLMs models, GPT-4o. In particular, we construct a Multimodal Visual Matching (MMVM) benchmark to fairly benchmark over 30 different MLLMs. The MMVM benchmark is built from 15 open-source datasets and Internet videos with manual annotation. We categorize the data samples of MMVM benchmark into eight aspects based on the required cues and capabilities to more comprehensively evaluate and analyze current MLLMs. In addition, we have designed an automatic annotation pipeline to generate the MMVM SFT dataset, including 220K visual matching data with reasoning annotation. To our knowledge, this is the first visual corresponding dataset and benchmark for the MLLM community. Finally, we present CoLVA, a novel contrastive MLLM with two novel technical designs: fine-grained vision expert with object-level contrastive learning and instruction augmentation strategy. The former learns instance discriminative tokens, while the latter further improves instruction following ability. CoLVA-InternVL2-4B achieves an overall accuracy (OA) of 49.80\% on the MMVM benchmark, surpassing GPT-4o and the best open-source MLLM, Qwen2VL-72B, by 7.15\% and 11.72\% OA, respectively. These results demonstrate the effectiveness of our MMVM SFT dataset and our novel technical designs. Code, benchmark, dataset, and models will be released. Yikang Zhou, Tao Zhang 0042, Shilin Xu 0001, Shihao Chen, Qianyu Zhou 0001, Yunhai Tong, Shunping Ji, Jiangning Zhang, Lu Qi 0001, Xiangtai Li |
ICCV | 7 |
| 2025 | SIU3R: Simultaneous Scene Understanding and 3D Reconstruction Beyond Feature AlignmentabstractSimultaneous understanding and 3D reconstruction plays an important role in developing end-to-end embodied intelligent systems.
To achieve this, recent approaches resort to 2D-to-3D feature alignment paradigm, which leads to limited 3D understanding capability and potential semantic information loss.
In light of this, we propose SIU3R, the first alignment-free framework for generalizable simultaneous understanding and 3D reconstruction from unposed images.
Specifically, SIU3R bridges reconstruction and understanding tasks via pixel-aligned 3D representation, and unifies multiple understanding tasks into a set of unified learnable queries, enabling native 3D understanding without the need of alignment with 2D models.
To encourage collaboration between the two tasks with shared representation, we further conduct in-depth analyses of their mutual benefits, and propose two lightweight modules to facilitate their interaction.
Extensive experiments demonstrate that our method achieves state-of-the-art performance not only on the individual tasks of 3D reconstruction and understanding, but also on the task of simultaneous understanding and 3D reconstruction, highlighting the advantages of our alignment-free framework and the effectiveness of the mutual benefit designs. Dongxu Wei, Lingzhe Zhao, Wenpu Li, Zhangchi Huang, Shunping Ji, Peidong Liu 0001 |
NeurIPS | 6 |
| 2025 | DVIS++: Improved Decoupled Framework for Universal Video SegmentationabstractWe present the Decoupled VIdeo Segmentation (DVIS) framework, a novel approach for the challenging task of universal video segmentation, including video instance segmentation (VIS), video semantic segmentation (VSS), and video panoptic segmentation (VPS). Unlike previous methods that model video segmentation in an end-to-end manner, our approach decouples video segmentation into three cascaded sub-tasks: segmentation, tracking, and refinement. This decoupling design allows for simpler and more effective modeling of the spatio-temporal representations of objects, especially in complex scenes and long videos. Accordingly, we introduce two novel components: the referring tracker and the temporal refiner. These components track objects frame by frame and model spatio-temporal representations based on pre-aligned features. To improve the tracking capability of DVIS, we propose a denoising training strategy and introduce contrastive learning, resulting in a more robust framework named DVIS++. The proposed decoupled framework efficiently handles universal and open-vocabulary object representations, allowing DVIS++ to conduct universal and open-vocabulary video segmentation. We conduct extensive experiments on six mainstream benchmarks, including the VIS, VSS, and VPS datasets. Using a unified architecture, DVIS++ significantly outperforms state-of-the-art specialized methods on these benchmarks in closed- and open-vocabulary settings. Tao Zhang 0042, Xingye Tian, Yikang Zhou, Shunping Ji, Xuebo Wang, Xin Tao 0001, Yuan Zhang 0020, Pengfei Wan 0001, Zhongyuan Wang 0006, Yu Wu 0011 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | P2PFormerV2: Improving Primitive-Based Regular Building Contour Extraction Methods via Contour Feature EnhancementabstractDeep learning methods have been widely used to map building contours automatically in high-resolution remote sensing images over recent years. However, most deep learning-based methods still require complex post-processing to generate regular contours. P2PFormer is an advanced method that can directly obtain the positions and sequences of general geometric primitives like points, lines, and angles (corners) of a building instance without complex post-processing. Nevertheless, due to the inherent limitations of ROI-Align, P2PFormer faces significant challenges in primitive detection. ROI-Align utilizes a uniform sampling approach for feature extraction, it inevitably introduces a high proportion of invalid sampling points into the extracted features. Resulting in missed and false detections of building primitives, ultimately affecting the accuracy of building contour extraction. To address these issues, we propose P2PFormerV2, which introduces a contour feature enhancer to improve P2PFormer. The contour feature enhancer increases the proportion of valid feature sampling points and enhances the extraction of contour-aware features, significantly improving the accuracy of primitive segmentation. This enhancer comprises three key components: the sparse feature extractor, the dense feature extractor, and the feature fusion module. The sparse feature extractor optimizes the sampling strategy to increase the proportion of valid feature sampling points; the dense feature extractor generates rich contour features and provides additional supervision signals; the feature fusion module integrates the outputs of the first two components to further enhance the extraction of contour features. Experimental results demonstrate that P2PFormerV2 achieves average precisions (AP) of 74.7%, 79.6%, and 64.2% in the WHU, CrowdAI, and WHU-Mix datasets, respectively, significantly outperforming the original P2PFormer and other existing advanced methods. Our findings about the shortcomings of ROI-Align and the importance of improving the effective feature extraction provides insights for future building extraction research. Wenling Yu, Tao Zhang 0042, Shunping Ji, Bo Liu 0068, Hua Liu 0002, Jianya Gong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Improving Video Segmentation via Dynamic Anchor Queries
Yikang Zhou, Tao Zhang 0042, Shunping Ji, Shuicheng Yan, Xiangtai Li |
ECCV (50) | 3 |
| 2024 | OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and UnderstandingabstractCurrent universal segmentation methods demonstrate strong capabilities in pixel-level image and video understanding. However, they lack reasoning abilities and cannot be controlled via text instructions. In contrast, large vision-language multimodal models exhibit powerful vision-based conversation and reasoning capabilities but lack pixel-level understanding and have difficulty accepting visual prompts for flexible user interaction. This paper proposes OMG-LLaVA, a new and elegant framework combining powerful pixel-level vision understanding with reasoning abilities. It can accept various visual and text prompts for flexible user interaction. Specifically, we use a universal segmentation method as the visual encoder, integrating image information, perception priors, and visual prompts into visual tokens provided to the LLM. The LLM is responsible for understanding the user's text instructions and providing text responses and pixel-level segmentation results based on the visual information. We propose perception prior embedding to better integrate perception priors with image features. OMG-LLaVA achieves image-level, object-level, and pixel-level reasoning and understanding in a single model, matching or surpassing the performance of specialized methods on multiple benchmarks. Rather than using LLM to connect each specialist, our work aims at end-to-end training on one encoder, one decoder, and one LLM. The code and model have been released for further research. Tao Zhang 0042, Xiangtai Li, Hao Fei 0001, Haobo Yuan, Shengqiong Wu, Shunping Ji, Chen Change Loy, Shuicheng Yan |
NeurIPS | 6 |
| 2024 | SAM-RSIS: Progressively Adapting SAM With Box Prompting to Remote Sensing Image Instance SegmentationabstractThe recent segment anything model (SAM) trained on massive close-range images has demonstrated impressive performance on general segmentation or specific segmentation tasks with manual prompts. However, the significant domain shift problem between remote sensing and close-range images should be tackled before introducing the pretrained SAM to remote sensing instance segmentation (RSIS). To address this and unlock the potential of SAM in RSIS, this article proposes a novel framework called SAM for remote sensing instance segmentation (SAM-RSIS), which overcomes the problems in a few recent works that only adapt a part of SAM to remote sensing. SAM-RSIS fine-tunes the vision transformer (ViT) backbone and mask decoder of SAM progressively on remote sensing data and uses automatic box prompting to eliminate the need for manual prompting. SAM-RSIS consists of an object detection stage and a mask generation stage. In object detection, we introduce an adapter to adapt knowledge embedded in the pretrained ViT backbone to remote sensing images and then build an object detector. In mask generation, using the detected bounding boxes as prompts, along with two learnable mask output tokens, and the two-layer high-resolution features from the adapter, we fine-tune the mask decoder of SAM to produce high-quality masks. Experimental results on the WHU, WHU-Mix, and NWPU datasets for binary and multiclass RSIS demonstrate the effectiveness and robustness of the proposed method, surpassing various derivative methods of SAM and achieving performance comparable to and even better than the specific state-of-the-art instance segmentation methods. Muying Luo, Tao Zhang 0042, Shiqing Wei, Shunping Ji |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | 3-D Building Instance Extraction From High-Resolution Remote Sensing Images and DSM With an End-to-End Deep Neural NetworkabstractThree-dimensional (3D) building models play a vital role in numerous applications including urban planning and smart cities. Recent 3D building modeling methods either rely heavily on available manaually-collected footprint reference or hardly reach real automation on par with manual editing. To approach the automated extraction of instance-level 3D buildings at Level of Detail (LoD) 1, we introduce an innovative end-to-end 3D building instance segmentation model. This model predicts accurate contours and heights of individual buildings simultaneously using ortho-rectified high-resolution remote sensing images and Digital Surface Models (DSMs), getting rid of additional reference data and impirical parameter settings. Firstly, we propose an Anchor-Free Multi-head building extraction network (AFM) tailored for extracting 2D building contours. AFM incorporates a full-resolution, long-range correlation boosted global mask prediction branch along with anchor-free bounding box generation, as well as a newly developed online hard sample mining (OHSM) training procedure based on uncertainty analysis to emphasize error-prone positions in locating building contours. Subsequently, we incorporate a height prediction component to AFM in order to derive accurate building height information, thus creating the comprehensive 3D building extraction model referred to as AFM-3D. The two-stage AFM-3D operates by initially predicting 3D cube proposals, followed by generating refined 3D prismatic models (LoD1 models) for each proposal. Thorough experimentation across different datasets demonstrates the superior performance of AFM and AFM-3D. A significant enhancement of 6.4% quality score is observed on the urban 3D dataset in comparison to recent methods. In addition to the proposed novel methodology, we compare anchor-based and anchor-free bounding box generation mechanisms for remote sensing data, explore pixel-based and contour-based segmentation strategies, evaluate learning-based and empirical height estimation methods, and discuss the indispensability of DSM data in 3D building instance extraction. These analyses yield valuable insights that contribute to the progression of 3D building extraction research. Dawen Yu, Shunping Ji, Shiqing Wei, Kourosh Khoshelham |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | P2PFormer: A Primitive-to-Polygon Method for Regular Building Contour Extraction From Remote Sensing ImagesabstractExtracting building contours from remote sensing imagery is a significant challenge due to buildings’ complex and diverse shapes, occlusions, and noise. Existing methods often struggle with irregular contours, rounded corners, and redundancy points, necessitating extensive postprocessing to produce regular polygonal building contours. To address these challenges, we introduce a novel, streamlined pipeline that generates regular building contours without postprocessing. Our approach begins with the segmentation of generic geometric primitives (which can include vertices, lines, and corners), followed by the prediction of their sequence. This allows for the direct construction of regular building contours by sequentially connecting the segmented primitives. Building on this pipeline, we developed primitive-to-polygon using transformer (P2PFormer), which uses a transformer-based architecture to segment geometric primitives and predict their order. To enhance the segmentation of primitives, we introduce a unique representation called group queries. This representation comprises a set of queries and a singular query position, which improve the focus on multiple midpoints of primitives and their efficient linkage. Furthermore, we propose an innovative implicit update strategy for the query position embedding aimed at sharpening the focus of queries on the correct positions and, consequently, enhancing the quality of primitive segmentation. Our experiments demonstrate that P2PFormer achieves new state-of-the-art (SOTA) performance on the WHU, CrowdAI, and WHU-Mix datasets, surpassing the previous SOTA PolyWorld by a margin of 2.7 AP and 6.5 AP75 on the largest CrowdAI dataset. We intend to make the code and trained weights publicly available to promote their use and facilitate further research. Tao Zhang 0042, Shiqing Wei, Yikang Zhou, Muying Luo, Wenling Yu, Shunping Ji |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | DVIS: Decoupled Video Instance Segmentation FrameworkabstractVideo instance segmentation (VIS) is a critical task with diverse applications, including autonomous driving and video editing. Existing methods often underperform on complex and long videos in real world, primarily due to two factors. Firstly, offline methods are limited by the tightly-coupled modeling paradigm, which treats all frames equally and disregards the interdependencies between adjacent frames. Consequently, this leads to the introduction of excessive noise during long-term temporal alignment. Secondly, online methods suffer from inadequate utilization of temporal information. To tackle these challenges, we propose a decoupling strategy for VIS by dividing it into three independent sub-tasks: segmentation, tracking, and refinement. The efficacy of the decoupling strategy relies on two crucial elements: 1) attaining precise long-term alignment outcomes via frame-by-frame association during tracking, and 2) the effective utilization of temporal information predicated on the aforementioned accurate alignment outcomes during refinement. We introduce a novel referring tracker and temporal refiner to construct the Decoupled VIS framework (DVIS). DVIS achieves new SOTA performance in both VIS and VPS, surpassing the current SOTA methods by 7.3 AP and 9.6 VPQ on the OVIS and VIPSeg datasets, which are the most challenging and realistic benchmarks. Moreover, thanks to the decoupling strategy, the referring tracker and temporal refiner are super light-weight (only 1.69% of the segmenter FLOPs), allowing for efficient training and inference on a single GPU with 11G memory. The code is available at https://github.com/zhang-tao-whu/DVIS. Tao Zhang 0042, Xingye Tian, Yu Wu 0011, Shunping Ji, Xuebo Wang, Yuan Zhang 0020, Pengfei Wan 0001 |
ICCV | 4 |
| 2023 | From Image Transfer to Object Transfer: Cross-Domain Instance Segmentation Based on Center Point Feature AlignmentabstractRemote sensing images can have significant appearance differences due to various factors such as atmospheric conditions, sensor types, seasons, and capture times. Therefore, when applying a pre-trained instance segmentation deep learning model to newly accessed remote sensing images, the model’s performance tends to decrease significantly. Current mainstream image-based or feature-based domain adaptation methods are not designed specifically for the cross-domain instance segmentation problem. These methods attempt to align the whole images, which may not be optimal for instance segmentation tasks. To address this issue, we propose a cross-domain instance segmentation method based on object-level alignment. Instead of aligning the entire images from both datasets, we only align the features of each object instance, particularly the representative center point features. Our approach mainly consists of an improved contour-based instance segmentation model for object-based domain adaptation, an object-pasting enhancement technique based on Fourier domain adaptation (FDA) that effectively reduces the gap between the source and target domains of the object instances, and a self-training strategy that dynamically generates pseudo-labels for iterative model training. Our experiments on cross-domain building instance segmentation demonstrate that the proposed method achieves a 9.5 intersection over union (IoU) improvement over the current best method. Additionally, experiments on a cross-domain close-range dataset involving transfer between simulated and real street images show that our method significantly outperforms the current best method by 6.5 mean average precision (mAP). These results on remote sensing and close-range datasets validate the universality of our approach. Shunping Ji, Tao Zhang 0042 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Long-Range Correlation Supervision for Land-Cover Classification From Remote Sensing ImagesabstractLong-range dependency modeling has been widely considered in modern deep learning-based semantic segmentation methods, especially those designed for large-size remote sensing images, to compensate the intrinsic locality of standard convolutions. However, in previous studies, the long-range dependency, modeled with an attention mechanism or transformer model, has been based on unsupervised learning, instead of explicit supervision from the objective ground truth (GT). In this article, we propose a novel supervised long-range correlation method for land-cover classification, called the supervised long-range correlation network (SLCNet), which is shown to be superior to the currently used unsupervised strategies. In SLCNet, pixels sharing the same category are considered highly correlated and those having different categories are less relevant, which can be easily supervised by the category consistency information available in the GT semantic segmentation map. Under such supervision, the recalibrated features are more consistent for pixels of the same category and more discriminative for pixels of other categories, regardless of their proximity. To complement the detailed information lacking in the global long-range correlation, we introduce an auxiliary adaptive receptive field feature extraction (ARFE) module, parallel to the long-range correlation module in the encoder, to capture finely detailed feature representations for multisize objects in multiscale remote sensing images. In addition, we apply multiscale side-output supervision and a hybrid loss function as local and global constraints to further boost the segmentation accuracy. Experiments were conducted on three public remote sensing datasets (the ISPRS Vaihingen dataset, the ISPRS Potsdam dataset, and the DeepGlobe dataset). Compared with the advanced segmentation methods from the computer vision, medicine, and remote sensing communities, the proposed SLCNet method achieved state-of-the-art performance on all the datasets. The code will be made available at gpcv.whu.edu.cn/data. Dawen Yu, Shunping Ji |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | E2EC: An End-to-End Contour-based Method for High-Quality High-Speed Instance SegmentationabstractContour-based instance segmentation methods have developed rapidly recently but feature rough and hand-crafted front-end contour initialization, which restricts the model performance, and an empirical and fixed backend predicted-label vertex pairing, which contributes to the learning difficulty. In this paper, we introduce a novel contour-based method, named E2EC, for high-quality instance segmentation. Firstly, E2EC applies a novel learnable contour initialization architecture instead of hand-crafted contour initialization. This consists of a contour initialization module for constructing more explicit learning goals and a global contour deformation module for taking advantage of all of the vertices' features better. Secondly, we propose a novel label sampling scheme, named multi-direction alignment, to reduce the learning difficulty. Thirdly, to improve the quality of the boundary details, we dynamically match the most appropriate predicted-ground truth vertex pairs and propose the corresponding loss function named dynamic matching loss. The experiments showed that E2EC can achieve a state-of-the-art performance on the KITTI INStance (KINS) dataset, the Semantic Boundaries Dataset (SBD), the Cityscapes and the COCO dataset. E2EC is also efficient for use in real-time applications, with an inference speed of 36 fps for$512\times 512$images on an NVIDIA A6000 GPU. Code will be released at https://github.com/zhang-tao-whu/e2ec. Tao Zhang 0042, Shiqing Wei, Shunping Ji |
CVPR | 3 |
| 2022 | Graph Convolutional Networks for the Automated Production of Building Vector Maps From Aerial ImagesabstractBuilding footprint delineation from remote sensing imagery is a basic task in surveying and mapping and geographic information system (GIS), which benefits many engineering applications but requires an enormous amount of human delineation. In this study, we aim to find a way to replace human delineation with automated algorithms. For this purpose, we designed a novel pipeline for the automated production of building vector maps from aerial images, which consists of a bounding-box generation module, a graph convolutional network (GCN)-based polygon prediction module, and an empirical polygon regularization module. First, we introduce the bounding box generation module based on region-based object detection, which along with our overlap cropping strategy is used to generate a bounding box for each building instance. The generated bounding boxes have a horizontal rectangle version and a rotated version, which are used to initialize the next stage of our method. Second, we propose a GCN-based method, which is the core of this study, to conduct the initial accurate building polygon prediction by integrating multiresolution optimization and multilevel loss constraints. Our proposed method, to the authors’ best knowledge, is the first of its kind that introduces GCN to building extraction. The final step is to apply a regularization algorithm to translate the predicted polygons into structured and highly accurate building footprints. We validated the proposed method on two large aerial building data sets, WHU data set and Inria data set, where it was shown to significantly outperform other state-of-the-art methods more than 10%. Specifically, with the ground-truth rotated bounding boxes of buildings, our method is able to automatically delineate 91% of the buildings in the WHU data set and with the predicted rotated bounding the percent reached human-level delineation. Shiqing Wei, Shunping Ji |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Scribble-Based Weakly Supervised Deep Learning for Road Surface Extraction From Remote Sensing ImagesabstractRoad surface extraction from remote sensing images using deep learning methods has achieved good performance, while most of the existing methods are based on fully supervised learning, which requires a large amount of training data with laborious per-pixel annotation. In this article, we propose a scribble-based weakly supervised road surface extraction method named ScRoadExtractor, which learns from easily accessible scribbles such as centerlines instead of densely annotated road surface ground truths. To propagate semantic information from sparse scribbles to unlabeled pixels, we introduce a road label propagation algorithm, which considers both the buffer-based properties of road networks and the color and spatial information of super-pixels, to produce a proposal mask with categories road, nonroad, and unknown. The proposal mask, along with the auxiliary boundary prior information detected from images, is utilized to train a dual-branch encoder–decoder network which we designed for precise road surface segmentation. We perform experiments on three diverse road data sets that are comprised of high-resolution remote sensing satellite and aerial images across the world. The results demonstrate that ScRoadExtractor exceeds the classic scribble-supervised segmentation method by 20% for the intersection over union (IoU) indicator and outperforms the state-of-the-art scribble-based weakly supervised methods at least 4%. Shunping Ji |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A Concentric Loop Convolutional Neural Network for Manual Delineation-Level Building Boundary Segmentation From Remote-Sensing ImagesabstractTo date, accurate building footprint delineation in the surveying, mapping, and geographic information system (GIS) communities has been dependent on human labor. In this article, to address this issue, we propose a concentric loop convolutional neural network (CLP-CNN) method for the automatic segmentation of building boundaries from remote-sensing images. The proposed method consists of three components: 1) a boundary detector to extract coarse polygonal boundaries of individual regions of interest; 2) a concentric loop-shaped convolutional network with bidirectional pairing loss to fine-tune the vertices of the polygons; and 3) a refinement block, which removes redundant vertices and regularizes the boundaries to polygons at the manual delineation level. We also demonstrate that the proposed CLP-CNN method is applicable to other generic objects in natural images. Experiments on two building datasets confirmed that more than 77%/67% of the building polygons predicted by the proposed method are on par with the manual delineation level, representing a significant saving in the labor cost of manual annotation. In generic object boundary delineation tests performed on the Semantic Boundaries Dataset (SBD), the proposed method outperformed the most recent state-of-the-art methods by at least 3.1% in average precision (AP). Furthermore, compared with other vertex matching methods, the learning process of the proposed method converges faster. The source code will be available athttp://gpcv.whu.edu.cn/data. Shiqing Wei, Tao Zhang 0042, Shunping Ji |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | A Combination of Convolutional and Graph Neural Networks for Regularized Road Surface ExtractionabstractRoad surface extraction from high-resolution remote sensing images has many engineering applications; however, extracting regularized and smooth road surface maps that reach the human delineation level is a very challenging task, and substantial and time-consuming manual work is usually unavoidable. In this article, to solve this problem, we propose a novel regularized road surface extraction framework by introducing a graph neural network (GNN) for processing the road graph that is preconstructed from the easily accessible road centerlines. The proposed framework formulates the road surface extraction problem as two-sided width inference of the road graph and consists of a convolutional neural network (CNN)-based feature extractor and a GNN model for vertex attribute adjustment. The CNN extracts the high-level abstract features of each vertex in the graph as the input of the GNN and also the road boundary features that allow us to distinguish roads from the background. The GNN propagates and aggregates the features of the vertices in the graph to achieve global optimization of the regression of the regularized widths of the vertices. At the same time, a biased centerline map can also be corrected based on the width prediction result. To the best of the authors’ knowledge, this is the first study to have introduced a GNN to regularized human-level road surface extraction. The proposed method was evaluated on four diverse datasets, and the results show that the proposed method comprehensively outperforms the recent CNN-based segmentation methods and other regularization methods in the intersection over union (IoU) and smoothness score, and a visual check shows that a majority of the prediction results of the proposed method approach the human delineation level. Shunping Ji |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | A New Spatial-Oriented Object Detection Framework for Remote Sensing ImagesabstractAlthough the orientation and scale properties of the objects in remote sensing images have been widely considered in the modern deep learning-based object detection methods, the spatial distribution property of objects has rarely been investigated. There is a distinct spatial distribution difference between close-range objects and remote sensing objects: the former may exhibit extensive mutual occlusion and overlap, whereas the latter rarely overlap. A current remote sensing object detection algorithm that ignores the spatial distribution difference may unnecessarily apply the massive anchor-based proposal bounding box generation and nonmaximum suppression (NMS) operations. In this article, considering the unique spatial distribution of remote sensing objects, and also the other spatial properties, we propose a novel, compact, and spatial-oriented object detection framework for remote sensing images. The proposed two-stage convolutional neural network (CNN) framework, which we call the Remote-sensing Spatial Adaptation DETector (RSADet), considers the spatial distribution, scale, and orientation/shape varieties of the objects in remote sensing images. In the first stage, each object instance is inferred on the scale-attention boosted CNN heatmaps to generate candidate bounding boxes, instead of using the anchor-based proposal box generation and NMS. In the second stage, deformable convolutions are introduced to adapt to the geometric variations of different object instances and to avoid the impact of complex and changeable backgrounds. A new bounding box confidence (IoU score) prediction branch is introduced as a convenient constraint for eliminating unreliable boxes and improving performance. Experiments were conducted on a large single-class remote sensing object detection dataset (the Ningbo Pylon dataset) built as part of this study and an open-source extraordinarily large multiclass dataset (the object DetectIon in Optical Remote sensing image (DIOR) dataset). Compared with the advanced detectors from both the computer vision and remote sensing communities, the proposed RSADet achieved state-of-the-art performance on both datasets. Dawen Yu, Shunping Ji |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Earthquake Crack Detection From Aerial Images Using a Deformable Convolutional Neural NetworkabstractDetecting the terrain surface cracks caused by earthquakes, which are termed coseismic ruptures, has important significance for discovering concealed faults, monitoring their movements, and forecasting possible follow-on earthquakes. On May 22, 2021, Maduo County in Qinghai province, China, suffered an earthquake with a magnitude of 7.4, which created densely distributed cracks. In this study, we designed an automatic crack detection framework based on remote sensing technology. With the use of low-altitude unmanned aerial vehicles (UAVs), we obtained very high-resolution aerial images of the area affected by the earthquake, which were further processed by photogrammetric software to produce digital orthophoto maps (DOMs). We then designed a novel terrain surface crack detection neural network, which differs from the previous methods that focus on detecting cracks in man-made object surfaces such as flat roads. We investigated the spatial property of the sinuous linear cracks and handled this by introducing adaptive deformable convolutions with a context-channel-space boosted mechanism. The feature extraction stage, feature optimization stage, and upsampling stage were embedded with the deformable convolutions to form a compact and powerful crack detector, named Crack-CADNet (the Context-chAnnel-space boosted Deformable convolutional neural network for crack detection). The postprocessing included filtering out the nontectonic cracks, aided by annotations from experts, and grouping and vectorizing the generated binary segmentation map as crack polygons, which were evaluated at the instance level. In addition to the first in-depth investigation of detecting earthquake cracks with aerial remote sensing and a deep learning based process, the crack detection network we propose outperformed the recent convolutional neural network (CNN)-based methods designed for general semantic segmentation and crack detection. Source code and the Maduo earthquake crack dataset will be available at http://gpcv.whu.edu.cn/data/. Dawen Yu, Shunping Ji, Xue Li 0032, Zhaode Yuan, Chaoyong Shen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Rational Polynomial Camera Model Warping for Deep Learning Based Satellite Multi-View Stereo MatchingabstractSatellite multi-view stereo (MVS) imagery is particularly suited for large-scale Earth surface reconstruction. Differing from the perspective camera model (pin-hole model) that is commonly used for close-range and aerial cameras, the cubic rational polynomial camera (RPC) model is the mainstream model for push-broom linear-array satellite cameras. However, the homography warping used in the prevailing learning based MVS methods is only applicable to pin-hole cameras. In order to apply the SOTA learning based MVS technology to the satellite MVS task for large-scale Earth surface reconstruction, RPC warping should be considered. In this work, we propose, for the first time, a rigorous RPC warping module. The rational polynomial coefficients are recorded as a tensor, and the RPC warping is formulated as a series of tensor transformations. Based on the RPC warping, we propose the deep learning based satellite MVS (SatMVS) framework for large-scale and wide depth range Earth surface reconstruction. We also introduce a large-scale satellite image dataset consisting of 519 5120×5120 images, which we call the TLC SatMVS dataset. The satellite images were acquired from a three-line camera (TLC) that catches triple-view images simultaneously, forming a valuable supplement to the existing open-source WorldView-3 datasets with single-scanline images. Experiments show that the proposed RPC warping module and the SatMVS framework can achieve a superior reconstruction accuracy compared to the pin-hole fitting method and conventional MVS methods. Code and data are available at https://github.com/WHU-GPCV/SatMVS. Jian Gao 0009, Shunping Ji |
ICCV | 3 |
| 2021 | Cascaded Deep Neural Networks for Predicting Biases Between Building Polygons in Vector Maps and New Remote Sensing ImagesabstractTremendous historical and recent building vector maps are available from private surveying and mapping departments or open-source platforms such as Open Street Map (OSM). However, there invariably exist offsets between the building vectors and new remote sensing images, caused by geometric registration errors and parallax from hypsography. Consequently, these vector maps cannot be directly used as labels for the up-to-date GIS productions, as well as popular supervised machine learning. As the first work of this kind, this paper proposes an offset correction method, based on deep learning, to predict the bias between a vector building map and a new remote sensing image. The method is based on a dense regression model to predict the offsets between the image and the vector map at the pixel level, which are then processed to retrieve the offsets of each polygon. Two datasets consist of remote sensing images and biased building polygons, the WHU change detection and London datasets, are prepared to test the effectiveness of our approach. In the WHU change detection dataset, approaching 80% polygons can be correctly registered within the tolerance of three pixels, in the London dataset, more than 60% polygons can be corrected by a pre-trained model on an available open-source dataset, both of which demonstrated that our method can contribute to the GIS map updating and more efficient preparation of training samples for a deep learning model. Mingyang Hu, Shunping Ji |
IGARSS | 3 |
| 2021 | Simultaneous Cloud Detection and Removal From Bitemporal Remote Sensing Images Using Cascade Convolutional Neural NetworksabstractClouds and cloud shadows heavily affect the quality of the remote sensing images and their application potential. Algorithms have been developed for detecting, removing, and reconstructing the shaded regions with the information from the neighboring pixels or multisource data. In this article, we propose an integrated cloud detection and removal framework using cascade convolutional neural networks, which provides accurate cloud and shadow masks and repaired images. First, a novel fully convolutional network (FCN), embedded with multiscale aggregation and the channel-attention mechanism, is developed for detecting clouds and shadows from a cloudy image. Second, another FCN, with the masks of the detected cloud and shadow, the cloudy image, and a temporal image as the input, is used for the cloud removal and missing-information reconstruction. The reconstruction is realized through a self-training strategy that is designed to learn the mapping between the clean-pixel pairs of the bitemporal images, which bypasses the high demand of manual labels. Experiments showed that our proposed framework can simultaneously detect and remove the clouds and shadows from the images and the detection accuracy surpassed several recent cloud-detection methods; the effects of image restoring outperform the mainstream methods in every indicator by a large margin. The data set used for cloud detection and removal is made open. Shunping Ji, Peiyu Dai, Yongjun Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Generative Adversarial Network-Based Full-Space Domain Adaptation for Land Cover Classification From Multiple-Source Remote Sensing ImagesabstractThe accuracy of remote sensing image segmentation and classification is known to dramatically decrease when the source and target images are from different sources; while deep learning-based models have boosted performance, they are only effective when trained with a large number of labeled source images that are similar to the target images. In this article, we propose a generative adversarial network (GAN) based domain adaptation for land cover classification using new target remote sensing images that are enormously different from the labeled source images. In GANs, the source and target images are fully aligned in the image space, feature space, and output space domains in two stages via adversarial learning. The source images are translated to the style of the target images, which are then used to train a fully convolutional network (FCN) for semantic segmentation to classify the land cover types of the target images. The domain adaptation and segmentation are integrated to form an end-to-end framework. The experiments that we conducted on a multisource data set covering more than 3500 km2with 51 560 256×256 high-resolution satellite images in Wuhan city and a cross-city data set with 11 383 256×256 aerial images in Potsdam and Vaihingen demonstrated that our method exceeded the recent GAN-based domain adaptation methods by at least 6.1% and 4.9% in the mean intersection over union (mIoU) and overall accuracy (OA) indexes, respectively. We also proved that our GAN is a generic framework that can be implemented for other domain transfer methods to boost their performance. Shunping Ji, Dingpan Wang, Muying Luo |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | A Novel Recurrent Encoder-Decoder Structure for Large-Scale Multi-View Stereo Reconstruction From an Open Aerial DatasetabstractA great deal of research has demonstrated recently that multi-view stereo (MVS) matching can be solved with deep learning methods. However, these efforts were focused on close-range objects and only a very few of the deep learning-based methods were specifically designed for large-scale 3D urban reconstruction due to the lack of multi-view aerial image benchmarks. In this paper, we present a synthetic aerial dataset, called the WHU dataset, we created for MVS tasks, which, to our knowledge, is the first large-scale multi-view aerial dataset. It was generated from a highly accurate 3D digital surface model produced from thousands of real aerial images with precise camera parameters. We also introduce in this paper a novel network, called RED-Net, for wide-range depth inference, which we developed from a recurrent encoder-decoder structure to regularize cost maps across depths and a 2D fully convolutional network as framework. RED-Net's low memory requirements and high performance make it suitable for large-scale and highly accurate 3D Earth surface reconstruction. Our experiments confirmed that not only did our method exceed the current state-of-the-art MVS methods by more than 50% mean absolute error (MAE) with less memory and computational cost, but its efficiency as well. It outperformed one of the best commercial software programs based on conventional methods, improving their efficiency 16 times over. Moreover, we proved that our RED-Net model pre-trained on the synthetic WHU dataset can be efficiently transferred to very different multi-view aerial image datasets without any fine-tuning. Dataset and code are available at http://gpcv.whu.edu.cn/data. Shunping Ji |
CVPR | 2 |
| 2020 | Toward Automatic Building Footprint Delineation From Aerial Images Using CNN and RegularizationabstractThis study proposes an automatic building footprint extraction framework that consists of a convolutional neural network (CNN)-based segmentation and an empirical polygon regularization that transforms segmentation maps into structured individual building polygons. The framework attempts to replace part of the manual delineation of building footprints that are involved in surveying and mapping field with algorithms. First, we develop a scale robust fully convolutional network (FCN) by introducing multiple scale aggregation of feature pyramids from convolutional layers. Two postprocessing strategies are introduced to refine the segmentation maps from the FCN. The refined segmentation maps are vectorized and polygonized. Then, we propose a polygon regularization algorithm consisting of a coarse and fine adjustment, to translate the initial polygons into structured footprints. Experiments on a large open building data set including 181 000 buildings showed that our algorithm reached a high automation level where at least 50% of individual buildings in the test area could be delineated to replace manual work. Experiments on different data sets demonstrated that our FCN-based segmentation method outperformed several most recent segmentation methods, and our polygon regularization algorithm is robust in challenging situations with different building styles, image resolutions, and even low-quality segmentation. Shiqing Wei, Shunping Ji |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Simultaneous Road Surface and Centerline Extraction From Large-Scale Remote Sensing Images Using CNN-Based Segmentation and TracingabstractAccurate and up-to-date road maps are of great importance in a wide range of applications. Unfortunately, automatic road extraction from high-resolution remote sensing images remains challenging due to the occlusion of trees and buildings, discriminability of roads, and complex backgrounds. To address these problems, especially road connectivity and completeness, in this article, we introduce a novel deep learning-based multistage framework to accurately extract the road surface and road centerline simultaneously. Our framework consists of three steps: boosting segmentation, multiple starting points tracing, and fusion. The initial road surface segmentation is achieved with a fully convolutional network (FCN), after which another lighter FCN is applied several times to boost the accuracy and connectivity of the initial segmentation. In the multiple starting points tracing step, the starting points are automatically generated by extracting the road intersections of the segmentation results, which then are utilized to track consecutive and complete road networks through an iterative search strategy embedded in a convolutional neural network (CNN). The fusion step aggregates the semantic and topological information of road networks by combining the segmentation and tracing results to produce the final and refined road segmentation and centerline maps. We evaluated our method utilizing three data sets covering various road situations in more than 40 cities around the world. The results demonstrate the superior performance of our proposed framework. Specifically, our method's performance exceeded the other methods by 7% and 40% for the connectivity indicator for road surface segmentation and for the completeness indicator for centerline extraction, respectively. Shunping Ji |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Road Network Extraction from Satellite Images Using CNN Based Segmentation and TracingabstractDrawing and updating road networks are both time-consuming and labor-intensive. Deep learning technology and high-resolution remote sensing images have provided opportunities for automatic road extraction. However, recent convolutional neural network (CNN) based segmentation methods have shown serious problems on connectivity; road tracing methods with single starting point perform well in connectivity but often result in part areas unreached. We propose a multiple starting points tracer which benefits from both segmentation and tracing methods. We compare our approach with most recent tracing methods on satellite images of global cities and find that our method achieves 8% improvement on IoU. Shunping Ji |
IGARSS | 3 |
| 2019 | Fully Convolutional Networks for Multisource Building Extraction From an Open Aerial and Satellite Imagery Data SetabstractThe application of the convolutional neural network has shown to greatly improve the accuracy of building extraction from remote sensing imagery. In this paper, we created and made open a high-quality multisource data set for building detection, evaluated the accuracy obtained in most recent studies on the data set, demonstrated the use of our data set, and proposed a Siamese fully convolutional network model that obtained better segmentation accuracy. The building data set that we created contains not only aerial images but also satellite images covering 1000 km2with both raster labels and vector maps. The accuracy of applying the same methodology to our aerial data set outperformed several other open building data sets. On the aerial data set, we gave a thorough evaluation and comparison of most recent deep learning-based methods, and proposed a Siamese U-Net with shared weights in two branches, and original images and their down-sampled counterparts as inputs, which significantly improves the segmentation accuracy, especially for large buildings. For multisource building extraction, the generalization ability is further evaluated and extended by applying a radiometric augmentation strategy to transfer pretrained models on the aerial data set to the satellite data set. The designed experiments indicate our data set is accurate and can serve multiple purposes including building instance segmentation and change detection; our result shows the Siamese U-Net outperforms current building extraction methods and could provide valuable reference. Shunping Ji, Shiqing Wei |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | Fusion of a panoramic camera and 2D laser scanner data for constrained bundle adjustment in GPS-denied environments
Shunping Ji, Xiaowei Shao, Peng Yang 0005, Zhongchao Shi, Ryosuke Shibasaki |
Image Vis. Comput. | 2 |