EDBT 2026 Demo / reviewers in the wild / expert
Shiqing Wei
dblp:232/4600
· DBLP profile ↗
11ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0001-9130-4584ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A One-Shot Pine Tree Disease Segmentation Model Integrating Interclass Relations and Prior Contour AwarenessabstractDespite the proven effectiveness of deep learning technology in pine tree disease segmentation, acquiring a large volume of labeled data remains challenging and inefficient. Few-shot segmentation (FSS) uses a small amount of labeled data to guide the segmentation of unknown categories, further evolving into one-shot segmentation (OSS), which utilizes a single labeled sample to perform segmentation under conditions of extreme data scarcity. However, these methods are mostly applicable to natural images with clear boundaries and have not yet been applied to segmenting pine tree disease in autonomous aerial vehicle (AAV) remote sensing images. For this reason, we have designed the OSS model C2Net for the first time, which includes two main modules: 1) a prior contour awareness module (PCAM) that first generates a query image prior mask with contour response and then uses an iterative feature refinement unit (FRU) to refine features and accurately delineate the segmentation boundaries of pine tree disease and 2) an interclass relationship module (ICRM), which studies the vegetation index features of the support and query images, constructing importance weights that reflect the differences between categories, solving the visual similarity issue. Our experiments on field-collected and publicly available datasets demonstrate that C2Net excels in challenging OSS tasks, showing its ability to generalize across different sensor domains and various disease categories. Especially, on the field acquisition dataset, using just a single labeled pine tree disease image achieves an intersection over union (IoU) of 55.24% and an$F_{1}$of 71.24%. Hui Sheng, Shiqing Wei, Ke Hou, Mingming Xu 0001, Shanwei Liu, Cunhui Zhang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | SAM-RSIS: Progressively Adapting SAM With Box Prompting to Remote Sensing Image Instance SegmentationabstractThe recent segment anything model (SAM) trained on massive close-range images has demonstrated impressive performance on general segmentation or specific segmentation tasks with manual prompts. However, the significant domain shift problem between remote sensing and close-range images should be tackled before introducing the pretrained SAM to remote sensing instance segmentation (RSIS). To address this and unlock the potential of SAM in RSIS, this article proposes a novel framework called SAM for remote sensing instance segmentation (SAM-RSIS), which overcomes the problems in a few recent works that only adapt a part of SAM to remote sensing. SAM-RSIS fine-tunes the vision transformer (ViT) backbone and mask decoder of SAM progressively on remote sensing data and uses automatic box prompting to eliminate the need for manual prompting. SAM-RSIS consists of an object detection stage and a mask generation stage. In object detection, we introduce an adapter to adapt knowledge embedded in the pretrained ViT backbone to remote sensing images and then build an object detector. In mask generation, using the detected bounding boxes as prompts, along with two learnable mask output tokens, and the two-layer high-resolution features from the adapter, we fine-tune the mask decoder of SAM to produce high-quality masks. Experimental results on the WHU, WHU-Mix, and NWPU datasets for binary and multiclass RSIS demonstrate the effectiveness and robustness of the proposed method, surpassing various derivative methods of SAM and achieving performance comparable to and even better than the specific state-of-the-art instance segmentation methods. Muying Luo, Tao Zhang 0042, Shiqing Wei, Shunping Ji |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | 3-D Building Instance Extraction From High-Resolution Remote Sensing Images and DSM With an End-to-End Deep Neural NetworkabstractThree-dimensional (3D) building models play a vital role in numerous applications including urban planning and smart cities. Recent 3D building modeling methods either rely heavily on available manaually-collected footprint reference or hardly reach real automation on par with manual editing. To approach the automated extraction of instance-level 3D buildings at Level of Detail (LoD) 1, we introduce an innovative end-to-end 3D building instance segmentation model. This model predicts accurate contours and heights of individual buildings simultaneously using ortho-rectified high-resolution remote sensing images and Digital Surface Models (DSMs), getting rid of additional reference data and impirical parameter settings. Firstly, we propose an Anchor-Free Multi-head building extraction network (AFM) tailored for extracting 2D building contours. AFM incorporates a full-resolution, long-range correlation boosted global mask prediction branch along with anchor-free bounding box generation, as well as a newly developed online hard sample mining (OHSM) training procedure based on uncertainty analysis to emphasize error-prone positions in locating building contours. Subsequently, we incorporate a height prediction component to AFM in order to derive accurate building height information, thus creating the comprehensive 3D building extraction model referred to as AFM-3D. The two-stage AFM-3D operates by initially predicting 3D cube proposals, followed by generating refined 3D prismatic models (LoD1 models) for each proposal. Thorough experimentation across different datasets demonstrates the superior performance of AFM and AFM-3D. A significant enhancement of 6.4% quality score is observed on the urban 3D dataset in comparison to recent methods. In addition to the proposed novel methodology, we compare anchor-based and anchor-free bounding box generation mechanisms for remote sensing data, explore pixel-based and contour-based segmentation strategies, evaluate learning-based and empirical height estimation methods, and discuss the indispensability of DSM data in 3D building instance extraction. These analyses yield valuable insights that contribute to the progression of 3D building extraction research. Dawen Yu, Shunping Ji, Shiqing Wei, Kourosh Khoshelham |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | P2PFormer: A Primitive-to-Polygon Method for Regular Building Contour Extraction From Remote Sensing ImagesabstractExtracting building contours from remote sensing imagery is a significant challenge due to buildings’ complex and diverse shapes, occlusions, and noise. Existing methods often struggle with irregular contours, rounded corners, and redundancy points, necessitating extensive postprocessing to produce regular polygonal building contours. To address these challenges, we introduce a novel, streamlined pipeline that generates regular building contours without postprocessing. Our approach begins with the segmentation of generic geometric primitives (which can include vertices, lines, and corners), followed by the prediction of their sequence. This allows for the direct construction of regular building contours by sequentially connecting the segmented primitives. Building on this pipeline, we developed primitive-to-polygon using transformer (P2PFormer), which uses a transformer-based architecture to segment geometric primitives and predict their order. To enhance the segmentation of primitives, we introduce a unique representation called group queries. This representation comprises a set of queries and a singular query position, which improve the focus on multiple midpoints of primitives and their efficient linkage. Furthermore, we propose an innovative implicit update strategy for the query position embedding aimed at sharpening the focus of queries on the correct positions and, consequently, enhancing the quality of primitive segmentation. Our experiments demonstrate that P2PFormer achieves new state-of-the-art (SOTA) performance on the WHU, CrowdAI, and WHU-Mix datasets, surpassing the previous SOTA PolyWorld by a margin of 2.7 AP and 6.5 AP75 on the largest CrowdAI dataset. We intend to make the code and trained weights publicly available to promote their use and facilitate further research. Tao Zhang 0042, Shiqing Wei, Yikang Zhou, Muying Luo, Wenling Yu, Shunping Ji |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Mutual Information Based Fusion Model (MIBFM): Mild Depression Recognition Using EEG and Pupil Area SignalsabstractThe detection of mild depression is conducive to the early intervention and treatment of depression. This study explored the fusion of electroencephalography (EEG) and pupil area signals to build an effective and convenient mild depression recognition model. We proposed Mutual Information Based Fusion Model (MIBFM), which innovatively used pupil area signals to select EEG electrodes based on mutual information. Then we extracted features from EEG and pupil area signals in different bands, and fused bimodal features using the denoising autoencoder. Experimental results showed that MIBFM could obtain the highest accuracy of 87.03%. And MIBFM exhibited better performance than other existing methods. Our findings validate the effectiveness of the use of pupil area as signals, which makes eye movement signals can be easily obtained using high resolution camera, and the EEG electrode selection scheme based on mutual information is also proved to be an applicable solution for data dimension reduction and multimodal complementary information screening. This study casts a new light for mild depression recognition using multimodal data of EEG and pupil area signals, and provides a theoretical basis for the development of portable and universal application systems. Jing Zhu 0003, Changlin Yang, Xiannian Xie, Shiqing Wei, Xiaowei Li 0005, Bin Hu 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2022 | Hybrid fusion model based on DBN and secondary classifier: Multimodal mild depression recognition using EEG and eye movementabstractIn recent years, depression recognition using physiological signals has achieved certain progress, but mild depression recognition is still in its infancy. Early detection can prevent the development of depression, and combining multiple modalities for analysis has been proved effective in the domain of mental disorders detection. Electroencephalogram (EEG) and eye movements (EM) are widely used to identify depression. However, the problem of using physiological signals to detect mental illness is that the generalization ability of the model is not strong, which is caused by individual differences. In view of the above problems, this paper proposes a hybrid fusion model based on deep belief network (DBN) and secondary classifier, called HFMBDSC, which first uses unsupervised DBN to fuse EEG features (linear, nonlinear features and network features) at the feature level. DBN transforms EEG features into another form to mitigate the effects of individual differences in EEG, and obtains DBN features that can more comprehensively represent EEG information. In the next step of decision level fusion, DBN features and EM features jointly make the final decision using the three classifiers that perform best when using single modality. The results show that DBN can improve the classification accuracy and effectively reduce the influence of individual differences in EEG features through visual analysis of feature space. The classification performance of HFMBDSC is significantly improved compared to the traditional single modality results. The highest accuracy rate of 89.54% is obtained under 10-fold cross-validation. These results suggest that mild depression recognition based on HFMBDSC is promising. Jing Zhu 0003, Xiannian Xie, Changlin Yang, Shiqing Wei, Xiaowei Li 0005, Bin Hu 0001 |
BIBM | 4 |
| 2022 | E2EC: An End-to-End Contour-based Method for High-Quality High-Speed Instance SegmentationabstractContour-based instance segmentation methods have developed rapidly recently but feature rough and hand-crafted front-end contour initialization, which restricts the model performance, and an empirical and fixed backend predicted-label vertex pairing, which contributes to the learning difficulty. In this paper, we introduce a novel contour-based method, named E2EC, for high-quality instance segmentation. Firstly, E2EC applies a novel learnable contour initialization architecture instead of hand-crafted contour initialization. This consists of a contour initialization module for constructing more explicit learning goals and a global contour deformation module for taking advantage of all of the vertices' features better. Secondly, we propose a novel label sampling scheme, named multi-direction alignment, to reduce the learning difficulty. Thirdly, to improve the quality of the boundary details, we dynamically match the most appropriate predicted-ground truth vertex pairs and propose the corresponding loss function named dynamic matching loss. The experiments showed that E2EC can achieve a state-of-the-art performance on the KITTI INStance (KINS) dataset, the Semantic Boundaries Dataset (SBD), the Cityscapes and the COCO dataset. E2EC is also efficient for use in real-time applications, with an inference speed of 36 fps for$512\times 512$images on an NVIDIA A6000 GPU. Code will be released at https://github.com/zhang-tao-whu/e2ec. Tao Zhang 0042, Shiqing Wei, Shunping Ji |
CVPR | 2 |
| 2022 | Graph Convolutional Networks for the Automated Production of Building Vector Maps From Aerial ImagesabstractBuilding footprint delineation from remote sensing imagery is a basic task in surveying and mapping and geographic information system (GIS), which benefits many engineering applications but requires an enormous amount of human delineation. In this study, we aim to find a way to replace human delineation with automated algorithms. For this purpose, we designed a novel pipeline for the automated production of building vector maps from aerial images, which consists of a bounding-box generation module, a graph convolutional network (GCN)-based polygon prediction module, and an empirical polygon regularization module. First, we introduce the bounding box generation module based on region-based object detection, which along with our overlap cropping strategy is used to generate a bounding box for each building instance. The generated bounding boxes have a horizontal rectangle version and a rotated version, which are used to initialize the next stage of our method. Second, we propose a GCN-based method, which is the core of this study, to conduct the initial accurate building polygon prediction by integrating multiresolution optimization and multilevel loss constraints. Our proposed method, to the authors’ best knowledge, is the first of its kind that introduces GCN to building extraction. The final step is to apply a regularization algorithm to translate the predicted polygons into structured and highly accurate building footprints. We validated the proposed method on two large aerial building data sets, WHU data set and Inria data set, where it was shown to significantly outperform other state-of-the-art methods more than 10%. Specifically, with the ground-truth rotated bounding boxes of buildings, our method is able to automatically delineate 91% of the buildings in the WHU data set and with the predicted rotated bounding the percent reached human-level delineation. Shiqing Wei, Shunping Ji |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | A Concentric Loop Convolutional Neural Network for Manual Delineation-Level Building Boundary Segmentation From Remote-Sensing ImagesabstractTo date, accurate building footprint delineation in the surveying, mapping, and geographic information system (GIS) communities has been dependent on human labor. In this article, to address this issue, we propose a concentric loop convolutional neural network (CLP-CNN) method for the automatic segmentation of building boundaries from remote-sensing images. The proposed method consists of three components: 1) a boundary detector to extract coarse polygonal boundaries of individual regions of interest; 2) a concentric loop-shaped convolutional network with bidirectional pairing loss to fine-tune the vertices of the polygons; and 3) a refinement block, which removes redundant vertices and regularizes the boundaries to polygons at the manual delineation level. We also demonstrate that the proposed CLP-CNN method is applicable to other generic objects in natural images. Experiments on two building datasets confirmed that more than 77%/67% of the building polygons predicted by the proposed method are on par with the manual delineation level, representing a significant saving in the labor cost of manual annotation. In generic object boundary delineation tests performed on the Semantic Boundaries Dataset (SBD), the proposed method outperformed the most recent state-of-the-art methods by at least 3.1% in average precision (AP). Furthermore, compared with other vertex matching methods, the learning process of the proposed method converges faster. The source code will be available athttp://gpcv.whu.edu.cn/data. Shiqing Wei, Tao Zhang 0042, Shunping Ji |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Toward Automatic Building Footprint Delineation From Aerial Images Using CNN and RegularizationabstractThis study proposes an automatic building footprint extraction framework that consists of a convolutional neural network (CNN)-based segmentation and an empirical polygon regularization that transforms segmentation maps into structured individual building polygons. The framework attempts to replace part of the manual delineation of building footprints that are involved in surveying and mapping field with algorithms. First, we develop a scale robust fully convolutional network (FCN) by introducing multiple scale aggregation of feature pyramids from convolutional layers. Two postprocessing strategies are introduced to refine the segmentation maps from the FCN. The refined segmentation maps are vectorized and polygonized. Then, we propose a polygon regularization algorithm consisting of a coarse and fine adjustment, to translate the initial polygons into structured footprints. Experiments on a large open building data set including 181 000 buildings showed that our algorithm reached a high automation level where at least 50% of individual buildings in the test area could be delineated to replace manual work. Experiments on different data sets demonstrated that our FCN-based segmentation method outperformed several most recent segmentation methods, and our polygon regularization algorithm is robust in challenging situations with different building styles, image resolutions, and even low-quality segmentation. Shiqing Wei, Shunping Ji |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Fully Convolutional Networks for Multisource Building Extraction From an Open Aerial and Satellite Imagery Data SetabstractThe application of the convolutional neural network has shown to greatly improve the accuracy of building extraction from remote sensing imagery. In this paper, we created and made open a high-quality multisource data set for building detection, evaluated the accuracy obtained in most recent studies on the data set, demonstrated the use of our data set, and proposed a Siamese fully convolutional network model that obtained better segmentation accuracy. The building data set that we created contains not only aerial images but also satellite images covering 1000 km2with both raster labels and vector maps. The accuracy of applying the same methodology to our aerial data set outperformed several other open building data sets. On the aerial data set, we gave a thorough evaluation and comparison of most recent deep learning-based methods, and proposed a Siamese U-Net with shared weights in two branches, and original images and their down-sampled counterparts as inputs, which significantly improves the segmentation accuracy, especially for large buildings. For multisource building extraction, the generalization ability is further evaluated and extended by applying a radiometric augmentation strategy to transfer pretrained models on the aerial data set to the satellite data set. The designed experiments indicate our data set is accurate and can serve multiple purposes including building instance segmentation and change detection; our result shows the Siamese U-Net outperforms current building extraction methods and could provide valuable reference. Shunping Ji, Shiqing Wei |
IEEE Trans. Geosci. Remote. Sens. | 2 |