VLDB 2026 Research / reviewers in the wild / expert
Mi Zhang 0004
dblp:84/2519-4
· DBLP profile ↗
7ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0003-4949-979XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DiffVector: Boosting Diffusion Framework for Building Vector Extraction From Remote Sensing ImagesabstractBuilding vector maps play an essential role in many remote sensing (RS) applications, thereby boosting the deep learning (DL)-based automatic building vector extraction methods. These approaches have achieved pleasant overall accuracy, but their predict-style framework struggles with perceiving subtle details within a tiny area, such as corners and adjacent walls. In this study, we introduce a denoising diffusion framework called DiffVector to generate representations for direct building vector extraction from the RS images. First, we develop a hierarchical diffusion transformer (HiDiT) to conditionally generate robust representations for detecting nodes and extracting corresponding features. The conditions of HiDiT are multilevel boundary attentive maps encoded from input RS images through a topology-concentrated Swin Transformer (TCSwin). Subsequently, an edge-biased graph diffusion transformer (EGDiT) takes extracted node features as conditions to produce new visual descriptors for the adjacency matrix prediction. In EGDiT, we replace the standard self-attention (SA) operation with an edge-biased attention (EBA) to inject edge information for training stabilization. Furthermore, given typical challenges of training difficulty and weak perceptive ability in convectional diffusion paradigms, we conduct an isomorphic training strategy (ITS), ensuring that the training procedures of both HiDiT and EGDiT precisely mirror the inference phase. Quantitative and qualitative experiments have evidently demonstrated that DiffVector can achieve competitive performance compared with existing modern approaches, especially in the metrics assessing topology quality. Bingnan Yang, Mi Zhang 0004, Yuanxin Zhao, Xiangyun Hu, Jianya Gong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | SegAssess: Panoramic Quality Mapping for Robust and Transferable Unsupervised Segmentation AssessmentabstractHigh-quality image segmentation is fundamental to pixel-level geospatial analysis in remote sensing, necessitating robust segmentation quality assessment (SQA), particularly in unsupervised settings lacking ground truth. Although recent deep learning (DL) based unsupervised SQA methods show potential, they often suffer from coarse evaluation granularity, incomplete assessments, and poor transferability. To overcome these limitations, this paper introduces Panoramic Quality Mapping (PQM) as a new paradigm for comprehensive, pixel-wise SQA, and presents SegAssess, a novel deep learning framework realizing this approach. SegAssess distinctively formulates SQA as a fine-grained, four-class panoramic segmentation task, classifying pixels within a segmentation mask under evaluation into true positive (TP), false positive (FP), true negative (TN), and false negative (FN) categories, thereby generating a complete quality map. Leveraging an enhanced Segment Anything Model (SAM) architecture, SegAssess uniquely employs the input mask as a prompt for effective feature integration via cross-attention. Key innovations include an Edge Guided Compaction (EGC) branch with an Aggregated Semantic Filter (ASF) module to refine predictions near challenging object edges, and an Augmented Mixup Sampling (AMS) training strategy integrating multi-source masks to significantly boost cross-domain robustness and zero-shot transferability. Comprehensive experiments demonstrate that SegAssess achieves state-of-the-art (SOTA) performance and exhibits remarkable zero-shot transferability to unseen masks. The code is available at https://github.com/Yangbn97/SegAssess. Bingnan Yang, Mi Zhang 0004, Yuanxin Zhao, Xiangyun Hu, Jianya Gong |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Faster Interactive Segmentation of Identical-Class Objects With One Mask in High-Resolution Remotely Sensed ImageryabstractInteractive segmentation (IS) using minimal prompts like points and bounding boxes facilitates rapid image annotation, which is crucial for enhancing data-driven deep learning methods. Traditional IS methods, however, process only one target per interaction, leading to inefficiency when annotating multiple identical-class objects in remote sensing imagery (RSI). To address this issue, we present a new task—identical-class object detection (ICOD) for rapid IS in RSI. This task aims to only identify and detect all identical-class targets within an image, guided by a specific category target in the image with its mask. For this task, we propose an ICOD network (ICODet) with a two-stage object detection framework, which consists of a backbone, feature similarity analysis module (S3QFM), and an identical-class object detector. In particular, the S3QFM analyzes feature similarities from images and support objects at both feature-space and semantic levels, generating similarity maps. These maps are processed by a region proposal network (RPN) to extract target-level features, which are then refined through a simple feature comparison module and classified to precisely identify identical-class targets. To evaluate the effectiveness of this method, we construct two datasets for the ICOD task: one containing a diverse set of buildings and another containing multicategory RSI objects. Experimental results show that our method outperforms the compared methods on both datasets. This research introduces a new method for rapid IS of RSI and advances the development of fast interaction modes, offering significant practical value for data production and fundamental applications in the remote sensing community. Jiabo Xu, Xiangyun Hu, Bingnan Yang, Mi Zhang 0004 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Prototypical Unknown-Aware Multiview Consistency Learning for Open-Set Cross-Domain Remote Sensing Image ClassificationabstractDeveloping a cross-domain classification model for remote sensing images has drawn significant attention in the literature. By leveraging the open-set unsupervised domain adaptation (UDA) technique, the generalization performance of deep learning models has been improved with the capability to recognize unknown categories. However, it remains challenging to explore distribution patterns in the target domain using uncertain category-wise supervision from unlabeled datasets while reducing negative transfer caused by unknown samples. To develop a robust open-set UDA framework, this article presents prototypical unknown-aware multiview consistency learning (PUMCL) designed for remote sensing scene classification across heterogeneous domains. Specifically, it employs a consistency learning scheme with multiview and multilevel perturbations to improve feature learning from unlabeled target samples. An entropy separation strategy is utilized to facilitate open-set detection and recognition during adaptation, enabling unknown-aware feature alignment. Furthermore, the introduction of prototypical constraints optimizes pseudo-label generation through online denoising and promotes a compact category-wise feature subspace for improved class separation across domains. Experiments conducted on six cross-domain scenarios using AID, NWPU, and UCMD datasets demonstrate the method’s superior performance compared to nine state-of-the-art approaches, achieving a gain of 4.5% to 21.2% in mIoU. More importantly, it shows promising class separability with clear boundaries between different classes and compact clustering of unknown samples in the feature space. The source code will be available athttps://github.com/zxk688. Wanjing Wu, Mi Zhang 0004, Weikang Yu, Pedram Ghamisi |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | TopDiG: Class-agnostic Topological Directional Graph Extraction from Remote Sensing ImagesabstractRapid development in automatic vector extraction from remote sensing images has been witnessed in recent years. However, the vast majority of existing works concentrate on a specific target, fragile to category variety, and hardly achieve stable performance crossing different categories. In this work, we propose an innovative class-agnostic model, namely TopDiG, to directly extract topological directional graphs from remote sensing images and solve these issues. Firstly, TopDiG employs a topology-concentrated node detector (TCND) to detect nodes and obtain compact perception of topological components. Secondly, we propose a dynamic graph supervision (DGS) strategy to dynamically generate adjacency graph labels from unordered nodes. Finally, the directional graph (DiG) generator module is designed to construct topological directional graphs from predicted nodes. Experiments on the Inria, CrowdAI, GID, GF2 and Massachusetts datasets empirically demonstrate that TopDiG is class-agnostic and achieves competitive performance on all datasets. Bingnan Yang, Mi Zhang 0004, Xiangyun Hu |
CVPR | 2 |
| 2023 | Luojia-AI: A Full-Stack Cloud Computing Infrastructure for Remote Sensing Intellignet InterpretationabstractThe rapid processing, analysis, and mining of remote sensing big data using intelligent interpretation technology on remote sensing cloud computing platforms (RS-CCPs) have emerged as a new trend. However, existing RS-CCPs primarily focus on optimizing data storage and intelligent computing for common visual representation, overlooking key characteristics of remote sensing data such as large image size, large-scale change, multiple data channels, and geographic knowledge embedding. This oversight hinders computational efficiency and accuracy in remote sensing image interpretation. To address this, we have developed the LuoJia-AI platform, comprising the LuoJiaSET standard large-scale sample database and the dedicated deep learning framework, LuoJiaNET. This platform achieves state-of-the-art performance on five crucial remote sensing interpretation tasks: scene classification, object detection, land-use classification, change detection, and multi-view 3D reconstruction. LuoJia-AI bridges the gap between the sample database and the deep learning framework, exhibiting significant potential for high-precision remote sensing mapping applications. Mi Zhang 0004, Jianya Gong, Xiangyun Hu, Liangcun Jiang, Jiansi Yang |
IGARSS | 2 |
| 2015 | Line-based Multi-Label Energy Optimization for fisheye image rectification and calibrationabstractFisheye image rectification and estimation of intrinsic parameters for real scenes have been addressed in the literature by using line information on the distorted images. In this paper, we propose an easily implemented fisheye image rectification algorithm with line constrains in the undistorted perspective image plane. A novel Multi-Label Energy Optimization (MLEO) method is adopted to merge short circular arcs sharing the same or the approximately same circular parameters and select long circular arcs for camera rectification. Further we propose an efficient method to estimate intrinsic parameters of the fisheye camera by automatically selecting three properly arranged long circular arcs from previously obtained circular arcs in the calibration procedure. Experimental results on a number of real images and simulated data show that the proposed method can achieve good results and outperforms the existing approaches and the commercial software in most cases. Mi Zhang 0004, Jian Yao 0002, Menghan Xia, Kai Li 0015 |
CVPR | 1 |