EDBT 2026 Demo / reviewers in the wild / expert
Ming Zhao 0009
dblp:39/6844-9
· DBLP profile ↗
11ranked-venue papers
9as first author
7since 2021 · last 2026
0000-0002-1310-6766ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Collaborative feature alignment with global-local fusion for fine-grained sketch-based image retrieval
Ming Zhao 0009, Lixiang Ma |
Pattern Recognit. Lett. | 2 |
| 2025 | An Arbitrary-Scale Super-Resolution Network for Multi-Contrast MRI With Permuted Cross-AttentionabstractIn magnetic resonance imaging (MRI), low-resolution (LR) images often hamper clinical diagnosis and research due to constraints in imaging conditions and technology limitations. Recent studies in super-resolution (SR) reconstruction of multi-contrast MRI have shown promise by leveraging the complementary information from different MRI contrasts. However, existing multi-contrast MRI SR techniques face several challenges: 1) a lack of pre-alignment precision can result in distorted reconstructions; 2) prevailing transformer network structures, with their smaller windows (e.g., 8 × 8), struggle to effectively capture long-range dependencies and lack the ability to interact between different windows; and 3) current methods are limited to fixed integer scaling (e.g., 2 ×, 3 ×, 4 ×), which limits flexibility and increases complexity in training and storage. To address these challenges, we propose a novel arbitrary-scale SR network for multi-contrast MRI. Specifically, our approach compensates for spatial misalignment between modalities through deformable registration module and employs permuted cross-attention transformer in MR images. In addition, we introduce a ref-scale ensemble implicit attention module that better integrates high-frequency information from reference images and enables arbitrary-scale upsampling. Extensive experiments on two publicly available MRI datasets validate the superiority of our method in multi-contrast MRI SR, demonstrating its significant potential in clinical applications. Ming Zhao 0009, Jia Fang, Boyang Chen 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | A Semi-Supervised Image Registration Framework Based on Multimodal Cross-AttentionabstractRegistration of multimodal image pairs is a fundamental task in many remote sensing applications. In order to achieve accurate and low-cost remote sensing image registration, we propose a semi-supervised image registration framework based on multimodal cross-attention, which consists of the encoder for feature extraction, multimodal cross-attention module, and detection/descriptor decoders. We adopt positional encoding for feature maps to enhance the features with spatial contexts, especially for remote sensing images with large geometrical deformations. In order to learn common features that independent of modalities between multimodal images, we proposed multimodal cross-attention module to extract cross modal features, which helps the detectors to extract more reliable matching keypoints. The network is trained in a semi-supervised manner, which requires only a small dataset of incompletely labeled images. In order to learn reliable keypoints from image pairs with inconsistent intensity and geomitrical deformations, we randomly establish different geometrical mappings for the multimodal image pairs during training, and then enrich the keypoint labels by continuously adding reliable keypoints extracted by the detection decoder in each epoch. Experimental results show that the proposed method achieves more comprehensive and accurate registration than the state-of-the-art methods for multimodal remote sensing images. Ming Zhao 0009 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | Cross-Modal Prealigned Method With Global and Local Information for Remote Sensing Image and Text RetrievalabstractIn recent years, remote sensing cross-modal text-image retrieval (RSCTIR) has attracted considerable attention owing to its convenience and information mining capabilities. However, two significant challenges persist: effectively integrating global and local information during feature extraction due to substantial variations in remote sensing imagery, and the failure of existing methods to adequately consider feature prealignment before modal fusion, resulting in complex modal interactions that adversely impact retrieval accuracy and efficiency. To address these challenges, we propose a cross-modal prealigned method with global and local information (CMPAGL) for remote sensing imagery. Specifically, we design a global-Swin (Gswin) Transformer block, which introduces a global information window on top of the local window attention mechanism, synergistically combining local window self-attention and global-local window cross-attention to effectively capture multiscale features of remote sensing images. In addition, our approach incorporates a prealignment mechanism to mitigate the training difficulty of modal fusion, thereby enhancing retrieval accuracy. Moreover, we propose a similarity matrix reweighting (SMR) reranking algorithm to deeply exploit information from the similarity matrix during the retrieval process. This algorithm combines forward and backward ranking, extreme difference ratio, and other factors to reweight the similarity matrix, thereby further enhancing retrieval accuracy. Finally, we optimize the triplet loss function by introducing an intraclass distance term for matched image-text pairs, not only focusing on the relative distance between matched and unmatched pairs but also minimizing the distance within matched pairs. Experiments on four public remote sensing text-image datasets, including RSICD, RSITMD, UCM-Captions, and Sydney-Captions, demonstrate the effectiveness of our proposed method, achieving improvements over state-of-the-art methods, such as a 2.28% increase in mean Recall (mR) on the RSITMD dataset and a significant 4.65% improvement in R@1. The code is available athttps://github.com/ZbaoSun/CMPAGL. Zengbao Sun, Ming Zhao 0009, Gaorui Liu, André Kaup |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | A Self-Adaptive Object Detection Network for Aerial Images Based on Feature EnhancementabstractObject detection in aerial images has received increasing attention for its widely applications. However, it is still a challenging task when dealing with difficult cases, such as the large variations of object sizes, the complex backgrounds in the large-scale views, as well as small targets packed in dense. In this regard, we propose a self-adaptive object detection network for aerial images based on feature enhancement, including region crop module with soft-attention (RCP), feature enhancement module (FET), self adaptive feature extraction module (SAE). Firstly, RCP module with soft-attention is explored to roughly crop the dense subregion and sparse subregion into patches accoridng to the variance of feature maps for the following feature extraction and detection. Secondly, FET module is proposed to acquire more semantic details by feature enhancement, which makes up the information loss during downsampling, especially for small objects in dense regions. Finally, SAE module is explored to effectively identify multiple target regions and single target regions in the dense patches. The similarity of adjacent dense areas in dense patches is calcluated, and the search range is gradually narrowed by continuously merging the areas with the largest similarity. The teacher-network and student-network are used to extract features for multiple target regions and single target regions, respectively. The proposed design improves the accuracy of real-time detection under the drone’s perspective. A large number of experiments and comprehensive evaluations on the VisDrone2019-DET dataset have shown the effectiveness and adaptability of the proposed method. Our source codes have been available at https://github.com/zhaokai152. Ming Zhao 0009 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2023 | Multitask Learning for SAR Ship Detection With Gaussian-Mask Joint SegmentationabstractDetecting ships from synthetic aperture radar (SAR) images is inherently subject to the limitations of SAR’s imaging mechanism. SAR object detection technology has rapidly advanced in recent years due to deep learning based techniques for detecting objects from optical images. SAR ship detection still faces some challenges due to the strong speckle noise, complex surroundings, and variety of scales. This paper proposes a multitask learning framework for object detection (MLDet) detect ships in SAR images. The proposed end-to-end framework consists of object detection task, speckle supression task and target segmentation task. Firstly, an angle classification loss with aspect ratio weighting is explored during object detection to improve the accuracy by making the detector sensitive to the periodicity of angular and the aspect ratio of objects. Secondly, the speckle supression task employs a dual-feature fusion attention mechanism to suppress noisy background information and fuse shallow features and denoising features, which helps MLDet be more robust to speckle noise. Thirdly, the target segmentation task with rotated Gaussian-mask is explored to further help the detection network to extract the regions of intersect from the cluttered background, as well as improving the detection efficiency through pixel-by-pixel prediction. The rotated Gaussian-mask for ship modeling ensures that the center of a ship has the highest probablities to be labeled as an object, and the probablities of the remaining regions are gradually reduced under a Gaussian distribution. Assisted by these two subtasks, the shallow level features are robust to speckle noise and reliably support deep level feature learning. In addition, the weighted rotated boxes fusion (WRBF) strategy is adopted to combine the predictions of multi-direction anchors for rotated objects, and eliminate the anchors beyond the boundary as well as the anchors with high overlap rates but low scores. A large number of experiments and comprehensive evaluations on SAR ship detection datasets SSDD+ and HRSID have shown the effectiveness and superiority of the proposed method. The code is available from https://github.com/zx152/MLDet. Ming Zhao 0009, André Kaup |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Orientation-Aware Feature Fusion Network for Ship Detection in SAR ImagesabstractRecently, deep learning methods have been successfully applied to the ship detection in synthetic aperture radar (SAR) images. It is still a great challenge to detect SAR ships, due to the extremely poor image quality and complex background. To solve the problems, a novel method named orientation-aware feature fusion network (OFF-Net) for ship detection in SAR images is proposed in this letter. OFF-Net consists of global context path aggregation (GCPA) module and local rotated contrast enhance (LRCE) module, which fuses the global and local information in feature extraction. First, GCPA module is explored to integrate the global context block with path aggregation network (PAN) to learn the global background information. Second, by designing a rotation scheme based on feature map cyclic shift with four directions, LRCE module is developed to enhance the targets and suppress the background clutters in SAR images. Finally, a decoupled orientation-aware head is proposed to handle the arbitrarily rotated ships more robustly and alleviate the conflict between classification and regression tasks during detection. In addition, a high-resolution SAR-ship detection dataset (OBB-HRSDD) with rotatable bounding boxes is provided. The detection results on the SAR ship detection dataset (SSDD+) and OBB-HRSDD illustrate that our method outperforms all the compared methods. The code and OBB-HRSDD are released athttps://github.com/SJX152/papercode Ming Zhao 0009, Jiaxian Shi |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2018 | Automatic Registration of Images With Inconsistent Content Through Line-Support Region Segmentation and Geometrical Outlier RemovalabstractThe implementation of automatic image registration is still difficult in various applications. In this paper, an automatic image registration approach through line-support region segmentation and geometrical outlier removal is proposed. This new approach is designed to address the problems associated with the registration of images with affine deformations and inconsistent content, such as remote sensing images with different spectral content or noise interference, or map images with inconsistent annotations. To begin with, line-support regions, namely a straight region whose points share roughly the same image gradient angle, are extracted to address the issues of inconsistent content existing in images. To alleviate the incompleteness of line segments, an iterative strategy with multi-resolution is employed to preserve global structures that are masked at full resolution by image details or noise. Then, geometrical outlier removal is developed to provide reliable feature point matching, which is based on affine-invariant geometrical classifications for corresponding matches initialized by scale invariant feature transform. The candidate outliers are selected by comparing the disparity of accumulated classifications among all matches, instead of conventional methods which only rely on local geometrical relations. Various image sets have been considered in this paper for the evaluation of the proposed approach, including aerial images with simulated affine deformations, remote sensing optical and synthetic aperture radar images taken at different situations (multispectral, multisensor, and multitemporal), and map images with inconsistent annotations. Experimental results demonstrate the superior performance of the proposed method over the existing approaches for the whole data set. Ming Zhao 0009, Yongpeng Wu 0001, Shengda Pan, Bowen An, André Kaup |
IEEE Trans. Image Process. | 1 |
| 2017 | RFVTM: A Recovery and Filtering Vertex Trichotomy Matching for Remote Sensing Image RegistrationabstractReliable feature point matching is a vital yet challenging process in feature-based image registration. In this paper, a robust feature point matching algorithm, which is called recovery and filtering vertex trichotomy matching, is proposed to remove outliers and retain sufficient inliers for remote sensing images. A novel affine-invariant descriptor, which is called the vertex trichotomy descriptor, is proposed on the basis of that geometrical relations between any of vertices and lines are preserved after affine transformations, which is constructed by mapping each vertex into trichotomy sets. The outlier removals in vertex trichotomy matching (VTM) are implemented by iteratively comparing the disparity of the corresponding vertex trichotomy descriptors. Some inliers mistakenly validated by a large number of outliers are removed in VTM iterations, and several residual outliers that are close to the correct locations cannot be excluded with the same graph structures. Therefore, a recovery and filtering strategy is designed to recover some inliers based on identical vertex trichotomy descriptors and restricted transformation errors. Assisted with the additional recovered inliers, residual outliers can be also filtered out during the process of reaching identical graphs for the expanded vertex sets. Experimental results demonstrate the superior performance on precision and stability of this algorithm under various conditions, such as remote sensing images with large transformations, duplicated patterns, or inconsistent spectral content. Ming Zhao 0009, Bowen An, Yongpeng Wu 0001, Huynh Van Luong, André Kaup |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2015 | A Robust Delaunay Triangulation Matching for Multispectral/Multidate Remote Sensing Image RegistrationabstractA novel dual-graph-based matching method is proposed in this letter particularly for the multispectral/multidate images with low overlapping areas, similar patterns, or large transformations. First, scale invariant feature transform based matching is improved by normalizing gradient orientations and maximizing the scale ratio similarity of all corresponding points. Next, Delaunay graphs are generated for outlier removal, and the candidate outliers are selected by comparing the distinction of Delaunay graph structures. In order to bring back the inliers removed in Delaunay triangulation matching iterations and to exclude the remaining outliers, the recovery strategy equipped with the dual graph of Delaunay is explored. Inliers located in the corresponding Voronoi cells are recovered to the residual sets. The experimental results demonstrate the accuracy and robustness of the proposed algorithm for various representative remote sensing images. Ming Zhao 0009, Bowen An, Yongpeng Wu 0001, Shengli Sun |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2013 | Bi-SOGC: A Graph Matching Approach Based on Bilateral KNN Spatial Orders Around Geometric Centers for Remote Sensing Image RegistrationabstractIn this letter, Bilateral K Nearest Neighbors Spatial Orders around Geometric Centers (Bi-SOGC) is presented to match feature points for remote sensing images with large affine transformation, similar patterns or multispectral images. In Bi-SOGC, both the bilateral adjacent relations and the spatial angular orders are considered. Bilateral K Nearest Neighbors (BiKNN) descriptors are proposed to describe the adjacent information. The vertices with maximum BiKNN difference are deemed as candidate outliers. The invariant spatial angular orders for affine transformation are used to deal with outliers in pseudo isomorphic structures, geometric centers are taken as the reference points. To increase the correct matching points and eliminate stubborn outliers, a recovery strategy utilizes the addition of fresh inliers to break down the stabilized pseudo graphs of the residual sets. Experimental results demonstrate the superior performance of this algorithm under various conditions for remote sensing images. Ming Zhao 0009, Bowen An, Yongpeng Wu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |