VLDB 2026 Research / reviewers in the wild / expert
Lei Ma 0004
dblp:20/6534-4
· DBLP profile ↗
25ranked-venue papers
15as first author
18since 2021 · last 2026
0000-0002-8298-3703ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 7 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Class-agnostic and semantic-aware fusing network with optimal transport for weakly supervised object localization
Lei Ma 0004, Hongbo Wen, Hanyu Hong, Fanman Meng, Qingbo Wu 0001 |
Expert Syst. Appl. | 1 |
| 2026 | Unsupervised deep hashing based on multi-scale aggregation and optimal transport matching for image retrieval
Lei Ma 0004, Hao Pei, Lei Wang 0068, Ying Zhu 0002, Yu Shi 0004, Hanyu Hong, Xinyu Dai, Fanman Meng, Qingbo Wu 0001 |
Neurocomputing | 1 |
| 2026 | L-MedViT: A hierarchical hybrid network for fine-grained diabetic retinopathy grading
Sanli Yi, Lei Ma 0004 |
Image Vis. Comput. | 3 |
| 2025 | Optimal Transport Quantization Based on Cross-X Semantic Hypergraph Learning for Fine-Grained Image RetrievalabstractLarge-scale fine-grained image retrieval aims to learn compact discriminative feature representations based on mining the subtle distinctions between visually similar objects. However, existing fine-grained image retrieval methods focus on enhancing the attention to the discriminative regions within single images, which barely exploit the high-order relational information between the global features and local region features across different images. Thus, the over-fitting problem of complex personalized differences cannot be effectively solved. In addition, existing unconstrained vector quantization methods tend to assign unquantized feature vectors to a few major codewords, which are unable to effectively distinguish the quantized features and reduce the redundant information. To address these issues, we propose a novel optimal transport quantization method based on cross-X semantic hypergraph learning for large-scale fine-grained image retrieval. Specifically, we first introduce a cross-layer multi-scale aggregation module to extract the global features and local region features. Subsequently, we build a semantic hypergraph to model the high-order correlations between the global features and local region features extracted from different layers, different scales and different images, which can alleviate the over-fitting problem of complex personalized differences by suppressing sample-level and background noise. Moreover, we introduce an error regularization term into the progressive asymmetric quantization loss to reduce the quantization errors and preserve the semantic similarity. Finally, we attempt to introduce the code balance and uncorrelated constraints into the multi-codebook quantization framework to improve the utilization efficiency of codewords and reduce the redundant information, which can be approximated by solving the optimal transport problem. Experimental results on several fine-grained image datasets demonstrate that the proposed method outperforms the state-of-the-art fine-grained image retrieval methods. Lei Ma 0004, Yu Shi 0004, Fanman Meng, Qingbo Wu 0001, Hanyu Hong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Progressive Learning-Based Jitter Distortion Correction for Remote Sensing Images of Time Delay and Integration CameraabstractThe widespread use of time delay and integration charge-coupled device (TDI CCD) technology in high-resolution spaceborne optical cameras has made high-frequency jitter effects a common issue, resulting in different levels of distortion in images. Current methods mostly concentrate on correction of obviously high levels of geometric distortion. Focusing on low levels of geometric distortion, which are more difficult to accurately detect, this paper proposes a progressive learning-based correction method for high-frequency jitter distortion in remote sensing images from spaceborne TDI CCD cameras, utilizing a Generative Adversarial Network (GAN). First, a distorted dataset with diverse jitter levels for progressive training is generated through jitter simulation model by adjusting the parameters. Then, a GAN model is employed for the correction task. The generator consists of the Distortion Net for geometric distortion correction and the Detail Enhancement Net for image detail restoration. Finally, a progressive learning strategy is used to gradually enhance the ability of network to correct minor geometric distortion. The proposed method is validated using simulated images and real-world satellite images. Experimental results demonstrate that the proposed method outperforms existing restoration methods both in simulated datasets and practical scenarios. Ying Zhu 0002, Mi Wang, Jun Pan 0001, Hanyu Hong, Lei Ma 0004, Lei Wang 0068 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Generative Adversarial Network-Based Jitter Distortion Correction for High Resolution Spaceborne ImagesabstractThis paper presents a Generative Adversarial Network (GAN)-based jitter distortion correction method for spaceborne images of Time Delay Integration (TDI) Charge-Coupled Device (CCD) camera. This method leverages the advantages of GANs and combines content loss, adversarial loss, and perceptual loss to effectively repair distorted images while preserving image details, which does not rely on jitter information captured by high-frequency attitude sensors, nor depends on the analysis of overlapping areas between different bands in multispectral images. The experimental results show that the proposed method achieves automated correction of geometric distortions and has shown promising restoration results on real distorted images captured by Yaogan-26 satellite and GaoFen satellite, which achieves better results than other blind restoration methods. Ying Zhu 0002, Lei Wang 0068, Lei Ma 0004, Jinmeng Wu |
IGARSS | 4 |
| 2024 | WDTSNet: Wavelet Decomposition Two-Stage Network for Infrared Thermal Radiation Effect CorrectionabstractRecently, infrared thermal radiation effect correction methods are dominated by removing bias field in spatial domain. Since they do not consider the low-frequency characteristics of thermal radiation bias field and the high-frequency information of image content, these methods often fail in the enhancement of contrast and details. To address this problem, we propose a novel wavelet decomposition two-stage network for infrared thermal radiation effect correction, named WDTSNet. Through wavelet decomposition, we construct a low-frequency thermal radiation effect coarse correction subnetwork (LFCCSN) and a high-frequency detail enhancement fine correction subnetwork (HFFCSN), respectively. Firstly, we take the small size low-frequency component of the degraded image after discrete wavelet transformation (DWT) as the input of the first stage LFCCSN and propose an intra-block multiscale residual dense module (IMRDM) to complete the coarse correction and contrast enhancement through different scales of receptive fields and intra-block channel information interaction. Secondly, we perform inverse discrete wavelet transformation (IDWT) to obtain the input of the second stage HFFCSN, and build a high-frequency gated residual module (HGRM) in HFFCSN to remove residual thermal radiation bias field and acquire the enhanced high-frequency information. In addition, we further design dual-branch cross-scale attention fusion module (DCAFM) between encoders and decoders to effectively aggregate the cross-scale information flow. Extensive experiments on simulated and real infrared images demonstrate that the proposed WDTSNet performs well on enhancing contrast and details than existing methods. The code will be publicly available upon acceptance. Yu Shi 0004, Yixin Zhou, Lei Ma 0004, Lei Wang 0068, Hanyu Hong |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Deep Progressive Asymmetric Quantization Based on Causal Intervention for Fine-Grained Image RetrievalabstractIn the field of computer vision, fine-grained image retrieval is an extremely challenging task due to the inherently subtle intra-class object variations. In addition, the high-dimensional real-valued features extracted from large-scale fine-grained image datasets slow the retrieval speed and increase the storage cost. To solve above issues, existing fine-grained image retrieval methods mainly focus on finding more discriminative local regions for generating discriminative and compact hash codes, which achieve limited fine-grained image retrieval performance due to the large quantization errors and the confounding granularities and context of discriminative parts, i.e., the correct recognition of fine-grained objects mainly attribute to the discriminative parts and their context. To learn robust causal features and reduce the quantization errors, we propose a deep progressive asymmetric quantization (DPAQ) method based on causal intervention to learn compact and robust descriptions for fine-grained image retrieval task. Specifically, we introduce a structural causal model to learn robust casual features via causal intervention for fine-grained visual recognition. Subsequently, we design a progressive asymmetric quantization layer in the feature embedding space, which can preserve the semantic information and reduce the quantization errors sufficiently. Finally, we incorporate both the fine-grained image classification and retrieval tasks into an end-to-end deep learning architecture for generating robust and compact descriptions. Experimental results on several fine-grained image retrieval datasets demonstrate that the proposed DPAQ method performs the best for fine-grained image retrieval task and surpasses the state-of-the art fine-grained hashing methods by a large margin. Lei Ma 0004, Hanyu Hong, Fanman Meng, Qingbo Wu 0001, Jinmeng Wu |
IEEE Trans. Multim. | 1 |
| 2024 | Logit Variated Product Quantization Based on Parts Interaction and Metric Learning With Knowledge Distillation for Fine-Grained Image RetrievalabstractImage retrieval with fine-grained categories is an extremely challenging task due to the high intraclass variance and low interclass variance. Most previous works have focused on localizing discriminative image regions in isolation, but have rarely exploited correlations across the different discriminative regions to alleviate intraclass differences. In addition, the intraclass compactness of embedding features is ensured by extra regularization terms that only exist during the training phase, which appear to generalize less well in the inference phase. Finally, the information granularity of the distance measure should distinguish subtle visual differences and the correlation between the embedding features and the quantized features should be maximized sufficiently. To address the above issues, we propose a logit variated product quantization method based on part interaction and metric learning with knowledge distillation for fine-grained image retrieval. Specifically, we introduce a causal context module into the deep navigator to generate discriminative regions and utilize a channelwise cross-part fusion transformer to model the part correlations while alleviating intraclass differences. Subsequently, we design a logit variation module based on a weighted sum scheme to further reduce the intraclass variance of the embedding features directly and enhance the learning power of the quantization model. Finally, we propose a novel product quantization loss based on metric learning and knowledge distillation to enhance the correlation between the embedding features and the quantized features and allow the quantization features to learn more knowledge from the embedding features. The experimental results on several fine-grained datasets demonstrate that the proposed method is superior to state-of-the-art fine-grained image retrieval methods. Lei Ma 0004, Hanyu Hong, Fanman Meng, Qingbo Wu 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Unsupervised Encoder-Decoder Model for Anomaly Prediction Task
Jinmeng Wu, Pengcheng Shu, Hanyu Hong, Xingxun Li, Lei Ma 0004, Yaozong Zhang, Ying Zhu 0002, Lei Wang 0068 |
MMM (2) | 5 |
| 2023 | Scribble-attention hierarchical network for weakly supervised salient object detection in optical remote sensing images
Lei Ma 0004, Hanyu Hong, Yaozong Zhang, Lei Wang 0068, Jinmeng Wu |
Appl. Intell. | 1 |
| 2023 | Joint ordinal regression and multiclass classification for diabetic retinopathy grading with transformers and CNNs fusion network
Lei Ma 0004, Qihang Xu, Hanyu Hong, Yu Shi 0004, Ying Zhu 0002, Lei Wang 0068 |
Appl. Intell. | 1 |
| 2023 | DDABNet: a dense Do-conv residual network with multisupervision and mixed attention for image deblurring
Yu Shi 0004, Zhigao Huang, Jisong Chen, Lei Ma 0004, Lei Wang 0068, Hanyu Hong |
Appl. Intell. | 4 |
| 2023 | Question-aware dynamic scene graph of local semantic representation learning for visual question answering
Jinmeng Wu, Fulin Ge, Hanyu Hong, Yu Shi 0004, Yanbin Hao, Lei Ma 0004 |
Pattern Recognit. Lett. | 6 |
| 2023 | Complementary Parts Contrastive Learning for Fine-Grained Weakly Supervised Object Co-LocalizationabstractThe aim of weakly supervised object co-localization is to locate different objects of the same superclass in a dataset. Recent methods achieve impressive co-localization performance by multiple instance learning and self-supervised learning. However, these methods ignore the common part information shared by fine-grained objects and the influence of the complementary parts on the co-localization of the fine-grained objects. To solve these issues, we propose a complementary parts contrastive learning method for fine-grained weakly supervised object co-localization. The proposed method follows such an assumption that fine-grained object parts with the same/different semantic meaning should have similar/dissimilar feature representations in the feature space. The proposed method tackles two critical issues in this task:$i)$how to spread the model’s attention and suppress the complex background noise, and$ii)$how to leverage the cross-category common parts information to mitigate the context co-occurrence problem. To address$i)$, we attempt to integrate local and context cues via three types of attention including self-supervised attention, channel, and spatial attention to spread the model’s attention toward automatically identifying and localizing most discriminative parts of objects in the fine-grained images. To solve$ii)$, we propose a cross-category object complementarity part contrastive learning module to identify the extracted part regions with different semantic information by pulling the same part features closer and pushing different part features away, which can mitigate the confounding bias caused by the co-occurrence surroundings within specific classes. Extensive qualitative and quantitative evaluations demonstrate the effectiveness of the proposed method on four fine-grained co-localization datasets: CUB-200–2011, Stanford Cars, FGVC-Aircraft, and Stanford Dogs. Code and models are available athttps://github.com/Zhao-fan/CPCL. Lei Ma 0004, Hanyu Hong, Lei Wang 0068, Ying Zhu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | PolSAR-SSN: An End-to-End Superpixel Sampling Network for PolSAR Image ClassificationabstractPolarimetric synthetic aperture radar (PolSAR) image classification is one of the fundamental research areas in remote sensing. Superpixels can provide boundary constraint information and are widely used in PolSAR image interpretation. However, traditional machine learning superpixel algorithms have many limitations for PolSAR image interpretation. Pseudo-color images are usually used as the superpixel algorithm inputs, and the loss of polarimetric information will decrease the performance. In addition, the superpixel algorithms are difficult to incorporate into state-of-the-art deep learning models and cannot be trained in an end-to-end manner. In this letter, a trainable end-to-end deep superpixel network is proposed for PolSAR image classification. The inputs of the proposed method can be any low/middle-level polarimetric features of a PolSAR image and the rich polarimetric feature representation can be learned. The produced superpixels of the proposed method are more concentrated near the land cover boundaries and can significantly improve the performance of PolSAR image classification. Experimental results show that the overall accuracies of the proposed method are approximately 2.57% and 1.44% higher than traditional superpixel algorithms on two PolSAR datasets and surpass some well-known deep learning methods. Lei Wang 0068, Hanyu Hong, Yaozong Zhang, Jinmeng Wu, Lei Ma 0004, Ying Zhu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2022 | Progressive Attention-Based Feature Recovery With Scribble Supervision for Saliency Detection in Optical Remote Sensing ImageabstractSalient object detection (SOD) task for optical remote sensing images (RSIs) plays an important role in many remote sensing applications. Most of the existing methods train their networks depending on a large amount of pixel-wise datasets. However, such expensive and time-consuming training setting prevents the approaches becoming flexible and scalable solutions. To this end, we explore efficient SOD for optical RSIs based on easily accessible weak supervision source. In this work, we propose a novel end-to-end progressive attention-based feature recovery framework with scribble supervision. Specifically, to better locate challenging salient objects in optical RSIs, an object position module (OPM) is proposed to capture and enhance the long-range semantic dependence of objects’ position information, which depends on the complementary attention mechanism. And to restore the entire salient objects, a context refinement module (CRM) is proposed, which extract local contextual information for better propagating high-level semantics to low-level details. Moreover, to improve the adaptability of the network to the changing scenarios of optical RSIs, we propose a salient region correcting (SRC) mechanism to help the predicted salient regions rectify their saliency values by constraining the saliency relationship between predictions from different augmentation models. In addition, due to the lack of dataset for weakly supervised SOD for optical RSIs, we relabeled an existing large-scale optical RSIs dataset with scribbles, namely EORSSD-S. Experimental results on benchmark datasets demonstrate that the proposed method can outperform other weakly supervised SOD methods. And the proposed method even outperformed some fully supervised methods. https://github.com/melonless/PAFR. Lei Ma 0004, Zhenghua Huang, Haiwen Yuan |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Learning discrete class-specific prototypes for deep semantic hashing
Lei Ma 0004, Yu Shi 0004, Likun Huang, Zhenghua Huang, Jinmeng Wu |
Neurocomputing | 1 |
| 2020 | Discriminative deep metric learning for asymmetric discrete hashing
Lei Ma 0004, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan |
Neurocomputing | 1 |
| 2020 | Parametric Deformable Exponential Linear Units for deep neural networks
Qishang Cheng, Hongliang Li 0001, Qingbo Wu 0001, Lei Ma 0004, King Ngi Ngan |
Neural Networks | 4 |
| 2020 | Correlation Filtering-Based Hashing for Fine-Grained Image RetrievalabstractThe low storage and strong representation capabilities of hash codes for image retrievalhas made hashing technologies very popular. Several existing deep hashing methods focuson the task of general image retrieval, while neglecting the task of fine-grained image retrieval. Recently, some fine-grained hashing methods have been proposed to capture the subtle differences, which mainly utilize the single-modality visual features to solve the discriminative region localization while ignoring the semantic information. In this letter, we propose a correlation filtering hashing (CFH) method to learn discrete binary codes, which can adequately take advantage of the cross-modal correlation between the semantic information and the visual features for discriminative region localization. Specifically, we utilize a feature pyramid network to learn multi-level visual features. Subsequently, the label vector is embedded into the visual space, which can be used as a correlation filter on the feature maps to capture the latent location of objects. Finally, weperform global average pooling over the output maps and concatenate the features of different levels to produce the hash codes of query images. Extensive experiments on two fine-grained datasets show that the proposed CFH outperforms the state-of-the-art hashing methods. Lei Ma 0004, Yu Shi 0004, Jinmeng Wu |
IEEE Signal Process. Lett. | 1 |
| 2018 | Multi-task Learning for Deep Semantic HashingabstractDeep learning to hash has emerged as a popular technique for large-scale image retrieval. Existing deep learning to hash methods seek to solve the single retrieval task within one stream framework or jointly solve the retrieval task and the classification task within two stream framework. Consequently, the semantic information is not fully exploited to generate compact and discriminative hash codes. In this paper, we propose a multi-task learning architecture for deep semantic hashing (MLDH), which incorporates the retrieval task and the classification task within one-stream framework. Specifically, we introduce a COCO loss to learn compact binary codes for the classification task. For the retrieval task, we introduce a pairwise loss to learn discriminative binary codes. Finally, these two tasks are investigated into one-stream deep learning framework. Extensive experiments show that MLDH can outperform state-of-the-art methods on benchmark datasets. Lei Ma 0004, Hongliang Li 0001, Qingbo Wu 0001, Chao Shang 0001, King Ngi Ngan |
VCIP | 1 |
| 2018 | Global and local semantics-preserving based deep hashing for cross-modal retrieval
Lei Ma 0004, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan |
Neurocomputing | 1 |
| 2017 | Manifold-ranking embedded order preserving hashing for image semantic retrieval
Lei Ma 0004, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Linfeng Xu 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2017 | Learning Efficient Binary Codes From High-Level Feature Representations for Multilabel Image RetrievalabstractDue to the efficiency and effectiveness of hashing technologies, they have become increasingly popular in large-scale image semantic retrieval. However, existing hash methods suppose that the data distributions satisfy the manifold assumption that semantic similar samples tend to lie on a low-dimensional manifold, which will be weakened due to the large intraclass variation. Moreover, these methods learn hash functions by relaxing the discrete constraints on binary codes to real value, which will introduce large quantization loss. To tackle the above problems, this paper proposes a novel unsupervised hashing algorithm to learn efficient binary codes from high-level feature representations. More specifically, we explore nonnegative matrix factorization for learning high-level visual features. Ultimately, binary codes are generated by performing binary quantization in the high-level feature representations space, which will map images with similar (visually or semantically) high-level feature representations to similar binary codes. To solve the corresponding optimization problem involving nonnegative and discrete variables, we develop an efficient optimization algorithm to reduce quantization loss with guaranteed convergence in theory. Extensive experiments show that our proposed method outperforms the state-of-the-art hashing methods on several multilabel real-world image datasets. Lei Ma 0004, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, King Ngi Ngan |
IEEE Trans. Multim. | 1 |