VLDB 2026 Research / reviewers in the wild / expert
Peirong Ma
dblp:243/8892
· DBLP profile ↗
12ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0001-6391-7527ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Align-then-generate: An effective cross-modal generation paradigm for multi-label zero-shot learning
Peirong Ma, Wu Ran, Yanhui Gu, Huaqiu Chen, Zhiquan He, Hong Lu 0001 |
Pattern Recognit. | 1 |
| 2025 | ReasonAlign: A Prompt-Based Framework for Zero-Shot Schema Alignment Across Data Sources
Jiutao Zhou, Peirong Ma, Weiguang Qu, Masaru Kitsuregawa, Yanhui Gu |
IEEE Big Data | 3 |
| 2025 | Implicit Retinex Decomposition with Chromaticity Disentanglement for Low-Light Image Enhancement
Mufan Liu, Wu Ran, Zhiquan He, Zuojie Xie, Hong Lu 0001, Peirong Ma |
ACM Multimedia | 6 |
| 2025 | Unleashing the Potential of Hierarchical Region Clues for Open-Vocabulary Multi-Label ClassificationabstractOpen-vocabulary multi-label classification (OV-MLC) aims to leverage the rich multi-modal knowledge from Vision-language pre-training (VLP) models to further improve the recognition ability for unseen (novel) classes beyond the training set in multi-label scenarios. Existing OV-MLC methods only perform predictions on single hierarchical regions, and aggregate the prediction scores of these regions through simpletop-kmean pooling. This fails to unleash the potential of rich hierarchical region clues in multi-label images and does not fully exploit the discriminative information from all regions in the image, resulting in sub-optimal performance. In this work, we propose a novel OV-MLC framework to fully harness the power of multiple hierarchical region clues. Specifically, we first design a hierarchical clue gathering (HCG) module to gather different hierarchical clues, enabling more precise recognition of multiple object categories with different sizes in a multi-label image. Then, by viewing multi-label classification as single-label classification of each region within the image, we present a novel hierarchical score aggregation (HSA) approach, thereby better utilizing the predictions of each image region for each class. We also utilize a well-designed region selection strategy (RSS) to eliminate noise or background regions in an image that are irrelevant to classification, achieving higher multi-label classification accuracy. In addition, we propose a hybrid prompt learning (HPL) strategy to enhance visual-semantic consistency while preserving the generalization capability of label embeddings for unseen classes. Extensive experiments on public benchmark datasets demonstrate that our method significantly outperforms the current state-of-the-art. Peirong Ma, Wu Ran, Zhiquan He, Jian Pu, Hong Lu 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Harnessing Joint Rain-/Detail-aware Representations to Eliminate Intricate RainsabstractRecent advances in image deraining have focused on training powerful models on mixed multiple datasets comprising diverse rain types and backgrounds. However, this approach tends to overlook the inherent differences among rainy images, leading to suboptimal results. To overcome this limitation, we focus on addressing various rainy images by delving into meaningful representations that encapsulate both the rain and background components. Leveraging these representations as instructive guidance, we put forth a Context-based Instance-level Modulation (CoI-M) mechanism adept at efficiently modulating CNN- or Transformer-based models. Furthermore, we devise a rain-/detail-aware contrastive learning strategy to help extract joint rain-/detail-aware representations. By integrating CoI-M with the rain-/detail-aware Contrastive learning, we develop [CoIC](https://github.com/Schizophreni/CoIC), an innovative and potent algorithm tailored for training models on mixed datasets. Moreover, CoIC offers insight into modeling relationships of datasets, quantitatively assessing the impact of rain and details on restoration, and unveiling distinct behaviors of models given diverse inputs. Extensive experiments validate the efficacy of CoIC in boosting the deraining ability of CNN and Transformer models. CoIC also enhances the deraining prowess remarkably when real-world dataset is included. Wu Ran, Peirong Ma, Zhiquan He, Hao Ren 0002, Hong Lu 0001 |
ICLR | 2 |
| 2024 | Rainmer: Learning Multi-view Representations for Comprehensive Image Deraining and BeyondabstractWe address image deraining under complex backgrounds, diverse rain scenarios, and varying illumination conditions, representing a highly practical and challenging problem. Our approach utilizes synthetic, real-world, and nighttime datasets, wherein rich backgrounds, multiple degradation types, and diverse illumination conditions coexist. The primary challenge in training models on these datasets arises from the discrepancies among them, potentially leading to conflicts or competition during the training period. To address this issue, we first align the distribution of synthetic, real-world and nighttime datasets. Then we propose a novel contrastive learning strategy to extract multi-view (multiple) representations that effectively capture image details, degradations, and illuminations, thereby facilitating training across all datasets. Regarding multiple representations as profitable prompts for deraining, we devise a prompting strategy to integrate them into the decoding process. This contributes to a potent deraining model, dubbed Rainmer. Additionally, a spatial-channel interaction module is introduced to fully exploit cues when extracting multi-view representations. Extensive experiments on synthetic, real-world, and nighttime datasets demonstrate that Rainmer outperforms current representative methods. Moreover, Rainmer achieves superior performance on the All-in-One image restoration dataset, underscoring its effectiveness. Furthermore, quantitative results reveal that Rainmer significantly improves object detection performance on both daytime and nighttime rainy datasets. These observations substantiate the potential of Rainmer for practical applications. Wu Ran, Peirong Ma, Zhiquan He, Hong Lu 0001 |
ACM Multimedia | 2 |
| 2024 | A Transferable Generative Framework for Multi-Label Zero-Shot LearningabstractMulti-label zero-shot learning (MLZSL) is a more realistic and challenging task than single-label zero-shot learning (SLZSL), which aims to recognize multiple unseen classes in a single image. To adapt generative models to the MLZSL task and better recognize multiple unseen object categories in an image, this paper proposes a Transferable Generative Framework (TGF), which consists of a Multi-Label Semantic Embedding Autoencoders (SEAs), a Semantic-Related Multi-Label Feature Transformation Network (FTN) and a Multi-Label Feature Generation Networks (FGNs). First, SEAs adaptively encodes the class-level word vectors corresponding to each sample containing different number of classes into sample-level semantic embeddings with the same dimension. Then, FTN transforms global features extracted by a CNN pre-trained on single-label images into features that are semantic-related and more suitable for multi-label classification. Finally, FGNs generates both global and local features to better recognize the dominant and minor object categories in a multi-label image, respectively. Extensive experiments on three benchmark datasets show that TGF significantly outperforms state-of-the-arts. Specifically, compared with the previous best generative MLZSL method (i.e., Gen-MLZSL), TGF improves the mAP of the ZSL (GZSL) task by 5.4% (6.9%), 20.5% (27.9%), and 2.4% (3.9%) on NUS-WIDE, Open Images, and MS-COCO datasets, respectively. Peirong Ma, Zhiquan He, Wu Ran, Hong Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | TRNR: Task-Driven Image Rain and Noise Removal With a Few Images Based on Patch AnalysisabstractThe recent success of learning-based image rain and noise removal can be attributed primarily to well-designed neural network architectures and large labeled datasets. However, we discover that current image rain and noise removal methods result in low utilization of images. To alleviate the reliance of deep models on large labeled datasets, we propose the task-driven image rain and noise removal (TRNR) based on a patch analysis strategy. The patch analysis strategy samples image patches with various spatial and statistical properties for training and can increase image utilization. Furthermore, the patch analysis strategy encourages us to introduce the N-frequency-K-shot learning task for the task-driven approach TRNR. TRNR allows neural networks to learn from numerous N-frequency-K-shot learning tasks, rather than from a large amount of data. To verify the effectiveness of TRNR, we build a Multi-Scale Residual Network (MSResNet) for both image rain removal and Gaussian noise removal. Specifically, we train MSResNet for image rain removal and noise removal with a few images (for example, 20.0% train-set of Rain100H). Experimental results demonstrate that TRNR enables MSResNet to learn more effectively when data is scarce. TRNR has also been shown in experiments to improve the performance of existing methods. Furthermore, MSResNet trained with a few images using TRNR outperforms most recent deep learning methods trained data-driven on large labeled datasets. These experimental results have confirmed the effectiveness and superiority of the proposed TRNR. The source code is available on https://github.com/Schizophreni/MSResNet-TRNR. Wu Ran, Bohong Yang, Peirong Ma, Hong Lu 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Semantic-Related Feature Generation for Generalized Zero-Shot LearningabstractGeneralized Zero-Shot Learning (GZSL) is a challenging task. Although no visual samples of unseen classes are provided during training, the classifier must learn to recognize all classes (i.e. both seen and unseen classes). Due to the ability to generate unseen classes samples, generative models have been widely used in GZSL. However, these generative models only learn from the seen classes, so the discriminability of the unseen class features they generate is usually poor, resulting in low unseen class classification accuracy. To solve this problem, this paper proposes a novel semantic-related feature generative (SRFG) model to improve visual-semantic consistency and alleviate seen-unseen bias effectively. SRFG can generate any number of semantic-related discriminative features for both seen and unseen classes. Extensive experiments on four benchmark datasets show that the proposed model significantly outperforms the state of the arts. Peirong Ma, Wu Ran, Hong Lu 0001 |
ICME | 1 |
| 2022 | GAN-MVAE: A discriminative latent feature generation framework for generalized zero-shot learning
Peirong Ma, Hong Lu 0001, Bohong Yang, Wu Ran |
Pattern Recognit. Lett. | 1 |
| 2020 | A Variational Autoencoder with Deep Embedding Model for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning (GZSL) is a challenging task that aims to recognize not only unseen classes unavailable during training, but also seen classes used at training stage. It is achieved by transferring knowledge from seen classes to unseen classes via a shared semantic space (e.g. attribute space). Most existing GZSL methods usually learn a cross-modal mapping between the visual feature space and the semantic space. However, the mapping model learned only from the seen classes will produce an inherent bias when used in the unseen classes. In order to tackle such a problem, this paper integrates a deep embedding network (DE) and a modified variational autoencoder (VAE) into a novel model (DE-VAE) to learn a latent space shared by both image features and class embeddings. Specifically, the proposed model firstly employs DE to learn the mapping from the semantic space to the visual feature space, and then utilizes VAE to transform both original visual features and the features obtained by the mapping into latent features. Finally, the latent features are used to train a softmax classifier. Extensive experiments on four GZSL benchmark datasets show that the proposed model significantly outperforms the state of the arts. Peirong Ma |
AAAI | 1 |
| 2019 | Face hallucination from low quality images using definition-scalable inference
Peirong Ma, Zhuohao Mai, Shao-Hu Peng, Zhao Yang 0001, Li Wang 0067 |
Pattern Recognit. | 2 |