VLDB 2026 Research / reviewers in the wild / expert
Pengna Li
dblp:246/5858
· DBLP profile ↗
8ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0002-8477-8340ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Think before Go: Hierarchical Reasoning for Image-goal NavigationabstractPengna Li, Kangyi Wu, Shaoqing Xu, Fang Li, Lin Zhao, Long Chen, Zhi-Xin Yang, Nanning Zheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Pengna Li, Kangyi Wu, Shaoqing Xu, Long Chen 0005, Nanning Zheng 0001 |
ACL (1) | 1 |
| 2025 | REGNav: Room Expert Guided Image-Goal NavigationabstractImage-goal navigation aims to steer an agent towards the goal location specified by an image. Most prior methods tackle this task by learning a navigation policy, which extracts visual features of goal and observation images, compares their similarity and predicts actions. However, if the agent is in a different room from the goal image, it's extremely challenging to identify their similarity and infer the likely goal location, which may result in the agent wandering around. Intuitively, when humans carry out this task, they may roughly compare the current observation with the goal image, having an approximate concept of whether they are in the same room before executing the actions. Inspired by this intuition, we try to imitate human behaviour and propose a Room Expert Guided Image-Goal Navigation model~(REGNav) to equip the agent with the ability to analyze whether goal and observation images are taken in the same room. Specifically, we first pre-train a room expert with an unsupervised learning technique on the self-collected unlabelled room images. The expert can extract the hidden room style information of goal and observation images and predict their relationship about whether they belong to the same room. In addition, two different fusion approaches are explored to efficiently guide the agent navigation with the room relation knowledge. Extensive experiments show that our REGNav surpasses prior state-of-the-art works on three popular benchmarks. Pengna Li, Kangyi Wu, Jingwen Fu, Sanping Zhou |
AAAI | 1 |
| 2025 | Event-Equalized Dense Video CaptioningabstractDense video captioning aims to localize and caption all events in arbitrary untrimmed videos. Although previous methods have achieved appealing results, they still face the issue of temporal bias, i.e, models tend to focus more on events with certain temporal characteristics. Specifically, 1) the temporal distribution of events in training datasets is uneven. Models trained on these datasets will pay less attention to out-of-distribution events. 2) long-duration events have more frame features than short ones and will attract more attention. To address this, we argue that events, with varying temporal characteristics, should be treated equally when it comes to dense video captioning. Intuitively, different events tend to have distinct visual differences due to varied camera views, backgrounds, or subjects. Inspired by that, we intend to utilize visual features to have an approximate perception of possible events and pay equal attention to them. In this paper, we introduce a simple but effective framework, called Event-Equalized Dense Video Captioning (E2DVC) to overcome the temporal bias and treat all possible events equally. Experimental results on ActivityNet Captions and YouCook2 dataset validate the effectiveness of the proposed methods and show State-of-the-art (SOTA) performance on dense video captioning. Kangyi Wu, Pengna Li, Jingwen Fu, Yang Wu 0001, Yuhan Liu 0006, Jinjun Wang, Sanping Zhou |
CVPR | 2 |
| 2025 | Mind the Gap: Aligning Vision Foundation Models to Image Feature MatchingabstractLeveraging the vision foundation models has emerged as a mainstream paradigm that improves the performance of image feature matching. However, previous works have ignored the misalignment when introducing the foundation models into feature matching. The misalignment arises from the discrepancy between the foundation models focusing on single-image understanding and the cross-image understanding requirement of feature matching. Specifically, 1) the embeddings derived from commonly used foundation models exhibit discrepancies with the optimal embeddings required for feature matching; 2) lacking an effective mechanism to leverage the single-image understanding ability into cross-image understanding. A significant consequence of the misalignment is they struggle when addressing multi-instance feature matching problems. To address this, we introduce a simple but effective framework, called IMD (Image feature Matching with a pre-trained Diffusion model) with two parts: 1) Unlike the dominant solutions employing contrastive-learning based foundation models that emphasize global semantics, we integrate the generative-based diffusion models to effectively capture instance-level details. 2) We leverage the prompt mechanism in generative model as a natural tunnel, propose a novel cross-image interaction prompting module to facilitate bidirectional information interaction between image pairs. To more accurately measure the misalignment, we propose a new benchmark called IMIM, which focuses on multi-instance scenarios. Our proposed IMD establishes a new state-of-the-art in commonly evaluated benchmarks, and the superior improvement 12% in IMIM indicates our method efficiently mitigates the misalignment. Yuhan Liu 0006, Jingwen Fu, Yang Wu 0001, Kangyi Wu, Pengna Li, Jiayi Wu 0002, Sanping Zhou, Jingmin Xin |
ICCV | 5 |
| 2025 | Meta channel masking for cross-domain few-shot image classificationabstractCross-domain Few-shot Learning (CD-FSL) aims to address the challenges of FSL where significant domain gaps exist between source and target image datasets. Unlike many existing CD-FSL methods that utilize an auxiliary target dataset with a few labeled target images to enhance model generalization, our approach directly tackles the limitations imposed by the reliance on source-specific knowledge. We observe that models trained on unbalanced datasets tend to overfit to source-specific features, which, while effective in the source domain, generalize poorly to the target image domain. To address this, we introduce a novel dropout-based framework named Meta Channel Masking (MCM). This framework attenuates the learning of model channels on the source domain by dynamically masking source feature channels during training. In contrast to traditional dropout techniques that manually set masking probabilities based on statistical assumptions about the source data, our MCM framework employs a meta-learning process that automatically adjusts channel mask probabilities. This adjustment is informed by auxiliary target data, effectively minimizing few-shot loss on the auxiliary target dataset and thereby enhancing the model’s generalization capabilities in the target domain. Our extensive experiments across various image classification benchmark datasets demonstrate that our framework outperforms state-of-the-art methods. Siqi Hui, Sanping Zhou, Ye Deng 0005, Pengna Li, Jinjun Wang |
Neurocomputing | 4 |
| 2024 | Semantic-aware Representation Learning for Homography EstimationabstractHomography estimation is the task of determining the transformation from an image pair. Our approach focuses on employing detector-free feature matching methods to address this issue. Previous work has underscored the importance of incorporating semantic information, however there still lacks an efficient way to utilize semantic information. Previous methods suffer from treating the semantics as a pre-processing, causing the utilization of semantics overly coarse-grained and lack adaptability when dealing with different tasks. In our work, we seek another way to use the semantic information, that is semantic-aware feature representation learning framework. Based on this, we propose SRMatcher, a new detector-free feature matching method, which encourages the network to learn integrated semantic feature representation. Specifically, to capture precise and rich semantics, we leverage the capabilities of recently popularized vision foundation models (VFMs) trained on extensive datasets. Then, a cross-images Semantic-aware Fusion Block (SFB) is proposed to integrate its fine-grained semantic features into the feature representation space. In this way, by reducing errors stemming from semantic inconsistencies in matching pairs, our proposed SRMatcher is able to deliver more accurate and realistic outcomes. Extensive experiments show that SRMatcher surpasses solid baselines and attains SOTA results on multiple real-world datasets. Compared to the previous SOTA approach GeoFormer, SRMatcher increases the area under the cumulative curve (AUC) by about 11% on HPatches. Additionally, the SRMatcher could serve as a plug-and-play framework for other matching methods like LoFTR, yielding substantial precision improvement. Yuhan Liu 0006, Qianxin Huang, Siqi Hui, Jingwen Fu, Sanping Zhou, Kangyi Wu, Pengna Li, Jinjun Wang |
ACM Multimedia | 7 |
| 2023 | Pseudo Labels Refinement with Intra-Camera Similarity for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (Re-ID) aims to retrieve person images across cameras without any identity labels. Most clustering-based methods roughly divide image features into clusters and neglect the feature distribution noise caused by domain shifts among different cameras, leading to inevitable performance degradation. To address this challenge, we propose a novel label refinement framework with clustering intra-camera similarity. Intra-camera feature distribution pays more attention to the appearance of pedestrians and labels are more reliable. We conduct intra-camera training to get local clusters in each camera, respectively, and refine inter-camera clusters with local results. We hence train the Re-ID model with refined reliable pseudo labels in a self-paced way. Extensive experiments demonstrate that the proposed method surpasses state-of-the-art performance. Code is available at https://github.com/leeBooMla/ICSR. Pengna Li, Kangyi Wu, Sanping Zhou, Qianxin Huang, Jinjun Wang |
ICIP | 1 |
| 2019 | Low-Shot Palmprint Recognition Based on Meta-Siamese NetworkabstractPalmprint is one of the discriminant biometrical features of humans. Recognizing palmprints in complex environments is a significant multimedia task, which is highly suitable for applications in information security and forensics. Recently, deep learning-based recognition methods have improved the accuracy and robustness of recognition results to a new level. However, obtaining the required large amount of training data and labels is impracticable in practical scenarios. Therefore, in this paper, we exploit few-shot learning for palmprint recognition. We propose Meta-Siamese network based on Siamese network. Specifically, we train this network episodically with a more flexible framework to learn both the feature embedding and the deep similarity metric function. Moreover, we extend our model to zero-shot recognition tasks based on deep hashing network. Experiment result shows competitive improvements compared to baseline methods in eight different datasets. Xuefeng Du, Dexing Zhong, Pengna Li |
ICME | 3 |