Peipei Zhu

dblp:147/3534 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
3since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Vision and language · 82% Image recognition and object detection · 18%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
image captioning
1.422024
Prompt-Based Learning for Unpaired Image Captioning · IEEE Trans. Multim. 2024
Unpaired Image Captioning by Image-Level Weakly-Supervised Visual Concept Recognition · IEEE Trans. Multim. 2023
Computer vision › Vision and language › image captioning
unpaired image captioning
1.422024
Prompt-Based Learning for Unpaired Image Captioning · IEEE Trans. Multim. 2024
Unpaired Image Captioning by Image-Level Weakly-Supervised Visual Concept Recognition · IEEE Trans. Multim. 2023
Computer vision › Image recognition and object detection › object recognition
weakly supervised object recognition
0.712023
Unpaired Image Captioning by Image-Level Weakly-Supervised Visual Concept Recognition · IEEE Trans. Multim. 2023
Computer vision › Vision and language › vision-language model
vision-language pre-trained model
0.212024
Prompt-Based Learning for Unpaired Image Captioning · IEEE Trans. Multim. 2024

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 0.8prompt learning · 0.8adversarial learning · 0.8weakly supervised learning · 0.7graph neural network · 0.7
YearPublicationVenuePosition
2026 Rule-Semantic Generative Calibration Blur Detection for UAV Imagery
Yihan Wen, Zhuo Zhang 0020, Xianping Ma, Peipei Zhu, Jinglei Li, Guanchong Niu, Qiguang Miao
IEEE Trans. Circuits Syst. Video Technol.4
2024 Prompt-Based Learning for Unpaired Image Captioning
abstract
Unpaired Image Captioning (UIC) has been developed to learn image descriptions from unaligned vision-language sample pairs. Existing works usually tackle this task using adversarial learning and visual concept reward based on reinforcement learning. However, these existing works were only able to learn limited cross-domain information in vision and language domains, which restrains the captioning performance of UIC. Inspired by the success of Vision-Language Pre-Trained Models (VL-PTMs) in this research, we attempt to infer the cross-domain cue information about a given image from the large VL-PTMs for the UIC task. This research is also motivated by recent successes of prompt learning in many downstream multi-modal tasks, including image-text retrieval and vision question answering. In this work, a semantic prompt is introduced and aggregated with visual features for more accurate caption prediction under the adversarial learning framework. In addition, a metric prompt is designed to select high-quality pseudo image-caption samples obtained from the basic captioning model and refine the model in an iterative manner. Extensive experiments on the COCO and Flickr30 K datasets validate the promising captioning ability of the proposed model. We expect that the proposed prompt-based UIC model will stimulate a new line of research for the VL-PTMs based captioning.
Peipei Zhu, Xiao Wang 0014, Lin Zhu 0012, Zhenglong Sun 0001, Wei-Shi Zheng 0001, Yaowei Wang 0001, Chang Wen Chen
IEEE Trans. Multim.1
2023 Unpaired Image Captioning by Image-Level Weakly-Supervised Visual Concept Recognition
abstract
The goal of unpaired image captioning (UIC) is to describe images without using image-caption pairs in the training phase. Although challenging, we expect the task can be accomplished by leveraging images aligned with visual concepts. Most existing studies use off-the-shelf algorithms to obtain the visual concepts because the Bounding Box (BBox) labels or relationship-triplet labels used for training are expensive to acquire. To avoid exhaustive annotations, we propose a novel approach to achieve cost-effective UIC. Specifically, we adopt image-level labels to optimize the UIC model in a weakly-supervised manner. For each image, we assume that only the image-level labels are available without specific locations and numbers. The image-level labels are utilized to train a weakly-supervised object recognition model to extract object information (e.g., instance), and the extracted instances are adopted to infer the relationships among different objects using an enhanced graph neural network (GNN). The proposed approach achieves comparable or even better performance compared with previous methods without expensive annotations. Furthermore, we design an unrecognized object (UnO) loss to improve the alignment of the inferred object and relationship information with the images. It can effectively alleviate the issue encountered by existing UIC models when generating sentences with nonexistent objects. To the best of our knowledge, this is the first attempt to address the problem of Weakly-Supervised visual concept recognition for UIC (WS-UIC) based only on image-level labels. Extensive experiments demonstrate that the proposed method achieves inspiring results on the COCO dataset while significantly reducing the labeling cost.
Peipei Zhu, Xiao Wang 0014, Yong Luo 0002, Zhenglong Sun 0001, Wei-Shi Zheng 0001, Yaowei Wang 0001, Chang Wen Chen
IEEE Trans. Multim.1
2020 Event Detection with Document Structure and Graph Modelling
Peipei Zhu, Hongling Wang, Shoushan Li, Guodong Zhou 0001
NLPCC (1)1