Pengyue Lin

dblp:283/3269 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
9since 2021 · last 2026
0009-0001-6193-6983ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Diffusion-Assisted Progressive Learning for Weakly Supervised Phrase Localization
abstract
Weakly supervised phrase localization (WSPL) aims to localize visual objects mentioned by given phrases, but it learns without human-annotated bounding boxes. Previous works struggle in multi-object scenarios where objects in the background often appear simultaneously with the target objects. To this end, we propose a Diffusion-Assisted PrOgressive learning framework (i.e., DAPO) for WSPL task in this paper. Specifically, we score the difficulty of training samples based on the quantity of objects and the level of semantic alignment. These samples are then used progressively during training, in an order by their difficulty scores. To address the sample imbalance problem, we propose a Generation-Assisted Tuning (GAT) method for the grounding network. First, to enrich the samples from few-object scenarios, we leverage Stable Diffusion (SD) to generate images with phrases. Second, we introduce an attention-driven scheme to direct SD's attention on the mentioned objects. Finally, we design a diffusion-guided loss, which helps the grounding network learn the objects' layouts. Extensive experiments show that our DAPO framework outperforms the strong baselines on benchmark datasets.
Pengyue Lin, Yanyang Hu, Xinjing Liu, Wenqi Jia 0006, Fangxiang Feng, Ruifan Li
AAAI1
2026 OX-MABSR: A Benchmark for Open-domain Explainable Multimodal Aspect-Based Sentiment Reasoning
abstract
Multimodal Aspect-Based Sentiment Analysis (MABSA) involves extracting aspect terms from text-image pairs and identifying their sentiments. Most existing tasks consider one fixed sentiment category with explicitly mentioned aspects. However, these tasks seldom consider expressive sentiment categories, implicit aspects, and explainability. To this end, we introduce a novel task of Open-domain Explainable Multimodal Aspect-Based Sentiment Reasoning (OX-MABSR). This task enables the prediction of open-vocabulary aspect-sentiment pairs, together with the generation of sentiment explanations and reasoning paths. To benchmark OX-MABSR task, we construct OX-MABSR-Bench, a dataset annotated with explicit and implicit aspects, expressive sentiment categories, as well as perceptual and cognitive two-level explanations. The explanations capture visual and textual cues, including aesthetics, facial expressions, scenes, and textual semantics, together with background and situational knowledge. In addition, we annotate the reasoning paths that trace how the sentiment evolves from surface cues to a deeper contextual understanding. To address OX-MABSR task, we propose MABSR-LLM. Extensive experimental results show our MABSR-LLM outperforms strong baselines. To the best of our knowledge, we are the first to provide a unified framework for open-domain and explainable MABSR.
Xinjing Liu, Zixin Xue, Pengyue Lin, Xinyu Tu, Siwei Xu, Ruifan Li
AAAI3
2026 Towards balancing the efficiency and effectiveness: a unified edit-based framework for automatic image captioning
Ruifan Li, Siwei Xu, Pengyue Lin, Fangxiang Feng, Zhangyu Ma
Multim. Syst.3
2026 A fine-grained entity understanding network for weakly supervised phrase grounding
Pengyue Lin, Ruifan Li, Fangxiang Feng, Lun Ke, Zhanyu Ma, Xiaojie Wang 0006
Pattern Recognit.1
2025 SDG-MLLM: Injecting Structured Dialogue Graphs into MLLM for Multimodal Conversational Aspect-Based Sentiment Analysis
abstract
Multimodal Conversational Aspect-based Sentiment Analysis (MCA BSA) is a challenging task for multimodal dialogue understanding. Existing works often treat the entire dialogue as a flat sequence and feed it into Large Language Models (LLMs) for pipeline-style generation. However, these methods sometimes accumulate errors and overlook critical discourse structure and fine-grained inter-word relations that are essential for accurate sentiment reasoning. To address these limitations, we propose SDG-MLLM, a unified generative framework that integrates Structured Dialogue Graphs into Multimodal LLM (MLLM) for an end-to-end MCABSA. Specifically, we construct heterogeneous dialogue graphs that capture diverse structural relations, including syntactic dependencies, coreference links, speaker turns, reply flow, semantic role labeling, and sentiment propagation paths. These graphs are encoded using a heterogeneous dialogue graph encoder, and the resulting structure-aware graph features are injected into the embedding layer of LLM. Furthermore, SDG-MLLM incorporates aligned multimodal features such as image, audio, and video cues at the utterance level to enable unified and context-aware multimodal reasoning. Experiments on the MCABSA dataset show that SDG-MLLM significantly outperforms strong baselines across multiple tasks. In addition, our method also achieved top performance in the ACM MM 2025 Grand Challenge of MCABSA. Our code is available at https://github.com/Liuxj-Anya/SDG-MLLM.
Xinjing Liu, Pengyue Lin, Xinyu Tu, Wenqi Jia 0006, Ruifan Li
ACM Multimedia2
2024 Visual Prompt Tuning for Weakly Supervised Phrase Grounding
abstract
Previous works on the task of weakly supervised phrase grounding (WSG) rely heavily on object detectors providing RoIs for the localization. However, such methods cannot be applied effectively to real-world scenarios largely because that the detectors are trained with limited categories. In this paper, we propose a refinement-based approach to WSG through fine-tuning a detector-free phrase grounding model with a visual prompt. This visual prompt is extracted from the text-related representations in CLIP. Furthermore, we combine the visual prompt with learnable features and then fine-tune the grounding network. Our experimental results significantly outperform state-of-the-art methods on the WSG task and shows the effectiveness of our method.
Pengyue Lin, Zhihan Yu, Mingcong Lu, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006
ICASSP1
2024 Triple Alignment Strategies for Zero-shot Phrase Grounding under Weak Supervision
abstract
Phrase Grounding, i.e., PG aims to locate objects referred by noun phrases. Recently, PG under weak supervision (i.e., grounding without region-level annotations) and zero-shot PG (i.e., grounding from seen categories to unseen ones) are proposed, respectively. However, for real-world applications these two approaches are limited due to slight annotations and numerable categories during training. In this paper, we propose a framework of zero-shot PG under weak supervision. Specifically, our PG framework is built on triple alignment strategies. Firstly, we propose a region-text alignment (RTA) strategy to build region-level attribute associations via CLIP. Secondly, we propose a domain alignment (DomA) strategy by minimizing the difference between distributions of seen classes in the training and those of the pre-training. Thirdly, we propose a category alignment (CatA) strategy by considering both category semantics and region-category relations. Extensive experimental results show that our proposed PG framework outperforms previous zero-shot methods and weakly-supervised methods. Our code is available at https://github.com/LinPengyue/ZS-WSG.
Pengyue Lin, Ruifan Li, Yuzhe Ji, Zhihan Yu, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006
ACM Multimedia1
2021 CFR-GAN: A Generative Model for Craniofacial Reconstruction
abstract
Craniofacial reconstruction is to reconstruct the face from the skull based on the relationship between the skull and the face to help recognition. This paper proposes a deep generative model for craniofacial reconstruction: CFR-GAN, which avoids the disadvantages of traditional methods of insufficient deep information learning ability of craniofacial data and insufficient ability to express specific features of the dataset. The model is divided into two steps: rough reconstruction and refinement reconstruction. Rough reconstruction rebuilds the overall structural content of the corresponding human head through the skull, and refinement reconstruction restorate facial feature contours. This paper constructs a dataset of 2210 two-dimensional images with craniofacial depth information, which is used to train a CFRGAN model to realize facial reconstruction of skull images. Experiments are conducted from the perspectives of qualitative analysis and quantitative analysis. The results show that CFRGAN generated image retains more identity information, and the similarity between the reconstructed face image and the real face image reaches 94%, which is better than the existing methods. In summary, CFR-GAN proposed in this paper has ability to generate high-fidelity images and is efficient at craniofacial reconstructing.
Pengyue Lin, Wen Yang 0003, Siyuan Xia, Xiaoning Liu 0001, Guohua Geng
BIBM1
2021 ANINet: a deep neural network for skull ancestry estimation
abstract
BACKGROUND: Ancestry estimation of skulls is under a wide range of applications in forensic science, anthropology, and facial reconstruction. This study aims to avoid defects in traditional skull ancestry estimation methods, such as time-consuming and labor-intensive manual calibration of feature points, and subjective results. RESULTS: This paper uses the skull depth image as input, based on AlexNet, introduces the Wide module and SE-block to improve the network, designs and proposes ANINet, and realizes the ancestry classification. Such a unified model architecture of ANINet overcomes the subjectivity of manually calibrating feature points, of which the accuracy and efficiency are improved. We use depth projection to obtain the local depth image and the global depth image of the skull, take the skull depth image as the object, use global, local, and local + global methods respectively to experiment on the 95 cases of Han skull and 110 cases of Uyghur skull data sets, and perform cross-validation. The experimental results show that the accuracies of the three methods for skull ancestry estimation reached 98.21%, 98.04% and 99.03%, respectively. Compared with the classic networks AlexNet, Vgg-16, GoogLenet, ResNet-50, DenseNet-121, and SqueezeNet, the network proposed in this paper has the advantages of high accuracy and small parameters; compared with state-of-the-art methods, the method in this paper has a higher learning rate and better ability to estimate. CONCLUSIONS: In summary, skull depth images have an excellent performance in estimation, and ANINet is an effective approach for skull ancestry estimation.
Pengyue Lin, Siyuan Xia, Jiang Yi, Wen Yang 0003, Xiaoning Liu 0001, Guohua Geng
BMC Bioinform.1
2020 Ancestry Estimation of Skull in Chinese Population Based on Improved Convolutional Neural Network
abstract
The estimation of ancestry is an essential benchmark for positive identification of heavily decomposed bodies that are recovered in a variety of death and crime scenes. Aiming at the problem of skull ancestry estimation, this paper proposes an improved convolutional neural network method to realize ancestry estimation. We use the six-angle images of the skull as the input of the network. By improving the basic model LeNet5 of the convolutional neural network, we preserve the depth semantics and content information of the image, reduce the number of parameters, and ensure the learning ability of network features. In the experiment, 156 yellow skulls from northern China and 178 white skulls from Xinjiang were used as subjects, 80% of skull samples were used as training sets and 20% as test sets. Experiments on the training set and test set show that the improved CNN network architecture achieves 95.88% accuracy on the training set and 95.52% accuracy on the test set. In addition, we also designed experiments on the contribution of various parts of the skull to ancestor identification. The experimental results show that each region of the skull is useful for ancestor identification, but the effect is different. Compared with other networks, the network structure of this paper has the highest accuracy and better performance.
Wen Yang 0003, Pengyue Lin, Guohua Geng, Xiaoning Liu 0001, Kang Li 0005
BIBM3