VLDB 2026 Research / reviewers in the wild / expert
Jiawei Peng 0001
dblp:244/7365-1
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0005-1751-239XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Vision and language · 41% Language models and text generation · 24% Efficient and distributed learning · 15% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
visual question answering |
1.5 | 2 | 2024 | LIVE: Learnable In-Context Vector for Visual Question Answering · NeurIPS 2024 How to Configure Good In-Context Sequence for Visual Question Answering · CVPR 2024 |
Natural language and speech › Language models and text generation
in-context learning |
1.0 | 2 | 2024 | LIVE: Learnable In-Context Vector for Visual Question Answering · NeurIPS 2024 How to Configure Good In-Context Sequence for Visual Question Answering · CVPR 2024 |
Machine learning › Efficient and distributed learning › data selection
coreset selection |
0.9 | 1 | 2025 | Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization · ACM Multimedia 2025 |
Computer vision › Image recognition and object detection
image classification |
0.9 | 1 | 2025 | Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization · ACM Multimedia 2025 |
Natural language and speech › Language models and text generation › in-context learning
in-context vectors |
0.8 | 1 | 2024 | LIVE: Learnable In-Context Vector for Visual Question Answering · NeurIPS 2024 |
Computer vision › Vision and language
image captioning |
0.7 | 1 | 2023 | Transforming Visual Scene Graphs to Image Captions · ACL (1) 2023 |
Computer vision › Vision and language › image captioning › structured image captioning
scene graph-based captioning |
0.7 | 1 | 2023 | Transforming Visual Scene Graphs to Image Captions · ACL (1) 2023 |
Computer vision › Segmentation and scene understanding › scene graph
visual scene graph |
0.7 | 1 | 2023 | Transforming Visual Scene Graphs to Image Captions · ACL (1) 2023 |
Machine learning › Efficient and distributed learning
inference efficiency |
0.2 | 1 | 2024 | LIVE: Learnable In-Context Vector for Visual Question Answering · NeurIPS 2024 |
Computer vision › Vision and language
caption generation |
0.2 | 1 | 2023 | Transforming Visual Scene Graphs to Image Captions · ACL (1) 2023 |
Methods — techniques the papers use, named apart from their topics
large vision-language model · 0.9in-context learning · 0.9coreset selection · 0.9retrieval · 0.8learnable in-context vector · 0.8knowledge distillation · 0.8demonstration manipulation · 0.8transformer · 0.7scene graph transformation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Multimodal In-Context Learning for Image Classification through Coreset OptimizationabstractIn-context learning (ICL) enables Large Vision-Language Models (LVLMs) to adapt to new tasks without parameter updates, using a few demonstrations from a large support set. However, selecting informative demonstrations leads to high computational and memory costs. While some methods explore selecting a small and representative coreset in the text classification, evaluating all support set samples remains costly, and discarded samples lead to unnecessary information loss. These methods may also be less effective for image classification due to differences in feature spaces. Given these limitations, we propose Key-based Coreset Optimization (KeCO), a novel framework that leverages untapped data to construct a compact and informative coreset. We introduce visual features as keys within the coreset, which serve as the anchor for identifying samples to be updated through different selection strategies. By leveraging untapped samples from the support set, we update the keys of selected coreset samples, enabling the randomly initialized coreset to evolve into a more informative coreset under low computational cost. Through extensive experiments on coarse-grained and fine-grained image classification benchmarks, we demonstrate that KeCO effectively enhances ICL performance for image classification task, achieving an average improvement of more than 20%. Notably, we evaluate KeCO under a simulated online scenario, and the strong performance in this scenario highlights the practical value of our framework for resource-constrained real-world scenarios. Huiyi Chen, Jiawei Peng 0001, Kaihua Tang, Xin Geng 0001, Xu Yang 0021 |
ACM Multimedia | 2 |
| 2024 | How to Configure Good In-Context Sequence for Visual Question AnsweringabstractInspired by the success of Large Language Models in dealing with new tasks via In-Context Learning (ICL) in NLP, researchers have also developed Large Vision-Language Models (LVLMs) with ICL capabilities. However, when implementing ICL using these LVLMs, researchers usually resort to the simplest way like random sampling to configure the in-context sequence, thus leading to sub-optimal results. To enhance the ICL performance, in this study, we use Visual Question Answering (VQA) as case study to explore diverse in-context configurations to find the powerful ones. Additionally, through observing the changes of the LVLM outputs by altering the in-context sequence, we gain insights into the inner properties of LVLMs, improving our understanding of them. Specifically, to explore in-context configurations, we design diverse retrieval methods and employ different strategies to manipulate the retrieved demonstrations. Through exhaustive experiments on three VQA datasets: VQAv2, VizWiz, and OK-VQA, we uncover three important inner properties of the applied LVLM and demonstrate which strategies can consistently improve the ICL VQA performance. Our code is provided in: https://github.com/GaryJiajia/OFv2_ICL_VQA. Jiawei Peng 0001, Huiyi Chen, Chongyang Gao, Xu Yang 0021 |
CVPR | 2 |
| 2024 | LIVE: Learnable In-Context Vector for Visual Question AnsweringabstractAs language models continue to scale, Large Language Models (LLMs) have exhibited emerging capabilities in In-Context Learning (ICL), enabling them to solve language tasks by prefixing a few in-context demonstrations (ICDs) as context. Inspired by these advancements, researchers have extended these techniques to develop Large Multimodal Models (LMMs) with ICL capabilities. However, applying ICL usually faces two major challenges: 1) using more ICDs will largely increase the inference time and 2) the performance is sensitive to the selection of ICDs. These challenges are further exacerbated in LMMs due to the integration of multiple data types and the combinational complexity of multimodal ICDs. Recently, to address these challenges, some NLP studies introduce non-learnable In-Context Vectors (ICVs) which extract useful task information from ICDs into a single vector and then insert it into the LLM to help solve the corresponding task. However, although useful in simple NLP tasks, these non-learnable methods fail to handle complex multimodal tasks like Visual Question Answering (VQA). In this study, we propose \underline{\textbf{L}}earnable \underline{\textbf{I}}n-Context \underline{\textbf{Ve}}ctor (LIVE) to distill essential task information from demonstrations, improving ICL performance in LMMs. Experiments show that LIVE can significantly reduce computational costs while enhancing accuracy in VQA tasks compared to traditional ICL and other non-learnable ICV methods. Yingzhe Peng, Chenduo Hao, Xinting Hu, Jiawei Peng 0001, Xin Geng 0001, Xu Yang 0021 |
NeurIPS | 4 |
| 2023 | Transforming Visual Scene Graphs to Image CaptionsabstractXu Yang, Jiawei Peng, Zihua Wang, Haiyang Xu, Qinghao Ye, Chenliang Li, Songfang Huang, Fei Huang, Zhangzikang Li, Yu Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Xu Yang 0021, Jiawei Peng 0001, Zihua Wang, Haiyang Xu 0001, Qinghao Ye, Chenliang Li 0003, Songfang Huang, Fei Huang 0002, Zhangzikang Li, Yu Zhang 0004 |
ACL (1) | 2 |