VLDB 2026 Research / reviewers in the wild / expert
Yanqi Dai
dblp:322/2311
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Learning paradigms · 41% Vision and language · 38% Optimization for machine learning · 20% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning paradigms
multi-task learning |
1.0 | 1 | 2026 | Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task Difficulty · WWW 2026 |
Machine learning › Optimization for machine learning › multi-task optimization
task balancing |
1.0 | 1 | 2026 | Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task Difficulty · WWW 2026 |
Machine learning › Learning paradigms › multi-task learning
task weighting |
1.0 | 1 | 2026 | Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task Difficulty · WWW 2026 |
Computer vision › Vision and language › vision-language model › multimodal large language model
visual instruction tuning |
1.0 | 1 | 2026 | Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task Difficulty · WWW 2026 |
Computer vision › Vision and language
multimodal dialogue |
0.9 | 1 | 2025 | MMRole: A Comprehensive Framework for Developing and Evaluating Multimodal Role-Playing Agents · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
reward model · 1.7multimodal large language model · 1.7validation-performance-based task balancing · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Task Balancing for Visual Instruction Tuning via Inter-Task Contribution and Intra-Task DifficultyabstractVisual instruction tuning is a key training stage of large multimodal models. However, when learning multiple visual tasks simultaneously, this approach often results in suboptimal and imbalanced overall performance due to latent knowledge conflicts across tasks. To mitigate this issue, we propose a novel Adaptive Task Balancing approach tailored for visual instruction tuning (VisATB). Specifically, we measure two critical dimensions for visual task balancing based on validation performance: (1) Inter-Task Contribution, the mechanism where learning one task enhances the performance on others owing to shared knowledge across tasks, and (2) Intra-Task Difficulty, which denotes the inherent learning difficulty of a single task. Furthermore, we propose prioritizing three categories of tasks with greater weight: those that offer substantial contributions to others, those that receive minimal contributions from others, and those that present high learning difficulties. Among these three task weighting strategies, the first and third focus on improving overall performance, and the second targets the mitigation of performance imbalance. Extensive experiments on three benchmarks demonstrate that our VisATB approach consistently achieves superior and more balanced overall performance in visual instruction tuning. The data, code, and models are available at https://github.com/YanqiDai/VisATB. Yanqi Dai, Zebin You, Dong Jing, Xiangxiang Chu, Zhiwu Lu 0001 |
WWW | 1 |
| 2025 | MMRole: A Comprehensive Framework for Developing and Evaluating Multimodal Role-Playing AgentsabstractRecently, Role-Playing Agents (RPAs) have garnered increasing attention for their potential to deliver emotional value and facilitate sociological research.
However, existing studies are primarily confined to the textual modality, unable to simulate humans' multimodal perceptual capabilities.
To bridge this gap, we introduce the concept of Multimodal Role-Playing Agents (MRPAs), and propose a comprehensive framework, MMRole, for their development and evaluation, which comprises a personalized multimodal dataset and a robust evaluation approach.
Specifically, we construct a large-scale, high-quality dataset, MMRole-Data, consisting of 85 characters, 11K images, and 14K single or multi-turn dialogues.
Additionally, we present a robust evaluation approach, MMRole-Eval, encompassing eight metrics across three dimensions, where a reward model is designed to score MRPAs with the constructed ground-truth data for comparison.
Moreover, we develop the first specialized MRPA, MMRole-Agent.
Extensive evaluation results demonstrate the improved performance of MMRole-Agent and highlight the primary challenges in developing MRPAs, emphasizing the need for enhanced multimodal understanding and role-playing consistency.
The data, code, and models are all available at https://github.com/YanqiDai/MMRole. Yanqi Dai, Huanran Hu 0001, Lei Wang 0198, Shengjie Jin, Xu Chen 0017, Zhiwu Lu 0001 |
ICLR | 1 |
| 2025 | CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual WorldsabstractLei Wang, Jianxun Lian, Yi Huang, Yanqi Dai, Haoxuan Li, Xu Chen, Xing Xie, Ji-Rong Wen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Lei Wang 0198, Jianxun Lian, Yanqi Dai, Haoxuan Li 0001, Xu Chen 0017, Xing Xie 0001, Ji-Rong Wen |
NAACL (Long Papers) | 4 |
| 2024 | VEMO: A Versatile Elastic Multi-modal Model for Search-Oriented Multi-task Learning
Nanyi Fei, Hao Jiang 0022, Haoyu Lu, Jinqiang Long, Yanqi Dai, Tuo Fan, Zhao Cao, Zhiwu Lu 0001 |
ECIR (1) | 5 |
| 2023 | Improvable Gap Balancing for Multi-Task LearningabstractIn multi-task learning (MTL), gradient balancing has recently attracted more research interest than loss balancing since it often leads to better performance. However, loss balancing is much more efficient than gradient balancing, and thus it is still worth further exploration in MTL. Note that prior studies typically ignore that there exist varying improvable gaps across multiple tasks, where the improvable gap per task is defined as the distance between the current training progress and desired final training progress. Therefore, after loss balancing, the performance imbalance still arises in many cases. In this paper, following the loss balancing framework, we propose two novel improvable gap balancing (IGB) algorithms for MTL: one takes a simple heuristic, and the other (for the first time) deploys deep reinforcement learning for MTL. Particularly, instead of directly balancing the losses in MTL, both algorithms choose to dynamically assign task weights for improvable gap balancing. Moreover, we combine IGB and gradient balancing to show the complementarity between the two types of algorithms. Extensive experiments on two benchmark datasets demonstrate that our IGB algorithms lead to the best results in MTL via loss balancing and achieve further improvements when combined with gradient balancing. Code is available at https://github.com/YanqiDai/IGB4MTL. Yanqi Dai, Nanyi Fei, Zhiwu Lu 0001 |
UAI | 1 |
| 2021 | TRGAN: Text to Image Generation Through Optimizing Initial Image
Liang Zhao 0005, Pingda Huang, Zhikui Chen, Yanqi Dai |
ICONIP (5) | 5 |