VLDB 2026 Research / reviewers in the wild / expert
Yuzhen Xiao
dblp:367/3390
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0003-4253-525XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 64% Question answering and dialogue systems · 13% Efficient and distributed learning · 13% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
retrieval-augmented generation |
1.7 | 2 | 2025 | Parenting: Optimizing Knowledge Selection of Retrieval-Augmented Language Models with Parameter Decoupling and Tailored Tuning · ACL (1) 2025 KnowPO: Knowledge-Aware Preference Optimization for Controllable Knowledge Selection in Retrieval-Augmented Language Models · AAAI 2025 |
Natural language and speech › Language models and text generation › retrieval-augmented generation
knowledge conflict |
0.9 | 1 | 2025 | KnowPO: Knowledge-Aware Preference Optimization for Controllable Knowledge Selection in Retrieval-Augmented Language Models · AAAI 2025 |
Natural language and speech › Question answering and dialogue systems › knowledge-grounded dialogue
knowledge selection |
0.9 | 1 | 2025 | Parenting: Optimizing Knowledge Selection of Retrieval-Augmented Language Models with Parameter Decoupling and Tailored Tuning · ACL (1) 2025 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.9 | 1 | 2025 | Parenting: Optimizing Knowledge Selection of Retrieval-Augmented Language Models with Parameter Decoupling and Tailored Tuning · ACL (1) 2025 |
Natural language and speech › Language models and text generation
preference optimization |
0.9 | 1 | 2025 | KnowPO: Knowledge-Aware Preference Optimization for Controllable Knowledge Selection in Retrieval-Augmented Language Models · AAAI 2025 |
Natural language and speech › Language models and text generation
prompt tuning |
0.9 | 1 | 2025 | Parenting: Optimizing Knowledge Selection of Retrieval-Augmented Language Models with Parameter Decoupling and Tailored Tuning · ACL (1) 2025 |
Machine learning › Deep learning architectures and training
data augmentation |
0.8 | 1 | 2024 | ProtoMix: Augmenting Health Status Representation Learning via Prototype-based Mixup · KDD 2024 |
Medical and health informatics
electronic health records |
0.2 | 1 | 2024 | ProtoMix: Augmenting Health Status Representation Learning via Prototype-based Mixup · KDD 2024 |
Medical and health informatics
healthcare prediction |
0.2 | 1 | 2024 | ProtoMix: Augmenting Health Status Representation Learning via Prototype-based Mixup · KDD 2024 |
Methods — techniques the papers use, named apart from their topics
mixup · 1.5generative adversarial network · 1.5diffusion model · 1.5preference optimization · 0.9parameter decoupling · 0.9instruction tuning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | KnowPO: Knowledge-Aware Preference Optimization for Controllable Knowledge Selection in Retrieval-Augmented Language ModelsabstractBy integrating external knowledge, Retrieval-Augmented Generation (RAG) has become an effective strategy for mitigating the hallucination problems that large language models (LLMs) encounter when dealing with knowledge-intensive tasks. However, in the process of integrating external non-parametric supporting evidence with internal parametric knowledge, inevitable knowledge conflicts may arise, leading to confusion in the model's responses. To enhance the knowledge selection of LLMs in various contexts, some research has focused on refining their behavior patterns through instruction-tuning. Nonetheless, due to the absence of explicit negative signals and comparative objectives, models fine-tuned in this manner may still exhibit undesirable behaviors such as contextual ignorance and contextual overinclusion. To this end, we propose a Knowledge-aware Preference Optimization strategy, dubbed KnowPO, aimed at achieving adaptive knowledge selection based on contextual relevance in real retrieval scenarios. Concretely, we proposed a general paradigm for constructing knowledge conflict datasets, which comprehensively cover various error types and learn how to avoid these negative signals through preference optimization methods. Simultaneously, we proposed a rewriting strategy and data ratio optimization strategy to address preference imbalances. Experimental results show that KnowPO outperforms previous methods for handling knowledge conflicts by over 37%, while also exhibiting robust generalization across various out-of-distribution datasets. Ruizhe Zhang 0013, Yongxin Xu, Yuzhen Xiao, Runchuan Zhu, Xinke Jiang, Junfeng Zhao 0001, Yasha Wang |
AAAI | 3 |
| 2025 | Parenting: Optimizing Knowledge Selection of Retrieval-Augmented Language Models with Parameter Decoupling and Tailored TuningabstractYongxin Xu, Ruizhe Zhang, Xinke Jiang, Yujie Feng, Yuzhen Xiao, Xinyu Ma, Runchuan Zhu, Xu Chu, Junfeng Zhao, Yasha Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yongxin Xu, Ruizhe Zhang 0013, Xinke Jiang, Yuzhen Xiao, Runchuan Zhu, Junfeng Zhao 0001, Yasha Wang |
ACL (1) | 5 |
| 2024 | ProtoMix: Augmenting Health Status Representation Learning via Prototype-based MixupabstractWith the widespread adoption of electronic health records (EHR) data, deep learning techniques have been broadly utilized for various health prediction tasks. Nevertheless, the labeled data scarcity issue restricts the prediction power of these deep models. To enhance the generalization capability of deep learning models when faced with such situations, a common trend is to train generative adversarial networks (GANs) or diffusion models for data augmentation. However, due to limitations in sample size and potential label imbalance issues, these methods are prone to mode collapse problems. This results in the generation of new samples that fail to preserve the subtype structure within EHR data, thereby limiting their practicality in health prediction tasks that generally require detailed patient phenotyping. Aiming at the above problems, we propose a Prototype-based Mixup method, dubbed ProtoMix, which combines prior knowledge of intrinsic data features from subtype centroids (i.e., prototypes) to guide the synthesis of new samples. Specifically, ProtoMix employs a prototype-guided mixup training task to shift the decision boundary away from the subtypes. Then, ProtoMix optimizes the sampling weights in different areas of the data manifold via a prototype-guided mixup sampling strategy. Throughout the training process, ProtoMix dynamically expands the training distribution using an adaptive mixing coefficient computation method. Experimental evaluations on three real-world datasets demonstrate the efficacy of ProtoMix. Yongxin Xu, Xinke Jiang, Yuzhen Xiao, Chaohe Zhang, Hongxin Ding, Junfeng Zhao 0001, Yasha Wang |
KDD | 4 |