VLDB 2026 Research / reviewers in the wild / expert
Ruiyao Xu
dblp:343/3377
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Language models and text generation · 67% Reinforcement learning · 33% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
1.0 | 1 | 2026 | CoAct: Co-Active LLM Preference Learning with Human-AI Synergy · ACL (1) 2026 |
Natural language and speech › Language models and text generation › alignment
preference alignment |
1.0 | 1 | 2026 | CoAct: Co-Active LLM Preference Learning with Human-AI Synergy · ACL (1) 2026 |
Machine learning › Reinforcement learning
preference learning |
1.0 | 1 | 2026 | CoAct: Co-Active LLM Preference Learning with Human-AI Synergy · ACL (1) 2026 |
Methods — techniques the papers use, named apart from their topics
self-rewarding · 1.0self-consistency · 1.0active learning · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoAct: Co-Active LLM Preference Learning with Human-AI SynergyabstractLearning from preference-based feedback has become an effective approach for aligning LLMs across diverse tasks.However, highquality human-annotated preference data remains expensive and scarce.Existing methods address this challenge through either selfrewarding, which scales by using purely AIgenerated labels but risks unreliability, or active learning, which ensures quality through oracle annotation but cannot fully leverage unlabeled data.In this paper, we present COACT, a novel framework that synergistically combines selfrewarding and active learning through strategic human-AI collaboration.COACT leverages self-consistency to identify both reliable selflabeled data and samples that are requiring oracle verification.Additionally, oracle feedback guides the model to generate new instructions within its solvable capability.Evaluated on three reasoning benchmarks across two model families, COACT achieves average improvements of +13.25% on GSM8K, +8.19% on MATH, and +13.16% on WebInstruct, consistently outperforming all baselines.1 Ruiyao Xu, Mihir Parmar, Tiankai Yang 0001, Zhengyu Hu, Yue Zhao 0016, Kaize Ding |
ACL (1) | 1 |
| 2022 | Population-Based Hierarchical Non-Negative Matrix Factorization for Survey DataabstractMotivated by the problem of identifying potential hierarchical population structure on modern survey data containing a wide range of complex data types, we introduce population-based hierarchical non-negative matrix factorization (PHNMF). PHNMF is a variant of hierarchical non-negative matrix factorization based on feature similarity. As such, it enables an automatic and interpretable approach for identifying and understanding hierarchical structure in a data matrix constructed from a wide range of data types. Our numerical experiments on synthetic and real survey data demonstrate that PHNMF can recover latent hierarchical population structure in complex data with high accuracy. Moreover, the recovered subpopulation structure is meaningful and can be useful for improving downstream inference. Xiaofu Ding, Olivia McGough, Chenxin Shen, Annie Ulichney, Ruiyao Xu, William Swartworth, Jocelyn T. Chi, Deanna Needell |
BDCAT | 6 |