Ruiyao Xu

dblp:343/3377 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 67% Reinforcement learning · 33%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
alignment
1.012026
CoAct: Co-Active LLM Preference Learning with Human-AI Synergy · ACL (1) 2026
Natural language and speech › Language models and text generation › alignment
preference alignment
1.012026
CoAct: Co-Active LLM Preference Learning with Human-AI Synergy · ACL (1) 2026
Machine learning › Reinforcement learning
preference learning
1.012026
CoAct: Co-Active LLM Preference Learning with Human-AI Synergy · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

self-rewarding · 1.0self-consistency · 1.0active learning · 1.0
YearPublicationVenuePosition
2026 CoAct: Co-Active LLM Preference Learning with Human-AI Synergy
abstract
Learning from preference-based feedback has become an effective approach for aligning LLMs across diverse tasks.However, highquality human-annotated preference data remains expensive and scarce.Existing methods address this challenge through either selfrewarding, which scales by using purely AIgenerated labels but risks unreliability, or active learning, which ensures quality through oracle annotation but cannot fully leverage unlabeled data.In this paper, we present COACT, a novel framework that synergistically combines selfrewarding and active learning through strategic human-AI collaboration.COACT leverages self-consistency to identify both reliable selflabeled data and samples that are requiring oracle verification.Additionally, oracle feedback guides the model to generate new instructions within its solvable capability.Evaluated on three reasoning benchmarks across two model families, COACT achieves average improvements of +13.25% on GSM8K, +8.19% on MATH, and +13.16% on WebInstruct, consistently outperforming all baselines.1
Ruiyao Xu, Mihir Parmar, Tiankai Yang 0001, Zhengyu Hu, Yue Zhao 0016, Kaize Ding
ACL (1)1
2022 Population-Based Hierarchical Non-Negative Matrix Factorization for Survey Data
abstract
Motivated by the problem of identifying potential hierarchical population structure on modern survey data containing a wide range of complex data types, we introduce population-based hierarchical non-negative matrix factorization (PHNMF). PHNMF is a variant of hierarchical non-negative matrix factorization based on feature similarity. As such, it enables an automatic and interpretable approach for identifying and understanding hierarchical structure in a data matrix constructed from a wide range of data types. Our numerical experiments on synthetic and real survey data demonstrate that PHNMF can recover latent hierarchical population structure in complex data with high accuracy. Moreover, the recovered subpopulation structure is meaningful and can be useful for improving downstream inference.
Xiaofu Ding, Olivia McGough, Chenxin Shen, Annie Ulichney, Ruiyao Xu, William Swartworth, Jocelyn T. Chi, Deanna Needell
BDCAT6