Xiao Yu 0011

dblp:89/2407-11 · DBLP profile ↗
← Back
7ranked-venue papers
6as first author
7since 2021 · last 2025
0009-0001-0924-6432ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 30% Question answering and dialogue systems · 24% Planning, search and constraint satisfaction · 22%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
1.522025
ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning · ICLR 2025
Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy Planning · EMNLP 2023
Natural language and speech › Question answering and dialogue systems
task-oriented dialogue
1.322023
KRLS: Improving End-to-End Response Generation in Task Oriented Dialog with Reinforced Keywords Learning · EMNLP 2023
Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy Planning · EMNLP 2023
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
agent planning
0.912025
ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning · ICLR 2025
Natural language and speech › Language models and text generation › test-time scaling
inference-time search
0.912025
ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning · ICLR 2025
Machine learning › Representation and self-supervised learning
self-learning
0.912025
ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning · ICLR 2025
Natural language and speech › Language models and text generation
alignment
0.812024
LIONs: An Empirically Optimized Approach to Align Language Models · EMNLP 2024
Machine learning › Reinforcement learning
preference learning
0.812024
LIONs: An Empirically Optimized Approach to Align Language Models · EMNLP 2024
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning
0.812024
LIONs: An Empirically Optimized Approach to Align Language Models · EMNLP 2024
Natural language and speech › Question answering and dialogue systems › dialogue management
dialogue policy planning
0.712023
Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy Planning · EMNLP 2023
Natural language and speech › Question answering and dialogue systems › dialogue generation
dialogue response generation
0.712023
KRLS: Improving End-to-End Response Generation in Task Oriented Dialog with Reinforced Keywords Learning · EMNLP 2023
Machine learning › Reinforcement learning
offline reinforcement learning
0.712023
KRLS: Improving End-to-End Response Generation in Task Oriented Dialog with Reinforced Keywords Learning · EMNLP 2023
Computer vision › Vision and language
vision-language model
0.312025
ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning · ICLR 2025
Natural language and speech › Language models and text generation › LLM agents
web agents
0.312025
ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning · ICLR 2025
Natural language and speech › Language models and text generation
instruction following
0.212024
LIONs: An Empirically Optimized Approach to Align Language Models · EMNLP 2024
Natural language and speech › Language models and text generation
prompting
0.212023
Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy Planning · EMNLP 2023

Methods — techniques the papers use, named apart from their topics

monte carlo tree search · 1.5self-learning · 0.9multi-agent debate · 0.9sequence packing · 0.8online preference learning · 0.8loss masking · 0.8direct preference optimization · 0.8user simulation · 0.7large language model prompting · 0.7keyword reward · 0.7
YearPublicationVenuePosition
2025 ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning
abstract
Autonomous agents have demonstrated significant potential in automating complex multistep decision-making tasks. However, even state-of-the-art vision-language models (VLMs), such as GPT-4o, still fall short of human-level performance, particularly in intricate web environments and long-horizon planning tasks. To address these limitations, we introduce Reflective Monte Carlo Tree Search (R-MCTS), a novel test-time algorithm designed to enhance the ability of AI agents, e.g., powered by GPT-4o, to explore decision space on the fly. R-MCTS extends traditional MCTS by 1) incorporating contrastive reflection, allowing agents to learn from past interactions and dynamically improve their search efficiency; and 2) using multi-agent debate to provide reliable state evaluation. Moreover, we improve the agent's performance by fine-tuning GPT-4o through self-learning, using R-MCTS generated tree traversals without any human-provided labels. On the challenging VisualWebArena benchmark, our GPT-4o-based R-MCTS agent achieves a 6% to 30% relative improvement across various tasks compared to the previous state-of-the-art. Additionally, we show that the knowledge gained from test-time search can be effectively transferred back to GPT-4o via fine-tuning. The fine-tuned GPT-4o matches 97\% of R-MCTS's performance while reducing compute usage by a factor of four at test time. Furthermore, qualitative results reveal that the fine-tuned GPT-4o model demonstrates the ability to explore the environment, evaluate a state, and backtrack to viable ones when it detects that the current state cannot lead to success. Moreover, our work demonstrates the compute scaling properties in both training - data collection with R-MCTS - and testing time. These results suggest a promising research direction to enhance VLMs' reasoning and planning capabilities for agentic applications via test-time search and self-learning.
Xiao Yu 0011, Baolin Peng, Vineeth Vajipey, Hao Cheng 0002, Michel Galley, Jianfeng Gao 0001, Zhou Yu 0005
ICLR1
2024 LIONs: An Empirically Optimized Approach to Align Language Models
abstract
Alignment is a crucial step to enhance the instruction-following and conversational abilities of language models.Despite many recent work proposing new algorithms, datasets, and training pipelines, there is a lack of comprehensive studies measuring the impact of various design choices throughout the whole training process.We first conduct a rigorous analysis over a three-stage training pipeline consisting of supervised fine-tuning, offline preference learning, and online preference learning.We have found that using techniques like sequence packing, loss masking in SFT, increasing the preference dataset size in DPO, and online DPO training can significantly improve the performance of language models.We then train from Gemma-2b-base and LLama-3-8b-base, and find that our best models exceed the performance of the official instruct models tuned with closed-source data and algorithms.Our code and models can be found at https://github.
Xiao Yu 0011, Qingyang Wu, Yu Li 0013, Zhou Yu 0005
EMNLP1
2024 Teaching Language Models to Self-Improve through Interactive Demonstrations
abstract
Xiao Yu, Baolin Peng, Michel Galley, Jianfeng Gao, Zhou Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Xiao Yu 0011, Baolin Peng, Michel Galley, Jianfeng Gao 0001, Zhou Yu 0005
NAACL-HLT1
2024 ConFit: Improving Resume-Job Matching using Data Augmentation and Contrastive Learning
abstract
A reliable resume-job matching system helps a company find suitable candidates from a pool of resumes, and helps a job seeker find relevant jobs from a list of job posts. However, since job seekers apply only to a few jobs, interaction records in resume-job datasets are sparse. Different from many prior work that use complex modeling techniques, we tackle this sparsity problem using data augmentations and a simple contrastive learning approach. ConFit first formulates resume-job datasets as a sparse bipartite graph, and creates an augmented dataset by paraphrasing specific sections in a resume or a job post. Then, ConFit finetunes pre-trained encoders with contrastive learning to further increase training samples from B pairs per batch to <?TeX $\mathcal {O}(B^2)$?> Math 1 per batch. We evaluate ConFit on two real-world datasets and find it outperforms prior methods (including BM25 and OpenAI text-ada-002) by up to 19% and 31% absolute in nDCG@10 for ranking jobs and ranking resumes, respectively. We believe ConFit’s simple yet highly performant approach lays a strong foundation for future research in modeling person-job fit.1
Xiao Yu 0011, Jinzhong Zhang 0002, Zhou Yu 0005
RecSys1
2023 FastKASSIM: A Fast Tree Kernel-Based Syntactic Similarity Metric
abstract
Syntax is a fundamental component of language, yet few metrics have been employed to capture syntactic similarity or coherence at the utterance-and document-level.The existing standard document-level syntactic similarity metric is computationally expensive and performs inconsistently when faced with syntactically dissimilar documents.To address these challenges, we present FastKASSIM, a metric for utterance-and document-level syntactic similarity which pairs and averages the most similar constituency parse trees between a pair of documents based on tree kernels.FastKAS-SIM is more robust to syntactic dissimilarities and runs up to to 5.32 times faster than its predecessor over documents in the r/ChangeMyView corpus.FastKASSIM's improvements allow us to examine hypotheses in two settings with large documents.We find that syntactically similar arguments on r/ChangeMyView tend to be more persuasive, and that syntax is predictive of authorship attribution in the Australian High Court Judgment corpus. * denotes equal contribution.Utterance 1: When we hate, we always move away from the grace of God.When we become resentful and unforgiving, the world around us seems spiteful and meaningless.Utterance 2: How can you be skiing if you are already swimming?FastKASSIM Score: 0.219 CASSIM Score: 0.838 LSM Score: 0.623 Utterance 1: I like swimming because it is cool.Utterance 2: I love running because it is fun.
Maximillian Chen 0001, Caitlyn Chen, Xiao Yu 0011, Zhou Yu 0005
EACL3
2023 Prompt-Based Monte-Carlo Tree Search for Goal-oriented Dialogue Policy Planning
abstract
Planning for goal-oriented dialogue often requires simulating future dialogue interactions and estimating task progress.Many approaches thus consider training neural networks to perform look-ahead search algorithms such as A* search and Monte Carlo Tree Search (MCTS).However, this training often requires abundant annotated data, which creates challenges when faced with noisy annotations or low-resource settings.We introduce GDP-ZERO, an approach using Open-Loop MCTS to perform goal-oriented dialogue policy planning without any model training.GDP-ZERO prompts a large language model to act as a policy prior, value function, user simulator, and system model during the tree search.We evaluate GDP-ZERO on the goal-oriented task Persua-sionForGood, and find that its responses are preferred over ChatGPT up to 59.32% of the time, and are rated more persuasive than Chat-GPT during interactive evaluations 1 .
Xiao Yu 0011, Maximillian Chen 0001, Zhou Yu 0005
EMNLP1
2023 KRLS: Improving End-to-End Response Generation in Task Oriented Dialog with Reinforced Keywords Learning
abstract
In task-oriented dialogs (TOD), reinforcement learning (RL) algorithms train a model to directly optimize response for task-related metrics.However, RL needs to perform exploration, which can be time-consuming due to the slow auto-regressive sequence generation process.We investigate an approach to create a more efficient RL-based algorithm to improve TOD performance in an offline setting.First, we use a faster generation procedure that samples from independent next-word distributions after training the language model (LM) with supervised learning.We then introduce a finegrained reward function to help the model focus on learning key information in a dialog, by measuring the importance and semantic closeness of each generated token.Experiments on the MultiWoZ dataset show our new training algorithm, Keywords Reinforcement Learning with Next-word Sampling (KRLS), achieves state-of-the-art performance on the end-to-end response generation task, with a 15% training time reduction compared to a standard RL algorithm using auto-regressive generation 1 .
Xiao Yu 0011, Qingyang Wu, Kun Qian 0016, Zhou Yu 0005
EMNLP1