Chloe Gu

dblp:439/3945 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Planning, search and constraint satisfaction · 77% Vision and language · 23%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › agent planning
embodied planning
1.012026
On-policy Reinforcement Fine-tuning with Offline reward for Multi-step Embodied Planning · ACL (1) 2026
Computer vision › Vision and language
vision-language model
0.312026
On-policy Reinforcement Fine-tuning with Offline reward for Multi-step Embodied Planning · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

reinforcement fine-tuning · 1.0offline reward · 1.0
YearPublicationVenuePosition
2026 On-policy Reinforcement Fine-tuning with Offline reward for Multi-step Embodied Planning
abstract
Embodied planning requires agents to make coherent multi-step decisions based on dynamic visual observations and verbal goals.While recent vision-language models (VLMs) excel at static perception tasks, they struggle in interactive environments.Reinforcement learning (RL) offers a natural way to address this limitation, yet online RL approaches suffer from costly interaction and sparse rewards in embodied settings.This paper introduces ORBIT, an On-policy Reinforcement finetuning (RFT) framework with offline rewards for EmBodIed Task Planning, that preserves the generalization benefits of RFT while addressing the challenges of costly interaction and sparse rewards, supported by solid theoretical guarantees.Our approach is evaluated on EmbodiedBench, a recent benchmark for interactive embodied tasks, covering both in-domain and out-of-domain scenarios.Experimental results show that ORBIT achieves SOTA performance on EB-ALFRED, outperforming all closed-source and online-RL-based methods, while being substantially more efficient in training speed and computational cost, remaining robust to sub-optimal expert trajectories, and exhibiting strong generalization to unseen environments.We released all code and data at https://github.com/mail-taii/Reinforced- Reasoning-for
Chloe Gu, Guanbo Wang, Wenhao Li 0001, Bo Jin 0003
ACL (1)3