VLDB 2026 Research / reviewers in the wild / expert
Yuqian Jiang
dblp:14/10243
· DBLP profile ↗
9ranked-venue papers
4as first author
4since 2021 · last 2025
0000-0003-0091-6871ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 4 since 2021Systems, architecture and hardware · 5 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Reinforcement learning · 57% Planning, search and constraint satisfaction · 28% Robot manipulation · 15% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-robot interaction · 100% | |
| Theoretical computer science
1 paper |
Logic in computer science · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › markov decision process
average-reward reinforcement learning |
0.5 | 1 | 2021 | Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning Tasks · AAAI 2021 |
Machine learning › Reinforcement learning › reward design
reward shaping |
0.5 | 1 | 2021 | Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning Tasks · AAAI 2021 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
task planning |
0.5 | 1 | 2021 | Planning and Reinforcement Learning for General-Purpose Service Robots · IJCAI 2021 |
Human-robot interaction › robot communication
interactive clarification |
0.4 | 1 | 2019 | Improving Grounded Natural Language Understanding through Human-Robot Dialog · ICRA 2019 |
Human-robot interaction
natural language understanding |
0.4 | 1 | 2019 | Improving Grounded Natural Language Understanding through Human-Robot Dialog · ICRA 2019 |
Robotics › Robot manipulation
service robot |
0.1 | 1 | 2021 | Planning and Reinforcement Learning for General-Purpose Service Robots · IJCAI 2021 |
Logic in computer science
temporal logic |
0.1 | 1 | 2021 | Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning Tasks · AAAI 2021 |
Robotics › Robot manipulation › grasping
pick-and-place |
0.1 | 1 | 2019 | Improving Grounded Natural Language Understanding through Human-Robot Dialog · ICRA 2019 |
Methods — techniques the papers use, named apart from their topics
temporal logic translation · 1.0reward shaping · 1.0end-to-end learning · 0.8symbolic planning · 0.5reinforcement learning · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | L3M+P: Lifelong Planning with Large Language ModelsabstractBy combining classical planning methods with large language models (LLMs), recent research such as LLM+P has enabled agents to plan for general tasks given in natural language. However, scaling these methods to general-purpose service robots remains challenging: (1) classical planning algorithms generally require a detailed and consistent specification of the environment, which is not always readily available; and (2) existing frameworks mainly focus on isolated planning tasks, whereas robots are often meant to serve in long-term continuous deployments, and therefore must maintain a dynamic memory of the environment which can be updated with multi-modal inputs and extracted as planning knowledge for future tasks. To address these two issues, this paper introduces L3M+P (Lifelong LLM+P), a framework that uses an external knowledge graph as a representation of the world state. The graph can be updated from multiple sources of information, including sensory input and natural language interactions with humans. L3M+P enforces rules for the expected format of the absolute world state graph to maintain consistency between graph updates. At planning time, given a natural language description of a task, L3M+P retrieves context from the knowledge graph and generates a problem definition for classical planners. Evaluated on household robot simulators and on a real-world service robot, L3M+P achieves significant improvement over baseline methods both on accurately registering natural language state changes and on correctly generating plans, thanks to the knowledge graph retrieval and verification. Krish Agarwal, Yuqian Jiang, Jiaheng Hu, Bo Liu 0042, Peter Stone 0001 |
IROS | 2 |
| 2023 | Symbolic State Space Optimization for Long Horizon Mobile Manipulation PlanningabstractIn existing task and motion planning (TAMP) research, it is a common assumption that experts manually specify the state space for task-level planning. A well-developed state space enables the desirable distribution of limited computational resources between task planning and motion planning. However, developing such task-level state spaces can be non-trivial in practice. In this paper, we consider a long horizon mobile manipulation domain including repeated navigation and manipulation. We propose Symbolic State Space Optimization (S3O) for computing a set of abstracted locations and their 2D geometric groundings for generating task-motion plans in such domains. Our approach has been extensively evaluated in simulation and demonstrated on a real mobile manipulator working on clearing up dining tables. Results show the superiority of the proposed method over TAMP baselines in task completion rate and execution time. Xiaohan Zhang 0002, Yan Ding 0002, Yuqian Jiang, Yuke Zhu, Peter Stone 0001, Shiqi Zhang 0001 |
IROS | 4 |
| 2021 | Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning TasksabstractIn continuing tasks, average-reward reinforcement learning may be a more appropriate problem formulation than the more common discounted reward formulation. As usual, learning an optimal policy in this setting typically requires a large amount of training experiences. Reward shaping is a common approach for incorporating domain knowledge into reinforcement learning in order to speed up convergence to an optimal policy. However, to the best of our knowledge, the theoretical properties of reward shaping have thus far only been established in the discounted setting. This paper presents the first reward shaping framework for average-reward learning and proves that, under standard assumptions, the optimal policy under the original reward function can be recovered. In order to avoid the need for manual construction of the shaping function, we introduce a method for utilizing domain knowledge expressed as a temporal logic formula. The formula is automatically translated to a shaping function that provides additional reward throughout the learning process. We evaluate the proposed method on three continuing tasks. In all cases, shaping speeds up the average-reward learning rate without any reduction in the performance of the learned policy compared to relevant baselines. Yuqian Jiang, Suda Bharadwaj, Bo Wu 0005, Rishi Shah, Ufuk Topcu, Peter Stone 0001 |
AAAI | 1 |
| 2021 | Planning and Reinforcement Learning for General-Purpose Service RobotsabstractDespite recent progress in AI and robotics research, especially learned robot skills, there remain significant challenges in building robust, scalable, and general-purpose systems for service robots. This Ph.D. research aims to combine symbolic planning and reinforcement learning to reason about high-level robot tasks and adapt to the real world. We will introduce task planning algorithms that adapt to the environment and other agents, as well as reinforcement learning methods that are practical for service robot systems. Taken together, this work will make a significant step towards creating general-purpose service robots. Yuqian Jiang |
IJCAI | 1 |
| 2020 | Deep R-Learning for Continual Area SweepingabstractCoverage path planning is a well-studied problem in robotics in which a robot must plan a path that passes through every point in a given area repeatedly, usually with a uniform frequency. To address the scenario in which some points need to be visited more frequently than others, this problem has been extended to non-uniform coverage planning. This paper considers the variant of non-uniform coverage in which the robot does not know the distribution of relevant events beforehand and must nevertheless learn to maximize the rate of detecting events of interest. This continual area sweeping problem has been previously formalized in a way that makes strong assumptions about the environment, and to date only a greedy approach has been proposed. We generalize the continual area sweeping formulation to include fewer environmental constraints, and propose a novel approach based on reinforcement learning in a Semi-Markov Decision Process. This approach is evaluated in an abstract simulation and in a high fidelity Gazebo simulation. These evaluations show significant improvement upon the existing approach in general settings, which is especially relevant in the growing area of service robotics. We also present a video demonstration on a real service robot. Rishi Shah, Yuqian Jiang, Justin W. Hart, Peter Stone 0001 |
IROS | 2 |
| 2020 | Jointly Improving Parsing and Perception for Natural Language Commands through Human-Robot DialogabstractIn this work, we present methods for using human-robot dialog to improve language understanding for a mobile robot agent. The agent parses natural language to underlying semantic meanings and uses robotic sensors to create multi-modal models of perceptual concepts like red and heavy. The agent can be used for showing navigation routes, delivering objects to people, and relocating objects from one location to another. We use dialog clari_cation questions both to understand commands and to generate additional parsing training data. The agent employs opportunistic active learning to select questions about how words relate to objects, improving its understanding of perceptual concepts. We evaluated this agent on Amazon Mechanical Turk. After training on data induced from conversations, the agent reduced the number of dialog questions it asked while receiving higher usability ratings. Additionally, we demonstrated the agent on a robotic platform, where it learned new perceptual concepts on the y while completing a real-world task. Jesse Thomason, Aishwarya Padmakumar, Jivko Sinapov, Nick Walker 0001, Yuqian Jiang, Harel Yedidsion, Justin W. Hart, Peter Stone 0001, Raymond J. Mooney |
J. Artif. Intell. Res. | 5 |
| 2019 | Improving Grounded Natural Language Understanding through Human-Robot DialogabstractNatural language understanding for robotics can require substantial domain- and platform-specific engineering. For example, for mobile robots to pick-and-place objects in an environment to satisfy human commands, we can specify the language humans use to issue such commands, and connect concept words like red can to physical object properties. One way to alleviate this engineering for a new domain is to enable robots in human environments to adapt dynamically-continually learning new language constructions and perceptual concepts. In this work, we present an end-to-end pipeline for translating natural language commands to discrete robot actions, and use clarification dialogs to jointly improve language parsing and concept grounding. We train and evaluate this agent in a virtual setting on Amazon Mechanical Turk, and we transfer the learned agent to a physical robot platform to demonstrate it in the real world. Jesse Thomason, Aishwarya Padmakumar, Jivko Sinapov, Nick Walker 0001, Yuqian Jiang, Harel Yedidsion, Justin W. Hart, Peter Stone 0001, Raymond J. Mooney |
ICRA | 5 |
| 2019 | Task-Motion Planning with Reinforcement Learning for Adaptable Mobile Service RobotsabstractTask-motion planning (TMP) addresses the problem of efficiently generating executable and low-cost task plans in a discrete space such that the (initially unknown) action costs are determined by motion plans in a corresponding continuous space. A task-motion plan for a mobile service robot that behaves in a highly dynamic domain can be sensitive to domain uncertainty and changes, leading to suboptimal behaviors or execution failures. In this paper, we propose a novel framework, TMP-RL, which is an integration of TMP and reinforcement learning (RL), to solve the problem of robust TMP in dynamic and uncertain domains. The robot first generates a low-cost, feasible task-motion plan by iteratively planning in the discrete space and updating relevant action costs evaluated by the motion planner in continuous space. During execution, the robot learns via model-free RL to further improve its task-motion plans. RL enables adaptability to the current domain, but can be costly with regards to experience; using TMP, which does not rely on experience, can jump-start the learning process before executing in the real world. TMP-RL is evaluated in a mobile service robot domain where the robot navigates in an office area, showing significantly improved adaptability to unseen domain dynamics over TMP and task planning (TP)-RL methods. Yuqian Jiang, Fangkai Yang, Shiqi Zhang 0001, Peter Stone 0001 |
IROS | 1 |
| 2019 | Task planning in robotics: an empirical comparison of PDDL- and ASP-based systemsabstractRobots need task planning algorithms to sequence actions toward accomplishing goals that are impossible through individual actions. Off-the-shelf task planners can be used by intelligent robotics practitioners to solve a variety of planning problems. However, many different planners exist, each with different strengths and weaknesses, and there are no general rules for which planner would be best to apply to a given problem. In this study, we empirically compare the performance of state-of-the-art planners that use either the planning domain description language (PDDL) or answer set programming (ASP) as the underlying action language. PDDL is designed for task planning, and PDDL-based planners are widely used for a variety of planning problems. ASP is designed for knowledge-intensive reasoning, but can also be used to solve task planning problems. Given domain encodings that are as similar as possible, we find that PDDL-based planners perform better on problems with longer solutions, and ASP-based planners are better on tasks with a large number of objects or tasks in which complex reasoning is required to reason about action preconditions and effects. The resulting analysis can inform selection among general-purpose planning systems for particular robot task planning domains. Yuqian Jiang, Shiqi Zhang 0001, Piyush Khandelwal, Peter Stone 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |