Edward S. Hu

dblp:245/4627 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 5 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 29% Motion planning and robot control · 28% Robot manipulation · 16%

Topics — the 13 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Motion planning and robot control
robot learning
2.232025
Real-World Reinforcement Learning of Active Perception Behaviors · NeurIPS 2025
Privileged Sensing Scaffolds Reinforcement Learning · ICLR 2024
Know Thyself: Transferable Visual Control Policies Through Robot-Awareness · ICLR 2022
Robotics › Robot navigation and mapping
active perception
0.912025
Real-World Reinforcement Learning of Active Perception Behaviors · NeurIPS 2025
Natural language and speech › Language models and text generation
text generation
0.912025
The Belief State Transformer · ICLR 2025
Machine learning › Deep learning architectures and training
transformer
0.912025
The Belief State Transformer · ICLR 2025
Machine learning › Reinforcement learning
exploration
0.712023
Planning Goals for Exploration · ICLR 2023
Machine learning › Reinforcement learning › exploration › autonomous exploration › mobile robot exploration
exploration planning
0.712023
Planning Goals for Exploration · ICLR 2023
Machine learning › Reinforcement learning › exploration › directed exploration
goal-directed exploration
0.712023
Planning Goals for Exploration · ICLR 2023
Robotics › Motion planning and robot control
robot control
0.612022
Know Thyself: Transferable Visual Control Policies Through Robot-Awareness · ICLR 2022
Machine learning › Reinforcement learning › hierarchical reinforcement learning
skill learning
0.412019
Composing Complex Skills by Learning Transition Policies · ICLR (Poster) 2019
Robotics › Robot navigation and mapping
sensor fusion
0.212024
Privileged Sensing Scaffolds Reinforcement Learning · ICLR 2024
Machine learning › Reinforcement learning
imitation learning
0.112021
IKEA Furniture Assembly Environment for Long-Horizon Complex Manipulation Tasks · ICRA 2021
Robotics › Motion planning and robot control › motion planning › manipulation planning
long-horizon manipulation
0.112021
IKEA Furniture Assembly Environment for Long-Horizon Complex Manipulation Tasks · ICRA 2021
Robotics › Motion planning and robot control › robot control
low-level control
0.112021
IKEA Furniture Assembly Environment for Long-Horizon Complex Manipulation Tasks · ICRA 2021

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.7imitation learning · 1.4privileged value function · 0.9model-based control · 0.9demonstration bootstrapping · 0.9asymmetric advantage weighted regression · 0.9world model · 0.8reward estimator · 0.8privileged sensing · 0.8robot-aware representation · 0.6
YearPublicationVenuePosition
2025 The Belief State Transformer
abstract
We introduce the "Belief State Transformer", a next-token predictor that takes both a prefix and suffix as inputs, with a novel objective of predicting both the next token for the prefix and the previous token for the suffix. The Belief State Transformer effectively learns to solve challenging problems that conventional forward-only transformers struggle with, in a domain-independent fashion. Key to this success is learning a compact belief state that captures all relevant information necessary for accurate predictions. Empirical ablations show that each component of the model is essential in difficult scenarios where standard Transformers fall short. For the task of story writing with known prefixes and suffixes, our approach outperforms the Fill-in-the-Middle method for reaching known goals and demonstrates improved performance even when the goals are unknown. Altogether, the Belief State Transformer enables more efficient goal-conditioned decoding, better test-time inference, and high-quality text representations on small scale problems. Website: https://edwhu.github.io/bst-website
Edward S. Hu, Kwangjun Ahn, Manan Tomar, Ada Langford, Dinesh Jayaraman, Alex Lamb, John Langford 0001
ICLR1
2025 The Value of Sensory Information to a Robot
abstract
A decision-making agent, such as a robot, must observe and react to any new task-relevant information that becomes available from its environment. We seek to study a fundamental scientific question: what value does sensory information hold to an agent at various moments in time during the execution of a task? Towards this, we empirically study agents of varying architectures, generated with varying policy synthesis approaches (imitation, RL, model-based control), on diverse robotics tasks. For each robotic agent, we characterize its regret in terms of performance degradation when state observations are withheld from it at various task states for varying lengths of time. We find that sensory information is surprisingly rarely task-critical in many commonly studied task setups. Task characteristics such as stochastic dynamics largely dictate the value of sensory information for a well-trained robot; policy architectures such as planning vs. reactive control generate more nuanced second-order effects. Further, sensing efficiency is curiously correlated with task proficiency: in particular, fully trained high-performing agents are more robust to sensor loss than novice agents early in their training. Overall, our findings characterize the tradeoffs between sensory information and task performance in practical sequential decision making tasks, and pave the way towards the design of more resource-efficient decision-making agents.
Arjun Krishna, Edward S. Hu, Dinesh Jayaraman
ICLR2
2025 Real-World Reinforcement Learning of Active Perception Behaviors
abstract
A robot's instantaneous sensory observations do not always reveal task-relevant state information. Under such partial observability, optimal behavior typically involves explicitly acting to gain the missing information. Today's standard robot learning techniques struggle to produce such active perception behaviors. We propose a simple real-world robot learning recipe to efficiently train active perception policies. Our approach, asymmetric advantage weighted regression (AAWR), exploits access to "privileged" extra sensors at training time. The privileged sensors enable training high-quality privileged value functions that aid in estimating the advantage of the target policy. Bootstrapping from a small number of potentially suboptimal demonstrations and an easy-to-obtain coarse policy initialization, AAWR quickly acquires active perception behaviors and boosts task performance. In evaluations on 8 manipulation tasks on 3 robots spanning varying degrees of partial observability, AAWR synthesizes reliable active perception behaviors that outperform all prior approaches. When initialized with a "generalist" robot policy that struggles with active perception tasks, AAWR efficiently generates information-gathering behaviors that allow it to operate under severe partial observability for manipulation tasks. Website: https://penn-pal-lab.github.io/aawr/
Edward S. Hu, Xingfang Yuan, Fiona Luo, Muyao Li, Gaspard Lambrechts, Oleh Rybkin, Dinesh Jayaraman
NeurIPS1
2024 Privileged Sensing Scaffolds Reinforcement Learning
abstract
We need to look at our shoelaces as we first learn to tie them but having mastered this skill, can do it from touch alone. We call this phenomenon “sensory scaffolding”: observation streams that are not needed by a master might yet aid a novice learner. We consider such sensory scaffolding setups for training artificial agents. For example, a robot arm may need to be deployed with just a low-cost, robust, general-purpose camera; yet its performance may improve by having privileged training-time-only access to informative albeit expensive and unwieldy motion capture rigs or fragile tactile sensors. For these settings, we propose “Scaffolder”, a reinforcement learning approach which effectively exploits privileged sensing in critics, world models, reward estimators, and other such auxiliary components that are only used at training time, to improve the target policy. For evaluating sensory scaffolding agents, we design a new “S3” suite of ten diverse simulated robotic tasks that explore a wide range of practical sensor setups. Agents must use privileged camera sensing to train blind hurdlers, privileged active visual perception to help robot arms overcome visual occlusions, privileged touch sensors to train robot hands, and more. Scaffolder easily outperforms relevant prior baselines and frequently performs comparably even to policies that have test-time access to the privileged sensors. Website: https://penn-pal-lab.github.io/scaffolder/
Edward S. Hu, James Springer, Oleh Rybkin, Dinesh Jayaraman
ICLR1
2023 Planning Goals for Exploration
Edward S. Hu, Oleh Rybkin, Dinesh Jayaraman
ICLR1
2022 Know Thyself: Transferable Visual Control Policies Through Robot-Awareness
Edward S. Hu, Oleh Rybkin, Dinesh Jayaraman
ICLR1
2021 IKEA Furniture Assembly Environment for Long-Horizon Complex Manipulation Tasks
abstract
The IKEA Furniture Assembly Environment is one of the first benchmarks for testing and accelerating the automation of long-horizon and hierarchical manipulation tasks. The environment is designed to advance reinforcement learning and imitation learning from simple toy tasks to complex tasks requiring both long-term planning and sophisticated low-level control. Our environment features 60 furniture models, 6 robots, photorealistic rendering, and domain randomization. We evaluate reinforcement learning and imitation learning methods on the proposed environment. Our experiments show furniture assembly is a challenging task due to its long horizon and sophisticated manipulation requirements, which provides ample opportunities for future research. The environment is publicly available at https://clvrai.com/furniture.
Youngwoon Lee, Edward S. Hu, Joseph J. Lim
ICRA2
2019 Composing Complex Skills by Learning Transition Policies
Youngwoon Lee, Shao-Hua Sun, Sriram Somasundaram, Edward S. Hu, Joseph J. Lim
ICLR (Poster)4