EDBT 2026 Demo / reviewers in the wild / expert
Edward S. Hu
dblp:245/4627
· DBLP profile ↗
8ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Reinforcement learning · 29% Motion planning and robot control · 28% Robot manipulation · 16% |
Topics — the 13 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Robotics › Motion planning and robot control
robot learning |
2.2 | 3 | 2025 | Real-World Reinforcement Learning of Active Perception Behaviors · NeurIPS 2025 Privileged Sensing Scaffolds Reinforcement Learning · ICLR 2024 Know Thyself: Transferable Visual Control Policies Through Robot-Awareness · ICLR 2022 |
Robotics › Robot navigation and mapping
active perception |
0.9 | 1 | 2025 | Real-World Reinforcement Learning of Active Perception Behaviors · NeurIPS 2025 |
Natural language and speech › Language models and text generation
text generation |
0.9 | 1 | 2025 | The Belief State Transformer · ICLR 2025 |
Machine learning › Deep learning architectures and training
transformer |
0.9 | 1 | 2025 | The Belief State Transformer · ICLR 2025 |
Machine learning › Reinforcement learning
exploration |
0.7 | 1 | 2023 | Planning Goals for Exploration · ICLR 2023 |
Machine learning › Reinforcement learning › exploration › autonomous exploration › mobile robot exploration
exploration planning |
0.7 | 1 | 2023 | Planning Goals for Exploration · ICLR 2023 |
Machine learning › Reinforcement learning › exploration › directed exploration
goal-directed exploration |
0.7 | 1 | 2023 | Planning Goals for Exploration · ICLR 2023 |
Robotics › Motion planning and robot control
robot control |
0.6 | 1 | 2022 | Know Thyself: Transferable Visual Control Policies Through Robot-Awareness · ICLR 2022 |
Machine learning › Reinforcement learning › hierarchical reinforcement learning
skill learning |
0.4 | 1 | 2019 | Composing Complex Skills by Learning Transition Policies · ICLR (Poster) 2019 |
Robotics › Robot navigation and mapping
sensor fusion |
0.2 | 1 | 2024 | Privileged Sensing Scaffolds Reinforcement Learning · ICLR 2024 |
Machine learning › Reinforcement learning
imitation learning |
0.1 | 1 | 2021 | IKEA Furniture Assembly Environment for Long-Horizon Complex Manipulation Tasks · ICRA 2021 |
Robotics › Motion planning and robot control › motion planning › manipulation planning
long-horizon manipulation |
0.1 | 1 | 2021 | IKEA Furniture Assembly Environment for Long-Horizon Complex Manipulation Tasks · ICRA 2021 |
Robotics › Motion planning and robot control › robot control
low-level control |
0.1 | 1 | 2021 | IKEA Furniture Assembly Environment for Long-Horizon Complex Manipulation Tasks · ICRA 2021 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.7imitation learning · 1.4privileged value function · 0.9model-based control · 0.9demonstration bootstrapping · 0.9asymmetric advantage weighted regression · 0.9world model · 0.8reward estimator · 0.8privileged sensing · 0.8robot-aware representation · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | The Belief State TransformerabstractWe introduce the "Belief State Transformer", a next-token predictor that takes both a prefix and suffix as inputs, with a novel objective of predicting both the next token for the prefix and the previous token for the suffix. The Belief State Transformer effectively learns to solve challenging problems that conventional forward-only transformers struggle with, in a domain-independent fashion. Key to this success is learning a compact belief state that captures all relevant information necessary for accurate predictions.
Empirical ablations show that each component of the model is essential in difficult scenarios where standard Transformers fall short.
For the task of story writing with known prefixes and suffixes, our approach outperforms the Fill-in-the-Middle method for reaching known goals and demonstrates improved performance even when the goals are unknown. Altogether, the Belief State Transformer enables more efficient goal-conditioned decoding, better test-time inference, and high-quality text representations on small scale problems. Website: https://edwhu.github.io/bst-website Edward S. Hu, Kwangjun Ahn, Manan Tomar, Ada Langford, Dinesh Jayaraman, Alex Lamb, John Langford 0001 |
ICLR | 1 |
| 2025 | The Value of Sensory Information to a RobotabstractA decision-making agent, such as a robot, must observe and react to any new task-relevant information that becomes available from its environment. We seek to study a fundamental scientific question: what value does sensory information hold to an agent at various moments in time during the execution of a task? Towards this, we empirically study agents of varying architectures, generated with varying policy synthesis approaches (imitation, RL, model-based control), on diverse robotics tasks. For each robotic agent, we characterize its regret in terms of performance degradation when state observations are withheld from it at various task states for varying lengths of time. We find that sensory information is surprisingly rarely task-critical in many commonly studied task setups. Task characteristics such as stochastic dynamics largely dictate the value of sensory information for a well-trained robot; policy architectures such as planning vs. reactive control generate more nuanced second-order effects. Further, sensing efficiency is curiously correlated with task proficiency: in particular, fully trained high-performing agents are more robust to sensor loss than novice agents early in their training. Overall, our findings characterize the tradeoffs between sensory information and task performance in practical sequential decision making tasks, and pave the way towards the design of more resource-efficient decision-making agents. Arjun Krishna, Edward S. Hu, Dinesh Jayaraman |
ICLR | 2 |
| 2025 | Real-World Reinforcement Learning of Active Perception BehaviorsabstractA robot's instantaneous sensory observations do not always reveal task-relevant state information. Under such partial observability, optimal behavior typically involves explicitly acting to gain the missing information.
Today's standard robot learning techniques struggle to produce such active perception behaviors.
We propose a simple real-world robot learning recipe to efficiently train active perception policies. Our approach, asymmetric advantage weighted regression (AAWR), exploits access to "privileged" extra sensors at training time. The privileged sensors enable training high-quality privileged value functions that aid in estimating the advantage of the target policy. Bootstrapping from a small number of potentially suboptimal demonstrations and an easy-to-obtain coarse policy initialization, AAWR quickly acquires active perception behaviors and boosts task performance. In evaluations on 8 manipulation tasks on 3 robots spanning varying degrees of partial observability, AAWR synthesizes reliable active perception behaviors that outperform all prior approaches. When initialized with a "generalist" robot policy that struggles with active perception tasks, AAWR efficiently generates information-gathering behaviors that allow it to operate under severe partial observability for manipulation tasks. Website:
https://penn-pal-lab.github.io/aawr/ Edward S. Hu, Xingfang Yuan, Fiona Luo, Muyao Li, Gaspard Lambrechts, Oleh Rybkin, Dinesh Jayaraman |
NeurIPS | 1 |
| 2024 | Privileged Sensing Scaffolds Reinforcement LearningabstractWe need to look at our shoelaces as we first learn to tie them but having mastered this skill, can do it from touch alone. We call this phenomenon “sensory scaffolding”: observation streams that are not needed by a master might yet aid a novice learner. We consider such sensory scaffolding setups for training artificial agents. For example, a robot arm may need to be deployed with just a low-cost, robust, general-purpose camera; yet its performance may improve by having privileged training-time-only access to informative albeit expensive and unwieldy motion capture rigs or fragile tactile sensors. For these settings, we propose “Scaffolder”, a reinforcement learning approach which effectively exploits privileged sensing in critics, world models, reward estimators, and other such auxiliary components that are only used at training time, to improve the target policy. For evaluating sensory scaffolding agents, we design a new “S3” suite of ten diverse simulated robotic tasks that explore a wide range of practical sensor setups. Agents must use privileged camera sensing to train blind hurdlers, privileged active visual perception to help robot arms overcome visual occlusions, privileged touch sensors to train robot hands, and more. Scaffolder easily outperforms relevant prior baselines and frequently performs comparably even to policies that have test-time access to the privileged sensors. Website: https://penn-pal-lab.github.io/scaffolder/ Edward S. Hu, James Springer, Oleh Rybkin, Dinesh Jayaraman |
ICLR | 1 |
| 2023 | Planning Goals for Exploration
Edward S. Hu, Oleh Rybkin, Dinesh Jayaraman |
ICLR | 1 |
| 2022 | Know Thyself: Transferable Visual Control Policies Through Robot-Awareness
Edward S. Hu, Oleh Rybkin, Dinesh Jayaraman |
ICLR | 1 |
| 2021 | IKEA Furniture Assembly Environment for Long-Horizon Complex Manipulation TasksabstractThe IKEA Furniture Assembly Environment is one of the first benchmarks for testing and accelerating the automation of long-horizon and hierarchical manipulation tasks. The environment is designed to advance reinforcement learning and imitation learning from simple toy tasks to complex tasks requiring both long-term planning and sophisticated low-level control. Our environment features 60 furniture models, 6 robots, photorealistic rendering, and domain randomization. We evaluate reinforcement learning and imitation learning methods on the proposed environment. Our experiments show furniture assembly is a challenging task due to its long horizon and sophisticated manipulation requirements, which provides ample opportunities for future research. The environment is publicly available at https://clvrai.com/furniture. Youngwoon Lee, Edward S. Hu, Joseph J. Lim |
ICRA | 2 |
| 2019 | Composing Complex Skills by Learning Transition Policies
Youngwoon Lee, Shao-Hua Sun, Sriram Somasundaram, Edward S. Hu, Joseph J. Lim |
ICLR (Poster) | 4 |