Arunkumar Byravan

dblp:151/9400 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 4 since 2021Systems, architecture and hardware · 7 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Reinforcement learning · 22% Motion planning and robot control · 20% Deep learning architectures and training · 19%

Topics — the 26 heaviest of 30, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
positional encoding
0.912025
Learning the RoPEs: Better 2D and 3D Position Encodings with STRING · ICML 2025
Machine learning › Deep learning architectures and training › positional encoding
rotary position embedding
0.912025
Learning the RoPEs: Better 2D and 3D Position Encodings with STRING · ICML 2025
Robotics › Motion planning and robot control
robot learning
0.832021
Representation Matters: Improving Perception and Exploration for Robotics · ICRA 2021
Graph-Based Inverse Optimal Control for Robot Manipulation · IJCAI 2015
Prospection: Interpretable plans from language by predicting the future · ICRA 2019
Robotics › Legged, aerial and field robots › legged robots › legged robot locomotion
bipedal locomotion
0.712023
NeRF2Real: Sim2real Transfer of Vision-guided Bipedal Motion Skills using Neural Radiance Fields · ICRA 2023
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
0.712023
NeRF2Real: Sim2real Transfer of Vision-guided Bipedal Motion Skills using Neural Radiance Fields · ICRA 2023
Machine learning › Reinforcement learning
continuous control
0.612022
Evaluating Model-Based Planning and Planner Amortization for Continuous Control · ICLR 2022
Machine learning › Reinforcement learning
model-based reinforcement learning
0.612022
Evaluating Model-Based Planning and Planner Amortization for Continuous Control · ICLR 2022
Robotics › Motion planning and robot control
robot control
0.522022
SE3-Pose-Nets: Structured Deep Dynamics Models for Visuomotor Control · ICRA 2018
Evaluating Model-Based Planning and Planner Amortization for Continuous Control · ICLR 2022
Machine learning › Reinforcement learning
representation learning for control
0.512021
Representation Matters: Improving Perception and Exploration for Robotics · ICRA 2021
Machine learning › Reinforcement learning › function approximation
representation learning for reinforcement learning
0.512021
Representation Matters: Improving Perception and Exploration for Robotics · ICRA 2021
Robotics › Robot manipulation
language-conditioned task planning
0.412019
Prospection: Interpretable plans from language by predicting the future · ICRA 2019
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
task planning
0.412019
Prospection: Interpretable plans from language by predicting the future · ICRA 2019
Machine learning › Deep learning architectures and training › sequence modeling
deep dynamics model
0.312018
SE3-Pose-Nets: Structured Deep Dynamics Models for Visuomotor Control · ICRA 2018
Robotics › Motion planning and robot control › robot control
model-based control
0.312018
SE3-Pose-Nets: Structured Deep Dynamics Models for Visuomotor Control · ICRA 2018
Robotics › Robot manipulation
visuomotor control
0.312018
SE3-Pose-Nets: Structured Deep Dynamics Models for Visuomotor Control · ICRA 2018
Computer vision › 3D vision
rigid body motion
0.312017
SE3-nets: Learning rigid body motion using deep neural networks · ICRA 2017
Robotics › Motion planning and robot control
robot controller
0.312025
Learning the RoPEs: Better 2D and 3D Position Encodings with STRING · ICML 2025
Machine learning › Reinforcement learning › imitation learning › inverse reinforcement learning
inverse optimal control
0.212015
Graph-Based Inverse Optimal Control for Robot Manipulation · IJCAI 2015
Computer vision › 3D vision
neural radiance field
0.212023
NeRF2Real: Sim2real Transfer of Vision-guided Bipedal Motion Skills using Neural Radiance Fields · ICRA 2023
Computer vision › 3D vision
novel view synthesis
0.212023
NeRF2Real: Sim2real Transfer of Vision-guided Bipedal Motion Skills using Neural Radiance Fields · ICRA 2023
Machine learning › Learning theory › computational complexity
time-space tradeoffs
0.212014
Space-time functional gradient optimization for motion planning · ICRA 2014
Robotics › Motion planning and robot control
trajectory optimization
0.212014
Space-time functional gradient optimization for motion planning · ICRA 2014
Computer vision › 3D vision
pose estimation
0.112018
SE3-Pose-Nets: Structured Deep Dynamics Models for Visuomotor Control · ICRA 2018
Robotics › Robot manipulation › robot vision
object motion prediction
0.112017
SE3-nets: Learning rigid body motion using deep neural networks · ICRA 2017
Computer vision › 3D vision
point cloud
0.112017
SE3-nets: Learning rigid body motion using deep neural networks · ICRA 2017
Robotics › Motion planning and robot control
motion planning
0.112014
Space-time functional gradient optimization for motion planning · ICRA 2014

Methods — techniques the papers use, named apart from their topics

vision transformer · 0.9sim2real transfer · 0.7physics simulation · 0.7neural radiance field · 0.7planner amortization · 0.6model-based planning · 0.6reinforcement learning · 0.5auxiliary tasks · 0.5natural language processing · 0.4future prediction · 0.4
YearPublicationVenuePosition
2025 Learning the RoPEs: Better 2D and 3D Position Encodings with STRING
abstract
We introduce $\textbf{STRING}$: Separable Translationally Invariant Position Encodings. STRING extends Rotary Position Encodings, a recently proposed and widely used algorithm in large language models, via a unifying theoretical framework. Importantly, STRING still provides $\textbf{exact}$ translation invariance, including token coordinates of arbitrary dimensionality, whilst maintaining a low computational footprint. These properties are especially important in robotics, where efficient 3D token representation is key. We integrate STRING into Vision Transformers with RGB(-D) inputs (color plus optional depth), showing substantial gains, e.g. in open-vocabulary object detection and for robotics controllers. We complement our experiments with a rigorous mathematical analysis, proving the universality of our methods. Videos of STRING-based robotics controllers can be found here: https://sites.google.com/view/string-robotics.
Connor Schenck, Isaac Reid, Mithun George Jacob, Alex Bewley, Joshua Ainslie, David Rendleman, Deepali Jain, Mohit Sharma 0001, Avinava Dubey, Ayzaan Wahid, Sumeet Singh, René Wagner, Tianli Ding, Chuyuan Fu, Arunkumar Byravan, Jake Varley, Alexey A. Gritsenko, Matthias Minderer, Dmitry Kalashnikov, Jonathan Tompson, Vikas Sindhwani, Krzysztof Choromanski
ICML15
2023 NeRF2Real: Sim2real Transfer of Vision-guided Bipedal Motion Skills using Neural Radiance Fields
abstract
We present a system for applying sim2real approaches to “in the wild” scenes with realistic visuals, and to policies which rely on active perception using RGB cameras. Given a short video of a static scene collected using a generic phone, we learn the scene's contact geometry and a function for novel view synthesis using a Neural Radiance Field (NeRF). We augment the NeRF rendering of the static scene by overlaying the rendering of other dynamic objects (e.g. the robot's own body, a ball). A simulation is then created using the rendering engine in a physics simulator which computes contact dynamics from the static scene geometry (estimated from the NeRF vol-ume density) and the dynamic objects' geometry and physical properties (assumed known). We demonstrate that we can use this simulation to learn vision-based whole body navigation and ball pushing policies for a 20 degree-of-freedom humanoid robot with an actuated head-mounted RGB camera, and we successfully transfer these policies to a real robot.
Arunkumar Byravan, Jan Humplik, Leonard Hasenclever, Arthur Brussee, Francesco Nori, Tuomas Haarnoja, Ben Moran, Steven Bohez, Fereshteh Sadeghi, Bojan Vujatovic, Nicolas Heess
ICRA1
2022 Evaluating Model-Based Planning and Planner Amortization for Continuous Control
Arunkumar Byravan, Leonard Hasenclever, Piotr Trochim, Mehdi Mirza, Alessandro Davide Ialongo, Yuval Tassa, Jost Tobias Springenberg, Abbas Abdolmaleki, Nicolas Heess, Josh Merel, Martin A. Riedmiller
ICLR1
2021 Representation Matters: Improving Perception and Exploration for Robotics
abstract
Projecting high-dimensional environment observations into lower-dimensional structured representations can considerably improve data-efficiency for reinforcement learning in domains with limited data such as robotics. Can a single generally useful representation be found? In order to answer this question, it is important to understand how the representation will be used by the agent and what properties such a good representation should have. In this paper we systematically evaluate a number of common learnt and hand-engineered representations in the context of three robotics tasks: lifting, stacking and pushing of 3D blocks. The representations are evaluated in two use-cases: as input to the agent, or as a source of auxiliary tasks. Furthermore, the value of each representation is evaluated in terms of three properties: dimensionality, observability and disentanglement. We can significantly improve performance in both use-cases and demonstrate that some representations can perform commensurate to simulator states as agent inputs. Finally, our results challenge common intuitions by demonstrating that: 1) dimensionality strongly matters for task generation, but is negligible for inputs, 2) observability of task-relevant aspects mostly affects the input representation use-case, and 3) disentanglement leads to better auxiliary tasks, but has only limited benefits for input representations. This work serves as a step towards a more systematic understanding of what makes a good representation for control in robotics, enabling practitioners to make more informed choices for developing new learned or hand-engineered representations.
Markus Wulfmeier, Arunkumar Byravan, Tim Hertweck, Irina Higgins, Tejas Kulkarni, Malcolm Reynolds, Denis Teplyashin, Roland Hafner, Thomas Lampe, Martin A. Riedmiller
ICRA2
2019 Prospection: Interpretable plans from language by predicting the future
abstract
High-level human instructions often correspond to behaviors with multiple implicit steps. In order for robots to be useful in the real world, they must be able to to reason over both motions and intermediate goals implied by human instructions. In this work, we propose a framework for learning representations that convert from a natural-language command to a sequence of intermediate goals for execution on a robot. A key feature of this framework is prospection, training an agent not just to correctly execute the prescribed command, but to predict a horizon of consequences of an action before taking it. We demonstrate the fidelity of plans generated by our framework when interpreting real, crowd-sourced natural language commands for a robot in simulated scenes.
Chris Paxton 0001, Yonatan Bisk, Jesse Thomason, Arunkumar Byravan, Dieter Fox
ICRA4
2018 SE3-Pose-Nets: Structured Deep Dynamics Models for Visuomotor Control
abstract
In this work, we present an approach to deep visuomotor control using structured deep dynamics models. Our model, a variant of SE3-Nets, learns a low-dimensional pose embedding for visuomotor control via an encoder-decoder structure. Unlike prior work, our model is structured: given an input scene, our network explicitly learns to segment salient parts and predict their pose embedding and motion, modeled as a change in the pose due to the applied actions. We train our model using a pair of point clouds separated by an action and show that given supervision only through point-wise data associations between the frames our network is able to learn a meaningful segmentation of the scene along with consistent poses. We further show that our model can be used for closed-loop control directly in the learned low-dimensional pose space, where the actions are computed by minimizing pose error using gradient-based methods, similar to traditional model-based control. We present results on controlling a Baxter robot from raw depth data in simulation and RGBD data in the real world and compare against two baseline deep networks. We also test the robustness and generalization performance of our controller under changes in camera pose, lighting, occlusion, and motion. Our method is robust, runs in real-time, achieves good prediction of scene dynamics, and outperforms baselines on multiple control runs. Video results can be found at: https://rse-lab.cs.washington.edu/se3-structured-deep-ctrl/.
Arunkumar Byravan, Felix Leeb, Franziska Meier, Dieter Fox
ICRA1
2017 SE3-nets: Learning rigid body motion using deep neural networks
abstract
We introduce SE3-Nets which are deep neural networks designed to model and learn rigid body motion from raw point cloud data. Based only on sequences of depth images along with action vectors and point wise data associations, SE3-Nets learn to segment effected object parts and predict their motion resulting from the applied force. Rather than learning point wise flow vectors, SE3-Nets predict SE(3) transformations for different parts of the scene. Using simulated depth data of a table top scene and a robot manipulator, we show that the structure underlying SE3-Nets enables them to generate a far more consistent prediction of object motion than traditional flow based networks. Additional experiments with a depth camera observing a Baxter robot pushing objects on a table show that SE3-Nets also work well on real data.
Arunkumar Byravan, Dieter Fox
ICRA1
2015 Graph-Based Inverse Optimal Control for Robot Manipulation
Arunkumar Byravan, Mathew Monfort, Brian D. Ziebart, Byron Boots, Dieter Fox
IJCAI1
2014 Learning predictive models of a depth camera & manipulator from raw execution traces
abstract
In this paper, we attack the problem of learning a predictive model of a depth camera and manipulator directly from raw execution traces. While the problem of learning manipulator models from visual and proprioceptive data has been addressed before, existing techniques often rely on assumptions about the structure of the robot or tracked features in observation space. We make no such assumptions. Instead, we formulate the problem as that of learning a high-dimensional controlled stochastic process. We leverage recent work on nonparametric predictive state representations to learn a generative model of the depth camera and robotic arm from sequences of uninterpreted actions and observations. We perform several experiments in which we demonstrate that our learned model can accurately predict future depth camera observations in response to sequences of motor commands.
Byron Boots, Arunkumar Byravan, Dieter Fox
ICRA2
2014 Space-time functional gradient optimization for motion planning
abstract
Functional gradient algorithms (e.g. CHOMP) have recently shown great promise for producing locally optimal motion for complex many degree-of-freedom robots. A key limitation of such algorithms is the difficulty in incorporating constraints and cost functions that explicitly depend on time. We present T-CHOMP, a functional gradient algorithm that overcomes this limitation by directly optimizing in space-time. We outline a framework for joint space-time optimization, derive an efficient trajectory-wide update for maintaining time monotonicity, and demonstrate the significance of T-CHOMP over CHOMP in several scenarios. By manipulating time, T-CHOMP produces lower-cost trajectories leading to behavior that is meaningfully different from CHOMP.
Arunkumar Byravan, Byron Boots, Siddhartha S. Srinivasa, Dieter Fox
ICRA1