Kevin Stone

dblp:249/9395 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
3since 2021 · last 2023
0000-0001-9842-7755ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 since 2021Systems, architecture and hardware · 3 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
3D vision · 28% Reinforcement learning · 24% Motion planning and robot control · 22%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
policy learning
0.822023
Learning Compiler Pass Orders using Coreset and Normalized Value Prediction · ICML 2023
Learning Rope Manipulation Policies Using Dense Object Descriptors Trained on Synthetic Depth Data · ICRA 2020
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
masked modeling
0.712023
Masked Trajectory Models for Prediction, Representation, and Control · ICML 2023
Machine learning › Reinforcement learning
offline reinforcement learning
0.712023
Masked Trajectory Models for Prediction, Representation, and Control · ICML 2023
Machine learning › Reinforcement learning
trajectory modeling
0.712023
Masked Trajectory Models for Prediction, Representation, and Control · ICML 2023
Robotics › Motion planning and robot control
trajectory representation
0.712023
Masked Trajectory Models for Prediction, Representation, and Control · ICML 2023
Computer vision › 3D vision
3d shape reconstruction
0.612022
CenterSnap: Single-Shot Multi-Object 3D Shape Reconstruction and Categorical 6D Pose and Size Estimation · ICRA 2022
Computer vision › 3D vision › object pose estimation › 6d object pose estimation
category-level pose and size estimation
0.612022
CenterSnap: Single-Shot Multi-Object 3D Shape Reconstruction and Categorical 6D Pose and Size Estimation · ICRA 2022
Computer vision › 3D vision
object pose estimation
0.612022
CenterSnap: Single-Shot Multi-Object 3D Shape Reconstruction and Categorical 6D Pose and Size Estimation · ICRA 2022
Robotics › Robot manipulation
deformable object manipulation
0.412020
Learning Rope Manipulation Policies Using Dense Object Descriptors Trained on Synthetic Depth Data · ICRA 2020
Computer vision › 3D vision › local feature descriptor › descriptor learning
dense object descriptors
0.412020
Learning Rope Manipulation Policies Using Dense Object Descriptors Trained on Synthetic Depth Data · ICRA 2020
Robotics › Motion planning and robot control › robot control › compliant motion control
hybrid position/force control
0.412020
A Mobile Manipulation System for One-Shot Teaching of Complex Tasks in Homes · ICRA 2020
Robotics › Robot manipulation
mobile manipulation
0.412020
A Mobile Manipulation System for One-Shot Teaching of Complex Tasks in Homes · ICRA 2020
Robotics › Robot manipulation › deformable object manipulation
rope manipulation
0.412020
Learning Rope Manipulation Policies Using Dense Object Descriptors Trained on Synthetic Depth Data · ICRA 2020
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › plan representation
task graph
0.412020
A Mobile Manipulation System for One-Shot Teaching of Complex Tasks in Homes · ICRA 2020
Robotics › Motion planning and robot control › robot learning
task learning
0.412020
A Mobile Manipulation System for One-Shot Teaching of Complex Tasks in Homes · ICRA 2020
Robotics › Motion planning and robot control
whole-body control
0.412020
A Mobile Manipulation System for One-Shot Teaching of Complex Tasks in Homes · ICRA 2020
Computer vision › 3D vision › 3d reconstruction
single-view 3d reconstruction
0.212022
CenterSnap: Single-Shot Multi-Object 3D Shape Reconstruction and Categorical 6D Pose and Size Estimation · ICRA 2022
Computer vision › 3D vision › 3d scene modeling
scene representation
0.112020
A Mobile Manipulation System for One-Shot Teaching of Complex Tasks in Homes · ICRA 2020

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.3normalized value prediction · 1.3coreset · 1.3self-supervised learning · 0.7masked modeling · 0.7single-stage detection · 0.6per-pixel representation · 0.6virtual reality demonstration · 0.4parameterized primitives · 0.4dense visual embeddings · 0.4
YearPublicationVenuePosition
2023 Learning Compiler Pass Orders using Coreset and Normalized Value Prediction
abstract
Finding the optimal pass sequence of compilation can lead to a significant reduction in program size. Prior works on compilation pass ordering have two major drawbacks. They either require an excessive budget (in terms of the number of compilation passes) at compile time or fail to generalize to unseen programs. In this work, instead of predicting passes sequentially, we directly learn a policy on the pass sequence space, which outperforms the default -Oz flag by an average of 4.5% over a large collection (4683) of unseen code repositories from diverse domains across 14 datasets. To achieve this, we first identify a small set (termed coreset) of pass sequences that generally optimize the size of most programs. Then, a policy is learned to pick the optimal sequences by predicting the normalized values of the pass sequences in the coreset. Our results demonstrate that existing human-designed compiler passes can be improved with a simple yet effective technique that leverages pass sequence space which contains dense rewards, while approaches operating on the individual pass space may suffer from issues of sparse reward, and do not generalize well to held-out programs from different domains. Website: https://rlcompopt.github.io.
Youwei Liang, Kevin Stone, Ali Shameli, Chris Cummins, Mostafa Elhoushi, Jiadong Guo, Benoit Steiner, Pengtao Xie, Hugh Leather, Yuandong Tian
ICML2
2023 Masked Trajectory Models for Prediction, Representation, and Control
abstract
We introduce Masked Trajectory Models (MTM) as a generic abstraction for sequential decision making. MTM takes a trajectory, such as a state-action sequence, and aims to reconstruct the trajectory conditioned on random subsets of the same trajectory. By training with a highly randomized masking pattern, MTM learns versatile networks that can take on different roles or capabilities, by simply choosing appropriate masks at inference time. For example, the same MTM network can be used as a forward dynamics model, inverse dynamics model, or even an offline RL agent. Through extensive experiments in several continuous control tasks, we show that the same MTM network -- i.e. same weights -- can match or outperform specialized networks trained for the aforementioned capabilities. Additionally, we find that state representations learned by MTM can significantly accelerate the learning speed of traditional RL algorithms. Finally, in offline RL benchmarks, we find that MTM is competitive with specialized offline RL algorithms, despite MTM being a generic self-supervised learning method without any explicit RL components. Code is available at https://github.com/facebookresearch/mtm.
Philipp Wu, Arjun Majumdar, Kevin Stone, Igor Mordatch, Pieter Abbeel, Aravind Rajeswaran
ICML3
2022 CenterSnap: Single-Shot Multi-Object 3D Shape Reconstruction and Categorical 6D Pose and Size Estimation
abstract
This paper studies the complex task of simultaneous multi-object 3D reconstruction, 6D pose and size estimation from a single-view RGB-D observation. In contrast to instance- level pose estimation, we focus on a more challenging problem where CAD models are not available at inference time. Existing approaches mainly follow a complex multi-stage pipeline which first localizes and detects each object instance in the image and then regresses to either their 3D meshes or 6D poses. These approaches suffer from high-computational cost and low performance in complex multi-object scenarios, where occlusions can be present. Hence, we present a simple one- stage approach to predict both the 3D shape and estimate the 6D pose and size jointly in a bounding-box free manner. In particular, our method treats object instances as spatial centers where each center denotes the complete shape of an object along with its 6D pose and size. Through this per- pixel representation, our approach can reconstruct in real- time (40 FPS) multiple novel object instances and predict their 6D pose and sizes in a single-forward pass. Through extensive experiments, we demonstrate that our approach significantly outperforms all shape completion and categorical 6D pose and size estimation baselines on multi-object ShapeNet and NOCS datasets respectively with a 12.6% absolute improvement in mAP for 6D pose for novel real-world object instances.
Muhammad Zubair Irshad, Thomas Kollar, Michael Laskey, Kevin Stone, Zsolt Kira
ICRA4
2020 A Mobile Manipulation System for One-Shot Teaching of Complex Tasks in Homes
abstract
We describe a mobile manipulation hardware and software system capable of autonomously performing complex human-level tasks in real homes, after being taught the task with a single demonstration from a person in virtual reality. This is enabled by a highly capable mobile manipulation robot, whole-body task space hybrid position/force control, teaching of parameterized primitives linked to a robust learned dense visual embeddings representation of the scene, and a task graph of the taught behaviors. We demonstrate the robustness of the approach by presenting results for performing a variety of tasks, under different environmental conditions, in multiple real homes. Our approach achieves 85% overall success rate on three tasks that consist of an average of 45 behaviors each. The video is available at: https://youtu.be/HSyAGMGikLk.
Max Bajracharya, James Borders, Daniel M. Helmick, Thomas Kollar, Michael Laskey, John Leichty, Jeremy Ma, Umashankar Nagarajan, Akiyoshi Ochiai, Josh Petersen, Krishna Shankar, Kevin Stone, Yutaka Takaoka
ICRA12
2020 Learning Rope Manipulation Policies Using Dense Object Descriptors Trained on Synthetic Depth Data
abstract
Robotic manipulation of deformable 1D objects such as ropes, cables, and hoses is challenging due to the lack of high-fidelity analytic models and large configuration spaces. Furthermore, learning end-to-end manipulation policies directly from images and physical interaction requires significant time on a robot and can fail to generalize across tasks. We address these challenges using interpretable deep visual representations for rope, extending recent work on dense object descriptors for robot manipulation. This facilitates the design of interpretable and transferable geometric policies built on top of the learned representations, decoupling visual reasoning and control. We present an approach that learns point-pair correspondences between initial and goal rope configurations, which implicitly encodes geometric structure, entirely in simulation from synthetic depth images. We demonstrate that the learned representation — dense depth object descriptors (DDODs) — can be used to manipulate a real rope into a variety of different arrangements either by learning from demonstrations or using interpretable geometric policies. In 50 trials of a knot-tying task with the ABB YuMi Robot, the system achieves a 66% knot-tying success rate from previously unseen configurations. See https://tinyurl.com/rope-learning for supplementary material and videos.
Priya Sundaresan, Jennifer Grannen, Brijen Thananjeyan, Ashwin Balakrishna, Michael Laskey, Kevin Stone, Joseph Gonzalez 0001, Kenneth Y. Goldberg
ICRA6