Jacob Arkin

dblp:188/0992 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
6since 2021 · last 2024
0000-0002-1074-9248ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 6 since 2021Systems, architecture and hardware · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2024 PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based Sampling
abstract
Prompt optimization aims to find the best prompt to a large language model (LLM) for a given task. LLMs have been successfully used to help find and improve prompt candidates for single-step tasks. However, realistic tasks for agents are multi-step and introduce new challenges: (1) Prompt content is likely to be more extensive and complex, making it more difficult for LLMs to analyze errors, (2) the impact of an individual step is difficult to evaluate, and (3) different people may have varied preferences about task execution. While humans struggle to optimize prompts, they are good at providing feedback about LLM outputs; we therefore introduce a new LLM-driven discrete prompt optimization framework PROMST that incorporates human-designed feedback rules to automatically offer direct suggestions for improvement. We also use an extra learned heuristic model that predicts prompt performance to efficiently sample from prompt candidates. This approach significantly outperforms both human-engineered prompts and several other prompt optimization methods across 11 representative multi-step tasks (an average 10.6%-29.3% improvement to current best methods on five LLMs respectively). We believe our work can serve as a benchmark for automatic prompt optimization for LLM-driven multi-step tasks.
Yongchao Chen, Jacob Arkin, Yilun Hao, Yang Zhang 0001, Nicholas Roy, Chuchu Fan
EMNLP2
2024 AutoTAMP: Autoregressive Task and Motion Planning with LLMs as Translators and Checkers
abstract
For effective human-robot interaction, robots need to understand, plan, and execute complex, long-horizon tasks described by natural language. Recent advances in large language models (LLMs) have shown promise for translating natural language into robot action sequences for complex tasks. However, existing approaches either translate the natural language directly into robot trajectories or factor the inference process by decomposing language into task sub-goals and relying on a motion planner to execute each sub-goal. When complex environmental and temporal constraints are involved, inference over planning tasks must be performed jointly with motion plans using traditional task-and-motion planning (TAMP) algorithms, making factorization into subgoals untenable. Rather than using LLMs to directly plan task sub-goals, we instead perform few-shot translation from natural language task descriptions to an intermediate task representation that can then be consumed by a TAMP algorithm to jointly solve the task and motion plan. To improve translation, we automatically detect and correct both syntactic and semantic errors via autoregressive re-prompting, resulting in significant improvements in task completion. We show that our approach outperforms several methods using LLMs as planners in complex task domains. See our project website§for prompts, videos, and code.
Yongchao Chen, Jacob Arkin, Charles Dawson 0001, Yang Zhang 0001, Nicholas Roy, Chuchu Fan
ICRA2
2024 Scalable Multi-Robot Collaboration with Large Language Models: Centralized or Decentralized Systems?
abstract
A flurry of recent work has demonstrated that pre-trained large language models (LLMs) can be effective task planners for a variety of single-robot tasks. The planning performance of LLMs is significantly improved via prompting techniques, such as in-context learning or re-prompting with state feedback, placing new importance on the token budget for the context window. An under-explored but natural next direction is to investigate LLMs as multi-robot task planners. However, long-horizon, heterogeneous multi-robot planning introduces new challenges of coordination while also pushing up against the limits of context window length. It is therefore critical to find token-efficient LLM planning frameworks that are also able to reason about the complexities of multi-robot coordination. In this work, we compare the task success rate and token efficiency of four multi-agent communication frameworks (centralized, decentralized, and two hybrid) as applied to four coordination-dependent multi-agent 2D task scenarios for increasing numbers of agents. We find that a hybrid framework achieves better task success rates across all four tasks and scales better to more agents. We further demonstrate the hybrid frameworks in 3D simulations where the vision-to-text problem and dynamical errors are considered. See our project website4for prompts, videos, and code.
Yongchao Chen, Jacob Arkin, Yang Zhang 0001, Nicholas Roy, Chuchu Fan
ICRA2
2023 Language Guided Temporally Adaptive Perception for Efficient Natural Language Grounding in Cluttered Dynamic Worlds
abstract
As robots operate alongside humans in shared spaces, such as homes and offices, it is essential to have an effective mechanism for interacting with them. Natural language offers an intuitive interface for communicating with robots, but most of the recent approaches to grounded language understanding reason only in the context of an instantaneous state of the world. Though this allows for interpreting a variety of utterances in the current context of the world, these models fail to interpret utterances which require the knowledge of past dynamics of the world, thereby hindering effective human-robot collaboration in dynamic environments. Constructing a comprehensive model of the world that tracks the dynamics of all objects in the robot's workspace is computationally expensive and difficult to scale with increasingly complex environments. To address this challenge, we propose a learned model of language and perception that facilitates the construction of temporally compact models of dynamic worlds through closed-loop grounding and perception. Our experimental results on the task of grounding referring expressions demonstrate more accurate interpretation of robot instructions in cluttered and dynamic table-top environments without a significant increase in runtime as compared to an open-loop baseline.
Siddharth Patki, Jacob Arkin, Nikola Raicevic, Thomas M. Howard
IROS2
2022 An Efficient Algorithm for Visualization and Interpretation of Grounded Language Models
abstract
Contemporary approaches to grounded language communication accept an utterance and current world representation as input and produce symbols representing the meaning as output. Since modern approaches to language understanding for human-robot interaction use techniques rooted in machine learning, the quality or sensitivity of the solution is often opaque relative to small changes in input. Although it is possible to sample and visualize solutions over a large space of inputs, naïve application of current techniques is often prohibitively expensive for real-time feedback. In this paper we address this problem by reformulating the inference process of Distributed Correspondence Graphs to only recompute subsets of spatially dependent constituent features over a space of sampled environment models. We quantitatively evaluate the speed of inference in physical experiments involving a tabletop robot manipulation scenario. We demonstrate the ability to visualize configurations of the environment where symbol grounding produces consistent solutions in real-time and illustrate how these techniques can be used to identify and repair gaps or inaccuracies in training data.
Jacob Arkin, Siddharth Patki, Joshua D. Rosser, Thomas M. Howard
RO-MAN1
2021 Discrete Optimization of Adaptive State Lattices for Iterative Motion Planning on Unmanned Ground Vehicles
abstract
Robust motion planners for unmanned ground vehicles must minimize risk while obeying vehicle mobility constraints. Algorithms such as the State Lattice (SL) utilize offline computation to generate expressive control sets which form recombinant search spaces, enabling the use of heuristic search to efficiently produce feasible motion plans online. The Adaptive State Lattice (ASL) demonstrated that local optimizations of the continuous states explored by heuristic search can produce lower-cost solutions in less time than more densely sampled unadapted lattices in sufficiently complex environments. However, the computational cost of this online adaptation limits the application of ASL for mobile robot navigation. We present the Efficiently Adaptive State Lattice (EASL), a novel formalism for online discrete ASL adaptation to overcome this limitation. By discretizing the space of states considered during adaptation, EASL limits the set of feasible motions which could arise during search. This permits the precomputation of an approximation of all motions that could be expressed by an ASL. This approximation removes the online trajectory generation component of the ASL while retaining the benefits of lattice adaptation and enables the use of precomputed swaths for evaluating edge costs. Experimental results demonstrate how an EASL-based planner can generate lower-cost paths than a SL-based planner in roughly equal to or less than the same amount of time.
Benned Hedegaard, Ethan Fahnestock, Jacob Arkin, Ashwin Menon, Thomas M. Howard
IROS3
2017 Grounding Abstract Spatial Concepts for Language Interaction with Robots
abstract
Our goal is to develop models that allow a robot to understand or ``ground" natural language instructionsin the context of its world model. Contemporary approaches estimate correspondences between an instruction and possible candidate groundings such as objects, regions and goals for a robot's action. However, these approaches are unable to reason about abstract or hierarchical concepts such as rows, columns and groups that are relevant in a manipulation domain. We introduce a probabilistic model that incorporates an expressive space of abstract spatial concepts as well as notions of cardinality and ordinality. Abstract concepts are introduced as explicit hierarchical symbols correlated with concrete groundings. Crucially, the abstract groundings form a Markov boundary over concrete groundings, effectively de-correlating them from the remaining variables in the graph which reduces the complexity of training and inference in the model. Empirical evaluation demonstrates accurate grounding of abstract concepts embedded in complex natural language instructions commanding a robot manipulator. The proposed inference method leads to significant efficiency gains compared to the baseline, with minimal trade-off in accuracy.
Rohan Paul, Jacob Arkin, Nicholas Roy, Thomas M. Howard
IJCAI2
2017 Contextual awareness: Understanding monologic natural language instructions for autonomous robots
abstract
Today, there are many examples of humans and robots regularly interacting in a variety of domains, such as manufacturing, coordinated assembly, and rehabilitation. A resulting demand for more generally accessible communication interfaces has motivated several recent independent research efforts focused on providing robotic systems with a robust natural language interface. Natural language interfaces enable intuitive interaction for untrained and non-expert users. However, achieving real-time performance is particularly challenging, yet essential, to enable flexible, efficient communication. The length of the language input directly impacts the run-time performance and quickly becomes a practical issue when the input is a sequence of multiple sentences, or a monologue. In this work, we propose a variant of a contemporary probabilistic graphical model for language understanding that introduces novel segmentation of the input into a sequence of sentences to be labeled in order. We introduce the notion of a continuously updated prior context that retains the meaning of previous sentences as the inference process proceeds. This prior context serves as evidence during future sentence evaluations. We evaluate our model on two natural language corpora, and demonstrate its utility on a Clearpath Husky A200 mobile manipulator and a simulated Rethink Robotics Baxter Robot.
Jacob Arkin, Matthew R. Walter, Adrian Boteanu, Michael E. Napoli, Harel Biggie, Hadas Kress-Gazit, Thomas M. Howard
RO-MAN1
2016 A model for verifiable grounding and execution of complex natural language instructions
abstract
Current methods of grounding natural language instructions do not include reactive or temporal components, making these methods unsuitable for instructions describing tasks as sets of conditional instructions. We introduce the Verifiable Distributed Correspondence Graph (V-DCG) model, which enables the validation of natural language instructions by using Linear Temporal Logic (LTL) specifications together with physical world groundings. We demonstrate the V-DCG model on a physical robot and provide examples of the output our system produces for natural language instructions.
Adrian Boteanu, Thomas M. Howard, Jacob Arkin, Hadas Kress-Gazit
IROS3