Rishi Shah

dblp:195/8203 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Human-computer interaction and pervasive computing
1 paper
Human-robot interaction · 91% Collaborative and social computing · 9%
Artificial intelligence
3 papers
Reinforcement learning · 53% Graph learning · 27% Learning paradigms · 20%
Theoretical computer science
2 papers
Graph algorithms and graph theory · 84% Logic in computer science · 16%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Human-robot interaction › service robot
delivery robot
0.912025
Look Further: Socially-Compliant Navigation System in Residential Buildings · HRI 2025
Human-robot interaction
mobile robot
0.912025
Look Further: Socially-Compliant Navigation System in Residential Buildings · HRI 2025
Human-robot interaction
robot navigation
0.912025
Look Further: Socially-Compliant Navigation System in Residential Buildings · HRI 2025
Machine learning › Graph learning
graph neural network
0.812024
NeuroCut: A Neural Approach for Robust Graph Partitioning · KDD 2024
Graph algorithms and graph theory
graph partitioning
0.812024
NeuroCut: A Neural Approach for Robust Graph Partitioning · KDD 2024
Machine learning › Reinforcement learning › markov decision process
average-reward reinforcement learning
0.512021
Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning Tasks · AAAI 2021
Machine learning › Reinforcement learning › reward design
reward shaping
0.512021
Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning Tasks · AAAI 2021
Machine learning › Learning paradigms › curriculum learning
automatic curriculum generation
0.312017
Automatic Curriculum Graph Generation for Reinforcement Learning Agents · AAAI 2017
Machine learning › Learning paradigms
curriculum learning
0.312017
Automatic Curriculum Graph Generation for Reinforcement Learning Agents · AAAI 2017
Machine learning › Reinforcement learning
transfer learning in reinforcement learning
0.312017
Automatic Curriculum Graph Generation for Reinforcement Learning Agents · AAAI 2017
Collaborative and social computing › collaborative virtual environments › social virtual reality
personal space
0.312025
Look Further: Socially-Compliant Navigation System in Residential Buildings · HRI 2025
Logic in computer science
temporal logic
0.112021
Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning Tasks · AAAI 2021

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.5positional features · 1.5graph neural network · 1.5temporal logic translation · 1.0reward shaping · 1.0user study · 0.9transfer potential metric · 0.3directed acyclic graph curriculum · 0.3
YearPublicationVenuePosition
2025 Look Further: Socially-Compliant Navigation System in Residential Buildings
abstract
The distance at which a mobile robot reacts to a person strongly impacts various qualities of the human-robot interaction. In this paper, we focus on the navigation of a mobile delivery robot platform in a residential indoor hallway environment. Social navigation methods typically focus on avoiding uncomfortable human-robot interactions, such as when a robot encroaches on someone's personal space. Since personal space has been shown to be in the range of just a few meters, social navigation methods typically focus on deconflicting and resolving these short-range interactions. In this work, however, we demonstrate that by extending the reaction distance to over eight meters, far beyond the typical interaction distance, we can improve the human's perception of the robot's motion. We introduce the Proactive Lane-Changing (PLC) motion pattern and a navigation system that leverages it to react to people at an increased distance. This pattern consists of changing the robot's lateral position as it navigates down the hallway from the center to the side at an eight-meter distance from an oncoming person. We conducted a user study with 42 participants to assess their impressions of the delivery robot based on three service objectives: safety, smoothness, and politeness. In the straight hallway scenario (Frontal Approach), results showed significant improvement in each of these three objectives compared to typical motion patterns found in the literature: slowing down, stopping, and reactive collision avoidance in the proximity of a person. In contrast, in the intersection (Blind Corner) scenarios, none of the approaches performed significantly better than any other, with participants having a diverse range of preferences among robot motion patterns.
Akira Shiba, Marina Obata, Nathan Kau, Zoltán Beck, Rishi Shah, Michael Sudano, Sabrina Lee
HRI5
2024 NeuroCut: A Neural Approach for Robust Graph Partitioning
abstract
Graph partitioning aims to divide a graph into 𝑘 disjoint subsets while optimizing a specific partitioning objective.The majority of formulations related to graph partitioning exhibit NP-hardness due to their combinatorial nature.Conventional methods, like approximation algorithms or heuristics, are designed for distinct partitioning objectives and fail to achieve generalization across other important partitioning objectives.Recently machine learning-based methods have been developed that learn directly from data.Further, these methods have a distinct advantage of utilizing node features that carry additional information.However, these methods assume differentiability of target partitioning objective functions and cannot generalize for an unknown number of partitions, i.e., they assume the number of partitions is provided in advance.In this study, we develop NeuroCUT with two key innovations over previous methodologies.First, by leveraging a reinforcement learning-based framework over node representations derived from a graph neural network and positional features, NeuroCUT can accommodate any optimization objective, even those with non-differentiable functions.Second, we decouple the parameter space and the partition count making NeuroCUT inductive to any unseen number of partition, which is provided at query time.Through empirical evaluation, we demonstrate that NeuroCUT excels in identifying high-quality partitions, showcases strong generalization across a wide spectrum of partitioning objectives, and exhibits strong generalization to unseen partition count.
Rishi Shah, Krishnanshu Jain, Sahil Manchanda, Sourav Medya, Sayan Ranu
KDD1
2021 Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning Tasks
abstract
In continuing tasks, average-reward reinforcement learning may be a more appropriate problem formulation than the more common discounted reward formulation. As usual, learning an optimal policy in this setting typically requires a large amount of training experiences. Reward shaping is a common approach for incorporating domain knowledge into reinforcement learning in order to speed up convergence to an optimal policy. However, to the best of our knowledge, the theoretical properties of reward shaping have thus far only been established in the discounted setting. This paper presents the first reward shaping framework for average-reward learning and proves that, under standard assumptions, the optimal policy under the original reward function can be recovered. In order to avoid the need for manual construction of the shaping function, we introduce a method for utilizing domain knowledge expressed as a temporal logic formula. The formula is automatically translated to a shaping function that provides additional reward throughout the learning process. We evaluate the proposed method on three continuing tasks. In all cases, shaping speeds up the average-reward learning rate without any reduction in the performance of the learned policy compared to relevant baselines.
Yuqian Jiang, Suda Bharadwaj, Bo Wu 0005, Rishi Shah, Ufuk Topcu, Peter Stone 0001
AAAI4
2020 Deep R-Learning for Continual Area Sweeping
abstract
Coverage path planning is a well-studied problem in robotics in which a robot must plan a path that passes through every point in a given area repeatedly, usually with a uniform frequency. To address the scenario in which some points need to be visited more frequently than others, this problem has been extended to non-uniform coverage planning. This paper considers the variant of non-uniform coverage in which the robot does not know the distribution of relevant events beforehand and must nevertheless learn to maximize the rate of detecting events of interest. This continual area sweeping problem has been previously formalized in a way that makes strong assumptions about the environment, and to date only a greedy approach has been proposed. We generalize the continual area sweeping formulation to include fewer environmental constraints, and propose a novel approach based on reinforcement learning in a Semi-Markov Decision Process. This approach is evaluated in an abstract simulation and in a high fidelity Gazebo simulation. These evaluations show significant improvement upon the existing approach in general settings, which is especially relevant in the growing area of service robotics. We also present a video demonstration on a real service robot.
Rishi Shah, Yuqian Jiang, Justin W. Hart, Peter Stone 0001
IROS1
2018 PRISM: Pose Registration for Integrated Semantic Mapping
abstract
Many robotics applications involve navigating to positions specified in terms of their semantic significance. A robot operating in a hotel may need to deliver room service to a named room. In a hospital, it may need to deliver medication to a patient's room. The Building-Wide Intelligence Project at UT Austin has been developing a fleet of autonomous mobile robots, called BWIBots, which perform tasks in the computer science department. Tasks include guiding a person, delivering a message, or bringing an object to a location such as an office, lecture hall, or classroom. The process of constructing a map that a robot can use for navigation has been simplified by modern SLAM algorithms. The attachment of semantics to map data, however, remains a tedious manual process of labeling locations in otherwise automatically generated maps. This paper introduces a system called PRISM to automate a step in this process by enabling a robot to localize door signs - a semantic markup intended to aid the human occupants of a building - and to annotate these locations in its map.
Justin W. Hart, Rishi Shah, Sean Kirmani, Nick Walker 0001, Kathryn Baldauf, Nathan John, Peter Stone 0001
IROS2
2017 Automatic Curriculum Graph Generation for Reinforcement Learning Agents
abstract
In recent years, research has shown that transfer learning methods can be leveraged to construct curricula that sequence a series of simpler tasks such that performance on a final target task is improved. A major limitation of existing approaches is that such curricula are handcrafted by humans that are typically domain experts. To address this limitation, we introduce a method to generate a curriculum based on task descriptors and a novel metric of transfer potential. Our method automatically generates a curriculum as a directed acyclic graph (as opposed to a linear sequence as done in existing work). Experiments in both discrete and continuous domains show that our method produces curricula that improve the agent's learning performance when compared to the baseline condition of learning on the target task from scratch.
Maxwell Svetlik, Matteo Leonetti, Jivko Sinapov, Rishi Shah, Nick Walker 0001, Peter Stone 0001
AAAI4