VLDB 2026 Research / reviewers in the wild / expert
Deval Shah
dblp:217/0997 · also Deval A. Shah
· DBLP profile ↗
9ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-1537-1686ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Neural Encodings for Energy-Efficient Motion PlanningabstractNeural motion planners can increase motion planning quality and, by reducing collision detection computations, improve runtime. However, when profiled on an accelerator-rich hardware system, neural planning contributes to more than 50% of the runtime, and 33% of the computation energy consumption, motivating the design of compute- and energy-efficient neural planners. In this work, we propose a neural planner using Binary Encoded Labels (BEL), where a set of binary classifiers are used instead of a typical regression network. Compared to conventional regression-based neural planners, the proposed BEL neural planner reduces neural planning (inference) computation and collision detection checks while maintaining equal or higher motion planning success rate across various motion planning benchmarks. This computation reduction can improve the computation energy efficiency of neural planning by$1.4 \times-21.4 \times$. Finally, we demonstrate the trade-offs between collision detection and neural planning computation to maximize energy efficiency for different hardware configurations. Jocelyn Zhao, Deval Shah, Tor M. Aamodt |
ICRA | 2 |
| 2024 | Collision Prediction for Robotics AcceleratorsabstractMotion planning in dynamic environments is an important task for autonomous robotics. Emerging approaches employ neural networks that can learn by observing (e.g., human) experts. Such motion planners react to the environment by continually proposing candidate paths to reach a goal. Some of these candidate paths may be unsafe-i.e., cause collisions. Hence, proposed paths must be checked for safety using collision detection. We observe that $25 \%-41 \%$ of the resulting collision detection queries can be eliminated if we can anticipate which queries will return an unsafe result. We leverage this observation to propose a mechanism, COORD, to predict whether a given robot position (pose) along a proposed path will result in a collision. By prioritizing the detailed evaluation of predicted collisions, COORD enables quickly eliminating invalid paths proposed by neural network and other sampling based motion planners. COORD does this by exploiting the physical spatial locality of different robot poses and using simple hashing and saturating counters. We demonstrate the potential of collision prediction on different computation platforms, including CPU, GPU, and ASIC. We further propose a hardware collision prediction unit (COPU), and integrate it with an existing collision detection accelerator. This results in an average $17.2 \%-32.1 \%$ decrease in number of collision detection queries across different motion planning algorithms and robots. When applied to a state-of-the-art neural motion planner [41], COORD improves performance/watt by $1.23 \times$ on average for motion planning queries of varying difficulty levels. Further, we find that the benefits of collision prediction grow as the compute complexity of motion planning queries increases and provides $1.30 \times \mathrm{im}-$ provement in performance/watt in narrow passages and cluttered environments. Deval Shah, Tor M. Aamodt |
ISCA | 1 |
| 2024 | Characterizing and Improving Resilience of Accelerators to Memory Errors in Autonomous RobotsabstractMotion planning is a computationally intensive and well-studied problem in autonomous robots. However, motion planning hardware accelerators (MPA) must be soft-error resilient for deployment in safety-critical applications, and blanket application of traditional mitigation techniques is ill suited due to cost, power, and performance overheads. We propose Collision Exposure Factor (CEF), a novel metric to assess the failure vulnerability of circuits processing spatial relationships, including motion planning. CEF is based on the insight that the safety violation probability increases with the surface area of the physical space exposed by a bit-flip. We evaluate CEF on four MPAs. We demonstrate empirically that CEF is correlated with safety violation probability and that CEF-aware selective error mitigation provides 12.3×, 9.6×, and 4.2× lower dangerous Failures-In-Time rate on average for the same amount of protected memory compared to uniform, bit-position, and access-frequency-aware selection of critical data. Furthermore, we show how to employ CEF to enable fault characterization using 23,000× fewer fault injection (FI) experiments than exhaustive FI and evaluate our FI approach on different robots and MPAs. We demonstrate that CEF-aware FI can provide insights on vulnerable bits in an MPA while taking the same amount of time as uniform statistical FI. Finally, we use the CEF to formulate guidelines for designing soft-error resilient MPAs. Deval Shah, Zi Yu Xue, Karthik Pattabiraman, Tor M. Aamodt |
ACM Trans. Cyber Phys. Syst. | 1 |
| 2023 | Learning Label Encodings for Deep Regression
Deval Shah, Tor M. Aamodt |
ICLR | 1 |
| 2023 | Energy-Efficient Realtime Motion PlanningabstractMotion planning is a fundamental problem in autonomous robotics with real-time and low-energy requirements for safe navigation through a dynamic environment. More than 90% of computation time in motion planning is spent on collision detection between the robot and the environment. Several motion planning approaches, such as deep learning-based motion planning, have shown significant improvements in motion planning quality and runtime with ample parallelism available in collision detection. However, naive parallelization of collision detection queries significantly increases computation compared to sequential execution. In this work, we investigate the sources of redundant computations in coarsegrained (inter-collision detection) and fine-grained (intracollision detection) parallelism. We find that the physical spatial locality of obstacles results in redundant computation in coarse-grained parallelism. We further show that the primary sources of redundant computation in fine-grained parallelism are easy cases where objects are far apart or significantly overlapping. Based on these insights, we propose MPAccel to improve the energy efficiency of parallelization in motion planning. MPAccel consists of SAS, a Spatially Aware Scheduler for coarse-grained parallelism, and CECDUs, Cascaded Early-exit Collision Detection Units for fine-grained parallelism. SAS results in 7× speedup using 8× parallelization with 6% increase in the computation compared to 3.7× speedup with 83% increase in computation for naive parallelization. CECDU can perform collision detection in 46 -- 154 cycles for a robot with 6 degrees of freedom. We evaluate MPAccel to execute a state-of-the-art learning-based motion planning algorithm. Our simulations suggest MPAccel can achieve real-time motion planning for a robot with 7 degrees of freedom in 0.014ms-0.49ms with an average latency of 0.099ms compared to 1.42ms on a CPU-GPU system. Deval Shah, Ningfeng Yang, Tor M. Aamodt |
ISCA | 1 |
| 2022 | Label Encoding for Regression Networks
Deval Shah, Zi Yu Xue, Tor M. Aamodt |
ICLR | 1 |
| 2021 | Classroom Digital Twins with Instrumentation-Free Gaze TrackingabstractClassroom sensing is an important and active area of research with great potential to improve instruction. Complementing professional observers – the current best practice – automated pedagogical professional development systems can attend every class and capture fine-grained details of all occupants. One particularly valuable facet to capture is class gaze behavior. For students, certain gaze patterns have been shown to correlate with interest in the material, while for instructors, student-centered gaze patterns have been shown to increase approachability and immediacy. Unfortunately, prior classroom gaze-sensing systems have limited accuracy and often require specialized external or worn sensors. In this work, we developed a new computer-vision-driven system that powers a 3D “digital twin” of the classroom and enables whole-class, 6DOF head gaze vector estimation without instrumenting any of the occupants. We describe our open source implementation, and results from both controlled studies and real-world classroom deployments. Karan Ahuja, Deval Shah, Sujeath Pareddy, Franceska Xhakaj, Amy Ogan, Yuvraj Agarwal, Chris Harrison 0001 |
CHI | 2 |
| 2019 | EDGE: Event-Driven GPU ExecutionabstractGPUs are known to benefit structured applications with ample parallelism, such as deep learning in a datacenter. Recently, GPUs have shown promise for irregular streaming network tasks. However, the GPU's co-processor dependence on a CPU for task management, inefficiencies with fine-grained tasks, and limited multiprogramming capabilities introduce challenges with efficiently supporting latency-sensitive streaming tasks. This paper proposes an event-driven GPU execution model, EDGE, that enables non-CPU devices to directly launch preconfigured tasks on a GPU without CPU interaction. Along with freeing up the CPU to work on other tasks, we estimate that EDGE can reduce the kernel launch latency by 4.4xcompared to the baseline CPU-launched approach. This paper also proposes a warp-level preemption mechanism to further reduce the end-to-end latency of fine-grained tasks in a shared GPU environment. We evaluate multiple optimizations that reduce the average warp preemption latency by 35.9x over waiting for a preempted warp to naturally flush the pipeline. When compared to waiting for the first available resources, we find that warp-level preemption reduces the average and tail warp scheduling latencies by 2.6x and 2.9x, respectively, and improves the average normalized turnaround time by 1.4x. Tayler H. Hetherington, Maria Lubeznov, Deval Shah, Tor M. Aamodt |
PACT | 3 |
| 2019 | Analyzing Machine Learning Workloads Using a Detailed GPU SimulatorabstractMachine learning (ML) has recently emerged as an important application driving future architecture design. Traditionally, architecture research has used detailed simulators to model and measure the impact of proposed changes. However, current open-source, publicly available simulators lack support for running a full ML stack like PyTorch. High-confidence, cycle-accurate simulations are crucial for architecture research and without them, it is difficult to rapidly prototype new ideas. In this paper, we describe changes we made to GPGPU-Sim, a popular, widely used GPU simulator, to run ML applications that use cuDNN and PyTorch, two widely used frameworks for running Deep Neural Networks (DNNs). This work has the potential to enable significant microarchitectural research into GPUs for DNNs. Our results show that the modified simulator, which has been made publicly available with this paper1Source code available at https://github.com/gpgpu-sim/gpgpu-sim_distribution (dev branch), provides execution time results within 18% of real hardware. We further use it to study other ML workloads and demonstrate how the simulator identifies opportunities for architectural optimization that prior tools are unable to provide. Jonathan S. Lew, Deval Shah, Suchita Pati, Shaylin Cattell, Mengchi Zhang, Amruth Sandhupatla, Christopher Ng, Negar Goli, Matthew D. Sinclair, Timothy G. Rogers, Tor M. Aamodt |
ISPASS | 2 |