Sandeep Chinchali

dblp:85/8366 · also Sandeep P. Chinchali · DBLP profile ↗
← Back
42ranked-venue papers
5as first author
32since 2021 · last 2026
0000-0002-0601-3633ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 4 first-author · 25 since 2021Systems, architecture and hardware · 8 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 7 since 2021Computer networks · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021
YearPublicationVenuePosition
2026 NeuS-QA: Grounding Long-Form Video Understanding in Temporal Logic and Neuro-Symbolic Reasoning
abstract
While vision-language models (VLMs) excel at tasks involving single images or short videos, they still struggle with Long Video Question Answering (LVQA) due to its demand for complex multi-step temporal reasoning. Vanilla approaches, which simply sample frames uniformly and feed them to a VLM along with the question, incur significant token overhead. This forces aggressive downsampling of long videos, causing models to miss fine-grained visual structure, subtle event transitions, and key temporal cues. Recent works attempt to overcome these limitations through heuristic approaches; however, they lack explicit mechanisms for encoding temporal relationships and fail to provide any formal guarantees that the sampled context actually encodes the compositional or causal logic required by the question. To address these foundational gaps, we introduce NeuS-QA, a training-free, plug-and-play neuro-symbolic pipeline for LVQA. NeuS-QA first translates a natural language question into a logic specification that models the temporal relationship between frame-level events. Next, we construct a video automaton to model the video's frame-by-frame event progression, and finally employ model checking to compare the automaton against the specification to identify all video segments that satisfy the question's logical requirements. Only these logic-verified segments are submitted to the VLM, thus improving interpretability, reducing hallucinations, and enabling compositional reasoning without modifying or fine-tuning the model. Experiments on the LongVideoBench and CinePile benchmarks show that NeuS-QA significantly improves performance by over 10%, particularly on questions involving event ordering, causality, and multi-step reasoning.
Sahil Shah, S. P. Sharan, Harsh Goel, Minkyu Choi 0001, Mustafa Munir, Manvik Pasula, Radu Marculescu, Sandeep Chinchali
AAAI8
2025 PPGWeaver: Diffusion-Augmented Models for Real-Time Heart Rate Estimation on Microcontrollers
abstract
We present a compact, deployable heart rate (HR) estimation system using photoplethysmography (PPG) and inertial measurement unit (IMU) data, combining TimeWeaver, a conditional diffusion model for metadata-aware synthetic augmentation, with progressive structured pruning of Temporal Convolutional Networks (TCNs). Our smallest model, with 1.56k parameters, achieves a mean absolute error (MAE) of 4.92 BPM on the PPG-DaLiA dataset and supports real-time inference ($<{40 ms}$latency) on a 64 MHz ARM Cortex-M4F microcontroller (MCU) without requiring quantization. Synthetic data conditioned on subject metadata, HR, and activity type significantly enhances model generalization, enabling pruned models to match or exceed the accuracy of larger baselines, achieving over a 23% improvement compared to training on real data alone. Our work establishes a new Pareto frontier for real-time, on-device HR monitoring using diffusion-augmented training and sub-2 k parameter models.
Vinayak Narasimhan, Raimi Shah, Shubhankar Agarwal, Sai Shankar Narasimhan, Shailabh Kumar, Sang Kyu Kim, Sandeep Chinchali, Radwanul Hasan Siddique
BSN7
2025 Neuro-Symbolic Evaluation of Text-to-Video Models using Formal Verification
abstract
Recent advancements in text-to-video models such as Sora, Gen-3, MovieGen, and CogVideoX are pushing the boundaries of synthetic video generation, with adoption seen in fields like robotics, autonomous driving, and entertainment. As these models become prevalent, various metrics and benchmarks have emerged to evaluate the quality of the generated videos. However, these metrics emphasize visual quality and smoothness, neglecting temporal fidelity and text-to-video alignment, which are crucial for safety-critical applications. To address this gap, we introduce NeuS-V, a novel synthetic video evaluation metric that rigorously assesses text-to-video alignment using neuro-symbolic formal verification techniques. Our approach first converts the prompt into a formally defined Temporal Logic (TL) specification and translates the generated video into an automaton representation. Then, it evaluates the text-to-video alignment by formally checking the video automaton against the TL specification. Furthermore, we present a dataset of temporally extended prompts to evaluate state-of-the-art video generation models against our benchmark. We find that NeuS-V demonstrates a higher correlation by over 5× with human evaluations when compared to existing metrics. Our evaluation further reveals that current video generation models perform poorly on these temporally complex prompts, highlighting the need for future work in improving text-to-video generation capabilities. We open-source our benchmark, code, and dataset at utaustin-swarmlab.github.io/neusv.
S. P. Sharan, Minkyu Choi 0001, Sahil Shah, Harsh Goel, Mohd. Omama, Sandeep Chinchali
CVPR6
2025 Fair Resource Allocation for Fleet Intelligence
Oguzhan Baser, Kaan Kale, Po-han Li, Sandeep Chinchali
GLOBECOM4
2025 CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features
abstract
Multimodal encoders like CLIP excel in tasks such as zero-shot image classification and cross-modal retrieval. However, they require excessive training data. We propose canonical similarity analysis (CSA), which uses two unimodal encoders to replicate multimodal encoders using limited data. CSA maps unimodal features into a multimodal space, using a new similarity score to retain only the multimodal information. CSA only involves the inference of unimodal encoders and a cubic-complexity matrix decomposition, eliminating the need for extensive GPU-based model training. Experiments show that CSA outperforms CLIP while requiring $50$,$000\times$ fewer multimodal data pairs to bridge the modalities given pre-trained unimodal encoders on ImageNet classification and misinformative news caption detection. CSA surpasses the state-of-the-art method to map unimodal features to multimodal features. We also demonstrate the ability of CSA with modalities beyond image and text, paving the way for future modality pairs with limited paired multimodal data but abundant unpaired unimodal data, such as LiDAR and text.
Po-han Li, Sandeep Chinchali, Ufuk Topcu
ICLR2
2025 Exploiting Distribution Constraints for Scalable and Efficient Image Retrieval
abstract
Image retrieval is crucial in robotics and computer vision, with downstream applications in robot place recognition and vision-based product recommendations. Modern retrieval systems face two key challenges: scalability and efficiency. State-of-the-art image retrieval systems train specific neural networks for each dataset, an approach that lacks scalability. Furthermore, since retrieval speed is directly proportional to embedding size, existing systems that use large embeddings lack efficiency. To tackle scalability, recent works propose using off-the-shelf foundation models. However, these models, though applicable across datasets, fall short in achieving performance comparable to that of dataset-specific models. Our key observation is that, while foundation models capture necessary subtleties for effective retrieval, the underlying distribution of their embedding space can negatively impact cosine similarity searches. We introduce Autoencoders with Strong Variance Constraints (AE-SVC), which, when used for projection, significantly improves the performance of foundation models. We provide an in-depth theoretical analysis of AE-SVC. Addressing efficiency, we introduce Single-Shot Similarity Space Distillation ((SS)2D), a novel approach to learn embeddings with adaptive sizes that offers a better trade-off between size and performance. We conducted extensive experiments on four retrieval datasets, including Stan- ford Online Products (SoP) and Pittsburgh30k, using four different off-the-shelf foundation models, including DinoV2 and CLIP. AE-SVC demonstrates up to a 16% improvement in retrieval performance, while (SS)2D shows a further 10% improvement for smaller embedding sizes.
Mohd. Omama, Po-han Li, Sandeep Chinchali
ICLR3
2025 R3DM: Enabling Role Discovery and Diversity Through Dynamics Models in Multi-agent Reinforcement Learning
abstract
Multi-agent reinforcement learning (MARL) has achieved significant progress in large-scale traffic control, autonomous vehicles, and robotics. Drawing inspiration from biological systems where roles naturally emerge to enable coordination, role-based MARL methods have been proposed to enhance cooperation learning for complex tasks. However, existing methods exclusively derive roles from an agent’s past experience during training, neglecting their influence on its future trajectories. This paper introduces a key insight: an agent’s role should shape its future behavior to enable effective coordination. Hence, we propose Role Discovery and Diversity through Dynamics Models (R3DM), a novel role-based MARL framework that learns emergent roles by maximizing the mutual information between agents’ roles, observed trajectories, and expected future behaviors. R3DM optimizes the proposed objective through contrastive learning on past trajectories to first derive intermediate roles that shape intrinsic rewards to promote diversity in future behaviors across different roles through a learned dynamics model. Benchmarking on SMAC and SMACv2 environments demonstrates that R3DM outperforms state-of-the-art MARL approaches, improving multi-agent coordination to increase win rates by up to 20%. The code is available at https://github.com/UTAustin-SwarmLab/R3DM.
Harsh Goel, Mohd. Omama, Behdad Chalaki, Vaishnav Tadiparthi, Ehsan Moradi-Pari, Sandeep Chinchali
ICML6
2025 Human-Agent Coordination in Games under Incomplete Information via Multi-Step Intent
Shenghui Chen, Ruihan Zhao 0001, Sandeep Chinchali, Ufuk Topcu
AAMAS3
2025 WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
abstract
Speech embeddings often retain sensitive attributes such as speaker identity, accent, or demographic information, posing risks in biased model training and privacy leakage. We propose WavShape, an information-theoretic speech representation learning framework that optimizes embeddings for fairness and privacy while preserving task-relevant information. We leverage mutual information (MI) estimation using the Donsker-Varadhan formulation to guide an MI-based encoder that systematically filters sensitive attributes while maintaining speech content essential for downstream tasks. Experimental results on three known datasets show that WavShape reduces MI between embeddings and sensitive attributes by up to 81% while retaining 97% of task-relevant information. By integrating information theory with self-supervised speech models, this work advances the development of fair, privacy-aware, and resource-efficient speech systems.
Oguzhan Baser, Ahmet Ege Tanriverdi, Kaan Kale, Sandeep Chinchali, Sriram Vishwanath
INTERSPEECH4
2025 PhonemeFake: Redefining Deepfake Realism with Language-Driven Segmental Manipulation and Adaptive Bilevel Detection
abstract
Deepfake (DF) attacks pose a growing threat as generative models become increasingly advanced. However, our study reveals that existing DF datasets fail to deceive human perception, unlike real DF attacks that influence public discourse. It highlights the need for more realistic DF attack vectors. We introduce PhonemeFake (PF), a DF attack that manipulates critical speech segments using language reasoning, significantly reducing human perception by up to 42% and benchmark accuracies by up to 94%. We release an easy-to-use PF dataset on HuggingFace and open-source bilevel DF segment detection model that adaptively prioritizes compute on manipulated regions. Our extensive experiments across three known DF datasets reveal that our detection model reduces EER by 91% while achieving up to 90% speed-up, with minimal compute overhead and precise localization beyond existing models as a scalable solution.
Oguzhan Baser, Ahmet Ege Tanriverdi, Sriram Vishwanath, Sandeep Chinchali
INTERSPEECH4
2025 VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR
abstract
Many decision-making tasks, where both accuracy and efficiency matter, still require human supervision. For example, tasks like traffic officers reviewing hour-long dashcam footage or researchers screening conference videos can benefit from concise summaries that reduce cognitive load and save time. Yet current vision-language models (VLMs) often produce verbose, redundant outputs that hinder task performance. Existing video caption evaluation depends on costly human annotations and overlooks the summaries' utility in downstream tasks. We address these gaps with $\underline{\textbf{V}}$ideo-to-text $\underline{\textbf{I}}$nformation $\underline{\textbf{B}}$ottleneck $\underline{\textbf{E}}$valuation (VIBE), an annotation-free method that scores VLM outputs using two metrics: $\textit{grounding}$ (how well the summary aligns with visual content) and $\textit{utility}$ (how informative it is for the task). VIBE selects from randomly sampled VLM outputs by ranking them according to the two scores to support effective human decision-making. Human studies on $\texttt{LearningPaper24}$, $\texttt{SUTD-TrafficQA}$, and $\texttt{LongVideoBench}$ show that summaries selected by VIBE consistently improve performance—boosting task accuracy by up to $61.23$% and reducing response time by $75.77$% compared to naive VLM summaries or raw video.
Shenghui Chen, Po-han Li, Sandeep Chinchali, Ufuk Topcu
NeurIPS3
2025 Constrained Posterior Sampling: Time Series Generation with Hard Constraints
abstract
Generating realistic time series samples is crucial for stress-testing models and protecting user privacy by using synthetic data. In engineering and safety-critical applications, these samples must meet certain hard constraints that are domain-specific or naturally imposed by physics or nature. Consider, for example, generating electricity demand patterns with constraints on peak demand times. This can be used to stress-test the functioning of power grids during adverse weather conditions. Existing approaches for generating constrained time series are either not scalable or degrade sample quality. To address these challenges, we introduce Constrained Posterior Sampling (CPS), a diffusion-based sampling algorithm that aims to project the posterior mean estimate into the constraint set after each denoising update. Notably, CPS scales to a large number of constraints ($\sim100$) without requiring additional training. We provide theoretical justifications highlighting the impact of our projection step on sampling. Empirically, CPS outperforms state-of-the-art methods in sample quality and similarity to real time series by around 70\% and 22\%, respectively, on real-world stocks, traffic, and air quality datasets.
Sai Shankar Narasimhan, Shubhankar Agarwal, Litu Rout, Sanjay Shakkottai, Sandeep Chinchali
NeurIPS5
2025 Synthetic data-augmented explainable Vision Transformer for colorectal cancer diagnosis via surface tactile imaging
Siddhartha Kapuria, Naruhiko Ikoma, Sandeep Chinchali, Farshid Alambeigi
Eng. Appl. Artif. Intell.3
2025 CoordLight: Learning Decentralized Coordination for Network-Wide Traffic Signal Control
abstract
Adaptive traffic signal control (ATSC) is crucial in alleviating congestion, maximizing throughput and promoting sustainable mobility in ever-expanding cities. Multi-Agent Reinforcement Learning (MARL) has recently shown significant potential in addressing complex traffic dynamics, but the intricacies of partial observability and coordination in decentralized environments still remain key challenges in formulating scalable and efficient control strategies. To address these challenges, we present CoordLight, a MARL-based framework designed to improve intra-neighborhood traffic by enhancing decision-making at individual junctions (agents), as well as coordination with neighboring agents, thereby scaling up to network-level traffic optimization. Specifically, we introduce the Queue Dynamic State Encoding (QDSE), a novel state representation based on vehicle queuing models, which strengthens the agents’ capability to analyze, predict, and respond to local traffic dynamics. We further propose an advanced MARL algorithm, named Neighbor-aware Policy Optimization (NAPO). It integrates an attention mechanism that discerns the state and action dependencies among adjacent agents, aiming to facilitate more coordinated decision-making, and to improve policy learning updates through robust advantage calculation. This enables agents to identify and prioritize crucial interactions with influential neighbors, thus enhancing the targeted coordination and collaboration among agents. Through comprehensive evaluations against state-of-the-art traffic signal control methods over three real-world traffic datasets composed of up to 196 intersections, we empirically show that CoordLight consistently exhibits superior performance across diverse traffic networks with varying traffic flows.
Harsh Goel, Peizhuo Li, Mehul Damani, Sandeep Chinchali, Guillaume Sartoretti
IEEE Trans. Intell. Transp. Syst.5
2024 Reduce, Reuse, Recycle: Categories for Compositional Reinforcement Learning
abstract
In reinforcement learning, conducting task composition by forming cohesive, executable sequences from multiple tasks remains challenging. However, the ability to (de)compose tasks is a linchpin in developing robotic systems capable of learning complex behaviors. Yet, compositional reinforcement learning is beset with difficulties, including the high dimensionality of the problem space, scarcity of rewards, and absence of system robustness after task composition. To surmount these challenges, we view task composition through the prism of category theory—a mathematical discipline exploring structures and their compositional relationships. The categorical properties of Markov decision processes untangle complex tasks into manageable sub-tasks, allowing for strategical reduction of dimensionality, facilitating more tractable reward structures, and bolstering system robustness. Experimental results support the categorical theory of reinforcement learning by enabling skill reduction, reuse, and recycling when learning complex robotic arm tasks.
Georgios Bakirtzis, Michail Savvas, Ruihan Zhao 0001, Sandeep Chinchali, Ufuk Topcu
ECAI4
2024 Towards Neuro-Symbolic Video Understanding
Minkyu Choi 0001, Harsh Goel, Mohd. Omama, Yunhao Yang, Sahil Shah, Sandeep Chinchali
ECCV (78)6
2024 Time Weaver: A Conditional Time Series Generation Model
abstract
Imagine generating a city’s electricity demand pattern based on weather, the presence of an electric vehicle, and location, which could be used for capacity planning during a winter freeze. Such real-world time series are often enriched with paired heterogeneous contextual metadata (e.g., weather and location). Current approaches to time series generation often ignore this paired metadata. Additionally, the heterogeneity in metadata poses several practical challenges in adapting existing conditional generation approaches from the image, audio, and video domains to the time series domain. To address this gap, we introduce TIME WEAVER, a novel diffusion-based model that leverages the heterogeneous metadata in the form of categorical, continuous, and even time-variant variables to significantly improve time series generation. Additionally, we show that naive extensions of standard evaluation metrics from the image to the time series domain are insufficient. These metrics do not penalize conditional generation approaches for their poor specificity in reproducing the metadata-specific features in the generated time series. Thus, we innovate a novel evaluation metric that accurately captures the specificity of conditional generation and the realism of the generated time series. We show that TIME WEAVER outperforms state-of-the-art benchmarks, such as Generative Adversarial Networks (GANs), by up to 30% in downstream classification tasks on real-world energy, medical, air quality, and traffic datasets.
Sai Shankar Narasimhan, Shubhankar Agarwal, Oguzhan Akcin, Sujay Sanghavi, Sandeep Chinchali
ICML5
2024 SecureSpectra: Safeguarding Digital Identity from Deep Fake Threats via Intelligent Signatures
Oguzhan Baser, Kaan Kale, Sandeep Chinchali
INTERSPEECH3
2024 Robot-Enabled Machine Learning-Based Diagnosis of Gastric Cancer Polyps Using Partial Surface Tactile Imaging
abstract
In this paper, to collectively address the existing limitations on endoscopic diagnosis of Advanced Gastric Cancer (AGC) Tumors, for the first time, we propose (i) utilization and evaluation of our recently developed Vision-based Tactile Sensor (VTS), and (ii) a complementary Machine Learning (ML) algorithm for classifying tumors using their textural features. Leveraging a seven DoF robotic manipulator and unique custom-designed and additively-manufactured realistic AGC tumor phantoms, we demonstrated the advantages of automated data collection using the VTS addressing the problem of data scarcity and biases encountered in traditional ML-based approaches. Our synthetic-data-trained ML model was successfully evaluated and compared with traditional ML models utilizing various statistical metrics even under mixed morphological characteristics and partial sensor contact.
Siddhartha Kapuria, Jeff Bonyun, Yash Kulkarni, Naruhiko Ikoma, Sandeep Chinchali, Farshid Alambeigi
IROS5
2024 PEERNet: An End-to-End Profiling Tool for Real-Time Networked Robotic Systems
abstract
Networked robotic systems balance compute, power, and latency constraints in applications such as self-driving vehicles, drone swarms, and teleoperated surgery. A core problem in this domain is deciding when to offload a computationally expensive task to the cloud, a remote server, at the cost of communication latency. Task offloading algorithms often rely on precise knowledge of system-specific performance metrics, such as sensor data rates, network bandwidth, and machine learning model latency. While these metrics can be modeled during system design, uncertainties in connection quality, server load, and hardware conditions introduce real-time performance variations, hindering overall performance. We introduce PEERNet, an end-to-end and real-time profiling tool for cloud robotics. PEERNet enables performance monitoring on heterogeneous hardware through targeted yet adaptive profiling of system components such as sensors, networks, deep-learning pipelines, and devices. We showcase PEERNet’s capabilities through networked robotics tasks, such as image-based teleoperation of a Franka Emika Panda arm and querying vision language models using an Nvidia Jetson Orin. PEERNet reveals non-intuitive behavior in robotic systems, such as asymmetric network transmission and bimodal language model output. Our evaluation underscores the effectiveness and importance of benchmarking in networked robotics, demonstrating PEERNet’s adaptability. Our code is open-source and available at github.com/UTAustin-SwarmLab/PEERNet.
Aditya Narayanan, Pranav Kasibhatla, Minkyu Choi 0001, Po-han Li, Ruihan Zhao 0001, Sandeep Chinchali
IROS6
2024 Joint learning of reward machines and policies in environments with partially known semantics
abstract
We study the problem of reinforcement learning for a task encoded by a reward machine. The task is defined over a set of properties in the environment, called atomic propositions, and represented by Boolean variables. One unrealistic assumption commonly used in the literature is that the truth values of these propositions are accurately known. In real situations, however, these truth values are uncertain since they come from sensors that suffer from imperfections. At the same time, reward machines can be difficult to model explicitly, especially when they encode complicated tasks. We develop a reinforcement-learning algorithm that infers a reward machine that encodes the underlying task while learning how to execute it, despite the uncertainties of the propositions' truth values. In order to address such uncertainties, the algorithm maintains a probabilistic estimate about the truth value of the atomic propositions; it updates this estimate according to new sensory measurements that arrive from exploration of the environment. Additionally, the algorithm maintains a hypothesis reward machine, which acts as an estimate of the reward machine that encodes the task to be learned. As the agent explores the environment, the algorithm updates the hypothesis reward machine according to the obtained rewards and the estimate of the atomic propositions' truth value. Finally, the algorithm uses a Q-learning procedure for the states of the hypothesis reward machine to determine an optimal policy that accomplishes the task. We prove that the algorithm successfully infers the reward machine and asymptotically learns a policy that accomplishes the respective task.
Christos K. Verginis, Cevahir Köprülü, Sandeep Chinchali, Ufuk Topcu
Artif. Intell.3
2024 Forecaster-Aided User Association and Load Balancing in Multi-Band Mobile Networks
abstract
Cellular networks are becoming increasingly heterogeneous with higher base station (BS) densities and ever more frequency bands, making BS selection and band assignment key decisions in terms of user service rate and coverage. In this paper, we decompose the mobility-aware user association task into (i) forecasting of user data rate and then (ii) convex utility maximization for user association accounting for the effects of BS load and handover overheads. Using a linear combination of normalized mean-squared error (NMSE) and normalized discounted cumulative gain (NDCG) as a novel loss function, a recurrent deep neural network is trained to reliably forecast the mobile users’ future data rates. Based on the forecast, the controller optimizes the association decisions to maximize the service rate-based network utility using our computationally efficient (speed up of 100× versus generic convex solver) algorithm based on the Frank-Wolfe method. Using an industry-grade network simulator developed by Meta, we show that the proposed model predictive control (MPC) approach improves the 5th percentile service rate by 3.5× compared to the traditional signal strength-based association, reduces the median number of handovers by 7× compared to a handover agnostic strategy, and achieves service rates close to a genie-aided scheme. Furthermore, our model-based approach is significantly more sample-efficient (needs 100× less training data) compared to model-free reinforcement learning (RL), and generalizes well across different user drop scenarios.
Manan Gupta, Sandeep Chinchali, Paul Parayil Varkey, Jeffrey G. Andrews
IEEE Trans. Wirel. Commun.2
2023 Learning-Based Model Predictive Control for User Association in Multi-Band Mobile Networks
abstract
As cellular networks embrace heterogeneity with higher base station (BS) densities and ever more frequency bands, BS selection and band assignment become increasingly key decisions in terms of rate and coverage optimization. In this paper, we propose a novel learning-based model predictive control (MPC) approach for BS selection and band assignment while accounting for user mobility. We formulate a convex utility maximization problem that accounts for the effects of BS load and handover overheads on the user's service rate. Using a linear combination of normalized mean-squared error (NMSE) and$\mathbf{top}-m$loss as a novel loss function, a recurrent deep neural network is trained to reliably forecast the mobile users' future rates. The MPC controller then uses this forecast to optimize the association decisions to maximize the service rate-based network utility. Using an industry-grade network simulator developed by Meta, we show that the proposed approach improves the 5th percentile service rate by$2.7\times$compared to the traditional signal strength-based association and its performance approaches that of a genie-aided scheme in terms of the achieved service rate and the number of handovers triggered.
Manan Gupta, Sandeep Chinchali, Paul Parayil Varkey, Jeffrey G. Andrews
ICC2
2023 Robust Forecasting for Robotic Control: A Game-Theoretic Approach
abstract
Modern robots require accurate forecasts to make optimal decisions in the real world. For example, self-driving cars need an accurate forecast of other agents' future actions to plan safe trajectories. Current methods rely heavily on historical time series to accurately predict the future. However, relying entirely on the observed history is problematic since it could be corrupted by noise, have outliers, or not completely represent all possible outcomes. To solve this problem, we propose a novel framework for generating robust forecasts for robotic control. In order to model real-world factors affecting future forecasts, we introduce the notion of an adversary, which perturbs observed historical time series to increase a robot's ultimate control cost. Specifically, we model this interaction as a zero-sum two-player game between a robot's forecaster and this hypothetical adversary. We show that our proposed game may be solved to a local Nash equilibrium using gradient-based optimization techniques. Furthermore, we show that a forecaster trained with our method performs 30.14% better on out-of-distribution real-world lane change data than baselines.
Shubhankar Agarwal, David Fridovich-Keil, Sandeep Chinchali
ICRA3
2023 Task-aware Network Coding over Butterfly Network
abstract
Network coding allows distributed information sources such as sensors to efficiently compress and transmit data to distributed receivers across a bandwidth-limited network. Classical network coding is largely task-agnostic – the coding schemes mainly aim to faithfully reconstruct data at the receivers, regardless of what ultimate task the received data is used for. In this paper, we analyze a new task-driven network coding problem, where distributed receivers pass transmitted data through machine learning (ML) tasks, which provides an opportunity to improve efficiency by transmitting salient task-relevant data representations. Specifically, we formulate a task-aware network coding problem over a butterfly network in real-coordinate space, where lossy analog compression through principal component analysis (PCA) can be applied. A lower bound for the total loss function for the formulated problem is given, and necessary and sufficient conditions for achieving this lower bound are also provided. We introduce ML algorithms to solve the problem in the general case, and our evaluation demonstrates the effectiveness of task-aware network coding on distributed image classification tasks on MNIST, CIFAR10, CIFAR100, and satellite imagery datasets.
Jiangnan Cheng, Sandeep Chinchali, Ao Tang
ISIT2
2023 Co-Designing Socio-Technical Interventions with Skilled Trade Workers
abstract
This paper lays out an approach to co-designing Artificial Intelligence (AI)-enhanced smart hand tools with skilled trade workers employed at a local municipality. Skilled trade workers contribute to society by building the infrastructure upon which the public depends. In addition, these technical interventions offer an opportunity for workers to benefit from the data they generate via smart hand tools, potentially creating a new empowerment dynamic with employers. Therefore, we consider technologies that support skilled trade workers in performing their work effectively, safely, and with increased levels of autonomy to be considered Public Interest Technologies (PIT). Interdisciplinary research is underway that aligns these approaches with Public Interest Design (PID) principles which informs researchers' desire to explore how emerging technologies - data cooperatives, blockchain, smart contracts, and data dividends - can further smart hand tools' empowerment dynamic. Future participatory design efforts may deliver additional insights and further impact technology deployment.
Chelsea Collier, Kenneth R. Fleischmann, Tina Lassiter, Sherri R. Greenberg, Raul G. Longoria, Sandeep Chinchali
ISTAS6
2023 Task-aware Distributed Source Coding under Dynamic Bandwidth
abstract
Efficient compression of correlated data is essential to minimize communication overload in multi-sensor networks. In such networks, each sensor independently compresses the data and transmits them to a central node. A decoder at the central node decompresses and passes the data to a pre-trained machine learning-based task model to generate the final output. Due to limited communication bandwidth, it is important for the compressor to learn only the features that are relevant to the task. Additionally, the final performance depends heavily on the total available bandwidth. In practice, it is common to encounter varying availability in bandwidth. Since higher bandwidth results in better performance, it is essential for the compressor to dynamically take advantage of the maximum available bandwidth at any instant. In this work, we propose a novel distributed compression framework composed of independent encoders and a joint decoder, which we call neural distributed principal component analysis (NDPCA). NDPCA flexibly compresses data from multiple sources to any available bandwidth with a single model, reducing compute and storage overhead. NDPCA achieves this by learning low-rank task representations and efficiently distributing bandwidth among sensors, thus providing a graceful trade-off between performance and bandwidth. Experiments show that NDPCA improves the success rate of multi-view robotic arm manipulation by 9% and the accuracy of object detection tasks on satellite imagery by 14% compared to an autoencoder with uniform bandwidth allocation.
Po-han Li, Sravan Kumar Ankireddy, Ruihan Zhao 0001, Hossein Nourkhiz Mahjoub, Ehsan Moradi-Pari, Ufuk Topcu, Sandeep Chinchali, Hyeji Kim
NeurIPS7
2022 Task-aware Privacy Preservation for Multi-dimensional Data
abstract
Local differential privacy (LDP) can be adopted to anonymize richer user data attributes that will be input to sophisticated machine learning (ML) tasks. However, today’s LDP approaches are largely task-agnostic and often lead to severe performance loss – they simply inject noise to all data attributes according to a given privacy budget, regardless of what features are most relevant for the ultimate task. In this paper, we address how to significantly improve the ultimate task performance with multi-dimensional user data by considering a task-aware privacy preservation problem. The key idea is to use an encoder-decoder framework to learn (and anonymize) a task-relevant latent representation of user data. We obtain an analytical near-optimal solution for the linear setting with mean-squared error (MSE) task loss. We also provide an approximate solution through a gradient-based learning algorithm for general nonlinear cases. Extensive experiments demonstrate that our task-aware approach significantly improves ultimate task accuracy compared to standard benchmark LDP approaches with the same level of privacy guarantee.
Jiangnan Cheng, Ao Tang, Sandeep Chinchali
ICML3
2022 Drift Reduced Navigation with Deep Explainable Features
abstract
Modern autonomous vehicles (AVs) often rely on vision, LIDAR, and even radar-based simultaneous localization and mapping (SLAM) frameworks for precise localization and navigation. However, modern SLAM frameworks often lead to unacceptably high levels of drift (i.e., localization error) when AVs observe few visually distinct features or encounter occlusions due to dynamic obstacles. This paper argues that minimizing drift must be a key desiderata in AV motion planning, which requires an AV to take active control decisions to move towards feature-rich regions while also minimizing conventional control cost. To do so, we first introduce a novel data-driven perception module that observes LIDAR point clouds and estimates which features/regions an AV must navigate towards for drift minimization. Then, we introduce an interpretable model predictive controller (MPC) that moves an AV toward such feature-rich regions while avoiding visual occlusions and gracefully trading off drift and control cost. Our experiments on challenging, dynamic scenarios in the state-of-the-art CARLA simulator indicate our method reduces drift up to 76.76% compared to benchmark approaches.
Mohd. Omama, Sundar Sripada V. S., Sandeep Chinchali, Arun Kumar Singh 0001, K. Madhava Krishna
IROS3
2022 Class-Aware Adversarial Transformers for Medical Image Segmentation
abstract
Transformers have made remarkable progress towards modeling long-range dependencies within the medical image analysis domain. However, current transformer-based models suffer from several disadvantages: (1) existing methods fail to capture the important features of the images due to the naive tokenization scheme; (2) the models suffer from information loss because they only consider single-scale feature representations; and (3) the segmentation label maps generated by the models are not accurate enough without considering rich semantic contexts and anatomical textures. In this work, we present CASTformer, a novel type of adversarial transformers, for 2D medical image segmentation. First, we take advantage of the pyramid structure to construct multi-scale representations and handle multi-scale variations. We then design a novel class-aware transformer module to better learn the discriminative regions of objects with semantic structures. Lastly, we utilize an adversarial training strategy that boosts segmentation accuracy and correspondingly allows a transformer-based discriminator to capture high-level semantically correlated contents and low-level anatomical features. Our experiments demonstrate that CASTformer dramatically outperforms previous state-of-the-art transformer-based approaches on three benchmarks, obtaining 2.54%-5.88% absolute improvements in Dice over previous models. Further qualitative experiments provide a more detailed picture of the model’s inner workings, shed light on the challenges in improved transparency, and demonstrate that transfer learning can greatly improve performance and reduce the size of medical image datasets in training, making CASTformer a strong starting point for downstream medical image analysis tasks.
Chenyu You, Ruihan Zhao 0001, Siyuan Dong, Sandeep Chinchali, Ufuk Topcu, Lawrence H. Staib, James S. Duncan
NeurIPS5
2021 Interpretable Trade-offs Between Robot Task Accuracy and Compute Efficiency
abstract
A robot can invoke heterogeneous computation resources such as CPUs, cloud GPU servers, or even human computation for achieving a high-level goal. The problem of invoking an appropriate computation model so that it will successfully complete a task while keeping its compute and energy costs within a budget is called a model selection problem. In this paper, we present an optimal solution to the model selection problem with two compute models, the first being fast but less accurate, and the second being slow but more accurate. The main insight behind our solution is that a robot should invoke the slower compute model only when the benefits from the gain in accuracy outweigh the computational costs. We show that such cost-benefit analysis can be performed by leveraging the statistical correlation between the accuracy of fast and slow compute models. We demonstrate the broad applicability of our approach to diverse problems such as perception using neural networks and safe navigation of a simulated Mars rover.
Bineet Ghosh, Sandeep Chinchali, Parasara Sridhar Duggirala
IROS2
2021 Data Sharing and Compression for Cooperative Networked Control
abstract
Sharing forecasts of network timeseries data, such as cellular or electricity load patterns, can improve independent control applications ranging from traffic scheduling to power generation. Typically, forecasts are designed without knowledge of a downstream controller's task objective, and thus simply optimize for mean prediction error. However, such task-agnostic representations are often too large to stream over a communication network and do not emphasize salient temporal features for cooperative control. This paper presents a solution to learn succinct, highly-compressed forecasts that are co-designed with a modular controller's task objective. Our simulations with real cellular, Internet-of-Things (IoT), and electricity load data show we can improve a model predictive controller's performance by at least 25% while transmitting 80% less data than the competing method. Further, we present theoretical compression results for a networked variant of the classical linear quadratic regulator (LQR) control problem.
Jiangnan Cheng, Marco Pavone 0001, Sachin Katti, Sandeep Chinchali, Ao Tang
NeurIPS4
2020 Multi-agent Reinforcement Learning for Networked System Control
Sandeep Chinchali, Sachin Katti
ICLR2
2018 Cellular Network Traffic Scheduling With Deep Reinforcement Learning
abstract
Modern mobile networks are facing unprecedented growth in demand due to a new class of traffic from Internet of Things (IoT) devices such as smart wearables and autonomous cars. Future networks must schedule delay-tolerant software updates, data backup, and other transfers from IoT devices while maintaining strict service guarantees for conventional real-time applications such as voice-calling and video. This problem is extremely challenging because conventional traffic is highly dynamic across space and time, so its performance is significantly impacted if all IoT traffic is scheduled immediately when it originates. In this paper, we present a reinforcement learning (RL) based scheduler that can dynamically adapt to traffic variation, and to various reward functions set by network operators, to optimally schedule IoT traffic. Using 4 weeks of real network data from downtown Melbourne, Australia spanning diverse traffic patterns, we demonstrate that our RL scheduler can enable mobile networks to carry 14.7% more data with minimal impact on existing traffic, and outpeforms heuristic schedulers by more than 2x. Our work is a valuable step towards designing autonomous, "self-driving" networks that learn to manage themselves from past data.
Sandeep Chinchali, Pan Hu 0003, Manu Sharma, Manu Bansal, Rakesh Misra, Marco Pavone 0001, Sachin Katti
AAAI1
2018 Neural Networks Meet Physical Networks: Distributed Inference Between Edge Devices and the Cloud
abstract
We believe that most future video uploaded over the network will be consumed by machines for sensing tasks such as automated surveillance and mapping rather than for human consumption. Today's systems typically collect raw data from distributed sensors, such as drones, with the computer vision logic implemented in the cloud using deep neural networks (DNNs). They use standard video encoding techniques, send it over the network, and then decompress it at the cloud before using the vision DNN. In other words, data encoding and distribution is decoupled from the sensing goal. This is bandwidth inefficient because video encoding schemes, such as MPEG4, might send data tailored for human perception but irrelevant for the overall sensing goal.
Sandeep Chinchali, Eyal Cidon, Evgenya Pergament, Sachin Katti
HotNets1
2017 Multi-objective Optimal Control for Proactive Decision Making with Temporal Logic Models
Sandeep Chinchali, Scott C. Livingston, Marco Pavone 0001
ISRR1
2016 Simultaneous model identification and task satisfaction in the presence of temporal logic constraints
abstract
Recent proliferation of cyber-physical systems, ranging from autonomous cars to nuclear hazard inspection robots, has exposed several challenging research problems on automated fault detection and recovery. This paper considers how recently developed formal synthesis and model verification techniques may be used to automatically generate information-seeking trajectories for anomaly detection. In particular, we consider the problem of how a robot could select its actions so as to maximally disambiguate between different model hypotheses that govern the environment it operates in or its interaction with other agents whose prime motivation is a priori unknown. The identification problem is posed as selection of the most likely model from a set of candidates, where each candidate is an adversarial Markov decision process (MDP) together with a linear temporal logic (LTL) formula that constrains robot-environment interaction. An adversarial MDP is an MDP in which transitions depend on both a (controlled) robot action and an (uncontrolled) adversary action. States are labeled, thus allowing interpretation of satisfaction of LTL formulae, which have a special form admitting satisfaction decisions in bounded time. An example where a robotic car must discern whether neighboring vehicles are following its trajectory for a surveillance operation is used to demonstrate our approach.
Sandeep Chinchali, Scott C. Livingston, Marco Pavone 0001, Joel W. Burdick
ICRA1
2016 NUMFabric: Fast and Flexible Bandwidth Allocation in Datacenters
abstract
We present xFabric, a novel datacenter transport design that provides flexible and fast bandwidth allocation control. xFabric is flexible: it enables operators to specify how bandwidth is allocated amongst contending flows to optimize for different service-level objectives such as minimizing flow completion times, weighted allocations, different notions of fairness, etc. xFabric is also very fast, it converges to the specified allocation one-to-two order of magnitudes faster than prior schemes. Underlying xFabric, is a novel distributed algorithm that uses in-network packet scheduling to rapidly solve general network utility maximization problems for bandwidth allocation. We evaluate xFabric using realistic datacenter topologies and highly dynamic workloads and show that it is able to provide flexibility and fast convergence in such stressful environments.
Kanthi Nagaraj, Dinesh Bharadia, Hongzi Mao, Sandeep Chinchali, Mohammad Alizadeh, Sachin Katti
SIGCOMM4
2016 Erosion of Conserved Binding Sites in Personal Genomes Points to Medical Histories
abstract
Although many human diseases have a genetic component involving many loci, the majority of studies are statistically underpowered to isolate the many contributing variants, raising the question of the existence of alternate processes to identify disease mutations. To address this question, we collect ancestral transcription factor binding sites disrupted by an individual's variants and then look for their most significant congregation next to a group of functionally related genes. Strikingly, when the method is applied to five different full human genomes, the top enriched function for each is invariably reflective of their very different medical histories. For example, our method implicates "abnormal cardiac output" for a patient with a longstanding family history of heart disease, "decreased circulating sodium level" for an individual with hypertension, and other biologically appealing links for medical histories spanning narcolepsy to axonal neuropathy. Our results suggest that erosion of gene regulation by mutation load significantly contributes to observed heritable phenotypes that manifest in the medical history. The test we developed exposes a hitherto hidden layer of personal variants that promise to shed new light on human disease penetrance, expressivity and the sensitivity with which we can detect them.
Harendra Guturu, Sandeep Chinchali, Shoa L. Clarke, Gill Bejerano
PLoS Comput. Biol.2
2013 Scheduling packets over multiple interfaces while respecting user preferences
abstract
Now that our smartphones have multiple interfaces (WiFi, 3G, 4G, etc.), we have preferences for which interfaces an application may use. We may prefer to stream video over WiFi because it is fast, but VoIP over 3G because it gives continued connectivity. We also have relative preferences, such as giving Netflix twice as much capacity as Dropbox. This means our mobile devices need to schedule packets in keeping with our preferences while making use of all the capacity available. This is the natural domain of fair queuing, and this paper is about the design of a packet scheduler to meet these requirements. We show that traditional fair queueing schedulers cannot take into account a user's preferences for some interfaces over others. We present a novel packet scheduler called miDRR that meets our needs by generalizing DRR for multiple interfaces. We demonstrate a prototype running in Linux and show that it works correctly and can easily run at the speeds we need.
Kok-Kiong Yap, Te-Yuan Huang, Yiannis Yiakoumis, Sandeep Chinchali, Nick McKeown, Sachin Katti
CoNEXT4
2012 Towards formal synthesis of reactive controllers for dexterous robotic manipulation
abstract
In robotic finger gaiting, fingers continuously manipulate an object until joint limitations or mechanical limitations periodically force a switch of grasp. Current approaches to gait planning and control are slow, lack formal guarantees on correctness, and are generally not reactive to changes in object geometry. To address these issues, we apply advances in formal methods to model a gait subject to external perturbations as a two-player game between a finger controller and its adversarial environment. High-level specifications are expressed in linear temporal logic (LTL) and low-level control primitives are designed for continuous kinematics. Simulations of planar manipulation with our synthesized correct-by-construction gait controller demonstrate the benefits of this approach.
Sandeep Chinchali, Scott C. Livingston, Ufuk Topcu, Joel W. Burdick, Richard M. Murray
ICRA1
2010 Axel rover paddle wheel design, efficiency, and sinkage on deformable terrain
abstract
This paper presents the Axel robotic rover which has been designed to provide robust and flexible access to extreme extra-planetary terrains. Axel is a lightweight 2-wheeled vehicle that can access steep slopes and negotiate relatively large obstacles due to its actively managed tether and novel wheel design. This paper reviews the Axel system and focuses on its novel paddle wheel characteristics. We show that the paddle design has superior rock climbing ability. We also adapt basic terramechanics principles to estimate the sinkage of paddle wheels on loose sand. Experimental comparisons between the transport efficiency of mountain bike wheels and paddle wheels are summarized. Finally, we present an unfolding wheel prototype which allows Axel to be compacted for efficient transport.
Pablo Abad-Manterola, Joel W. Burdick, Issa A. D. Nesnas, Sandeep Chinchali, Christine Fuller, Xuecheng Zhou
ICRA4