Gaurav Jain

dblp:40/5058 · DBLP profile ↗
← Back
19ranked-venue papers
8as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 6 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Security and privacy · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1
YearPublicationVenuePosition
2026 SceneScout: Towards AI-Driven Access to Street Level Imagery for Blind Users
abstract
People who are blind or have low-vision (BLV) may hesitate to travel independently in unfamiliar environments due to uncertainty about the physical landscape. While most tools focus on in-situ navigation assistance, those supporting pre-travel assistance typically provide information about only landmarks and turn-by-turn instructions, lacking detailed visual context. Street level imagery, which contains rich visual information and has the potential to reveal numerous environmental details, remains inaccessible to BLV people. In this work, we present SceneScout, a multimodal large language model (MLLM)-driven prototype that enables accessible interactions with street level imagery. SceneScout supports two modes: (1) Route Preview, enabling users to familiarize themselves with visual details along a route, and (2) Virtual Exploration, enabling free, user-driven movement within street level imagery. Our user study (N = 10) demonstrates that SceneScout helps BLV users uncover visual information otherwise unavailable through existing means. An initial analysis of AI-generated descriptions suggests that the majority are accurate and describe stable visual elements even in older imagery, though occasional subtle and plausible errors make them difficult to verify without sight. We discuss future opportunities and challenges of street level imagery-based navigation experiences.
Gaurav Jain, Leah Findlater, Cole Gleason
CHI1
2025 SEE++: Evolving Snowpark Execution Environment for Modern Workloads
Gaurav Jain, Brandon Baker, Joe Yin, Chenwei Xie, Sidh Kulkarni, Sara Abdelrahman, Nova Qi, Urjeet Shrestha, Mike Halcrow, Dave Bailey, Yuxiong He
IEEE Big Data1
2025 Dialogue Without Limits: Constant-Sized KV Caches for Extended Response in LLMs
abstract
Autoregressive Transformers rely on Key-Value (KV) caching to accelerate inference. However, the linear growth of the KV cache with context length leads to excessive memory consumption and bandwidth constraints. Existing methods drop distant tokens or compress states in a lossy manner, sacrificing accuracy by discarding vital context or introducing bias. We propose ${MorphKV}$, an inference-time technique that maintains a constant-sized KV cache while preserving accuracy. MorphKV balances long-range dependencies and local coherence during text generation. It eliminates early-token bias while retaining high-fidelity context by adaptively ranking tokens through correlation-aware selection. Unlike heuristic retention or lossy compression, MorphKV iteratively refines the KV cache via lightweight updates guided by attention patterns of recent tokens. This approach captures inter-token correlation with greater accuracy, which is crucial for tasks like content creation and code generation. Our studies on long-response tasks show 52.9% memory savings and 18.2% higher accuracy on average compared to state-of-the-art prior works, enabling efficient deployment.
Ravi Ghadia, Avinash Kumar 0008, Gaurav Jain, Prashant J. Nair, Poulami Das 0005
ICML3
2025 Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies
abstract
Accelerating the inference of large language models (LLMs) is a critical challenge in generative AI. Speculative decoding (SD) methods offer substantial efficiency gains by generating multiple tokens using a single target forward pass. However, existing SD approaches require the drafter and target models to share the same vocabulary, thus limiting the pool of possible drafters, often necessitating the training of a drafter from scratch. We present three new SD methods that remove this shared-vocabulary constraint. All three methods preserve the target distribution (i.e., they are lossless) and work with off-the-shelf models without requiring additional training or modifications. Empirically, on summarization, programming, and long-context tasks, our algorithms demonstrate significant speedups of up to 2.8x over standard autoregressive decoding. By enabling any off-the-shelf model to serve as a drafter and requiring no retraining, this work substantially broadens the applicability of the SD framework in practice.
Nadav Timor, Jonathan Mamou, Daniel Korat, Moshe Berchansky, Gaurav Jain, Oren Pereg, Moshe Wasserblat, David Harel
ICML5
2024 StreetNav: Leveraging Street Cameras to Support Precise Outdoor Navigation for Blind Pedestrians
abstract
Blind and low-vision (BLV) people rely on GPS-based systems for outdoor navigation. GPS’s inaccuracy, however, causes them to veer off track, run into obstacles, and struggle to reach precise destinations. While prior work has made precise navigation possible indoors via hardware installations, enabling this outdoors remains a challenge. Interestingly, many outdoor environments are already instrumented with hardware such as street cameras. In this work, we explore the idea of repurposing existing street cameras for outdoor navigation. Our community-driven approach considers both technical and sociotechnical concerns through engagements with various stakeholders: BLV users, residents, business owners, and Community Board leadership. The resulting system, StreetNav, processes a camera’s video feed using computer vision and gives BLV pedestrians real-time navigation assistance. Our evaluations show that StreetNav guides users more precisely than GPS, but its technical performance is sensitive to environmental occlusions and distance from the camera. We discuss future implications for deploying such systems at scale.
Gaurav Jain, Basel Hindi, Koushik Srinivasula, Mingyu Xie, Mahshid Ghasemi, Daniel Weiner, Sophie Ana Paris, Xin Yi Therese Xu, Michael C. Malcolm, Mehmet Kerem Türkcan, Javad Ghaderi, Zoran Kostic, Gil Zussman, Brian A. Smith 0001
UIST1
2023 Towards Street Camera-based Outdoor Navigation for Blind Pedestrians
abstract
Blind and low-vision (BLV) people use GPS-based systems for outdoor navigation assistance, which provide instructions to get from one place to another. However, such systems do not provide users with real-time, precise information about their location and surroundings which is crucial for safe navigation. In this work, we investigate whether street cameras can be used to address aspects of navigation that BLV people still find challenging with existing GPS-based assistive technologies. We conducted formative interviews with six BLV participants to identify specific challenges they face in outdoor navigation. We discovered three main challenges: anticipating environment layouts, avoiding obstacles while following directions, and crossing noisy street intersections. To address these challenges, we are currently developing a street camera-based navigation system that provides real-time auditory feedback to help BLV users avoid obstacles, know exactly when to cross the street, and understand the overall layout of the environment. We close by discussing our evaluation plan.
Gaurav Jain, Basel Hindi, Mingyu Xie, Koushik Srinivasula, Mahshid Ghasemi, Daniel Weiner, Xin Yi Therese Xu, Sophie Ana Paris, Chloe Tedjo, Josh Bassin, Michael C. Malcolm, Mehmet Kerem Türkcan, Javad Ghaderi, Zoran Kostic, Gil Zussman, Brian A. Smith 0001
ASSETS1
2023 Front Row: Automatically Generating Immersive Audio Representations of Tennis Broadcasts for Blind Viewers
abstract
Blind and low-vision (BLV) people face challenges watching sports due to the lack of accessibility of sports broadcasts. Currently, BLV people rely on descriptions from TV commentators, radio announcers, or their friends to understand the game. These descriptions, however, do not allow BLV viewers to visualize the action by themselves. We present Front Row, a system that automatically generates an immersive audio representation of sports broadcasts, specifically tennis, allowing BLV viewers to more directly perceive what is happening in the game. Front Row first recognizes gameplay from the video feed using computer vision, then renders players’ positions and shots via spatialized (3D) audio cues. User evaluations with 12 BLV participants show that Front Row gives BLV viewers a more accurate understanding of the game compared to TV and radio, enabling viewers to form their own opinions on players’ moods and strategies. We discuss future implications of Front Row and illustrate several applications, including a Front Row plug-in for video streaming platforms to enable BLV people to visualize the action in sports videos across the Web.
Gaurav Jain, Basel Hindi, Connor Courtien, Xin Yi Therese Xu, Conrad Wyrick, Michael C. Malcolm, Brian A. Smith 0001
UIST1
2023 Polarised social media discourse during COVID-19 pandemic: evidence from YouTube
abstract
The onset of the COVID-19 pandemic has attracted significant attention on social media platforms as these platforms provide users unparalleled access to ‘information’ from around the globe. In spite of demographic differences, people have been expressing and shaping their opinions using social media on topics ranging from the plight of migrant workers to vaccine development. However, the social media induced polarisation owing to selective online exposure to information during the COVID-19 pandemic has been a major cause of concern for countries across the world. In this paper, we analyse the temporal dynamics of polarisation in online discourse related to the COVID-19. We use random network theory-based simulation to investigate the evolution of opinion formation in comments posted on different COVID-19-related YouTube videos. Our findings reveal that as the pandemic unfolded, the extent of polarisation in the online discourse increased with time. We validate our experimental model using real-world complex networks and compare consensus formation on these networks with equivalent random networks. This study has several implications as polarisation around socio-cultural issues in crises such as pandemic can exacerbate the social divide. The framework proposed in this study can aid regulatory agencies to take required actions and mitigate social media-induced polarisation.
Samrat Gupta, Gaurav Jain, Amit Anand Tiwari
Behav. Inf. Technol.2
2023 "I Want to Figure Things Out": Supporting Exploration in Navigation for People with Visual Impairments
abstract
Navigation assistance systems (NASs) aim to help visually impaired people (VIPs) navigate unfamiliar environments. Most of today's NASs support VIPs via turn-by-turn navigation, but a growing body of work highlights the importance of exploration as well. It is unclear, however, how NASs should be designed to help VIPs explore unfamiliar environments. In this paper, we perform a qualitative study to understand VIPs' information needs and challenges with respect to exploring unfamiliar environments to inform the design of NASs that support exploration. Our findings reveal the types of spatial information that VIPs need as well as factors that affect VIPs' information preferences. We also discover specific challenges that VIPs face that future NASs can address, such as orientation and mobility education and collaborating effectively with others. We present design implications for NASs that support exploration, and we identify specific research opportunities and discuss open socio-technical challenges for making such NASs possible. We conclude by reflecting on our study procedure to inform future approaches in research on ethical considerations that may be adopted while interacting with the broader VIP community.
Gaurav Jain, Yuanyang Teng, Dong Heon Cho, Yunhao Xing, Maryam Aziz, Brian A. Smith 0001
Proc. ACM Hum. Comput. Interact.1
2023 Kepler: Robust Learning for Parametric Query Optimization
abstract
Most existing parametric query optimization (PQO) techniques rely on traditional query optimizer cost models, which are often inaccurate and result in suboptimal query performance. We propose Kepler, an end-to-end learning-based approach to PQO that demonstrates significant speedups in query latency over a traditional query optimizer. Central to our method is Row Count Evolution (RCE), a novel plan generation algorithm based on perturbations in the sub-plan cardinality space. While previous approaches require accurate cost models, we bypass this requirement by evaluating candidate plans via actual execution data and training anML model to predict the fastest plan given parameter binding values. Our models leverage recent advances in neural network uncertainty in order to robustly predict faster plans while avoiding regressions in query performance. Experimentally, we show that Kepler achieves significant improvements in query runtime on multiple datasets on PostgreSQL.
Lyric Doshi, Vincent Zhuang, Gaurav Jain, Ryan Marcus, Deniz Altinbüken, Eugene Brevdo, Campbell Fraser
Proc. ACM Manag. Data3
2021 SketchFormer: transformer-based approach for sketch recognition using vector images
Anil Singh Parihar, Gaurav Jain, Shivang Chopra, Suransh Chopra
Multim. Tools Appl.2
2020 TransSketchNet: Attention-Based Sketch Recognition Using Transformers
Gaurav Jain, Shivang Chopra, Suransh Chopra, Anil Singh Parihar
ECAI1
2020 Adaptive Weighted Graph Approach to Generate Multimodal Cancelable Biometric Templates
abstract
Multimodal biometric systems offer numerous advantages over unimodal counterparts and are being used extensively in diverse applications. However, fusion of biometric data is a non-trivial task and curtail employability of multimodal systems for a varying set of biometric characteristics with different type and dimension. Moreover, comprehensive solutions against adversary attacks that ensure template protection and prevent presentation attacks are not in place. In this article, a secure multimodal cancelable biometric system is proposed to address these concerns. This approach introduces key images based generic feature extraction technique which reduces feature dimension and achieves revocability. The non-invertibility and unlinkability are ensured through cross-diffusion of complementary information from different modalities. A new feature fusion method based on an adaptive graph is proposed to generate multimodal cancelable biometric templates. Robustness against presentation attack is accomplished through quality based adaptation of features. Extensive experimentation is performed on benchmark databases for fingerprint, face, and iris, to illustrate the efficacy of multimodal cancelable templates. The proposed approach is shown to perform favorably against state-of-the-art feature fusion methods. Furthermore, the resilience of the proposed approach against security and privacy attacks is demonstrated.
Gurjit Singh Walia, Gaurav Jain, Nipun Bansal, Kuldeep Singh 0002
IEEE Trans. Inf. Forensics Secur.2
2019 MEG: A RISCV-Based System Simulation Infrastructure for Exploring Memory Optimization Using FPGAs and Hybrid Memory Cube
abstract
Emerging 3D memory technologies, such as the Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM), provide increased bandwidth and massive memory-level parallelism. Efficiently integrating emerging memories into existing system pose new challenges and require detailed evaluation in a real computing environment. In this paper, we propose MEG, an open-source, configurable, cycle-exact, and RISC-V based full system simulation infrastructure using FPGA and HMC. MEG has three highly configurable design components: (i) a HMC adaptation module that not only enables communication between the HMC device and the processor cores but also can be extended to fit other memories (e.g., HBM, nonvolatile memory) with minimal effort, (ii) a reconfigurable memory controller along with its OS support that can be effectively leveraged by system designers to perform software-hardware co-optimization, and (iii) a performance monitor module that effectively improves the observability and debuggability of the system to guide performance optimization. We provide a prototype implementation of MEG on Xilinx VCU110 board and demonstrate its capability, fidelity, and flexibility on real-world benchmark applications. We hope that our open-source release of MEG fills a gap in the space of publicly-available FPGA-based full system simulation infrastructures specifically targeting memory system and inspires further collaborative software/hardware innovations.
Gaurav Jain, Yue Zha, Jonathan Ta, Jing Jane Li
FCCM3
2018 Enhanced multi-RAT support for 5G
abstract
5G is one of the most sought after technology for supporting massive connectivity, reduced latency, higher throughput, D2D communication, Dual Connectivity, LTE-Wifi aggregation and many other services. Multi-Radio Access Technology (Multi-RATs) carrier aggregation (CA), also known as multi-flow CA allows different RATs to be aggregated and allocated to the UE. So, an optimized Multi-RAT support is very much desired in 5G Environment. This paper covers the 5G architecture facilitating an optimized plug and play model for providing dedicated Multi-RAT services through simulation results and studies.
Arjun Nanjundappa, Sukhdeep Singh, Gaurav Jain
CCNC3
2015 Efficient Path Rescheduling of Heterogeneous Mobile Data Collectors for Dynamic Events in Shanty Town Emergency Response
abstract
To investigate reported emergency incidents and provide better situational awareness during an emergency response effort in a shanty town, we envision the use of volunteers with networked sensing devices employed as Mobile Data Collectors (MDCs). These MDCs are heterogeneous depending upon the type of roads they can access. They gather information about events reported dynamically at random and relay it to central command center. We consider the problem of minimizing the Travel Time of such heterogeneous volunteer MDCs and maximizing the gathering of event data before its expiry time. We model this problem as a Dynamic Vehicle Routing Problem with Time Windows (DVRPTW), which reduces to a Combinatorial Optimization Problem and is NP-Hard to solve. In this paper, we developed two algorithms, Minimum Deviated Walk and Ortho Walk, to dynamically route or reroute the path of these MDCs to capture the data efficiently. We tested the effectiveness of these algorithms with three different classes of MDCs on simulated non- deterministic random sets of events applied to a real road map of Dharavi, a shanty town in Mumbai, India. We show that both these algorithms are capable of capturing 20% more data than a naive algorithm as well as more than 90% of the events generated within a specified time.
Ranga Raj, Sarath Babu 0001, Kyle E. Benson, Gaurav Jain, B. S. Manoj 0001, Nalini Venkatasubramanian
GLOBECOM4
2014 Use of Enterprise Clinical Decision Support Infrastructure to Implement Pharmacogenomics at the Point of Care
Pedro J. Caraballo, David Blair, Michelle Elliott, Robert R. Bleimeyer, John Crooks, Donald B. Gabrielson, Gaurav Jain, Wayne T. Nicholson, Charles Pugh, Padma S. Rao, Cloann Schultz, Lynn Summerlin, Joseph Sutton, Carolyn Rohrer Vitek, Kelly Wix, John L. Black, Mark A. Parkulo
AMIA7
2002 Service level agreements in IP networks
abstract
Internet services are being deployed over an infrastructure that involves co‐operation between multiple organizations and systems. This has necessitated the need for standard means to share the information between the service providers and their customers. This information essentially pertains to the service level obligations between the service provider and their customers so that the customers can ensure the quality of service (QoS) that they are able to achieve at their end. A service level agreement (SLA) essentially quantifies the level of service as it includes the metrics that define the quality of service. The research undertaken identifies the QoS dimensions, which are required to define the multimedia services. Each application used by the user will involve different values of the QoS dimensions in order to maintain an expected level of service. The QoS requirement for a particular application will also depend upon the provisioning of the network resources depending on the client and server side CPU and memory available for processing. The relationship between the system resources, QoS dimensions and the SLA has been depicted in the form of a general model of SLA and as an example taken for the video conferencing application.
Gaurav Jain, Deepali Singh, Shekhar Verma
Inf. Manag. Comput. Secur.1
2002 Indexing for local appearance-based recognition of planar objects
Gaurav Jain, Santanu Chaudhury
Pattern Recognit. Lett.2