Divyansh Pradhan

dblp:429/1354 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Human-computer interaction and pervasive computing
2 papers
Immersive interaction · 36% Interaction techniques and input · 28% Human-robot interaction · 28%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Immersive interaction › augmented reality
augmented reality guidance
1.012026
From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented Reality · VR 2026
Human-robot interaction › robot perception
object disambiguation
1.012026
From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented Reality · VR 2026
Interaction techniques and input
voice interaction
1.012026
SpeechLess: Micro-utterance with Personalized Spatial Memory-aware Assistant in Everyday Augmented Reality · VR 2026
Collaborative and social computing › remote collaboration
remote assistance
0.312026
From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented Reality · VR 2026

Methods — techniques the papers use, named apart from their topics

speech parsing · 1.0relational graph grounding · 1.0intent inference · 1.0in-the-wild study · 1.0formative study · 1.0
YearPublicationVenuePosition
2026 SpeechLess: Micro-utterance with Personalized Spatial Memory-aware Assistant in Everyday Augmented Reality
abstract
Speaking aloud to a wearable AR assistant in public can be socially awkward, and re-articulating the same requests every day creates unnecessary effort. We present SpeechLess, a wearable AR assistant that introduces a speech-based intent granularity control paradigm grounded in personalized spatial memory. SpeechLess helps users "speak less," while still obtaining the information they need, and supports gradual explicitation of intent when more complex expression is required. SpeechLess binds prior interactions to multimodal personal context–space, time, activity, and referents–to form spatial memories, and leverages them to extrapolate missing intent dimensions from under-specified user queries. This enables users to dynamically adjust how explicitly they express their informational needs, from full-utterance to micro/zero-utterance interaction. We motivate our design through a week-long formative study using a commercial smart glasses platform, revealing discomfort with public voice use, frustration with repetitive speech, and hardware constraints. Building on these insights, we design SpeechLess, and evaluate it through controlled lab and in-the-wild studies. Our results indicate that regulated speech-based interaction, can improve everyday information access, reduce articulation effort, and support socially acceptable use without substantially degrading perceived usability or intent resolution accuracy across diverse everyday environments.
Yoonsang Kim, Devshree Jadeja, Divyansh Pradhan, Yalong Yang 0001, Arie E. Kaufman
VR3
2026 From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented Reality
abstract
We introduce Speech-to-Spatial, a referent disambiguation framework that converts verbal remote-assistance instructions into spatially grounded AR guidance. Unlike prior systems that rely on additional cues (e.g., gesture, gaze) or manual expert annotations, Speech-to-Spatial infers the intended target solely from spoken references (speech input). Motivated by our formative study of speech referencing patterns, we characterize recurring ways people specify targets (Direct Attribute, Relational, Remembrance, and Chained) and ground them to our object-centric relational graph. Given an utterance, referent cues are parsed and rendered as persistent in-situ AR visual guidance, reducing iterative micro-guidance ("a bit more to the right", "now, stop.") during remote guidance. We demonstrate the use cases of our system with remote guided assistance and intent disambiguation scenarios. Our evaluation shows that Speech-to-Spatial improves task efficiency, reduces cognitive load, and enhances usability compared to a conventional voice-only baseline, transforming disembodied verbal instruction into visually explainable, actionable guidance on a live shared view.
Yoonsang Kim, Divyansh Pradhan, Devshree Jadeja, Arie E. Kaufman
VR2