Shubhangi S. R. Garnaik

dblp:334/5697 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0005-3211-1703ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Virtual and augmented reality · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Virtual and augmented reality › augmented reality › augmented reality applications
AR storytelling
0.912025
Artspeak: An Interactive AR Application for Lifelike Speaking with Art Portraits · ISMAR 2025
Virtual and augmented reality › augmented reality
augmented reality applications
0.912025
Artspeak: An Interactive AR Application for Lifelike Speaking with Art Portraits · ISMAR 2025
Information retrieval › retrieval models › neural retrieval
embedding-based retrieval
0.312025
Artspeak: An Interactive AR Application for Lifelike Speaking with Art Portraits · ISMAR 2025
Information retrieval
retrieval models
0.312025
Artspeak: An Interactive AR Application for Lifelike Speaking with Art Portraits · ISMAR 2025

Methods — techniques the papers use, named apart from their topics

cosine similarity · 1.7GPT-based embeddings · 1.7FAQ retrieval · 1.7
YearPublicationVenuePosition
2025 Artspeak: An Interactive AR Application for Lifelike Speaking with Art Portraits
abstract
Museum visits often lack personalized and interactive experiences, limiting visitor engagement with art and historical artifacts. To address this, we present ArtSpeak, a standalone augmented reality (AR) application that transforms traditional art viewing into an interactive storytelling experience. When users point their mobile cameras at an artwork, the system responds to their questions with lifelike, talking-head video narratives generated from historical portraits. However, generating such talking-head videos at runtime is computationally expensive, often requiring over a minute per response. To address this challenge, ArtSpeak introduces two major contributions. First, it employs a collection of frequently asked questions (FAQ) to generate a set of lifelike video responses for various art portraits. Second, it introduces a novel retrieval-based approach that uses GPT-based embeddings and cosine similarity to select the most relevant response. As a result, the system dynamically presents the video reply that best aligns with the user's inquiry, reducing computational overhead and ensuring a real-time, low-latency experience. More precisely, ArtSpeak achieves over 30 x lower latency and reduces energy consumption by approximately 81 % compared to the real-time video generation method. User studies further validate the system's effectiveness, with 85 % of participants rating the retrieved responses as relevant to their queries and 90 % reporting smooth video playback. These results highlight the efficiency and user satisfaction enabled by our retrieval-based approach.
Shubhangi S. R. Garnaik, Aruna Balasubramanian, Niranjan Balasubramanian, Jihoon Ryoo
ISMAR1
2025 CLOUD-CODEC: A New Way of Storing Traffic Camera Footage at Scale
abstract
Storing large volumes of traffic video content in cloud storage is an expensive undertaking, given the limited capacity of cloud storage and its inability to store data beyond a few weeks. To address this issue, this article introduces CLOUD-CODEC , a novel video encoding approach tailored specifically for traffic monitoring video. CLOUD-CODEC offers three key advantages: (i) real-time encoding without any delay, (ii) near-perfect video quality upon decoding, and (iii) one-fifth the storage size of traditional encoding methods. CLOUD-CODEC is generally applicable to traffic cameras under various weather and lighting conditions. The encoding algorithm is a lightweight DNN-based object detection and box-shaped segmentation approach. The method can uniquely detect and segment cars, pedestrians, and moving objects with the marginal box-shaped contours. Periodic object detection makes it possible for CLOUD-CODEC to operate in real-time and estimate the movement of objects between predictions. Proof-of-concept evaluations using a massive dataset indicate that CLOUD-CODEC reduces video size by 80%—surpassing AV1 (34.9%), CloudSeg (58.4%), Detection (76.9%), Segmentation (73.1%), and Segm&Sort (69.5%). It achieves a frame rate of 95.8 when encoding and a VMAF score of 72.54 after decoding, with a storage size that is one-fifth of traditional methods. Field-testing of CLOUD-CODEC on metropolitan traffic cameras demonstrates its ability to extend storage time by 74.92%.
Hoyoung Kim, Azimbek Khudoyberdiev, Shubhangi S. R. Garnaik, Arani Bhattacharya, Jihoon Ryoo
ACM Trans. Multim. Comput. Commun. Appl.3