Puneet Jain

dblp:12/7491 · DBLP profile ↗
← Back
22ranked-venue papers
12as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 12 · 6 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 When Assistive Technologies become Provocations: Unpacking Access in HCI practices using Crip Technoscience, Mouth Interfaces, and XR
abstract
XR technologies often presume hand-based interaction as universal, marginalizing people with limited mobility. In close collaboration with two disabled artists, we developed a mouth interface that reimagines XR interaction, foregrounding disabled practices. Rather than positioning as an accessibility “fix,” we stage the interface as an activist provocation against hand-centric design. Using a narrative-flip method, we asked early-stage HCI researchers to use the interface without context, later revealing its disability-led, activist origins. This before/after framing exposed how the participants imagined accessibility, shifting from awkwardness and discomfort to empathy and political reflection − yet sliding back into solutionist product thinking. Through analysis of participants’ reflections, enriched by our disabled collaborator’s feedback, we reveal how accessibility was framed as a fix, as political, and as agency. We argue that accessibility research in HCI can move beyond solutionist fixes: assistive technologies as provocations that disrupt able-bodied assumptions and expand how access is imagined.
Puneet Jain, Ayush Sharma, Sidharth Chaudhary, Vivek Rawat, Akhilesh Kumar Bhagat, Kratika Jain, Christian Bayerlein, Christopher L. Salter, Gowdham Prabhakar
CHI1
2026 KC-Agent: A Dual-Process Cognitive Architecture for Efficient ML Model Improvement
abstract
Data drift poses significant challenges for machine learning systems in production, requiring continuous model updates to maintain performance. We present KC-Agent, a dual-process cognitive architecture for automated ML model improvement that combines fast pattern recognition (System 1) with deliberate incremental updates (System 2). Our approach implements structured memory systems enabling System 1 to leverage successful solutions previously discovered by System 2, achieving efficient pattern-based responses without costly re-computation. KC-Agent incorporates atomic change principles and rollback capabilities to ensure reliable, verifiable updates in production environments. We evaluate our method on five datasets including real-world NASA turbofan data with authentic temporal degradation and synthetic datasets with controlled drift scenarios. KC-Agent achieves state-of-the-art performance (76.8% accuracy) while maintaining optimal efficiency (13.2s execution time), outperforming established cognitive architectures: CodeAct (+2.4%), Tree of Thoughts (+3.6%), ReAct (+8.0%), and Reflexion (+8.9%). Consensus evaluation by a panel of state-of-the-art LLMs confirms superior strategic efficacy (8.33/10 Smartness score), significantly outperforming baseline agents. The knowledge consolidation mechanism delivers 91% speedup over the slow variant while maintaining higher accuracy. Our approach demonstrates both theoretical foundations and practical viability for cognitive-inspired automated ML improvement systems capable of handling complex real-world data drift scenarios.
Gusseppe Bravo Rocca, Jordi Guitart, Ajay Dholakia, David Ellison, Puneet Jain
COMPSAC5
2025 Intelligent Sampling for Predicting the Performance of Hub-based Swarms
abstract
This paper presents an inductive learning algorithm to predict the performance of hub-based swarms solving the best-of-N problem. Since a major constraint in learning swarm behavior is the high computational cost of obtaining sample data, it is desirable to ensure the right samples are used to train the models. The paper’s main contribution is formulating and comparing various sampling techniques to improve performance prediction using manageable amounts of training data. We compare random sampling with in-distribution sampling and out-of-distribution sampling, and then apply the lessons learned to modify random sampling to improve sampling. Results show that in-distribution sampling has the best F1 across different sampling techniques for classifying slow versus fast convergence. Model performance indicates that an informed combination of in-distribution and out-of-distribution sampling produces the highest classification accuracy of the swarm’s time-to-converge.
Puneet Jain, Chaitanya Dwivedi, Michael A. Goodrich
RO-MAN1
2023 Designing and Predicting the Performance of Agent-based Models for Solving Best-of-N
abstract
Biological inspiration from honeybees, insects, and other animals has been used to create interesting implementations of multi-robot swarms. When the robots in a swarm are completely distributed, that is they lack any form of centralized control, the swarm acts as an agent-based model (ABM) wherein each agent implements its own controller and collective behavior emerges from the interactions between agents. Differential equation and graph-based models of some types of swarms have been used to guarantee collective behavior, but guaranteeing or predicting outcomes for hub-based agent colonies with finite numbers of robots remains an open problem. This paper presents a case study of designing an agent-based, hub-based swarm that solves the best-of-N problem with predictable success rates and completion times. The key innovation is modifying a tripartite graph formulation (TGF) from previous work so that it acts as a graph schema which abstracts an ABM into a simplified four state model, which in turn leads to a large discrete time Markov chain (DTMC) that describes how the collective state evolves over time. The DTMC can be used to compute success rates and completion times, which act as predictions for the ABM. Deviations between observed ABM outcomes and DTMC predictions lead to modifications in the ABM so that the swarm becomes more predictable.
Puneet Jain, Michael A. Goodrich
SMC1
2023 Navigating in VR using free-hand gestures and embodied controllers: A comparative evaluation
abstract
While natural body-based movements are essential features for immersive VR (Virtual Reality) experiences, most of the available input techniques for navigation in VR involve the use of hand-held controllers. Alternatively, while body-based input for VR navigation has previously been explored in HCI using external tracking devices, there is little to no work that utilizes the in-built tracking functionalities of the predominant VR headsets (such as Meta Quest 2) for gesture-based navigation in VR. This paper addresses this research gap by proposing five free-hand gestures for 3-D navigation in VR using internal gesture-tracking functionality of Quest 2 headset. Additionally, a qualitative and quantitative comparison is presented between free-hand and controller-based navigation in VR using a custom designed task (with 10 users). Overall, the findings from the task-analysis indicate that while in-built tracking functionalities in VR headsets open doors for inexpensive gesture-based VR navigation, the mid-air hand-gestures result into greater fatigue as compared to using controllers for navigation in VR.
Puneet Jain
VRST1
2023 Dissociable default-mode subnetworks subserve childhood attention and cognitive flexibility: Evidence from deep learning and stereotactic electroencephalography
Nebras M. Warsi, Simeon M. Wong, Jürgen Germann, Alexandre Boutet, Olivia N. Arski, Lauren Erdman, Hrishikesh Suresh, Flavia Venetucci Gouveia, Aaron Loh, Gavin J. B. Elias, Elizabeth Kerr, Mary Lou Smith, Ayako Ochi, Hiroshi Otsubo, Roy Sharma, Puneet Jain, Elizabeth Donner, Andres M. Lozano, O. Carter Snead, George M. Ibrahim
Neural Networks18
2022 Adapted Metrics for Measuring Competency and Resilience for Autonomous Robot Systems in Discrete Time Markov Chains
abstract
Autonomous robot systems are often designed to achieve specific goals. This paper restricts attention to a specific type of goal, namely reaching a desired state within a certain time bound. For such goals, a robot system’s competency and resilience can be defined as the probability of reaching the desired state as a function of the time bound under a nominal unperturbed condition and under known perturbation conditions, respectively. Two metrics taken from prior work for measuring competency and resilience, power and efficiency, are modified so that they do not require subjective parameters. This paper formalizes the adapted metrics for discrete time Markov chains. The adapted metrics are applied to a best-of-N case study that is solved by a graph-based approach and modeled as a discrete time Markov chain. The case study demonstrates that the modified metrics allow power-efficiency trade-offs to be more easily visualized than the cluttered visualizations produced by the original metrics.
Xuan Cao, Puneet Jain, Michael A. Goodrich
SMC2
2021 Ideation via Critic-Based Exploration of Generator Latent Space
Puneet Jain, Najma Mathema, Jonathan Skaggs, Dan Ventura
ICCC1
2020 Egocentric Analysis of Dash-Cam Videos for Vehicle Forensics
abstract
Video acquisition using dashboard-mounted cameras has recently achieved massive popularity around the world. One of the major developments following the dash-cam's popularity is that videos captured by them can be used as testimony during scenarios, like traffic violations and accidents. The widespread deployment of dash-cams brings new problems ranging from the compromise of privacy by uploading these videos on public websites using videos captured from other cars for making fraudulent claims. Therefore, there is a compelling need to address the problems associated with the usage of dash-cam videos. In this paper, we discuss and highlight the importance of the emerging area of multimedia vehicle forensics. We propose an algorithm for linking a dash-cam video to a specific car. The proposed algorithm is useful for various applications, for example, insurance companies can authenticate the origin of video before processing the claim. In a different scenario of illegitimate video upload on the Web, the video can be traced back to the car it originated from. To this end, we make use of motion blur extracted from dash-cam videos for generating a discriminative feature. We observe that the subtle motion pattern of every vehicle can serve as its unique signature. We extract motion blur from dash-cam videos and use random forest trees for classifying the vehicle correctly. The experimental results on thousands of frames obtained from dash-cam videos of several cars show the effectiveness of our approach. We further investigate the process of forging the signature of a car and propose a counter forensics method to detect such forgery. Also, we discuss the application of our technique to other potential platforms where the camera can be mounted, for example, on the chest of a person. We believe that ours is the first work that describes this new area of research.
Ambuj Mehrish, Puneet Jain, A. Venkata Subramanyam, Mohan Kankanhalli
IEEE Trans. Circuits Syst. Video Technol.3
2019 Towards Scalable Video Analytics at the Edge
abstract
Breakthroughs in deep learning, GPUs, and edge computing have paved the way for always-on, live video analytics. However, to achieve real-time performance, a GPU needs to be dedicated amongst a few video feeds. But, GPUs are expensive resources and a large-scale deployment requires supporting hundreds of video cameras - exorbitant cost prohibits widespread adoption. To ease this burden, we propose Tetris, a system comprising of several optimization techniques from computer vision and deep-learning literature blended in a synergistic manner. Tetris is designed to maximize the parallel processing of video feeds on a single GPU, with a marginal drop in inference accuracy. Tetris performs CPU-based tiling of active regions to combine activities across video feeds. resulting in a condensed input volume. It then runs the deep learning model on this condensed volume instead of individual feeds, which significantly improves the GPU utilization. Our evaluation on Duke MTMC dataset reveals that Tetris can process 4x video feeds in parallel compared to any of the existing methods used in isolation.
Theodore Stone, Nathaniel Stone, Puneet Jain, Yurong Jiang, Kyu-Han Kim, Srihari Nelakuditi
SECON3
2018 CoDrive: Improving Automobile Positioning via Collaborative Driving
abstract
An increasing number of depth sensors and surrounding-aware cameras are being installed in the new generation of cars. For example, Tesla Motors uses a forward radar, a front-facing camera, and multiple ultrasonic sensors to enable its Autopilot feature. Meanwhile, older or legacy cars are expected to be around in volumes, for at least the next 10 to 15 years. Legacy car drivers rely on traditional GPS for navigation services, whose accuracy varies 5 to 10 meters in a clear line-of-sight and degrades up to 30 meters in a downtown environment. At the same time, a sensor-rich car achieves better accuracy due to high-end sensing capabilities. To bridge this gap, we propose CoDrive, a system to provide a sensor-rich car's accuracy to a legacy car. We achieve this by correcting GPS errors of a legacy car on an opportunistic encounter with a sensor-rich car. CoDrive uses smartphone GPS of all participating cars, RGB-D sensors of sensor-rich cars, and road boundaries of a traffic scene to generate optimization constraints. Our algorithm collectively reduces GPS errors, resulting in accurate reconstruction of a traffic scene's aerial view. CoDrive does not require stationary landmarks or 3D maps. We empirically evaluate CoDrive which is shown to achieve a 90% and a 30% reduction in cumulative GPS error for legacy and sensor-rich cars respectively, while preserving the shape of the traffic.
Soteris Demetriou, Puneet Jain, Kyu-Han Kim
INFOCOM2
2018 TAR: Enabling Fine-Grained Targeted Advertising in Retail Stores
abstract
Mobile advertisements influence customers' in-store purchases and boost in-store sales for brick-and-mortar retailers. Targeting mobile ads has become significantly important to compete with online shopping. The key to enabling targeted mobile advertisement and service is to learn shoppers' interest during their stay in the store. Precise shopper tracking and identification are essential to gain the insights. However, existing sensor-based or vision-based solutions are neither practical nor accurate; no commercial solutions today can be readily deployed in a large store. On the other hand, we recognize that most retail stores have the installation of surveillance cameras, and most shoppers carry Bluetooth-enabled smartphones. Thus, in this paper, we propose TAR to learn shoppers' in-store interest via accurate multi-camera people tracking and identification. TAR leverages widespread camera deployment and Bluetooth proximity information to accurately track and identify shoppers in the store. TAR is composed of four novel design components: (1) a deep neural network (DNN) based visual tracking, (2) a user trajectory estimation by using shopper visual and BLE proximity trace, (3) an identity matching and assignment to recognize shopper's identity, and (4) a cross-camera calibration algorithm. TAR carefully combines these components to track and identify shoppers in real-time. TAR achieves 90% accuracy in two different real-life deployments, which is 20% better than the state-of-the-art solution.
Yurong Jiang, Puneet Jain, Kyu-Han Kim
MobiSys3
2017 CamForensics: Understanding Visual Privacy Leaks in the Wild
abstract
Many mobile apps, including augmented-reality games, bar-code readers, and document scanners, digitize information from the physical world by applying computer-vision algorithms to live camera data. However, because camera permissions for existing mobile operating systems are coarse (i.e., an app may access a camera's entire view or none of it), users are vulnerable to visual privacy leaks. An app violates visual privacy if it extracts information from camera data in unexpected ways. For example, a user might be surprised to find that an augmented-reality makeup app extracts text from the camera's view in addition to detecting faces. This paper presents results from the first large-scale study of visual privacy leaks in the wild. We build CamForensics to identify the kind of information that apps extract from camera data. Our extensive user surveys determine what kind of information users expected an app to extract. Finally, our results show that camera apps frequently defy users' expectations based on their descriptions.
Animesh Srivastava, Puneet Jain, Soteris Demetriou, Landon P. Cox, Kyu-Han Kim
SenSys2
2016 Low Bandwidth Offload for Mobile AR
abstract
Environmental fingerprinting has been proposed as a key enabler to immersive, highly contextualized mobile computing applications, especially augmented reality. While fingerprints can be constructed in many domains (e.g., wireless RF, magnetic field, and motion patterns), visual fingerprinting is especially appealing due to the inherent heterogeneity in many indoor spaces. This visual diversity, however, is also its Achilles' heel -- matching a unique visual signature against a database of millions requires either impractical computation for a mobile device, or to upload large quantities of visual data for cloud offload. Further, most visual "features" tend to be low entropy -- e.g., homogeneous repetitions of floor and ceiling tiles. Our system VisualPrint, proposes a means to offload only the most distinctive visual data, that is, only those visual signatures which stand a good chance to yield a unique match. VisualPrint enables cloud-offloaded visual fingerprinting with efficacy comparable to using whole images, but with an order reduction in network transfer.
Puneet Jain, Justin Manweiler, Romit Roy Choudhury
CoNEXT1
2015 Poster: User Location Fingerprinting at Scale
abstract
Many emerging mobile computing applications are continuous vision based. The primary challenge these applications face is computation partitioning between the phone and cloud. The indoor location information is one metadata that can help these applications in making this decision. In this extended-abstract, we propose a vision based scheme to uniquely fingerprint an environment which can in turn be used to identify user's location from the uploaded visual features. Our approach takes into account that the opportunity to identify location is fleeting and the phones are resource constrained -- therefore minimal yet sufficient computation needs to be performed to make the offloading decision. Our work aims to achieve near real-time performance while scaling to buildings of arbitrary sizes. The current work is in preliminary stages but holds promise for the future -- may apply to many applications in this area.
Puneet Jain, Justin Manweiler, Romit Roy Choudhury
MobiCom1
2015 OverLay: Practical Mobile Augmented Reality
abstract
The idea of augmented reality - the ability to look at a physical object through a camera and view annotations about the object - is certainly not new. Yet, this apparently feasible vision has not yet materialized into a precise, fast, and comprehensively usable system. This paper asks: What does it take to enable augmented reality (AR) on smartphones today? To build a ready-to-use mobile AR system, we adopt a top-down approach cutting across smartphone sensing, computer vision, cloud offloading, and linear optimization. Our core contribution is in a novel location-free geometric representation of the environment - from smartphone sensors - and using this geometry to prune down the visual search space. Metrics of success include both accuracy and latency of object identification, coupled with the ease of use and scalability in uncontrolled environments. Our converged system, OverLay, is currently deployed in the engineering building and open for use to regular public; ongoing work is focussed on campus-wide deployment to serve as a "historical tour guide" of UIUC. Performance results and user responses thus far have been promising, to say the least.
Puneet Jain, Justin Manweiler, Romit Roy Choudhury
MobiSys1
2014 Scalable Social Analytics for Live Viral Event Prediction
Puneet Jain, Justin Manweiler, Arup Acharya, Romit Roy Choudhury
ICWSM1
2014 Demo: real-time object tagging and retrieval
abstract
We propose an augmented reality system on off-the-shelves smartphones which allows random physical object tagging. At later times, such tags could be retrieved from different locations and orientations. Our approach does not require any additional infrastructure support, localization scheme, specialized camera, or modification to smartphone's operating system. Designed and developed for current generation smartphones, our application shows promising initial results with retrieval accuracy of 82% in indoor environment without noticeable impact on the user experience.
Puneet Jain, Romit Roy Choudhury
MobiSys1
2013 FOCUS: clustering crowdsourced videos by line-of-sight
abstract
Crowdsourced video often provides engaging and diverse perspectives not captured by professional videographers. Broad appeal of user-uploaded video has been widely confirmed: freely distributed on YouTube, by subscription on Vimeo, and to peers on Facebook/Google+. Unfortunately, user-generated multimedia can be difficult to organize; these services depend on manual "tagging" or machine-mineable viewer comments. While manual indexing can be effective for popular, well-established videos, newer content may be poorly searchable; live video need not apply. We envisage video-sharing services for live user video streams, indexed automatically and in realtime, especially by shared content. We propose FOCUS, for Hadoop-on-cloud video-analytics. FOCUS uniquely leverages visual, 3D model reconstruction and multimodal sensing to decipher and continuously track a video's line-of-sight. Through spatial reasoning on the relative geometry of multiple video streams, FOCUS recognizes shared content even when viewed from diverse angles and distances. In a 70-volunteer user study, FOCUS' clustering correctness is roughly comparable to humans.
Puneet Jain, Justin Manweiler, Arup Acharya, Kirk A. Beaty
SenSys1
2013 FOCUS: clustering crowdsourced videos by line-of-sight
abstract
We present a demonstration of FOCUS [1], a system to appear in the SenSys 2013 main conference. FOCUS is a video-clustering service for live user video streams, indexed automatically and in realtime by shared content. FOCUS uniquely leverages visual, 3D model reconstruction and multimodal sensing to decipher and continuously track a video's line-of-sight. Through spatial reasoning on the relative geometry of multiple video streams, FOCUS recognizes shared content even when viewed from diverse angles and distances. We believe FOCUS can enable a new family of applications, such as instant replay, augmented reality, citizen journalism, security breach detection, and disaster assessment.
Puneet Jain, Justin Manweiler, Arup Acharya, Kirk A. Beaty
SenSys1
2012 Satellites in our pockets: an object positioning system using smartphones
abstract
This paper attempts to solve the following problem: can a distant object be localized by looking at it through a smartphone. As an example use-case, while driving on a highway entering New York, we want to look at one of the skyscrapers through the smartphone camera, and compute its GPS location. While the problem would have been far more difficult five years back, the growing number of sensors on smartphones, combined with advances in computer vision, have opened up important opportunities. We harness these opportunities through a system called Object Positioning System (OPS) that achieves reasonable localization accuracy. Our core technique uses computer vision to create an approximate 3D structure of the object and camera, and applies mobile phone sensors to scale and rotate the structure to its absolute configuration. Then, by solving (nonlinear) optimizations on the residual (scaling and rotation) error, we ultimately estimate the object's GPS position.
Justin Manweiler, Puneet Jain, Romit Roy Choudhury
MobiSys2
2012 Demo: satellites in our pockets: an object positioning system using smartphones
abstract
We attempt to localize a distant object by looking at it through a smartphone. As an example use-case, while driving on a highway entering New York, we want to look at one of the skyscrapers through the smartphone camera, and compute its GPS location. While the problem would have been far more difficult five years back, the growing number of sensors on smartphones, combined with advances in computer vision, have opened up important opportunities. We harness these opportunities through a system called Object Positioning System (OPS) [1] that achieves reasonable localization accuracy. Our core technique uses computer vision to create an approximate 3D structure of the object and camera, and applies mobile phone sensors to scale and rotate the structure to its absolute configuration. Then, by solving (nonlinear) optimizations on the residual (scaling and rotation) error, we ultimately estimate the object's GPS position.
Justin Manweiler, Puneet Jain, Romit Roy Choudhury
MobiSys2