VLDB 2026 Research / reviewers in the wild / expert
Karan Ahuja
dblp:185/0960
· DBLP profile ↗
31ranked-venue papers
13as first author
21since 2021 · last 2026
0000-0003-2497-0530ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 24 · 9 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SurfaceXR: Fusing Smartwatch IMUs and Egocentric Hand Pose for Seamless Surface InteractionsabstractMid-air gestures in Extended Reality (XR) often lead to fatigue, discomfort and imprecision, limiting their suitability for extended use. Surface-based interactions offer a compelling alternative, providing improved accuracy, speed, and comfort. However, current egocentric vision-based methods struggle with reliable surface inputs due to challenges in hand tracking and surface plane estimation from oblique and occluded viewing angles. To this extent, we introduce SurfaceXR, a novel sensor fusion approach that combines headset based hand tracking with micro-vibration data sampled from commodity smartwatch IMUs to enable precise and robust inputs on everyday surfaces. Our system is designed with flexibility in mind - it can function using only hand tracking, only IMU sensing, or optimally with both modalities combined, and remains robust even without explicit surface calibration. Our key insight is that these modalities are complementary - hand tracking provides 3D positional data of hand joints, whereas IMUs supply high-frequency wrist/hand motion data. Our user study across 21 participants validates SurfaceXR's effectiveness in augmenting surface touch tracking and 8 class hand-surface gesture recognition, demonstrating significant improvements over single-modality approaches. Enabled by SurfaceXR, we demonstrate a series of interactive apps for both AR and VR, ranging from on-surface sketching, text entry and gesture-based navigation. Vasco Xu, Eric J. Gonzalez, Andrea Colaco, Henry Hoffmann, Mar González-Franco, Karan Ahuja |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | Text Entry for XR Trove (TEXT): Collecting and Analyzing Techniques for Text Input in XRabstractText entry for extended reality (XR) is far from perfect, and a variety of text entry techniques (TETs) have been proposed to fit various contexts of use. However, comparing between TETs remains challenging due to the lack of a consolidated collection of techniques, and limited understanding of how interaction attributes of a technique (e.g., presence of visual feedback) impact user performance. To address these gaps, this paper examines the current landscape of XR TETs by creating a database of 176 different techniques. We analyze this database to highlight trends in the design of these techniques, the metrics used to evaluate them, and how various interaction attributes impact these metrics. We discuss implications for future techniques and present TEXT: Text Entry for XR Trove, an interactive online tool to navigate our database. Arpit Bhatia, Moaaz Hudhud Mughrabi, Diar Abdlkarim, Massimiliano Di Luca, Mar González-Franco, Karan Ahuja, Hasti Seifi |
CHI | 6 |
| 2025 | Online-EYE: Multimodal Implicit Eye Tracking Calibration for XRabstractUnlike other inputs for extended reality (XR) that work out of the box, eye tracking typically requires custom calibration per user or session. We present a multimodal inputs approach for implicit calibration of eye tracker in VR, leveraging UI interaction for continuous, background calibration. Our method analyzes gaze data alongside controller interaction with UI elements, and employing ML techniques it continuously refines the calibration matrix without interrupting users from their current tasks. Potentially eliminating the need for explicit calibration. We demonstrate the accuracy and effectiveness of this implicit approach across various tasks and real time applications achieving comparable eye tracking accuracy to native, explicit calibration. While our evaluation focuses on VR and controller-based interactions, we anticipate the broader applicability of this approach to various XR devices and input modalities. Baosheng James Hou, Lucy Abramyan, Prasanthi Gurumurthy, Haley Adams, Ivana Tosic Rodgers, Eric J. Gonzalez, Khushman Patel, Andrea Colaco, Ken Pfeuffer, Hans-Werner Gellersen, Karan Ahuja, Mar González-Franco |
CHI | 11 |
| 2025 | WatchHAR: Real-time On-device Human Activity Recognition System for SmartwatchesabstractDespite advances in practical and multimodal fine-grained Human Activity Recognition (HAR), a system that runs entirely on smartwatches in unconstrained environments remains elusive. We present WatchHAR, an audio and inertial-based HAR system that operates fully on smartwatches, addressing privacy and latency issues associated with external data processing. By optimizing each component of the pipeline, WatchHAR achieves compounding performance gains. We introduce a novel architecture that unifies sensor data preprocessing and inference into an end-to-end trainable module, achieving 5x faster processing while maintaining over 90% accuracy across more than 25 activity classes. WatchHAR outperforms state-of-the-art models for event detection and activity classification while running directly on the smartwatch, achieving 9.3 ms processing time for activity event detection and 11.8 ms for multimodal activity classification. This research advances on-device activity recognition, realizing smartwatches' potential as standalone, privacy-aware, and minimally-invasive continuous activity tracking devices. Taeyoung Yeon, Vasco Xu, Henry Hoffmann, Karan Ahuja |
ICMI | 4 |
| 2025 | EI-Lite: Electrical Impedance Sensing for Micro-gesture Recognition and Pinch Force Estimation
Junyi Zhu 0001, Tianyu Xu 0008, Emily Guan, JaeYoung Moon, Stiven Morvan, D. Shin, Andrea Colaco, Stefanie Mueller 0001, Karan Ahuja, Yiyue Luo, Ishan Chatterjee |
UIST | 10 |
| 2025 | EmBARDiment: an Embodied AI Agent for Productivity in XRabstractXR devices running chat-bots powered by Large Language Models (LLMs) have the to become always-on agents that enable much better productivity scenarios. Current screen based chat-bots do not take advantage of the the full-suite of natural inputs available in XR, including inward facing sensor data, instead they over-rely on explicit voice or text prompts, sometimes paired with multi-modal data dropped as part of the query. We propose a solution that leverages an attention framework that derives context implicitly from user actions, eye-gaze, and contextual memory within the XR environment. Our work minimizes the need for engineered explicit prompts, fostering grounded and intuitive interactions that glean user insights for the chat-bot. Riccardo Bovo, Steven Abreu, Karan Ahuja, Eric J. Gonzalez, Li-Te Cheng, Mar González-Franco |
VR | 3 |
| 2024 | MotionTrace: IMU-based Trajectory Prediction for Smartphone AR InteractionsabstractSmartphone powered mobile Augmented Reality (AR) applications require precise and accurate 6DoF localization of the phone within a 3D space. This typically requires camera-based localization systems that are power intensive and prone to motion blurs, thereby limiting their usage. In such cases, IMU-based methods are a promising alternative for low power localization sensing. In this paper, we present MotionTrace, a method for predicting a user's handheld phone trajectory using a smartphone's inertial sensor. We evaluated MotionTrace over future hand positions at 50, 100, 200, 400, and 800ms time horizons using the large motion capture (AMASS) and smartphone-based full-body pose estimation (Pose-on-the-Go) datasets. Our results show that MotionTrace can estimate the future phone position of the user with an average MSE between 0.11 – 143.62 mm across different time horizons. Rahul Islam, Vasco Miguel Liang Xu, Karan Ahuja |
BSN | 3 |
| 2024 | EITPose: Wearable and Practical Electrical Impedance Tomography for Continuous Hand Pose EstimationabstractReal-time hand pose estimation has a wide range of applications spanning gaming, robotics, and human-computer interaction. In this paper, we introduce EITPose, a wrist-worn, continuous 3D hand pose estimation approach that uses eight electrodes positioned around the forearm to model its interior impedance distribution during pose articulation. Unlike wrist-worn systems relying on cameras, EITPose has a slim profile (12 mm thick sensing strap) and is power-efficient (consuming only 0.3 W of power), making it an excellent candidate for integration into consumer electronic devices. In a user study involving 22 participants, EITPose achieves with a within-session mean per joint positional error of 11.06 mm. Its camera-free design prioritizes user privacy, yet it maintains cross-session and cross-user accuracy levels comparable to camera-based wrist-worn systems, thus making EITPose a promising technology for practical hand pose estimation. Alexander Kyu, Hongyu Mao, Junyi Zhu 0001, Mayank Goel, Karan Ahuja |
CHI | 5 |
| 2024 | Geometry Fidelity for Spherical Images
Anders Christensen, Nooshin Mojab, Khushman Patel, Karan Ahuja, Zeynep Akata, Ole Winther, Mar González-Franco, Andrea Colaco |
ECCV (80) | 4 |
| 2024 | Augmented Object Intelligence with XR-ObjectsabstractSeamless integration of physical objects as interactive digital entities remains a challenge for spatial computing. This paper explores Augmented Object Intelligence (AOI) in the context of XR, an interaction paradigm that aims to blur the lines between digital and physical by equipping real-world objects with the ability to interact as if they were digital, where every object has the potential to serve as a portal to digital functionalities. Our approach utilizes real-time object segmentation and classification, combined with the power of Multimodal Large Language Models (MLLMs), to facilitate these interactions without the need for object pre-registration. We implement the AOI concept in the form of XR-Objects, an open-source prototype system that provides a platform for users to engage with their physical environment in contextually relevant ways using object-based context menus. This system enables analog objects to not only convey information but also to initiate digital actions, such as querying for details or executing tasks. Our contributions are threefold: (1) we define the AOI concept and detail its advantages over traditional AI assistants, (2) detail the XR-Objects system’s open-source design and implementation, and (3) show its versatility through various use cases and a user study. Mustafa Doga Dogan, Eric J. Gonzalez, Karan Ahuja, Ruofei Du, Andrea Colaco, Johnny Lee, Mar González-Franco, David Kim 0002 |
UIST | 3 |
| 2024 | MobilePoser: Real-Time Full-Body Pose Estimation and 3D Human Translation from IMUs in Mobile Consumer DevicesabstractThere has been a continued trend towards minimizing instrumentation for full-body motion capture, going from specialized rooms and equipment, to arrays of worn sensors and recently sparse inertial pose capture methods. However, as these techniques migrate towards lower-fidelity IMUs on ubiquitous commodity devices, like phones, watches, and earbuds, challenges arise including compromised online performance, temporal consistency, and loss of global translation due to sensor noise and drift. Addressing these challenges, we introduce MobilePoser, a real-time system for full-body pose and global translation estimation using any available subset of IMUs already present in these consumer devices. MobilePoser employs a multi-stage deep neural network for kinematic pose estimation followed by a physics-based motion optimizer, achieving state-of-the-art accuracy while remaining lightweight. We conclude with a series of demonstrative applications to illustrate the unique potential of MobilePoser across a variety of fields, such as health and wellness, gaming, and indoor navigation to name a few. Vasco Xu, Chenfeng Gao, Henry Hoffmann, Karan Ahuja |
UIST | 4 |
| 2023 | IMUPoser: Full-Body Pose Estimation using IMUs in Phones, Watches, and EarbudsabstractTracking body pose on-the-go could have powerful uses in fitness, mobile gaming, context-aware virtual assistants, and rehabilitation. However, users are unlikely to buy and wear special suits or sensor arrays to achieve this end. Instead, in this work, we explore the feasibility of estimating body pose using IMUs already in devices that many users own — namely smartphones, smartwatches, and earbuds. This approach has several challenges, including noisy data from low-cost commodity IMUs, and the fact that the number of instrumentation points on a user’s body is both sparse and in flux. Our pipeline receives whatever subset of IMU data is available, potentially from just a single device, and produces a best-guess pose. To evaluate our model, we created the IMUPoser Dataset, collected from 10 participants wearing or holding off-the-shelf consumer devices and across a variety of activity contexts. We provide a comprehensive evaluation of our system, benchmarking it on both our own and existing IMU datasets. Vimal Mollyn, Riku Arakawa, Mayank Goel, Chris Harrison 0001, Karan Ahuja |
CHI | 5 |
| 2022 | ControllerPose: Inside-Out Body Capture with VR Controller CamerasabstractWe present a new and practical method for capturing user body pose in virtual reality experiences: integrating cameras into handheld controllers, where batteries, computation and wireless communication already exist. By virtue of the hands operating in front of the user during many VR interactions, our controller-borne cameras can capture a superior view of the body for digitization. Our pipeline composites multiple camera views together, performs 3D body pose estimation, uses this data to control a rigged human model with inverse kinematics, and exposes the resulting user avatar to end user applications. We developed a series of demo applications illustrating the potential of our approach and more leg-centric interactions, such as balancing games and kicking soccer balls. We describe our proof-of-concept hardware and software, as well as results from our user study, which point to imminent feasibility. Karan Ahuja, Vivian Shen, Cathy Mengying Fang, Nathan Riopelle, Andy Kong, Chris Harrison 0001 |
CHI | 1 |
| 2022 | TriboTouch: Micro-Patterned Surfaces for Low Latency TouchscreensabstractTouchscreen tracking latency, often 80ms or more, creates a rubber-banding effect in everyday direct manipulation tasks such as dragging, scrolling, and drawing. This has been shown to decrease system preference, user performance, and overall realism of these interfaces. In this research, we demonstrate how the addition of a thin, 2D micro-patterned surface with 5 micron spaced features can be used to reduce motor-visual touchscreen latency. When a finger, stylus, or tangible is translated across this textured surface frictional forces induce acoustic vibrations which naturally encode sliding velocity. This acoustic signal is sampled at 192kHz using a conventional audio interface pipeline with an average latency of 28ms. When fused with conventional low-speed, but high-spatial-accuracy 2D touch position data, our machine learning model can make accurate predictions of real time touch location. Craig D. Shultz, Daehwa Kim, Karan Ahuja, Chris Harrison 0001 |
CHI | 3 |
| 2022 | RGBDGaze: Gaze Tracking on Smartphones with RGB and Depth DataabstractTracking a user’s gaze on smartphones offers the potential for accessible and powerful multimodal interactions. However, phones are used in a myriad of contexts and state-of-the-art gaze models that use only the front-facing RGB cameras are too coarse and do not adapt adequately to changes in context. While prior research has showcased the efficacy of depth maps for gaze tracking, they have been limited to desktop-grade depth cameras, which are more capable than the types seen in smartphones, that must be thin and low-powered. In this paper, we present a gaze tracking system that makes use of today’s smartphone depth camera technology to adapt to the changes in distance and orientation relative to the user’s face. Unlike prior efforts that used depth sensors, we do not constrain the users to maintain a fixed head position. Our approach works across different use contexts in unconstrained mobile settings. The results show that our multimodal ML model has a mean gaze error of 1.89 cm; a 16.3% improvement over using RGB data alone (2.26 cm error). Our system and dataset offer the first benchmark of gaze tracking on smartphones using RGB+Depth data under different use contexts. Riku Arakawa, Mayank Goel, Chris Harrison 0001, Karan Ahuja |
ICMI | 4 |
| 2021 | Vid2Doppler: Synthesizing Doppler Radar Data from Videos for Training Privacy-Preserving Activity RecognitionabstractMillimeter wave (mmWave) Doppler radar is a new and promising sensing approach for human activity recognition, offering signal richness approaching that of microphones and cameras, but without many of the privacy-invading downsides. However, unlike audio and computer vision approaches that can draw from huge libraries of videos for training deep learning models, Doppler radar has no existing large datasets, holding back this otherwise promising sensing modality. In response, we set out to create a software pipeline that converts videos of human activities into realistic, synthetic Doppler radar data. We show how this cross-domain translation can be successful through a series of experimental results. Overall, we believe our approach is an important stepping stone towards significantly reducing the burden of training such as human sensing systems, and could help bootstrap uses in human-computer interaction. Karan Ahuja, Yue Jiang 0002, Mayank Goel, Chris Harrison 0001 |
CHI | 1 |
| 2021 | Pose-on-the-Go: Approximating User Pose with Smartphone Sensor Fusion and Inverse KinematicsabstractWe present Pose-on-the-Go, a full-body pose estimation system that uses sensors already found in today’s smartphones. This stands in contrast to prior systems, which require worn or external sensors. We achieve this result via extensive sensor fusion, leveraging a phone’s front and rear cameras, the user-facing depth camera, touchscreen, and IMU. Even still, we are missing data about a user’s body (e.g., angle of the elbow joint), and so we use inverse kinematics to estimate and animate probable body poses. We provide a detailed evaluation of our system, benchmarking it against a professional-grade Vicon tracking system. We conclude with a series of demonstration applications that underscore the unique potential of our approach, which could be enabled on many modern smartphones with a simple software update. Karan Ahuja, Sven Mayer, Mayank Goel, Chris Harrison 0001 |
CHI | 1 |
| 2021 | Classroom Digital Twins with Instrumentation-Free Gaze TrackingabstractClassroom sensing is an important and active area of research with great potential to improve instruction. Complementing professional observers – the current best practice – automated pedagogical professional development systems can attend every class and capture fine-grained details of all occupants. One particularly valuable facet to capture is class gaze behavior. For students, certain gaze patterns have been shown to correlate with interest in the material, while for instructors, student-centered gaze patterns have been shown to increase approachability and immediacy. Unfortunately, prior classroom gaze-sensing systems have limited accuracy and often require specialized external or worn sensors. In this work, we developed a new computer-vision-driven system that powers a 3D “digital twin” of the classroom and enables whole-class, 6DOF head gaze vector estimation without instrumenting any of the occupants. We describe our open source implementation, and results from both controlled studies and real-world classroom deployments. Karan Ahuja, Deval Shah, Sujeath Pareddy, Franceska Xhakaj, Amy Ogan, Yuvraj Agarwal, Chris Harrison 0001 |
CHI | 1 |
| 2021 | PrivacyMic: Utilizing Inaudible Frequencies for Privacy Preserving Daily Activity RecognitionabstractSound presents an invaluable signal source that enables computing systems to perform daily activity recognition. However, microphones are optimized for human speech and hearing ranges: capturing private content, such as speech, while omitting useful, inaudible information that can aid in acoustic recognition tasks. We simulated acoustic recognition tasks using sounds from 127 everyday household/workplace objects, finding that inaudible frequencies can act as a substitute for privacy-sensitive frequencies. To take advantage of these inaudible frequencies, we designed a Raspberry Pi-based device that captures inaudible acoustic frequencies with settings that can remove speech or all audible frequencies entirely. We conducted a perception study, where participants “eavesdropped’’ on PrivacyMic’s filtered audio and found that none of our participants could transcribe speech. Finally, PrivacyMic’s real-world activity recognition performance is comparable to our simulated results, with over 95% classification accuracy across all environments, suggesting immediate viability in performing privacy-preserving daily activity recognition. Yasha Iravantchi, Karan Ahuja, Mayank Goel, Chris Harrison 0001, Alanson P. Sample |
CHI | 2 |
| 2021 | EyeMU Interactions: Gaze + IMU Gestures on Mobile DevicesabstractAs smartphone screens have grown in size, single-handed use has become more cumbersome. Interactive targets that are easily seen can be hard to reach, particularly notifications and upper menu bar items. Users must either adjust their grip to reach distant targets, or use their other hand. In this research, we show how gaze estimation using a phone’s user-facing camera can be paired with IMU-tracked motion gestures to enable a new, intuitive, and rapid interaction technique on handheld phones. We describe our proof-of-concept implementation and gesture set, built on state-of-the-art techniques and capable of self-contained execution on a smartphone. In our user study, we found a mean euclidean gaze error of 1.7 cm and a seven-class motion gesture classification accuracy of 97.3%. Andy Kong, Karan Ahuja, Mayank Goel, Chris Harrison 0001 |
ICMI | 2 |
| 2021 | TouchPose: Hand Pose Prediction, Depth Estimation, and Touch Classification from Capacitive ImagesabstractToday’s touchscreen devices commonly detect the coordinates of user input through capacitive sensing. Yet, these coordinates are the mere 2D manifestations of the more complex 3D configuration of the whole hand—a sensation that touchscreen devices so far remain oblivious to. In this work, we introduce the problem of reconstructing a 3D hand skeleton from capacitive images, which encode the sparse observations captured by touch sensors. These low-resolution images represent intensity mappings that are proportional to the distance to the user’s fingers and hands. Karan Ahuja, Paul Streli, Christian Holz 0001 |
UIST | 1 |
| 2020 | Gaze-based Screening of Autistic Traits for Adolescents and Young Adults using Prosaic VideosabstractAutism Spectrum Disorder (ASD) is a universal and often lifelong neuro-developmental disorder. Individuals with ASD often present comorbidities such as epilepsy, depression, and anxiety. In the United States, in 2014, 1 out of 68 people was affected by autism, but worldwide, the number of affected people drops to 1 in 160. This disparity is primarily due to underdiagnosis and unreported cases in resource-constrained environments. Wiggins et al. 1 found that, in the US, children of color are under-identified with ASD. Missing a diagnosis is not without consequences; approximately 26% of adults with ASD are under-employed, and are under-enrolled in higher education. Karan Ahuja, Abhishek Bose, Kuntal Dey, Anil Joshi, Krishnaveni Achary, Blessin Varkey, Chris Harrison 0001, Mayank Goel |
COMPASS | 1 |
| 2020 | Direction-of-Voice (DoV) Estimation for Intuitive Speech Interaction with Smart Devices EcosystemsabstractFuture homes and offices will feature increasingly dense ecosystems of IoT devices, such as smart lighting, speakers, and domestic appliances. Voice input is a natural candidate for interacting with out-of-reach and often small devices that lack full-sized physical interfaces. However, at present, voice agents generally require wake-words and device names in order to specify the target of a spoken command (e.g., 'Hey Alexa, kitchen lights to full bright-ness'). In this research, we explore whether speech alone can be used as a directional communication channel, in much the same way visual gaze specifies a focus. Instead of a device's microphones simply receiving and processing spoken commands, we suggest they also infer the Direction of Voice (DoV). Our approach innately enables voice commands with addressability (i.e., devices know if a command was directed at them) in a natural and rapid manner. We quantify the accuracy of our implementation across users, rooms, spoken phrases, and other key factors that affect performance and usability. Taken together, we believe our DoV approach demonstrates feasibility and the promise of making distributed voice interactions much more intuitive and fluid. Karan Ahuja, Andy Kong, Mayank Goel, Chris Harrison 0001 |
UIST | 1 |
| 2019 | MeCap: Whole-Body Digitization for Low-Cost VR/AR HeadsetsabstractLow-cost, smartphone-powered VR/AR headsets are becoming more popular. These basic devices - little more than plastic or cardboard shells - lack advanced features, such as controllers for the hands, limiting their interactive capability. Moreover, even high-end consumer headsets lack the ability to track the body and face. For this reason, interactive experiences like social VR are underdeveloped. We introduce MeCap, which enables commodity VR headsets to be augmented with powerful motion capture ("MoCap") and user-sensing capabilities at very low cost (under $5). Using only a pair of hemi-spherical mirrors and the existing rear-facing camera of a smartphone, MeCap provides real-time estimates of a wearer's 3D body pose, hand pose, facial expression, physical appearance and surrounding environment - capabilities which are either absent in contemporary VR/AR systems or which require specialized hardware and controllers. We evaluate the accuracy of each of our tracking features, the results of which show imminent feasibility. Karan Ahuja, Chris Harrison 0001, Mayank Goel, Robert Xiao |
UIST | 1 |
| 2019 | LightAnchors: Appropriating Point Lights for Spatially-Anchored Augmented Reality InterfacesabstractAugmented reality requires precise and instant overlay of digital information onto everyday objects. We present our work on LightAnchors, a new method for displaying spatially-anchored data. We take advantage of pervasive point lights - such as LEDs and light bulbs - for both in-view anchoring and data transmission. These lights are blinked at high speed to encode data. We built a proof-of-concept ap-plication that runs on iOS without any hardware or software modifications. We also ran a study to characterize the performance of LightAnchors and built eleven example demos to highlight the potential of our approach. Karan Ahuja, Sujeath Pareddy, Robert Xiao, Mayank Goel, Chris Harrison 0001 |
UIST | 1 |
| 2018 | Ubicoustics: Plug-and-Play Acoustic Activity RecognitionabstractDespite sound being a rich source of information, computing devices with microphones do not leverage audio to glean useful insights about their physical and social context. For example, a smart speaker sitting on a kitchen countertop cannot figure out if it is in a kitchen, let alone know what a user is doing in a kitchen - a missed opportunity. In this work, we describe a novel, real-time, sound-based activity recognition system. We start by taking an existing, state-of-the-art sound labeling model, which we then tune to classes of interest by drawing data from professional sound effect libraries traditionally used in the entertainment industry. These well-labeled and high-quality sounds are the perfect atomic unit for data augmentation, including amplitude, reverb, and mixing, allowing us to exponentially grow our tuning data in realistic ways. We quantify the performance of our approach across a range of environments and device categories and show that microphone-equipped computing devices already have the requisite capability to unlock real-time activity recognition comparable to human accuracy. Gierad Laput, Karan Ahuja, Mayank Goel, Chris Harrison 0001 |
UIST | 2 |
| 2017 | OptiDwell: Intelligent Adjustment of Dwell Click TimeabstractGaze based navigation with digital screens offer a hands-free and touchless interaction, which is often useful in providing a hygienic interaction experience in a public kiosk scenario. The goodness of such a navigation system depends not only on the accuracy of detecting the eye gaze but also on the ability to determine whether a user is interested in clicking a button or is just looking at the button. The time for which a user needs to gaze at a particular button before it is considered as a click action is called the dwell time. In this paper, we explore intelligent adjustment of dwell times, where mouse click events on the buttons of a given application are emulated with user gaze. A constant dwell-time for all buttons and for all users may not provide an efficient and intuitive interface. We thereby propose a model to dynamically adjust dwell-time values used to emulate user mouse click events, exploiting the user's experience with different portions of a given application. The adjustment happens at a per-user, per-button granularity, as a function of the user's (a) prior usage experience of the given button within the application and (b) Midas touch characteristics for the given button. We propose OptiDwell, inspired by the action-value method based solutions to the Multi-Armed Bandits problem, for dwell click time adaptation. We experiment OptiDwell using an interactive TV channel browsing interface application, constituting of a mix of text and image buttons, over 10 computer-savvy users generating over 9000 click tasks. We observe significant improvement of user comfort level over the sessions, quantified by (a) improved (reduced) dwell times and (b) reduced number of Midas touches in spite of faster dwell-clicks, as high as 10-fold reduction in the best case. Our work is useful for creating an interface, with accurate, fast and comfortable dwell-clicks for each interface element (e.g., buttons), and each user. Aanand Nayyar, Utkarsh Dwivedi, Karan Ahuja, Nitendra Rajput, Seema Nagar, Kuntal Dey |
IUI | 3 |
| 2017 | Convolutional neural networks for ocular smartphone-based biometrics
Karan Ahuja, Rahul Islam, Ferdous A. Barbhuiya, Kuntal Dey |
Pattern Recognit. Lett. | 1 |
| 2016 | ISURE: User authentication in mobile devices using ocular biometrics in visible spectrumabstractIn this paper, we propose a supervised learning based model for ocular biometrics. Using Speeded-Up Robust Features (SURF) for detecting local features of the eye region, we create a local feature descriptor vector of each image. We cluster these feature vectors, representing an image as a normalized histogram of membership to various clusters, thereby creating a bag-of-visual-words model. We conduct a multiphase training, first performing a fast Multinomial Naïve Bayes learning, and subsequently using a pyramid-up topology to use the top k% results (based upon confidence scores) thus predicted and perform Dense SIFT for nearest neighbor matching. Contrary to traditional ocular biometric systems, our proposed approach does not rely highly accurate iris pattern segmentation, allowing less constrained image acquisition conditions such as from mobile devices. Our method identifies the individuals with an identification accuracy varying from 48.76% to 79.49%, across different lighting conditions and phone handset data sources, while testing on the given data. Karan Ahuja, Abhishek Bose, Seema Nagar, Kuntal Dey, Ferdous A. Barbhuiya |
ICIP | 1 |
| 2016 | Eye center localization and detection using radial mappingabstractWe propose a geometrical method, applied over eye-specific features, to improve the accuracy of the art of eye-center localization. Our solution is built upon: (a) checking radially constrained gradient vectors, (b) adding weightage to iris specific features and (c) considering bi-directional image gradients to eliminate errors due to reflection on pupil. Our system outperforms the state of the art methods, when compared collectively across multiple benchmark databases, such as BioID and FERET. Our process is lightweight, robust and significantly fast: achieving 50-60 fps for eye center localization, using a single threaded approach on a 2.4 GHz CPU with no GPU. This makes it practicable for real-life applications. Karan Ahuja, Ruchika Banerjee, Seema Nagar, Kuntal Dey, Ferdous A. Barbhuiya |
ICIP | 1 |
| 2016 | A preliminary study of CNNs for iris and periocular verification in the visible spectrumabstractOcular biometrics in the visible spectrum has emerged as an area of significant research activity. In this paper, we propose two convolution-based models for verifying a pair of periocular images containing the iris, and compare the two approaches amongst each other as well as with a baseline model. In the first approach, we perform deep learning in an unsupervised manner using a stacked convolutional architecture, using external models learned a-priori on external facial and periocular data, on top of the baseline model applied on the provided data, and apply different score fusion models. In the second approach, we again use a stacked convolution architecture; but here, we learn the feature vector in a supervised manner. We obtain an AUROC of 0.946 and 0.981, and EER of 0.092 and 0.066, for the two models respectively. We further combine the two models, and observe the combined model to deliver the best performance in case the both the images arise from the same device type, but not necessarily so otherwise, obtaining a AUROC of 0.985 and EER of 0.057. Given the significant performance our methodology yields, our system can be used in real-life applications with minimal error. Karan Ahuja, Rahul Islam, Ferdous A. Barbhuiya, Kuntal Dey |
ICPR | 1 |