Natasha Kholgade

dblp:69/7590 · also Natasha Banerjee, Natasha Kholgade Banerjee · DBLP profile ↗
← Back
32ranked-venue papers
1as first author
20since 2021 · last 2026
0000-0001-7730-7754ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 13 · 11 since 2021Human-computer interaction and ubiquitous computing · 11 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Cross-system forecasting-based user authentication for virtual reality (VR)
abstract
As virtual reality (VR) systems gain acceptance into critical domains, such as healthcare, education, banking, and military, sensitive user data has to be protected from malicious users. Existing security measures such as usernames/passwords, PINs, or multi-factor approaches do not provide any protection when the attacker gains access to the credentials or when the user intentionally hands over their credentials to an ally, for instance to take an exam on their behalf. Recognizing these challenges, a large body of work has emerged over the past decade on behavioral biometrics for single VR systems. However, cross-system behavioral biometrics for VR remains at a nascent stage. Cross-system behavioral biometrics is necessary to enable users to seamlessly transition between their office, home, job site, or clinic issued system. Early work in cross-system behavioral biometrics for VR showed that while deep learning techniques such as Siamese neural networks were effective, they required near complete trajectories for high assurance user authentication or identification. The emergence of motion forecasting approaches for VR biometrics enabled the use of limited data from the start of the user’s action, thereby preventing the attacker from gaining access to the full user trajectory. However, such forecasting models were designed for a single VR system that did not enable cross-system authentication. In recent work, we showed that cross-system forecasting-based authentication for VR biometrics can be performed using an Informer-based model to train the forecasting component and a fully convolutional network to train the authenticator. Using a publicly available dataset of 41 users performing a ball throwing task using the Meta Quest, HTC Vive, and HTC Vive Cosmos, we showed that in comparison to non-forecasted Siamese networks, our approach reduces the equal error rate (EER) by an average of 53.16% across all VR system combinations over prior cross-system authentication work. In this paper, we compare the performance of the Informer-based cross-system forecasting model, which operates on a point-wise input and treats each timestamp as a separate token, against a patching architecture, namely PatchTST, that that provides lower runtime by splitting the input time series into individual channels, operates on each channel independently, and encodes channel features into patches. Using the same ball throwing dataset, we show that PatchTST shows an average EER reduction of 30.30% over prior cross-system authentication. Using PatchTST, we obtain a speedup of more than 3 times compared to Informer when using a GPU with half the core count. Link to code: https://github.com/Terascale-All-sensing-Research-Studio/Cross_System_Authentication_using_Forecasting/ .
Natasha Kholgade, Sean Banerjee
Comput. Graph.2
2026 Understanding User Perceptions Toward Queue Cutting in Immersive VR Environments
abstract
ABSTRACT Queue cutting is a frustrating experience in the real‐world when a person is waiting to receive service as it violates the normative behavior of first‐in‐first‐out ordering of queues. As social experiences move to virtual worlds users may experience human and non‐playable characters (NPCs) attempting to cut the queue, thereby adding to the negative experiences of waiting. While normative behavior for activities such as joining an existing conversational group or walking between agents engaged in conversation have been studied in virtual reality, perceptions toward queue cutters have not been studied. We conduct a study on understanding how users perceive a queue‐cutting NPC in a virtual doctor's office reception area. During the wait, a queue‐cutting NPC requests an in‐queue NPC to cut in ahead of them and is either allowed or denied entry into the queue. Using data on subjective responses from 45 participants we find significant differences in perception of time, frustration, and likelihood of exit when the queue‐cutting NPC is allowed in. Our work enables future research on detecting situations capable of generating user frustration, and providing appropriate intervention via the VR environment, mitigating negative experiences, and ensuring timely service.
Elza Ibragimov, Natasha Kholgade, Sean Banerjee, Ashutosh Shivakumar
Comput. Animat. Virtual Worlds2
2025 Cross-System Virtual Reality (VR) Authentication Using Transformer-Based Trajectory Forecasting
Natasha Kholgade, Sean Banerjee
EuroXR2
2025 Safety and Naturalness Perceptions of Robot-to-Human Handovers Performed by Data-Driven Robotic Mimicry of Human Givers
abstract
We study human perceptions of a robot that performs robot-to-human (R2H) handovers controlled to grasp, transport, and transfer 34 objects by mimicking human givers in human-human (H2H) handover data. Recognizing the importance of human-like robotic behavior for successful collaboration, R2H studies use models of human behavior or observations of H2H data to plan robot giver motion. However, R2H studies have been limited in object counts. In this work, we use the Human-Object-Human (HOH) dataset, consisting of H2H interactions performed by 20 giver-receiver pairs with 136 objects, to conduct an R2H study with 34 objects. We teleoperate a Kinova Gen3 manipulator to grip an object as grasped by an HOH human giver, and program it to automatically transport and orient the object to a participant by mimicking the HOH giver's trajectory and transfer pose. We survey participants on safety, naturalness, and preferred choice over linear trajectory and random orientation baselines. We find that transfer pose influences perceptions of naturalness, with HOH poses showing higher naturalness ratings. Participants prefer handovers with HOH end poses when asked to pick their preferred interaction.
Ava Megyeri, Noah Wiederhold, Yu Liu 0069, Sean Banerjee, Natasha Kholgade
ICRA5
2025 Analysis of Pre-Handover Peak Speed Timing and Patterns for Human Givers and Receivers
abstract
In this work, we study patterns in the movement times of human participants engaging in object handover during the pre-handover phase, to inform on conducting fluent human-robot handovers. Fluency of object handover between two agents is critical to ensure success of shared overall collaborative goals in the context of larger tasks involving the transferred objects. Human givers and receivers often coordinate their pre-handover movements to ensure fluent transfer, with receivers showing anticipatory behavior and proactive response. Since typical timing patterns of human and robot movements for short-range tasks demonstrate a speedup to a peak speed, and a slowdown to the end point, we analyze relationships between time differences of peak speed attainment relative to start and end (i.e., reach) times, across the giver and receiver. Our work helps inform the design of motion planning algorithms for robots that embody timing patterns during movement that are predictable to the partner.
Ava Megyeri, Sean Banerjee, Maria Kyrarini, Natasha Kholgade
RO-MAN4
2025 HI-Grasp: Human-Inspired Grasp Network for Intuitive and Stable Robotic Grasp
abstract
As robots increasingly integrate into human environments, they must interact safely, intuitively, and effectively. We present HI-Grasp, a novel human-inspired grasp network that generates robotic grasps that closely mimic human grasp behaviors observed in human-to-human object handovers while ensuring stability. Our approach combines a deep grasp prediction network with Transformer-based modules to predict human-like, stable grasps. We contribute the dataset HOH-Grasps to train and evaluate HI-Grasp. HOH-Grasps consists of interaction data from the HOH handover dataset annotated with grasp labels, and is available at https://huggingface.co/datasets/tars-home/HOH-Grasps. Using HOH-Grasps, HI-Grasp learns and reproduces human grasp preferences while maintaining grasp stability. Extensive experiments on HOH-Grasps, along with real-world robot tests, show that HI-Grasp outperforms baselines and ablated variants in terms of human alignment and stability. We propose HI-Grasp-Lift, a robot-to-human object handover strategy built on HI-Grasp's predictions, to showcase the practical applicability of our approach.
Xinchao Song 0001, Ava Megyeri, Noah Wiederhold, Sean Banerjee, Natasha Kholgade
RO-MAN5
2025 Studying Disagreement in Grasp Intentions of Givers and Receivers Engaging in Human-Human Handover
abstract
Successful handover of objects between two agents—two humans, or a human and a robot—plays an important role in ensuring success of larger-scale collaborative physical interactions between the agents. Handovers are likely to be seen as successful if the intentions of the agents participating in the handover are met, e.g., if the receiver intends to hold the object at the location presented by the giver. In this work, we analyze participant responses on giver and receiver comfort, giver’s intention being met on receiver’s receipt grasp, and receiver’s intentions being met on giving and receipt grasp from a dataset of 1,632 handovers performed by 48 human giver-receiver dyads using 204 objects. Our findings show evidence of misalignment in intentions in unprompted handovers. In 19.47% or nearly a fifth of the analyzed grasps, the receiver would not have performed the giving grasp where the giver did. In 8.66% and 10.20%, i.e., nearly an eleventh and a tenth of cases, the receipt grasp was not in alignment with the giver and the receiver’s intentions respectively. Our findings show the need for grasp and manipulation algorithms for robots engaging in handover to be aware of object structure and expected behavior to minimize misalignment in intentions for successful handover.
Noah Wiederhold, Sean Banerjee, Natasha Kholgade
RO-MAN3
2025 What should an industrial robot record? Understanding worker perceptions toward privacy
abstract
Modern robots found in the workplace include high-resolution color cameras, depth cameras, thermal sensors, and microphones, as well as onboard computing that enables generation of human body pose and facial landmarks that can be used to infer the physiological state of the worker. Prior research reveals that privacy concerns in the workplace tend higher towards robots than towards humans, and that robots that include privacy controllers are perceived to be more trustworthy. However, prior research lacks a systematic understanding of how workers in lower-income positions in industries such as warehousing and manufacturing perceive sensing technologies, such as color cameras and microphones, and inferred human data, such as body pose, in their concerns for privacy. In this paper, we use survey responses from 530 workers across four countries and six job domains to understand how workers perceive robots that record human activity through video, audio, or body pose. Our studies reveal that privacy concern grows with the amount of personally identifiable data that is available, i.e., workers show higher concern for video rather than body pose inferred from video. We find that Millennials show more concern about being recorded than Gen Z individuals or than Gen X individuals and older. We find that privacy concerns vary by the worker’s country with lowest concerns in the United States and highest in Canada, and that worker concern is lowest in large-size cities and highest in suburban areas.
Tyler Yankee, Maria Kyrarini, Natasha Kholgade, Sean Banerjee
RO-MAN3
2025 Motion Forecasting Attacks on Behavioral Biometric Authentication Systems in Virtual Reality
abstract
Inspired by behavioral biometrics for keystroke and touch-based systems, a large body of work has emerged over the past decade on using user behavior in VR applications as a signature of the genuine user. Recent work on forecasting approaches for behavioral biometrics for VR helps address a key challenge in existing approaches where complete user movement signatures are needed to authenticate the user. Forecasting-based approaches enable VR authentication systems to use limited user behavior data and forecast future movement trajectories. However, forecasting-based approaches present a new concern where malicious users can exploit the predictability of user motions to launch an attack. In this paper, we present the first forecasting-based attack model against VR authentication systems that rely on behavioral biometrics. We propose a two-phase approach to assess authentication performance and adversarial risk. Phase 1 develops a Fully Convolutional Network for authentication using VR motion data, evaluating stochastic gradient descent (SGD) and Adam optimizers with Equal Error Rate (EER) as the primary metric. Phase 2 introduces a forecasting attack, where partial motion sequences of an impostor’s motion are fed to a Transformer model to generate future trajectories that represent genuine user behavior for an authenticator enabling an impostor to deceive the authentication system. Experimental results demonstrate the attack’s effectiveness, achieving an EER as low as 0.0346, exposing security risks in motion-based authentication. These findings underscore the urgent need for robust countermeasures to defend against predictive motion attacks in VR environments. Our code is shared at: https://bit.ly/4n1GtxG.
Ashutosh Shivakumar, Natasha Kholgade, Sean Banerjee
VRST3
2024 Analyzing Perceptions on Barriers to Safety in the Workforce and Expectations on Human and Robotic Assistance
abstract
Prior research on the fear of robots in blue-collar workplaces shows that perception is influenced by minority status, education level, job domains, age, and location. However, the perception of robots is not universally negative, as workers exhibit positive views when robots can improve workplace safety, enhance existing skills, or take over tasks that are viewed as less desirable. Currently, an understanding is lacking of the types of barriers to safety that a worker faces or how the job context, e.g. hours spent in heavy lifts, or demographics, e.g. weight or height, contribute to the perceptions of safety, need for and type of assistance from a co-worker, and perception of an assistive robot. Through a survey of 530 blue-collar workers, we find that 85% of the barriers to safety in the workplace are physical, psychological, and a combination of physical and psychological factors. Our survey shows that the time spent in performing heavy lifts is a significant predictor for barriers to safety, while the interaction between weight and time is a marginal predictor. We find that the barriers to safety significantly influence perception of job longevity. We observe that the worker’s height is a significant predictor to comfort levels in receiving assistance from co-workers, while weight and the interaction between height and weight are marginally significant. We find that the worker’s weight and hours spent in lifting are significant predictors for the type of assistance expected. Our findings show that the number of hours spent in lifting is a significant predictor on the type of assistance expected from a robot. Based on our results, future collaborative robots should be designed to show awareness of physical and physiological needs and cognizance to the worker’s physical traits, in particular height and weight, as well as the amount of time the worker has spent on lifting tasks.
Yu Liu 0069, Natasha Kholgade, Sean Banerjee
RO-MAN2
2024 Using Human-Human Handover Data to Analyze Giver and Receiver Timing Relationships During the Pre-Handover Phase
abstract
The fluency of handover between two agents is important to ensure safety and success of handover. In this work, we study the relationships between the timings of giver and receiver motions in human-human handover interactions, in order to inform human-robot handover. We use giver and receiver hand trajectories from the Human-Object-Human (HOH) handover dataset to study movement during the pre-handover phase, prior to the point of transfer. We find that human receivers adopt a largely proactive behavior, and plan and start motion early in the pre-handover phase. We also find that human receivers spend much of their motion moving in coordination with the giver, rather than after the giver has reached the transfer point. Further, we find that human receivers may predict future movement of the giver from early giver motion, and adjust their start times accordingly to ensure coordinated grasp at transfer. Our findings suggest that robot receivers should adopt a predictive giver-aware approach to plan motion early, and robot givers should recognize that human receivers may expect giver behavior to be human-like and predictable.
Ava Megyeri, Sean Banerjee, Maria Kyrarini, Natasha Kholgade
RO-MAN4
2024 GraspPC: Generating Diverse Hand Grasp Point Clouds on Objects
abstract
We present GraspPC, an approach to perform learning-based synthesis of multiple human hand grasps as point clouds from point clouds of objects. GraspPC benefits human-robot handover approaches by providing hypotheses of human grasp on objects to inform robotic manipulation algorithms on how to bias robotic grasp for safe handover. Existing learning-based approaches to conduct hand grasp prediction require datasets to contain annotated articulated hand models, making them difficult to train on datasets that lack hand model annotations. GraspPC treats the problem of hand point cloud generation from object point clouds as a set-to-set translation problem. We contribute a Transformer architecture to synthesize point clouds via GraspPC. To generate diverse hand grasps, we generate multiple object-dependent queries and train the network using a winner-takes-gradient strategy. We show results of diverse grasps by training and testing on a variety of real-world datasets. We demonstrate how human grasps generated by GraspPC can be used to filter robotic grasp candidates to inform human-robot handover. Our code is available at: https://github.com/Terascale-All-sensing-Research-Studio/GraspPC.
Ava Megyeri, Noah Wiederhold, Maria Kyrarini, Sean Banerjee, Natasha Kholgade
RO-MAN5
2023 Fantastic Breaks: A Dataset of Paired 3D Scans of Real-World Broken Objects and Their Complete Counterparts
abstract
Automated shape repair approaches currently lack access to datasets that describe real-world damaged geometry. We present Fantastic Breaks (and Where to Find Them: https://terascale-all-sensing-research-studio.github.io/FantasticBreaks), a dataset containing scanned, waterproofed, and cleaned 3D meshes for 150 broken objects, paired and geometrically aligned with complete counterparts. Fantastic Breaks contains class and material labels, proxy repair parts that join to broken meshes to generate complete meshes, and manually annotated fracture boundaries. Through a detailed analysis of fracture geometry, we reveal differences between Fantastic Breaks and synthetic fracture datasets generated using geometric and physics-based methods. We show experimental shape repair evaluation with Fantastic Breaks using multiple learning-based approaches pre-trained with synthetic datasets and re-trained with subset of Fantastic Breaks.
Nikolas Lamb, Cameron Palmer, Benjamin Molloy, Sean Banerjee, Natasha Kholgade
CVPR5
2023 HOH: Markerless Multimodal Human-Object-Human Handover Dataset with Large Object Count
abstract
We present the HOH (Human-Object-Human) Handover Dataset, a large object count dataset with 136 objects, to accelerate data-driven research on handover studies, human-robot handover implementation, and artificial intelligence (AI) on handover parameter estimation from 2D and 3D data of two-person interactions. HOH contains multi-view RGB and depth data, skeletons, fused point clouds, grasp type and handedness labels, object, giver hand, and receiver hand 2D and 3D segmentations, giver and receiver comfort ratings, and paired object metadata and aligned 3D models for 2,720 handover interactions spanning 136 objects and 20 giver-receiver pairs—40 with role-reversal—organized from 40 participants. We also show experimental results of neural networks trained using HOH to perform grasp, orientation, and trajectory prediction. As the only fully markerless handover capture dataset, HOH represents natural human-human handover interactions, overcoming challenges with markered datasets that require specific suiting for body tracking, and lack high-resolution hand tracking. To date, HOH is the largest handover dataset in terms of object count, participant count, pairs with role reversal accounted for, and total interactions captured.
Noah Wiederhold, Ava Megyeri, DiMaggio Paris, Sean Banerjee, Natasha Kholgade
NeurIPS5
2022 DeepMend: Learning Occupancy Functions to Represent Shape for Repair
Nikolas Lamb, Sean Banerjee, Natasha Kholgade
ECCV (3)3
2022 Combining Real-World Constraints on User Behavior with Deep Neural Networks for Virtual Reality (VR) Biometrics
abstract
Deep networks have demonstrated enormous potential for identification and authentication using behavioral biometrics in virtual reality (VR). However, existing VR behavioral biometrics datasets have small sample sizes which can make it challenging for deep networks to automatically learn features that characterize real-world user behavior and that may enable high success, e.g., high-level spatial relationships between headset and hand controller devices and underlying smoothness of trajectories despite noise. We provide an approach to perform behavioral biometrics using deep networks while incorporating spatial and smoothing constraints on input data to represent real-world behavior. We represent the input data to neural networks as a combination of scale- and translation-invariant device-centric position and orientation features, and displacement vectors representing spatial relationships between device pairs. We assess identification and authentication by including spatial relationships and by performing Gaussian smoothing of the position features. We evaluate our approach against baseline methods that use the raw data directly and that perform a global normalization of the data. By using displacement vectors, our work shows higher success over baseline methods in 36 out of 42 cases of analysis done by varying user sets and pairings of VR systems and sessions.
Natasha Kholgade, Sean Banerjee
VR2
2022 Temporal Effects in Motion Behavior for Virtual Reality (VR) Biometrics
abstract
Using the motion behavior of users in virtual reality (VR) as a biometric signature has the potential to enable continuous identification and authentication of users without compromising VR applications if traditional passwords are acquired by malicious agents. Users exhibit natural variabilities in behavior over time that influence their body motions and can alter the trajectories of VR devices such as the headset and the controllers. Behavior variabilities may negatively impact the success rate of VR biometrics. In this work, we evaluate how deep learning approaches to match input and enrollment trajectories are influenced by user behavior variation over varying time scales. We demonstrate that over short timescales on the order of seconds to minutes, no statistically significant relationship is found in the temporal placement of enrollment trajectories and their matches to input trajectories. We find that on medium-scale separation between enrollment and input trajectories, on the order of days to weeks, median accuracy is similar within users who provide input close and distant to enrollment data. Over long timescales on the order of 7 to 18 months, we obtain optimal performance for short and long enrollment/input separations by using training sets from users providing long-timescale data, as these sets encompass coarse and fine-scale changes in behavior.
Natasha Kholgade, Sean Banerjee
VR2
2022 MendNet: Restoration of Fractured Shapes Using Learned Occupancy Functions
abstract
Abstract We provide a novel approach to perform fully automated generation of restorations for fractured shapes using learned implicit shape representations in the form of occupancy functions. Our approach lays the groundwork to perform automated object repair via additive manufacturing. Existing approaches for restoration of fractured shapes either require prior knowledge of object structure such as symmetries between the restoration and the fractured object, or predict restorations as voxel outputs that are impractical for repair at current resolutions. By leveraging learned occupancy functions for restoration prediction, our approach overcomes the curse of dimensionality with voxel approaches, while providing plausible restorations. Given a fractured shape, we fit a function to occupancy samples from the shape to infer a latent code. We apply a learned transformation to the fractured shape code to predict a corresponding code for restoration generation. To ensure physical validity and well‐constrained shape estimation, we contribute a loss that models feasible occupancy values for fractured shapes, restorations, and complete shapes obtained by joining fractured and restoration shapes. Our work overcomes deficiencies of shape completion approaches adapted for repair, and enables consumer‐driven object repair and cultural heritage object restoration. We share our code and a synthetic dataset of fractured meshes from 8 ShapeNet classes at: https://github.com/Terascale‐All‐sensing‐Research‐Studio/MendNet .
Nikolas Lamb, Sean Banerjee, Natasha Kholgade
Comput. Graph. Forum3
2022 DeepJoin: Learning a Joint Occupancy, Signed Distance, and Normal Field Function for Shape Repair
abstract
We introduce DeepJoin, an automated approach to generate high-resolution repairs for fractured shapes using deep neural networks. Existing approaches to perform automated shape repair operate exclusively on symmetric objects, require a complete proxy shape, or predict restoration shapes using low-resolution voxels which are too coarse for physical repair. We generate a high-resolution restoration shape by inferring a corresponding complete shape and a break surface from an input fractured shape. We present a novel implicit shape representation for fractured shape repair that combines the occupancy function, signed distance function, and normal field. We demonstrate repairs using our approach for synthetically fractured objects from ShapeNet, 3D scans from the Google Scanned Objects dataset, objects in the style of ancient Greek pottery from the QP Cultural Heritage dataset, and real fractured objects. We outperform six baseline approaches in terms of chamfer distance and normal consistency. Unlike existing approaches and restorations generated using subtraction, DeepJoin restorations do not exhibit surface artifacts and join closely to the fractured region of the fractured shape. Our code is available at: https://github.com/Terascale-All-sensing-Research-Studio/DeepJoin.
Nikolas Lamb, Sean Banerjee, Natasha Kholgade
ACM Trans. Graph.3
2021 Using Siamese Neural Networks to Perform Cross-System Behavioral Authentication in Virtual Reality
abstract
In this paper, we provide an approach on using behavioral biometrics to perform cross-system high-assurance authentication of users in virtual reality (VR) environments. VR is currently being explored as a critical tool to ensure seamless delivery of essential services, such as education, healthcare, and personal finance, while enabling users to work from home environments. Due to the sensitive nature of personal data generated, VR applications for essential services need to provide secure access. Traditional PIN or password-based credentials can be breached by malicious impostors, or be handed over by an intended user of a VR system to a confederate to assist the intended user in completing a task, e.g., an exam or a physical therapy routine. Existing approaches that use the behavior of the user in VR as a biometric signature fail when users provide enrollment and use-time data on different VR systems. We use Siamese neural networks to learn a distance function that characterizes the systematic differences between data provided across pairs of dissimilar VR systems. Our approach provides average equal error rates (EERs) ranging from 1.38% to 3.86% for authentication using a benchmark dataset that consists of 41 users performing a ball-throwing task with 3 VR systems-an Oculus Quest, an HTC Vive, and an HTC Vive Cosmos. To compare to prior approaches in VR biometrics, we also obtain average accuracies for the task of identification, where given an input user's trajectory in a use-time VR system, we use Siamese networks to return the user with the top matching trajectory in an enrollment VR system as the label. We report identification results ranging from 87.82% to 98.53% with average improvements of 29.78%±8.58% and 30.78%±3.68% over existing approaches that use generic distance matching and fully convolutional networks on the enrollment dataset respectively.
Natasha Kholgade, Sean Banerjee
VR2
2020 Automatic Material Classification Using Thermal Finger Impression
Jacob Gately, Matthew Kolessar Wright, Natasha Kholgade, Sean Banerjee, Soumyabrata Dey
MMM (1)4
2019 CNN-Based Non-contact Detection of Food Level in Bottles from RGB Images
Yijun Jiang, Elim Schenck, Spencer Kranz, Sean Banerjee, Natasha Kholgade
MMM (1)5
2019 Task-Driven Biometric Authentication of Users in Virtual Reality (VR) Environments
Alexander Kupin, Benjamin Moeller, Yijun Jiang, Natasha Kholgade, Sean Banerjee
MMM (1)4
2018 Hardware Synchronization of Multiple Kinects and Microphones for 3D Audiovisual Spatiotemporal Data Capture
abstract
We present an approach that uses linear timecode to temporally synchronize visual data captured at 30 fps from multiple Kinect v2 sensors and audio data captured from external microphones. Existing approaches suffer from reduced frame rate due to latencies in relaying synchronization signals over a network, have inaccuracies due to ambient sound contamination when using audio for synchronization, or show multi-frame delays when using networking protocols to synchronize capture computer clocks. We align audio and color frame times to a timecode signal injected using a custom-designed hardware board to the audio processing board of each Kinect. Our approach provides synchronized capture unaffected by ambient noise or network latencies, with maximum synchronization offsets within a single frame for any number of devices. We show 3D point cloud reconstruction results with audio for a variety of single-and multi-person interactions captured using four Kinect v2 sensors and two external microphones synchronized by our approach.
Yijun Jiang, Timothy Godisart, Natasha Kholgade, Sean Banerjee
ICME4
2018 Spatiotemporal 3D Models of Aging Fruit from Multi-view Time-Lapse Videos
Lintao Guo, Hunter Quant, Nikolas Lamb, Benjamin Lowit, Sean Banerjee, Natasha Kholgade
MMM (1)6
2018 Multi-camera Microenvironment to Capture Multi-view Time-Lapse Videos for 3D Analysis of Aging Objects
Lintao Guo, Hunter Quant, Nikolas Lamb, Benjamin Lowit, Natasha Kholgade, Sean Banerjee
MMM (2)5
2018 Programmatic 3D Printing of a Revolving Camera Track to Automatically Capture Dense Images for 3D Scanning of Objects
Nikolas Lamb, Natasha Kholgade, Sean Banerjee
MMM (2)2
2018 A Virtual Reality Interface for Interactions with Spatiotemporal 3D Data
Hunter Quant, Sean Banerjee, Natasha Kholgade
MMM (2)3
2018 User-Independent Detection of Swipe Pressure Using a Thermal Camera for Natural Surface Interaction
abstract
In this paper, we use a thermal camera to distinguish hard and soft swipes performed by a user interacting with a natural surface by detecting differences in the thermal signature of the surface due to heat transferred by the user. Unlike prior work, our approach provides swipe pressure classifiers that are user-agnostic, i.e., that recognize the swipe pressure of a novel user not present in the training set, enabling our work to be ported into natural user interfaces without user-specific calibration. Our approach generates average classification accuracy of 76% using random forest classifiers trained on a test set of 9 subjects interacting with paper and wood, with 8 hard and 8 soft test swipes per user. We compare results of the user-agnostic classification to user-aware classification with classifiers trained by including training samples from the user. We obtain average user-aware classification accuracy of 82% by adding up to 8 hard and 8 soft training swipes for each test user. Our approach enables seamless adaptation of generic pressure classification systems based on thermal data to the specific behavior of users interacting with natural user interfaces.
Tim Dunn, Sean Banerjee, Natasha Kholgade
MMSP3
2016 Water Fixture Identification in Smart Housing: A Domain Knowledge Based Case Study
abstract
In current practice, smart housing environments often over-install smart sensors on every fixture and log data from them at high sampling rates, resulting in more data being collected than is necessary. Fixture identification offers a possible alternative to reduce the number of sensors installed and the amount of data collected in smart housing. Fixture identification applies classifiers to label utility consumption data aggregated at the apartment level by the specific fixture that actually contributes the data, such as the shower or the kitchen sink. Successful fixture identification can be used to educate tenants, optimize the resource supply strategy, and offer a smart solution for detecting abnormal usage activities. In this paper, we report a case study of water fixture identification by using support vector machines (SVMs) to perform fixture classification. We use the Smart Housing Dataset from Clarkson University, which comprises of one academic year of tenant activities from 12 student apartments. Our results show that the proposed approach achieves an average accuracy between 78% to 87.8% for identifying hot water fixtures including kitchen sink, bathroom sink and shower. As a result, the number of smart meters per apartment is reduced from 7 to 3, one for hot water, one for cold water, and the third for toilet. The novelty of our study lies in the feature selection process, which is guided by our domain knowledge of water fixture characteristics and the correlation between water fixture usage and other user behavior in the apartments. We describe our proposed features, their rationale, and their effect on classification performance.
Yan Gao 0006, Daqing Hou, Natasha Kholgade, Sean Banerjee
ICMLA3
2016 Real-time scalable 6DOF pose estimation for textureless objects
abstract
Real-time recognition of the 6DOF pose of textureless objects is a fundamental and challenging problem in robotics. We present a novel approach to perform real-time estimation of the viewpoint, scale, and translation of an object in RGB and RGB-D image captures. In this work, we use a 3D model to render example poses of a textureless object, and find the nearest match to the input image using a GPU implementation. To achieve invariance to illumination and appearance across an object, we transform images to the Laplacian of Gaussian space. To perform real-time matching, we introduce a novel reshaping of the template set and the image, and we restructure the traditional normalized cross-correlation operation to leverage the GPU for fast matrix-matrix multiplication. We provide further speed up of large-scale template matching by contributing a dimensionality reduction approach using principal component analysis, and a candidate elimination method. Our method achieves state-of-the-art performance as shown by qualitative results and quantitative comparisons to pre-existing methods.
Zhe Cao 0003, Yaser Sheikh, Natasha Kholgade
ICRA3
2014 3D object manipulation in a single photograph using stock 3D models
abstract
Photo-editing software restricts the control of objects in a photograph to the 2D image plane. We present a method that enables users to perform the full range of 3D manipulations, including scaling, rotation, translation, and nonrigid deformations, to an object in a photograph. As 3D manipulations often reveal parts of the object that are hidden in the original photograph, our approach uses publicly available 3D models to guide the completion of the geometry and appearance of the revealed areas of the object. The completion process leverages the structure and symmetry in the stock 3D model to factor out the effects of illumination, and to complete the appearance of the object. We demonstrate our system by producing object manipulations that would be impossible in traditional 2D photo-editing programs, such as turning a car over, making a paper-crane flap its wings, or manipulating airplanes in a historical photograph to change its story.
Natasha Kholgade, Tomas Simon, Alexei A. Efros, Yaser Sheikh
ACM Trans. Graph.1