Eakta Jain

dblp:17/8357 · DBLP profile ↗
← Back
37ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0001-5131-3355ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 4 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 15 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 3 since 2021Computer networks · 1Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Toward Multimodal Privacy in XR: Design and Evaluation of Composite Privatization Methods for Gaze and Body Tracking Data
abstract
As extended reality (XR) systems become increasingly immersive and sensor-rich, they enable the collection of behavioral signals such as eye and body telemetry. These signals support personalized and responsive experiences and may also contain unique patterns that can be linked back to individuals. However, privacy mechanisms that naively pair unimodal mechanisms (e.g., independently apply privacy mechanisms for eye and body privatization) are often ineffective at preventing re-identification in practice. In this work, we systematically evaluate real-time privacy mechanisms for XR, both individually and in pair, across eye and body modalities. We assess privacy through re-identification rates and evaluate utility using numerical performance thresholds derived from existing literature to ensure real-time interaction requirements are met. We evaluated four eye and ten body mechanisms across multiple datasets, comprising up to 407 participants. Our results show that when carefully paired, multimodal mechanisms reduce re-identification rate from 80.3% to 26.3% in casual XR applications (e.g., VRChat and Job Simulator) and from 84.8% to 26.1 % in competitive XR applications (e.g., Beat Saber and Synth Riders), all while maintaining acceptable performance based on established thresholds. To facilitate adoption, we additionally release XR Privacy SDK, an open-source toolkit enabling developers to integrate the privacy mechanisms into XR applications for real-time use. These findings underscore the potential of modality-specific and context-aware privacy strategies for protecting behavioral data in XR environments.
Azim Ibragimov, Ethan Wilson, Kevin R. B. Butler, Eakta Jain
IEEE Trans. Vis. Comput. Graph.4
2025 Eye-Tracked Virtual Reality: A Comprehensive Survey on Methods and Privacy Challenges
abstract
The latest developments in computer hardware, sensor technologies, and artificial intelligence can make virtual reality (VR) and virtual spaces an important part of human everyday life. Eye tracking offers not only a hands-free way of interaction but also the possibility of a deeper understanding of human visual attention and cognitive processes in VR. Despite these possibilities, eye-tracking data also reveal users’ privacy-sensitive attributes when combined with the information about the presented stimulus. To address all, this survey first covers major works in eye tracking, VR, and privacy areas between 2012 and 2022. While eye tracking in VR part covers the computational eye-tracking pipeline from pupil detection and gaze estimation to offline data analysis, for privacy and security, we focus on eye-based authentication as well as computational methods to preserve the privacy of individuals and their eye-tracking data in VR. Later, we outline three main directions by focusing on privacy. In summary, this survey presents an extensive literature review of the utmost possibilities of eye tracking in VR and their privacy implications.
Efe Bozkir, Süleyman Özdel, Mengdi Wang 0002, Brendan David-John, Hong Gao 0008, Kevin R. B. Butler, Eakta Jain, Enkelejda Kasneci
Proc. IEEE7
2024 "I Had Sort of a Sense that I Was Always Being Watched...Since I Was": Examining Interpersonal Discomfort From Continuous Location-Sharing Applications
abstract
Continuous location sharing (CLS) applications are widely used for safety and social convenience. However, these applications have privacy concerns that can be used for control and harm. To understand user concerns, we performed the largest user study of CLS application usage performed to date, with 1500 of 3000 users indicating they use CLS applications and 896 of these users completing surveys. From survey responses, we conducted 23 interviews with participants who had uncomfortable experiences. With these interviews, we perform thematic analysis grounded by sociological frameworks of power dynamics and social exchange theory. We observe that CLS application users face discomfort related to three primary categories that build on each other: (1) overstepped boundaries, (2) continued discomfort, and (3) lifestyle-impacting behaviors. With this foundational understanding, we suggest features that aim to reduce relationship imbalances that CLS applications enable. Our resulting study demonstrates that CLS applications contribute to interpersonal discomfort, highlighting the need for design changes.
Kevin Childs, Cassidy Gibson, Anna Crowder, Kevin Warren, Carson Stillman, Elissa M. Redmiles, Eakta Jain, Patrick Traynor, Kevin R. B. Butler
CCS7
2024 Towards mitigating uncann(eye)ness in face swaps via gaze-centric loss terms
Ethan Wilson, Frédérick Shic, Sophie Jörg, Eakta Jain
Comput. Graph.4
2024 Privacy-Preserving Gaze Data Streaming in Immersive Interactive Virtual Reality: Robustness and User Experience
abstract
Eye tracking is routinely being incorporated into virtual reality (VR) systems. Prior research has shown that eye tracking data, if exposed, can be used for re-identification attacks [14]. The state of our knowledge about currently existing privacy mechanisms is limited to privacy-utility trade-off curves based on data-centric metrics of utility, such as prediction error, and black-box threat models. We propose that for interactive VR applications, it is essential to consider user-centric notions of utility and a variety of threat models. We develop a methodology to evaluate real-time privacy mechanisms for interactive VR applications that incorporate subjective user experience and task performance metrics. We evaluate selected privacy mechanisms using this methodology and find that re-identification accuracy can be decreased to as low as 14% while maintaining a high usability score and reasonable task performance. Finally, we elucidate three threat scenarios (black-box, black-box with exemplars, and white-box) and assess how well the different privacy mechanisms hold up to these adversarial scenarios. This work advances the state of the art in VR privacy by providing a methodology for end-to-end assessment of the risk of re-identification attacks and potential mitigating solutions. f.
Ethan Wilson, Azim Ibragimov, Michael J. Proulx, Sai Deep Tetali, Kevin R. B. Butler, Eakta Jain
IEEE Trans. Vis. Comput. Graph.6
2023 Horse as Teacher: How human-horse interaction informs human-robot interaction
abstract
Robots are entering our lives and workplaces as companions and teammates. Though much research has been done on how to interact with robots, teach robots and improve task performance, an open frontier for HCI/HRI research is how to establish a working relationship with a robot in the first place. Studies that explore the early stages of human-robot interaction are an emerging area of research. Simultaneously, there is resurging interest in how human-animal interaction could inform human-robot interaction. We present a first examination of early stage human-horse interaction through the lens of human-robot interaction, thus connecting these two areas. Following Strauss’ approach, we conduct a thematic analysis of data from three sources gathered over a year of field work: observations, interviews and journal entries. We contribute design guidelines based on our analyses and findings.
Eakta Jain, Christina Gardner-McCune
CHI1
2023 Introducing Explicit Gaze Constraints to Face Swapping
abstract
Face swapping combines one face’s identity with another face’s non-appearance attributes (expression, head pose, lighting) to generate a synthetic face. This technology is rapidly improving, but falls flat when reconstructing some attributes, particularly gaze. Image-based loss metrics that consider the full face do not effectively capture the perceptually important, yet spatially small, eye regions. Improving gaze in face swaps can improve naturalness and realism, benefiting applications in entertainment, human computer interaction, and more. Improved gaze will also directly improve Deepfake detection efforts, serving as ideal training data for classifiers that rely on gaze for classification. We propose a novel loss function that leverages gaze prediction to inform the face swap model during training and compare against existing methods. We find all methods to significantly benefit gaze in resulting face swaps.
Ethan Wilson, Frédérick Shic, Eakta Jain
ETRA3
2023 Real-Time Conversational Gaze Synthesis for Avatars
abstract
Eye movement plays an important role in face-to-face communication. In this work, we present a deep learning approach for synthesizing the eye movements of avatars for two-party conversations and evaluate viewer perception of different types of eye motions. We aim to synthesize believable gaze behavior based on head motions and audio features as they would typically be available in virtual reality applications. To this end, we captured the head motion, eye motion, and audio of several two-party conversations and trained an RNN-based model to predict where an avatar looks in a two-person conversational scenario. We evaluated our approach with a user study on the perceived quality of the eye animation and compared our method with other eye animation methods. While our model was not rated highest, our model and our user study lead to a series of insights on model features, viewer perception, and study design that we present.
Ryan Canales, Eakta Jain, Sophie Jörg
MIG2
2023 Privacy-preserving datasets of eye-tracking samples with applications in XR
abstract
Virtual and mixed-reality (XR) technology has advanced significantly in the last few years and will enable the future of work, education, socialization, and entertainment. Eye-tracking data is required for supporting novel modes of interaction, animating virtual avatars, and implementing rendering or streaming optimizations. While eye tracking enables many beneficial applications in XR, it also introduces a risk to privacy by enabling re-identification of users. We applied privacy definitions of k-anonymity and plausible deniability (PD) to datasets of eye-tracking samples and evaluated them against the state-of-the-art differential privacy (DP) approach. Two VR datasets were processed to reduce identification rates while minimizing the impact on the performance of trained machine-learning models. Our results suggest that both PD and DP mechanisms produced practical privacy-utility trade-offs with respect to re-identification and activity classification accuracy, while k-anonymity performed best at retaining utility for gaze prediction.
Brendan David-John, Kevin R. B. Butler, Eakta Jain
IEEE Trans. Vis. Comput. Graph.3
2022 For Your Eyes Only: Privacy-preserving eye-tracking datasets
abstract
Eye-tracking is a critical source of information for understanding human behavior and developing future mixed-reality technology. Eye-tracking enables applications that classify user activity or predict user intent. However, eye-tracking datasets collected during common virtual reality tasks have also been shown to enable unique user identification, which creates a privacy risk. In this paper, we focus on the problem of user re-identification from eye-tracking features. We adapt standardized privacy definitions of k-anonymity and plausible deniability to protect datasets of eye-tracking features, and evaluate performance against re-identification by a standard biometric identification model on seven VR datasets. Our results demonstrate that re-identification goes down to chance levels for the privatized datasets, even as utility is preserved to levels higher than 72% accuracy in document type classification.
Brendan David-John, Kevin R. B. Butler, Eakta Jain
ETRA3
2022 Is the avatar scared? Pupil as a perceptual cue
abstract
Abstract The importance of eyes for virtual characters stems from the intrinsic social cues in a person's eyes. While previous work on computer generated eyes has considered realism and naturalness, there has been little investigation into how details in the eye animation impact the perception of an avatar's internal emotional state. We present three large scale experiments (N≈500) that investigate the extent to which viewers can identify if an avatar is scared. We find that participants can identify a scared avatar with accuracy using cues in the eyes including pupil size variation, gaze, and blinks. Because eye trackers return pupil diameter in addition to gaze, our experiments inform practitioners that animating the pupil correctly will add expressiveness to a virtual avatar with negligible additional cost. These findings also have implications for creating expressive eyes in intelligent conversational agents and social robots.
Yuzhu Dong, Sophie Jörg, Eakta Jain
Comput. Animat. Virtual Worlds3
2022 Fast Foveating Cameras for Dense Adaptive Resolution
abstract
Traditional cameras field of view (FOV) and resolution predetermine computer vision algorithm performance. These trade-offs decide the range and performance in computer vision algorithms. We present a novel foveating camera whose viewpoint is dynamically modulated by a programmable micro-electromechanical (MEMS) mirror, resulting in a natively high-angular resolution wide-FOV camera capable of densely and simultaneously imaging multiple regions of interest in a scene. We present calibrations, novel MEMS control algorithms, a real-time prototype, and comparisons in remote eye-tracking performance against a traditional smartphone, where high-angular resolution and wide-FOV are necessary, but traditionally unavailable.
Brevin Tilmon, Eakta Jain, Silvia Ferrari, Sanjeev J. Koppal
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Introduction to the Special Issue on SAP 2021
abstract
No abstract available.
Eakta Jain, Anne-Hélène Olivier
ACM Trans. Appl. Percept.1
2021 A privacy-preserving approach to streaming eye-tracking data
abstract
Eye-tracking technology is being increasingly integrated into mixed reality devices. Although critical applications are being enabled, there are significant possibilities for violating user privacy expectations. We show that there is an appreciable risk of unique user identification even under natural viewing conditions in virtual reality. This identification would allow an app to connect a user's personal ID with their work ID without needing their consent, for example. To mitigate such risks we propose a framework that incorporates gatekeeping via the design of the application programming interface and via software-implemented privacy mechanisms. Our results indicate that these mechanisms can reduce the rate of identification from as much as 85% to as low as 30%. The impact of introducing these mechanisms is less than 1.5° error in gaze position for gaze prediction. Gaze data streams can thus be made private while still allowing for gaze prediction, for example, during foveated rendering. Our approach is the first to support privacy-by-design in the flow of eye-tracking data within mixed reality use cases.
Brendan David-John, Diane Hosfelt, Kevin R. B. Butler, Eakta Jain
IEEE Trans. Vis. Comput. Graph.4
2020 FoveaCam: A MEMS Mirror-Enabled Foveating Camera
abstract
Most cameras today photograph their entire visual field. In contrast, decades of active vision research have proposed foveating camera designs, which allow for selective scene viewing. However, active vision's impact is limited by slow options for mechanical camera movement. We propose a new design, called FoveaCam, and which works by capturing reflections off a tiny, fast moving mirror. FoveaCams can obtain high resolution imagery on multiple regions of interest, even if these are at different depths and viewing directions. We first discuss our prototype and optical calibration strategies. We then outline a control algorithm for the mirror to track target pairs. Finally, we demonstrate a practical application of the full system to enable eye tracking at a distance for frontal faces.
Brevin Tilmon, Eakta Jain, Silvia Ferrari, Sanjeev J. Koppal
ICCP2
2020 Adult2child: Motion Style Transfer using CycleGANs
abstract
Child characters are commonly seen in leading roles in top-selling video games. Previous studies have shown that child motions are perceptually and stylistically different from those of adults. Creating motion for these characters by motion capturing children is uniquely challenging because of confusion, lack of patience and regulations. Retargeting adult motion, which is much easier to record, onto child skeletons, does not capture the stylistic differences. In this paper, we propose that style translation is an effective way to transform adult motion capture data to the style of child motion. Our method is based on CycleGAN, which allows training on a relatively small number of sequences of child and adult motions that do not even need to be temporally aligned. Our adult2child network converts short sequences of motions called motion words from one domain to the other. The network was trained using a motion capture database collected by our team containing 23 locomotion and exercise motions. We conducted a perception study to evaluate the success of style translation algorithms, including our algorithm and recently presented style translation neural networks. Results show that the translated adult motions are recognized as child motions significantly more often than adult motions.
Yuzhu Dong, Andreas Aristidou, Ariel Shamir, Moshe Mahler, Eakta Jain
MIG5
2020 The Security-Utility Trade-off for Iris Authentication and Eye Animation for Social Virtual Avatars
abstract
The gaze behavior of virtual avatars is critical to social presence and perceived eye contact during social interactions in Virtual Reality. Virtual Reality headsets are being designed with integrated eye tracking to enable compelling virtual social interactions. This paper shows that the near infra-red cameras used in eye tracking capture eye images that contain iris patterns of the user. Because iris patterns are a gold standard biometric, the current technology places the user's biometric identity at risk. Our first contribution is an optical defocus based hardware solution to remove the iris biometric from the stream of eye tracking images. We characterize the performance of this solution with different internal parameters. Our second contribution is a psychophysical experiment with a same-different task that investigates the sensitivity of users to a virtual avatar's eye movements when this solution is applied. By deriving detection threshold values, our findings provide a range of defocus parameters where the change in eye movements would go unnoticed in a conversational setting. Our third contribution is a perceptual study to determine the impact of defocus parameters on the perceived eye contact, attentiveness, naturalness, and truthfulness of the avatar. Thus, if a user wishes to protect their iris biometric, our approach provides a solution that balances biometric protection while preventing their conversation partner from perceiving a difference in the user's virtual avatar. This work is the first to develop secure eye tracking configurations for VR/AR/XR applications and motivates future work in the area.
Brendan David-John, Sophie Jörg, Sanjeev J. Koppal, Eakta Jain
IEEE Trans. Vis. Comput. Graph.4
2019 EyeVEIL: degrading iris authentication in eye tracking headsets
abstract
Mixed reality headsets are being designed with integrated eye trackers: cameras that image the user's eye to infer gaze location and pupil diameter. While the intent is to improve the quality of experience, built-in eye trackers create a security vulnerability for hackers - high resolution images of the user's iris. Anyone stealing an iris image has effectively captured a gold standard biometric, relied on for secure authentication in applications such as banking and voting. We present a low cost solution to degrade iris authentication while still permitting the utility of gaze tracking with acceptable accuracy. By demonstrating this solution on a commodity eye tracker, this paper urges the community to think about iris based authentication as a byproduct of eye tracking, and create solutions that empower a user to control this biometric.
Brendan David-John, Sanjeev J. Koppal, Eakta Jain
ETRA3
2019 Differential privacy for eye-tracking data
abstract
As large eye-tracking datasets are created, data privacy is a pressing concern for the eye-tracking community. De-identifying data does not guarantee privacy because multiple datasets can be linked for inferences. A common belief is that aggregating individuals' data into composite representations such as heatmaps protects the individual. However, we analytically examine the privacy of (noise-free) heatmaps and show that they do not guarantee privacy. We further propose two noise mechanisms that guarantee privacy and analyze their privacy-utility tradeoff. Analysis reveals that our Gaussian noise mechanism is an elegant solution to preserve privacy for heatmaps. Our results have implications for interdisciplinary research to create differentially private mechanisms for eye tracking.
Ao Liu 0001, Lirong Xia, Andrew T. Duchowski, Reynold J. Bailey, Kenneth Holmqvist, Eakta Jain
ETRA6
2019 Eye tracking and virtual reality
abstract
Virtual Reality has the potential to transform the way we work, rest and play. We are seeing use cases as diverse as education and pain management, with new applications being imagined every day. VR technology comes with new challenges, and many obstacles need to be overcome to ensure good user experience. Recently many new Virtual Reality systems with integrated eye tracking have become available. This course presents timely, relevant information on how Virtual Reality (VR) can leverage eye-tracking data to optimize the user experience and to alleviate usability issues surrounding many challenges in immersive VEs. The integration of eye tracking allows us to determine where the viewer is focusing their attention. If we, as the content creators and world builders, need the user to focus on another area of the VE we can use techniques to attract attention to these regions and also we can confirm we are doing so successful as we continually track the users gaze. Advancing these approaches could make the VR experience more comfortable, safe and effective for the user.
Ann McNamara, Eakta Jain
SIGGRAPH Asia2
2018 Deepcomics: saliency estimation for comics
abstract
A key requirement for training deep learning saliency models is large training eye tracking datasets. Despite the fact that the accessibility of eye tracking technology has greatly increased, collecting eye tracking data on a large scale for very specific content types is cumbersome, such as comic images, which are different from natural images such as photographs because text and pictorial content is integrated. In this paper, we show that a deep network trained on visual categories where the gaze deployment is similar to comics outperforms existing models and models trained with visual categories for which the gaze deployment is dramatically different from comics. Further, we find that it is better to use a computationally generated dataset on visual category close to comics one than real eye tracking data of a visual category that has different gaze deployment. These findings hold implications for the transference of deep networks to different domains.
Kévin Bannier, Eakta Jain, Olivier Le Meur
ETRA2
2018 How many words is a picture worth?: attention allocation on thumbnails versus title text regions
abstract
Cognitive scientists and psychologists have long noted the "picture superiority effect", that is, pictorial content is more likely to be remembered and more likely to lead to an increased understanding of the material. We investigated the relative importance of pictorial regions versus textual regions on a website where pictures and text co-occur in a very structured manner: video content sharing websites. We tracked participants' eye movements as they performed a casual browsing task, that is, selecting a video to watch. We found that participants allocated almost twice as much attention to thumbnails as to title text regions. They also tended to look at the thumbnail images before the title text, as predicted by the picture superiority effect. These results have implications for both user experience designers as well as video content creators.
Chaitra Yangandul, Sachin Paryani, Madison Le, Eakta Jain
ETRA4
2018 An evaluation of pupillary light response models for 2D screens and VR HMDs
abstract
Pupil diameter changes have been shown to be indicative of user engagement and cognitive load for various tasks and environments. However, it is still not the preferred physiological measure for applied settings. This reluctance to leverage the pupil as an index of user engagement stems from the problem that in scenarios where scene brightness cannot be controlled, the pupil light response confounds the cognitive-emotional response. What if we could predict the light response of an individual's pupil, thus creating the opportunity to factor it out of the measurement? In this work, we lay the groundwork for this research by evaluating three models of pupillary light response in 2D, and in a virtual reality (VR) environment. Our results show that either a linear or an exponential model can be fit to an individual participant with an easy-to-use calibration procedure. This work opens several new research directions in VR relating to performance analysis and inspires the use of eye tracking beyond gaze as a pointer and foveated rendering.
Brendan David-John, Pallavi Raiturkar, Arunava Banerjee, Eakta Jain
VRST4
2017 Adult2Child: dynamic scaling laws to create child-like motion
abstract
Child characters are widely used in animations and games; however, child motion capture databases are less easily available than those involving adult actors. Previous studies have shown that there is a perceivable difference in adult and child motion based on point light displays, so it may not be appropriate to just use adult motion data on child characters. Due to the costs associated with motion capture of child actors, it would be beneficial if we could create a child motion corpus by translating adult motion into child-like motion. Previous works have proposed dynamic scaling laws to transfer motion from one character to its scaled version. In this paper, we conduct a perception study to understand if this procedure can be applied to translate adult motion into child-like motion. Viewers were shown three types of point light display videos: adult motion, child motion, and dynamically scaled adult motion and asked to identify if the translated motion belongs to a child or an adult. We found that the use of dynamic scaling led to an increase in the number of people identifying the motion as belonging to a child compared to the original adult motion. Our findings suggest that although the dynamic scaling method is not a final solution to translate adult motion into child-like motion, it is nevertheless an intermediate step in the right direction. To better illustrate the original and dynamically scaled motions for the purposes of this paper, we rendered the dynamically scaled motion on an androgynous manikin character.
Yuzhu Dong, Aishat Aloba, Sachin Paryani, Lisa Anthony, Neha Rana, Eakta Jain
MIG6
2017 Creating Segments and Effects on Comics by Clustering Gaze Data
abstract
Traditional comics are increasingly being augmented with digital effects, such as recoloring, stereoscopy, and animation. An open question in this endeavor is identifying where in a comic panel the effects should be placed. We propose a fast, semi-automatic technique to identify effects-worthy segments in a comic panel by utilizing gaze locations as a proxy for the importance of a region. We take advantage of the fact that comic artists influence viewer gaze towards narrative important regions. By capturing gaze locations from multiple viewers, we can identify important regions and direct a computer vision segmentation algorithm to extract these segments. The challenge is that these gaze data are noisy and difficult to process. Our key contribution is to leverage a theoretical breakthrough in the computer networks community towards robust and meaningful clustering of gaze locations into semantic regions, without needing the user to specify the number of clusters. We present a method based on the concept of relative eigen quality that takes a scanned comic image and a set of gaze points and produces an image segmentation. We demonstrate a variety of effects such as defocus, recoloring, stereoscopy, and animations. We also investigate the use of artificially generated gaze locations from saliency models in place of actual gaze locations.
Ishwarya Thirunarayanan, Khimya Khetarpal, Sanjeev J. Koppal, Olivier Le Meur, John M. Shea, Eakta Jain
ACM Trans. Multim. Comput. Commun. Appl.6
2016 Is the motion of a child perceivably different from the motion of an adult?
abstract
No abstract available.
Eakta Jain, Lisa Anthony, Aishat Aloba, Amanda Castonguay, Isabella Cuba, Julia Woodward
SAP1
2016 Measuring viewers' heart rate response to environment conservation videos
abstract
Digital media, particularly pictures and videos, have long been used to influence a person's cognition as well as her consequent actions. Previous work has shown that physiological indices such as heart rate variability can be used to measure emotional arousal. We measure heart rate variability as participants watch environment conservation videos. We compare the heart rate response against the pleasantness rating recorded during an independent Internet survey.
Pallavi Raiturkar, Susan Jacobson, Beida Chen, Kartik Chaturvedi, Isabella Cuba, Melissa Franklin, Julian Tolentino, Nia Haynes, Rebecca Soodeen, Eakta Jain
SAP11
2016 Decoupling light reflex from pupillary dilation to measure emotional arousal in videos
abstract
Predicting the exciting portions of a video is a widely relevant problem because of applications such as video summarization, searching for similar videos, and recommending videos to users. Researchers have proposed the use of physiological indices such as pupillary dilation as a measure of emotional arousal. The key problem with using the pupil to measure emotional arousal is accounting for pupillary response to brightness changes. We propose a linear model of pupillary light reflex to predict the pupil diameter of a viewer based only on incident light intensity. The residual between the measured pupillary diameter and the model prediction is attributed to the emotional arousal corresponding to that scene. We evaluate the effectiveness of this method of factoring out pupillary light reflex for the particular application of video summarization. The residual is converted into an exciting-ness score for each frame of a video. We show results on a variety of videos, and compare against ground truth as reported by three independent coders.
Pallavi Raiturkar, Andrea Kleinsmith, Andreas Keil, Arunava Banerjee, Eakta Jain
SAP5
2016 Scan path and movie trailers for implicit annotation of videos
abstract
Affective annotation of videos is important for video understanding, ranking, retrieval, and summarization. We present an approach that uses excerpts that appeared in the official trailers of movies, as training data. Total scan path is computed as a metric for emotional arousal, based on previous eye tracking research. Arousal level on trailer excerpts is modeled as a Gaussian distribution, and signed distance from the mean of this distribution is used to separate out exemplars of high and low emotional arousal in movies.
Pallavi Raiturkar, Eakta Jain
SAP3
2016 Leveraging gaze data for segmentation and effects on comics
abstract
In this work, we present a semi-automatic method based on gaze data to identify the objects in comic images on which digital effects will look best. Our key contribution is a robust technique to cluster the noisy gaze data without having to specify the number of clusters as input. We also present an approach to segment the identified object of interest.
Ishwarya Thirunarayanan, Sanjeev J. Koppal, John M. Shea, Eakta Jain
SAP4
2016 Is the Motion of a Child Perceivably Different from the Motion of an Adult?
abstract
Artists and animators have observed that children’s movements are quite different from adults performing the same action. Previous computer graphics research on human motion has primarily focused on adult motion. There are open questions as to how different child motion actually is, and whether the differences will actually impact animation and interaction. We report the first explicit study of the perception of child motion (ages 5 to 9 years old), compared to analogous adult motion. We used markerless motion capture to collect an exploratory corpus of child and adult motion, and conducted a perceptual study with point light displays to discover whether naive viewers could identify a motion as belonging to a child or an adult. We find that people are generally successful at this task. This work has implications for creating more engaging and realistic avatars for games, online social media, and animated videos and movies.
Eakta Jain, Lisa Anthony, Aishat Aloba, Amanda Castonguay, Isabella Cuba, Julia Woodward
ACM Trans. Appl. Percept.1
2015 Gaze-Driven Video Re-Editing
abstract
Given the current profusion of devices for viewing media, video content created at one aspect ratio is often viewed on displays with different aspect ratios. Many previous solutions address this problem by retargeting or resizing the video, but a more general solution would re-edit the video for the new display. Our method employs the three primary editing operations: pan, cut, and zoom. We let viewers implicitly reveal what is important in a video by tracking their gaze as they watch the video. We present an algorithm that optimizes the path of a cropping window based on the collected eyetracking data, finds places to cut, and computes the size of the cropping window. We present results on a variety of video clips, including close-up and distant shots, and stationary and moving cameras. We conduct two experiments to evaluate our results. First, we eyetrack viewers on the result videos generated by our algorithm, and second, we perform a subjective assessment of viewer preference. These experiments show that viewer gaze patterns are similar on our result videos and on the original video clips, and that viewers prefer our results to an optimized crop-and-warp algorithm.
Eakta Jain, Yaser Sheikh, Ariel Shamir, Jessica K. Hodgins
ACM Trans. Graph.1
2013 Predicting Primary Gaze Behavior Using Social Saliency Fields
abstract
We present a method to predict primary gaze behavior in a social scene. Inspired by the study of electric fields, we posit "social charges"-latent quantities that drive the primary gaze behavior of members of a social group. These charges induce a gradient field that defines the relationship between the social charges and the primary gaze direction of members in the scene. This field model is used to predict primary gaze behavior at any location or time in the scene. We present an algorithm to estimate the time-varying behavior of these charges from the primary gaze behavior of measured observers in the scene. We validate the model by evaluating its predictive precision via cross-validation in a variety of social scenes.
Hyun Soo Park, Eakta Jain, Yaser Sheikh
ICCV2
2013 ERELT: a faster alternative to the list-based interfaces for tree exploration and searching in mobile devices
abstract
This paper presents ERELT (Enhanced Radial Edgeless Tree), a tree visualization approach on modern mobile devices. ERELT is designed to offer a clear visualization of any tree structure with intuitive interaction. We are interested in both the observation and navigation of such structures. Such visualization can assist users in interacting with a hierarchical structure such as a media collection, file system, etc.
Abhishek P. Chhetri, Kang Zhang 0001, Eakta Jain
VINCI3
2012 Inferring artistic intention in comic art through viewer gaze
abstract
Comics are a compelling, though complex, visual storytelling medium. Researchers are interested in the process of comic art creation to be able to automatically tell new stories, and also, summarize videos and catalog large collections of photographs for example. A primary organizing principle used by artists to lay out the components of comic art (panels, word bubbles, objects inside each panel) is to lead the viewer's attention along a deliberate visual route that reveals the narrative. If artists are successful in leading viewer attention, then their intended visual route would be accessible through recorded viewer attention, i.e., eyetracking data. In this paper, we conduct an experiment to verify if artists are successful in their goal of leading viewer gaze. We eyetrack viewers on images taken from comic books, as well as photographs taken by experts, amateur photographers and a robot. Our data analyses show that there is increased consistency in viewer gaze for comic pictures versus photographs taken by a robot and by amateur photographers, thus confirming that comic artists do indeed direct the flow of viewer attention.
Eakta Jain, Yaser Sheikh, Jessica K. Hodgins
SAP1
2012 3D Social Saliency from Head-mounted Cameras
Hyun Soo Park, Eakta Jain, Yaser Sheikh
NIPS2
2012 Three-dimensional proxies for hand-drawn characters
abstract
Drawing shapes by hand and manipulating computer-generated objects are the two dominant forms of animation. Though each medium has its own advantages, the techniques developed for one medium are not easily leveraged in the other medium because hand animation is two-dimensional, and inferring the third dimension is mathematically ambiguous. A second challenge is that the character is a consistent three-dimensional (3D) object in computer animation while hand animators introduce geometric inconsistencies in the two-dimensional (2D) shapes to better convey a character's emotional state and personality. In this work, we identify 3D proxies to connect hand-drawn animation and 3D computer animation. We present an integrated approach to generate three levels of 3D proxies: single-points, polygonal shapes, and a full joint hierarchy. We demonstrate how this approach enables one medium to take advantage of techniques developed for the other; for example, 3D physical simulation is used to create clothes for a hand-animated character, and a traditionally trained animator is able to influence the performance of a 3D character while drawing with paper and pencil.
Eakta Jain, Yaser Sheikh, Moshe Mahler, Jessica K. Hodgins
ACM Trans. Graph.1