Qi Sun 0003

dblp:05/4187-3 · DBLP profile ↗
← Back
39ranked-venue papers
3as first author
32since 2021 · last 2026
0000-0002-3094-5844ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 3 first-author · 27 since 2021Human-computer interaction and ubiquitous computing · 8 · 8 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 HOICraft: In-Situ VLM-based Authoring Tool for Part-Level Hand-Object Interaction Design in VR
abstract
Hand–Object Interaction (HOI) is a key interaction component in Virtual Reality (VR). However, designing HOI still requires manual efforts to decide how object should be selected and manipulated, while also considering user abilities, which leads to time-consuming refinements. We present HOICraft, a VLM-based in-situ HOI authoring tool that enables part-level interaction design in VR. Here, HOICraft assists designers by recommending interactable elements from 3D objects, customizing HOI design properties, and mapping hand movement with virtual object behavior. We conducted a formative study with three expert VR designers to identify five representative HOI designs to support diverse user experiences. Building upon preference data from 20 participants, we develop an HOI mapping module with in-context learning. In a user study with 12 VR interaction designers, HOI mapping from HOICraft significantly reduced trial-and-error iterations compared to manual authoring. Finally, we assessed the usability of HOICraft, demonstrating its effectiveness for HOI design in VR.
Dohui Lee, Qi Sun 0003, Sang Ho Yoon
CHI2
2026 GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts
abstract
Video generated by the current state-of-the-art generative models contain undesirable artifacts. We introduce GeneVA, the first large-scale dataset of human-annotated artifact bounding boxes in AI-generated videos. The dataset consists of 16,356 AI-generated videos, each labeled by a human annotator with per-frame artifact bounding boxes, their labels and descriptions, and video quality ratings. A custom data collection pipeline was developed in Prolific, and a novel taxonomy for spatio-temporal artifacts present in AI-generated videos was defined. The videos were from the VidProM [41] dataset, with text prompts from this dataset then used to generate an additional subset of videos using Sora. We trained an artifact detector and caption generator using a pre-trained image-based model, and a custom temporal fusion module. The dataset can be found at https://www.immersivecomputinglab.org/publication/geneva. We hope that datasets like GeneVA will encourage improvements in artifact detection in AI-generated video towards applications such as deepfake detection.
Jenna Kang, Maria Beatriz Silva, Patsorn Sangkloy, Kenneth Chen, Niall L. Williams, Qi Sun 0003
WACV6
2026 Perceptually Guided 3DGS Streaming and Rendering for Mixed Reality
abstract
Recent breakthroughs in radiance fields, particularly 3D Gaussian Splatting (3DGS), have unlocked real-time, high-quality rendering of complex environments, enabling a wide range of applications. However, the stringent requirements of mixed reality (MR) rendering, such as rapid refresh rates, high-resolution stereo viewing, and constrained computing budgets, remain out of reach for current 3DGS techniques. Nevertheless, the wide field-of-view design of MR displays, which mimics human vision, presents a unique opportunity to exploit human visual system’s own perceptual limitations to reduce computational overhead while not compromising user-perceived rendering quality.To this end, we propose a perception-guided, continuous level-of-detail (LOD) framework for 3DGS that maximizes perceived quality under given compute resources. We distill a visual quality metric, which encodes the spatial, temporal, and peripheral characteristics of human visual perception, into a lightweight, gaze-contingent model that predicts and adaptively modulates rendering LOD across the visual field based on each region’s contributions to perceptual quality. This budget-driven LOD modulation, guided by both scene content and gaze behavior, enables significant computation reduction with minimal loss in perceived quality. To support low-power, untethered MR setups, we design an edge-cloud collaborative rendering framework to partially offload computation to the cloud, further reducing overhead on the edge MR devices. Objective metrics and MR user study evidence that, compared to vanilla and foveated LOD baselines, our method achieves superior trade-offs between computational efficiency and user-perceived visual quality.
Sai Harsha Mupparaju, Kenneth Chen, Jenna Kang, Maito Omori, Kazuyuki Arimatsu, Qi Sun 0003
WACV8
2026 ML-PEA: Machine Learning-Based Perceptual Algorithms for Display Power Optimization
abstract
Abstract Image processing techniques can be used to modulate the pixel intensities of an image to reduce the power consumption of the display device. A simple example of this consists of uniformly dimming the entire image. Such algorithms should strive to minimize the impact on image quality while maximizing power savings. Techniques based on heuristics or human perception have been proposed, both for traditional flat panel displays and modern display modalities such as virtual and augmented reality (VR/AR). In this paper, we focus on developing and evaluating display power‐saving techniques that use machine learning (ML) in VR displays. We developed a U‐Net‐based technique paired with perceptual and power optimization loss functions that generates spatially varying dimming maps. These dimming maps are used to modulate input images, per‐pixel, to generate a power‐efficient image. Our pipeline was validated via quantitative analysis using image quality metrics and through a subjective study. Our subjective validation provides results scaled in perceptual just‐objectionable‐difference (JOD) units. This data, when rescaled, allows for comparisons of our technique with recent studies on VR display power optimization. Our results show that participants prefer our technique over a uniform dimming baseline for high target power saving conditions. This model and study serve as a template and baseline for future applications of deep learning to display power optimization. Model training code and data can be found at kenchen10.github.io/projects/mlpea/index.html .
Kenneth Chen, Nathan Matsuda, Thomas Wan, Ajit Ninan, Alexandre Chapiro, Qi Sun 0003
Comput. Graph. Forum6
2026 Overdriving Visual Depth Perception via Sound Modulation in VR
abstract
Our ability to perceive and navigate the spatial world is a cornerstone of human experience, relying on the integration of visual and auditory cues to form a coherent sense of depth and distance. In stereoscopic 3D vision, depth perception requires fixation of both eyes on a target object, which is achieved through vergence movements, with convergence for near objects and divergence for distant ones. In contrast, auditory cues provide complementary depth information through variations in loudness, interaural differences (IAD), and the frequency spectrum. We investigate the interaction between visual and auditory cues and examine how contradictory auditory information can overdrive visual depth perception in virtual reality (VR). When a new visual target appears, we introduce a spatial discrepancy between the visual and auditory cues: the visual target is shifted closer to the previously fixated object, while the corresponding sound localization is displaced in the opposite direction. By integrating these conflicting cues through multimodal processing, the resulting percept is biased toward the intended depth location. This audiovisual fusion counteracts depth compression, thus reducing the required vergence magnitude and enabling faster gaze retargeting. Such audio-driven depth enhancement may further help mitigate the vergence-accommodation conflict (VAC) in scenarios where physical depth must be compressed. In a series of psychophysical studies, we first assess the efficiency of depth overdriving for various VR-relevant combinations of initial fixations and shifted target locations, considering different scenarios of audio displacements and their loudness and frequency parameters. Next, we quantify the resulting speedup in gaze retargeting for target shifts that can be successfully overdriven by sound manipulations. Finally, we apply our method in a naturalistic VR scenario where user interface interactions with the scene show an extended perceptual depth.
Daniel Jiménez Navarro, Colin Groth, Jorge Pina, Qi Sun 0003, Praneeth Chakravarthula, Karol Myszkowski, Hans-Peter Seidel, Ana Serrano
IEEE Trans. Vis. Comput. Graph.5
2025 Process Only Where You Look: Hardware and Algorithm Co-optimization for Efficient Gaze-Tracked Foveated Rendering in Virtual Reality
abstract
Virtual reality (VR) plays a crucial role in advancing immersive, interactive experiences that transform learning, work, and entertainment by enhancing user engagement and expanding possibilities across various fields.Image rendering is one of the most crucial application in VR, as it produces high-quality, realistic visuals that are vital for maintaining immersive user experiences and preventing visual discomfort or motion sickness.However, the cost of image rendering in VR environment is considerable, primarily due to the demands of high-quality visual experiences from users.This challenge is even greater in real-time applications, where maintaining low latency further increases the complexity of the rendering process.On the other hand, VR devices, such as head-mounted displays (HMDs), are intrinsically linked to human behavior, using insights from perception and cognition to enhance user experience.In this work, we aim to reduce the high computational costs of the rendering process in VR by leveraging natural human eye dynamics and focusing on processing only where you look (POLO).This involves co-optimizing AI algorithms with underlying hardware for greater efficiency.We introduce POLONet, an efficient multitask deep learning framework designed to track human eye movements with minimal latency.Integrated with the POLO accelerator as a plug-in for VR HMD SoCs, this approach significantly lowers image rendering costs, achieving up to a 3.9× reduction in end-to-end latency compared to the latest gaze tracking methods.
Wenxuan Liu 0006, Kenneth Chen, Qi Sun 0003, Sai Qian Zhang
ISCA4
2025 Audiovisual Disparities in VR: Impact on Spatial Perception
abstract
Virtual reality (VR) experiences often leverage rich and spatialized multimodal environments to increase immersion and engagement. This demands a consistent spatial perception of audiovisual stimuli, since perceived discrepancies can disrupt the sense of presence. In this work, we investigate the consequences of two types of spatial audiovisual disparities: true disparity, where there is a measurable spatial offset between auditory and visual cues, and perceptual disparity, where users report misalignment despite cues being colocated. Unlike most previous studies that employed controlled but simplified experimental setups, our research focuses on complex, realistic VR environments, allowing us to assess the actual implications for VR content design. Our experiments indicate that users are highly sensitive to true audiovisual disparities in controlled environments, detecting even minor misalignments. However, when engaged in additional tasks within realistic settings, their ability to notice such discrepancies diminishes significantly. We also observed that previously found perceptual disparities persist in complex audiovisual environments. However, we identify self-initiated head rotations as a key factor; its absence prevents the effect entirely. We hope our findings offer practical insights for designing more immersive and perceptually coherent VR experiences.
Edurne Bernal-Berdun, Mateo Vallejo, Qi Sun 0003, Ana Serrano, Diego Gutierrez
ISMAR3
2025 Why Slow Feels Fast and Fast Feels Slow: Evaluating and Predicting Speed Misperception
abstract
Human perception of speed is largely driven by visual cues. However, our subjective estimations of speed are influenced by several factors that can lead to deceptive cues and speed misperception. While some prior studies have explored individual effects on speed perception, such as contrast, spatial frequency, and temporal frequency, their combined influence remains underexamined, particularly in immersive VR environments. In this work, we systematically investigate the influence and interplay of four visual factors—contrast, spatial frequency, temporal frequency, and eccentricity—on human perception of speed. To this end, we conduct a psychophysical study measuring subjective speed judgments across controlled stimuli and reveal significant perceptual biases induced by these factors. Based on our collected data, we learn a model to predict the underestimation or overestimation of perceived speed from visual scene properties. We apply and validate our findings in three immersive environments and demonstrate their influence on common VR scenarios. Finally, we discuss how understanding the factors that shape speed perception can drive the design of perceptually aligned virtual environments, with potential future applications such as correcting speed misperception and conceivably mitigating visual-vestibular conflicts by modulating perceived speed.
Colin Groth, Daniel Jiménez Navarro, Zihao Zou, Ana Serrano, Karol Myszkowski, Qi Sun 0003, Praneeth Chakravarthula
ISMAR8
2025 Performance Analysis of Catch-Up Eye Movements in Visual Tracking
abstract
In graphics applications featuring dynamically moving visual targets – such as film and gaming – we have to rotate our eyes to follow objects as they move across the screen. Because target motion is often unpredictable and ever-changing, we must rapidly respond to motion cues and adjust eye movements to maintain the target within the fovea, a process known as catch-up. This catch-up behavior reflects how efficiently the eyes react to and compensate for sudden changes in motion, making it a critical indicator for both task performance and the overall visual experience. In this work, we study and measure the eye catch-up performance during visual tracking. In particular, we present a behavioral analysis that predicts users’ reaction latency to abrupt target motion based on target visibility. Our numerical analysis and human subject studies evidence the effectiveness and generalizability. We further show how the catch-up metric can be applied to evaluate video quality, adjust game difficulty, and optimize display configurations for enhanced user performance. We envision this research to create a computational link between human perception and behavioral performance in dynamic graphics contexts.
Jenna Kang, Budmonde Duinkharjav, Niall L. Williams, Qi Sun 0003
SIGGRAPH Asia4
2025 Perceptually-Guided Acoustic "Foveation"
abstract
Realistic spatial audio rendering improves immersion in virtual environments. However, the computational complexity of acoustic propagation increases linearly with the number of sources. Consequently, real-time accurate acoustic rendering becomes challenging in highly dynamic scenarios such as virtual and augmented reality (VR/AR). Exploiting the fact that human spatial sensitivity of acoustic sources is not equal at azimuth eccentricities in the horizontal plane, we introduce a perceptually-aware acoustic "foveation" guidance model to the audio rendering pipeline, which can integrate audio sources that are not spatially resolvable by human listeners. To this end, we first conduct a series of psychophysical studies to measure the minimum resolvable audible angular distance under various spatial and background conditions. We leverage this data to derive an azimuth-characterized real-time acoustic foveation algorithm. Numerical analysis and subjective user studies in VR environments demonstrate our method’s effectiveness in significantly reducing acoustic rendering workload, without compromising users’ spatial perception of audio sources. We believe that the presented research will motivate future investigation into the new frontier of modeling and leveraging human multimodal perceptual limitations — beyond the extensively studied visual acuity — for designing efficient VR/AR systems.
Kenneth Chen, Irán R. Román, Juan Pablo Bello, Qi Sun 0003, Praneeth Chakravarthula
VR5
2025 HuBar: A Visual Analytics Tool to Explore Human Behavior Based on fNIRS in AR Guidance Systems
abstract
The concept of an intelligent augmented reality (AR) assistant has significant, wide-ranging applications, with potential uses in medicine, military, and mechanics domains. Such an assistant must be able to perceive the environment and actions, reason about the environment state in relation to a given task, and seamlessly interact with the task performer. These interactions typically involve an AR headset equipped with sensors which capture video, audio, and haptic feedback. Previous works have sought to facilitate the development of intelligent AR assistants by visualizing these sensor data streams in conjunction with the assistant's perception and reasoning model outputs. However, existing visual analytics systems do not focus on user modeling or include biometric data, and are only capable of visualizing a single task session for a single performer at a time. Moreover, they typically assume a task involves linear progression from one step to the next. We propose a visual analytics system that allows users to compare performance during multiple task sessions, focusing on non-linear tasks where different step sequences can lead to success. In particular, we design visualizations for understanding user behavior through functional near-infrared spectroscopy (fNIRS) data as a proxy for perception, attention, and memory as well as corresponding motion data (acceleration, angular velocity, and gaze). We distill these insights into embedding representations that allow users to easily select groups of sessions with similar behaviors. We provide two case studies that demonstrate how to use these visualizations to gain insights about task performance using data collected during helicopter copilot training tasks. Finally, we evaluate our approach through an in-depth examination of a think-aloud experiment with five domain experts.
Sonia Castelo Quispe, João Rulff, Parikshit Solunke, Erin McGowan, Guande Wu, Irán R. Román, Roque Lopez, Bea Steers, Qi Sun 0003, Juan Pablo Bello, Bradley Feest, Michael Middleton, Ryan McKendrick, Cláudio T. Silva
IEEE Trans. Vis. Comput. Graph.9
2025 Message from the ISMAR 2025 Science and Technology Paper Chairs and TVCG Guest Editors
abstract
In this special issue of IEEE Transactions on Visualization and Computer Graphics (TVCG), we are pleased to present the journal papers from the 24th IEEE International Symposium on Mixed and Augmented Reality (ISMAR 2025), which will be held between October 8 and 12, 2025, in Daejeon, South Korea. ISMAR continues the over twenty-year-long tradition of IWAR, ISMR, and ISAR, and is the world's premier conference for Mixed and Augmented Reality.
Ulrich Eck, Gun A. Lee, Alexander Plopski, Missie Smith, Qi Sun 0003, Markus Tatzgern
IEEE Trans. Vis. Comput. Graph.5
2025 FovealNet: Advancing AI-Driven Gaze Tracking Solutions for Efficient Foveated Rendering in Virtual Reality
abstract
Leveraging real-time eye tracking, foveated rendering optimizes hardware efficiency and enhances visual quality virtual reality (VR). This approach leverages eye-tracking techniques to determine where the user is looking, allowing the system to render high-resolution graphics only in the foveal region-the small area of the retina where visual acuity is highest, while the peripheral view is rendered at lower resolution. However, modern deep learning-based gaze-tracking solutions often exhibit a long-tail distribution of tracking errors, which can degrade user experience and reduce the benefits of foveated rendering by causing misalignment and decreased visual quality. This paper introduces FovealNet, an advanced AI-driven gaze tracking framework designed to optimize system performance by strategically enhancing gaze tracking accuracy. To further reduce the implementation cost of the gaze tracking algorithm, FovealNet employs an event-based cropping method that eliminates over 64.8% of irrelevant pixels from the input image. Additionally, it incorporates a simple yet effective token-pruning strategy that dynamically removes tokens on the fly without compromising tracking accuracy. Finally, to support different runtime rendering configurations, we propose a system performance-aware multi-resolution training strategy, allowing the gaze tracking DNN to adapt and optimize overall system performance more effectively. Evaluation results demonstrate that FovealNet achieves at least 1.42× speed up compared to previous methods and 13% increase in perceptual quality for foveated output. The code is available at https://github.com/wl3181/FovealNet.
Wenxuan Liu 0006, Budmonde Duinkharjav, Qi Sun 0003, Sai Qian Zhang
IEEE Trans. Vis. Comput. Graph.3
2024 Exploiting Human Color Discrimination for Memory- and Energy-Efficient Image Encoding in Virtual Reality
abstract
Virtual Reality (VR) has the potential of becoming the next ubiquitous computing platform. Continued progress in the burgeoning field of VR depends critically on an efficient computing substrate. In particular, DRAM access energy is known to contribute to a significant portion of system energy. Today's framebuffer compression system alleviates the DRAM traffic by using a numerically lossless compression algorithm. Being numerically lossless, however, is unnecessary to preserve perceptual quality for humans. This paper proposes a perceptually lossless, but numerically lossy, system to compress DRAM traffic. Our idea builds on top of long-established psychophysical studies that show that humans cannot discriminate colors that are close to each other. The discrimination ability becomes even weaker (i.e., more colors are perceptually indistinguishable) in our peripheral vision. Leveraging the color discrimination (in)ability, we propose an algorithm that adjusts pixel colors to minimize the bit encoding cost without introducing visible artifacts. The algorithm is coupled with lightweight architectural support that, in real-time, reduces the DRAM traffic by 66.9% and outperforms existing framebuffer compression mechanisms by up to 20.4%. Psychophysical studies on human participants show that our system introduce little to no perceptual fidelity degradation.
Nisarg Ujjainkar, Ethan Shahan, Kenneth Chen, Budmonde Duinkharjav, Qi Sun 0003, Yuhao Zhu 0001
ASPLOS (1)5
2024 Toward User-Aware Interactive Virtual Agents: Generative Multi-Modal Agent Behaviors in VR
abstract
Virtual agents serve as a vital interface within XR platforms. However, generating virtual agent behaviors typically rely on pre-coded actions or physics-based reactions. In this paper we present a learning-based multimodal agent behavior generation framework that adapts to users’ in-situ behaviors, similar to how humans interact with each other in the real world. By leveraging an in-house collected, dyadic conversational behavior dataset, we trained a conditional variational autoencoder (CVAE) model to achieve user-conditioned generation of virtual agents’ behaviors. Together with large language models (LLM), our approach can generate both the verbal and non-verbal reactive behaviors of virtual agents. Our comparative user study confirmed our method’s superiority over conventional animation graph-based baseline techniques, particularly regarding user-centric criteria. Thorough analyses of our results underscored the authentic nature of our virtual agents’ interactions and the heightened user engagement during VR interaction.
Bhasura S. Gunawardhana, Qi Sun 0003, Zhigang Deng 0001
ISMAR3
2024 May the Force Be with You: Dexterous Finger Force-Aware VR Interface
abstract
Advances in virtual reality (VR) have reduced experience differentials for users. However, gaps between reality and virtuality persist in tasks that require coupling users’ multimodal physical skills with virtual environments in delicate ways. User embodiment in VR easily breaks when physicality feels inauthentic, especially when users invoke their innate predilection to touch and manipulate things that they encounter. In this research, we examine the potential of forceaware VR interfaces for enabling natural connections to user physicality and evaluate them in high-finesse cases of touch. Combining surface electromyography (SEMG) with visual tracking, we develop an end-to-end learning-based system, ForceSense, to decode users’ dexterous finger forces from their forearm sEMG signals for direct usage in standard VR pipelines. This approach eliminates the need for hand-held tactile equipment, thereby promoting natural embodiment. A series of user studies on manipulation tasks in VR validate that ForceSense is more accurate, robust, and intuitive than alternative solutions. Two proofs-of-concept VR applications, calligraphy and piano playing, demonstrate that the good synergy between visual, auditory, and tactile modalities, as ForceSense affords, has the potential of enhancing users’ task learning performance in VR. Our source code and trained models are released at https://github.com/NYU-ICL/vr-force-aware-multimodal-interface.
Fengze Zhang, Sky Achitoff, Paul M. Torrens, Qi Sun 0003
ISMAR6
2024 GazeFusion: Saliency-Guided Image Generation
abstract
Diffusion models offer unprecedented image generation power given just a text prompt. While emerging approaches for controlling diffusion models have enabled users to specify the desired spatial layouts of the generated content, they cannot predict or control where viewers will pay more attention due to the complexity of human vision. Recognizing the significance of attention-controllable image generation in practical applications, we present a saliency-guided framework to incorporate the data priors of human visual attention mechanisms into the generation process. Given a user-specified viewer attention distribution, our control module conditions a diffusion model to generate images that attract viewers’ attention toward the desired regions. To assess the efficacy of our approach, we performed an eye-tracked user study and a large-scale model-based saliency analysis. The results evidence that both the cross-user eye gaze distributions and the saliency models’ predictions align with the desired attention distributions. Lastly, we outline several applications, including interactive design of saliency guidance, attention suppression in unwanted regions, and adaptive generation for varied display/viewing conditions.
Connor Z. Lin, Gordon Wetzstein, Qi Sun 0003
ACM Trans. Appl. Percept.5
2024 PEA-PODs: Perceptual Evaluation of Algorithms for Power Optimization in XR Displays
abstract
Display power consumption is an emerging concern for untethered devices. This goes double for augmented and virtual extended reality (XR) displays, which target high refresh rates and high resolutions while conforming to an ergonomically light form factor. A number of image mapping techniques have been proposed to extend battery usage. However, there is currently no comprehensive quantitative understanding of how the power savings provided by these methods compare to their impact on visual quality. We set out to answer this question. To this end, we present a perceptual evaluation of algorithms (PEA) for power optimization in XR displays (PODs). Consolidating a portfolio of six power-saving display mapping approaches, we begin by performing a large-scale perceptual study to understand the impact of each method on perceived quality in the wild. This results in a unified quality score for each technique, scaled in just-objectionable-difference (JOD) units. In parallel, each technique is analyzed using hardware-accurate power models. The resulting JOD-to-Milliwatt transfer function provides a first-of-its-kind look into tradeoffs offered by display mapping techniques, and can be directly employed to make architectural decisions for power budgets on XR displays. Finally, we leverage our study data and power models to address important display power applications like the choice of display primary, power implications of eye tracking, and more 1 .
Kenneth Chen, Thomas Wan, Nathan Matsuda, Ajit Ninan, Alexandre Chapiro, Qi Sun 0003
ACM Trans. Graph.6
2024 Evaluating Visual Perception of Object Motion in Dynamic Environments
abstract
Precisely understanding how objects move in 3D is essential for broad scenarios such as video editing, gaming, driving, and athletics. With screen-displayed computer graphics content, users only perceive limited cues to judge the object motion from the on-screen optical flow. Conventionally, visual perception is studied with stationary settings and singular objects. However, in practical applications, we---the observer---also move within complex scenes. Therefore, we must extract object motion from a combined optical flow displayed on screen, which can often lead to mis-estimations due to perceptual ambiguities. We measure and model observers' perceptual accuracy of object motions in dynamic 3D environments, a universal but under-investigated scenario in computer graphics applications. We design and employ a crowdsourcing-based psychophysical study, quantifying the relationships among patterns of scene dynamics and content, and the resulting perceptual judgments of object motion direction. The acquired psychophysical data underpins a model for generalized conditions. We then demonstrate the model's guidance ability to significantly enhance users' understanding of task object motion in gaming and animation design. With applications in measuring and compensating for object motion errors in video and rendering, we hope the research establishes a new frontier for understanding and mitigating perceptual errors caused by the gap between screen-displayed graphics and the physical world.
Budmonde Duinkharjav, Jenna Jiayi Kang, Gavin S. P. Miller, Chang Xiao 0001, Qi Sun 0003
ACM Trans. Graph.5
2024 Modeling the Impact of Head-Body Rotations on Audio-Visual Spatial Perception for Virtual Reality Applications
abstract
Humans perceive the world by integrating multimodal sensory feedback, including visual and auditory stimuli, which holds true in virtual reality (VR) environments. Proper synchronization of these stimuli is crucial for perceiving a coherent and immersive VR experience. In this work, we focus on the interplay between audio and vision during localization tasks involving natural head-body rotations. We explore the impact of audio-visual offsets and rotation velocities on users' directional localization acuity for various viewing modes. Using psychometric functions, we model perceptual disparities between visual and auditory cues and determine offset detection thresholds. Our findings reveal that target localization accuracy is affected by perceptual audio-visual disparities during head-body rotations, but remains consistent in the absence of stimuli-head relative motion. We then showcase the effectiveness of our approach in predicting and enhancing users' localization accuracy within realistic VR gaming applications. To provide additional support for our findings, we implement a natural VR game wherein we apply a compensatory audio-visual offset derived from our measured psychometric functions. As a result, we demonstrate a substantial improvement of up to 40% in participants' target localization accuracy. We additionally provide guidelines for content creation to ensure coherent and seamless VR experiences.
Edurne Bernal-Berdun, Mateo Vallejo, Qi Sun 0003, Ana Serrano, Diego Gutierrez
IEEE Trans. Vis. Comput. Graph.3
2024 : Visualization of AI-Assisted Task Guidance in AR
abstract
The concept of augmented reality (AR) assistants has captured the human imagination for decades, becoming a staple of modern science fiction. To pursue this goal, it is necessary to develop artificial intelligence (AI)-based methods that simultaneously perceive the 3D environment, reason about physical tasks, and model the performer, all in real-time. Within this framework, a wide variety of sensors are needed to generate data across different modalities, such as audio, video, depth, speech, and time-of-flight. The required sensors are typically part of the AR headset, providing performer sensing and interaction through visual, audio, and haptic feedback. AI assistants not only record the performer as they perform activities, but also require machine learning (ML) models to understand and assist the performer as they interact with the physical world. Therefore, developing such assistants is a challenging task. We propose ARGUS, a visual analytics system to support the development of intelligent AR assistants. Our system was designed as part of a multi-year-long collaboration between visualization researchers and ML and AR experts. This co-design process has led to advances in the visualization of ML in AR. Our system allows for online visualization of object, action, and step detection as well as offline analysis of previously recorded AR sessions. It visualizes not only the multimodal sensor data streams but also the output of the ML models. This allows developers to gain insights into the performer activities as well as the ML models, helping them troubleshoot, improve, and fine-tune the components of the AR assistant.
Sonia Castelo Quispe, João Rulff, Erin McGowan, Bea Steers, Guande Wu, Shaoyu Chen, Irán R. Román, Roque Lopez, Ethan Brewer, Chen Zhao 0013, Kyunghyun Cho, He He 0001, Qi Sun 0003, Huy T. Vo, Juan Pablo Bello, Michael Krone, Cláudio T. Silva
IEEE Trans. Vis. Comput. Graph.14
2024 Measuring and Predicting Multisensory Reaction Latency: A Probabilistic Model for Visual-Auditory Integration
abstract
Virtual/augmented reality (VR/AR) devices offer both immersive imagery and sound. With those wide-field cues, we can simultaneously acquire and process visual and auditory signals to quickly identify objects, make decisions, and take action. While vision often takes precedence in perception, our visual sensitivity degrades in the periphery. In contrast, auditory sensitivity can exhibit an opposite trend due to the elevated interaural time difference. What occurs when these senses are simultaneously integrated, as is common in VR applications such as 360° video watching and immersive gaming? We present a computational and probabilistic model to predict VR users' reaction latency to visual-auditory multisensory targets. To this aim, we first conducted a psychophysical experiment in VR to measure the reaction latency by tracking the onset of eye movements. Experiments with numerical metrics and user studies with naturalistic scenarios showcase the model's accuracy and generalizability. Lastly, we discuss the potential applications, such as measuring the sufficiency of target appearance duration in immersive video playback, and suggesting the optimal spatial layouts for AR interface design.
Daniel Jiménez Navarro, Ana Serrano, Karol Myszkowski, Qi Sun 0003
IEEE Trans. Vis. Comput. Graph.6
2023 The Shortest Route is Not Always the Fastest: Probability-Modeled Stereoscopic Eye Movement Completion Time in VR
abstract
Speed and consistency of target-shifting play a crucial role in human ability to perform complex tasks. Shifting our gaze between objects of interest quickly and consistently requires changes both in depth and direction. Gaze changes in depth are driven by slow, inconsistent vergence movements which rotate the eyes in opposite directions, while changes in direction are driven by ballistic, consistent movements called saccades , which rotate the eyes in the same direction. In the natural world, most of our eye movements are a combination of both types. While scientific consensus on the nature of saccades exists, vergence and combined movements remain less understood and agreed upon. We eschew the lack of scientific consensus in favor of proposing an operationalized computational model which predicts the completion time of any type of gaze movement during target-shifting in 3D. To this end, we conduct a psychophysical study in a stereo VR environment to collect more than 12,000 gaze movement trials, analyze the temporal distribution of the observed gaze movements, and fit a probabilistic model to the data. We perform a series of objective measurements and user studies to validate the model. The results demonstrate its predictive accuracy, generalization, as well as applications for optimizing visual performance by altering content placement. Lastly, we leverage the model to measure differences in human target-changing time relative to the natural world, as well as suggest scene-aware projection depth. By incorporating the complexities and randomness of human oculomotor control, we hope this research will support new behavior-aware metrics for VR/AR display design, interface layout, and gaze-contingent rendering.
Budmonde Duinkharjav, Benjamin Liang, Anjul Patney, Rachel Brown, Qi Sun 0003
ACM Trans. Graph.5
2022 Dually Noted: Layout-Aware Annotations with Smartphone Augmented Reality
abstract
Sharing annotations encourages feedback, discussion, and knowledge passing among readers and can be beneficial for personal and public use. Prior augmented reality (AR) systems have expanded these benefits to both digital and printed documents. However, despite smartphone AR now being widely available, there is a lack of research about how to use AR effectively for interactive document annotation. We propose Dually Noted, a smartphone-based AR annotation system that recognizes the layout of structural elements in a printed document for real-time authoring and viewing of annotations. We conducted experience prototyping with eight users to elicit potential benefits and challenges within smartphone AR, and this informed the resulting Dually Noted system and annotation interactions with the document elements. AR annotation is often unwieldy, but during a 12-user empirical study our novel structural understanding component allows Dually Noted to improve precise highlighting and annotation interaction accuracy by 13%, increase interaction speed by 42%, and significantly lower cognitive load over a baseline method without document layout understanding. Qualitatively, participants commented that Dually Noted was a swift and portable annotation experience. Overall, our research provides new methods and insights for how to improve AR annotations for physical documents.
Qi Sun 0003, Curtis Wigington, Han L. Han, Tong Sun 0005, Jennifer A. Healey, James Tompkin 0001, Jeff Huang 0002
CHI2
2022 Image features influence reaction time: a learned probabilistic perceptual model for saccade latency
abstract
We aim to ask and answer an essential question " how quickly do we react after observing a displayed visual target?" To this end, we present psychophysical studies that characterize the remarkable disconnect between human saccadic behaviors and spatial visual acuity. Building on the results of our studies, we develop a perceptual model to predict temporal gaze behavior, particularly saccadic latency, as a function of the statistics of a displayed image. Specifically, we implement a neurologically-inspired probabilistic model that mimics the accumulation of confidence that leads to a perceptual decision. We validate our model with a series of objective measurements and user studies using an eye-tracked VR display. The results demonstrate that our model prediction is in statistical alignment with real-world human behavior. Further, we establish that many sub-threshold image modifications commonly introduced in graphics pipelines may significantly alter human reaction timing, even if the differences are visually undetectable. Finally, we show that our model can serve as a metric to predict and alter reaction latency of users in interactive computer graphics applications, thus may improve gaze-contingent rendering, design of virtual experiences, and player performance in e-sports. We illustrate this with two examples: estimating competition fairness in a video game with two different team colors, and tuning display viewing distance to minimize player reaction time.
Budmonde Duinkharjav, Praneeth Chakravarthula, Rachel Brown, Anjul Patney, Qi Sun 0003
ACM Trans. Graph.5
2022 Color-Perception-Guided Display Power Reduction for Virtual Reality
abstract
Battery life is an increasingly urgent challenge for today's untethered VR and AR devices. However, the power efficiency of head-mounted displays is naturally at odds with growing computational requirements driven by better resolution, refresh rate, and dynamic ranges, all of which reduce the sustained usage time of untethered AR/VR devices. For instance, the Oculus Quest 2, under a fully-charged battery, can sustain only 2 to 3 hours of operation time. Prior display power reduction techniques mostly target smartphone displays. Directly applying smartphone display power reduction techniques, however, degrades the visual perception in AR/VR with noticeable artifacts. For instance, the "power-saving mode" on smartphones uniformly lowers the pixel luminance across the display and, as a result, presents an overall darkened visual perception to users if directly applied to VR content. Our key insight is that VR display power reduction must be cognizant of the gaze-contingent nature of high field-of-view VR displays. To that end, we present a gaze-contingent system that, without degrading luminance, minimizes the display power consumption while preserving high visual fidelity when users actively view immersive video sequences. This is enabled by constructing 1) a gaze-contingent color discrimination model through psychophysical studies, and 2) a display power model (with respect to pixel color) through real-device measurements. Critically, due to the careful design decisions made in constructing the two models, our algorithm is cast as a constrained optimization problem with a closed-form solution, which can be implemented as a real-time, image-space shader. We evaluate our system using a series of psychophysical studies and large-scale analyses on natural images. Experiment results show that our system reduces the display power by as much as 24% (14% on average) with little to no perceptual fidelity degradation.
Budmonde Duinkharjav, Kenneth Chen, Abhishek Tyagi, Yuhao Zhu 0001, Qi Sun 0003
ACM Trans. Graph.6
2022 Joint neural phase retrieval and compression for energy- and computation-efficient holography on the edge
abstract
Recent deep learning approaches have shown remarkable promise to enable high fidelity holographic displays. However, lightweight wearable display devices cannot afford the computation demand and energy consumption for hologram generation due to the limited onboard compute capability and battery life. On the other hand, if the computation is conducted entirely remotely on a cloud server, transmitting lossless hologram data is not only challenging but also result in prohibitively high latency and storage. In this work, by distributing the computation and optimizing the transmission, we propose the first framework that jointly generates and compresses high-quality phase-only holograms. Specifically, our framework asymmetrically separates the hologram generation process into high-compute remote encoding (on the server), and low-compute decoding (on the edge) stages. Our encoding enables light weight latent space data, thus faster and efficient transmission to the edge device. With our framework, we observed a reduction of 76% computation and consequently 83% in energy cost on edge devices, compared to the existing hologram generation methods. Our framework is robust to transmission and decoding errors, and approach high image fidelity for as low as 2 bits-per-pixel, and further reduced average bit-rates and decoding time for holographic videos.
Praneeth Chakravarthula, Qi Sun 0003, Baoquan Chen
ACM Trans. Graph.3
2022 Force-Aware Interface via Electromyography for Natural VR/AR Interaction
abstract
While tremendous advances in visual and auditory realism have been made for virtual and augmented reality (VR/AR), introducing a plausible sense of physicality into the virtual world remains challenging. Closing the gap between real-world physicality and immersive virtual experience requires a closed interaction loop: applying user-exerted physical forces to the virtual environment and generating haptic sensations back to the users. However, existing VR/AR solutions either completely ignore the force inputs from the users or rely on obtrusive sensing devices that compromise user experience. By identifying users' muscle activation patterns while engaging in VR/AR, we design a learning-based neural interface for natural and intuitive force inputs. Specifically, we show that lightweight electromyography sensors, resting non-invasively on users' forearm skin, inform and establish a robust understanding of their complex hand activities. Fuelled by a neural-network-based model, our interface can decode finger-wise forces in real-time with 3.3% mean error, and generalize to new users with little calibration. Through an interactive psychophysical study, we show that human perception of virtual objects' physical properties, such as stiffness, can be significantly enhanced by our interface. We further demonstrate that our interface enables ubiquitous control via finger tapping. Ultimately, we envision our findings to push forward research towards more realistic physicality in future VR/AR.
Benjamin Liang, Boyuan Chen 0004, Paul M. Torrens, Seyed Farokh Atashzar, Dahua Lin, Qi Sun 0003
ACM Trans. Graph.7
2022 Instant Reality: Gaze-Contingent Perceptual Optimization for 3D Virtual Reality Streaming
abstract
Media streaming, with an edge-cloud setting, has been adopted for a variety of applications such as entertainment, visualization, and design. Unlike video/audio streaming where the content is usually consumed passively, virtual reality applications require 3D assets stored on the edge to facilitate frequent edge-side interactions such as object manipulation and viewpoint movement. Compared to audio and video streaming, 3D asset streaming often requires larger data sizes and yet lower latency to ensure sufficient rendering quality, resolution, and latency for perceptual comfort. Thus, streaming 3D assets faces remarkably additional than streaming audios/videos, and existing solutions often suffer from long loading time or limited quality. To address this challenge, we propose a perceptually-optimized progressive 3D streaming method for spatial quality and temporal consistency in immersive interactions. On the cloud-side, our main idea is to estimate perceptual importance in 2D image space based on user gaze behaviors, including where they are looking and how their eyes move. The estimated importance is then mapped to 3D object space for scheduling the streaming priorities for edge-side rendering. Since this computational pipeline could be heavy, we also develop a simple neural network to accelerate the cloud-side scheduling process. We evaluate our method via subjective studies and objective analysis under varying network conditions (from 3G to 5G) and edge devices (HMD and traditional displays), and demonstrate better visual quality and temporal consistency than alternative solutions.
Shaoyu Chen, Budmonde Duinkharjav, Xin Sun 0014, Li-Yi Wei, Stefano Petrangeli, Jose Echevarria, Cláudio T. Silva, Qi Sun 0003
IEEE Trans. Vis. Comput. Graph.8
2022 FoV-NeRF: Foveated Neural Radiance Fields for Virtual Reality
abstract
Virtual Reality (VR) is becoming ubiquitous with the rise of consumer displays and commercial VR platforms. Such displays require low latency and high quality rendering of synthetic imagery with reduced compute overheads. Recent advances in neural rendering showed promise of unlocking new possibilities in 3D computer graphics via image-based representations of virtual or physical environments. Specifically, the neural radiance fields (NeRF) demonstrated that photo-realistic quality and continuous view changes of 3D scenes can be achieved without loss of view-dependent effects. While NeRF can significantly benefit rendering for VR applications, it faces unique challenges posed by high field-of-view, high resolution, and stereoscopic/egocentric viewing, typically causing low quality and high latency of the rendered images. In VR, this not only harms the interaction experience but may also cause sickness. To tackle these problems toward six-degrees-of-freedom, egocentric, and stereo NeRF in VR, we present the first gaze-contingent 3D neural representation and view synthesis method. We incorporate the human psychophysics of visual- and stereo-acuity into an egocentric neural representation of 3D scenery. We then jointly optimize the latency/performance and visual quality while mutually bridging human perception and neural scene synthesis to achieve perceptually high-quality immersive interaction. We conducted both objective analysis and subjective studies to evaluate the effectiveness of our approach. We find that our method significantly reduces latency (up to 99% time reduction compared with NeRF) without loss of high-fidelity rendering (perceptually identical to full-resolution ground truth). The presented approach may serve as the first step toward future VR/AR systems that capture, teleport, and visualize remote environments in real-time.
Nianchen Deng, Zhenyi He, Jiannan Ye, Budmonde Duinkharjav, Praneeth Chakravarthula, Xubo Yang, Qi Sun 0003
IEEE Trans. Vis. Comput. Graph.7
2021 Tailored Reality: Perception-aware Scene Restructuring for Adaptive VR Navigation
abstract
In virtual reality (VR), the virtual scenes are pre-designed by creators. Our physical surroundings, however, comprise significantly varied sizes, layouts, and components. To bridge the gap and further enable natural navigation, recent solutions have been proposed to redirect users or recreate the virtual content. However, they suffer from either interrupted experience or distorted appearances. We present a novel VR-oriented algorithm that automatically restructures a given virtual scene for a user’s physical environment. Different from the previous methods, we introduce neither interrupted walking experience nor curved appearances. Instead, a perception-aware function optimizes our retargeting technique to preserve the fidelity of the virtual scene that appears in VR head-mounted displays. Besides geometric and topological properties, it emphasizes the unique first-person view perceptual factors in VR, such as dynamic visibility and objectwise relationships. We conduct both analytical experiments and subjective studies. The results demonstrate our system’s versatile capability and practicability for natural navigation in VR: It reduces the virtual space by 40% without statistical loss of perceptual identicality.
Zhichao Dong 0001, Wenming Wu 0001, Zenghao Xu, Qi Sun 0003, Guan-Jie Yuan, Ligang Liu 0001, Xiao-Ming Fu 0001
ACM Trans. Graph.4
2021 Gaze-Contingent Retinal Speckle Suppression for Perceptually-Matched Foveated Holographic Displays
abstract
Computer-generated holographic (CGH) displays show great potential and are emerging as the next-generation displays for augmented and virtual reality, and automotive heads-up displays. One of the critical problems harming the wide adoption of such displays is the presence of speckle noise inherent to holography, that compromises its quality by introducing perceptible artifacts. Although speckle noise suppression has been an active research area, the previous works have not considered the perceptual characteristics of the Human Visual System (HVS), which receives the final displayed imagery. However, it is well studied that the sensitivity of the HVS is not uniform across the visual field, which has led to gaze-contingent rendering schemes for maximizing the perceptual quality in various computer-generated imagery. Inspired by this, we present the first method that reduces the "perceived speckle noise" by integrating foveal and peripheral vision characteristics of the HVS, along with the retinal point spread function, into the phase hologram computation. Specifically, we introduce the anatomical and statistical retinal receptor distribution into our computational hologram optimization, which places a higher priority on reducing the perceived foveal speckle noise while being adaptable to any individual's optical aberration on the retina. Our method demonstrates superior perceptual quality on our emulated holographic display. Our evaluations with objective measurements and subjective studies demonstrate a significant reduction of the human perceived noise.
Praneeth Chakravarthula, Zhan Zhang 0009, Okan Tarhan Tursun, Piotr Didyk, Qi Sun 0003, Henry Fuchs
IEEE Trans. Vis. Comput. Graph.5
2020 Deep Multi Depth Panoramas for View Synthesis
Kai-En Lin, Zexiang Xu, Ben Mildenhall, Pratul P. Srinivasan, Yannick Hold-Geoffroy, Stephen DiVerdi, Qi Sun 0003, Kalyan Sunkavalli, Ravi Ramamoorthi
ECCV (13)7
2020 DiffTaichi: Differentiable Programming for Physical Simulation
Yuanming Hu, Luke Anderson 0001, Tzu-Mao Li, Qi Sun 0003, Nathan Carr 0001, Jonathan Ragan-Kelley, Frédo Durand
ICLR4
2019 Learning to Reconstruct 3D Manhattan Wireframes From a Single Image
abstract
From a single view of an urban environment, we propose a method to effectively exploit the global structural regularities for obtaining a compact, accurate, and intuitive 3D wireframe representation. Our method trains a single convolutional neural network to simultaneously detect salient junctions and straight lines, as well as predict their 3D depth and vanishing points. Compared with state-of-the-art learning-based wireframe detection methods, our network is much simpler and more unified, leading to better 2D wireframe detection. With a global structural prior (such as Manhattan assumption), our method further reconstructs a full 3D wireframe model, a compact vector representation suitable for a variety of high-level vision tasks such as AR and CAD. We conduct extensive evaluations of our method on a large new synthetic dataset of urban scenes as well as real images. Our code and datasets will be published along with the paper.
Yichao Zhou 0003, Haozhi Qi, Yuexiang Zhai, Qi Sun 0003, Li-Yi Wei, Yi Ma 0001
ICCV4
2019 Reducing simulator sickness with perceptual camera control
abstract
Virtual-reality provides an immersive environment but can induce cybersickness due to the discrepancy between visual and vestibular cues. To avoid this problem, the movement of the virtual camera needs to match the motion of the user in the real world. Unfortunately, this is usually difficult due to the mismatch between the size of the virtual environments and the space available to the users in the physical domain. The resulting constraints on the camera movement significantly hamper the adoption of virtual-reality headsets in many scenarios and make the design of the virtual environments very challenging. In this work, we study how the characteristics of the virtual camera movement (e.g., translational acceleration and rotational velocity) and the composition of the virtual environment (e.g., scene depth) contribute to perceived discomfort. Based on the results from our user experiments, we devise a computational model for predicting the magnitude of the discomfort for a given scene and camera trajectory. We further apply our model to a new path planning method which optimizes the input motion trajectory to reduce perceptual sickness. We evaluate the effectiveness of our method in improving perceptual comfort in a series of user studies targeting different applications. The results indicate that our method can reduce the perceived discomfort while maintaining the fidelity of the original navigation, and perform better than simpler alternatives.
Ping Hu 0003, Qi Sun 0003, Piotr Didyk, Li-Yi Wei, Arie E. Kaufman
ACM Trans. Graph.2
2018 Towards virtual reality infinite walking: dynamic saccadic redirection
abstract
Redirected walking techniques can enhance the immersion and visual-vestibular comfort of virtual reality (VR) navigation, but are often limited by the size, shape, and content of the physical environments. We propose a redirected walking technique that can apply to small physical environments with static or dynamic obstacles. Via a head- and eye-tracking VR headset, our method detects saccadic suppression and redirects the users during the resulting temporary blindness. Our dynamic path planning runs in real-time on a GPU, and thus can avoid static and dynamic obstacles, including walls, furniture, and other VR users sharing the same physical space. To further enhance saccadic redirection, we propose subtle gaze direction methods tailored for VR perception. We demonstrate that saccades can significantly increase the rotation gains during redirection without introducing visual distortions or simulator sickness. This allows our method to apply to large open virtual spaces and small physical environments for room-scale VR. We evaluate our system via numerical simulations and real user studies.
Qi Sun 0003, Anjul Patney, Li-Yi Wei, Omer Shapira, Jingwan Lu, Paul Asente, Suwen Zhu, Morgan McGuire, David P. Luebke, Arie E. Kaufman
ACM Trans. Graph.1
2017 Perceptually-guided foveation for light field displays
abstract
A variety of applications such as virtual reality and immersive cinema require high image quality, low rendering latency, and consistent depth cues. 4D light field displays support focus accommodation, but are more costly to render than 2D images, resulting in higher latency. The human visual system can resolve higher spatial frequencies in the fovea than in the periphery. This property has been harnessed by recent 2D foveated rendering methods to reduce computation cost while maintaining perceptual quality. Inspired by this, we present foveated 4D light fields by investigating their effects on 3D depth perception. Based on our psychophysical experiments and theoretical analysis on visual and display bandwidths, we formulate a content-adaptive importance model in the 4D ray space. We verify our method by building a prototype light field display that can render only 16% -- 30% rays without compromising perceptual quality.
Qi Sun 0003, Fu-Chung Huang, Joohwan Kim, Li-Yi Wei, David P. Luebke, Arie E. Kaufman
ACM Trans. Graph.1
2016 Mapping virtual and physical reality
abstract
Real walking offers higher immersive presence for virtual reality (VR) applications than alternative locomotive means such as walking-in-place and external control gadgets, but needs to take into consideration different room sizes, wall shapes, and surrounding objects in the virtual and real worlds. Despite perceptual study of impossible spaces and redirected walking, there are no general methods to match a given pair of virtual and real scenes. We propose a system to match a given pair of virtual and physical worlds for immersive VR navigation. We first compute a planar map between the virtual and physical floor plans that minimizes angular and distal distortions while conforming to the virtual environment goals and physical environment constraints. Our key idea is to design maps that are globally surjective to allow proper folding of large virtual scenes into smaller real scenes but locally injective to avoid locomotion ambiguity and intersecting virtual objects. From these maps we derive altered rendering to guide user navigation within the physical environment while retaining visual fidelity to the virtual environment. Our key idea is to properly warp the virtual world appearance into real world geometry with sufficient quality and performance. We evaluate our method through a formative user study, and demonstrate applications in gaming, architecture walkthrough, and medical imaging.
Qi Sun 0003, Li-Yi Wei, Arie E. Kaufman
ACM Trans. Graph.1