Zhenyi He

dblp:156/1146 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2023 Aware-Transformer: A Novel Pure Transformer-Based Model for Remote Sensing Image Captioning
Yukun Cao, Jialuo Yan, Yijia Tang, Zhenyi He, Kangle Xu
CGI (1)4
2023 Dialogue-Clues: Dual-channel Dialogue Clues Embedding Context Perception Network for Emotion Recognition in Conversations
abstract
In recent years, emotion recognition in conversations has gained widespread attention due to its extensive applications. Many recent studies have focused on perceiving conversational context from the perspective of capturing dialogue clues. However, these studies often use the entire conversation sequence to represent the dialogue clues, which may result in insufficient representation of the speaker's emotional dynamics in multi-turn conversation. To address these issues, we propose a novel dual-channel dialogue clues embedding context perception network (Dialogue-Clues) that integrates dialogue clues information into global conversation context modeling. We also introduce a dual-channel dialogue clues perception architecture that captures and reinforces both static and dynamic dialogue clues in conversations, and bidirectionally reinforces both types of dialogue clues information. To better represent dynamic dialogue clues, we construct a novel speaker-interaction isomorphic graph structure for the dual-channel dialogue clues perception architecture. Through extensive comparisons with ten existing methods on four public datasets, we confirm the effectiveness of the proposed method. The results demonstrate that integrating Dialogue-Clues information can improve the ability of context modeling.
Yukun Cao, Zhenyi He, Yijia Tang, Jialuo Yan, Kangle Xu
SMC2
2022 TapGazer: Text Entry with Finger Tapping and Gaze-directed Word Selection
abstract
While using VR, efficient text entry is a challenge: users cannot easily locate standard physical keyboards, and keys are often out of reach, e.g. when standing. We present TapGazer, a text entry system where users type by tapping their fingers in place. Users can tap anywhere as long as the identity of each tapping finger can be detected with sensors. Ambiguity between different possible input words is resolved by selecting target words with gaze. If gaze tracking is unavailable, ambiguity is resolved by selecting target words with additional taps. We evaluated TapGazer for seated and standing VR: seated novice users using touchpads as tap surfaces reached 44.81 words per minute (WPM), 79.17% of their QWERTY typing speed. Standing novice users tapped on their thighs with touch-sensitive gloves, reaching 45.26 WPM (71.91%). We analyze TapGazer with a theoretical performance model and discuss its potential for text input in future AR scenarios.
Zhenyi He, Christof Lutteroth, Ken Perlin
CHI1
2022 FoV-NeRF: Foveated Neural Radiance Fields for Virtual Reality
abstract
Virtual Reality (VR) is becoming ubiquitous with the rise of consumer displays and commercial VR platforms. Such displays require low latency and high quality rendering of synthetic imagery with reduced compute overheads. Recent advances in neural rendering showed promise of unlocking new possibilities in 3D computer graphics via image-based representations of virtual or physical environments. Specifically, the neural radiance fields (NeRF) demonstrated that photo-realistic quality and continuous view changes of 3D scenes can be achieved without loss of view-dependent effects. While NeRF can significantly benefit rendering for VR applications, it faces unique challenges posed by high field-of-view, high resolution, and stereoscopic/egocentric viewing, typically causing low quality and high latency of the rendered images. In VR, this not only harms the interaction experience but may also cause sickness. To tackle these problems toward six-degrees-of-freedom, egocentric, and stereo NeRF in VR, we present the first gaze-contingent 3D neural representation and view synthesis method. We incorporate the human psychophysics of visual- and stereo-acuity into an egocentric neural representation of 3D scenery. We then jointly optimize the latency/performance and visual quality while mutually bridging human perception and neural scene synthesis to achieve perceptually high-quality immersive interaction. We conducted both objective analysis and subjective studies to evaluate the effectiveness of our approach. We find that our method significantly reduces latency (up to 99% time reduction compared with NeRF) without loss of high-fidelity rendering (perceptually identical to full-resolution ground truth). The presented approach may serve as the first step toward future VR/AR systems that capture, teleport, and visualize remote environments in real-time.
Nianchen Deng, Zhenyi He, Jiannan Ye, Budmonde Duinkharjav, Praneeth Chakravarthula, Xubo Yang, Qi Sun 0003
IEEE Trans. Vis. Comput. Graph.2
2021 GazeChat: Enhancing Virtual Conferences with Gaze-aware 3D Photos
abstract
Communication software such as Clubhouse and Zoom has evolved to be an integral part of many people’s daily lives. However, due to network bandwidth constraints and concerns about privacy, cameras in video conferencing are often turned off by participants. This leads to a situation in which people can only see each others’ profile images, which is essentially an audio-only experience. Even when switched on, video feeds do not provide accurate cues as to who is talking to whom. This paper introduces GazeChat, a remote communication system that visually represents users as gaze-aware 3D profile photos. This satisfies users’ privacy needs while keeping online conversations engaging and efficient. GazeChat uses a single webcam to track whom any participant is looking at, then uses neural rendering to animate all participants’ profile images so that participants appear to be looking at each other. We have conducted a remote user study (N=16) to evaluate GazeChat in three conditions: audio conferencing with profile photos, GazeChat, and video conferencing. Based on the results of our user study, we conclude that GazeChat maintains the feeling of presence while preserving more privacy and requiring lower bandwidth than video conferencing, provides a greater level of engagement than to audio conferencing, and helps people to better understand the structure of their conversation.
Zhenyi He, Keru Wang, Brandon Yushan Feng, Ruofei Du, Ken Perlin
UIST1
2021 Render-based factorization for additive light field display
abstract
Abstract Augmented Reality (AR) and Virtual Reality (VR) applications enable viewers to experience 3D graphics immersively. However, current hardware either fails to provide multiple viewpoints like projectors or cause conflicts between vergence and accommodation like consumable headsets. Prior researchers have designed multi‐view displays to solve the problem by enabling view‐dependent images and focus cues. Nevertheless, it requires minutes to calculate one frame, which is critical for real‐time AR/VR applications. In this paper, we propose a new render‐based factorization for additive light field display and further improve the performance by optimizing the initialization of layers. We next compare our work with the state of the art and the results show that our solution has competitive results while the calculation takes less than 20 ms for each frame.
Nianchen Deng, Zhenyi He, Xubo Yang
Comput. Animat. Virtual Worlds2
2020 CollaboVR: A Reconfigurable Framework for Creative Collaboration in Virtual Reality
abstract
Writing or sketching on whiteboards is an essential part of collaborative discussions in business meetings, reading groups, design sessions, and interviews. However, prior work in collaborative virtual reality (VR) systems has rarely explored the design space of multi-user layouts and interaction modes with virtual whiteboards. In this paper, we present CollaboVR, a reconfigurable framework for both co-located and geographically dispersed multi-user communication in VR. Our system unleashes users’ creativity by sharing freehand drawings, converting 2D sketches into 3D models, and generating procedural animations in real-time. To minimize the computational expense for VR clients, we leverage a cloud architecture in which the computational expensive application (Chalktalk) is hosted directly on the servers, with results being simultaneously streamed to clients. We have explored three custom layouts – integrated, mirrored, and projective – to reduce visual clutter, increase eye contact, or adapt different use cases. To evaluate CollaboVR, we conducted a within-subject user study with 12 participants. Our findings reveal that users appreciate the custom configurations and real-time interactions provided by CollaboVR. We have open sourced CollaboVR at https://github.com/snowymo/CollaboVR to facilitate future research and development of natural user interfaces and real-time collaborative systems in virtual and augmented reality.
Zhenyi He, Ruofei Du, Ken Perlin
ISMAR1