Duotun Wang

dblp:225/8349 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0005-4393-5230ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Direct vs. Score-based Selection: Understanding the Heisenberg Effect in Target Acquisition Across Input Modalities in Virtual Reality
abstract
Target selection is a fundamental interaction in virtual reality (VR). But the act of confirming a selection, such as a button press or pinch, can disturb the tracked pose and shift the intended target, which is referred to as the Heisenberg Effect. Prior research has mainly investigated controller input. However, it remains unclear how the effect manifests in the bare-hand input and how score-based techniques may mitigate the effect in different spatial variations. To fill the gap, we conduct a within-subject study to examine the Heisenberg Effect across two input modalities (i.e., controller and hand) and two selection mechanisms (i.e., direct and score-based). Our results show that hand input is more susceptible to the Heisenberg Effect, with direct selection more influenced by target width and score-based selection more sensitive to target density. Based on previous vote-oriented technique and our temporal analysis, we introduce weighted VOTE, a history-based intention accuracy model for target voting, that reweights recent interaction intent to counteract input disturbances. Our evaluation shows the method improves selection accuracy compared to baseline techniques. Finally, we discuss future directions for adaptive selection methods.
Linjie Qiu, Duotun Wang, Boyu Li 0007, Jiawei Li 0009, Yulin Shen 0001, Zeyu Wang 0003, Mingming Fan 0001
IEEE Trans. Vis. Comput. Graph.2
2025 DEGAS: Detailed Expressions on Full-Body Gaussian Avatars
abstract
Although neural rendering has made significant ad-vances in creating lifelike, animatable full-body and head avatars, incorporating detailed expressions into full-body avatars remains largely unexplored. We present DEGAS, the first 3D Gaussian Splatting (3DGS)-based modeling method for full-body avatars with rich facial expressions. Trained on multiview videos of a given subject, our method learns a conditional variational autoencoder that takes both the body motion and facial expression as driving signals to generate Gaussian maps in the UV layout. To drive the facial expressions, instead of the commonly used 3D Mor-phable Models (3DMMs) in 3D head avatars, we propose to adopt the expression latent space trained solely on 2D portrait images, bridging the gap between 2D talking faces and 3D avatars. Leveraging the rendering capability of 3DGS and the rich expressiveness of the expression latent space, the learned avatars can be reenacted to reproduce photo-realistic rendering images with subtle and accurate facial expressions. Experiments on an existing dataset and our newly proposed dataset offull-body talking avatars demonstrate the efficacy of our method. We also propose an audio-driven extension of our method with the help of 2D talking faces, opening new possibilities for interactive AI agents. Project page: https://initialneil.github.io/DEGAS.
Zhijing Shao, Duotun Wang, Qing-Yao Tian, Yao-Dong Yang, Hengyu Meng, Yu Zhang 0166, Kang Zhang 0001, Zeyu Wang 0003
3DV2
2025 HeadEvolver: Text to Head Avatars via Expressive and Attribute-Preserving Mesh Deformation
abstract
Current text-to-avatar methods often rely on implicit representations (e.g., NeRF, SDF, and DMTet), leading to 3D content that artists cannot easily edit and animate in graphics software. This paper introduces a novel framework for generating stylized head avatars from text guidance, which leverages locally learnable mesh deformation and 2D diffusion priors to achieve high-quality digital assets for attribute-preserving manipulation. Given a template mesh, our method represents mesh deformation with perface Jacobians and adaptively modulates local deformation using a learnable vector field. This vector field enables anisotropic scaling while preserving the rotation of vertices, which can better express identity and geometric details. We also employ landmark- and contour-based regularization terms to balance the expressiveness and plausibility of generated head avatars from multiple views without relying on any specific shape prior. Our framework can generate realistic shapes and textures that can be further edited via text, while supporting seamless editing using the preserved attributes from the template mesh, such as 3DMM parameters, blendshapes, and UV coordinates. Extensive experiments demonstrate that our framework can generate diverse and expressive head avatars with high-quality meshes that artists can easily manipulate in 3D graphics software, facilitating downstream applications such as efficient asset creation and animation with preserved attributes.
Duotun Wang, Hengyu Meng, Zhijing Shao, Qianxi Liu, Lin Wang 0025, Mingming Fan 0001, Xiaohang Zhan, Zeyu Wang 0003
3DV1
2025 Text2VDM: Text to Vector Displacement Maps for Expressive and Interactive 3D Sculpting
abstract
Professional 3D asset creation often requires diverse sculpting brushes to add surface details and geometric structures. Despite recent progress in 3D generation, producing reusable sculpting brushes compatible with artists' workflows remains an open and challenging problem. These sculpting brushes are typically represented as vector displacement maps (VDMs), which existing models cannot easily generate compared to natural images. This paper presents Text2VDM, a novel framework for text-to-VDM brush generation through the deformation of a dense planar mesh guided by score distillation sampling (SDS). The original SDS loss is designed for generating full objects and struggles with generating desirable sub-object structures from scratch in brush generation. We refer to this issue as semantic coupling, which we address by introducing weighted blending of prompt tokens to SDS, resulting in a more accurate target distribution and semantic guidance. Experiments demonstrate that Text2VDM can generate diverse, high-quality VDM brushes for sculpting surface details and geometric structures. Our generated brushes can be seamlessly integrated into mainstream modeling software, enabling various applications such as mesh stylization and real-time interactive modeling.
Hengyu Meng, Duotun Wang, Zhijing Shao, Zeyu Wang 0003
ICCV2
2025 DesignMemo: Integrating Discussion Context into Online Collaboration with Enhanced Design Rationale Tracking
abstract
Remote collaborative design has become increasingly popular, but current design tools often overlook the importance of contextual communication during synchronized design activities, which is critical for understanding the rationale and decisions behind design choices. In this paper, we introduce DesignMemo, a proof-of-concept system that integrates the verbal context of remote discussions into visual design history tracking. The system automatically labels the visual elements with an annotation, which is linked to a certain transcript of the meeting, so that the user can easily recall the context of the visual design by clicking the element. The system also integrates an LLM agent for annotation-oriented summarization based on global context tracking, so users can quickly follow the rationale of the design without reading the lengthy transcript. Our user study with 24 participants suggests that the ability to track communication context makes the iterative design process smoother and more efficient.
Boyu Li 0007, Linjie Qiu, Duotun Wang, Qianxi Liu, Ryo Suzuki 0001, Mingming Fan 0001, Zeyu Wang 0003
Proc. ACM Hum. Comput. Interact.3
2025 FocalSelect: Improving Occluded Objects Acquisition with Heuristic Selection and Disambiguation in Virtual Reality
abstract
In recent years, various head-worn virtual reality (VR) techniques have emerged to enhance object selection for occluded or distant targets. However, many approaches focus solely on ray-casting inputs, restricting their use with other input methods, such as bare hands. Additionally, some techniques speed up selection by changing the user's perspective or modifying the scene context, which may complicate interactions when users plan to resume or manipulate the scene afterward. To address these challenges, we present FocalSelect, a heuristic selection technique that builds 3D disambiguation through head-hand coordination and scoring-based functions. Our interaction design adheres to the principle that the intended selection range is a small sector of the headset's viewing frustum, allowing optimal targets to be identified within this scope. We also introduce a density-aware adjustable occlusion plane for effective depth culling of rendered objects. Two experiments are conducted to assess the adaptability of FocalSelect across different input modalities and its performance against five selection techniques. The results indicate that FocalSelect enhances selection experiences in occluded and remote scenarios while preserving the spatial context among objects. This preservation helps maintain users' understanding of the original scene and facilitates further manipulation. We also explore potential applications and enhancements to demonstrate more practical implementations of FocalSelect.
Duotun Wang, Linjie Qiu, Boyu Li 0007, Qianxi Liu, Xiaoying Wei, Jianhao Chen 0002, Zeyu Wang 0003, Mingming Fan 0001
IEEE Trans. Vis. Comput. Graph.1
2024 Toward Making Virtual Reality (VR) More Inclusive for Older Adults: Investigating Aging Effect on Target Selection and Manipulation Tasks in VR
abstract
Recent studies show the promise of VR in improving physical, cognitive, and emotional health of older adults. However, prior work on optimizing object selection and manipulation performance in VR was mostly conducted among younger adults. It remains unclear how older adults would perform such tasks compared to younger adults and the challenges they might face. To fill in this gap, we conducted two studies with both older and younger adults to understand their performances and user experiences of object selection and manipulation in VR respectively. Based on the results, we delineated interaction difficulties that older adults exhibited in VR and identified multiple factors, such as headset-related neck fatigue, extra head movements from out-of-view interactions, and slow spatial perceptions, that significantly decreased the motor performance of older adults. We further proposed design recommendations for improving the accessibility of direct interaction experiences in VR for older adults.
Zhiqing Wu, Duotun Wang, Shumeng Zhang, Yuru Huang, Zeyu Wang 0003, Mingming Fan 0001
CHI2
2024 SplattingAvatar: Realistic Real-Time Human Avatars With Mesh-Embedded Gaussian Splatting
abstract
We present SplattingAvatar, a hybrid 3D representation of photorealistic human avatars with Gaussian Splatting em-bedded on a triangle mesh, which renders over 300 FPS on a modern GPU and 30 FPS on a mobile device. We disentangle the motion and appearance of a virtual human with explicit mesh geometry and implicit appearance modeling with Gaus-sian Splatting. The Gaussians are defined by barycentric coordinates and displacement on a triangle mesh as Phong surfaces. We extend lifted optimization to simultaneously op-timize the parameters of the Gaussians while walking on the triangle mesh. SplattingAvatar is a hybrid representation of virtual humans where the mesh represents low-frequency motion and surface deformation, while the Gaussians take over the high-frequency geometry and detailed appearance. Un-like existing deformation methods that rely on an MLP-based linear blend skinning (LBS) field for motion, we control the rotation and translation of the Gaussians directly by mesh, which empowers its compatibility with various animation techniques, e.g., skeletal animation, blend shapes, and mesh editing. Trainable from monocular videos for both full-body and head avatars, SplattingAvatar shows state-of-the-art ren-dering quality across multiple datasets. Code and data are available at https://github.com/initialneil/SplattingAvatar.
Zhijing Shao, Duotun Wang, Xiangru Lin, Yu Zhang 0166, Mingming Fan 0001, Zeyu Wang 0003
CVPR4
2018 Spatially Perturbed Collision Sounds Attenuate Perceived Causality in 3D Launching Events
abstract
When a moving object collides with an object at rest, people immediately perceive a causal event: i.e., the first object has launched the second object forwards. However, when the second object's motion is delayed, or is accompanied by a collision sound, causal impressions attenuate and strengthen. Despite a rich literature on causal perception, researchers have exclusively utilized 2D visual displays to examine the launching effect. It remains unclear whether people are equally sensitive to the spatiotemporal properties of observed collisions in the real world. The present study first examined whether previous findings in causal perception with audiovisual inputs can be extended to immersive 3D virtual environments. We then investigated whether perceived causality is influenced by variations in the spatial position of an auditory collision indicator. We found that people are able to localize sound positions based on auditory inputs in VR environments, and spatial discrepancy between the estimated position of the collision sound and the visually observed impact location attenuates perceived causality.
Duotun Wang, James Kubricht, Yixin Zhu 0001, Wei Liang 0008, Song-Chun Zhu, Chenfanfu Jiang, Hongjing Lu
VR1