VLDB 2026 Research / reviewers in the wild / expert
Yan Zhang 0101
dblp:04/3348-101
· DBLP profile ↗
14ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-7549-4725ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PunctVR: VR Training for Image-Guided Needle Puncture with a Scaffolded, Self-Directed FrameworkabstractImage-guided percutaneous needle puncture is a critical yet challenging clinical procedure, constrained by the high cognitive demand of mental model construction and manipulation of 3D anatomy via scrolling through 2D cross-sectional images. While virtual reality (VR) simulators provide a risk-free training platform, many focus on simulation fidelity but lack structured, self-directed learning frameworks. In this paper, we present PunctVR, a VR system that incorporates the instructional principles of scaffolding. PunctVR features a training mode employing a phased subgoal workflow and instructional guidance scaffolding, and an assessment-only test mode where both the workflow and enhanced 3D visualization are removed. We conducted a between-subject experiment with 16 physicians, comparing training using a baseline multiplanar reconstruction (MPR) view with a combined MPR + 3D visualization across two difficulty levels. Our test mode results indicate that all trainees significantly improved their performance after training. Furthermore, those who trained with the integrated 3D visualization achieved a greater reduction in puncture time in both easy and hard cases. These findings suggest that PunctVR effectively enhances procedural efficiency in simulated needle puncture training and provides important insights into how learning scaffolding can accelerate skill acquisition and retention for image-guided interventions. Wenqing Liu, Yan Zhang 0101, Hangyu Zhou, Zixuan Guo 0003, Aixi Guo, Ziang Qi, Jiannan Ye, Qishan Tong, Xubo Yang |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | LIVE-GS: LLM Powers Interactive VR Experience with Physics-Aware Gaussian SplattingabstractAs 3D Gaussian Splatting (3DGS) emerges as a leading approach for novel view synthesis and scene reconstruction, its potential in digital asset creation has gained significant attention. An increasing number of asset libraries based on GS are being established. However, generating physics-based dynamic assets remains a time-consuming and expertise-intensive task, especially for non-experts. In this paper, we propose LIVE-GS, a highly realistic Virtual Reality (VR) system powered by Large Language Models (LLMs), which enables rapid creation of dynamic Gaussian assets and real-time VR interactions. To inform our system design, we conducted interviews to examine challenges faced by current GS-based VR systems and the specific demands of users. Based on these insights, we employed GPT-4o to analyze key physical properties of objects that significantly impact user interactions, ensuring physics-based interactions in VR align with real-world phenomena. A key innovation of LIVE-GS is its ability to predict reasonable parameters in just 10 seconds from static Gaussian assets while maintaining high-quality VR interactions. To validate our approach, we invited participants experienced in physical simulation to manually adjust physical parameters, providing a baseline for comparison in both asset quality and authoring efficiency. We also conducted a comprehensive user study to evaluate system usability and user satisfaction. Experimental results demonstrate that LIVE-GS, leveraging LLMs' scene understanding capabilities, can achieve efficient physical scene creation and natural interactions without requiring manual design or annotation. Haotian Mao, Hangyu Zhou, Zhuoxiong Xu, Siyue Wei, Yule Quan, Yan Zhang 0101, Zixuan Guo 0003, Nianchen Deng, Xubo Yang |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2026 | Temporal Foveated Fluid Animation in Virtual RealityabstractSimulating realistic fluids in virtual reality (VR) is computationally demanding, often limiting the scale and complexity of immersive environments. Existing foveated fluid simulation approaches primarily focus on spatial adaptivity. In this paper, we introduce a gaze-contingent fluid simulation system from a temporal perspective. We conduct a perceptual study to quantify the relationship between gaze eccentricity, fluid density deviation, and the perceptual threshold for simulation timesteps. Based on these findings, we fit a perceptual model that predicts the timestep requirements for maintaining perceptual realism in VR fluid animation. To exploit this model, we propose an asynchronous position-based fluids (PBF) algorithm that assigns fluid particles different local timesteps according to their visual importance and density deviation, ensuring both physical stability and perceptual validity. Our solver performs high-frequency updates in perceptually critical regions while progressively reducing updates elsewhere. A validation user study shows that our method remains perceptually indistinguishable from a high-fidelity, uniform-timestep PBF simulation. Objective evaluations further confirm that our approach improves efficiency while maintaining perceptual quality. Runtime experiments demonstrate speed-ups of up to 1.52× across diverse fluid scenarios, enabling more complex and larger-scale fluid phenomena in real-time VR. Our findings extend the paradigm of foveated fluid animation into the temporal domain, providing a perceptually grounded framework for fluid simulation in VR. Yue Wang 0136, Yan Zhang 0101, Xubo Yang |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2026 | Mask Balancing: Perception-Driven Dynamic Visibility Enhancement for Occlusion-Capable Optical See-Through Head-Mounted DisplaysabstractThe poor transparency of occlusion-capable optical see-through head-mounted displays (OC-OSTHMDs) deteriorates the visibility of the real scene, hindering the practical application of the devices. Previous works mitigate the issue by upgrading the transmittance of the spatial light modulator (SLM). However, the strategy soon reaches a limit because further optimization requires improving the transmittance of all optical elements, e.g., lenses and beam splitters. Moreover, pixelated occlusion usually relies on polarizing the real scene light, inevitably cutting the input optical power by half. To overcome this limitation, we propose a mask balancing method that improves real-scene brightness through polarization blending. Specifically, the s-polarized component, which passes through the optical system to provide occlusion-capable vision, is blended with the p-polarized component, which bypasses the system to preserve the raw view of the real scene. The blending is realized by simply modulating the cross-angle between a polarizing beam splitter and a linear polarizer, benefiting the robustness and versatility of the proposed method. We introduce a perception-driven blending approach, where the cross-angle is optimized in real-time to balance the visibility of the real scene and the texture and lighting of the virtual object. A benchtop prototype is built. A user study with 12 participants is conducted to quantify the visibility threshold of the texture and lighting of virtual objects. Then, a user study with 12 participants proves that the proposed method improves the visibility of the real scene while keeping a good appearance of the virtual object. We believe the proposed method is an important step toward developing practical solutions for OC-OSTHMDs. Yan Zhang 0101, Rundong Chu, Qingtai Dong, Xiaodan Hu, Keyao You, Zixuan Guo 0003, Hangyu Zhou, Kiyoshi Kiyokawa, Xubo Yang |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2026 | AdaptiController: VR-Enhanced Fine Motor Assistance Through Finger Pressure ModulationabstractThis paper explores finger pressure as a continuous implicit input modality to enhance interaction precision in virtual reality (VR). While motion controllers are widely adopted, their limitations in delicate operations remain a critical challenge. We investigate whether finger pressure signals from conventional VR controllers could offer advantages over traditional kinematic metrics for precision interaction.Through empirical studies, we demonstrate a robust relationship between pressure dynamics and task precision requirements, leading to a lightweight sigmoid-based model that leverages detected pressure to infer desired control granularity. In a comparative evaluation of video-scrubbing tasks, our adaptive method outperforms static sensitivity baselines in both task performance and subjective preference, without elevating cognitive load. Further validation via a VR sketching application demonstrates that our technique maintains task performance while reducing mental demand compared to manual control. Our findings reveal the untapped potential of pressure-based input to bridge coarse and fine-grained VR interactions, offering a path toward more versatile and intuitive input systems. Hangyu Zhou, Haotian Mao, Zixuan Guo 0003, Yushi Wei, Yan Zhang 0101, Xubo Yang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | X-Mask: Improving Soft-Edge Occlusion in Optical See-Through Displays with Cross-Shaped PinholesabstractPlacing a transparent liquid crystal display (LCD) into the light path is a simple approach to create occlusion-capable optical seethrough head-mounted displays (OST-HMDs) that suffers from defocused (soft-edge) occlusion where the mask leakage partially occludes surrounding content as well. Creating a focused (hard-edge) occlusion that does not suffer from mask leakage requires complicated, bulky optical setups. We present X-Mask, a pinhole-arraybased OST-HMD that creates a sharp occlusion mask without the need for a bulky setup requiring only two transparent LCD layers. By rendering a pinhole array on the layer closer to the user's eye, our system functions as a programmable aperture layer that extends the effective depth of field and improves the sharpness of the occlusion mask rendered on the second LCD layer. Utilizing a conventional circular pinhole would result in non-uniform brightness and contrast. By changing the pinhole shape to a cross enables nearoptimal retinal tiling with reduced overlaps and gaps. To accommodate pupil size variation, focus distance, and gaze direction, our system design allows for gaze-contingent adjustment of both LCD layers. We validate X-Mask in simulations and a physical prototype showing improved occlusion sharpness and visual uniformity. Xiaodan Hu, Christoph Ebner, Yan Zhang 0101, Kiyoshi Kiyokawa, Alexander Plopski |
ISMAR | 3 |
| 2025 | Color Correction for Occlusion-Capable Optical See-Through Head-Mounted Displays by Using Phase-ModulationabstractOcclusion-capable optical see-through head-mounted displays (OC-OSTHMDs) overcome the deficiency of semi-transparent virtual images by selectively cutting off light emitted from the physical background, considerably improving the graphics performance of augmented reality (AR). Existing OC-OSTHMDs achieve compact form factors by compressing the optical system based on the modulation of light polarization. However, the wavelength sensitivity of polarizing optical elements (POEs) causes color aberration in the see-through view. In this paper, we propose the spectrum- tuning method that mitigates color aberration of the see-through view caused by the wavelength sensitivity of OC-OSTHMDs. The methods operate on a spectrum-based color perception model that formulates the variation of the visible spectrum through OC-OSTHMDs. The optimization is performed globally, requiring minimal computation at runtime. A bench-top prototype of the OC-OSTHMD was built to validate these methods. Experimental results demonstrate that the spectrum-tuning method reduces the color difference by 18.1%. Additionally, the advantages of OC-OSTHMDs in presenting fluid animations in AR scenarios are demonstrated based on the prototype and a multi-buffer mask synthesis method. Yan Zhang 0101, Shulin Hong, Weike Qian, Keyao You, Hangyu Zhou, Kiyoshi Kiyokawa, Xubo Yang |
VR | 1 |
| 2025 | Perception-Driven Soft-Edge Occlusion for Optical See-Through Head-Mounted DisplaysabstractSystems with occlusion capabilities, such as those used in vision augmentation, image processing, and optical see-through head-mounted display (OST-HMD), have gained popularity. Achieving precise (hard-edge) occlusion in these systems is challenging, often requiring complex optical designs and bulky volumes. On the other hand, utilizing a single transparent liquid crystal display (LCD) is a simple approach to create occlusion masks. However, the generated mask will appear defocused (soft-edge) resulting in insufficient blocking or occlusion leakage. In our work, we delve into the perception of soft-edge occlusion by the human visual system and present a preference-based optimal expansion method that minimizes perceived occlusion leakage. In a user study involving 20 participants, we made a noteworthy observation that the human eye perceives a sharper edge blur of the occlusion mask when individuals see through it and gaze at a far distance, in contrast to the camera system's observation. Moreover, our study revealed significant individual differences in the perception of soft-edge masks in human vision when focusing. These differences may lead to varying degrees of demand for mask size among individuals. Our evaluation demonstrates that our method successfully accounts for individual differences and achieves optimal masking effects at arbitrary distances and pupil sizes. Xiaodan Hu, Yan Zhang 0101, Alexander Plopski, Yuta Itoh 0001, Monica Perusquía-Hernández, Naoya Isoyama, Hideaki Uchiyama, Kiyoshi Kiyokawa |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Scene-Based Foveated Fluid Animation in Virtual RealityabstractPhysically-based fluid animation in Virtual Reality (VR) significantly enhances the user experience through visually engaging flow motions. Nonetheless, such simulations are often limited by their substantial computational demands. A tailored adaptive simulation algorithm is important for high-performance VR fluid simulations, which dynamically allocate degrees of freedom (DoF) while accounting for user perception in VR. This paper proposes a novel scene-based gaze-contingent fluid simulation system for VR, featuring a highly adaptive fluid simulator integrated with a VR perceptual model that accounts for the foveation and geometry of fluid. Our method leverages an eccentricity and curvature-dependent perceptual model to dynamically allocate computational resources, improving the efficiency and maintaining spatio-temporal stability of fluid animation in VR. A user study was conducted to measure the simulation resolution thresholds for fluid animations in VR, considering various levels of eccentricity and curvature. Our findings indicate notable differences in perceptual thresholds based on these metrics. By incorporating these insights into our adaptive fluid simulator as a unified sizing function, we maintain perceptually optimal particle resolution, achieving up to a 3.62× performance improvement while delivering superior perceptual realism and user experience, as validated by a subjective evaluation study. Yue Wang 0136, Yan Zhang 0101, Xuanhui Yang, Hui Wang 0045, Xubo Yang |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Foveated Fluid Animation in Virtual RealityabstractLarge-scale fluid simulation is widely useful in various Virtual Reality (VR) applications. While physics-based fluid animation holds the promise of generating highly realistic fluid details, it often imposes significant computational demands, particularly when simulating high-resolution fluid for VR. In this paper, we propose a novel foveated fluid simulation method that enhances both the visual quality and computational efficiency of physics-based fluid simulation in VR. To leverage the natural foveation feature of human vision, we divide the visible domain of the fluid simulation into foveal, peripheral, and boundary regions. Our foveated fluid system dynamically allocates computational resources, striking a balance between simulation accuracy and computational efficiency. We implement this approach using a multi-scale method. To evaluate the effectiveness of our approach, we have conducted subjective studies. Our findings show a significant reduction in computational resource requirements, resulting in a speedup of up to 2.27 times. It is crucial to note that our method preserves the visual quality of fluid animations at a level that is perceptually identical to full-resolution outcomes. Additionally, we investigate the impact of various metrics, including particle radius and viewing distance, on the visual effects of fluid animations. Our work provides new techniques and evaluations tailored to facilitate real-time foveated fluid simulation in VR, which can enhance the efficiency and realism of fluids in VR applications. Yue Wang 0136, Yan Zhang 0101, Xuanhui Yang, Hui Wang 0045, Xubo Yang |
VR | 2 |
| 2024 | Retinotopic Foveated RenderingabstractFoveated rendering (FR) improves the rendering performance of virtual reality (VR) by allocating fewer computational loads in the peripheral field of view (FOV). Existing FR techniques are built based on the radially symmetric regression model of human visual acuity. However, horizontal-vertical asymmetry (HVA) and vertical meridian asymmetry (VMA) in the cortical magnification factor (CMF) of the human visual system have been evidenced by retinotopy research of neuroscience, suggesting the radially asymmetric regression of visual acuity. In this paper, we begin with functional magnetic resonance imaging (fMRI) data, construct an anisotropic CMF model of the human visual system, and then introduce the first radially asymmetric regression model of the rendering precision for FR applications. We conducted a pilot experiment to adapt the proposed model to VR head-mounted displays (HMDs). A user study demonstrates that retinotopic foveated rendering (RFR) provides participants with perceptually equal image quality compared to typical FR methods while reducing fragments shading by 27.2% averagely, leading to the acceleration of 1/6 for graphics rendering. We anticipate that our study will enhance the rendering performance of VR by bridging the gap between retinotopy research in neuroscience and computer graphics in VR. Yan Zhang 0101, Keyao You, Xiaodan Hu, Hangyu Zhou, Kiyoshi Kiyokawa, Xubo Yang |
VR | 1 |
| 2023 | Redirected Placement: Evaluating the Redirection of Passive Props during Reach-to-Place in Virtual RealityabstractHand redirection is an effective technique that can provide users with haptic feedback in virtual reality (VR) when a disparity exists between virtual objects and their physical counterparts. Psychophysiological research has revealed the distinct motion profiles of different kinematic phases when people operate hand-object interaction. In this paper, we proposed the Redirected Placement (RP), which determines the new placement of a physical prop using a constrained optimization problem. The visual illusion is used during the "reach-to-place" kinematic phase in the proposed RP method rather than the "reach-to-grasp" phase in the typical Redirected Reach (RR) method. We conducted two experiments based on the proposed RP method. Our first experiment showed that detection thresholds are generally higher with the proposed method compared to the RR method. The second experiment evaluated the embodiment experience with hand redirection using RR-only, RP-only, and RR&RP methods. The results report an enhanced sense of embodiment with the combined use of both RR and RP techniques. Our study further indicates that a 1:1 combination ratio of RR&RP resulted in the closest subjective experience to the baseline. Xuanhui Yang, Yan Zhang 0101, Xubo Yang |
VRST | 2 |
| 2023 | Add-on Occlusion: Turning Off-the-Shelf Optical See-through Head-mounted Displays Occlusion-capableabstractThe occlusion-capable optical see-through head-mounted display (OC-OSTHMD) is actively developed in recent years since it allows mutual occlusion between virtual objects and the physical world to be correctly presented in augmented reality (AR). However, implementing occlusion with the special type of OSTHMDs prevents the appealing feature from the wide application. In this paper, a novel approach for realizing mutual occlusion for common OSTHMDs is proposed. A wearable device with per-pixel occlusion capability is designed. OSTHMD devices are upgraded to be occlusion-capable by attaching the device before optical combiners. A prototype with HoloLens 1 is built. The virtual display with mutual occlusion is demonstrated in real-time. A color correction algorithm is proposed to mitigate the color aberration caused by the occlusion device. Potential applications, including the texture replacement of real objects and the more realistic semi-transparent objects display, are demonstrated. The proposed system is expected to realize a universal implementation of mutual occlusion in AR. Yan Zhang 0101, Xiaodan Hu, Kiyoshi Kiyokawa, Xubo Yang |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | Super Wide-view Optical See-through Head Mounted Displays with Per-pixel Occlusion CapabilityabstractAugmented reality (AR) has been widely used that combines human with the digital world tightly at an unprecedented level, and various types of optical see-through head-mounted displays (OSTHMDs) have been actively developed to meet the requirement of everyday AR use. Correct mutual occlusion between real and virtual objects is often necessary for displaying realistic virtual images for users in the AR application scenarios. Some optical designs have been proposed to realize mutual occlusion by means of an OSTHMD in recent years. However, all of them support a limited field of view (FOV) that is much narrower than that of a natural human. The main limiting factor of the FOV of general OSTHMDs is the limited numerical aperture (NA) of the lenses. To address the problem, we propose an OSTHMD based on the double ellipsoidal mirror structure to avoid a stack of lenses and to achieve a wide FOV close to that of the naked eye with small distortion. A pair of imaging lenses are carefully arranged, then assembled between the two ellipsoidal mirrors with a pinhole mask to improve image quality. An experiment using a monocular prototype shows that a sharp see through view is achieved which remains clear regardless of the focus of the eye-simulating camera. An image undistortion algorithm is also developed to obtain a full-view display, enabling a virtual image to be displayed spanning a super wide FOV of H160°×V74°. Finally, per-pixel mutual occlusion of a wide FOV of H122°×V74° is realized by placing a spatial light modulator (SLM) in front of the entrance pupil. Yan Zhang 0101, Naoya Isoyama, Nobuchika Sakata, Kiyoshi Kiyokawa, Hong Hua |
ISMAR | 1 |