VLDB 2026 Research / reviewers in the wild / expert
Xuehuai Shi
dblp:213/5223
· DBLP profile ↗
23ranked-venue papers
7as first author
22since 2021 · last 2026
0000-0002-6671-0553ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 7 first-author · 20 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Constrained and directional ensemble attention for facial action unit detection
Zhiwen Shao, Bikuan Chen, Yong Zhou 0003, Xuehuai Shi, Canlin Li, Lizhuang Ma, Dit-Yan Yeung |
Pattern Recognit. | 4 |
| 2026 | TextRSR: Enhanced Arbitrary-Shaped Scene Text Representation via Robust Subspace RecoveryabstractIn recent years, scene text detection research has increasingly focused on arbitrary-shaped texts, where text representation is a fundamental problem. However, most existing methods still struggle to separate adjacent or overlapping texts due to ambiguous spatial positions of points or segmentation masks. Besides, the time efficiency of the entire pipeline is often neglected, resulting in sub-optimal inference speed. To tackle these problems, we first propose a novel text representation method based on robust subspace recovery, which robustly represents complex text shapes by combining orthogonal basis vectors learned from labeled text contours. These basis vectors capture basis contour patterns with distinct information, enabling clearer boundaries even in densely populated text scenarios. Moreover, we propose a dynamic sparse assignment scheme for positive samples that adaptively adjusts their weights during training, which not only accelerates inference speed by eliminating redundant predictions but also enhances feature learning by providing sufficient supervision signals. Building on these innovations, we present TextRSR, an accurate and efficient scene text detection network. Extensive experiments on challenging benchmarks demonstrate the superior accuracy and efficiency of TextRSR compared to state-of-the-art methods. Particularly, TextRSR achieves an F-measure of 88.5% at 37.8 frames per second (FPS) for CTW1500 dataset and an F-measure of 89.1% at 23.1 FPS for Total-Text dataset. Zhiwen Shao, Shengtian Jiang, Hancheng Zhu, Xuehuai Shi, Canlin Li, Lizhuang Ma, Dit-Yan Yeung |
IEEE Trans. Multim. | 4 |
| 2026 | MOA: Efficient Scene-Aware Multi-Object Arrangement in VRabstract3D multi-object arrangement is a fundamental task in VR that relies on accurate and natural initial selection alongside rapid and convenient subsequent manipulation to ensure high efficiency. However, existing methods fail to support efficient multi-object arrangement in highly occluded scenes with densely packed candidate objects through controller-free natural interactions. In this article, we propose an efficient, scene-aware multi-object arrangement method (MOA) designed for fast, precise, and convenient object arrangement. First, MOA introduces an importance-driven multi-object initial selection algorithm that assigns higher spatiotemporally correlated object importance (IMP) to target objects, establishing a natural multi-object initial selection mode that enables quick and accurate selection of high-IMP objects. Subsequently, it presents an auxiliary-structure-guided multi-object manipulation algorithm that constructs an auxiliary manipulation structure to assist subsequent multi-object manipulation, alongside a multi-modal interaction mode that facilitates swift and natural manipulation. Compared to state-of-the-art controller-free and controller-based methods, MOA significantly improves task performance, reduces task load, and enhances convenience in complex multi-object arrangement scenes involving hundreds of highly occluded objects need to be arranged. Xuehuai Shi, Yuhan Duan, Ziteng Wang 0002, Jian Wu 0033, Zhiwen Shao, Jieming Yin, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2026 | Automatic Formation Generation Based on Scene Awareness for Guided Group Navigation in VRabstractGroup navigation is a virtual reality (VR) technology designed to replace personal navigation, enhancing efficiency and aiding guided tours in path planning for other users. Group navigation techniques require users to have a clear understanding of navigation and travel modalities. Tour guides must ensure that users do not intersect with the environment or other users' models and must provide reasonable tour formations. To address these needs, we propose a scene-aware automatic formation generation method for fast and easy-to-use guided tours. First, we generate a limited number of candidate points for the exhibition and initialize the visiting group formation. Next, we optimize the formation to maximize view quality using our proposed viewpoint observation score. Finally, we match the visitors to the optimized formation to ensure a minimal view deflection angle. Moreover, we conducted a user study to evaluate the performance of our approach. Compared to the current method, our approach significantly increased navigation efficiency and view quality. Additionally, it substantially decreased the task load and improved the system usability for both tour guides and visitors. Jian Wu 0033, Lili Wang 0006, Zhikai Wen, Yanzhou Chen, Xuehuai Shi |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | FlexAcc: Accelerating Batch Normalization through GPU-FPGA IntegrationabstractConvolutional neural networks are fundamental to deep learning, especially in computer vision. However, their computational demands, particularly during batch normalization, create significant inefficiencies due to excessive data movement between memory and processing units. To address this, we propose FlexAcc, a novel architecture that integrates GPUs and FPGAs to offload BN computations to FPGAs, reducing data movement and improving hardware utilization. FlexAcc accelerates end-to-end training performance by up to 1.1× across various models. This approach bridges the performance gap between convolutional and non-convolutional layers, advancing deep learning model deployment. Haishuai Zhang, Pengyang Li, Xuehuai Shi, Xiaobai Chen, Jieming Yin |
ISCAS | 4 |
| 2025 | MMG: Manipulation-Aware Holistic Human Motion Generation from Sparse Tracking SignalsabstractGenerating realistic avatar motion via sparse tracking signals through VR devices is essential for enhancing the immersive user experience. Human-object manipulation behaviors not only affect hand motion but also significantly impact body motion. However, existing motion generation methods for human-object interactions overlook the coordinated coupling between body and hand motions during manipulations. Due to the diversity and complexity of holistic motion (body and hand motions simultaneously) in the latent motion space, generating physically plausible and temporally consistent holistic motion in real time, via the joint constraints imposed by sparse tracking signals and manipulation content, is a major challenge in the human motion generation task. We propose the manipulation-aware holistic human motion generation method (MMG) to help resolve this issue. In MMG, first, we construct a manipulation-aware holistic human motion generation framework that serially compresses the latent motion space distribution of the body and hand to generate realistic holistic human motion with object manipulation enabled. Second, to enhance the impact of object manipulation on holistic motion generation, MMG designs a novel object manipulation representation to extract effective manipulation features. Third, MMG is trained by an elaborate progressive manipulation-guided training algorithm to improve motion generation robustness and inference performance. Compared to state-of-the-art methods, MMG achieves up to a 39% improvement in the generated holistic motion quality with a 3.55 × speedup in generation performance. In manipulation-enabled scenes, MMG generates holistic motion in real time ($\geq 24 f p s$). Compared to the state-of-the-art methods, its perceived quality is significantly improved, and the task performance of holistic motion-required VR manipulation is high-significantly improved. This paper's code is at https://github.com/XRZ-BUAA/MMG. Xuehuai Shi, Renzhi Xiao, Yilun Sheng, Xiaobai Chen, Jieming Yin, Qingshan Liu 0001 |
ISMAR | 1 |
| 2025 | Mirror Detection via Multi-Directional Similarity Perception and Spectral Saliency EnhancementabstractMirror detection is a challenging task, due to the reflective properties of mirrors. Most existing approaches rely on exploiting the relationship between the content inside the mirror and the surrounding environment to aid in locating mirrors. A typical solution is to utilize contextual contrasted features. However, the discontinuity in content at the edges of mirrors may not always be prominent. To overcome this limitation, we propose a novel mirror detection framework called S2MD including two main modules, multi-directional similarity perception module (MSPM) and spectral saliency enhancement decoder module (SSEDM). Specifically, we employ a backbone network to extract multi-scale global information from images using a dual-path approach. Then, we feed these high-level dual-path features into MSPMs to generate direction-sensitive similarity-consistent features. MSPM utilizes active rotating filters and oriented response pooling to model the similarity relations in different orientations. Moreover, the SSEDM is utilized to enhance the spatial contextual contrasted features using feature spectral residuals and fuse the dual-path features to obtain the final predicted mirror mask. Extensive experiments demonstrate that our method achieves state-of-the-art performance on challenging MSD, PMD, and RGBD-Mirror benchmarks. The code is available at https://github.com/RuiChen-stack/M2SD. Zhiwen Shao, Xuehuai Shi, Bing Liu 0016, Canlin Li, Lizhuang Ma, Dit-Yan Yeung |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Micro-Expression Recognition via Fine-Grained Dynamic PerceptionabstractFacial micro-expression recognition (MER) is a challenging task, due to the transience, subtlety, and dynamics of micro-expressions (MEs). Most existing methods resort to hand-crafted features or deep networks, in which the former often additionally requires key frames, and the latter suffers from small-scale and low-diversity training data. In this article, we develop a novel fine-grained dynamic perception (FDP) framework for MER. We propose to rank frame-level features of a sequence of raw frames in chronological order, in which the rank process encodes the dynamic information of both ME appearances and motions. Specifically, a novel local-global feature-aware transformer is proposed for frame representation learning. A rank scorer is further adopted to calculate rank scores of each frame-level feature. Afterwards, the rank features from rank scorer are pooled in temporal dimension to capture dynamic representation. Finally, the dynamic representation is shared by a MER module and a dynamic image construction module, in which the former predicts the ME category, and the latter uses an encoder-decoder structure to construct the dynamic image. The design of dynamic image construction task is beneficial for capturing facial subtle actions associated with MEs and alleviating the data scarcity issue. Extensive experiments show that our method (i) significantly outperforms the state-of-the-art MER methods, and (ii) works well for dynamic image construction. Particularly, our FDP improves by 4.05%, 2.50%, 7.71%, and 2.11% over the previous best results in terms of F1-score on the CASME II, SAMM, CAS(ME) 2 , and CAS(ME) 3 datasets, respectively. The code is available at https://github.com/CYF-cuber/FDP . Zhiwen Shao, Xuehuai Shi, Canlin Li, Lizhuang Ma, Dit-Yan Yeung |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Fov-GS: Foveated 3D Gaussian Splatting for Dynamic ScenesabstractRendering quality and performance greatly affect the user's immersion in VR experiences. 3D Gaussian Splatting-based methods can achieve photo-realistic rendering with speeds of over 100 fps in static scenes, but the speed drops below 10 fps in monocular dynamic scenes. Foveated rendering provides a possible solution to accelerate rendering without compromising visual perceptual quality. However, 3DGS and foveated rendering are not compatible. In this paper, we propose Fov-GS, a foveated 3D Gaussian splatting method for rendering dynamic scenes in real time. We introduce a 3D Gaussian forest representation that represents the scene as a forest. To construct the 3D Gaussian forest, we propose a 3D Gaussian forest initialization method based on dynamic-static separation. Subsequently, we propose a 3D Gaussian forest optimization method based on deformation field and Gaussian decomposition to optimize the forest and deformation field. To achieve real-time dynamic scene rendering, we present a 3D Gaussian forest rendering method based on HVS models. Experiments demonstrate that our method not only achieves higher rendering quality in the foveal and salient regions compared to the SOTA methods but also dramatically improves rendering performance, achieving up to 11.33X speedup. We also conducted a user study, and the results prove that the perceptual quality of our method has a high visual similarity with the ground truth. Runze Fan, Jian Wu 0033, Xuehuai Shi, Lizhi Zhao, Qixiang Ma, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Audio-Visual Aware Foveated RenderingabstractWith the increasing complexity of geometry and rendering effects in virtual reality (VR) scenes, existing foveated rendering methods for VR head-mounted displays (HMDs) struggle to meet users' demands for VR scene rendering with high frame rates ($\geq 60fps$≥60fps for rendering binocular foveated images in VR scenes containing over 50 m triangles). Current research validates that auditory content affects the perception of the human visual system (HVS). However, existing foveated rendering methods primarily model the HVS's eccentricity-dependent visual perception ability on the visual content in VR while ignoring the impact of auditory content on the HVS's visual perception. In this article, we introduce an auditory-content-based perceived rendering quality analysis to quantify the impact of visual perception under different auditory conditions in foveated rendering. Based on the analysis results, we propose an audio-visual aware foveated rendering method (AvFR). AvFR first constructs an audio-visual feature-driven perception model that predicts the eccentricity-based visual perception in real time by combining the scene's audio-visual content, and then proposes a foveated rendering cost optimization algorithm to adaptively control the shading rate of different regions with the guidance of the perception model. In complex scenes with visual and auditory content containing over 1.17 m triangles, AvFR renders high-quality binocular foveated images at an average frame rate of 116$fps$fps. The results of the main user study and performance evaluation validate that AvFR achieves significant performance improvement (up to 1.4× speedup) without lowering the perceived visual quality compared with the state-of-the-art VR-HMD foveated rendering method. Xuehuai Shi, Jian Wu 0033, Jieming Yin, Xiaobai Chen, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | Scene-Aware Foveated Neural Radiance FieldsabstractFoveated rendering provides an idea for improving the image synthesis performance of neural radiance fields (NeRF) methods. In this article, we propose a scene-aware foveated neural radiance fields method to synthesize high-quality foveated images in complex VR scenes at high frame rates. First, we construct a multi-ellipsoidal neural representation to enhance the neural radiance field's representation capability in salient regions of complex VR scenes based on the scene content. Then, we introduce a uniform sampling based foveated neural radiance field framework to improve the foveated image synthesis performance with one-pass color inference, and improve the synthesis quality by leveraging the foveated scene-aware objective function. Our method synthesizes high-quality binocular foveated images at the average frame rate of 66 frames per second ($FPS$FPS) in complex scenes with high occlusion, intricate textures, and sophisticated geometries. Compared with the state-of-the-art foveated NeRF method, our method achieves significantly higher synthesis quality in both the foveal and peripheral regions with 1.41-1.46× speedup. We also conduct a user study to prove that the perceived quality of our method has a high visual similarity with the ground truth. Xuehuai Shi, Lili Wang 0006, Xinda Liu, Jian Wu 0033, Zhiwen Shao |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | PwP: Permutating with Probability for Efficient Group Selection in VRabstractGroup selection in virtual reality is an important means of multi-object selection, which allows users to quickly group multiple objects and can significantly improve the operation efficiency of multiple types of objects. In this paper, we propose a group selection method based on multiple rounds of probability permutation, in which the efficiency of group selection is substantially improved by making the object layout of the next round easier to be batch-selected through interactive selection, object grouping probability computation, and position rearrangement in each round of the selection process. We conducted ablation experiments to determine the algorithm coefficients and validate the effectiveness of the algorithm. In addition, an empirical user study was conducted to evaluate the ability of our method to significantly improve the efficiency of the group selection task in an immersive virtual reality environment. The reduced operations also indirectly reduce the user task load and improve usability. Jian Wu 0033, Weicheng Zhang, Handong Chen, Xuehuai Shi, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | ViP-Fluid: Visual Perception Driven Method for VR Fluid RenderingabstractThe demand for fluid simulation and rendering in virtual reality (VR) is increasing. However, achieving high visual quality while maintaining real-time efficiency remains a challenge. Traditional foveated rendering methods balance the simulation quality in the foveated region but neglect the physical realism in the peripheral areas, and fail to account for the perceptual degradation caused by frame rate fluctuations during adaptive updates. To address these challenges, we propose a novel visual perception driven fluid rendering method ViP-Fluid, which further enhances rendering quality while balancing efficiency. Our approach employs a spatiotemporal saliency model for multi-granularity simulation and rendering of Lagrangian fluid systems, and introduces a Perception Threshold for Physical Process Elapsing (PTPE) metric, which guides our temporal acceleration strategy. Through a series of objective experiments, we demonstrate the advantages of our method in rendering quality and performance efficiency. ViP-Fluid demonstrates superior metrics not only in the foveated region but also in the salient and overall regions, achieving up to 2.15 times speed-up compared to the high-resolution Position Based Fluids (PBF) benchmark. Subsequent user experiments further validate the visual perception advantages of ViP-Fluid over both traditional and state-of-the-art methods, confirming the spatiotemporal fidelity of our acceleration strategy as well as a user preference for our approach. Qixiang Ma, Jian Wu 0033, Runze Fan, Xuehuai Shi |
ISMAR | 5 |
| 2024 | Scene-aware Foveated RenderingabstractWe propose a new scene-aware foveated rendering method, which incorporates the scene awareness and characteristics of the human visual system into the mapping-based foveated rendering framework. First, we generate the conservative visual importance map that encodes the visual features of the scene, visual acuity, and gaze motion. Second, we construct the pixel size control map using a convolution kernel method. Third, we utilize the pixel size control map to guide the foveated rendering. At last, a temporal coherent refinement strategy is used to maintain the smooth foveated rendering for the adjacent frames. Compared to the state-of-the-art mapping-based foveated rendering methods using the same compression ratio, our method achieves smaller MSE, higher PSNR, and SSIM in the fovea, periphery, salient regions, and the whole image. We also conducted user studies, and the results proved that the perceptual quality of our method has a high visual similarity with the around truth rendered with the full resolution. Runze Fan, Xuehuai Shi, Kangyu Wang, Qixiang Ma, Lili Wang 0006 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | Locomotion-aware Foveated RenderingabstractOptimizing rendering performance improves the user's immersion in virtual scene exploration. Foveated rendering uses the features of the human visual system (HVS) to improve rendering performance without sacrificing perceptual visual quality. We collect and analyze the viewing motion of different locomotion methods, and describe the effects of these viewing motions on HVS's sensitivity, as well as the advantages of these effects that may bring to foveated rendering. Then we propose the locomotion-aware foveated rendering method (LaFR) to further accelerate foveated rendering by leveraging the advantages. In LaFR, we first introduce the framework of LaFR. Secondly, we propose an eccentricity-based shading rate controller that provides the shading rate control of the given region in foveated rendering. Thirdly, we propose a locomotion-aware log-polar mapping method, which controls the foveal average shading rate, the peripheral shading rate decrease speed, and the overall shading quantity with the locomotion-aware coefficients based on the eccentricity-based shading rate controller. LaFR achieves similar perceptual visual quality as the conventional foveated rendering while achieving up to 1.6× speedup. Compared with the full resolution rendering, LaFR achieves up to 3.8× speedup. Xuehuai Shi, Lili Wang 0006, Jian Wu 0033, Wei Ke 0001, Chan-Tong Lam |
VR | 1 |
| 2023 | Foveated rendering: A state-of-the-art surveyabstractRecently, virtual reality (VR) technology has been widely used in medical, military, manufacturing, entertainment, and other fields. These applications must simulate different complex material surfaces, various dynamic objects, and complex physical phenomena, increasing the complexity of VR scenes. Current computing devices cannot efficiently render these complex scenes in real time, and delayed rendering makes the content observed by the user inconsistent with the user’s interaction, causing discomfort. Foveated rendering is a promising technique that can accelerate rendering. It takes advantage of human eyes’ inherent features and renders different regions with different qualities without sacrificing perceived visual quality. Foveated rendering research has a history of 31 years and is mainly focused on solving the following three problems. The first is to apply perceptual models of the human visual system into foveated rendering. The second is to render the image with different qualities according to foveation principles. The third is to integrate foveated rendering into existing rendering paradigms to improve rendering performance. In this survey, we review foveated rendering research from 1990 to 2021. We first revisit the visual perceptual models related to foveated rendering. Subsequently, we propose a new foveated rendering taxonomy and then classify and review the research on this basis. Finally, we discuss potential opportunities and open questions in the foveated rendering field. We anticipate that this survey will provide new researchers with a high-level overview of the state-of-the-art in this field, furnish experts with up-to-date information, and offer ideas alongside a framework to VR display software and hardware designers and engineers. Lili Wang 0006, Xuehuai Shi |
Comput. Vis. Media | 2 |
| 2022 | Distant Object Manipulation with Adaptive Gains in Virtual RealityabstractObject Manipulation is a fundamental interaction in virtual reality (VR). The efficiency and accuracy of object manipulation are important to provide immersion to users. We propose a manipulation method with adaptive gains to improve the efficiency and accuracy of object manipulation in VR applications. First, we introduce manipulation gains. We then design an experiment to collect user behavior during manipulation to determine fitting functions for calculating manipulation gain. At last, we design a user study to evaluate the performance of our distant object manipulation method with adaptive gains. The results show that, compared with the state of the art methods, our method has a significant improvement in the completion time, and the manipulation accuracy of the tasks. Moreover, our method significantly increases usability and reduces task load. Lili Wang 0006, Shuai Luan, Xuehuai Shi, Xinda Liu |
ISMAR | 4 |
| 2022 | Label Guidance based Object Locating in Virtual RealityabstractObject locating in virtual reality (VR) has been widely used in many VR applications, such as virtual assembly, virtual repair, virtual remote coaching. However, when there are a large number of objects in the virtual environment(VE), the user cannot locate the target object efficiently and comfortably. In this paper, we propose a label guidance based object locating method for locating the target object efficiently in VR. Firstly, we introduce the label guidance based object locating pipeline to improve the efficiency of the object locating. It arranges the labels of all objects on the same screen, lets the user select the target labels first, and then uses the flying labels to guide the user to the target object. Then we summarize five principles for constructing the label layout for object locating and propose a two-level hierarchical sorted and orientated label layout based on the five principles for the user to select the candidate labels efficiently and comfortably. After that, we propose the view and gaze based label guidance method for guiding the user to locate the target object based on the selected candidate labels. It generates specific flying trajectories for candidate labels, updates the flying speed of candidate labels, keeps valid candidate labels, and removes the invalid candidate labels in real time during object locating with the guidance of the candidate labels. Compared with the traditional method, the user study results show that our method significantly improves efficiency and reduces task load for object locating. Xiaoheng Wei, Xuehuai Shi, Lili Wang 0006 |
ISMAR | 2 |
| 2022 | Interactive Mixed Reality Rendering on Holographic PyramidabstractCurrently, ray tracing and image-based lighting (IBL) have shortcomings when rendering the metallic virtual object displayed in the holographic pyramid in mixed reality. Ray tracing can hardly achieve the interactive frame rates, and IBL cannot accurately render the reflection result of the foreground near the virtual object. In this paper, we propose a mixed reality rendering method to render glossy and specular reflection effects on metallic virtual objects displayed in the holographic pyramid based on the surrounding real environment at interactive frame rates. First, we acquire the real environment data with four RGBD cameras and a panoramic camera; then, we introduce a foreground point cloud generation method to extract a temporally stable foreground point cloud from RGBD videos captured in real time; after this, we propose an efficient ray tracing method to render the dynamic glossy and specular reflections on the virtual objects that are displayed in the holographic pyramid. We test our method on several real and synthetic scenes. Compared with IBL and screen-space ray tracing, our method can generate the rendering results closer to the ground truth at the same time cost. Compared with Monte Carlo path tracing, our method is 2.5-4.5× faster in generating rendering results of the comparable quality. Danqing Dai, Xuehuai Shi, Lili Wang 0006 |
VR | 2 |
| 2022 | Efficient Flower Text Entry in Virtual RealityabstractText entry is a frequently used task in virtual reality (VR) applications, and controller is the most common interactive device in current VR systems. However, in terms of typing speed, there is still a gap between the existing controller-based text entry techniques and using a physical keyboard in reality, so it is important to improve the efficiency of the controller-based text entry. In this paper, we introduce Flower Text Entry, a single-controller text entry method based on a newly designed flower-shaped keyboard using hand 3D translation interaction for letters selection. We conduct user studies to optimize the keyboard design and the mapping between the interaction and selection, so as to evaluate our method. The results show that our method has high typing speed, lower error rate, and is very friendly to novices compared with the state-of-the-art controller-based text entry methods. After a short training, the novice group can type at 17.65 words per minute (WPM), and the potential expert group can type at 22.97 WPM. The highest typing speed is up to 30.80 WPM achieved by a potential expert participant. Jiaye Leng, Lili Wang 0006, Xuehuai Shi, Miao Wang 0004 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Foveated Stochastic LightcutsabstractFoveated rendering provides an idea for accelerating rendering algorithms without sacrificing the perceived rendering quality in virtual reality applications. In this paper, we propose a foveated stochastic lightcuts method to render high-quality many-lights illumination effects in high perception-sensitive regions. First, we introduce a spatiotemporal-luminance based lightcuts generation method to generate lightcuts with different accuracy for different visual perception-sensitive regions. Then we propose a multi-resolution light samples selection method to select the light sample for each node in the lightcuts more efficiently. Our method supports full-dynamic scenes containing over 250k dynamic light sources and dynamic diffuse/specular/glossy objects. It provides frame rates up to 110fps for high-quality many-lights illumination effects in high perception-sensitive regions of the HVS in VR HMDs. Compared with the state-of-the-art stochastic lightcuts method using the same rendering time, our method achieves smaller mean squared errors in the fovea and periphery. We also conduct user studies to prove that the perceived quality of our method has a high visual similarity with the results of the ground truth rendered by using the stochastic lightcuts with 2048 light samples per pixel. Xuehuai Shi, Lili Wang 0006, Jian Wu 0033, Runze Fan, Aimin Hao |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2021 | Foveated Photon MappingabstractVirtual reality (VR) applications require high-performance rendering algorithms to efficiently render 3D scenes on the VR head-mounted display, to provide users with an immersive and interactive virtual environment. Foveated rendering provides a solution to improve the performance of rendering algorithms by allocating computing resources to different regions based on the human visual acuity, and renders images of different qualities in different regions. Rasterization-based methods and ray tracing methods can be directly applied to foveated rendering, but rasterization-based methods are difficult to estimate global illumination (GI), and ray tracing methods are inefficient for rendering scenes that contain paths with low probability. Photon mapping is an efficient GI rendering method for scenes with different materials. However, since photon mapping cannot dynamically adjust the rendering quality of GI according to the human acuity, it cannot be directly applied to foveated rendering. In this paper, we propose a foveated photon mapping method to render realistic GI effects in the foveal region. We use the foveated photon tracing method to generate photons with high density in the foveal region, and these photons are used to render high-quality images in the foveal region. We further propose a temporal photon management to select and update the valid foveated photons of the previous frame for improving our method's performance. Our method can render diffuse, specular, glossy and transparent materials to achieve effects specifically related to GI, such as color bleeding, specular reflection, glossy reflection and caustics. Our method supports dynamic scenes and renders high-quality GI in the foveal region at interactive rates. Xuehuai Shi, Lili Wang 0006, Xiaoheng Wei, Lingqi Yan 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | Foveated Instant RadiosityabstractFoveated rendering distributes computational resources based on visual acuity, more in the foveal regions of our eyes and less in the periphery. The traditional rasterization method can be adapted into the foveated rendering framework in a quite straightforward way, but it's difficult for estimating global illumination. Instant Radiosity is an efficient global illumination method. It generates Virtual Point Lights (VPLs) on the surface of the virtual scenes from light sources and uses these VPLs to simulate light bounces. However, instant radiosity can not be adapted into the foveated rendering pipeline directly, and is too slow for virtual reality experience. What's more, instant radiosity does not consider temporal coherence, therefore it lacks temporal stability for dynamic scenes. In this paper, we propose a foveated rendering method for instant radiosity with more accurate global illumination effects in the foveal region and less accurate global illumination in the peripheral region. We define a foveated importance for each VPL, and use it to smartly distribute the VPLs to guarantee the rendering precision of the foveal region. Meanwhile, we propose a novel VPL reuse scheme, which updates only a small fraction of VPLs over frames, which ensures temporal coherence and improves time efficiency. Our method supports dynamic scenes and achieves high quality in the foveal regions at interactive frame rates. Lili Wang 0006, Xuehuai Shi, Lingqi Yan 0001 |
ISMAR | 3 |