Xubo Yang

dblp:04/3750 · DBLP profile ↗
← Back
78ranked-venue papers
4as first author
39since 2021 · last 2026
0000-0001-5378-4003ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 61 · 1 first-author · 31 since 2021Human-computer interaction and ubiquitous computing · 29 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Security and privacy · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Artificial intelligence for virtual reality: a review
Lili Wang 0006, Yebin Liu, Miao Wang 0004, Xubo Yang, Lan Xu 0003, Zhangyao Tan, Runze Fan, Hongwen Zhang 0001, Yijian Wen, Haozhong Yang, Jian Wu 0033, Jiahui Fan, Hui Wang 0045, Qixuan Zhang, Yongtian Wang, Qinping Zhao
Sci. China Inf. Sci.6
2026 DiffSurFlow: Efficient and Robust Differentiable Fluid Optimization via Surrogate Strategy on Flow Map
abstract
This paper presents a highly efficient and robust differentiable flu::id framework, centered on a novel surrogate gradient method that utilizes the flow map structural advantages. Our key insight reveals a significant misalignment between computational intensity and gradient importance during the backward pass. Specifically, we identify a physical duality within the adjoint process, revealing that the cross-step connections inherent in the flow map act as dominant gradient "highways" that propagate sensitivities over long horizons with high fidelity. Leveraging these insights, we develop a surrogate gradient model that retains these critical connections while pruning redundant adjoint computations in a physics-informed manner. Integrated with tailored acceleration techniques, our framework is successfully applied to diverse, challenging optimization tasks characterized by long time horizons and rich vorticity. Results demonstrate significant speedups and memory reductions while maintaining nearly-identical gradients compared to the full-gradient baseline.
Yuhao Quan, Hui Wang 0045, Weile Lian, Xubo Yang
ACM Trans. Graph.5
2026 PunctVR: VR Training for Image-Guided Needle Puncture with a Scaffolded, Self-Directed Framework
abstract
Image-guided percutaneous needle puncture is a critical yet challenging clinical procedure, constrained by the high cognitive demand of mental model construction and manipulation of 3D anatomy via scrolling through 2D cross-sectional images. While virtual reality (VR) simulators provide a risk-free training platform, many focus on simulation fidelity but lack structured, self-directed learning frameworks. In this paper, we present PunctVR, a VR system that incorporates the instructional principles of scaffolding. PunctVR features a training mode employing a phased subgoal workflow and instructional guidance scaffolding, and an assessment-only test mode where both the workflow and enhanced 3D visualization are removed. We conducted a between-subject experiment with 16 physicians, comparing training using a baseline multiplanar reconstruction (MPR) view with a combined MPR + 3D visualization across two difficulty levels. Our test mode results indicate that all trainees significantly improved their performance after training. Furthermore, those who trained with the integrated 3D visualization achieved a greater reduction in puncture time in both easy and hard cases. These findings suggest that PunctVR effectively enhances procedural efficiency in simulated needle puncture training and provides important insights into how learning scaffolding can accelerate skill acquisition and retention for image-guided interventions.
Wenqing Liu, Yan Zhang 0101, Hangyu Zhou, Zixuan Guo 0003, Aixi Guo, Ziang Qi, Jiannan Ye, Qishan Tong, Xubo Yang
IEEE Trans. Vis. Comput. Graph.12
2026 LIVE-GS: LLM Powers Interactive VR Experience with Physics-Aware Gaussian Splatting
abstract
As 3D Gaussian Splatting (3DGS) emerges as a leading approach for novel view synthesis and scene reconstruction, its potential in digital asset creation has gained significant attention. An increasing number of asset libraries based on GS are being established. However, generating physics-based dynamic assets remains a time-consuming and expertise-intensive task, especially for non-experts. In this paper, we propose LIVE-GS, a highly realistic Virtual Reality (VR) system powered by Large Language Models (LLMs), which enables rapid creation of dynamic Gaussian assets and real-time VR interactions. To inform our system design, we conducted interviews to examine challenges faced by current GS-based VR systems and the specific demands of users. Based on these insights, we employed GPT-4o to analyze key physical properties of objects that significantly impact user interactions, ensuring physics-based interactions in VR align with real-world phenomena. A key innovation of LIVE-GS is its ability to predict reasonable parameters in just 10 seconds from static Gaussian assets while maintaining high-quality VR interactions. To validate our approach, we invited participants experienced in physical simulation to manually adjust physical parameters, providing a baseline for comparison in both asset quality and authoring efficiency. We also conducted a comprehensive user study to evaluate system usability and user satisfaction. Experimental results demonstrate that LIVE-GS, leveraging LLMs' scene understanding capabilities, can achieve efficient physical scene creation and natural interactions without requiring manual design or annotation.
Haotian Mao, Hangyu Zhou, Zhuoxiong Xu, Siyue Wei, Yule Quan, Yan Zhang 0101, Zixuan Guo 0003, Nianchen Deng, Xubo Yang
IEEE Trans. Vis. Comput. Graph.9
2026 Consistent 3D Human Reconstruction From Monocular Video: Learning Correctable Appearance and Temporal Motion Priors
abstract
Recent advancements in rendering dynamic humans using NeRF and 3D Gaussian splatting have made significant progress, leveraging implicit geometry learning and image appearance rendering to create digital humans. However, in monocular video rendering, there are still challenges in rendering subtle and complex motion from different viewpoints and states, primarily due to the imbalance of viewpoints. Additionally, ensuring continuity between adjacent frames when rendering from novel and free viewpoints remains a difficult task. To address these challenges, we first propose a pixel-level motion correction module that adjusts the errors in the learned representation between different viewpoints. We also introduce a temporal information-based model to improve motion continuity by leveraging adjacent frames. Experimental results on dynamic human rendering, using the NeuMan, ZJU-Mocap, and People-Snapshot datasets, demonstrate that our method outperforms state-of-the-art techniques both quantitatively and qualitatively.
Cheng Shang, Liang An 0001, Jiajun Zhang 0012, Yuxiang Zhang 0006, Jidong Tian, Yebin Liu, Xubo Yang
IEEE Trans. Vis. Comput. Graph.8
2026 DanceAgent: Dance Movement Refinement With LLM Agent
abstract
Recent research on motion generation and text-to-motion synthesis focus on coarse-grained motion descriptions, neglecting fine-grained motion details and motion quality refinement. Additionally, current text-to-motion models, such as MotionGPT, lack multi-turn interaction capabilities, relying on single-turn and single-modality transformations, which limit their ability to integrate information from different modalities across interaction stages. These gaps leave critical questions, such as "How well is the motion performed" and "How can it be refined?" largely unaddressed. To address these issues, first, we introduce two fine-grained dance datasets-one focusing on jazz dance and the other on folk dance, which we have independently collected. Second, considering that dance motions are inherently complex and consist of long sequential actions, we introduce both global and local optimization during the motion encoding phase and employ Hidden Markov Model (HMM) temporal modeling to capture differential features between correct and incorrect movements, thereby optimizing the training process. Finally, we propose a multi-turn historical dialogue framework that enables three stages generation-motion assess, text instructions, and motion refinement-for input videos. This framework assists dance beginners by providing feedback on their movements, offering textual instructions, and delivering motion-based refinement. Experimental results on the jazz dance and folk dance datasets demonstrate that our method surpasses existing approaches in both quantitative and qualitative metrics, establishing a new benchmark for motion-text generation in the field of dance training.
Cheng Shang, Liang An 0001, Jiajun Zhang 0012, Yuxiang Zhang 0006, Yebin Liu, Xubo Yang
IEEE Trans. Vis. Comput. Graph.7
2026 SemanticAction: A Semantic-Driven and Behavior-Aware Password Framework for Adaptive VR Authentication Under Observation Attacks
abstract
As immersive Virtual Reality (VR) applications become increasingly widespread, ensuring secure and usable authentication is critical. Traditional knowledge-based methods (e.g., passwords, PINs) suffer from memorability issues and are highly vulnerable to observation attacks such as Man-in-the-Room (MITR). Meanwhile, biometric and behavioral approaches raise concerns regarding practicality, privacy, and cross-platform deployment. We present SemanticAction, a semantic-driven, behavior-augmented authentication framework that integrates knowledge-based passwords with gesture-based behavioral biometrics. Passwords are encoded as scene-anchored directions executed through intuitive hand gestures, enabling semantic meaning to guide user interactions. To counter observation attacks, SemanticAction employs randomized scene prompts and decoy scenes, while a dual-constraint verification mechanism adapts to both cold-start/few-shot conditions. Two user studies provide initial evidence that SemanticAction can mitigate MITR attacks while alleviating memorability challenges, maintaining favorable usability and security performance even with limited behavioral data. This work offers early insights and practical design considerations for behavior-aware authentication in VR.
Tingjie Wan, Yalin Deng, Zixuan Guo 0003, Xubo Yang, Boyu Gao 0003
IEEE Trans. Vis. Comput. Graph.6
2026 Temporal Foveated Fluid Animation in Virtual Reality
abstract
Simulating realistic fluids in virtual reality (VR) is computationally demanding, often limiting the scale and complexity of immersive environments. Existing foveated fluid simulation approaches primarily focus on spatial adaptivity. In this paper, we introduce a gaze-contingent fluid simulation system from a temporal perspective. We conduct a perceptual study to quantify the relationship between gaze eccentricity, fluid density deviation, and the perceptual threshold for simulation timesteps. Based on these findings, we fit a perceptual model that predicts the timestep requirements for maintaining perceptual realism in VR fluid animation. To exploit this model, we propose an asynchronous position-based fluids (PBF) algorithm that assigns fluid particles different local timesteps according to their visual importance and density deviation, ensuring both physical stability and perceptual validity. Our solver performs high-frequency updates in perceptually critical regions while progressively reducing updates elsewhere. A validation user study shows that our method remains perceptually indistinguishable from a high-fidelity, uniform-timestep PBF simulation. Objective evaluations further confirm that our approach improves efficiency while maintaining perceptual quality. Runtime experiments demonstrate speed-ups of up to 1.52× across diverse fluid scenarios, enabling more complex and larger-scale fluid phenomena in real-time VR. Our findings extend the paradigm of foveated fluid animation into the temporal domain, providing a perceptually grounded framework for fluid simulation in VR.
Yue Wang 0136, Yan Zhang 0101, Xubo Yang
IEEE Trans. Vis. Comput. Graph.5
2026 Mask Balancing: Perception-Driven Dynamic Visibility Enhancement for Occlusion-Capable Optical See-Through Head-Mounted Displays
abstract
The poor transparency of occlusion-capable optical see-through head-mounted displays (OC-OSTHMDs) deteriorates the visibility of the real scene, hindering the practical application of the devices. Previous works mitigate the issue by upgrading the transmittance of the spatial light modulator (SLM). However, the strategy soon reaches a limit because further optimization requires improving the transmittance of all optical elements, e.g., lenses and beam splitters. Moreover, pixelated occlusion usually relies on polarizing the real scene light, inevitably cutting the input optical power by half. To overcome this limitation, we propose a mask balancing method that improves real-scene brightness through polarization blending. Specifically, the s-polarized component, which passes through the optical system to provide occlusion-capable vision, is blended with the p-polarized component, which bypasses the system to preserve the raw view of the real scene. The blending is realized by simply modulating the cross-angle between a polarizing beam splitter and a linear polarizer, benefiting the robustness and versatility of the proposed method. We introduce a perception-driven blending approach, where the cross-angle is optimized in real-time to balance the visibility of the real scene and the texture and lighting of the virtual object. A benchtop prototype is built. A user study with 12 participants is conducted to quantify the visibility threshold of the texture and lighting of virtual objects. Then, a user study with 12 participants proves that the proposed method improves the visibility of the real scene while keeping a good appearance of the virtual object. We believe the proposed method is an important step toward developing practical solutions for OC-OSTHMDs.
Yan Zhang 0101, Rundong Chu, Qingtai Dong, Xiaodan Hu, Keyao You, Zixuan Guo 0003, Hangyu Zhou, Kiyoshi Kiyokawa, Xubo Yang
IEEE Trans. Vis. Comput. Graph.10
2026 AdaptiController: VR-Enhanced Fine Motor Assistance Through Finger Pressure Modulation
abstract
This paper explores finger pressure as a continuous implicit input modality to enhance interaction precision in virtual reality (VR). While motion controllers are widely adopted, their limitations in delicate operations remain a critical challenge. We investigate whether finger pressure signals from conventional VR controllers could offer advantages over traditional kinematic metrics for precision interaction.Through empirical studies, we demonstrate a robust relationship between pressure dynamics and task precision requirements, leading to a lightweight sigmoid-based model that leverages detected pressure to infer desired control granularity. In a comparative evaluation of video-scrubbing tasks, our adaptive method outperforms static sensitivity baselines in both task performance and subjective preference, without elevating cognitive load. Further validation via a VR sketching application demonstrates that our technique maintains task performance while reducing mental demand compared to manual control. Our findings reveal the untapped potential of pressure-based input to bridge coarse and fine-grained VR interactions, offering a path toward more versatile and intuitive input systems.
Hangyu Zhou, Haotian Mao, Zixuan Guo 0003, Yushi Wei, Yan Zhang 0101, Xubo Yang
IEEE Trans. Vis. Comput. Graph.6
2025 FaceCapGes: Real-Time Frame-by-Frame Gesture Generation from Audio, Facial Capture, and Head Pose
Jun Hanaizumi, Cheng Shang, Xubo Yang
CGI (3)3
2025 Dynamic Quadruple Optimization Based Transfer Learning for Animal Biometric Identification
abstract
With the progress of computer vision and machine learning, the research of object detection and pedestrian recognition has demonstrated significant performance. However, the identification studies in domestic animals, especially in the same species of domestic animals, remains a significant challenge. His study focuses on distinguishing cashmere and dairy goats, which share similar traits. Our contributions are: (1) Proposing a dynamic quadruple optimization algorithm to optimize goat images from local and global dimensions, enhancing network representation with a multi-branch structure; (2) Introducing a novel transfer learning algorithm based on goat granularity to preview dataset knowledge; (3) Validating our approach on our goat dataset and a public bird dataset. We achieved recognition accuracies of 95% for cashmere goats, 94.04% for dairy goats, and 82.48% on the public dataset, demonstrating the effectiveness of our methods for animal biometric identification.
Cheng Shang, Chong He, Xubo Yang, Yongliang Qiao, Meili Wang 0001
CSCWD5
2025 Color Correction for Occlusion-Capable Optical See-Through Head-Mounted Displays by Using Phase-Modulation
abstract
Occlusion-capable optical see-through head-mounted displays (OC-OSTHMDs) overcome the deficiency of semi-transparent virtual images by selectively cutting off light emitted from the physical background, considerably improving the graphics performance of augmented reality (AR). Existing OC-OSTHMDs achieve compact form factors by compressing the optical system based on the modulation of light polarization. However, the wavelength sensitivity of polarizing optical elements (POEs) causes color aberration in the see-through view. In this paper, we propose the spectrum- tuning method that mitigates color aberration of the see-through view caused by the wavelength sensitivity of OC-OSTHMDs. The methods operate on a spectrum-based color perception model that formulates the variation of the visible spectrum through OC-OSTHMDs. The optimization is performed globally, requiring minimal computation at runtime. A bench-top prototype of the OC-OSTHMD was built to validate these methods. Experimental results demonstrate that the spectrum-tuning method reduces the color difference by 18.1%. Additionally, the advantages of OC-OSTHMDs in presenting fluid animations in AR scenarios are demonstrated based on the prototype and a multi-buffer mask synthesis method.
Yan Zhang 0101, Shulin Hong, Weike Qian, Keyao You, Hangyu Zhou, Kiyoshi Kiyokawa, Xubo Yang
VR7
2025 Voice of artifacts: Evaluating user preferences for artifact voice in VR museums
Bingqing Chen, Wenqi Chu, Xubo Yang, Yue Li 0023
Comput. Graph.3
2025 Digital twin-based stress prediction for autonomous grasping of underwater robots with reinforcement learning
Xubo Yang, Jian Gao 0003, Yufeng Li 0006, Shengfa Wang, Jinglu Li
Expert Syst. Appl.1
2025 A Moving Least-Squares/Level-Set Particle Method for Bubble and Foam Simulation
abstract
We present a novel particle-grid scheme for simulating bubble and foam flow. At the core of our approach lies a particle representation that combines the computational nature of moving least-squares particles and particle level-set methods. Specifically, we assign a dedicated particle system to each individual bubble, enabling accurate tracking of its interface evolution and topological changes in a foaming fluid system. The particles within each bubble's particle system serve dual purposes. First, they function as a surface discretization, allowing for the solution of surfactant flow physics on the bubble's membrane. Additionally, these particles act as interface trackers, facilitating the evolution of the bubble's shape and topology within the multiphase fluid domain. The combination of particle systems from all bubbles contributes to the generation of an unsigned level-set field, further enhancing the simulation of coupled multiphase flow dynamics. By seamlessly integrating our particle representation into a multiphase, volumetric flow solver, our method enables the simulation of a broad range of intricate bubble and foam phenomena. These phenomena exhibit highly dynamic and complex structural evolution, as well as interfacial flow details.
Hui Wang 0045, Shulin Hong, Xubo Yang, Bo Zhu 0002
IEEE Trans. Vis. Comput. Graph.4
2025 Scene-Based Foveated Fluid Animation in Virtual Reality
abstract
Physically-based fluid animation in Virtual Reality (VR) significantly enhances the user experience through visually engaging flow motions. Nonetheless, such simulations are often limited by their substantial computational demands. A tailored adaptive simulation algorithm is important for high-performance VR fluid simulations, which dynamically allocate degrees of freedom (DoF) while accounting for user perception in VR. This paper proposes a novel scene-based gaze-contingent fluid simulation system for VR, featuring a highly adaptive fluid simulator integrated with a VR perceptual model that accounts for the foveation and geometry of fluid. Our method leverages an eccentricity and curvature-dependent perceptual model to dynamically allocate computational resources, improving the efficiency and maintaining spatio-temporal stability of fluid animation in VR. A user study was conducted to measure the simulation resolution thresholds for fluid animations in VR, considering various levels of eccentricity and curvature. Our findings indicate notable differences in perceptual thresholds based on these metrics. By incorporating these insights into our adaptive fluid simulator as a unified sizing function, we maintain perceptually optimal particle resolution, achieving up to a 3.62× performance improvement while delivering superior perceptual realism and user experience, as validated by a subjective evaluation study.
Yue Wang 0136, Yan Zhang 0101, Xuanhui Yang, Hui Wang 0045, Xubo Yang
IEEE Trans. Vis. Comput. Graph.5
2024 Neural Metameric Enhancement for Foveated Rendering
Jiannan Ye, Zhenkai Zhong, Xiaoxu Meng, Xubo Yang
CGI (2)4
2024 Free-view Rendering of Dynamic Human from Monocular Video Via Modeling Temporal Information Globally and Locally among Adjacent Frames
abstract
Recent research developments on rendering dynamic humans using neural radiance fields are remarkable. These methods often utilize learning implicit geometry and image appearance rendering for digital humans. However, keeping the complex and fast motions in detail, such as fingers, clothes, and faces, remains a challenge. Inspired by temporal information from human motion, we propose an architecture among adjacent frames by constructing a model on global and local levels. For the global level, we propose a hidden Markov model (HMM)based method to capture the global similarity among adjacent frames. At the local level, we introduce a module composed of a multi-head attention mechanism on a triplet canonical space structure for patch-level local temporal information. Experiments on two public datasets of dynamic human rendering (ZJU-MoCap and the People-Snapshot dataset) demonstrate that the proposed method outperforms advanced methods quantitatively and qualitatively.
Cheng Shang, Jidong Tian, Jiannan Ye, Xubo Yang
ICME4
2024 Visual Street Localization Refinement Method Using Differentiable Rendering
Jiannan Ye, Xiaoting Miao, Xubo Yang
ICXR3
2024 Foveated Fluid Animation in Virtual Reality
abstract
Large-scale fluid simulation is widely useful in various Virtual Reality (VR) applications. While physics-based fluid animation holds the promise of generating highly realistic fluid details, it often imposes significant computational demands, particularly when simulating high-resolution fluid for VR. In this paper, we propose a novel foveated fluid simulation method that enhances both the visual quality and computational efficiency of physics-based fluid simulation in VR. To leverage the natural foveation feature of human vision, we divide the visible domain of the fluid simulation into foveal, peripheral, and boundary regions. Our foveated fluid system dynamically allocates computational resources, striking a balance between simulation accuracy and computational efficiency. We implement this approach using a multi-scale method. To evaluate the effectiveness of our approach, we have conducted subjective studies. Our findings show a significant reduction in computational resource requirements, resulting in a speedup of up to 2.27 times. It is crucial to note that our method preserves the visual quality of fluid animations at a level that is perceptually identical to full-resolution outcomes. Additionally, we investigate the impact of various metrics, including particle radius and viewing distance, on the visual effects of fluid animations. Our work provides new techniques and evaluations tailored to facilitate real-time foveated fluid simulation in VR, which can enhance the efficiency and realism of fluids in VR applications.
Yue Wang 0136, Yan Zhang 0101, Xuanhui Yang, Hui Wang 0045, Xubo Yang
VR6
2024 Retinotopic Foveated Rendering
abstract
Foveated rendering (FR) improves the rendering performance of virtual reality (VR) by allocating fewer computational loads in the peripheral field of view (FOV). Existing FR techniques are built based on the radially symmetric regression model of human visual acuity. However, horizontal-vertical asymmetry (HVA) and vertical meridian asymmetry (VMA) in the cortical magnification factor (CMF) of the human visual system have been evidenced by retinotopy research of neuroscience, suggesting the radially asymmetric regression of visual acuity. In this paper, we begin with functional magnetic resonance imaging (fMRI) data, construct an anisotropic CMF model of the human visual system, and then introduce the first radially asymmetric regression model of the rendering precision for FR applications. We conducted a pilot experiment to adapt the proposed model to VR head-mounted displays (HMDs). A user study demonstrates that retinotopic foveated rendering (RFR) provides participants with perceptually equal image quality compared to typical FR methods while reducing fragments shading by 27.2% averagely, leading to the acceleration of 1/6 for graphics rendering. We anticipate that our study will enhance the rendering performance of VR by bridging the gap between retinotopy research in neuroscience and computer graphics in VR.
Yan Zhang 0101, Keyao You, Xiaodan Hu, Hangyu Zhou, Kiyoshi Kiyokawa, Xubo Yang
VR6
2024 Can manipulating control-display ratio dynamically really work in changing pseudo-haptic weight?
abstract
Abstract Due to the limited hardware available in virtual reality, such as VR controllers or simple motion capture devices, users will need more tactile feedback, like the weight of virtual objects. Pseudo‐tactile feedback, such as the control display (C/D) ratio manipulation method, is considered a standard method to simulate weight perception. In past studies, this method was basically used in a static environment, and the C/D ratio was usually determined initially. Can this method still work if the C/D ratio changes in dynamic usage scenarios? In a series of experiments, we tried to answer this question. We proved that the dynamic change of C/D ratio could simulate the weight change and improve the sense of embodiment through another hand redirection method.
Yonghua Chen, Xubo Yang
Comput. Animat. Virtual Worlds2
2024 Neural foveated super-resolution for real-time VR rendering
abstract
Abstract As virtual reality display technologies advance, resolutions and refresh rates continue to approach human perceptual limits, presenting a challenge for real‐time rendering algorithms. Neural super‐resolution is promising in reducing the computation cost and boosting the visual experience by scaling up low‐resolution renderings. However, the added workload of running neural networks cannot be neglected. In this article, we try to alleviate the burden by exploiting the foveated nature of the human visual system, in a way that we upscale the coarse input in a heterogeneous manner instead of uniform super‐resolution according to the visual acuity decreasing rapidly from the focal point to the periphery. With the help of dynamic and geometric information (i.e., pixel‐wise motion vectors, depth, and camera transformation) available inherently in the real‐time rendering content, we propose a neural accumulator to effectively aggregate the amortizedly rendered low‐resolution visual information from frame to frame recurrently. By leveraging a partition‐assemble scheme, we use a neural super‐resolution module to upsample the low‐resolution image tiles to different qualities according to their perceptual importance and reconstruct the final output adaptively. Perceptually high‐fidelity foveated high‐resolution frames are generated in real‐time, surpassing the quality of other foveated super‐resolution methods.
Jiannan Ye, Xiaoxu Meng, Daiyun Guo, Cheng Shang, Haotian Mao, Xubo Yang
Comput. Animat. Virtual Worlds6
2023 CCBA: Code Poisoning-Based Clean-Label Covert Backdoor Attack Against DNNs
Xubo Yang, Linsen Li 0002, Cunqing Hua, Changhao Yao
ICDF2C (1)1
2023 Neural Network Backdoor Attacks Fully Controlled by Composite Natural Utterance Fragments
Xubo Yang, Linsen Li 0002, Yenan Chen
ICICS1
2023 Redirected Placement: Evaluating the Redirection of Passive Props during Reach-to-Place in Virtual Reality
abstract
Hand redirection is an effective technique that can provide users with haptic feedback in virtual reality (VR) when a disparity exists between virtual objects and their physical counterparts. Psychophysiological research has revealed the distinct motion profiles of different kinematic phases when people operate hand-object interaction. In this paper, we proposed the Redirected Placement (RP), which determines the new placement of a physical prop using a constrained optimization problem. The visual illusion is used during the "reach-to-place" kinematic phase in the proposed RP method rather than the "reach-to-grasp" phase in the typical Redirected Reach (RR) method. We conducted two experiments based on the proposed RP method. Our first experiment showed that detection thresholds are generally higher with the proposed method compared to the RR method. The second experiment evaluated the embodiment experience with hand redirection using RR-only, RP-only, and RR&RP methods. The results report an enhanced sense of embodiment with the combined use of both RR and RP techniques. Our study further indicates that a 1:1 combination ratio of RR&RP resulted in the closest subjective experience to the baseline.
Xuanhui Yang, Yan Zhang 0101, Xubo Yang
VRST3
2023 Emotional Voice Puppetry
abstract
The paper presents emotional voice puppetry, an audio-based facial animation approach to portray characters with vivid emotional changes. The lips motion and the surrounding facial areas are controlled by the contents of the audio, and the facial dynamics are established by category of the emotion and the intensity. Our approach is exclusive because it takes account of perceptual validity and geometry instead of pure geometric processes. Another highlight of our approach is the generalizability to multiple characters. The findings showed that training new secondary characters when the rig parameters are categorized as eye, eyebrows, nose, mouth, and signature wrinkles is significant in achieving better generalization results compared to joint training. User studies demonstrate the effectiveness of our approach both qualitatively and quantitatively. Our approach can be applicable in AR/VR and 3DUI, namely, virtual reality avatars/self-avatars, teleconferencing and in-game dialogue.
Ruisi Zhang, Shengran Cheng, Shuai Tan 0002, Yu Ding 0001, Kenny Mitchell, Xubo Yang
IEEE Trans. Vis. Comput. Graph.7
2023 Add-on Occlusion: Turning Off-the-Shelf Optical See-through Head-mounted Displays Occlusion-capable
abstract
The occlusion-capable optical see-through head-mounted display (OC-OSTHMD) is actively developed in recent years since it allows mutual occlusion between virtual objects and the physical world to be correctly presented in augmented reality (AR). However, implementing occlusion with the special type of OSTHMDs prevents the appealing feature from the wide application. In this paper, a novel approach for realizing mutual occlusion for common OSTHMDs is proposed. A wearable device with per-pixel occlusion capability is designed. OSTHMD devices are upgraded to be occlusion-capable by attaching the device before optical combiners. A prototype with HoloLens 1 is built. The virtual display with mutual occlusion is demonstrated in real-time. A color correction algorithm is proposed to mitigate the color aberration caused by the occlusion device. Potential applications, including the texture replacement of real objects and the more realistic semi-transparent objects display, are demonstrated. The proposed system is expected to realize a universal implementation of mutual occlusion in AR.
Yan Zhang 0101, Xiaodan Hu, Kiyoshi Kiyokawa, Xubo Yang
IEEE Trans. Vis. Comput. Graph.4
2022 A Secure Approach for Human Computer Interaction Using Human Hand Action
abstract
Hand actions classification is an imperative field for acquiring smart functionality in modern electronic devices because hand actions classification offers interactive and innovative methods to communicate and interact. Therefore, we develop a novel architecture based on you only looking at coefficients (YOLACT), a real-time instance segmentation approach, and a temporal relation network (TRN) for hand actions understanding. In addition, our framework consists of a face recognition-based security network (FRB-SN) for user identification. We trained the YOLACT and the TRN models using the segmented version of the 20BN jester dataset composed of hand actions images and ground truths while the FRB-SN is trained using the VGGFace2 dataset. For testing, the YOLACT is used to segment the object from the given image sequence and then passed to the TRN-trained model to predict the corresponding action. Our experimental results showed that the accuracy and frame rate of the proposed framework are competitive.
Vachiraporn Ketsoi, Muhammad Raza, Haopeng Chen, Xubo Yang
SMC4
2022 Dta: An Integrative Approach For Human Action Understanding Based On Region Of Interest
abstract
Human action recognition (HAR) is a popular topic in developing a visual analysis system because of its tremendous potential in autonomous visual analysis. However, visual analysis is a sophisticated field in computer vision because an image sequence consists of various features that do not belong to a specific action. Therefore, we present a novel architecture approach for human action recognition and localization. We dubbed it DTA, an abbreviation of the detect, track, and analyze. It is inspired by yolov3, deep-sort, and 3D convolutional neural networks. Our framework is compact in analyzing human action, and the results showed that the proposed method outperforms previous state-of-the-art methods in various aspects. Moreover, the action recognition model is developed, trained, and tested using the ROI version of the KTH dataset. The experimental results showed the accuracy of the proposed model is superior compared to other traditional methods.
Muhammad Raza, Vachiraporn Ketsoi, Haopeng Chen, Xubo Yang
SMC4
2022 Rectangular Mapping-based Foveated Rendering
abstract
With the speedy increase of display resolution and the demand for interactive frame rate, rendering acceleration is becoming more critical for a wide range of virtual reality applications. Foveated rendering addresses this challenge by rendering with a non-uniform resolution for the display. Motivated by the non-linear optical lens equation, we present rectangular mapping-based foveated rendering (RMFR), a simple yet effective implementation of foveated rendering framework. RMFR supports varying level of foveation according to the eccentricity and the scene complexity. Compared with traditional foveated rendering methods, rectangular mapping-based foveated rendering provides a superior level of perceived visual quality while consuming minimal rendering cost.
Jiannan Ye, Anqi Xie, Susmija Jabbireddy, Yunchuan Li, Xubo Yang, Xiaoxu Meng
VR5
2022 SREFBN: Enhanced feature block network for single-image super-resolution
abstract
Abstract Deep learning has assisted the field of single‐image super‐resolution (SR) in achieving new heights. However, the task of restoring a high‐resolution (HR) image from a highly degraded low‐resolution (LR) image is sophisticated due to poor image restoration quality. A novel and effective lightweight SR method is presented as super‐resolution via an enhanced feature block network (SREFBN) that successfully reconstructs an HR image using a corresponding LR image with a purposed deep residual block. In addition, a novel shared parameters approach in the top‐down pathway among low‐level feature maps is introduced. The experimental results prove that SREFBN achieves remarkable performance. The presented framework requires lower computational cost and outperforms many state‐of‐the‐art methods. It is also highly adaptable with low‐end devices, requiring lower multiplication and adding operations. A trade‐off comparison between the number of parameters, execution time, and accuracies is given while also showing different variations of our approach to prove the effectiveness and reliability of the shared parameters. Most importantly, the results indicate that our framework has gained state‐of‐the‐art performance on larger scales 3 and 4. Code is available at https://github.com/curzii23/SREFBN .
Vachiraporn Ketsoi, Muhammad Raza, Haopeng Chen, Xubo Yang
IET Image Process.4
2022 FoV-NeRF: Foveated Neural Radiance Fields for Virtual Reality
abstract
Virtual Reality (VR) is becoming ubiquitous with the rise of consumer displays and commercial VR platforms. Such displays require low latency and high quality rendering of synthetic imagery with reduced compute overheads. Recent advances in neural rendering showed promise of unlocking new possibilities in 3D computer graphics via image-based representations of virtual or physical environments. Specifically, the neural radiance fields (NeRF) demonstrated that photo-realistic quality and continuous view changes of 3D scenes can be achieved without loss of view-dependent effects. While NeRF can significantly benefit rendering for VR applications, it faces unique challenges posed by high field-of-view, high resolution, and stereoscopic/egocentric viewing, typically causing low quality and high latency of the rendered images. In VR, this not only harms the interaction experience but may also cause sickness. To tackle these problems toward six-degrees-of-freedom, egocentric, and stereo NeRF in VR, we present the first gaze-contingent 3D neural representation and view synthesis method. We incorporate the human psychophysics of visual- and stereo-acuity into an egocentric neural representation of 3D scenery. We then jointly optimize the latency/performance and visual quality while mutually bridging human perception and neural scene synthesis to achieve perceptually high-quality immersive interaction. We conducted both objective analysis and subjective studies to evaluate the effectiveness of our approach. We find that our method significantly reduces latency (up to 99% time reduction compared with NeRF) without loss of high-fidelity rendering (perceptually identical to full-resolution ground truth). The presented approach may serve as the first step toward future VR/AR systems that capture, teleport, and visualize remote environments in real-time.
Nianchen Deng, Zhenyi He, Jiannan Ye, Budmonde Duinkharjav, Praneeth Chakravarthula, Xubo Yang, Qi Sun 0003
IEEE Trans. Vis. Comput. Graph.6
2021 Real-Time Fluid Simulation with Atmospheric Pressure Using Weak Air Particles
Tian Sang, Yitian Ma, Hui Wang 0045, Xubo Yang
CGI5
2021 Render-based factorization for additive light field display
abstract
Abstract Augmented Reality (AR) and Virtual Reality (VR) applications enable viewers to experience 3D graphics immersively. However, current hardware either fails to provide multiple viewpoints like projectors or cause conflicts between vergence and accommodation like consumable headsets. Prior researchers have designed multi‐view displays to solve the problem by enabling view‐dependent images and focus cues. Nevertheless, it requires minutes to calculate one frame, which is critical for real‐time AR/VR applications. In this paper, we propose a new render‐based factorization for additive light field display and further improve the performance by optimizing the initialization of layers. We next compare our work with the state of the art and the results show that our solution has competitive results while the calculation takes less than 20 ms for each frame.
Nianchen Deng, Zhenyi He, Xubo Yang
Comput. Animat. Virtual Worlds3
2021 Real-time simulation of violent boiling in concentrated sulfuric acid dilution
Tian Sang, Yitian Ma, Yuwei Xiao, Xubo Yang
Vis. Comput.7
2021 Towards Stereoscopic On-vehicle AR-HUD
Nianchen Deng, Jiannan Ye, Xubo Yang
Vis. Comput.4
2021 Data-driven simulation in fluids animation: A survey
abstract
The field of fluid simulation is developing rapidly, and data-driven methods provide many frameworks and techniques for fluid simulation. This paper presents a survey of data-driven methods used in fluid simulation in computer graphics in recent years. First, we provide a brief introduction of physicalbased fluid simulation methods based on their spatial discretization, including Lagrangian, Eulerian, and hybrid methods. The characteristics of these underlying structures and their inherent connection with datadriven methodologies are then analyzed. Subsequently, we review studies pertaining to a wide range of applications, including data-driven solvers, detail enhancement, animation synthesis, fluid control, and differentiable simulation. Finally, we discuss some related issues and potential directions in data-driven fluid simulation. We conclude that the fluid simulation combined with data-driven methods has some advantages, such as higher simulation efficiency, rich details and different pattern styles, compared with traditional methods under the same parameters. It can be seen that the data-driven fluid simulation is feasible and has broad prospects.
Yue Wang 0136, Hui Wang 0045, Xubo Yang
Virtual Real. Intell. Hardw.4
2020 Codimensional surface tension flow using moving-least-squares particles
abstract
We propose a new Eulerian-Lagrangian approach to simulate the various surface tension phenomena characterized by volume, thin sheets, thin filaments, and points using Moving-Least-Squares (MLS) particles. At the center of our approach is a meshless Lagrangian description of the different types of codimensional geometries and their transitions using an MLS approximation. In particular, we differentiate the codimension-1 and codimension-2 geometries on Lagrangian MLS particles to precisely describe the evolution of thin sheets and filaments, and we discretize the codimension-0 operators on a background Cartesian grid for efficient volumetric processing. Physical forces including surface tension and pressure across different codimensions are coupled in a monolithic manner by solving one single linear system to evolve the surface-tension driven Navier-Stokes system in a complex non-manifold space. The codimensional transitions are handled explicitly by tracking a codimension number stored on each particle, which replaces the tedious meshing operators in a conventional mesh-based approach. Using the proposed framework, we simulate a broad array of visually appealing surface tension phenomena, including the fluid chain, bell, polygon, catenoid, and dripping, to demonstrate the efficacy of our approach in capturing the complex fluid characteristics with mixed codimensions, in a robust, versatile, and connectivity-free manner.
Hui Wang 0045, Yongxu Jin, Anqi Luo, Xubo Yang, Bo Zhu 0002
ACM Trans. Graph.4
2020 An adaptive staggered-tilted grid for incompressible flow simulation
abstract
Enabling adaptivity on a uniform Cartesian grid is challenging due to its highly structured grid cells and axis-aligned grid lines. In this paper, we propose a new grid structure - the adaptive staggered-tilted (AST) grid - to conduct adaptive fluid simulations on a regular discretization. The key mechanics underpinning our new grid structure is to allow the emergence of a new set of tilted grid cells from the nodal positions on a background uniform grid. The original axis-aligned cells, in conjunction with the populated axis-tilted cells, jointly function as the geometric primitives to enable adaptivity on a regular spatial discretization. By controlling the states of the tilted cells both temporally and spatially, we can dynamically evolve the adaptive discretizations on an Eulerian domain. Our grid structure preserves almost all the computational merits of a uniform Cartesian grid, including the cache-coherent data layout, the easiness for parallelization, and the existence of high-performance numerical solvers. Further, our grid structure can be integrated into other adaptive grid structures, such as an Octree or a sparsely populated grid, to accommodate the T-junction-free hierarchy. We demonstrate the efficacy of our AST grid by showing examples of large-scale incompressible flow simulation in domains with irregular boundaries.
Yuwei Xiao, Szeyu Chan, Siqi Wang 0003, Bo Zhu 0002, Xubo Yang
ACM Trans. Graph.5
2020 A Novel CNN-Based Poisson Solver for Fluid Simulation
abstract
Solving a large-scale Poisson system is computationally expensive for most of the Eulerian fluid simulation applications. We propose a novel machine learning-based approach to accelerate this process. At the heart of our approach is a deep convolutional neural network (CNN), with the capability of predicting the solution (pressure) of a Poisson system given the discretization structure and the intermediate velocities as input. Our system consists of four main components, namely, a deep neural network to solve the large linear equations, a geometric structure to describe the spatial hierarchies of the input vector, a Principal Component Analysis (PCA) process to reduce the dimension of input in training, and a novel loss function to control the incompressibility constraint. We have demonstrated the efficacy of our approach by simulating a variety of high-resolution smoke and liquid phenomena. In particular, we have shown that our approach accelerates the projection step in a conventional Eulerian fluid simulator by two orders of magnitude. In addition, we have also demonstrated the generality of our approach by producing a diversity of animations deviating from the original datasets.
Xiangyun Xiao, Yanqing Zhou, Hui Wang 0045, Xubo Yang
IEEE Trans. Vis. Comput. Graph.4
2020 Dense feature pyramid network for cartoon dog parsing
Jerome Wan, Guillaume Mougeot, Xubo Yang
Vis. Comput.3
2020 Effects of Virtual-real fusion on immersion, presence, and learning performance in laboratory education
abstract
Virtual-reality (VR) fusion techniques have become increasingly popular in recent years, and several previous studies have applied them to laboratory education. However, without a basis for evaluating the effects of virtual-real fusion on VR in education, many developers have chosen to abandon this expensive and complex set of techniques. In this study, we experimentally investigate the effects of virtual-real fusion on immersion, presence, and learning performance. Each participant was randomly assigned to one of three conditions: a PC environment (PCE) operated by mouse; a VR environment (VRE) operated by controllers; or a VR environment running virtual-real fusion (VRVRFE), operated by real hands. The analysis of variance (ANOVA) and t-test results for presence and self-efficacy show significant differences between the PCE*VR-VRFE condition pair. Furthermore, the results show significant differences in the intrinsic value of learning performance for pairs PCE*VRVRFE and VRE*VR-VRFE, and a marginally significant difference was found for the immersion group. The results suggest that virtual-real fusion can offer improved immersion, presence, and selfefficacy compared to traditional PC environments, as well as a better intrinsic value of learning performance compared to both PC and VR environments. The results also suggest that virtual-real fusion offers a lower sense of presence compared to traditional VR environments.
Jingcheng Qian, Yancong Ma, Xubo Yang
Virtual Real. Intell. Hardw.4
2019 A CNN-based Flow Correction Method for Fast Preview
abstract
Abstract Eulerian‐based smoke simulations are sensitive to the initial parameters and grid resolutions. Due to the numerical dissipation on different levels of the grid and the nonlinearity of the governing equations, the differences in simulation resolutions will result in different results. This makes it challenging for artists to preview the animation results based on low‐resolution simulations. In this paper, we propose a learning‐based flow correction method for fast previewing based on low‐resolution smoke simulations. The main components of our approach lie in a deep convolutional neural network, a grid‐layer feature vector and a special loss function. We provide a novel matching model to represent the relationship between low‐resolution and high‐resolution smoke simulations and correct the overall shape of a low‐resolution simulation to closely follow the shape of a high‐resolution down‐sampled version. We introduce the grid‐layer concept to effectively represent the 3D fluid shape, which can also reduce the input and output dimensions. We design a special loss function for the fluid divergence‐free constraint in the neural network training process. We have demonstrated the efficacy and the generality of our approach by simulating a diversity of animations deviating from the original training set. In addition, we have integrated our approach into an existing fluid simulation framework to showcase its wide applications.
Xiangyun Xiao, Hui Wang 0045, Xubo Yang
Comput. Graph. Forum3
2018 A Calibration Method for On-Vehicle AR-HUD System Using Mixed Reality Glasses
abstract
Calibration is a key step for on-vehicle AR-HUD systems to ensure the augmented information to be correctly viewed by the driver. State-of-art calibration methods require setting up of spatial tracking devices or attaching markers on vehicles, which is time-consuming and error-prone. In this paper, we present a novel multi-viewpoints calibration method for AR-HUD using only a mixed reality glasses such as HoloLens. The full calibration process can be done in one minute and provides high precise calibration result, while no markers need to be attached on vehicle.
Nianchen Deng, Yanqing Zhou, Jiannan Ye, Xubo Yang
VR4
2018 Adaptive learning-based projection method for smoke simulation
abstract
Abstract Traditional Eulerian‐based fluid simulations require much time and computational resources to solve the projection step, especially the large linear system produced by the Poisson equation. In this paper, we propose an adaptive machine‐learning‐based projection method combining deep neural network and incremental learning technique. We provide two modes: Fast Mode and Normal Mode to solve the most time‐consuming projection step and deal with various simulation scenes largely different from the training data set. We introduce patch‐based feature vectors to represent the whole velocity field and a loss function to keep the divergence‐free constraint. We have demonstrated the acceleration and extrapolation ability of our method by testing various smoke scenes far different from our training data set.
Xiangyun Xiao, Xubo Yang
Comput. Animat. Virtual Worlds3
2017 Real-time high-quality surface rendering for large scale particle-based fluids
abstract
Particle-based methods like Smoothed Particle Hydrodynamics (SPH) are increasingly adopted for large scale fluid simulation in interactive computer graphics. However, surface rendering for such dynamic particle sets is challenging: current methods either produce coarse results, or time consuming. We introduce a novel approach to render high-quality fluid surface in screen space by an efficient combination of particle splatting, ray-casting and surface normal estimation techniques. We apply particle splatting to accelerate ray-casting process, and estimate surface normal using Principal Component Analysis (PCA). We adopt GPU technique to further accelerate our method. Our method can produce high-quality smooth surface while preserving thin and sharp details of large scale fluids. The computation and memory cost of our rendering step only depend on the image resolution. These advantages make our method very suitable for previewing or rendering hundreds of millions particles interactively. We demonstrate the efficiency and effectiveness of our method by rendering various fluid scenarios with different-sized particle sets.
Xiangyun Xiao, Xubo Yang
I3D3
2017 MagicToon: A 2D-to-3D creative cartoon modeling system with mobile AR
abstract
We present MagicToon, an interactive modeling system with mobile augmented reality (AR) that allows children to build 3D cartoon scenes creatively from their own 2D cartoon drawings on paper. Our system consists of two major components: an automatic 2D-to-3D cartoon model creator and an interactive model editor to construct more complicated AR scenes. The model creator can generate textured 3D cartoon models according to 2D drawings automatically and overlay them on the real world, bringing life to flat cartoon drawings. With our interactive model editor, the user can perform several optional operations on 3D models such as copying and animating in AR context through a touchscreen of a handheld device. The user can also author more complicated AR scenes by placing multiple registered drawings simultaneously. The results of our user study have shown that our system is easier to use compared with traditional sketch-based modeling systems and can give more play to children's innovations compared with AR coloring books.
Lele Feng, Xubo Yang, Shuangjiu Xiao
VR2
2017 A Schur Complement Preconditioner for Scalable Parallel Fluid Simulation
abstract
We present an algorithmically efficient and parallelized domain decomposition based approach to solving Poisson’s equation on irregular domains. Our technique employs the Schur complement method, which permits a high degree of parallel efficiency on multicore systems. We create a novel Schur complement preconditioner which achieves faster convergence, and requires less computation time and memory. This domain decomposition method allows us to apply different linear solvers for different regions of the flow. Subdomains with regular boundaries can be solved with an FFT-based Fast Poisson Solver. We can solve systems with 1,024 3 degrees of freedom, and demonstrate its use for the pressure projection step of incompressible liquid and gas simulations. The results demonstrate considerable speedup over preconditioned conjugate gradient methods commonly employed to solve such problems, including a multigrid preconditioned conjugate gradient method.
Jieyu Chu, Nafees Bin Zafar, Xubo Yang
ACM Trans. Graph.3
2016 Data-driven projection method in fluid simulation
abstract
Abstract Physically based fluid simulation requires much time in numerical calculation to solve Navier–Stokes equations. Especially in grid‐based fluid simulation, because of iterative computation, the projection step is much more time‐consuming than other steps. In this paper, we propose a novel data‐driven projection method using an artificial neural network to avoid iterative computation. Once the grid resolution is decided, our data‐driven method could obtain projection results in relatively constant time per grid cell, which is independent of scene complexity. Experimental results demonstrated that our data‐driven method drastically speeded up the computation in the projection step. With the growth of grid resolution, the speed‐up would increase strikingly. In addition, our method is still applicable in different fluid scenes with some alterations, when computational cost is more important than physical accuracy. Copyright © 2016 John Wiley & Sons, Ltd.
Xubo Yang, Xiangyun Xiao
Comput. Animat. Virtual Worlds2
2016 A Fast Iterated Orthogonal Projection Framework for Smoke Simulation
abstract
We present a fast iterated orthogonal projection (IOP) framework for smoke simulations. By modifying the IOP framework with a different means for convergence, our framework significantly reduces the number of iterations required to converge to the desired precision. Our new iteration framework adds a divergence redistributor component to IOP that can improve the impeded convergence logic of IOP. We tested Jacobi, GS and SOR as divergence redistributors and used the Multigrid scheme to generate a highly efficient Poisson solver. It provides a rapid convergence rate and requires less computation time. In all of our experiments, our method only requires 2-3 iterations to satisfy the convergence condition of 1e-5 and 5-7 iterations for 1e-10. Compared with the commonly used Incomplete Cholesky Preconditioned Conjugate Gradient(ICPCG) solver, our Poisson solver accelerates the overall speed to approximately 7- to 30-fold faster for grids ranging from 128(3) to 256(3). Our solver can accelerate more on larger grids because of the property that the iteration count required to satisfy the convergence condition is independent of the problem size. We use various experimental scenes and settings to demonstrate the efficiency of our method. In addition, we present a feasible method for both IOP and our fast IOP to support free surfaces.
Xubo Yang, Shuangcai Yang
IEEE Trans. Vis. Comput. Graph.2
2015 Position-based fluid control
abstract
We present a novel fluid control method that is capable of driving particle-based fluid simulation to match a rapidly changing target while keeping natural fluid-like motion. To achieve the desired behavior, we first generate control particles by sampling the target shape and then apply a non-linear constraint to each control particle, with its neighboring fluid particles keeping a constant fluid density within its influence region. This density constraint is highly in line with the incompressible nature of the fluid, which can drive the fluid to match the target shape in a natural way. In addition, to match a fast moving or deforming target, we add an adaptive spring for each fluid particle in the control region, connecting with its nearest control particle. The spring constraint takes effect only when the fluid particle is far from its corresponding control particle to avoid introducing artificial viscosity. Therefore, the fluid particles are well controlled even if the target shape changes rapidly. Furthermore, we integrate a velocity constraint to adjust the stiffness of the controlled fluid. All these three constraints are solved under position-based framework which enables our simulation fast, robust and well-suitable for interactive applications. We demonstrate the efficiency and effectiveness of our method in various scenarios in real time.
Xubo Yang, Ziqi Wu
I3D2
2014 Real-time Human Pose and Shape Estimation for Virtual Try-On Using a Single Commodity Depth Camera
abstract
We present a system that allows the user to virtually try on new clothes. It uses a single commodity depth camera to capture the user in 3D. Both the pose and the shape of the user are estimated with a novel real-time template-based approach that performs tracking and shape adaptation jointly. The result is then used to drive realistic cloth simulation, in which the synthesized clothes are overlayed on the input image. The main challenge is to handle missing data and pose ambiguities due to the monocular setup, which captures less than 50 percent of the full body. Our solution is to incorporate automatic shape adaptation and novel constraints in pose tracking. The effectiveness of our system is demonstrated with a number of examples.
Mao Ye 0005, Huamin Wang 0001, Nianchen Deng, Xubo Yang, Ruigang Yang
IEEE Trans. Vis. Comput. Graph.4
2013 General-purpose telepresence with head-worn optical see-through displays and projector-based lighting
abstract
In this paper we propose a general-purpose telepresence system design that can be adapted to a wide range of scenarios and present a framework for a proof-of-concept prototype. The prototype system allows users to see remote participants and their surroundings merged into the local environment through the use of an optical see-through head-worn display. Real-time 3D acquisition and head tracking allows the remote imagery to be seen from the correct point of view and with proper occlusion. A projector-based lighting control system permits the remote imagery to appear bright and opaque even in a lit room. Immersion can be adjusted across the VR continuum. Our approach relies only on commodity hardware; we also experiment with wider field of view custom displays.
Andrew Maimone, Xubo Yang, Nate Dierk, Andrei State, Mingsong Dou, Henry Fuchs
VR2
2013 A Novel Projection Technique with Detail Capture and Shape Correction for Smoke Simulation
abstract
Abstract Smoke simulation on a large grid is quite time consuming and most of the computation time is spent on the projection step. We present a novel projection method which produces quite similar visual results as those produced with the traditional projection method, but uses much less computation time. Our method includes two steps: detail‐capture and shape‐correction. The first step preserves most of the smoke details using an efficient DST (Discrete Sine Transformation) Poisson Solver with auxiliary boundary sweeping. The second step maintains the overall flow shape by solving a correcting Poisson equation on a coarse grid. Our algorithm is very fast and quite easy to implement. Experiments show that our projection is approximately 10–30 times faster than the traditional projection with PCG(Preconditioned Conjugate Gradient), while convincingly preserving both the flow details and the overall shape of the smoke.
Xiaoyue Wu, Xubo Yang
Comput. Graph. Forum2
2011 A multi-layer grid approach for fluid animation
Xubo Yang, ZhanXin Yang
Sci. China Inf. Sci.2
2010 Creating and Preserving Vortical Details in SPH Fluid
abstract
Abstract We present a new method to create and preserve the turbulent details generated around moving objects in SPH fluid. In our approach, a high‐resolution overlapping grid is bounded to each object and translates with the object. The turbulence formation is modeled by resolving the local flow around objects using a hybrid SPH‐FLIP method. Then these vortical details are carried on SPH particles flowing through the local region and preserved in the global field in a synthetic way. Our method provides a physically plausible way to model the turbulent details around both rigid and deformable objects in SPH fluid, and can efficiently produce animations of complex gaseous phenomena with rich visual details.
Bo Zhu 0002, Xubo Yang
Comput. Graph. Forum2
2009 Robust and Fast Keypoint Recognition Based on SE-FAST
abstract
In this paper, we present a key point recognition scheme, which consists of a novel feature detector and an efficient descriptor. Inspired by FAST (features from accelerated segment test), our feature detector is easy to compute and has high repeatability. Scale-invariance and optimized robustness are gained by extending traditional FAST to scale space.We combine this detector with an adapted version of SURF (speed up robust features) descriptor, providing the system with all means to do feature matching and object detection. Experimental evaluation and comparison with standard SURF using Hessian matrix-based detector are included in this paper, showing improvement in speed with comparable robustness.
Xueting Tan, Xubo Yang, Shuangjiu Xiao
DASC2
2009 Real-time horizon-based reflection occlusion
abstract
Reflection occlusion(RO) method can yield the self-occluded effect when using environment maps to produce a reflection scene. In most modern renderers the common approaches to achieve RO are the ray-trace method and the static reflection maps method. We propose screen-space reflection occlusion(SSRO) method to compute RO in screen-space like screen-space ambient occlusion in [Bavoil et al. 2008] and screen-space directional occlusion in [Ritschel et al. 2009]. Our different strategies for both calculating the occlusion value and sampling in the depth image enable SSRO to calculate RO in real-time and to deal with dynamic scenes, variational lighting environments and changing views.
Xubo Yang
SIGGRAPH ASIA Sketches2
2009 Physically-based fluid animation: A survey
Xubo Yang
Sci. China Ser. F Inf. Sci.2
2008 Efficient visibility projection on spherical polar coordinates for shadow rendering using geometry shader
abstract
Visibility precomputation has become essential for PRT-based rendering involving soft shadows. While graphic hardware rasterization has been increasingly used to compute visibility information for a mass of precomputation tasks, the linear feature of rasterization makes it unadaptive to sample visibility in a coordinates system, to which the Cartesian coordinates are non-linearly mapped. With the emergence of Geometry Shader in the framework of DirectX 10, this problem can be solved. In this paper, we present an efficient solution for visibility sampling on the spherical polar coordinates in the precomputation stage of PRT by means of the Geometry Shader within the DirectX 10 pipelines. In the mean time, a spherical reparameterization scheme is introduced for shadow rendering, under which spherical functions as lighting, BRDF and visibility are mapped onto the spherical polar coordinates.
Lingchun Li, Xubo Yang, Shuangjiu Xiao
ICME2
2008 Relighting with real incident light source
abstract
This paper presents an approach that can relight captured image with real incident light source. We first captured basis images of objects illuminated with a projector as the basis light source. We then established a pixel level mapping relationship between a real incident light source and the basis light source. Finally we developed a relighting algorithm that can simulate the relit effects with the real light source. The techniques can be applied to mixed reality.
Xubo Yang, Shuangjiu Xiao, Xiaodong Ding
ISMAR2
2008 A practical radiometric compensation method for projector-based augmentation
abstract
Radiometric compensation has made it possible for a projector to display on ordinary surface with colors and textures. For previous methods, itpsilas necessary to calibrate both the projector and camera at first. The calibration can be time-consuming and needs to be redone once the system settings change. We present a method that simplifies the calibration process. As a result, the system is more practicable for ad-hoc setups.
Xinli Chen, Xubo Yang, Shuangjiu Xiao
ISMAR2
2007 A Real-Time ProCam System for Interaction with Chinese Ink-and-Wash Cartoons
abstract
This poster describes our recently developed real-time projector-camera system for interaction with Chinese Ink-and-Wash Cartoons. We implement a real-time interactive water simulation under the acceleration of GPU, together with a sensing subsystem using computer vision techniques. Combined with Chinese stylized fish rendered in process, the system provides real-time interactions with traditional Chinese paintings. By stirring up the still water, fish and other essential elements of Chinese paintings, we hope to present new interaction techniques and more lively Chinese painting sceneries compared to those in traditional static settings.
Xubo Yang, Shuangjiu Xiao
CVPR3
2007 GPU-based rendering and animation for Chinese painting cartoon
abstract
This paper presents a real-time rendering system for generating a Chinese ink-and-wash cartoon. The objective is to free the animators from laboriously designing traditional Chinese painting appearance. The system constitutes a morphing animation framework and a rendering process. The whole rendering process is based on Graphic Process Unit (GPU), including interior shading, silhouette extracting and background shading. Moreover, the morphing framework is created to automatically generate Chinese painting cartoon from a set of surface mesh models. These techniques can be applied to real-time Chinese style entertainment application.
Manli Yuan, Xubo Yang, Shuangjiu Xiao
Graphics Interface2
2007 A Framework for Tangible User Interfaces within Projector-based Mixed Reality
abstract
This paper proposes a framework named TIPMR, for designing tangible user interfaces (TUIs) within projector-based mixed reality applications. The framework divides the target application into three parts: GUI-based application, TUI and an assistant. The assistant is employed as an adapter to translate between TUI operations and general GUI commands like mouse or keyboard events. This architecture makes it easier to focus on designing GUI-based applications and TUI separately. We built a tourist guidance system with two different tangible interaction modes based on TIPMR to demonstrate its usefulness and efficiency.
Xubo Yang, Shuangjiu Xiao
ISMAR2
2007 A Practical Framework for Virtual Viewing and Relighting
Qi Duan, Jianjun Yu, Xubo Yang, Shuangjiu Xiao
ICEC3
2007 Interactive Image Based Relighting with Physical Light Acquisition
Jianjun Yu, Xubo Yang, Shuangjiu Xiao
ICEC2
2007 The Role of 3-D Sound in Human Reaction and Performance in Augmented Reality Environments
abstract
Three-dimensional sound's effectiveness in virtual reality (VR) environments has been widely studied. However, due to the big differences between VR and augmented reality (AR) systems in registration, calibration, perceptual difference of immersiveness, navigation, and localization, it is important to develop new approaches to seamlessly register virtual 3-D sound in AR environments and conduct studies on 3-D sound's effectiveness in AR context. In this paper, we design two experimental AR environments to study the effectiveness of 3-D sound both quantitatively and qualitatively. Two different tracking methods are applied to retrieve the 3-D position of virtual sound sources in each experiment. We examine the impacts of 3-D sound on improving depth perception and shortening task completion time. We also investigate its impacts on immersive and realistic perception, different spatial objects identification, and subjective feeling of "human presence and collaboration". Our studies show that applying 3-D sound is an effective way to complement visual AR environments. It helps depth perception and task performance, and facilitates collaborations between users. Moreover, it enables a more realistic environment and more immersive feeling of being inside the AR environment by both visual and auditory means. In order to make full use of the intensity cues provided by 3-D sound, a process to scale the intensity difference of 3-D sound at different depths is designed to cater small AR environments. The user study results show that the scaled 3-D sound significantly increases the accuracy of depth judgments and shortens the searching task completion time. This method provides a necessary foundation for implementing 3-D sound in small AR environments. Our user study results also show that this process does not degrade the intuitiveness and realism of an augmented audio reality environment
Zhiying Zhou, Adrian David Cheok, Xubo Yang
IEEE Trans. Syst. Man Cybern. Part A4
2004 An experimental study on the role of software synthesized 3D sound in augmented reality environments
abstract
Investigation of augmented reality (AR) environments has become a popular research topic for engineers, computer and cognitive scientists. Although application oriented studies focused on audio AR environments have been published, little work has been done to vigorously study and evaluate the important research questions of the effectiveness of 3D sound in the AR context, and to what extent the addition of 3D sound would contribute to the AR experience. Thus, we have developed two AR environments and performed vigorous experiments with human subjects to study the effects of 3D sound in the AR context. The study concerns two scenarios. In the first scenario, one participant must use vision only and vision with 3D sound to judge the relative depth of augmented virtual objects. In the second scenario, two participants must co-operate to perform a joint task in a game-based AR environment. Hence, the goals of this study are (1) to access the impact of 3D sound on depth perception in a single-camera AR environment, (2) to study the impact of 3D sound on task performance and the feeling of ‘human presence and collaboration’, (3) to better understand the role of 3D sound in human-computer and human–human interactions, (4) to investigate if gender can affect the impact of 3D sound in AR environments. The outcomes of this research can have a useful impact on the development of audio AR systems which provide more immersive, realistic and entertaining experiences by introducing 3D sound. Our results suggest that 3D sound in AR environment significantly improves the accuracy of depth judgment and improves task performance. Our results also suggest that 3D sound contributes significantly to the feeling of ‘human presence and collaboration’ and helps the subjects to ‘identify spatial objects’.
Zhiying Zhou, Adrian David Cheok, Xubo Yang
Interact. Comput.3
2004 An experimental study on the role of 3D sound in augmented reality environment
abstract
Investigation of augmented reality (AR) environments has become a popular research topic for engineers, computer and cognitive scientists. Although application oriented studies focused on audio AR environments have been published, little work has been done to vigorously study and evaluate the important research questions of the effectiveness of three-dimensional (3D) sound in the AR context, and to what extent the addition of 3D sound would contribute to the AR experience. Thus, we have developed two AR environments and performed vigorous experiments with human subjects to study the effects of 3D sound in the AR context. The study concerns two scenarios. In the first scenario, one participant must use vision only and vision with 3D sound to judge the relative depth of augmented virtual objects. In the second scenario, two participants must cooperate to perform a joint task in a game-based AR environment. Hence, the goals of this study are (1) to access the impact of 3D sound on depth perception in a single-camera AR environment, (2) to study the impact of 3D sound on task performance and the feeling of ‘human presence and collaboration’, (3) to better understand the role of 3D sound in human–computer and human–human interactions, (4) to investigate if gender can affect the impact of 3D sound in AR environments. The outcomes of this research can have a useful impact on the development of audio AR systems, which provide more immersive, realistic and entertaining experiences by introducing 3D sound. Our results suggest that 3D sound in AR environment significantly improves the accuracy of depth judgment and improves task performance. Our results also suggest that 3D sound contributes significantly to the feeling of human presence and collaboration and helps the subjects to ‘identify spatial objects’.
Zhiying Zhou, Adrian David Cheok, Xubo Yang
Interact. Comput.3
2004 Human Pacman: a mobile, wide-area entertainment system based on physical, social, and ubiquitous computing
Adrian David Cheok, Kok Hwee Goh, Wei Liu 0009, Farzam Farbiz, Siew Wan Fong, Sze Lee Teo, Yu Li 0024, Xubo Yang
Pers. Ubiquitous Comput.8
2003 Human Pacman: A Mobile Entertainment System with Ubiquitous Computing and Tangible Interaction over a Wide Outdoor Area
Adrian David Cheok, Fong Siew Wan, Kok Hwee Goh, Xubo Yang, Wei Liu 0009, Farzam Farbiz, Yu Li 0024
Mobile HCI4
2002 Interactive Theatre Experience in Embodied + Wearable Mixed Reality Space
abstract
This paper presents an interactive theatre based on an embodied mixed reality space and wearable computers. Embodied computing mixed reality spaces integrate ubiquitous computing, tangible interaction and social computing within a mixed reality space, which enables intuitive interaction with physical world and virtual world. We believe it has potential advantages to support novel interactive theatre experiences. Therefore, we explored the novel interactive theatre experience supported in the embodied mixed reality space, and implemented live 3D characters to interact with user in such a system.
Adrian David Cheok, Xubo Yang, Simon Prince, Fong Siew Wan, Mark Billinghurst, Hirokazu Kato 0001
ISMAR3
2002 Interactive Theatre Experience in Embodied + Wearable Mixed Reality Space
Adrian David Cheok, Xubo Yang, Simon Prince, Fong Siew Wan, Mark Billinghurst, Hirokazu Kato 0001
ISMAR3
2002 Touch-Space: Mixed Reality Game Space Based on Ubiquitous, Tangible, and Social Computing
Adrian David Cheok, Xubo Yang, Zhiying Zhou, Mark Billinghurst, Hirokazu Kato 0001
Pers. Ubiquitous Comput.2
1997 GIVE: a general interactive visualization environment
abstract
We present a data flow based software platform called GIVE (General Interactive Visualization Environment). GIVE owns a data flow kernel, a visual programming interface and an extensible module library. It greatly enhances the efficiency of building visualization programs. GIVE has the following distinctive features: it introduces control nodes and procedure nodes into traditional data flow architecture, so it can support complex ViSC application: it is a well opened environment, supported by two efficient interactive tools-GIVE/ME (Module Extender) and GIVE/DE (Data Extender)-which enables users to extend modules and data types into the system freely and conveniently; and it has a module library composed of varieties of modules. We first introduce several concepts of the data flow architecture, then describe the designation of GIVE, and finally discuss the implementation of GIVE.
Xubo Yang, Wenli Cai, Jiaoying Shi
IV1