EDBT 2026 Demo / reviewers in the wild / expert
Yang Gao 0032
dblp:89/4402-32
· DBLP profile ↗
38ranked-venue papers
12as first author
31since 2021 · last 2026
0000-0002-9149-3554ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 11 first-author · 28 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What Makes a Virtual Celebrity Agent Trustworthy in VR? Exploring the Role of Stylization and VoiceabstractRecent advances in generative AI have accelerated the deployment of virtual celebrity agents for commercial endorsements. However, little is known about how their visual style and vocal type, especially when these characteristics are generated via advanced AI reconstruction or voice cloning techniques, affect user trust, perceived realism, familiarity, and social presence. We extracted a representative clip of Sheldon Cooper from The Big Bang Theory as a baseline and generated multiple Sheldon virtual agents that varied in visual style (hand-sculpted, AI-reconstructed, or non-stylized) and vocal type (cloned or synthesized). A 3 × 2 within-subjects experiment (N = 30) revealed that visual style significantly affected user trust, perceived realism, familiarity, and social presence, while vocal type affected only perceived realism and familiarity. Comparisons between the virtual agents and the video baseline confirmed that even cutting-edge 3D modeling still differs significantly from authentic video representations. Behavioral data indicate that the relationship between interpersonal-distance variations and trust levels is not a simple linear one. These findings provide actionable guidance for designers leveraging generative AI to create trustworthy virtual avatars or agents and delineate new research avenues for virtual celebrity agents. Yang Gao 0032, Yangbin Dai, Guangtao Zhang, Fariba Mostajeran, Frank Steinicke, Lin Li 0062, Tao Yu 0007 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2026 | Deep-Saliency Foveated Ray Tracing For Real-time VR RenderingabstractImmersive VR applications demand high resolutions and refresh rates, posing significant challenges for real-time rendering. Foveated rendering mitigates this cost by exploiting properties of the Human Visual System (HVS), but conventional approaches often rely on oversimplified heuristic models that neglect high-level attentional cues, resulting in artifacts in peripheral regions. To this end, we present a neural saliency-driven foveated ray tracing framework that overcomes these limitations. Our method introduces a motion-aware foveation model to capture temporal dynamics and employs a lightweight convolutional neural network to predict saliency maps that reflect complex attentional patterns derived from eye-gaze data. The combination of these guides adaptive path tracing and filtering, enabling perceptually optimized rendering with minimal artifacts. Experimental results show that our approach improves perceptual quality over prior methods while sustaining real-time performance. Yang Gao 0032, Wencan Li, Shiyu Liang, Weizichuan Feng, Qing Xia 0002, Shuai Li 0001, Aimin Hao |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2026 | Generating Audiovisual Synergy Fluid Animation for Highly Immersive VR ExperienceabstractGenerative content is increasingly applied in VR to provide immersive experiences, yet maintaining high generation quality remains challenging for audiovisual effects. Particularly in dynamic fluid phenomena, achieving realism and presence requires adherence to physical laws. To accomplish this objective, this work proposes an audiovisual synergy fluid animation generation framework, which enhances immersion by improving motion texture fidelity and audiovisual consistency. It comprises Detail-Enhanced Texture generator (DET) and Physics-Guided Audio generator (PGA). DET integrates Global-Local Physics guidance (GLP) and Temporal Texture Modeling (TTM) to produce video textures, explicitly optimizing dynamic details by leveraging local motion cues and assigned cumulative differences. PGA incorporates Visual Semantic Augmenter (VSA) and Rhythm Semantic Adapter (RSA) to synchronize audio by fusing static visual semantics with dynamic motion semantics to improve temporal coherence. By integrating DET and PGA, this framework strengthens audiovisual immersion in VR natural dynamic scenes from both visual and auditory perspectives. Quantitative and qualitative evaluations demonstrate that our approach surpasses most existing methods in terms of texture realism and audiovisual synchronization, offering new insights for advancing immersive experiences in dynamic VR phenomena. Xiangcheng Zhai, Yuxuan Qiu, Xiaohui Tan, Aimin Hao, Yang Gao 0032 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2026 | EmoPoseFace: Head Pose Aware Speech-Driven 3D Emotional Facial Animation Using Latent DiffusionabstractSpeech-driven 3D facial animation has notable applications in the VR domain, including virtual anchors and digital avatars, etc. However, producing facial animations that convey complex emotional expressions remains a substantial challenge. Existing methods struggle to simultaneously achieve accurate lip synchronization, natural facial expressions, and realistic emotional representation. Significantly, the impact of head pose on boosting facial emotional expressiveness has not been thoroughly investigated. To address these issues, we propose EmoPoseFace, a novel Diffusion-based network to generate speech-driven 3D emotional facial animations with synchronized head poses. Our method employs a dual-branch conditional generation architecture to separately model facial expressions and head poses, integrating emotion and head-pose conditions for coherent facial expression-pose control. In addition, we design the Global-local Facial Fine-grained Editing Module (GL-FFE), which achieves emotional enhancement of facial expressions and fine-grained facial modification, while maintains the naturalness and authenticity of facial movements. Extensive experiments demonstrate that our approach outperforms existing methods in lip-sync accuracy and emotional detail preservation. The introduction of head pose control and GL-FFE significantly expands the expressiveness of emotional virtual facial animation, and the fine-grained editing is widely approved in perceptual user studies. Xin Zhao 0025, Ju Dai, Feng Zhou 0007, Haofei Wang 0001, Aimin Hao, Hong Qin 0001, Yang Gao 0032 |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2026 | EmoDiffuser: emotional diffuser for speech-driven 3D facial animation
Xin Zhao 0025, Ju Dai, Feng Zhou 0007, Haofei Wang 0001, Aimin Hao, JunJun Pan, Yang Gao 0032 |
Vis. Comput. | 7 |
| 2025 | Effects of interaction modalities and emotional states on user's perceived empathy with an LLM-based embodied conversational agent
Yang Gao 0032, Yangbin Dai, Guangtao Zhang, Aimin Hao, Shuai Li 0001 |
Int. J. Hum. Comput. Stud. | 1 |
| 2025 | Trust in Virtual Agents: Exploring the Role of Stylization and VoiceabstractWith the continuous advancement of artificial intelligence technology, data-driven methods for reconstructing and animating virtual agents have achieved increasing levels of realism. However, there is limited research on how these novel data-driven methods, combined with voice cues, affect user perceptions. We use advanced data-driven methods to reconstruct stylized agents and combine them with synthesized voices to study their effects on users' trust and other perceptions (e.g. social presence and empathy). Through an experiment with 27 participants, our findings reveal that stylized virtual agents enhance user trust to a degree comparable to real style, while voice has a negligible effect on trust. Additionally, elder agents are more likely to be trusted. The style of the agents also plays a key role in participants' perceived realism, and audio-visual matching significantly enhances perceived empathy. These results provide new insights into designing trustworthy virtual agents and further support and validate the audio-visual integration theory. Yang Gao 0032, Yangbin Dai, Guangtao Zhang, Fariba Mostajeran, Binge Zheng, Tao Yu 0007 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | Saliency-Aware Foveated Path Tracing for Virtual Reality RenderingabstractFoveated rendering reduces computational load by distributing resources based on the human visual system. This enables the implementation of ray tracing in virtual reality applications, where a high frame rate is essential to achieve visual immersion. However, traditional foveation methods based solely on eccentricity cannot adequately account for the complex behavior of visual attention. This is one of the main reasons that leads to lower perceived quality compared to non-foveated techniques. In this study, we introduce a novel rendering pipeline that incorporates ocular attention through the use of visual saliency. Based on foveation saliency, our approach facilitates the real-time production of high-quality images utilizing path tracing by distributing samples according to saliency metrics derived from geometric and historical data. To further augment image quality, an adaptive filtering process, aligned with the saliency metrics, is employed to reduce visible artifacts in non-foveal regions. Our experiments prove that this novel approach can demonstrate superior performance compared to previous methods, both in terms of quantitative metrics and perceived visual quality. Yang Gao 0032, Wencan Li, Shiyu Liang, Aimin Hao, Xiaohui Tan |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | Efficient Photon Beam Diffusion for Directional Subsurface ScatteringabstractReal-time subsurface scattering techniques are widely used in translucent material rendering. Among advanced methods that rely on the bidirectional scattering-surface reflectance distribution function (BSSRDF), screen space algorithms exhibit limited translucency, while existing large-distance methods are inefficient and yield poor illumination details. To address these limitations for better large-distance scattering, we develop a novel algorithm by extending the photon beam diffusion (PBD) model within the light view and screen space. Unlike surface irradiance in prior methods, we incorporate the refracted beam in the medium into real-time scattering estimation, presenting a new consideration for photon beam utilization. Concretely, we store all photon beam samples in light view textures and utilize an adaptive sampling pattern for beam sample selection in large filtering kernel sizes. This can reduce the sample count based on surface attributes. In screen space, virtual sources are derived from samples to estimate PBD contributions, with an approximation that preserves boundary conditions. To avoid possible overestimation, we implement correction factors that scale contributions, effectively aligning our results with path-tracing references. Through these reformulations, our efficient PBD generates results closest to references among existing methods. The experiments accurately represent better front-face illumination details and backlit translucency effects, while significantly accelerating performance compared to previous large-distance methods. Shiyu Liang, Yang Gao 0032, Chonghao Hu, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Fluid Inverse Volumetric Modeling and Applications From Surface MotionabstractIn this study, we devise a framework for volumetrically reconstructing fluid from observable, measurable free surface motion. Our innovative method amalgamates the benefits of deep learning and conventional simulation to preserve the guiding motion and temporal coherence of the reproduced fluid. We infer surface velocities by encoding and decoding spatiotemporal features of surface sequences, and a 3D CNN is used to generate the volumetric velocity field, which is then combined with 3D labels of obstacles and boundaries. Concurrently, we employ a network to estimate the fluid's physical properties. To progressively evolve the flow field over time, we input the reconstructed velocity field and estimated parameters into the physical simulator as the initial state. Our approach yields promising results for both synthetic fluid generated by different fluid solvers and captured real fluid. The developed framework naturally lends itself to a variety of graphics applications, such as 1) effective reproductions of fluid behaviors visually congruent with the observed surface motion, and 2) physics-guided re-editing of fluid scenes. Extensive experiments affirm that our novel method surpasses state-of-the-art approaches for 3D fluid inverse modeling and animation in graphics. Xueguang Xie, Yang Gao 0032, Fei Hou 0001, Tianwei Cheng, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Motion Editing for Quadruped Characters via Latent Frequency EmbeddingabstractThe accurate and diversified generation of motion sequences for virtual characters poses both an enticing and challenging task within the domain of 3D animation and game content production. To achieve a natural and realistic full-body motion, the movements of virtual characters must adhere to a set of constraints, promoting reliable and seamless pose-changing. This study presents a two-stage model specifically designed to learn Inverse Kinematics (IK) constraints from the representative quadruped character poses. In the first stage, we employ frequency analysis to decompose motion poses into the base-level and style-level components. The base-level content encapsulates the global correlations in the dataset, while the style-level variation centers on distinguishing the local attributes in similar data elements. In order to construct data correlations among poses, we embed the decomposed pose feature into a latent space in the second stage. The kernel matrix of the embedding, which is refined from the original joint angles to the decomposed representation and the IK constraints, creates a more compact distribution of the pose similarity and also guarantees a plausible sampling result with certain IK constraints. Moreover, new motions from the edited IK constraints can also be generated by proposing a searching strategy to adapt to our latent embedding. Experimental results reveal that our method is competitive with the state-of-the-art synthetic approaches in terms of accuracy, highlighting our considerable potential for high efficiency in the animation production. JunJun Pan, Ju Dai, Yang Gao 0032, Junxuan Bai, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Real-time immersive haptic sculpting with elastoplastic virtual clay
Zhiyang Ji, Aimin Hao, Yang Gao 0032 |
Vis. Comput. | 4 |
| 2024 | Weakly Supervised Multimodal Affordance Grounding for Egocentric ImagesabstractTo enhance the interaction between intelligent systems and the environment, locating the affordance regions of objects is crucial. These regions correspond to specific areas that provide distinct functionalities. Humans often acquire the ability to identify these regions through action demonstrations and verbal instructions. In this paper, we present a novel multimodal framework that extracts affordance knowledge from exocentric images, which depict human-object interactions, as well as from accompanying textual descriptions that describe the performed actions. The extracted knowledge is then transferred to egocentric images. To achieve this goal, we propose the HOI-Transfer Module, which utilizes local perception to disentangle individual actions within exocentric images. This module effectively captures localized features and correlations between actions, leading to valuable affordance knowledge. Additionally, we introduce the Pixel-Text Fusion Module, which fuses affordance knowledge by identifying regions in egocentric images that bear resemblances to the textual features defining affordances. We employ a Weakly Supervised Multimodal Affordance (WSMA) learning approach, utilizing image-level labels for training. Through extensive experiments, we demonstrate the superiority of our proposed method in terms of evaluation metrics and visual results when compared to existing affordance grounding models. Furthermore, ablation experiments confirm the effectiveness of our approach. Code:https://github.com/xulingjing88/WSMA. Lingjing Xu, Yang Gao 0032, Wenfeng Song, Aimin Hao |
AAAI | 2 |
| 2024 | HOIAnimator: Generating Text-Prompt Human-Object Animations Using Novel Perceptive Diffusion ModelsabstractTo date, the quest to rapidly and effectively produce human-object interaction (HOI) animations directly from textual descriptions stands at the forefront of computer vision research. The underlying challenge demands both a discriminating interpretation of language and a comprehen-sive physics-centric model supporting real-world dynamics. To ameliorate, this paper advocates HOIAnimator, a novel and interactive diffusion model with perception ability and also ingeniously crafted to revolutionize the animation of complex interactions from linguistic narratives. The effectiveness of our model is anchored in two ground-breaking innovations: (1) Our Perceptive Diffusion Models (PDM) brings together two types of models: one focused on hu-man movements and the other on objects. This combination allows for animations where humans and objects move in concert with each other, making the overall motion more realistic. Additionally, we propose a Perceptive Message Passing (PMP) mechanism to enhance the communication bridging the two models, ensuring that the animations are smooth and unified; (2) We devise an Interaction Contact Field (ICF), a sophisticated model that implicitly captures the essence of HOls. Beyond mere predictive contact points, the ICF assesses the proximity of human and object to their respective environment, informed by a probabilistic distribution of interactions learned throughout the denoising phase. Our comprehensive evaluation showcases HOlani-mator's superior ability to produce dynamic, context-aware animations that surpass existing benchmarks in text-driven animation synthesis. Wenfeng Song, Shuai Li 0001, Yang Gao 0032, Aimin Hao, Xia Hau, Chenglizhao Chen, Hong Qin 0001 |
CVPR | 4 |
| 2024 | Detail Enhancement for Free Surface LBM Using Adaptive Sizing of Coupled Particles
Qingyue Qu, Shaonan Zhu, Aimin Hao, Yang Gao 0032 |
ICXR | 6 |
| 2024 | SIE-DepthNet: Semantic-Guided Monocular Depth Estimation for Dynamic Environment
Zilong Song, Yang Gao 0032, Sijia Dai, Shuai Li 0001, Aimin Hao, Shoulong Zhang |
ICXR | 2 |
| 2024 | ANFluid: Animate Natural Fluid Photos base on Physics-Aware Simulation and Dual-Flow Texture LearningabstractGenerating photorealistic animations from a single still photo represents a significant advancement in multimedia editing and artistic creation. While existing AIGC methods have reached milestone successes, they often struggle with maintaining consistency with real-world physical laws, particularly in fluid dynamics. To address this issue, this paper introduces ANFluid, a physics solver and data-driven coupled framework that combines physics-aware simulation (PAS) and dual-flow texture learning (DFTL) to animate natural fluid photos effectively. The PAS component of ANFluid ensures that motion guides adhere to physical laws, and can be automatically tailored with specific numerical solver to meet the diversities of different fluid scenes. Concurrently, DFTL focuses on enhancing texture prediction. It employs bidirectional self-supervised optical flow estimation and multi-scale wrapping to strengthen dynamic relationships and elevate the overall animation quality. Notably, despite being built on a transformer architecture, the innovative encoder-decoder design in DFTL does not increase the parameter count but rather enhances inference efficiency. Extensive quantitative experiments have shown that our ANFluid surpasses most current methods on the Holynski and CLAW datasets. User studies further confirm that animations produced by ANFluid maintain better physical and content consistency with the real world and the original input, respectively. Moreover, ANFluid supports interactive editing during the simulation process, enriching the animation content and broadening its application potential. Xiangcheng Zhai, Yingqi Jie, Xueguang Xie, Aimin Hao, Yang Gao 0032 |
ACM Multimedia | 6 |
| 2024 | CoupNeRF: Property-aware Neural Radiance Fields for Multi-Material Coupled Scenario ReconstructionabstractAbstract Neural Radiance Fields (NeRFs) have achieved significant recognition for their proficiency in scene reconstruction and rendering by utilizing neural networks to depict intricate volumetric environments. Despite considerable research dedicated to reconstructing physical scenes, rare works succeed in challenging scenarios involving dynamic, multi‐material objects. To alleviate, we introduce CoupNeRF, an efficient neural network architecture that is aware of multiple material properties. This architecture combines physically grounded continuum mechanics with NeRF, facilitating the identification of motion systems across a wide range of physical coupling scenarios. We first reconstruct specific‐material of objects within 3D physical fields to learn material parameters. Then, we develop a method to model the neighbouring particles, enhancing the learning process specifically in regions where material transitions occur. The effectiveness of CoupNeRF is demonstrated through extensive experiments, showcasing its proficiency in accurately coupling and identifying the behavior of complex physical scenes that span multiple physics domains. Jin Li 0068, Yang Gao 0032, Wenfeng Song, Yacong Li, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
Comput. Graph. Forum | 2 |
| 2024 | State of the Art in Efficient Translucent Material Rendering with BSSRDFabstractAbstract Sub‐surface scattering is always an important feature in translucent material rendering. When light travels through optically thick media, its transport within the medium can be approximated using diffusion theory, and is appropriately described by the bidirectional scattering‐surface reflectance distribution function (BSSRDF). BSSRDF methods rely on assumptions about object geometry and light distribution in the medium, which limits their applicability to general participating media problems. However, despite the high computational cost of path tracing, BSSRDF methods are often favoured due to their suitability for real‐time applications. We review these methods and discuss the most recent breakthroughs in this field. We begin by summarizing various BSSRDF models and then implement most of them in a 2D searchlight problem to demonstrate their differences. We focus on acceleration methods using BSSRDF, which we categorize into two primary groups: pre‐computation and texture methods. Then we go through some related topics, including applications and advanced areas where BSSRDF is used, as well as problems that are sometimes important yet are ignored in sub‐surface scattering estimation. In the end of this survey, we point out remaining constraints and challenges, which may motivate future work to facilitate sub‐surface scattering. Shiyu Liang, Yang Gao 0032, Chonghao Hu, Aimin Hao, Lili Wang 0006, Hong Qin 0001 |
Comput. Graph. Forum | 2 |
| 2024 | Dynamic ocean inverse modeling based on differentiable renderingabstractLearning and inferring underlying motion patterns of captured 2D scenes and then re-creating dynamic evolution consistent with the real-world natural phenomena have high appeal for graphics and animation. To bridge the technical gap between virtual and real environments, we focus on the inverse modeling and reconstruction of visually consistent and property-verifiable oceans, taking advantage of deep learning and differentiable physics to learn geometry and constitute waves in a self-supervised manner. First, we infer hierarchical geometry using two networks, which are optimized via the differentiable renderer. We extract wave components from the sequence of inferred geometry through a network equipped with a differentiable ocean model. Then, ocean dynamics can be evolved using the reconstructed wave components. Through extensive experiments, we verify that our new method yields satisfactory results for both geometry reconstruction and wave estimation. Moreover, the new framework has the inverse modeling potential to facilitate a host of graphics applications, such as the rapid production of physically accurate scene animation and editing guided by real ocean scenes. Xueguang Xie, Yang Gao 0032, Fei Hou 0001, Aimin Hao, Hong Qin 0001 |
Comput. Vis. Media | 2 |
| 2024 | Erratum to: Dynamic ocean inverse modeling based on differentiable renderingabstractThe authors apologize for a hidden error in the article. It is that the images in Figs. 14(a) and 14(d) were mistakenly presented as left–right mirror images. The authors have flipped them to ensure that the figures now correspond correctly with others in the subfigures (b, c, e, f). The accurate version of Fig. 14 is provided as below. Xueguang Xie, Yang Gao 0032, Fei Hou 0001, Aimin Hao, Hong Qin 0001 |
Comput. Vis. Media | 2 |
| 2024 | MPMNet: A Data-Driven MPM Framework for Dynamic Fluid-Solid InteractionabstractHigh-accuracy, high-efficiency physics-based fluid-solid interaction is essential for reality modeling and computer animation in online games or real-time Virtual Reality (VR) systems. However, the large-scale simulation of incompressible fluid and its interaction with the surrounding solid environment is either time-consuming or suffering from the reduced time/space resolution due to the complicated iterative nature pertinent to numerical computations of involved Partial Differential Equations (PDEs). In recent years, we have witnessed significant growth in exploring a different, alternative data-driven approach to addressing some of the existing technical challenges in conventional model-centric graphics and animation methods. This article showcases some of our exploratory efforts in this direction. One technical concern of our research is to address the central key challenge of how to best construct the numerical solver effectively and how to best integrate spatiotemporal/dimensional neural networks with the available MPM's pressure solvers. In particular, we devise the MPMNet, a hybrid data-driven framework supporting the popular and powerful MPM, to combine the comprehensive properties of MPM in numerically handling physical behaviors ranging from fluid to deformable solids and the high efficiency of data-driven models. At the architectural level, our MPMNet comprises three primary components: A data processing module to describe the physical properties by way of the input fields; A deep neural network group to learn the spatiotemporal features; And an iterative refinement process to continue to reduce possible numerical errors. The goal of these special technical developments is to aim at involved numerical acceleration while preserving physical accuracy, realizing efficient and accurate fluid-solid interactions in a data-driven fashion. The extensive experimental results verify that our MPMNet can tremendously speed up the computation compared with the popular numerical methods as the complexity of interaction scenes increases while better retaining the numerical accuracy. Jin Li 0068, Yang Gao 0032, Ju Dai, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | A Unified Particle-Based Solver for Non-Newtonian Behaviors SimulationabstractIn this article, we present a unified framework to simulate non-Newtonian behaviors. We combine viscous and elasto-plastic stress into a unified particle solver to achieve various non-Newtonian behaviors ranging from fluid-like to solid-like. Our constitutive model is based on a Generalized Maxwell model, which incorporates viscosity, elasticity and plasticity in one non-linear framework by a unified way. On the one hand, taking advantage of the viscous term, we construct a series of strain-rate dependent models for classical non-Newtonian behaviors such as shear-thickening, shear-thinning, Bingham plastic, etc. On the other hand, benefiting from the elasto-plastic model, we empower our framework with the ability to simulate solid-like non-Newtonian behaviors, i.e., visco-elasticity/plasticity. In addition, we enrich our method with a heat diffusion model to make our method flexible in simulating phase change. Through sufficient experiments, we demonstrate a wide range of non-Newtonian behaviors ranging from viscous fluid to deformable objects. We believe this non-Newtonian model will enhance the realism of physically-based animation, which has great potential for computer graphics. Yang Gao 0032, Tianwei Cheng, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Multimodal Physiological Analysis of Impact of Emotion on Cognitive Control in VRabstractCognitive control is often perplexing to elucidate and can be easily influenced by emotions. Understanding the individual cognitive control level is crucial for enhancing VR interaction and designing adaptive and self-correcting VR/AR applications. Emotions can reallocate processing resources and influence cognitive control performance. However, current research has primarily emphasized the impact of emotional valence on cognitive control tasks, neglecting emotional arousal. In this study, we comprehensively investigate the influence of emotions on cognitive control based on the arousal-valence model. A total of 26 participants are recruited, inducing emotions through VR videos with high ecological validity and then performing related cognitive control tasks. Leveraging physiological data including EEG, HRV, and EDA, we employ classification techniques such as SVM, KNN, and deep learning to categorize cognitive control levels. The experiment results demonstrate that high-arousal emotions significantly enhance users' cognitive control abilities. Utilizing complementary information among multi-modal physiological signal features, we achieve an accuracy of 84.52% in distinguishing between high and low cognitive control. Additionally, time-frequency analysis results confirm the existence of neural patterns related to cognitive control, contributing to a better understanding of the neural mechanisms underlying cognitive control in VR. Our research indicates that physiological signals measured from both the central and autonomic nervous systems can be employed for cognitive control classification, paving the way for novel approaches to improve VR/AR interactions. JunJun Pan, Yang Gao 0032, Hong Qin 0001, Yang Shen 0009 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Interactive Virtual Ankle Movement Controlled by Wrist sEMG Improves Motor Imagery: An Exploratory StudyabstractVirtual reality (VR) techniques can significantly enhance motor imagery training by creating a strong illusion of action for central sensory stimulation. In this article, we establish a precedent by using surface electromyography (sEMG) of contralateral wrist movement to trigger virtual ankle movement through an improved data-driven approach with a continuous sEMG signal for fast and accurate intention recognition. Our developed VR interactive system can provide feedback training for stroke patients in the early stages, even if there is no active ankle movement. Our objectives are to evaluate: 1) the effects of VR immersion mode on body illusion, kinesthetic illusion, and motor imagery performance in stroke patients; 2) the effects of motivation and attention when utilizing wrist sEMG as a trigger signal for virtual ankle motion; 3) the acute effects on motor function in stroke patients. Through a series of well-designed experiments, we have found that, compared to the 2D condition, VR significantly increases the degree of kinesthetic illusion and body ownership of the patients, and improves their motor imagery performance and motor memory. When compared to conditions without feedback, using contralateral wrist sEMG signals as trigger signals for virtual ankle movement enhances patients' sustained attention and motivation during repetitive tasks. Furthermore, the combination of VR and feedback has an acute impact on motor function. Our exploratory study suggests that the sEMG-based immersive virtual interactive feedback provides an effective option for active rehabilitation training for severe hemiplegia patients in the early stages, with great potential for clinical application. Yanqing Xiao, Hongming Bai, Yang Gao 0032, Ben Hu, XiaoE Cai, Jiasheng Rao, Aimin Hao |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Mem-Box: VR sandbox for adaptive working memory evaluation and training using physiological signals
Yang Gao 0032 |
Vis. Comput. | 3 |
| 2023 | Dual Temporal Transformers for Fine-Grained Dangerous Action RecognitionabstractRecognizing dangerous actions is a critical task in computer vision, especially for surveillance applications. While existing deep learning methods have been successful in confined environments, they struggle with the anomalous and salient variations of human postures in dangerous actions. Additionally, finer-grained dangerous actions require more discriminative cues, adding to the complexity of the task. To address these challenges, we propose a novel solution that models the intrinsic and invariant properties of dangerous actions at multiple temporal semantic levels. Concretely, we propose a Dual Temporal Transformers (DTT) to capture temporal interactions between distinct key points in the human body aggregation from shallow to deep layers, increasing the perception field from local to global, simultaneously. By doing so, our method avoids overfitting to unrelated or minor clues in videos and achieves a generalized representation of abnormal actions. We evaluate our approach on indoor and outdoor environments and found that DTT outperforms existing methods in terms of efficiency and accuracy. Our code and dataset are pubic available on https://github.com/AveryJohnsonJJ/DTT.git. Wenfeng Song, Xingliang Jin, Yang Gao 0032, Xia Hou |
ICIP | 4 |
| 2023 | PhyVR: Physics-based Multi-material and Free-hand Interaction in VRabstractThe realistic interaction with physical phenomena is a crucial aspect of human-computer interaction (HCI) in virtual reality (VR). However, the real-time performance of physical simulation, interactive computation, and rendering is the bottleneck of physics-based VR HCI. To address these challenges, we propose a novel physics-oriented framework for multi-material objects and free-hand interaction, termed PhyVR. This framework enables users to interact with diverse virtual phenomena dynamically. At the algorithm level, we develop a unified particle system to describe both the virtual multi-materials and the user’s avatar for the efficiency issue, optimize collision detection, and accelerate the HCI algorithms with a variable fine-coarse particle sampling scheme. At the rendering level, we introduce a hybrid particle-grid anisotropic algorithm for surface reconstruction, enabling real-time and visually convincing fluid rendering. Comprehensive experiments and user studies demonstrate that our framework effectively captures various physical interaction phenomena, providing an enhanced user experience and paving the way for expanding VR-related HCI applications. Hanchen Deng, Jin Li 0068, Yang Gao 0032, Xiaohui Liang 0001, Aimin Hao |
ISMAR | 3 |
| 2022 | Person Re-Identification in Panoramic Views Based on Bayesian TransformersabstractThe panoramic view cameras offer more broad perspectives and continuous information for person re-identification (ReID). However, the panoramic-view videos suffer from objects distortion and bring more occlusion due to the fixed or moving capture points. This paper proposes a novel Bayesian Transformer Network (BTN) to adaptively capture the occlusion clues as Bayesian prior to guide the discriminative pedestrian-related feature extraction in the high-occlusion scenes. The Bayesian prior is built via a pre-trained CNN, which could recognize different occluded scenarios based on the severeness of noisy backgrounds. Moreover, to fully explore the occlusion prior, we propose to embed the semantic labels into a well-designed transformer network. By fostering the collaborative occlusion clues between the person and background, our method could achieve outstanding performance on both public benchmarks and panoramic view videos, which verifies the advantages of our BTN framework over existing methods. Wenfeng Song, Yang Gao 0032, Aimin Hao, Xia Hou |
ICIP | 4 |
| 2022 | Neurophysiological and Subjective Analysis of VR Emotion Induction ParadigmabstractThe ecological validity of emotion-inducing scenarios is essential for emotion research. In contrast to the classical passive induction paradigm, immersive VR fully engages the psychological and physiological components of the subject, which is considered an ecologically valid paradigm for studying emotion. Several studies investigate the emotional responses to different VR tasks or games using subjective scales. However, little research regards VR as an eliciting material, especially when systematically analyzing emotional processes in VR from a neurophysiological perspective. To fill this gap and scientifically evaluate VR's ability to be used as an active method for emotion elicitation, we investigate the dynamic relationship between explicit information (subjective evaluations) and implicit information (objective neurophysiological data). A total of 28 participants are enlisted to watch eight VR videos while their SAM/IPQ scores and EEG data are recorded simultaneously. In ecologically valid scenarios, the subjective results demonstrate that VR has significant advantages for evoking emotion in arousal-valence. This conclusion is backed by our examination of objective neurophysiological evidence that VR videos effectively induce high-arousal emotions. In addition, we obtain features of critical channels and frequency oscillations associated with emotional valence, thereby validating previous research in more lifelike circumstances. In particular, we discover hemispheric asymmetry in the occipital region under high and low emotional arousal, which adds to our understanding of neural features and the dynamics of emotional arousal. As a result, we successfully integrate EEG and VR to demonstrate that VR is more pragmatic for evoking natural feelings and is beneficial for emotional research. Our research has set a precedent for new methodologies of using VR induction paradigms to acquire a more reliable explanation of affective computing. JunJun Pan, Yang Gao 0032, Yang Shen 0009, Ju Dai, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | Simulating Multi-Scale, Granular Materials and Their Transitions With a Hybrid Euler-Lagrange SolverabstractMulti-scale granular materials, such as powdered materials and mudslides, are pretty common in nature. Modeling such materials and their phase transitions remains challenging since this task involves the delicate representations of various ranges of particles with multiple scales that cause their property variations among liquid, granular solid (i.e., particles), and smoke-like materials. To effectively animate the complicated yet intriguing natural phenomena involving multi-scale granular materials and their phase transitions in graphics with high fidelity, this article advocates a hybrid Euler-Lagrange solver to handle the behaviors of involved discontinuous fluid-like materials faithfully. At the algorithmic level, we present a unified framework that tightly couples the affine particle-in-cell (APIC) solver with density field to achieve the transformation spanning across granular particles, dust cloud, powders, and their natural mixtures. For example, a part of the granular particles could be transformed into dust cloud while interacting with air and being represented by density field. Meanwhile, the velocity decrease of the involved materials could also result in the transit from the density-field-driven dust to powder particles. Besides, to further enhance our modeling and simulation power to broaden the range of multi-scale materials, we introduce a moisture property for granular particles to control the transitions between particles and viscous liquid. At the geometric level, we devise an additional surface-tracking procedure to simulate the viscous liquid phase. We can arrive at delicate viscous behaviors by controlling the corresponding yield conditions. Through various experiments with the different scenes design being conducted in our unified framework, we can validate the mixed multi-scale materials' mutual transformation processes. Our unified framework furnished with a hybrid solver can significantly enhance the modeling flexibility and the animation potential of the particle-grid hybrid materials in graphics. Yang Gao 0032, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | Accelerating Liquid Simulation With an Improved Data-Driven MethodabstractAbstract In physics‐based liquid simulation for graphics applications, pressure projection consumes a significant amount of computational time and is frequently the bottleneck of the computational efficiency. How to rapidly apply the pressure projection and at the same time how to accurately capture the liquid geometry are always among the most popular topics in the current research trend in liquid simulations. In this paper, we incorporate an artificial neural network into the simulation pipeline for handling the tricky projection step for liquid animation. Compared with the previous neural‐network‐based works for gas flows, this paper advocates new advances in the composition of representative features as well as the loss functions in order to facilitate fluid simulation with free‐surface boundary. Specifically, we choose both the velocity and the level‐set function as the additional representation of the fluid states, which allows not only the motion but also the boundary position to be considered in the neural network solver. Meanwhile, we use the divergence error in the loss function to further emulate the lifelike behaviours of liquid. With these arrangements, our method could greatly accelerate the pressure projection step in liquid simulation, while maintaining fairly convincing visual results. Additionally, our neutral network performs well when being applied to new scene synthesis even with varied boundaries or scales. Yang Gao 0032, Quancheng Zhang, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
Comput. Graph. Forum | 1 |
| 2020 | Dynamic particle partitioning SPH model for high-speed fluids simulation
Yang Gao 0032, Jin Li 0068, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
Graph. Model. | 1 |
| 2019 | A Hybrid Method for Powdered Materials ModelingabstractPowdered materials, such as sand and flour, are quite common in nature, whose properties always range from granular particles to smog materials under the air friction while throwing. This paper presents a hybrid method that tightly couples APIC solver with density field to accomplish the transformation of continuous powdered materials varying among granular particles, smog, powders and their natural mixtures. In our method, a part of the granular particles will be transformed to dust smog while interacting with air and represented by density field, then, as velocity decreases the density-based dust will deposit to powder particles. We construct a unified framework to imitate the mutual transformation process for the powdered materials of different scales, which greatly enhance the details of particle-based materials modeling. We have conducted extensive experiments to verify the performance of our model, and get satisfactory results in terms of stability, efficiency and visual authenticity as expected. Yang Gao 0032, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
VRST | 1 |
| 2019 | An efficient FLIP and shape matching coupled method for fluid-solid and two-phase fluid simulations
Yang Gao 0032, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
Vis. Comput. | 1 |
| 2019 | Real-time simulation of electrocautery procedure using meshfree methods in laparoscopic cholecystectomy
JunJun Pan, Yang Gao 0032, Hong Qin 0001, Yaqing Si |
Vis. Comput. | 3 |
| 2017 | A novel fluid-solid coupling framework integrating FLIP and shape matching methodsabstractPhysically-based fluid animation and solid deformation driven by numerical simulation have manifested their significance for many graphics applications during the past two decades. For example, the fluid implicit particle (FLIP) method and shape matching technique based on position based dynamics (PBD) have demonstrated their unique graphics strength in fluid and solid animation, respectively. We propose a novel integrated approach supporting the seamless unification of FLIP and shape matching. We devise new algorithms to tackle existing difficulties when handling new phenomena such as high-fidelity fluid-solid interaction and solid melting. The key innovation of this paper is a unified Lagrangian framework that seamlessly blends FLIP and PBD based shape matching constraint towards the natural yet strong coupling between fluid and deformable solid. Within our integrated framework, it enables many complicated fluid-solid phenomena with ease. We conduct various kinds of experiments. All the results demonstrate the advantages of our unified hybrid approach towards visual fidelity, efficiency, stability, and versatility. Yang Gao 0032, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
CGI | 1 |
| 2017 | An efficient heat-based model for solid-liquid-gas phase transition and dynamic interaction
Yang Gao 0032, Shuai Li 0001, Lipeng Yang, Hong Qin 0001, Aimin Hao |
Graph. Model. | 1 |