Aimin Hao

dblp:94/5679 · DBLP profile ↗
← Back
184ranked-venue papers
2as first author
97since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 132 · 2 first-author · 66 since 2021Artificial intelligence and machine learning · 44 · 32 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 10 since 2021Human-computer interaction and ubiquitous computing · 7 · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DECON: Reconstruction of Clothed-Geometric Multiple Humans from a Single Image via Geometry-Guided Decoupling
abstract
3D multi-human reconstruction from single images holds significant potential for advancing AR/VR applications. While remarkable progress has been made in single-human reconstruction, existing methods face challenges when reconstructing multiple humans. These challenges include: (1) severe inter-occlusion that disrupts individual body structures, and (2) the absence of physically plausible relative positioning among subjects. We present DECON, a novel DEcouple-and-reCONstruct framework that systematically addresses these limitations through two technical innovations: (1) a decouple-and-reconstruct framework with multi-view synthesis. It separates individuals and reconstructs detailed 3D bodies from a single image. (2) a Perspective-Aware Position Optimization (PAPO) approach. It ensures realistic positioning by fixing overlaps and gaps between subjects. Extensive experiments demonstrate our method's capability to reconstruct fully separated, anatomically complete 3D humans with clothed-geometric details and plausible interactions. Quantitative evaluations show a 54% reduction in Chamfer Distance and 35% in Point-to-Surface Distance compared to state-of-the-art methods.
Yiming Jiang 0018, Wenfeng Song, Shuai Li 0001, Aimin Hao
AAAI4
2026 IntentMotion: Learning Intent-Aware Human Motion from Language in 3D Scenes
abstract
Generating human motion in complex 3D scenes from text is a challenging task with broad applications. However, existing methods often overlook realistic physical contact, resulting in visually plausible but physically unrealistic motion, e.g., penetration. To alleviate this, we propose IntentMotion, a novel framework that generates human motion in 3D scenes from natural language instructions by explicitly modeling intent. We first introduce the Intention-Guided Contact Field (IGCF). This differentiable voxel-based contact region representation explicitly aligns parsed language roles with spatial contact regions through a hierarchical attention mechanism. IGCF is jointly trained with a diffusion-based motion generator, allowing contact predictions to adapt dynamically through gradient feedback. To improve the controllability and physics-aware motion, we further propose an Intention-Aware Diffusion Model (IADM), which decouples the high-level semantic planning from the low-level contact refinement in a coarse-to-fine process. The optimized contact cues are utilized to guide the synthesis of a coarse trajectory, followed by refining detailed pose sequences under IGCF supervision. Experiments on the HUMANISE and LINGO datasets demonstrate that our IntentMotion outperforms recent baselines in contact accuracy, semantic alignment, and generalization to unseen scenes.
Wenfeng Song, Shi Zheng, Xingliang Jin, Aimin Hao, Fei Hou 0001, Xia Hou, Shuai Li 0001
AAAI5
2026 Energy-based haptic rendering for real-time surgical simulation
Mingbo Hu, Wenli Xiu, Siming Zheng, Shuai Li 0001, Aimin Hao
Comput. Graph.8
2026 Weakly supervised visual-auditory fixation prediction with multigranularity perception
Guotao Wang 0004, Chenglizhao Chen, Deng-Ping Fan, Aimin Hao, Qinping Zhao
Sci. China Inf. Sci.4
2026 MedSAM-guided geometry-aware 2D-3D feature fusion for medical image registration
Yuanbo He, Aimin Hao
Neural Networks8
2026 Deep-Saliency Foveated Ray Tracing For Real-time VR Rendering
abstract
Immersive VR applications demand high resolutions and refresh rates, posing significant challenges for real-time rendering. Foveated rendering mitigates this cost by exploiting properties of the Human Visual System (HVS), but conventional approaches often rely on oversimplified heuristic models that neglect high-level attentional cues, resulting in artifacts in peripheral regions. To this end, we present a neural saliency-driven foveated ray tracing framework that overcomes these limitations. Our method introduces a motion-aware foveation model to capture temporal dynamics and employs a lightweight convolutional neural network to predict saliency maps that reflect complex attentional patterns derived from eye-gaze data. The combination of these guides adaptive path tracing and filtering, enabling perceptually optimized rendering with minimal artifacts. Experimental results show that our approach improves perceptual quality over prior methods while sustaining real-time performance.
Yang Gao 0032, Wencan Li, Shiyu Liang, Weizichuan Feng, Qing Xia 0002, Shuai Li 0001, Aimin Hao
IEEE Trans. Vis. Comput. Graph.7
2026 HFHuman: High-Fidelity Human Reconstruction From Single Image With Multi-Modality Fusion
abstract
Accurately reconstructing high-fidelity human models from single images is critical for virtual reality applications. Existing methods often rely on 3D features from the estimated parametric human model to provide geometric priors. This approach addresses challenges such as missing limbs or deformations, which often arise due to viewpoint limitations and self-occlusion. However, accurately predicting 3D features from monocular images remains a significant challenge. This limitation poses difficulties for the fusion of 2D and 3D information. In this paper, we introduce HFHuman, a novel approach for high-fidelity human reconstruction from a single image using multi-modality fusion. HFHuman effectively fuses multiple modalities, including geometric and depth, directly from images. Our method introduces three key innovations: (1) a depth and geometric parallel reconstruction framework that simultaneously handles whole-body geometry and detailed depth reconstruction, refining a parameterized 3D human model under progressive depth guidance; (2) a pixel-voxel feature fusion strategy that combines pixel-aligned features with voxel-aligned features using a multi-modality adaptor; and (3) a depth-refined technique that integrates RGB imagery with surface normals and depth mapping. By addressing the challenge of blending 2D and 3D modalities, HFHuman results in more accurate and realistic human reconstructions. Experimental results demonstrate that HFHuman outperforms state-of-the-art methods, setting a new standard for realistic 3D human body reconstruction.
Yiming Jiang 0018, Wenfeng Song, Shuai Li 0001, Aimin Hao
IEEE Trans. Vis. Comput. Graph.4
2026 Generating Audiovisual Synergy Fluid Animation for Highly Immersive VR Experience
abstract
Generative content is increasingly applied in VR to provide immersive experiences, yet maintaining high generation quality remains challenging for audiovisual effects. Particularly in dynamic fluid phenomena, achieving realism and presence requires adherence to physical laws. To accomplish this objective, this work proposes an audiovisual synergy fluid animation generation framework, which enhances immersion by improving motion texture fidelity and audiovisual consistency. It comprises Detail-Enhanced Texture generator (DET) and Physics-Guided Audio generator (PGA). DET integrates Global-Local Physics guidance (GLP) and Temporal Texture Modeling (TTM) to produce video textures, explicitly optimizing dynamic details by leveraging local motion cues and assigned cumulative differences. PGA incorporates Visual Semantic Augmenter (VSA) and Rhythm Semantic Adapter (RSA) to synchronize audio by fusing static visual semantics with dynamic motion semantics to improve temporal coherence. By integrating DET and PGA, this framework strengthens audiovisual immersion in VR natural dynamic scenes from both visual and auditory perspectives. Quantitative and qualitative evaluations demonstrate that our approach surpasses most existing methods in terms of texture realism and audiovisual synchronization, offering new insights for advancing immersive experiences in dynamic VR phenomena.
Xiangcheng Zhai, Yuxuan Qiu, Xiaohui Tan, Aimin Hao, Yang Gao 0032
IEEE Trans. Vis. Comput. Graph.5
2026 FCMD: Fine-Grained Text-Driven Cohesive Motion Generation With Diffusion Model
abstract
Generating continuous and expressive human motion from textual descriptions is a critical challenge in applications such as gaming and filmmaking. Existing methods often struggle to maintain global coherence, realistic frame continuity, and smooth transitions. To address these limitations, we propose FCMD, a novel diffusion-based model for generating cohesive motion sequences from fine-grained textual descriptions. FCMD introduces three key innovations: (1) Fine-grained Text Fusion, which integrates detailed textual cues with transitional narratives to enhance semantic consistency; (2) History Motion Guidance, ensuring motion accuracy and consistency across consecutive frames; and (3) Smooth Stitching Sampling, which leverages preceding and current motion information to achieve seamless transitions. Additionally, FCMD employs a large language model (LLM) to refine motion datasets by extracting fine-grained textual descriptions. Extensive experiments demonstrate that FCMD outperforms state-of-the-art methods in generating coherent, natural, and highly controllable motion sequences.
Shuai Li 0001, Wenfeng Song, Aimin Hao
IEEE Trans. Vis. Comput. Graph.5
2026 DynAvatar: Dynamic 3D Head Avatar Deformation With Expression Guided Gaussian Splatting
abstract
Generating high-fidelity, expressive, and realistic 3D head avatars remains a fundamental challenge for immersive applications such as virtual reality, gaming, and telepresence. This task requires not only precise modeling of non-rigid facial deformations but also semantically controllable expression synthesis under diverse viewpoints and motion contexts. We present DynAvatar, a novel framework that integrates expression-guided deformation into the 3D Gaussian splatting pipeline to produce photorealistic and emotionally resonant head avatars. Our method introduces two key innovations: (1) an expression-guided Gaussian deformation module that tightly couples geometric displacement with high-level semantic cues, enabling fine-grained and anatomically meaningful facial animation; and (2) a spatial context embedding mechanism that encodes the canonical position of each Gaussian to preserve semantic coherence and spatial consistency during expression generation. Extensive experiments on both controlled and in-the-wild datasets demonstrate that DynAvatar significantly outperforms state-of-the-art methods in terms of visual realism, expression fidelity, and rendering quality.
Wenfeng Song, Zhongyong Ye, Shuai Li 0001, Xia Hou, Aimin Hao
IEEE Trans. Vis. Comput. Graph.6
2026 EmoPoseFace: Head Pose Aware Speech-Driven 3D Emotional Facial Animation Using Latent Diffusion
abstract
Speech-driven 3D facial animation has notable applications in the VR domain, including virtual anchors and digital avatars, etc. However, producing facial animations that convey complex emotional expressions remains a substantial challenge. Existing methods struggle to simultaneously achieve accurate lip synchronization, natural facial expressions, and realistic emotional representation. Significantly, the impact of head pose on boosting facial emotional expressiveness has not been thoroughly investigated. To address these issues, we propose EmoPoseFace, a novel Diffusion-based network to generate speech-driven 3D emotional facial animations with synchronized head poses. Our method employs a dual-branch conditional generation architecture to separately model facial expressions and head poses, integrating emotion and head-pose conditions for coherent facial expression-pose control. In addition, we design the Global-local Facial Fine-grained Editing Module (GL-FFE), which achieves emotional enhancement of facial expressions and fine-grained facial modification, while maintains the naturalness and authenticity of facial movements. Extensive experiments demonstrate that our approach outperforms existing methods in lip-sync accuracy and emotional detail preservation. The introduction of head pose control and GL-FFE significantly expands the expressiveness of emotional virtual facial animation, and the fine-grained editing is widely approved in perceptual user studies.
Xin Zhao 0025, Ju Dai, Feng Zhou 0007, Haofei Wang 0001, Aimin Hao, Hong Qin 0001, Yang Gao 0032
IEEE Trans. Vis. Comput. Graph.6
2026 EmoDiffuser: emotional diffuser for speech-driven 3D facial animation
Xin Zhao 0025, Ju Dai, Feng Zhou 0007, Haofei Wang 0001, Aimin Hao, JunJun Pan, Yang Gao 0032
Vis. Comput.5
2025 CtrlAvatar: Controllable Avatars Generation via Disentangled Invertible Networks
abstract
As virtual experiences grow in popularity, the demand for realistic, personalized, and animatable human avatars increases. Traditional methods, relying on fixed templates, often produce costly avatars that lack expressiveness and realism. To overcome these challenges, we introduce Controllable Avatars generation via disentangled invertible networks (CtrlAvatar), a real-time framework for generating lifelike and customizable avatars. CtrlAvatar uses disentangled invertible networks to separate the deformation process into implicit body geometry and explicit texture components. This approach eliminates the need for repeated occupancy reconstruction, enabling detailed and coherent animations. The body geometry component ensures anatomical accuracy, while the texture component allows for complex, artifact-free clothing customization. This architecture ensures smooth integration between body movements and surface details. By optimizing transformations with position-varying offsets from the avatar’s initial Linear Blend Skinning vertices, CtrlAvatar achieves flexible, natural deformations that adapt to various scenarios. Extensive experiments show that CtrlAvatar outperforms other methods in quality, diversity, controllability, and cost-efficiency, marking a significant advancement in avatar generation.
Wenfeng Song, Fei Hou 0001, Shuai Li 0001, Aimin Hao, Xia Hou
AAAI5
2025 3D Dental Model Segmentation with Geometrical Boundary Preserving
abstract
3D intraoral scan mesh is widely used in digital dentistry diagnosis, segmenting 3D intraoral scan mesh is a critical preliminary task. Numerous approaches have been devised for precise tooth segmentation. Currently, the deep learning-based methods are capable of the high accuracy segmentation of crown. However, the segmentation accuracy at the junction between the crown and the gum is still below average. Existing down-sampling methods are unable to effectively preserve the geometric details at the junction. To address these problems, we propose CrossTooth, a boundary-preserving segmentation method that combines 3D mesh selective downsampling to retain more vertices at the tooth-gingiva area, along with cross-modal discriminative boundary features extracted from multi-view rendered images, enhancing the geometric representation of the segmentation network. Using a point network as a backbone and incorporating image complementary features, CrossTooth significantly improves segmentation accuracy, as demonstrated by experiments on a public intraoral scan dataset. The source code is available at https://github.com/XiShuFan/CrossTooth_CVPR2025
Shufan Xi, Zexian Liu, Junlin Chang, Aimin Hao
CVPR6
2025 Effects of interaction modalities and emotional states on user's perceived empathy with an LLM-based embodied conversational agent
Yang Gao 0032, Yangbin Dai, Guangtao Zhang, Aimin Hao, Shuai Li 0001
Int. J. Hum. Comput. Stud.5
2025 Advancing MRI segmentation with CLIP-driven semi-supervised learning and semantic alignment
Kexuan Li, Jingjuan Liu, Xuehao Wang, Yuanbo He, Huadan Xue, Aimin Hao, Shuai Li 0001
Neurocomputing9
2025 Saliency-Free and Aesthetic-Aware Panoramic Video Navigation
abstract
Most of the existing panoramic video navigation approaches are saliency-driven, whereby off-the-shelf saliency detection tools are directly employed to aid the navigation approaches in localizing video content that should be incorporated into the navigation path. In view of the dilemma faced by our research community, we rethink if the "saliency clues" are really appropriate to serve the panoramic video navigation task. According to our in-depth investigation, we argue that using "saliency clues" cannot generate a satisfying navigation path, failing to well represent the given panoramic video, and the views in the navigation path are also low aesthetics. In this paper, we present a brand-new navigation paradigm. Although our model is still trained on eye-fixations, our methodology can additionally enable the trained model to perceive the "meaningful" degree of the given panoramic video content. Outwardly, the proposed new approach is saliency-free, but inwardly, it is developed from saliency but biasing more to be "meaningful-driven"; thus, it can generate a navigation path with more appropriate content coverage. Besides, this paper is the first attempt to devise an unsupervised learning scheme to ensure all localized meaningful views in the navigation path have high aesthetics. Thus, the navigation path generated by our approach can also bring users an enjoyable watching experience. As a new topic in its infancy, we have devised a series of quantitative evaluation schemes, including objective verifications and subjective user studies. All these innovative attempts would have great potential to inspire and promote this research field in the near future.
Chenglizhao Chen, Guangxiao Ma, Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 WinDB: HMD-Free and Distortion-Free Panoptic Video Fixation Learning
abstract
To date, the widely adopted way to perform fixation collection in panoptic video is based on a head-mounted display (HMD), where users' fixations are collected while wearing a HMD to explore the given panoptic scene freely. However, this widely-used data collection method is insufficient for training deep models to accurately predict which regions in a given panoptic are most important when it contains intermittent salient events. The main reason is that there always exist "blind zooms" when using HMD to collect fixations since the users cannot keep spinning their heads to explore the entire panoptic scene all the time. Consequently, the collected fixations tend to be trapped in some local views, leaving the remaining areas to be the "blind zooms". Therefore, fixation data collected using HMD-based methods that accumulate local views cannot accurately represent the overall global importance - the main purpose of fixations - of complex panoptic scenes. To conquer, this paper introduces the auxiliary window with a dynamic blurring (WinDB) fixation collection approach for panoptic video, which doesn't need HMD and is able to well reflect the regional-wise importance degree. Using our WinDB approach, we have released a new PanopticVideo-300 dataset, containing 300 panoptic clips covering over 225 categories. Specifically, since using WinDB to collect fixations is blind zoom free, there exists frequent and intensive "fixation shifting" - a very special phenomenon that has long been overlooked by the previous research - in our new set. Thus, we present an effective fixation shifting network (FishNet) to conquer it. All these new fixation collection tool, dataset, and network could be very potential to open a new age for fixation-related research and applications in 360o environments.
Guotao Wang 0004, Chenglizhao Chen, Aimin Hao, Hong Qin 0001, Deng-Ping Fan
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Pixel is All You Need: Adversarial Spatio-Temporal Ensemble Active Learning for Salient Object Detection
abstract
Although weakly-supervised techniques can reduce the labeling effort, it is unclear whether a saliency model trained with weakly-supervised data (e.g., point annotation) can achieve the equivalent performance of its fully-supervised version. This paper attempts to answer this unexplored question by proving a hypothesis: there is a point-labeled dataset where saliency models trained on it can achieve equivalent performance when trained on the densely annotated dataset. To prove this conjecture, we proposed a novel yet effective adversarial spatio-temporal ensemble active learning. Our contributions are four-fold: 1) Our proposed adversarial attack triggering uncertainty can conquer the overconfidence of existing active learning methods and accurately locate these uncertain pixels. 2) Our proposed spatio-temporal ensemble strategy not only achieves outstanding performance but significantly reduces the model's computational cost. 3) Our proposed relationship-aware diversity sampling can conquer oversampling while boosting model performance. 4) We provide theoretical proof for the existence of such a point-labeled dataset. Experimental results show that our approach can find such a point-labeled dataset, where a saliency model trained on it obtained 98%-99% performance of its fully-supervised version with only ten annotated points per image.
Wei Wang 0169, Yacong Li, Fengmao Lv, Qing Xia 0002, Chenglizhao Chen, Aimin Hao, Shuo Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.8
2025 AttriDiffuser: Adversarially enhanced diffusion model for text-to-facial attribute image synthesis
Wenfeng Song, Zhongyong Ye, Xia Hou, Shuai Li 0001, Aimin Hao
Pattern Recognit.6
2025 Saliency-Aware Foveated Path Tracing for Virtual Reality Rendering
abstract
Foveated rendering reduces computational load by distributing resources based on the human visual system. This enables the implementation of ray tracing in virtual reality applications, where a high frame rate is essential to achieve visual immersion. However, traditional foveation methods based solely on eccentricity cannot adequately account for the complex behavior of visual attention. This is one of the main reasons that leads to lower perceived quality compared to non-foveated techniques. In this study, we introduce a novel rendering pipeline that incorporates ocular attention through the use of visual saliency. Based on foveation saliency, our approach facilitates the real-time production of high-quality images utilizing path tracing by distributing samples according to saliency metrics derived from geometric and historical data. To further augment image quality, an adaptive filtering process, aligned with the saliency metrics, is employed to reduce visible artifacts in non-foveal regions. Our experiments prove that this novel approach can demonstrate superior performance compared to previous methods, both in terms of quantitative metrics and perceived visual quality.
Yang Gao 0032, Wencan Li, Shiyu Liang, Aimin Hao, Xiaohui Tan
IEEE Trans. Vis. Comput. Graph.4
2025 Efficient Photon Beam Diffusion for Directional Subsurface Scattering
abstract
Real-time subsurface scattering techniques are widely used in translucent material rendering. Among advanced methods that rely on the bidirectional scattering-surface reflectance distribution function (BSSRDF), screen space algorithms exhibit limited translucency, while existing large-distance methods are inefficient and yield poor illumination details. To address these limitations for better large-distance scattering, we develop a novel algorithm by extending the photon beam diffusion (PBD) model within the light view and screen space. Unlike surface irradiance in prior methods, we incorporate the refracted beam in the medium into real-time scattering estimation, presenting a new consideration for photon beam utilization. Concretely, we store all photon beam samples in light view textures and utilize an adaptive sampling pattern for beam sample selection in large filtering kernel sizes. This can reduce the sample count based on surface attributes. In screen space, virtual sources are derived from samples to estimate PBD contributions, with an approximation that preserves boundary conditions. To avoid possible overestimation, we implement correction factors that scale contributions, effectively aligning our results with path-tracing references. Through these reformulations, our efficient PBD generates results closest to references among existing methods. The experiments accurately represent better front-face illumination details and backlit translucency effects, while significantly accelerating performance compared to previous large-distance methods.
Shiyu Liang, Yang Gao 0032, Chonghao Hu, Aimin Hao, Hong Qin 0001
IEEE Trans. Vis. Comput. Graph.4
2025 TalkingStyle: Personalized Speech-Driven 3D Facial Animation With Style Preservation
abstract
It is a challenging task to create realistic 3D avatars that accurately replicate individuals' speech and unique talking styles for speech-driven facial animation. Existing techniques have made remarkable progress but still struggle to achieve lifelike mimicry. This article proposes "TalkingStyle", a novel method to generate personalized talking avatars while retaining the talking style of the person. Our approach uses a set of audio and animation samples from an individual to create new facial animations that closely resemble their specific talking style, synchronized with speech. We disentangle the style codes from the motion patterns, allowing our method to associate a distinct identifier with each person. To manage each aspect effectively, we employ three separate encoders for style, speech, and motion, ensuring the preservation of the original style while maintaining consistent motion in our stylized talking avatars. Additionally, we propose a new style-conditioned transformer decoder, offering greater flexibility and control over the facial avatar styles. We comprehensively evaluate TalkingStyle through qualitative and quantitative assessments, as well as user studies demonstrating its superior realism and lip synchronization accuracy compared to current state-of-the-art methods.
Wenfeng Song, Xuan Wang 0024, Shi Zheng, Shuai Li 0001, Aimin Hao, Xia Hou
IEEE Trans. Vis. Comput. Graph.5
2025 Fluid Inverse Volumetric Modeling and Applications From Surface Motion
abstract
In this study, we devise a framework for volumetrically reconstructing fluid from observable, measurable free surface motion. Our innovative method amalgamates the benefits of deep learning and conventional simulation to preserve the guiding motion and temporal coherence of the reproduced fluid. We infer surface velocities by encoding and decoding spatiotemporal features of surface sequences, and a 3D CNN is used to generate the volumetric velocity field, which is then combined with 3D labels of obstacles and boundaries. Concurrently, we employ a network to estimate the fluid's physical properties. To progressively evolve the flow field over time, we input the reconstructed velocity field and estimated parameters into the physical simulator as the initial state. Our approach yields promising results for both synthetic fluid generated by different fluid solvers and captured real fluid. The developed framework naturally lends itself to a variety of graphics applications, such as 1) effective reproductions of fluid behaviors visually congruent with the observed surface motion, and 2) physics-guided re-editing of fluid scenes. Extensive experiments affirm that our novel method surpasses state-of-the-art approaches for 3D fluid inverse modeling and animation in graphics.
Xueguang Xie, Yang Gao 0032, Fei Hou 0001, Tianwei Cheng, Aimin Hao, Hong Qin 0001
IEEE Trans. Vis. Comput. Graph.5
2025 Real-time immersive haptic sculpting with elastoplastic virtual clay
Zhiyang Ji, Aimin Hao, Yang Gao 0032
Vis. Comput.3
2024 Weakly Supervised Multimodal Affordance Grounding for Egocentric Images
abstract
To enhance the interaction between intelligent systems and the environment, locating the affordance regions of objects is crucial. These regions correspond to specific areas that provide distinct functionalities. Humans often acquire the ability to identify these regions through action demonstrations and verbal instructions. In this paper, we present a novel multimodal framework that extracts affordance knowledge from exocentric images, which depict human-object interactions, as well as from accompanying textual descriptions that describe the performed actions. The extracted knowledge is then transferred to egocentric images. To achieve this goal, we propose the HOI-Transfer Module, which utilizes local perception to disentangle individual actions within exocentric images. This module effectively captures localized features and correlations between actions, leading to valuable affordance knowledge. Additionally, we introduce the Pixel-Text Fusion Module, which fuses affordance knowledge by identifying regions in egocentric images that bear resemblances to the textual features defining affordances. We employ a Weakly Supervised Multimodal Affordance (WSMA) learning approach, utilizing image-level labels for training. Through extensive experiments, we demonstrate the superiority of our proposed method in terms of evaluation metrics and visual results when compared to existing affordance grounding models. Furthermore, ablation experiments confirm the effectiveness of our approach. Code:https://github.com/xulingjing88/WSMA.
Lingjing Xu, Yang Gao 0032, Wenfeng Song, Aimin Hao
AAAI4
2024 Frequency-Guided Network for Low-contrast Staining-free Dental Plaque Segmentation
abstract
Traditional dental plaque detection relies on medical staining reagents and professional intervention. Deep learning-based automatic staining-free dental plaque segmentation provides an alternative for patients to perform plaque detection at home without staining reagents. However, existing methods still struggle with low-contrast visual features between unstained plaque and healthy teeth. To address this, we propose a Frequency-Guided Network (FGN) for low-contrast staining-free dental plaque segmentation. We observe that dental plaque tends to concentrate specifically near the junction between the teeth and the gingiva. This junction demonstrates abrupt changes in pixel values, indicating high-frequency regions in the image. In other words, dental plaque tends to appear near the high-frequency regions of oral endoscope images. Exploiting this characteristic, we employ a frequency-guided decoupling module to separate the image into high-frequency and low-frequency regions automatically and expand the high-frequency region to encompass nearby potential dental plaque. Then we supervise two regions individually to specifically focus on the expended high-frequency region for localizing nearby dental plaque. Additionally, we propose a high-to-low frequency multiple tasks framework. In the first phase, the network segments the teeth region, and then we input the teeth mask into the second phase. In the second stage, the teeth mask allows us to have a higher frequency at the junction between the teeth and gums, thereby enhancing the effectiveness of frequency-guided decoupling. Furthermore, FGN integrates a frequency-driven refinement module to enhance the guidance quality of the teeth mask for the second phase. Extensive evaluations of the oral endoscope dataset demonstrate that our method outperforms existing high-performance segmentation methods. User studies also confirm that our approach achieves superior results to experienced dentists. https://frequency-guided-network.github.io/
Yiming Jiang 0018, Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001
BIBM6
2024 SL-SFGR: Segmentation Learning Coupling Spatial-Frequency Structure Information Enhancement for Guiding Registration
abstract
The core of medical image registration lies in the alignment of corresponding structures. Hence, the effective construction of structure information in images is crucial for guiding registration. Some methods directly introduce explicit structure priors for assisting registration training, but obtaining high-precision structure priors is intrinsically challenging. As a segmentation model for obtaining structure priors, its capability is built upon the effective construction of structure information, thus its model learning can also be used to guide registration. However, existing registration-segmentation joint methods mostly only use segmentation results at the output level to constrain registration, neglecting the guidance of segmentation learning at the feature level for registration. Moreover, most existing methods only extract features in the image spatial domain for registration, overlooking the structure information in the frequency domain that is more easily captured to guide registration. To this end, this paper proposes an innovative registration method, namely Segmentation Learning Coupling Spatial-Frequency Structure Information Enhancement for Guiding Registration (SL-SFGR). Specifically, first, a semi-supervised segmentation learning network is constructed based on the registration deformation field to introduce structure features suitable for guiding registration. Second, an adaptive feature enhancement module is built in the spatial-frequency dual domain to further strengthen the inherent structure features. Finally, Dynamic Weight Average (DWA) is utilized for joint optimization of the model. The effectiveness of the proposed method has been verified on different brain MRI datasets. The related code is available at: https://github.com/goghfan/SL-SFGR/.
Yuanbo He, Shuai Li 0001, Aimin Hao, Desen Cao
BIBM4
2024 Semi-Supervised Medical Image Segmentation with Cross-View Consistency and Contrastive Learning
abstract
Medical image segmentation plays a crucial role in many clinical applications. To alleviate the dependency on massive annotations, semi-supervised learning has attracted increasing attention. However, these methods face significant intra-class and inter-class variation and do not fully utilize the critical multi-view information inherent in medical images. This study proposes a novel network, CV-Net, which integrates multi-view information for semi-supervised medical image segmentation. Concretely, the network is based on Mean-Teacher architecture which largely narrows the empirical distribution gap between labeled and unlabeled data. The proposed cross-view consistency regularization module incorporates a dual-branch attention architecture to integrate consistent semantics while focusing on details, enhancing feature extraction capabilities. The proposed bi-semantic contrastive learning module leverages limited labels and explore pseudo-labels to define semantically similar regions, enhancing the representation capacity. Experiments conducted on two datasets demonstrated the effectiveness of the proposed network. CV-Net showed significant improvements across four metrics, evident with both 5% and 10% labeled data. Specifically, with 5% labeled data, the mean Dice increased by 1.37%. Compared with previous state-of-the-art methods, CV-Net achieved the best results, notably reducing both intra-class and inter-class errors.
Kexuan Li, Jingjuan Liu, Xuehao Wang, Huadan Xue, Aimin Hao, Shuai Li 0001
BIBM7
2024 FaceCom: Towards High-fidelity 3D Facial Shape Completion via Optimization and Inpainting Guidance
abstract
We propose FaceCom, a method for 3D facial shape completion, which delivers high-fidelity results for incomplete facial inputs of arbitrary forms. Unlike end-to-end shape completion methods based on point clouds or voxels, our approach relies on a mesh-based generative network that is easy to optimize, enabling it to handle shape completion for irregular facial scans. We first train a shape generator on a mixed 3D facial dataset containing 2405 identities. Based on the incomplete facial input, we fit complete faces using an optimization approach under image inpainting guidance. The completion results are refined through a post-processing step. FaceCom demonstrates the ability to effectively and naturally complete facial scan data with varying missing regions and degrees of missing areas. Our method can be used in medical prosthetic fabrication and the registration of deficient scanning data. Our experimental results demonstrate that FaceCom achieves exceptional performance in fitting and shape completion tasks. The code is available at https://github.com/dragonylee/FaceCom.git.
Yinglong Li, Xiaogang Wang 0005, Qingzhao Qin, Yijiao Zhao, Aimin Hao
CVPR7
2024 Arbitrary Motion Style Transfer with Multi-Condition Motion Latent Diffusion Model
abstract
Computer animation's quest to bridge content and style has historically been a challenging venture, with previous efforts often leaning toward one at the expense of the other. This paper tackles the inherent challenge of content-style duality, ensuring a harmonious fusion where the core narrative of the content is both preserved and elevated through stylistic enhancements. We propose a novel Multi-condition Motion Latent Diffusion Model (MCM-LDM) for Arbitrary Motion Style Transfer (AMST). Our MCM-LDM significantly emphasizes preserving trajectories, recognizing their fundamental role in defining the essence and fluidity of motion content. Our MCM-LDM's cornerstone lies in its ability first to disentangle and then intricately weave together motion's tripartite components: motion trajectory, motion content, and motion style. The critical insight of MCM-LDM is to embed multiple conditions with distinct priorities. The content channel serves as the primary flow, guiding the overall structure and movement, while the trajectory and style channels act as auxiliary components and synchronize with the primary one dynamically. This mechanism ensures that multi-conditions can seamlessly integrate into the main flow, enhancing the overall animation without overshadowing the core content. Empirical evaluations underscore the model's proficiency in achieving fluid and authentic motion style transfers, setting a new benchmark in the realm of computer animation. The source code and model are available at https://github.com/XingliangJin/MCM-LDM.git.
Wenfeng Song, Xingliang Jin, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Xia Hou, Hong Qin 0001
CVPR5
2024 HOIAnimator: Generating Text-Prompt Human-Object Animations Using Novel Perceptive Diffusion Models
abstract
To date, the quest to rapidly and effectively produce human-object interaction (HOI) animations directly from textual descriptions stands at the forefront of computer vision research. The underlying challenge demands both a discriminating interpretation of language and a comprehen-sive physics-centric model supporting real-world dynamics. To ameliorate, this paper advocates HOIAnimator, a novel and interactive diffusion model with perception ability and also ingeniously crafted to revolutionize the animation of complex interactions from linguistic narratives. The effectiveness of our model is anchored in two ground-breaking innovations: (1) Our Perceptive Diffusion Models (PDM) brings together two types of models: one focused on hu-man movements and the other on objects. This combination allows for animations where humans and objects move in concert with each other, making the overall motion more realistic. Additionally, we propose a Perceptive Message Passing (PMP) mechanism to enhance the communication bridging the two models, ensuring that the animations are smooth and unified; (2) We devise an Interaction Contact Field (ICF), a sophisticated model that implicitly captures the essence of HOls. Beyond mere predictive contact points, the ICF assesses the proximity of human and object to their respective environment, informed by a probabilistic distribution of interactions learned throughout the denoising phase. Our comprehensive evaluation showcases HOlani-mator's superior ability to produce dynamic, context-aware animations that surpass existing benchmarks in text-driven animation synthesis.
Wenfeng Song, Shuai Li 0001, Yang Gao 0032, Aimin Hao, Xia Hau, Chenglizhao Chen, Hong Qin 0001
CVPR5
2024 Detail Enhancement for Free Surface LBM Using Adaptive Sizing of Coupled Particles
Qingyue Qu, Shaonan Zhu, Aimin Hao, Yang Gao 0032
ICXR4
2024 SIE-DepthNet: Semantic-Guided Monocular Depth Estimation for Dynamic Environment
Zilong Song, Yang Gao 0032, Sijia Dai, Shuai Li 0001, Aimin Hao, Shoulong Zhang
ICXR5
2024 ANFluid: Animate Natural Fluid Photos base on Physics-Aware Simulation and Dual-Flow Texture Learning
abstract
Generating photorealistic animations from a single still photo represents a significant advancement in multimedia editing and artistic creation. While existing AIGC methods have reached milestone successes, they often struggle with maintaining consistency with real-world physical laws, particularly in fluid dynamics. To address this issue, this paper introduces ANFluid, a physics solver and data-driven coupled framework that combines physics-aware simulation (PAS) and dual-flow texture learning (DFTL) to animate natural fluid photos effectively. The PAS component of ANFluid ensures that motion guides adhere to physical laws, and can be automatically tailored with specific numerical solver to meet the diversities of different fluid scenes. Concurrently, DFTL focuses on enhancing texture prediction. It employs bidirectional self-supervised optical flow estimation and multi-scale wrapping to strengthen dynamic relationships and elevate the overall animation quality. Notably, despite being built on a transformer architecture, the innovative encoder-decoder design in DFTL does not increase the parameter count but rather enhances inference efficiency. Extensive quantitative experiments have shown that our ANFluid surpasses most current methods on the Holynski and CLAW datasets. User studies further confirm that animations produced by ANFluid maintain better physical and content consistency with the real world and the original input, respectively. Moreover, ANFluid supports interactive editing during the simulation process, enriching the animation content and broadening its application potential.
Xiangcheng Zhai, Yingqi Jie, Xueguang Xie, Aimin Hao, Yang Gao 0032
ACM Multimedia4
2024 A Coupling Physics Model for Real-Time 4D Simulation of Cardiac Electromechanics
Jiahao Cui 0001, Shuai Li 0001, Aimin Hao
Comput. Aided Des.4
2024 Conditional room layout generation based on graph neural networks
Zhihan Yao, Jiahao Cui 0001, Shoulong Zhang, Shuai Li 0001, Aimin Hao
Comput. Graph.6
2024 CoupNeRF: Property-aware Neural Radiance Fields for Multi-Material Coupled Scenario Reconstruction
abstract
Abstract Neural Radiance Fields (NeRFs) have achieved significant recognition for their proficiency in scene reconstruction and rendering by utilizing neural networks to depict intricate volumetric environments. Despite considerable research dedicated to reconstructing physical scenes, rare works succeed in challenging scenarios involving dynamic, multi‐material objects. To alleviate, we introduce CoupNeRF, an efficient neural network architecture that is aware of multiple material properties. This architecture combines physically grounded continuum mechanics with NeRF, facilitating the identification of motion systems across a wide range of physical coupling scenarios. We first reconstruct specific‐material of objects within 3D physical fields to learn material parameters. Then, we develop a method to model the neighbouring particles, enhancing the learning process specifically in regions where material transitions occur. The effectiveness of CoupNeRF is demonstrated through extensive experiments, showcasing its proficiency in accurately coupling and identifying the behavior of complex physical scenes that span multiple physics domains.
Jin Li 0068, Yang Gao 0032, Wenfeng Song, Yacong Li, Shuai Li 0001, Aimin Hao, Hong Qin 0001
Comput. Graph. Forum6
2024 State of the Art in Efficient Translucent Material Rendering with BSSRDF
abstract
Abstract Sub‐surface scattering is always an important feature in translucent material rendering. When light travels through optically thick media, its transport within the medium can be approximated using diffusion theory, and is appropriately described by the bidirectional scattering‐surface reflectance distribution function (BSSRDF). BSSRDF methods rely on assumptions about object geometry and light distribution in the medium, which limits their applicability to general participating media problems. However, despite the high computational cost of path tracing, BSSRDF methods are often favoured due to their suitability for real‐time applications. We review these methods and discuss the most recent breakthroughs in this field. We begin by summarizing various BSSRDF models and then implement most of them in a 2D searchlight problem to demonstrate their differences. We focus on acceleration methods using BSSRDF, which we categorize into two primary groups: pre‐computation and texture methods. Then we go through some related topics, including applications and advanced areas where BSSRDF is used, as well as problems that are sometimes important yet are ignored in sub‐surface scattering estimation. In the end of this survey, we point out remaining constraints and challenges, which may motivate future work to facilitate sub‐surface scattering.
Shiyu Liang, Yang Gao 0032, Chonghao Hu, Aimin Hao, Lili Wang 0006, Hong Qin 0001
Comput. Graph. Forum5
2024 Dynamic ocean inverse modeling based on differentiable rendering
abstract
Learning and inferring underlying motion patterns of captured 2D scenes and then re-creating dynamic evolution consistent with the real-world natural phenomena have high appeal for graphics and animation. To bridge the technical gap between virtual and real environments, we focus on the inverse modeling and reconstruction of visually consistent and property-verifiable oceans, taking advantage of deep learning and differentiable physics to learn geometry and constitute waves in a self-supervised manner. First, we infer hierarchical geometry using two networks, which are optimized via the differentiable renderer. We extract wave components from the sequence of inferred geometry through a network equipped with a differentiable ocean model. Then, ocean dynamics can be evolved using the reconstructed wave components. Through extensive experiments, we verify that our new method yields satisfactory results for both geometry reconstruction and wave estimation. Moreover, the new framework has the inverse modeling potential to facilitate a host of graphics applications, such as the rapid production of physically accurate scene animation and editing guided by real ocean scenes.
Xueguang Xie, Yang Gao 0032, Fei Hou 0001, Aimin Hao, Hong Qin 0001
Comput. Vis. Media4
2024 Erratum to: Dynamic ocean inverse modeling based on differentiable rendering
abstract
The authors apologize for a hidden error in the article. It is that the images in Figs. 14(a) and 14(d) were mistakenly presented as left–right mirror images. The authors have flipped them to ensure that the figures now correspond correctly with others in the subfigures (b, c, e, f). The accurate version of Fig. 14 is provided as below.
Xueguang Xie, Yang Gao 0032, Fei Hou 0001, Aimin Hao, Hong Qin 0001
Comput. Vis. Media4
2024 Correction: Automatic Generation of 3D Scene Animation Based on Dynamic Knowledge Graphs and Contextual Encoding
Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001
Int. J. Comput. Vis.5
2024 Semantic and style based multiple reference learning for artistic and general image aesthetic assessment
Tengfei Shi, Chenglizhao Chen, Aimin Hao
Neurocomputing4
2024 Dynamic attention augmented graph network for video accident anticipation
Wenfeng Song, Shuai Li 0001, Tao Chang, Ke Xie 0005, Aimin Hao, Hong Qin 0001
Pattern Recognit.5
2024 Joints-Centered Spatial-Temporal Features Fused Skeleton Convolution Network for Action Recognition
abstract
Skeleton-based action recognition is crucial for natural human-computer interaction, dynamic behavior analysis, and behavior surveillance. The key challenge is to effectively capture the intrinsic local-global clues of the activity. However, it remains challenging to efficiently leverage multidimensional information related to joints' local visual appearances, global spatial relationships, and coherent temporal cues. To address this challenge, we propose a joints-centered spatial-temporal feature-fused framework for action recognition, which exploits skeleton-based graph diffusion and convolution. Specifically, we employ Partial Differential Equation (PDE) based skeleton graph diffusion to automatically activate and diffuse the salient appearance features of joints. This approach simultaneously integrates the joints' appearance clues and their hierarchical relationships at both the super-pixel level and structure level. The diffused appearance-related features of the joints are further fused with skeleton-related spatial-temporal features, and the resulting fused features are fed into a skeleton convolution network for action recognition. Our method was extensively evaluated on two public datasets (NTU-RGBD and UWA3D), and the results demonstrate the improved accuracy and effectiveness of our approach. Our code will be public.
Wenfeng Song, Tangli Chu, Shuai Li 0001, Nannan Li 0002, Aimin Hao, Hong Qin 0001
IEEE Trans. Multim.5
2024 CenterFormer: A Novel Cluster Center Enhanced Transformer for Unconstrained Dental Plaque Segmentation
abstract
Dental plaque segmentation is crucial for maintaining oral health. However, accurately segmenting dental plaque in unconstrained environments can be challenging due to its low contrast and high variability in appearance. While existing transformer-based networks rely on attention mechanisms for each pixel, they do not take into account the relationships between neighboring pixels. Consequently, feature extraction is limited, making it difficult to achieve accurate segmentation of low-contrast images. To address this issue, we propose a simple yet efficient cluster center transformer that improves dental plaque segmentation by clustering image pixels based on multiple levels of feature maps' intensity and texture information. By grouping similar pixels into regions, the proposed method enables the transformers to focus on the local contour and edge around the teeth regions, adapting to the low contrast and high variability of plaque appearance, leading to more accurate and efficient segmentation of dental plaque in dental images. Additionally, we designed Multiple Granularity Perceptions using a pyramid fusion mechanism to capture multiple scales of vision features, thereby enhancing the low-contrast vision features. The proposed method can benefit the dental diagnosis and treatment planning process by improving the accuracy and efficiency of dental plaque segmentation. Our proposed method achieved state-of-the-art results on the dental plaque dataset (Li et al., 2020), with intersection over union (IoU) of 60.91% and pixel accuracy (PA) of 76.81%, all of which were the highest among all methods, demonstrating its effectiveness in plaque segmentation in unconstrained environments.
Wenfeng Song, Xuan Wang 0024, Shuai Li 0001, Aimin Hao
IEEE Trans. Multim.6
2024 Improving Image Aesthetic Assessment via Multiple Image Joint Learning
abstract
Image Aesthetic Assessment (IAA) is an emerging paradigm that predicts aesthetic score as the popular aesthetic taste for an image. Previous IAA approaches take a single image as input to predict the aesthetic score of the image. However, we discover that most existing IAA methods fail dramatically to predict the images with a large variance of aesthetic voting distribution. Motivated by the practice that people consider similar experiences to improve the consistence of the voting result, we present a novel Multiple Image joint Learning Network (MILNet) to mimic this natural process. Our novelty is mainly three-fold: (a) Semantic-based retrieval method that constructs aesthetic similarity (the similarity of aesthetic attribution) to select reference images; (b) Graph network reasoning that initializes and updates the weight of intrinsic relationships among multiple images; (c) Adaptive Earth Mover’s Distance (AdaEMD) loss function that adjusts weight for easy and hard instances to mitigate unbalanced distribution of aesthetic datasets. Our evaluation with the benchmark AVA and TAD datasets demonstrates that the proposed MILNet outperforms state-of-the-art IAA methods. The code is available at https://github.com/flyingbird93/MILNet .
Tengfei Shi, Chenglizhao Chen, Aimin Hao, Yuming Fang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2024 MPMNet: A Data-Driven MPM Framework for Dynamic Fluid-Solid Interaction
abstract
High-accuracy, high-efficiency physics-based fluid-solid interaction is essential for reality modeling and computer animation in online games or real-time Virtual Reality (VR) systems. However, the large-scale simulation of incompressible fluid and its interaction with the surrounding solid environment is either time-consuming or suffering from the reduced time/space resolution due to the complicated iterative nature pertinent to numerical computations of involved Partial Differential Equations (PDEs). In recent years, we have witnessed significant growth in exploring a different, alternative data-driven approach to addressing some of the existing technical challenges in conventional model-centric graphics and animation methods. This article showcases some of our exploratory efforts in this direction. One technical concern of our research is to address the central key challenge of how to best construct the numerical solver effectively and how to best integrate spatiotemporal/dimensional neural networks with the available MPM's pressure solvers. In particular, we devise the MPMNet, a hybrid data-driven framework supporting the popular and powerful MPM, to combine the comprehensive properties of MPM in numerically handling physical behaviors ranging from fluid to deformable solids and the high efficiency of data-driven models. At the architectural level, our MPMNet comprises three primary components: A data processing module to describe the physical properties by way of the input fields; A deep neural network group to learn the spatiotemporal features; And an iterative refinement process to continue to reduce possible numerical errors. The goal of these special technical developments is to aim at involved numerical acceleration while preserving physical accuracy, realizing efficient and accurate fluid-solid interactions in a data-driven fashion. The extensive experimental results verify that our MPMNet can tremendously speed up the computation compared with the popular numerical methods as the complexity of interaction scenes increases while better retaining the numerical accuracy.
Jin Li 0068, Yang Gao 0032, Ju Dai, Shuai Li 0001, Aimin Hao, Hong Qin 0001
IEEE Trans. Vis. Comput. Graph.5
2024 A Unified Particle-Based Solver for Non-Newtonian Behaviors Simulation
abstract
In this article, we present a unified framework to simulate non-Newtonian behaviors. We combine viscous and elasto-plastic stress into a unified particle solver to achieve various non-Newtonian behaviors ranging from fluid-like to solid-like. Our constitutive model is based on a Generalized Maxwell model, which incorporates viscosity, elasticity and plasticity in one non-linear framework by a unified way. On the one hand, taking advantage of the viscous term, we construct a series of strain-rate dependent models for classical non-Newtonian behaviors such as shear-thickening, shear-thinning, Bingham plastic, etc. On the other hand, benefiting from the elasto-plastic model, we empower our framework with the ability to simulate solid-like non-Newtonian behaviors, i.e., visco-elasticity/plasticity. In addition, we enrich our method with a heat diffusion model to make our method flexible in simulating phase change. Through sufficient experiments, we demonstrate a wide range of non-Newtonian behaviors ranging from viscous fluid to deformable objects. We believe this non-Newtonian model will enhance the realism of physically-based animation, which has great potential for computer graphics.
Yang Gao 0032, Tianwei Cheng, Shuai Li 0001, Aimin Hao, Hong Qin 0001
IEEE Trans. Vis. Comput. Graph.6
2024 Expressive 3D Facial Animation Generation Based on Local-to-Global Latent Diffusion
abstract
3D Facial animations, crucial to augmented and mixed reality digital media, have evolved from mere aesthetic elements to potent storytelling media. Despite considerable progress in facial animation of neutral emotions, existing methods still struggle to capture the authenticity of emotions. This paper introduces a novel approach to capture fine facial expressions and generate facial animations using audio synchronization. Our method consists of two key components: First, the Local-to-global Latent Diffusion Model (LG-LDM) tailored for authentic facial expressions, which can integrate audio, time step, facial expressions, and other conditions towards possible encoding of emotionally rich yet latent features in response to possibly noisy raw audio signals. The core of LG-LDM is our carefully designed Facial Denoiser Model (FDM) for aligning the local-to-global animation feature with audio. Second, we redesign an Emotion-centric Vector Quantized-Variational AutoEncoder framework (EVQ-VAE) to finely decode the subtle differences under different emotions and reconstruct the final 3D facial geometry. Our work significantly contributes to the key challenges of emotionally realistic 3D facial animation for audio synchronization and enhances the immersive experience and emotional depth in augmented and mixed reality applications. We provide a reproducibility kit including our code, dataset, and detailed instructions for running the experiments. This kit is available at https://github.com/wangxuanx/Face-Diffusion-Model.
Wenfeng Song, Xuan Wang 0024, Yiming Jiang 0018, Shuai Li 0001, Aimin Hao, Xia Hou, Hong Qin 0001
IEEE Trans. Vis. Comput. Graph.5
2024 Interactive Virtual Ankle Movement Controlled by Wrist sEMG Improves Motor Imagery: An Exploratory Study
abstract
Virtual reality (VR) techniques can significantly enhance motor imagery training by creating a strong illusion of action for central sensory stimulation. In this article, we establish a precedent by using surface electromyography (sEMG) of contralateral wrist movement to trigger virtual ankle movement through an improved data-driven approach with a continuous sEMG signal for fast and accurate intention recognition. Our developed VR interactive system can provide feedback training for stroke patients in the early stages, even if there is no active ankle movement. Our objectives are to evaluate: 1) the effects of VR immersion mode on body illusion, kinesthetic illusion, and motor imagery performance in stroke patients; 2) the effects of motivation and attention when utilizing wrist sEMG as a trigger signal for virtual ankle motion; 3) the acute effects on motor function in stroke patients. Through a series of well-designed experiments, we have found that, compared to the 2D condition, VR significantly increases the degree of kinesthetic illusion and body ownership of the patients, and improves their motor imagery performance and motor memory. When compared to conditions without feedback, using contralateral wrist sEMG signals as trigger signals for virtual ankle movement enhances patients' sustained attention and motivation during repetitive tasks. Furthermore, the combination of VR and feedback has an acute impact on motor function. Our exploratory study suggests that the sEMG-based immersive virtual interactive feedback provides an effective option for active rehabilitation training for severe hemiplegia patients in the early stages, with great potential for clinical application.
Yanqing Xiao, Hongming Bai, Yang Gao 0032, Ben Hu, XiaoE Cai, Jiasheng Rao, Aimin Hao
IEEE Trans. Vis. Comput. Graph.9
2024 Efficient frictional contacts for soft body dynamics via ADMM
Siyan Zhu, Xiao Zhai, Aimin Hao, JunJun Pan
Vis. Comput.5
2023 Pixel Is All You Need: Adversarial Trajectory-Ensemble Active Learning for Salient Object Detection
abstract
Although weakly-supervised techniques can reduce the labeling effort, it is unclear whether a saliency model trained with weakly-supervised data (e.g., point annotation) can achieve the equivalent performance of its fully-supervised version. This paper attempts to answer this unexplored question by proving a hypothesis: there is a point-labeled dataset where saliency models trained on it can achieve equivalent performance when trained on the densely annotated dataset. To prove this conjecture, we proposed a novel yet effective adversarial trajectory-ensemble active learning (ATAL). Our contributions are three-fold: 1) Our proposed adversarial attack triggering uncertainty can conquer the overconfidence of existing active learning methods and accurately locate these uncertain pixels. 2) Our proposed trajectory-ensemble uncertainty estimation method maintains the advantages of the ensemble networks while significantly reducing the computational cost. 3) Our proposed relationship-aware diversity sampling algorithm can conquer oversampling while boosting performance. Experimental results show that our ATAL can find such a point-labeled dataset, where a saliency model trained on it obtained 97%-99% performance of its fully-supervised version with only 10 annotated points per image.
Wei Wang 0169, Qing Xia 0002, Chenglizhao Chen, Aimin Hao, Shuo Li 0001
AAAI6
2023 Diffusing Coupling High-Frequency-Purifying Structure Feature Extraction for Brain Multimodal Registration
abstract
The core of medical image registration is the alignment of corresponding structures. However, in multimodal image registration, substantial differences in appearance (intensity distribution) of the images often compel the registration model to prioritize intensity information over structure information, resulting in low accuracy of registration. Therefore, the disentangling structure information from intensity information is vital to improve the registration effectiveness. To this end, we propose a diffusing coupling high-frequency-purifying structure feature extraction for brain multimodal registration. Specifically, the denoising diffusion probabilistic models (DDPM) is firstly utilized to extract complete feature information from images. Then, the discrete cosine transform (DCT) is applied to purify high-frequency structure information from the complete feature information for registration. Furthermore, structure consistency constraint (SCC) is introduced based on purified structure information to emphasize the core position of the structure in registration. Through comprehensive comparisons with traditional and learning-based methods on the multimodal brain MRI dataset, our method demonstrates superior accuracy and stability in brain multimodal registration. Our code is available at https://github.com/goghfan/DDNet.
Yuanbo He, Shuai Li 0001, Aimin Hao, Desen Cao
BIBM4
2023 RGB and LUT based Cross Attention Network for Image Enhancement
Tengfei Shi, Chenglizhao Chen, Yuanbo He, Wenfeng Song, Aimin Hao
BMVC5
2023 Propose-and-Complete: Auto-regressive Semantic Group Generation for Personalized Scene Synthesis
Shoulong Zhang, Shuai Li 0001, Xinwei Huang, Wenchong Xu, Aimin Hao, Hong Qin 0001
BMVC5
2023 Sequential Texts Driven Cohesive Motions Synthesis with Natural Transitions
abstract
The intelligent synthesis/generation of daily-life motion sequences is fundamental and urgently needed for many VR/metaverse-related applications. However, existing approaches commonly focus on monotonic motion generation (e.g., walking, jumping, etc.) based on single instruction-like text, which is still not intelligent enough and can’t meet practical demands. To this end, we propose a cohesive human motion sequence synthesis framework based on free-form sequential texts while ensuring semantic connection and natural transitions between adjacent motions. At the technical level, we explore the local-to-global semantic features of previous and current texts to extract relevant information. This information is used to guide the framework in understanding the semantics of the current moment. Moreover, we propose learnable tokens to adaptively learn the influence range of the previous motions towards natural transitions. These tokens can be trained to encode the relevant information into well-designed transition loss. To demonstrate the efficacy of our method, we conduct extensive experiments and comprehensive evaluations on the public dataset as well as a new dataset produced by us. All the experiments confirm that our method outperforms the state-of-the-art methods in terms of semantic matching, realism, and transition fluency. Our project is public available. https://druthrie.github.io/sequential-texts-to-motion/
Shuai Li 0001, Sisi Zhuang, Wenfeng Song, Hejia Chen, Aimin Hao
ICCV6
2023 Joint Probability Distribution Regression for Image Cropping
abstract
Image cropping aims at locating a candidate (rectangle region) with the highest aesthetic quality in professional photography. One solution of the previous methods is to generate a large number of candidates and then filter them, which leads to low efficiency. Another idea directly regresses the candidate coordinates to speed up but ignores the aesthetic subjectivity of the candidate’s evaluation, limiting the model’s performance. In this paper, we present an Aesthetic and Composition joint Probability Distribution regression Network (ACPD-Net) to explicitly investigate the process of generating the candidate with a joint probability distribution paradigm to improve the performance of cropping results in an efficient way. The joint probability distribution paradigm between location and size branch can identify the subjective aesthetic region and satisfy the objective composition rules in an end-to-end manner. Our method has been tested on the FCDB and FLMS datasets, which shows the superiority of ACPD-Net. The code is available at https://github.com/flyingbird93/ACPD-Net.
Tengfei Shi, Chenglizhao Chen, Yuanbo He, Wenfeng Song, Aimin Hao
ICIP5
2023 PhyVR: Physics-based Multi-material and Free-hand Interaction in VR
abstract
The realistic interaction with physical phenomena is a crucial aspect of human-computer interaction (HCI) in virtual reality (VR). However, the real-time performance of physical simulation, interactive computation, and rendering is the bottleneck of physics-based VR HCI. To address these challenges, we propose a novel physics-oriented framework for multi-material objects and free-hand interaction, termed PhyVR. This framework enables users to interact with diverse virtual phenomena dynamically. At the algorithm level, we develop a unified particle system to describe both the virtual multi-materials and the user’s avatar for the efficiency issue, optimize collision detection, and accelerate the HCI algorithms with a variable fine-coarse particle sampling scheme. At the rendering level, we introduce a hybrid particle-grid anisotropic algorithm for surface reconstruction, enabling real-time and visually convincing fluid rendering. Comprehensive experiments and user studies demonstrate that our framework effectively captures various physical interaction phenomena, providing an enhanced user experience and paving the way for expanding VR-related HCI applications.
Hanchen Deng, Jin Li 0068, Yang Gao 0032, Xiaohui Liang 0001, Aimin Hao
ISMAR6
2023 Modality Profile - A New Critical Aspect to be Considered When Generating RGB-D Salient Object Detection Training Set
abstract
It is widely acknowledged that selecting appropriate training data is crucial for obtaining good results in real-world testing, more so than utilizing complex network architectures. However, in the field of RGB-D SOD research, researchers have primarily focused on enhancing network architectures and have given less consideration to the choice of training and testing datasets, which may not translate well in practical applications. This paper aims to address an existing issue - how can we automatically generate a data-driven RGB-D SOD training dataset? We propose that in addition to scene similarity, the concept of "modality profile'' should be taken into account. The term "modality profile'' refers to the complementary status of modalities within a given dataset. A training dataset with a modality profile similar to the test dataset can significantly improve performance. To address this, we present a viable solution for automatically generating a training dataset with any desired modality profile in a weakly supervised manner. Our method also provides high-quality pseudo-GTs for all RGB-D images obtained from the web, making it suitable for training RGB-D SOD models. Extensive quantitative evaluations demonstrate the significance of the proposed "modality profile'' and confirm the superiority of the newly constructed training set guided by our "modality profile''. All codes, datasets, and results are available at this link.
Xuehao Wang, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001
ACM Multimedia4
2023 ZetaDesign: an end-to-end deep learning method for protein sequence design and side-chain packing
abstract
Computational protein design has been demonstrated to be the most powerful tool in the last few years among protein designing and repacking tasks. In practice, these two tasks are strongly related but often treated separately. Besides, state-of-the-art deep-learning-based methods cannot provide interpretability from an energy perspective, affecting the accuracy of the design. Here we propose a new systematic approach, including both a posterior probability and a joint probability parts, to solve the two essential questions once for all. This approach takes the physicochemical property of amino acids into consideration and uses the joint probability model to ensure the convergence between structure and amino acid type. Our results demonstrated that this method could generate feasible, high-confidence sequences with low-energy side conformations. The designed sequences can fold into target structures with high confidence and maintain relatively stable biochemical properties. The side chain conformation has a significantly lower energy landscape without delegating to a rotamer library or performing the expensive conformational searches. Overall, we propose an end-to-end method that combines the advantages of both deep learning and energy-based methods. The design results of this model demonstrate high efficiency, and precision, as well as a low energy state and good interpretability.
Junyu Yan, Shuai Li 0001, Aimin Hao, Qinping Zhao
Briefings Bioinform.4
2023 Analyzing part functionality via multi-modal latent space embedding and interweaving
Jiahao Cui 0001, Shuai Li 0001, Fei Hou 0001, Aimin Hao, Hong Qin 0001
Comput. Graph.4
2023 Automatic Generation of 3D Scene Animation Based on Dynamic Knowledge Graphs and Contextual Encoding
Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001
Int. J. Comput. Vis.5
2023 Graph Diffusion Convolutional Network for Skeleton Based Semantic Recognition of Two-Person Actions
abstract
Graph Convolutional Networks (GCNs) have successfully boosted skeleton-based human action recognition. However, existing GCN-based methods mostly cast the problem as separated person's action recognition while ignoring the interaction between the action initiator and the action responder, especially for the fundamental two-person interactive action recognition. It is still challenging to effectively take into account the intrinsic local-global clues of the two-person activity. Additionally, message passing in GCN depends on adjacency matrix, but skeleton-based human action recognition methods tend to calculate the adjacency matrix with the fixed natural skeleton connectivity. It means that messages can only travel along a fixed path at different layers of the network or in different actions, which greatly reduces the flexibility of the network. To this end, we propose a novel graph diffusion convolutional network for skeleton based semantic recognition of two-person actions by embedding the graph diffusion into GCNs. At technical fronts, we dynamically construct the adjacency matrix based on practical action information, so that we can guide the message propagation in a more meaningful way. Simultaneously, we introduce the frame importance calculation module to conduct dynamic convolution, so that we can avoid the negative effect caused by the traditional convolution, wherein the shared weights may fail to capture key frames or be affected by noisy frames. Besides, we comprehensively leverage the multidimensional features related to joints' local visual appearances, global spatial relationship and temporal coherency, and for different features, different metrics are designed to measure the similarity underlying the corresponding real physical law of the motions. Moreover, extensive experiments and comprehensive evaluations on four public large-scale datasets (NTU-RGB+D 60, NTU-RGB+D 120, Kinetics-Skeleton 400, and SBU-Interaction) demonstrate that our method outperforms the state-of-the-art methods.
Shuai Li 0001, Xinxue He, Wenfeng Song, Aimin Hao, Hong Qin 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 SC-GAN: Subspace Clustering based GAN for Automatic Expression Manipulation
Shuai Li 0001, Wenfeng Song, Aimin Hao, Hong Qin 0001
Pattern Recognit.5
2023 An Intelligent Virtual Standard Patient for Medical Students Training Based on Oral Knowledge Graph
abstract
Virtual standard patient (VSP) is in high demand for medical students' diagnosis ability training in an efficient manner. Different from the traditional conversation system in medical dialogue generation, VSP needs a novel conversation paradigm to act as the patient instead of the doctor. However, existing conversation techniques still have limited ability in terms of generation of symptoms exhibited by patients with the personalized and knowledge-centered expressions. To alleviate these problems, we propose to construct a novel oral knowledge graph, which sufficiently provides medical clues of the certain disease. Accordingly, the VSP could accurately interact with the dentists for their underlying intention and express the symptoms characters in a natural style. To efficiently retrieve the related disease clues, the symptoms descriptions of the oral diseases are encoded into the oral knowledge graph, which could well organize the disease-centered symptom entities and speaking styles. Moreover, to transfer the common sense knowledge from existing large scale of medical knowledge graph to the specific oral knowledge graph, a coupled pre-trained Bert models is further designed to learn the related medical knowledge from coarse-level to fine-level hierarchically. Finally, a series of well-designed personalized templates are proposed to generate plausible and realistic answers in condition of the certain disease. We also conduct extensive user studies to demonstrate that the VSP satisfies the medical students' diagnosis practice requirement in terms of naturalness, realism, and topic relevance.
Wenfeng Song, Xia Hou, Shuai Li 0001, Chenglizhao Chen, Danyang Gao, Xian'e Wang, Yuzhe Sun, Jianxia Hou, Aimin Hao
IEEE Trans. Multim.9
2023 FineStyle: Semantic-Aware Fine-Grained Motion Style Transfer with Dual Interactive-Flow Fusion
abstract
We present FineStyle, a novel framework for motion style transfer that generates expressive human animations with specific styles for virtual reality and vision fields. It incorporates semantic awareness, which improves motion representation and allows for precise and stylish animation generation. Existing methods for motion style transfer have all failed to consider the semantic meaning behind the motion, resulting in limited controls over the generated human animations. To improve, FineStyle introduces a new cross-modality fusion module called Dual Interactive-Flow Fusion (DIFF). As the first attempt, DIFF integrates motion style features and semantic flows, producing semantic-aware style codes for fine-grained motion style transfer. FineStyle uses an innovative two-stage semantic guidance approach that leverages semantic clues to enhance the discriminative power of both semantic and style features. At an early stage, a semantic-guided encoder introduces distinct semantic clues into the style flow. Then, at a fine stage, both flows are further fused interactively, selecting the matched and critical clues from both flows. Extensive experiments demonstrate that FineStyle outperforms state-of-the-art methods in visual quality and controllability. By considering the semantic meaning behind motion style patterns, FineStyle allows for more precise control over motion styles. Source code and model are available on https://github.com/XingliangJin/Fine-Style.git.
Wenfeng Song, Xingliang Jin, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Xia Hou
IEEE Trans. Vis. Comput. Graph.5
2022 Person Re-Identification in Panoramic Views Based on Bayesian Transformers
abstract
The panoramic view cameras offer more broad perspectives and continuous information for person re-identification (ReID). However, the panoramic-view videos suffer from objects distortion and bring more occlusion due to the fixed or moving capture points. This paper proposes a novel Bayesian Transformer Network (BTN) to adaptively capture the occlusion clues as Bayesian prior to guide the discriminative pedestrian-related feature extraction in the high-occlusion scenes. The Bayesian prior is built via a pre-trained CNN, which could recognize different occluded scenarios based on the severeness of noisy backgrounds. Moreover, to fully explore the occlusion prior, we propose to embed the semantic labels into a well-designed transformer network. By fostering the collaborative occlusion clues between the person and background, our method could achieve outstanding performance on both public benchmarks and panoramic view videos, which verifies the advantages of our BTN framework over existing methods.
Wenfeng Song, Yang Gao 0032, Aimin Hao, Xia Hou
ICIP6
2022 Synthetic Data Supervised Salient Object Detection
abstract
Although deep salient object detection (SOD) has achieved remarkable progress, deep SOD models are extremely data-hungry, requiring large-scale pixel-wise annotations to deliver such promising results. In this paper, we propose a novel yet effective method for SOD, coined SODGAN, which can generate infinite high-quality image-mask pairs requiring only a few labeled data, and these synthesized pairs can replace the human-labeled DUTS-TR to train any off-the-shelf SOD model. Its contribution is three-fold. 1) Our proposed diffusion embedding network can address the manifold mismatch and is tractable for the latent code generation, better matching with the ImageNet latent space. 2) For the first time, our proposed few-shot saliency mask generator can synthesize infinite accurate image synchronized saliency masks with a few labeled data. 3) Our proposed quality-aware discriminator can select highquality synthesized image-mask pairs from noisy synthetic data pool, improving the quality of synthetic data. For the first time, our SODGAN tackles SOD with synthetic data directly generated from the generative model, which opens up a new research paradigm for SOD. Extensive experimental results show that the saliency model trained on synthetic data can achieve $98.4%$ F-measure of the saliency model trained on the DUTS-TR. Moreover, our approach achieves a new SOTA performance in semi/weakly-supervised methods, and even outperforms several fully-supervised SOTA methods. Code is available at https://github.com/wuzhenyubuaa/SODGAN
Wei Wang 0169, Tengfei Shi, Chenglizhao Chen, Aimin Hao, Shuo Li 0001
ACM Multimedia6
2022 Distribution-motivated 3D Style Characterization Based on Latent Feature Decomposition
Xinwei Huang, Shuai Li 0001, Shoulong Zhang, Aimin Hao, Hong Qin 0001
Comput. Aided Des.4
2022 Erratum to: Self-adjustable hyper-graphs for video pose estimation based on spatial-temporal subspace construction
Jizhou Ma, Shuai Li 0001, Hong Qin 0001, Aimin Hao, Qinping Zhao
Sci. China Inf. Sci.4
2022 Self-adjustable hyper-graphs for video pose estimation based on spatial-temporal subspace construction
Jizhou Ma, Shuai Li 0001, Hong Qin 0001, Aimin Hao, Qinping Zhao
Sci. China Inf. Sci.4
2022 Automatic image matting and fusing for portrait synthesis
Zhike Yi, Wenfeng Song, Shuai Li 0001, Aimin Hao
Sci. China Inf. Sci.4
2022 Multi-scale and multi-level shape descriptor learning via a hybrid fusion network
Xinwei Huang, Nannan Li 0002, Qing Xia 0002, Shuai Li 0001, Aimin Hao, Hong Qin 0001
Graph. Model.5
2022 Recursive multi-model complementary deep fusion for robust salient object detection via parallel sub-networks
Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001
Pattern Recognit.4
2022 Salient Object Detection via Dynamic Scale Routing
abstract
Recent research advances in salient object detection (SOD) could largely be attributed to ever-stronger multi-scale feature representation empowered by the deep learning technologies. The existing SOD deep models extract multi-scale features via the off-the-shelf encoders and combine them smartly via various delicate decoders. However, the kernel sizes in this commonly-used thread are usually "fixed". In our new experiments, we have observed that kernels of small size are preferable in scenarios containing tiny salient objects. In contrast, large kernel sizes could perform better for images with large salient objects. Inspired by this observation, we advocate the "dynamic" scale routing (as a brand-new idea) in this paper. It will result in a generic plug-in that could directly fit the existing feature backbone. This paper's key technical innovations are two-fold. First, instead of using the vanilla convolution with fixed kernel sizes for the encoder design, we propose the dynamic pyramid convolution (DPConv), which dynamically selects the best-suited kernel sizes w.r.t. the given input. Second, we provide a self-adaptive bidirectional decoder design to accommodate the DPConv-based encoder best. The most significant highlight is its capability of routing between feature scales and their dynamic collection, making the inference process scale-aware. As a result, this paper continues to enhance the current SOTA performance. Both the code and dataset are publicly available at https://github.com/wuzhenyubuaa/DPNet.
Shuai Li 0001, Chenglizhao Chen, Hong Qin 0001, Aimin Hao
IEEE Trans. Image Process.5
2022 Automatic Dental Plaque Segmentation Based on Local-to-Global Features Fused Self-Attention Network
abstract
The accurate detection of dental plaque at an early stage will definitely prevent periodontal diseases and dental caries. However, it remains difficult for the current dental examination to accurately recognize dental plaque without using medical dyeing reagent due to the low contrast between dental plaque and healthy teeth. To combat this problem, this paper proposes a novel network enhanced by a self-attention module for intelligent dental plaque segmentation. The key motivation is to directly utilize oral endoscope images (bypassing the need for dyeing reagent) and get accurate pixel-level dental plaque segmentation results. The algorithm needs to conduct self-attention at the super-pixel level and fuse the super-pixels' local-to-global features. Our newly-designed network architecture will afford the simultaneous fusion of multiple-scale complementary information guided by the powerful deep learning paradigm. The critical fused information includes the statistical distribution of the plaques color, the heat kernel signature (HKS) based local-to-global structure relationship, and the circle-LBP based local texture pattern in the nearby regions centering around the plaque area. To further refine the fuzed multiple-scale features, we devise an attention module based on CNN, which could focalize the regions of interest in plaque more easily, especially for many challenging cases. Extensive experiments and comprehensive evaluations confirm that, for a small-scale training dataset, our method could outperform the state-of-the-art methods. Meanwhile, the user studies verify the claim that our method is more accurate than conventional dental practice conducted by experienced dentists.
Shuai Li 0001, Zhennan Pang, Wenfeng Song, Aimin Hao, Hong Qin 0001
IEEE J. Biomed. Health Informatics5
2022 Deeper Look at Image Salient Object Detection: Bi-Stream Network With a Small Training Dataset
abstract
Compared with the conventional hand-crafted approaches, the deep learning based ISOD (image salient object detection) models have achieved tremendous performance improvements by training exquisitely crafted fancy networks over large-scale training sets. However, do we really need large-scale training set for ISOD? In this article, we provide a deeper insight into the interrelationship between the ISOD performance and the training data. To alleviate the conventional demands for large-scale training data, we provide a feasible way to construct a novel small-scale training set, which only contains 4 K images. To take full advantage of this new set, we propose a novel bi-stream network consisting of two different feature backbones. Benefit from the proposed gate control unit, this bi-stream network is able to achieve complementary fusion status for its subbranches. To our best knowledge, this is the first attempt to use a small-scale training set to compete with other large-scale ones; nevertheless, our method can still achieve the leading SOTA performance on all tested benchmark datasets. Both the code and dataset are publicly available athttps://github.com/wuzhenyubuaa/TSNet.
Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001
IEEE Trans. Multim.4
2022 Neurophysiological and Subjective Analysis of VR Emotion Induction Paradigm
abstract
The ecological validity of emotion-inducing scenarios is essential for emotion research. In contrast to the classical passive induction paradigm, immersive VR fully engages the psychological and physiological components of the subject, which is considered an ecologically valid paradigm for studying emotion. Several studies investigate the emotional responses to different VR tasks or games using subjective scales. However, little research regards VR as an eliciting material, especially when systematically analyzing emotional processes in VR from a neurophysiological perspective. To fill this gap and scientifically evaluate VR's ability to be used as an active method for emotion elicitation, we investigate the dynamic relationship between explicit information (subjective evaluations) and implicit information (objective neurophysiological data). A total of 28 participants are enlisted to watch eight VR videos while their SAM/IPQ scores and EEG data are recorded simultaneously. In ecologically valid scenarios, the subjective results demonstrate that VR has significant advantages for evoking emotion in arousal-valence. This conclusion is backed by our examination of objective neurophysiological evidence that VR videos effectively induce high-arousal emotions. In addition, we obtain features of critical channels and frequency oscillations associated with emotional valence, thereby validating previous research in more lifelike circumstances. In particular, we discover hemispheric asymmetry in the occipital region under high and low emotional arousal, which adds to our understanding of neural features and the dynamics of emotional arousal. As a result, we successfully integrate EEG and VR to demonstrate that VR is more pragmatic for evoking natural feelings and is beneficial for emotional research. Our research has set a precedent for new methodologies of using VR induction paradigms to acquire a more reliable explanation of affective computing.
JunJun Pan, Yang Gao 0032, Yang Shen 0009, Ju Dai, Aimin Hao, Hong Qin 0001
IEEE Trans. Vis. Comput. Graph.7
2022 Foveated Stochastic Lightcuts
abstract
Foveated rendering provides an idea for accelerating rendering algorithms without sacrificing the perceived rendering quality in virtual reality applications. In this paper, we propose a foveated stochastic lightcuts method to render high-quality many-lights illumination effects in high perception-sensitive regions. First, we introduce a spatiotemporal-luminance based lightcuts generation method to generate lightcuts with different accuracy for different visual perception-sensitive regions. Then we propose a multi-resolution light samples selection method to select the light sample for each node in the lightcuts more efficiently. Our method supports full-dynamic scenes containing over 250k dynamic light sources and dynamic diffuse/specular/glossy objects. It provides frame rates up to 110fps for high-quality many-lights illumination effects in high perception-sensitive regions of the HVS in VR HMDs. Compared with the state-of-the-art stochastic lightcuts method using the same rendering time, our method achieves smaller mean squared errors in the fovea and periphery. We also conduct user studies to prove that the perceived quality of our method has a high visual similarity with the results of the ground truth rendered by using the stochastic lightcuts with 2048 light samples per pixel.
Xuehuai Shi, Lili Wang 0006, Jian Wu 0033, Runze Fan, Aimin Hao
IEEE Trans. Vis. Comput. Graph.5
2021 Point Cloud Semantic Scene Completion from RGB-D Images
abstract
In this paper, we devise a novel semantic completion network, called point cloud semantic scene completion network (PCSSC-Net), for indoor scenes solely based on point clouds. Existing point cloud completion networks still suffer from their inability of fully recovering complex structures and contents from global geometric descriptions neglecting semantic hints. To extract and infer comprehensive information from partial input, we design a patch-based contextual encoder to hierarchically learn point-level, patch-level, and scene-level geometric and contextual semantic information with a divide-and-conquer strategy. Consider that the scene semantics afford a high-level clue of constituting geometry for an indoor scene environment, we articulate a semantics-guided completion decoder where semantics could help cluster isolated points in the latent space and infer complicated scene geometry. Given the fact that real-world scans tend to be incomplete as ground truth, we choose to synthesize scene dataset with RGB-D images and annotate complete point clouds as ground truth for the supervised training purpose. Extensive experiments validate that our new method achieves the state-of-the-art performance, in contrast with the current methods applied to our dataset.
Shoulong Zhang, Shuai Li 0001, Aimin Hao, Hong Qin 0001
AAAI3
2021 From Semantic Categories to Fixations: A Novel Weakly-Supervised Visual-Auditory Saliency Detection Approach
abstract
Thanks to the rapid advances in the deep learning techniques and the wide availability of large-scale training sets, the performances of video saliency detection models have been improving steadily and significantly. However, the deep learning based visual-audio fixation prediction is still in its infancy. At present, only a few visual-audio sequences have been furnished with real fixations being recorded in the real visual-audio environment. Hence, it would be neither efficiency nor necessary to re-collect real fixations under the same visual-audio circumstance. To address the problem, this paper advocate a novel approach in a weakly-supervised manner to alleviating the demand of large-scale training sets for visual-audio model training. By using the video category tags only, we propose the selective class activation mapping (SCAM), which follows a coarse-to-fine strategy to select the most discriminative regions in the spatial-temporal-audio circumstance. Moreover, these regions exhibit high consistency with the real human-eye fixations, which could subsequently be employed as the pseudo GTs to train a new spatial-temporal-audio (STA) network. Without resorting to any real fixation, the performance of our STA network is comparable to that of the fully supervised ones. Our code and results are publicly available at https://github.com/guotaowang/STANet.
Guotao Wang 0004, Chenglizhao Chen, Deng-Ping Fan, Aimin Hao, Hong Qin 0001
CVPR4
2021 Knowledge-inspired 3D Scene Graph Prediction in Point Cloud
abstract
Prior knowledge integration helps identify semantic entities and their relationships in a graphical representation, however, its meaningful abstraction and intervention remain elusive. This paper advocates a knowledge-inspired 3D scene graph prediction method solely based on point clouds. At the mathematical modeling level, we formulate the task as two sub-problems: knowledge learning and scene graph prediction with learned prior knowledge. Unlike conventional methods that learn knowledge embedding and regular patterns from encoded visual information, we propose to suppress the misunderstandings caused by appearance similarities and other perceptual confusion. At the network design level, we devise a graph auto-encoder to automatically extract class-dependent representations and topological patterns from the one-hot class labels and their intrinsic graphical structures, so that the prior knowledge can avoid perceptual errors and noises. We further devise a scene graph prediction model to predict credible relationship triplets by incorporating the related prototype knowledge with perceptual information. Comprehensive experiments confirm that, our method can successfully learn representative knowledge embedding, and the obtained prior knowledge can effectively enhance the accuracy of relationship predictions. Our thorough evaluations indicate the new method can achieve the state-of-the-art performance compared with other scene graph prediction methods.
Shoulong Zhang, Shuai Li 0001, Aimin Hao, Hong Qin 0001
NeurIPS3
2021 Correction to: Long-Short Temporal-Spatial Clues Excited Network for Robust Person Re-identification
Shuai Li 0001, Wenfeng Song, Zheng Fang 0008, Jiaying Shi, Aimin Hao, Qinping Zhao, Hong Qin 0001
Int. J. Comput. Vis.5
2021 Depth quality-aware selective saliency fusion for RGB-D image salient object detection
Xuehao Wang, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001
Neurocomputing4
2021 Improving video anomaly detection performance by mining useful data from unseen video frames
Renzhi Wu, Shuai Li 0001, Chenglizhao Chen, Aimin Hao
Neurocomputing4
2021 Hierarchical Object Relationship Constrained Monocular Depth Estimation
Shuai Li 0001, Jiaying Shi, Wenfeng Song, Aimin Hao, Hong Qin 0001
Pattern Recognit.4
2021 A Plug-and-Play Scheme to Adapt Image Saliency Deep Model for Video Data
abstract
With the rapid development of deep learning techniques, image saliency deep models trained solely by spatial information have occasionally achieved detection performance for video data comparable to that of the models trained by both spatial and temporal information. However, due to the lesser consideration of temporal information, the image saliency deep models may become fragile in the video sequences dominated by temporal information. Thus, the most recent video saliency detection approaches have adopted the network architecture starting with a spatial deep model that is followed by an elaborately designed temporal deep model. However, such methods easily encounter the performance bottleneck arising from the single stream learning methodology, so the overall detection performance is largely determined by the spatial deep model. In sharp contrast to the current mainstream methods, this paper proposes a novel plug-and-play scheme to weakly retrain a pretrained image saliency deep model for video data by using the newly sensed and coded temporal information. Thus, the retrained image saliency deep model will be able to maintain temporal saliency awareness, achieving much improved detection performance. Moreover, our method is simple yet effective for adapting any off-the-shelf pre-trained image saliency deep model to obtain high-quality video saliency detection. Additionally, both the data and source code of our method are publicly available.
Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001
IEEE Trans. Circuits Syst. Video Technol.4
2021 A Global-Local Self-Adaptive Network for Drone-View Object Detection
abstract
Directly benefiting from the deep learning methods, object detection has witnessed a great performance boost in recent years. However, drone-view object detection remains challenging for two main reasons: (1) Objects of tiny-scale with more blurs w.r.t. ground-view objects offer less valuable information towards accurate and robust detection; (2) The unevenly distributed objects make the detection inefficient, especially for regions occupied by crowded objects. Confronting such challenges, we propose an end-to-end global-local self-adaptive network (GLSAN) in this paper. The key components in our GLSAN include a global-local detection network (GLDN), a simple yet efficient self-adaptive region selecting algorithm (SARSA), and a local super-resolution network (LSRN). We integrate a global-local fusion strategy into a progressive scale-varying network to perform more precise detection, where the local fine detector can adaptively refine the target's bounding boxes detected by the global coarse detector via cropping the original images for higher-resolution detection. The SARSA can dynamically crop the crowded regions in the input images, which is unsupervised and can be easily plugged into the networks. Additionally, we train the LSRN to enlarge the cropped images, providing more detailed information for finer-scale feature extraction, helping the detector distinguish foreground and background more easily. The SARSA and LSRN also contribute to data augmentation towards network training, which makes the detector more robust. Extensive experiments and comprehensive evaluations on the VisDrone2019-DET benchmark dataset and UAVDT dataset demonstrate the effectiveness and adaptivity of our method. Towards an industrial application, our network is also applied to a DroneBolts dataset with proven advantages. Our source codes have been available at https://github.com/dengsutao/glsan.
Sutao Deng, Shuai Li 0001, Ke Xie 0005, Wenfeng Song, Xiao Liao, Aimin Hao, Hong Qin 0001
IEEE Trans. Image Process.6
2021 Rethinking Image Salient Object Detection: Object-Level Semantic Saliency Reranking First, Pixelwise Saliency Refinement Later
abstract
Human attention is an interactive activity between our visual system and our brain, using both low-level visual stimulus and high-level semantic information. Previous image salient object detection (SOD) studies conduct their saliency predictions via a multitask methodology in which pixelwise saliency regression and segmentation-like saliency refinement are conducted simultaneously. However, this multitask methodology has one critical limitation: the semantic information embedded in feature backbones might be degenerated during the training process. Our visual attention is determined mainly by semantic information, which is evidenced by our tendency to pay more attention to semantically salient regions even if these regions are not the most perceptually salient at first glance. This fact clearly contradicts the widely used multitask methodology mentioned above. To address this issue, this paper divides the SOD problem into two sequential steps. First, we devise a lightweight, weakly supervised deep network to coarsely locate the semantically salient regions. Next, as a postprocessing refinement, we selectively fuse multiple off-the-shelf deep models on the semantically salient regions identified by the previous step to formulate a pixelwise saliency map. Compared with the state-of-the-art (SOTA) models that focus on learning the pixelwise saliency in single images using only perceptual clues, our method aims at investigating the object-level semantic ranks between multiple images, of which the methodology is more consistent with the human attention mechanism. Our method is simple yet effective, and it is the first attempt to consider salient object detection as mainly an object-level semantic reranking problem.
Guangxiao Ma, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001
IEEE Trans. Image Process.4
2021 Data-Level Recombination and Lightweight Fusion Scheme for RGB-D Salient Object Detection
abstract
Existing RGB-D salient object detection methods treat depth information as an independent component to complement RGB and widely follow the bistream parallel network architecture. To selectively fuse the CNN features extracted from both RGB and depth as a final result, the state-of-the-art (SOTA) bistream networks usually consist of two independent subbranches: one subbranch is used for RGB saliency, and the other aims for depth saliency. However, depth saliency is persistently inferior to the RGB saliency because the RGB component is intrinsically more informative than the depth component. The bistream architecture easily biases its subsequent fusion procedure to the RGB subbranch, leading to a performance bottleneck. In this paper, we propose a novel data-level recombination strategy to fuse RGB with D (depth) before deep feature extraction, where we cyclically convert the original 4-dimensional RGB-D into DGB, RDB and RGD. Then, a newly lightweight designed triple-stream network is applied over these novel formulated data to achieve an optimal channel-wise complementary fusion status between the RGB and D, achieving a new SOTA performance.
Xuehao Wang, Shuai Li 0001, Chenglizhao Chen, Yuming Fang 0001, Aimin Hao, Hong Qin 0001
IEEE Trans. Image Process.5
2021 Simulating Multi-Scale, Granular Materials and Their Transitions With a Hybrid Euler-Lagrange Solver
abstract
Multi-scale granular materials, such as powdered materials and mudslides, are pretty common in nature. Modeling such materials and their phase transitions remains challenging since this task involves the delicate representations of various ranges of particles with multiple scales that cause their property variations among liquid, granular solid (i.e., particles), and smoke-like materials. To effectively animate the complicated yet intriguing natural phenomena involving multi-scale granular materials and their phase transitions in graphics with high fidelity, this article advocates a hybrid Euler-Lagrange solver to handle the behaviors of involved discontinuous fluid-like materials faithfully. At the algorithmic level, we present a unified framework that tightly couples the affine particle-in-cell (APIC) solver with density field to achieve the transformation spanning across granular particles, dust cloud, powders, and their natural mixtures. For example, a part of the granular particles could be transformed into dust cloud while interacting with air and being represented by density field. Meanwhile, the velocity decrease of the involved materials could also result in the transit from the density-field-driven dust to powder particles. Besides, to further enhance our modeling and simulation power to broaden the range of multi-scale materials, we introduce a moisture property for granular particles to control the transitions between particles and viscous liquid. At the geometric level, we devise an additional surface-tracking procedure to simulate the viscous liquid phase. We can arrive at delicate viscous behaviors by controlling the corresponding yield conditions. Through various experiments with the different scenes design being conducted in our unified framework, we can validate the mixed multi-scale materials' mutual transformation processes. Our unified framework furnished with a hybrid solver can significantly enhance the modeling flexibility and the animation potential of the particle-grid hybrid materials in graphics.
Yang Gao 0032, Shuai Li 0001, Aimin Hao, Hong Qin 0001
IEEE Trans. Vis. Comput. Graph.3
2021 Design and Evaluation of Personalized Percutaneous Coronary Intervention Surgery Simulation System
abstract
In recent years, medical simulators have been widely applied to a broad range of surgery training tasks. However, most of the existing surgery simulators can only provide limited immersive environments with a few pre-processed organ models, while ignoring the instant modeling of various personalized clinical cases, which brings substantive differences between training experiences and real surgery situations. To this end, we present a virtual reality (VR) based surgery simulation system for personalized percutaneous coronary intervention (PCI). The simulation system can directly take patient-specific clinical data as input and generate virtual 3D intervention scenarios. Specially, we introduce a fiber-based patient-specific cardiac dynamic model to simulate the nonlinear deformation among the multiple layers of the cardiac structure, which can well respect and correlate the atriums, ventricles and vessels, and thus gives rise to more effective visualization and interaction. Meanwhile, we design a tracking and haptic feedback hardware, which can enable users to manipulate physical intervention instruments and interact with virtual scenarios. We conduct quantitative analysis on deformation precision and modeling efficiency, and evaluate the simulation system based on the user studies from 16 cardiologists and 20 intervention trainees, comparing it to traditional desktop intervention simulators. The results confirm that our simulation system can provide a better user experience, and is a suitable platform for PCI surgery training and rehearsal.
Shuai Li 0001, Jiahao Cui 0001, Aimin Hao, Qinping Zhao
IEEE Trans. Vis. Comput. Graph.3
2021 Vectorized Painting with Temporal Diffusion Curves
abstract
This paper presents a vector painting system for digital artworks. We first propose Temporal Diffusion Curve (TDC), a new form of vector graphics, and a novel random-access solver for modeling the evolution of strokes. With the help of a procedural stroke processing function, the TDC strokes can achieve various shapes and effects for multiple art styles. Based on these, we build a painting system of great potential. Thanks to the random-access solver, our method has real-time performance regardless of the rendering resolution, provides straightforward editing possibilities on strokes both at runtime and afterward, and is effective and straightforward for art production. Compared with the previous Diffusion Curve, our method uses strokes as the basic graphics primitives, which are able to intersect each other and much more consistent with the intuition and painting habits of human. We finally demonstrate that professional artists can create multiple genres of artworks with our painting system.
Yingjia Li, Xiao Zhai, Fei Hou 0001, Aimin Hao, Hong Qin 0001
IEEE Trans. Vis. Comput. Graph.5
2021 Augmented reality-based visual-haptic modeling for thoracoscopic surgery training systems
abstract
Compared with traditional thoracotomy, video-assisted thoracoscopic surgery (VATS) has less minor trauma, faster recovery, higher patient compliance, but higher requirements for surgeons. Virtual surgery training simulation systems are important and have been widely used in Europe and America. Augmented reality (AR) in surgical training simulation systems significantly improve the training effect of virtual surgical training, although AR technology is still in its initial stage. Mixed reality has gained increased attention in technology-driven modern medicine but has yet to be used in everyday practice. This study proposed an immersive AR lobectomy within a thoracoscope surgery training system, using visual and haptic modeling to study the potential benefits of this critical technology. The content included immersive AR visual rendering, based on the cluster-based extended position-based dynamics algorithm of soft tissue physical modeling. Furthermore, we designed an AR haptic rendering systems, whose model architecture consisted of multi-touch interaction points, including kinesthetic and pressure-sensitive points. Finally, based on the above theoretical research, we developed an AR interactive VATS surgical training platform. Twenty-four volunteers were recruited from the First People's Hospital of Yunnan Province to evaluate the VATS training system. Face, content, and construct validation methods were used to assess the tactile sense, visual sense, scene authenticity, and simulator performance. The results of our construction validation demonstrate that the simulator is useful in improving novice and surgical skills that can be retained after a certain period of time. The video-assisted thoracoscopic system based on AR developed in this study is effective and can be used as a training device to assist in the development of thoracoscopic skills for novices.
Yonghang Tai, Junsheng Shi, JunJun Pan, Aimin Hao, Victor Chang 0001
Virtual Real. Intell. Hardw.4
2021 Interactive Hepatic Parenchymal Transection Simulation with Haptic Feedback
abstract
Liver resection involves surgical removal of a portion of the liver. It is used to treat liver tumors and liver injuries. The complexity and high-risk nature of this surgery prevents novice doctors from practicing it on real patients. Virtual surgery simulation was developed to simulate surgical procedures to enable medical professionals to be trained without requiring a patient, a cadaver, or an animal. Therefore, there is a strong need for the development of a liver resection surgery simulation system. We propose a real-time simulation system that provides realistic visual and tactile feedback for hepatic parenchymal transection. The tetrahedron structure and cluster-based shape matching are used for physical model construction, topology update of a three-dimensional liver model soft deformation simulation, and haptic rendering acceleration. During the liver parenchyma separation simulation, a tetrahedral mesh is used for surface triangle subdivision and surface generation of the surgical wound. The shape-matching cluster is separated via component detection on an undirected graph constructed using the tetrahedral mesh. In our system, cluster-based shape matching is implemented on a GPU, whereas haptic rendering and topology updates are implemented on a CPU. Experimental results show that haptic rendering can be performed at a high frequency (> 900 Hz), whereas mesh skinning and graphics rendering can be performed at 45 fps. The topology update can be executed at an interactive rate (> 10 Hz) on a single CPU thread. We propose an interactive hepatic parenchymal transection simulation method based on a tetrahedral structure. The tetrahedral mesh simultaneously supports physical model construction, topology update, and haptic rendering acceleration.
Haonan Yu, Aimin Hao
Virtual Real. Intell. Hardw.7
2021 Orthodontic simulation system with force feedback for training complete bracket placement procedures
abstract
A virtual system that simulates the complete process of orthodontic bracket placement can be used for pre-clinical skill training to help students gain confidence by performing the required tasks on a virtual patient. The hardware for the virtual simulation system is built using two force feedback devices to support bi-manual force feedback operation. A 3D mouse is used to adjust the position of the virtual patient. A multi-threaded computational methodology is adopted to satisfy the requirements of the frame rate. The computation threads mainly consist of the haptic thread running at a frequency of >1000Hz and the graphic thread at >30Hz. The graphic thread allows the graphics engine to effectively display the visual effects of biofilm removal and acid erosion through texture mapping. Using the haptic thread, the physics engine adopts the hierarchy octree collision-detection algorithm to simulate the multi-point and multi-region interaction between the tools and the virtual environment. Its high efficiency guarantees that the time cost can be controlled within 1 ms. The physics engine also performs collision detection between the tools and particles, making it possible to simulate paint and removal of colloids. The surface-contact constraints are defined in the system; this ensures that the bracket will not divorce from or embed into the tooth during the adjustment of the bracket. Therefore, the simulated adjustment is more realistic and natural. A virtual system to simulate the complete process of orthodontic bracket bonding was developed. In addition to bracket bonding and adjustment, the system simulates the necessary auxiliary steps such as smearing, acid etching, and washing. Furthermore, the system supports personalized case training. The system provides a new method for students to practice orthodontic skills.
Luwei Liu, Xiaohan Zhao, Aimin Hao
Virtual Real. Intell. Hardw.5
2020 Meta-RetinaNet for Few-shot Object Detection
Shaoqi Li, Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001
BMVC4
2020 Meta Transfer Learning for Adaptive Vehicle Tracking in UAV Videos
Wenfeng Song, Shuai Li 0001, Shaoqi Li, Aimin Hao, Hong Qin 0001, Qinping Zhao
MMM (1)5
2020 Cross-View Contextual Relation Transferred Network for Unsupervised Vehicle Tracking in Drone Videos
abstract
Recently CNN-centric object tracking methods have been gaining tremendous success in ground-view videos, however, it remains hard to cope with vehicle tracking in unmanned aerial vehicle (UAV) videos. The key difficulties mainly stem from lacking large-scale well-labeled training datasets and view-invariant appearance model for fast-moving drone-view vehicles. We enhance the vehicle's cross-view feature by exploring relations between the pivotal context and the target to facilitate unsupervised vehicle tracking. The relation is modeled as the relevance of the target and its contextual regions in the tracking task. Specifically, we propose a contextual relation actor-critic (CRAC) framework integrates an actor-critic agent with a dual GAN learning mechanism, which aims to dynamically search the related contextual regions and transfer the relations from ground-view to drone-view videos while retaining the discriminative features. We demonstrate that CRAC could be applied to several state-of-the-art trackers by extensive experiments and ablation studies on four public benchmarks. All the experiments confirm that, our CRAC can improve the performance of state-of-the-art methods in terms of accuracy, robustness, and versatility.
Wenfeng Song, Shuai Li 0001, Tao Chang, Aimin Hao, Qinping Zhao, Hong Qin 0001
WACV4
2020 Accelerating Liquid Simulation With an Improved Data-Driven Method
abstract
Abstract In physics‐based liquid simulation for graphics applications, pressure projection consumes a significant amount of computational time and is frequently the bottleneck of the computational efficiency. How to rapidly apply the pressure projection and at the same time how to accurately capture the liquid geometry are always among the most popular topics in the current research trend in liquid simulations. In this paper, we incorporate an artificial neural network into the simulation pipeline for handling the tricky projection step for liquid animation. Compared with the previous neural‐network‐based works for gas flows, this paper advocates new advances in the composition of representative features as well as the loss functions in order to facilitate fluid simulation with free‐surface boundary. Specifically, we choose both the velocity and the level‐set function as the additional representation of the fluid states, which allows not only the motion but also the boundary position to be considered in the neural network solver. Meanwhile, we use the divergence error in the loss function to further emulate the lifelike behaviours of liquid. With these arrangements, our method could greatly accelerate the pressure projection step in liquid simulation, while maintaining fairly convincing visual results. Additionally, our neutral network performs well when being applied to new scene synthesis even with varied boundaries or scales.
Yang Gao 0032, Quancheng Zhang, Shuai Li 0001, Aimin Hao, Hong Qin 0001
Comput. Graph. Forum4
2020 Spatiotemporal consistency-based adaptive hand-held video stabilization
Shuai Li 0001, Hong Qin 0001, Aimin Hao
Sci. China Inf. Sci.4
2020 Dynamic particle partitioning SPH model for high-speed fluids simulation
Yang Gao 0032, Jin Li 0068, Shuai Li 0001, Aimin Hao, Hong Qin 0001
Graph. Model.5
2020 Long-Short Temporal-Spatial Clues Excited Network for Robust Person Re-identification
Shuai Li 0001, Wenfeng Song, Zheng Fang 0008, Jiaying Shi, Aimin Hao, Qinping Zhao, Hong Qin 0001
Int. J. Comput. Vis.5
2020 Real-time suturing simulation for virtual reality medical training
abstract
Abstract At present, virtual reality (VR) ‐based medical simulators provide an efficient and cost‐effective alternative without exposing risk to the traditional training approaches. As an essential and indispensable task in fundamental surgical skills training, the research of suturing simulation still remains insufficient in the field of virtual surgery. In this paper, we present a real‐time suturing simulation framework which can handle the complex interactions between surgical instruments and soft tissue. The simulation consists of two stages: external interaction and internal coupling. External interaction involves the interplay between needle/suture and the soft tissue, which are both deformed by position‐based dynamics (PBD) with different constraints. At the internal coupling stage, once the force exceeds a threshold, the needle tip will puncture and penetrate into the soft tissue and generate a path. To guarantee the needle/suture accurately following the path inside the soft tissue, we propose a novel coupling method by matching and generating the constraints among needle, suture, and penetration path. We have applied this suturing simulation into a VR laparoscopic surgery simulator with haptic force. Our experimental results demonstrate that our approach can achieve real‐time performance with a high degree of visual realism and haptic fidelity.
JunJun Pan, Hong Qin 0001, Aimin Hao
Comput. Animat. Virtual Worlds4
2020 Adaptive appearance modeling via hierarchical entropy analysis over multi-type features
Jizhou Ma, Shuai Li 0001, Hong Qin 0001, Aimin Hao
Pattern Recognit.4
2020 Context-Interactive CNN for Person Re-Identification
abstract
Despite growing progresses in recent years, cross-scenario person re-identification remains challenging, mainly due to the pedestrians commonly surrounded by highly-complex environment contexts. In reality, the human perception mechanism could adaptively find proper contextualized spatial-temporal clues towards pedestrian recognition. However, conventional methods fall short in adaptively leveraging the long-term spatial-temporal information due to ever-increasing computational cost. Moreover, CNN-based deep learning methods are hard to conduct optimization due to the non-differentiable property of the built-in context search operation. To ameliorate, this paper proposes a novel Context-Interactive CNN (CI-CNN) to dynamically find both spatial and temporal contexts by embedding multi-task Reinforcement Learning (MTRL). The CI-CNN streamlines the multi-task reinforcement learning by using an actor-critic agent to capture the temporal-spatial context simultaneously, which comprises a context-policy network and a context-critic network. The former network learns policies to determine the optimal spatial context region and temporal sequence range. Based on the inferred temporal-spatial cues, the latter one focuses on the identification task and provides feedback for the policy network. Thus, CI-CNN can simultaneously zoom in/out the perception field in spatial and temporal domain for the context interaction with the environment. By fostering the collaborative interaction between the person and context, our method could achieve outstanding performance on various public benchmarks, which confirms the rationality of our hypothesis, and verifies the effectiveness of our CI-CNN framework.
Wenfeng Song, Shuai Li 0001, Tao Chang, Aimin Hao, Qinping Zhao, Hong Qin 0001
IEEE Trans. Image Process.4
2020 Accurate and Robust Video Saliency Detection via Self-Paced Diffusion
abstract
Conventional video saliency detection methods frequently follow the common bottom-up thread to estimate video saliency within the short-term fashion. As a result, such methods can not avoid the obstinate accumulation of errors when the collected low-level clues are constantly ill-detected. Also, being noticed that a portion of video frames, which are not nearby the current video frame over the time axis, may potentially benefit the saliency detection in the current video frame. Thus, we propose to solve the aforementioned problem using our newly-designed key frame strategy (KFS), whose core rationale is to utilize both the spatial-temporal coherency of the salient foregrounds and the objectness prior (i.e., how likely it is for an object proposal to contain an object of any class) to reveal the valuable long-term information. We could utilize all this newly-revealed long-term information to guide our subsequent “self-paced” saliency diffusion, which enables each key frame itself to determine its diffusion range and diffusion strength to correct those ill-detected video frames. At the algorithmic level, we first divide a video sequence into short-term frame batches, and the object proposals are obtained in a frame-wise manner. Then, for each object proposal, we utilize a pre-trained deep saliency model to obtain high-dimensional features in order to represent the spatial contrast. Since the contrast computation within multiple neighbored video frames (i.e., the non-local manner) is relatively insensitive to the appearance variation, those object proposals with high-quality low-level saliency estimation frequently exhibit strong similarity over the temporal scale. Next, the long-term common consistency (e.g., appearance models/movement patterns) of the salient foregrounds could be explicitly revealed via similarity analysis accordingly. We further boost the detection accuracy via long-term information guided saliency diffusion in a self-paced manner. We have conducted extensive experiments to compare our method with 16 state-of-the-art methods over 4 largest public available benchmarks, and all results demonstrate the superiority of our method in terms of both accuracy and robustness.
Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001
IEEE Trans. Multim.4
2020 Salient Object Detection via Multiple Instance Joint Re-Learning
abstract
In recent years deep neural networks have been widely applied to visual saliency detection tasks with remarkable detection performance improvements. As for the salient object detection in single image, the automatically computed convolutional features frequently demonstrate high discriminative power to distinguish salient foregrounds from its non-salient surroundings in most cases. Yet, the obstinate feature conflicts still persist, which naturally gives rise to the learning ambiguity, arriving at massive failure detections. To solve such problem, we propose to jointly re-learn common consistency of inter-image saliency and then use it to boost the detection performance. Its core rationale is to utilize the easy-to-detect cases to re-boost much harder ones. Compared with the conventional methods, which focus on their problem domain within the single image scope, our method attempts to utilize those beyond-scope information to facilitate the current salient object detection. To validate our new approach, we have conducted a comprehensive quantitative comparisons between our approach and 13 state-of-the-art methods over 5 publicly available benchmarks, and all the results suggest the advantage of our approach in terms of accuracy, reliability, and versatility.
Guangxiao Ma, Chenglizhao Chen, Shuai Li 0001, Chong Peng 0001, Aimin Hao, Hong Qin 0001
IEEE Trans. Multim.5
2020 Contextualized CNN for Scene-Aware Depth Estimation From Single RGB Image
abstract
Directly benefited from deep learning techniques, depth estimation from single image has gained great momentum in recent years. However, most of the existing approaches treat depth prediction as an isolated problem without taking into consideration high-level semantic context information, which results in inefficient utilization of training dataset and unavoidably requires a large number of captured depth data during the training phase. To ameliorate, this paper develops a novel scene-aware contextualized convolution neural network (CCNN), which characterizes the semantic context relationship at the class-level and refines depth at the pixel-level. Our newly-proposed CCNN is built upon the intrinsic exploitation of context-dependent depth association, including inner-object continuous depth and inter-object depth change priors nearby. Specifically, rather than conducting regression on depth in single CNN, we make the first attempt to integrate both class-level and pixel-level conditional random fields (CRFs) based probabilistic graphical model into the powerful CNN framework to simultaneously learn different-level features within the same CNN layer. With our CCNN, the former model will guide the latter one to learn the contextualized RGB-Depth mapping. Hence, CCNN has desirable properties in both class-level integrity and pixel-level discrimination, which makes it ideal to share such two-level convolutional features in parallel during the end-to-end training with the commonly-used back-propagation algorithm. We conduct extensive experiments and comprehensive evaluations on public benchmarks involving various indoor and outdoor scenes, and all the experiments confirm that, our method outperforms the state-of-the-art depth estimation methods, especially for the cases where only small-scale training data are readily available.
Wenfeng Song, Shuai Li 0001, Aimin Hao, Qinping Zhao, Hong Qin 0001
IEEE Trans. Multim.4
2020 Poisson Vector Graphics (PVG)
abstract
This paper presents Poisson vector graphics (PVG), an extension of the popular diffusion curves (DC), for generating smooth-shaded images. Armed with two new types of primitives, called Poisson curves and Poisson regions, PVG can easily produce photorealistic effects such as specular highlights, core shadows, translucency and halos. Within the PVG framework, the users specify color as the Dirichlet boundary condition of diffusion curves and control tone by offsetting the Laplacian of colors, where both controls are simply done by mouse click and slider dragging. PVG distinguishes itself from other diffusion based vector graphics for 3 unique features: 1) explicit separation of colors and tones, which follows the basic drawing principle and eases editing; 2) native support of seamless cloning in the sense that PCs and PRs can automatically fit into the target background; and 3) allowed intersecting primitives (except for DC-DC intersection) so that users can create layers. Through extensive experiments and a preliminary user study, we demonstrate that PVG is a simple yet powerful authoring tool that can produce photo-realistic vector graphics from scratch.
Fei Hou 0001, Qian Sun 0003, Zheng Fang 0008, Yong-Jin Liu 0001, Shi-Min Hu 0001, Hong Qin 0001, Aimin Hao, Ying He 0001
IEEE Trans. Vis. Comput. Graph.7
2020 Stage-wise Salient Object Detection in 360° Omnidirectional Image via Object-level Semantical Saliency Ranking
abstract
The 2D image based salient object detection (SOD) has been extensively explored, while the 360° omnidirectional image based SOD has received less research attention and there exist three major bottlenecks that are limiting its performance. Firstly, the currently available training data is insufficient for the training of 360° SOD deep model. Secondly, the visual distortions in 360° omnidirectional images usually result in large feature gap between 360° images and 2D images; consequently, the widely used stage-wise training-a widely-used solution to alleviate the training data shortage problem, becomes infeasible when conducing SOD in 360° omnidirectional images. Thirdly, the existing 360° SOD approach has followed a multi-task methodology that performs salient object localization and segmentation-like saliency refinement at the same time, being faced with extremely large problem domain, making the training data shortage dilemma even worse. To tackle all these issues, this paper divides the 360° SOD into a multi-staqe task, the key rationale of which is to decompose the original complex problem domain into sequential easy sub problems that only demand for small-scale training data. Meanwhile, we learn how to rank the "object-level semantical saliency", aiming to locate salient viewpoints and objects accurately. Specifically, to alleviate the training data shortage problem, we have released a novel dataset named 360-SSOD, containing 1,105 360° omnidirectional images with manually annotated object-level saliency ground truth, whose semantical distribution is more balanced than that of the existing dataset. Also, we have compared the proposed method with 13 SOTA methods, and all quantitative results have demonstrated the performance superiority.
Guangxiao Ma, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001
IEEE Trans. Vis. Comput. Graph.4
2020 Fluid Simulation with Adaptive Staggered Power Particles on GPUs
abstract
This paper extends the recently proposed power-particle-based fluid simulation method with staggered discretization, GPU implementation, and adaptive sampling, largely enhancing the efficiency and usability of the method. In contrast to the original formulation which uses co-located pressures and velocities, in this paper, a staggered scheme is adapted to the Power Particles to benefit visual details and computing efficiency. Meanwhile, we propose a novel facet-based power diagrams construction algorithm suitable for parallelization and explore its GPU implementation, achieving an order of magnitude boost in performance over the existing code library. In addition, to utilize the potential of Power Particles to control individual cell volume, we apply adaptive particle sampling to improve the detail level with varying resolution. The proposed method can be entirely carried out on GPUs, and our extensive experiments validate our method both in terms of efficiency and visual quality.
Xiao Zhai, Fei Hou 0001, Hong Qin 0001, Aimin Hao
IEEE Trans. Vis. Comput. Graph.4
2020 Compressing animated meshes with fine details using local spectral analysis and deformation transfer
Chengju Chen, Qing Xia 0002, Shuai Li 0001, Hong Qin 0001, Aimin Hao
Vis. Comput.5
2020 Personalized cardiovascular intervention simulation system
abstract
Background This study proposes a series of geometry and physics modeling methods for personalized cardiovascular intervention procedures, which can be applied to a virtual endovascular simulator. Methods Based on personalized clinical computed tomography angiography (CTA) data, mesh models of the cardiovascular system were constructed semi-automatically. By coupling 4D magnetic resonance imaging (MRI) sequences corresponding to a complete cardiac cycle with related physics models, a hybrid kinetic model of the cardiovascular system was built to drive kinematics and dynamics simulation. On that basis, the surgical procedures related to intervention instruments were simulated using specially-designed physics models. These models can be solved in real-time; therefore, the complex interactions between blood vessels and instruments can be well simulated. Additionally, X-ray imaging simulation algorithms and realistic rendering algorithms for virtual intervention scenes are also proposed. In particular, instrument tracking hardware with haptic feedback was developed to serve as the interaction interface of real instruments and the virtual intervention system. Finally, a personalized cardiovascular intervention simulation system was developed by integrating the techniques mentioned above. Results This system supported instant modeling and simulation of personalized clinical data and significantly improved the visual and haptic immersions of vascular intervention simulation. Conclusions It can be used in teaching basic cardiology and effectively satisfying the demands of intervention training, personalized intervention planning, and rehearsing.
Aimin Hao, Jiahao Cui 0001, Shuai Li 0001, Qinping Zhao
Virtual Real. Intell. Hardw.1
2019 Fine-Grained Thyroid Nodule Classification via Multi-Semantic Attention Network
abstract
Thyroid nodule classification in ultrasound images has gained great momentum based on deep convolutional neural networks in recent years. Nevertheless, it is still challenging to intelligently classify the fine-grained thyroid nodules, which is significant for the subsequent clinical treatments. The difficulties mainly stem from four aspects: few fine-grained training dataset, highly-variable appearances of intra-class nodules, overall-similar characteristics of inter-class nodules, and the low resolution and contrast degree of the ultrasonic images as well as the influence of intrinsic speckle noises. In this paper, we propose a multi-semantic attention networks (MSAN) for fine-grained thyroid nodule classification in ultrasound images. Specifically, we employ a main network branch for coarse granularity feature extraction, which only focuses on the benign and malignant characteristics, and simultaneously employ multi-semantic network branches to extract discriminative features from the fine-grained pathological categories. Meanwhile, we introduce an self-attention scheme together with global average pooling (GAP) in our network, which facilitates to learn from the dynamically-selected nodule regions ranging from local to global. Extensive experiments demonstrate that, our MSAN gives rise to significant improvement of classification accuracy and outperforms the state-of-the-art methods.
Shuai Li 0001, Wenfeng Song, Zhennan Pang, Aimin Hao, Hong Qin 0001
BIBM5
2019 Few-Shot Learning for Monocular Depth Estimation Based on Local Object Relationship
abstract
Monocular depth estimation has gained great momentum and achieved growing success recently. Nonetheless, due to the intrinsic difficulty associated with large-scale RGB-D data capture for training purpose and the inefficient utilization of existing training datasets, it is still challenging to accommodate flexibly-changing scenarios. To ameliorate, we propose a fewshot learning method for monocular depth estimation augmented by local object-object relationship. Our method is based on the insight that the depth changing between neighboring objects is relatively stable across diverse but similar scenarios. At the technical front, we first learn the object relationship based on the relative distance between single objects. Towards this goal, we design a CNN architecture to simultaneously encode the object spatial context into object-object relationship features and encode the original image into global context features. Hence we can complementally leverage few-shot dataset with only a few samples for depth estimation while preserving the global depth changing range and respecting the local object-object depth details. As a result, our novel approach could estimate depth from various indoor RGB images, which greatly alleviates the training dataset dependency in monocular depth estimation. Finally, we conduct extensive experiments and comprehensive evaluations on the widely-used public benchmarks, and all the experiments confirm that, our method outperforms the state-of-the-art depth estimation methods, especially for the cases where only smallscale training samples are available.
Shuai Li 0001, Jiaying Shi, Wenfeng Song, Aimin Hao, Hong Qin 0001
ICTAI4
2019 Context-Aware Network for 3D Human Pose Estimation from Monocular RGB Image
abstract
Convolutional Neural Network (CNN) has brought tremendous improvements in estimating 3D human pose from a monocular RGB image. However, the task of 3D human pose estimation still remains extremely challenging, especially when the task is geared towards estimating the depth of human body parts. Different from 2D human pose estimation, which focuses on the fusion of spatial information and context information, depth estimation demands more context information. Inspired by this, we build a Context-Aware Network (CAN) which can fully explore the context information to discover the underlying relationships among different body parts. The key ingredient of our network is High-Level Depth Estimation Module (HLDEM) designed to extract context information effectively. Additionally, multi-scale supervision is introduced in our network to extract context information at different scales. Experimental results show that our network achieves competitive performance compared with state-of-the-art methods on Human3.6M dataset.
Binyi Yin, Dongbo Zhang 0004, Shuai Li 0001, Aimin Hao, Hong Qin 0001
IJCNN4
2019 A Hybrid Method for Powdered Materials Modeling
abstract
Powdered materials, such as sand and flour, are quite common in nature, whose properties always range from granular particles to smog materials under the air friction while throwing. This paper presents a hybrid method that tightly couples APIC solver with density field to accomplish the transformation of continuous powdered materials varying among granular particles, smog, powders and their natural mixtures. In our method, a part of the granular particles will be transformed to dust smog while interacting with air and represented by density field, then, as velocity decreases the density-based dust will deposit to powder particles. We construct a unified framework to imitate the mutual transformation process for the powdered materials of different scales, which greatly enhance the details of particle-based materials modeling. We have conducted extensive experiments to verify the performance of our model, and get satisfactory results in terms of stability, efficiency and visual authenticity as expected.
Yang Gao 0032, Shuai Li 0001, Aimin Hao, Hong Qin 0001
VRST4
2019 Quantitative and flexible 3D shape dataset augmentation via latent space embedding and deformation learning
Jiarui Liu 0003, Qing Xia 0002, Shuai Li 0001, Aimin Hao, Hong Qin 0001
Comput. Aided Geom. Des.4
2019 Learning multi-view manifold for single image based modeling
Jiahao Cui 0001, Shuai Li 0001, Qing Xia 0002, Aimin Hao, Hong Qin 0001
Comput. Graph.4
2019 Efficient 4D shape completion from sparse samples via cubic spline fitting in linear rotation-invariant space
Qing Xia 0002, Chengju Chen, Jiarui Liu 0003, Shuai Li 0001, Aimin Hao, Hong Qin 0001
Comput. Graph.5
2019 Hybrid 4D cardiovascular modeling based on patient-specific clinical images for real-time PCI surgery simulation
Shuai Li 0001, Zhijun Xie, Qing Xia 0002, Aimin Hao, Hong Qin 0001
Graph. Model.4
2019 Bidirectional Optimization Coupled Lightweight Networks for Efficient and Robust Multi-Person 2D Pose Estimation
Shuai Li 0001, Zheng Fang 0008, Wenfeng Song, Aimin Hao, Hong Qin 0001
J. Comput. Sci. Technol.4
2019 Multitask learning on monocular water images: Surface reconstruction and image synthesis
abstract
Abstract In this paper, we present a new strategy, a joint deep learning architecture, for two classic tasks in computer graphics: water surface reconstruction and water image synthesis. Modeling water surfaces from single images can be regarded as the inverse of image rendering, which converts surface geometries into photorealistic images. On the basis of this fact, we therefore consider these two problems as a cycle image‐to‐image translation and propose to tackle them together using a pair of neural networks, with the three‐dimensional surface geometries being represented as two‐dimensional surface normal maps. Furthermore, we also estimate the imaging parameters from the existing water images with a subnetwork to reuse the lighting conditions when synthesizing new images. Experiments demonstrate that our method achieves an accurate reconstruction of surfaces from monocular images efficiently and produces visually plausible new images under variable lighting conditions.
Xueguang Xie, Xiao Zhai, Fei Hou 0001, Aimin Hao, Hong Qin 0001
Comput. Animat. Virtual Worlds4
2019 Multitask Cascade Convolution Neural Networks for Automatic Thyroid Nodule Detection and Recognition
abstract
Thyroid ultrasonography is a widely used clinical technique for nodule diagnosis in thyroid regions. However, it remains difficult to detect and recognize the nodules due to low contrast, high noise, and diverse appearance of nodules. In today's clinical practice, senior doctors could pinpoint nodules by analyzing global context features, local geometry structure, and intensity changes, which would require rich clinical experience accumulated from hundreds and thousands of nodule case studies. To alleviate doctors' tremendous labor in the diagnosis procedure, we advocate a machine learning approach to the detection and recognition tasks in this paper. In particular, we develop a multitask cascade convolution neural network (MC-CNN) framework to exploit the context information of thyroid nodules. It may be noted that our framework is built upon a large number of clinically confirmed thyroid ultrasound images with accurate and detailed ground truth labels. Other key advantages of our framework result from a multitask cascade architecture, two stages of carefully designed deep convolution networks in order to detect and recognize thyroid nodules in a pyramidal fashion, and capturing various intrinsic features in a global-to-local way. Within our framework, the potential regions of interest after initial detection are further fed to the spatial pyramid augmented CNNs to embed multiscale discriminative information for fine-grained thyroid recognition. Experimental results on 4309 clinical ultrasound images have indicated that our MC-CNN is accurate and effective for both thyroid nodules detection and recognition. For the correct diagnosis rate of malignant and benign thyroid nodules, its mean Average Precision (mAP) performance can achieve up to [Formula: see text] accuracy, which outperforms the common CNNs by [Formula: see text] on average. In addition, we conduct rigorous user studies to confirm that our MC-CNN outperforms experienced doctors, yet only consuming roughly [Formula: see text] ( 1/48) of doctors' examination time on average. Therefore, the accuracy and efficiency of our new method exhibit its great potential in clinical applications.
Wenfeng Song, Shuai Li 0001, Hong Qin 0001, Aimin Hao
IEEE J. Biomed. Health Informatics7
2019 An efficient FLIP and shape matching coupled method for fluid-solid and two-phase fluid simulations
Yang Gao 0032, Shuai Li 0001, Hong Qin 0001, Aimin Hao
Vis. Comput.5
2018 A Novel Radiogenomics Framework for Genomic and Image Feature Correlation using Deep Learning
Shuai Li 0001, Hongze Han, Dong Sui, Aimin Hao, Hong Qin 0001
BIBM4
2018 High-fidelity Compression of Dynamic Meshes with Fine Details using Piece-wise Manifold Harmonic Bases
abstract
Mesh-based animation, usually represented as dynamic meshes with fixed connectivity, is becoming more and more prevalent in movies, games and other graphics applications nowadays, and there is a growing need to compactly store and rapidly transmit these meshes for practical use, especially for those with high-quality geometric details. In this paper, we explore a novel key-frame based dynamic mesh compression method, wherein we apply pose-similarity with spectral techniques to define piece-wise manifold harmonic bases to reduce spatial-temporal redundancy. We first partition the sequence into several clusters with similar poses, and then decompose the meshes in each cluster into primary poses and geometric details using the manifold harmonic bases derived from the extracted key-frame in that cluster. The primary poses can be characterized as linear combinations of manifold harmonic bases, and the geometric details can be recovered by deformation transfer technique. Thus, we only need a small number of key-frames and a few coefficients for compressing dynamic meshes, which saves a significant amount of storage comparing to traditional methods in which bases are stored explicitly. Furthermore, we apply a second-order linear prediction coding to the harmonic coefficients to further reduce the temporal redundancy. Our extensive experiments and evaluations on various datasets have manifested that our novel method could obtain a high compression ratio while preserving high-fidelity geometry details and guaranteeing limited human perceived distortion rate simultaneously.
Chengju Chen, Qing Xia 0002, Shuai Li 0001, Hong Qin 0001, Aimin Hao
CGI5
2018 Automatic Beautification for Group-Photo Facial Expressions Using Novel Bayesian GANs
Shuai Li 0001, Wenfeng Song, Hong Qin 0001, Aimin Hao
ICANN (1)6
2018 Learning from Weakly-Labeled Clinical Data for Automatic Thyroid Nodule Classification in Ultrasound Images
abstract
This paper proposes a semi-supervised learning method based on weakly-labeled data to automatically classify ultrasound (US) thyroid nodules. Key to our new approach is the unification of multi-instance learning (MIL) with deep learning. Benefiting from that, our method can directly use off-the-shelf clinical data, which involves no labels to indicate nodule classes. To this end, we take the US images of a patient as a bag, and take the corresponding pathology report as the bag label. Specifically, we first propose a bag generating method, wherein the detected thyroid nodules are considered as instances corresponding to certain bag. After that, we design an effective EM algorithm to train a convolutional neural network (CNN) for nodule classification. We conduct extensive experiments and comprehensive evaluations on different datasets, and all the experiments confirm that, our method significantly outperforms state-of-the-art MIL algorithms, which exhibits great potential in clinical applications.
Jianxiong Wang, Shuai Li 0001, Wenfeng Song, Hong Qin 0001, Aimin Hao
ICIP6
2018 Feature-preserving, mesh-free empirical mode decomposition for point clouds and its applications
Lixin Guo 0003, Dongbo Zhang 0004, Hong Qin 0001, Aimin Hao
Comput. Aided Geom. Des.6
2018 Multi-scale geometry detail recovery on surfaces via Empirical Mode Decomposition
Dongbo Zhang 0004, Lixin Guo 0003, Hong Qin 0001, Aimin Hao
Comput. Graph.6
2018 Hybrid-feature-guided lung nodule type classification on CT images
Jingjing Yuan, Xinglong Liu, Fei Hou 0001, Hong Qin 0001, Aimin Hao
Comput. Graph.5
2018 Deep variance network: An iterative, improved CNN framework for unbalanced training datasets
Shuai Li 0001, Wenfeng Song, Hong Qin 0001, Aimin Hao
Pattern Recognit.4
2018 Multi-view multi-scale CNNs for lung nodule type classification from CT images
Xinglong Liu, Fei Hou 0001, Hong Qin 0001, Aimin Hao
Pattern Recognit.4
2018 A Novel Bottom-Up Saliency Detection Method for Video With Dynamic Background
abstract
After years of extensive studies, the salient motion detection problem has gained plausible performance improvement that was primarily propelled by the rapid development of self-adaptive top-down modeling techniques. Nevertheless, almost all the conventional solutions are still not robust enough to handle video sequences captured by hand-hold cameras. This is mainly due to the absence of the position alignment information that is indispensable for top-down background modeling. In contrast, the bottom-up video saliency detection methods, though achieving excellent salient motion detection in either stationary or nonstationary videos, still have rather poor detection performance in scenarios with massive dynamic background. In this letter, we explore a bottom-up saliency framework by introducing a novel spatial-temporal regional filter method to handle the dynamic background problem. Our key rationale is to assign large saliency value to those regions with stable spatial-temporal coherency while eliminating irregular, repeating dynamic background. As far as we know, this is the first work to address the dynamic background problem from the perspective of the bottom-up video saliency. We conduct massive quantitative evaluations over public available benchmarks to validate the effectiveness and robustness of our method.
Chenglizhao Chen, Shuai Li 0001, Hong Qin 0001, Aimin Hao
IEEE Signal Process. Lett.5
2018 Real-time dissection of organs via hybrid coupling of geometric metaballs and physics-centric mesh-free method
JunJun Pan, Shizeng Yan, Hong Qin 0001, Aimin Hao
Vis. Comput.4
2017 An Extended Type Cell Detection and Counting Method based on FCN
abstract
Cell detection and counting are critical and essential tasks for many biological and clinical studies. Traditionally, these tasks are usually performed by visual inspection, which is time consuming and prone to induce subjective bias. These make automatic cell counting and detection essential for large- scale and objective studies. Unfortunately, the hard examples such as cell blur, clutter, bleed-through and imaging noise make these tasks extremely challenging. Over the last few years, automatic cell detection and counting have evolved from earlier methods that are often based on filters to the current state-of- the-art deep learning methods. In this paper, we propose a novel efficient method for robust counting and detection task based on fully convolution networks (FCN). Our method is able to handle most of detection and counting problems from different kinds of cell datasets, and can cover most senior microscopy images, such as bright field, pathology stained material and electron. Extensive experiments on the public and private datasets demonstrate the effectiveness and reliability of our approach.
Runkai Zhu, Dong Sui, Hong Qin 0001, Aimin Hao
BIBE4
2017 A novel fluid-solid coupling framework integrating FLIP and shape matching methods
abstract
Physically-based fluid animation and solid deformation driven by numerical simulation have manifested their significance for many graphics applications during the past two decades. For example, the fluid implicit particle (FLIP) method and shape matching technique based on position based dynamics (PBD) have demonstrated their unique graphics strength in fluid and solid animation, respectively. We propose a novel integrated approach supporting the seamless unification of FLIP and shape matching. We devise new algorithms to tackle existing difficulties when handling new phenomena such as high-fidelity fluid-solid interaction and solid melting. The key innovation of this paper is a unified Lagrangian framework that seamlessly blends FLIP and PBD based shape matching constraint towards the natural yet strong coupling between fluid and deformable solid. Within our integrated framework, it enables many complicated fluid-solid phenomena with ease. We conduct various kinds of experiments. All the results demonstrate the advantages of our unified hybrid approach towards visual fidelity, efficiency, stability, and versatility.
Yang Gao 0032, Shuai Li 0001, Hong Qin 0001, Aimin Hao
CGI4
2017 Hessian-constrained detail-preserving 3D implicit reconstruction from raw volumetric dataset
Shuai Li 0001, Dehui Yan, Aimin Hao, Hong Qin 0001
Comput. Graph.4
2017 Inverse Modelling of Incompressible Gas Flow in Subspace
abstract
Abstract This paper advocates a novel method for modelling physically realistic flow from captured incompressible gas sequence via modal analysis in frequency‐constrained subspace. Our analytical tool is uniquely founded upon empirical mode decomposition (EMD) and modal reduction for fluids, which are seamlessly integrated towards a powerful, style‐controllable flow modelling approach. We first extend EMD, which is capable of processing 1D time series but has shown inadequacies for 3D graphics earlier, to fit gas flows in 3D. Next, frequency components from EMD are adopted as candidate vectors for bases of modal reduction. The prerequisite parameters of the Navier–Stokes equations are then optimized to inversely model the physically realistic flow in the frequency‐constrained subspace. The estimated parameters can be utilized for re‐simulation, or be altered toward fluid editing. Our novel inverse‐modelling technique produces real‐time gas sequences after precomputation, and is convenient to couple with other methods for visual enhancement and/or special visual effects. We integrate our new modelling tool with a state‐of‐the‐art fluid capturing approach, forming a complete pipeline from real‐world fluid to flow re‐simulation and editing for various graphics applications.
Xiao Zhai, Fei Hou 0001, Hong Qin 0001, Aimin Hao
Comput. Graph. Forum4
2017 A CADe system for nodule detection in thoracic CT images based on artificial neural network
Xinglong Liu, Fei Hou 0001, Hong Qin 0001, Aimin Hao
Sci. China Inf. Sci.4
2017 An efficient heat-based model for solid-liquid-gas phase transition and dynamic interaction
Yang Gao 0032, Shuai Li 0001, Lipeng Yang, Hong Qin 0001, Aimin Hao
Graph. Model.5
2017 Novel fluid detail enhancement based on multi-layer depth regression analysis and FLIP fluid simulation
abstract
Abstract In this paper, we propose a novel integrated method for effective modeling and realistic enhancement of scale‐sensitive fluid simulation details. The core of our method is the organic of multi‐layer depth image regression analysis and fluid implicit particle fluid simulation of which the regression analysis induces the criterion where the fluid details should be produced. First, we capture the depth buffer of the fluid surface dynamically from the top of scene. Second, we employ depth peeling technique to decompose the target fluid volume into multiple depth layers and conduct time‐space analysis over surface layers. Third, we propose a logistic regression‐based model to rigorously pinpoint the complex interacting regions, wherein multiple detail‐relevant factors are taken into account based on the captured multiple depth layers. Finally, details are enhanced by animating extra diffuse materials and augmenting the air‐fluid mixing phenomenon. It is evident that, with depth peeling technology, we can afford rigorous analysis not only across surface layers at different fluid depth but along the depth direction as well. After integrating the analysis results from these two sources, we are capable of performing detail enhancement both on the fluid surface and inside the fluid to obtain a great visual effect, even when large occlusion exists. Directly benefiting from the flexibility of image‐space‐dominant processing, our unified framework can be entirely implemented on graphics processing units and thus achieves interactive performance. For various fluid phenomena with different diffuse materials (e.g., spray, foam, and bubble), comprehensive experiments and evaluations have demonstrated its superiority in high‐fidelity fluid detail enhancement and its interaction with surrounding environment.
Yuxing Qiu, Lipeng Yang, Shuai Li 0001, Qing Xia 0002, Hong Qin 0001, Aimin Hao
Comput. Animat. Virtual Worlds6
2017 Video Saliency Detection via Spatial-Temporal Fusion and Low-Rank Coherency Diffusion
abstract
This paper advocates a novel video saliency detection method based on the spatial-temporal saliency fusion and low-rank coherency guided saliency diffusion. In sharp contrast to the conventional methods, which conduct saliency detection locally in a frame-by-frame way and could easily give rise to incorrect low-level saliency map, in order to overcome the existing difficulties, this paper proposes to fuse the color saliency based on global motion clues in a batch-wise fashion. And we also propose low-rank coherency guided spatial-temporal saliency diffusion to guarantee the temporal smoothness of saliency maps. Meanwhile, a series of saliency boosting strategies are designed to further improve the saliency accuracy. First, the original long-term video sequence is equally segmented into many short-term frame batches, and the motion clues of the individual video batch are integrated and diffused temporally to facilitate the computation of color saliency. Then, based on the obtained saliency clues, inter-batch saliency priors are modeled to guide the low-level saliency fusion. After that, both the raw color information and the fused low-level saliency are regarded as the low-rank coherency clues, which are employed to guide the spatial-temporal saliency diffusion with the help of an additional permutation matrix serving as the alternative rank selection strategy. Thus, it could guarantee the robustness of the saliency map's temporal consistence, and further boost the accuracy of the computed saliency map. Moreover, we conduct extensive experiments on five public available benchmarks, and make comprehensive, quantitative evaluations between our method and 16 state-of-the-art techniques. All the results demonstrate the superiority of our method in accuracy, reliability, robustness, and versatility.
Chenglizhao Chen, Shuai Li 0001, Yongguang Wang, Hong Qin 0001, Aimin Hao
IEEE Trans. Image Process.5
2017 Unsupervised Multi-Class Co-Segmentation via Joint-Cut Over L1 -Manifold Hyper-Graph of Discriminative Image Regions
abstract
This paper systematically advocates a robust and efficient unsupervised multi-class co-segmentation approach by leveraging underlying subspace manifold propagation to exploit the cross-image coherency. It can combat certain image co-segmentation difficulties due to viewpoint change, partial occlusion, complex background, transient illumination, and cluttering texture patterns. Our key idea is to construct a powerful hyper-graph joint-cut framework, which incorporates mid-level image regions-based intra-image feature representation and L1-manifold graph-based inter-image coherency exploration. For local image region generation, we propose a bi-harmonic distance distribution difference metric to govern the super-pixel clustering in a bottom-up way. It not only affords drastic data reduction but also gives rise to discriminative and structure meaningful feature representation. As for the inter-image coherency, we leverage multi-type features involved L1-graph to detect the underlying local manifold from cross-image regions. As a result, the implicit supervising information could be encoded into the unsupervised hyper-graph joint-cut framework. We conduct extensive experiments and make comprehensive evaluations with other state-of-the-art methods over various benchmarks, including iCoseg, MSRC, and Oxford flower. All the results demonstrate the superiorities of our method in terms of accuracy, robustness, efficiency, and versatility.
Jizhou Ma, Shuai Li 0001, Hong Qin 0001, Aimin Hao
IEEE Trans. Image Process.4
2017 Knot Optimization for Biharmonic B-splines on Manifold Triangle Meshes
abstract
Biharmonic B-splines, proposed by Feng and Warren, are an elegant generalization of univariate B-splines to planar and curved domains with fully irregular knot configuration. Despite the theoretic breakthrough, certain technical difficulties are imperative, including the necessity of Voronoi tessellation, the lack of analytical formulation of bases on general manifolds, expensive basis re-computation during knot refinement/removal, being applicable for simple domains only (e.g., such as euclidean planes, spherical and cylindrical domains, and tori). To ameliorate, this paper articulates a new biharmonic B-spline computing paradigm with a simple formulation. We prove that biharmonic B-splines have an equivalent representation, which is solely based on a linear combination of Green's functions of the bi-Laplacian operator. Consequently, without explicitly computing their bases, biharmonic B-splines can bypass the Voronoi partitioning and the discretization of bi-Laplacian, enable the computational utilities on any compact 2-manifold. The new representation also facilitates optimization-driven knot selection for constructing biharmonic B-splines on manifold triangle meshes. We develop algorithms for spline evaluation, data interpolation and hierarchical data decomposition. Our results demonstrate that biharmonic B-splines, as a new type of spline functions with theoretic and application appeal, afford progressive update of fully irregular knots, free of singularity, without the need of explicit parameterization, making it ideal for a host of graphics tasks on manifolds.
Fei Hou 0001, Ying He 0001, Hong Qin 0001, Aimin Hao
IEEE Trans. Vis. Comput. Graph.4
2016 Detail-Preserving 3D Shape Modeling from Raw Volumetric Dataset via Hessian-Constrained Local Implicit Surfaces Optimization
abstract
Massive routinely-acquired raw volumetric datasets are hard to be deeply exploited by cyber worlds related downstream applications due to the challenges in accurate and efficient shape modeling. This paper systematically advocates an interactive 3D shape modeling framework for raw volumetric datasets by iteratively optimizing Hessian-constrained local implicit surfaces. The key idea is to incorporate contour based interactive segmentation into the generalized local implicit surface reconstruction. Our framework allows a user to flexibly define derivative constraints up to the second order via intuitively placing contours on the cross sections of volumetric images and fine-tuning the eigenvector frame of Hessian matrix. It enables detail-preserving local implicit representation while combating certain difficulties due to ambiguous image regions, low-quality irregular data, close sheets, and massive coefficients involved extra computing burden. Moreover, we conduct extensive experiments on some volumetric images with blurry object boundaries, and make comprehensive, quantitative performance evaluation between our method and the state-of-the-art radial basis function based techniques. All the results demonstrate our method's advantages in the accuracy, detail-preserving, efficiency, and versatility of shape modeling.
Shuai Li 0001, Dehui Yan, Aimin Hao, Hong Qin 0001
CW4
2016 Automatic extraction of generic focal features on 3D shapes via random forest regression analysis of geodesics-in-heat
Qing Xia 0002, Shuai Li 0001, Hong Qin 0001, Aimin Hao
Comput. Aided Geom. Des.4
2016 Coupling time-varying modal analysis and FEM for real-time cutting simulation of objects with multi-material sub-domains
Chen Yang 0002, Shuai Li 0001, Lili Wang 0006, Aimin Hao, Hong Qin 0001
Comput. Aided Geom. Des.5
2016 Haptics-equiped interactive PCI simulation for patient-specific surgery training and rehearsing
Shuai Li 0001, Qing Xia 0002, Aimin Hao, Hong Qin 0001, Qinping Zhao
Sci. China Inf. Sci.3
2016 Automatic non-parametric image parsing via hierarchical semantic voting based on sparse-dense reconstruction and spatial-contextual cues
Xinyi An, Shuai Li 0001, Hong Qin 0001, Aimin Hao
Neurocomputing4
2016 Robust salient motion detection in non-stationary videos via novel integrated strategies of spatio-temporal coherency clues and low-rank analysis
Chenglizhao Chen, Shuai Li 0001, Hong Qin 0001, Aimin Hao
Pattern Recognit.4
2016 Super-Resolution of Multi-Observed RGB-D Images Based on Nonlocal Regression and Total Variation
abstract
There is growing demand for accuracy in image processing and visualization, and the super-resolution (SR) technique for multi-observed RGB-D images has become popular, because it provides space-redundant information and produces a detailed reconstruction even with a large magnification factor. This technique has been thoroughly investigated in recent years. Nevertheless, technical challenges remain, such as finding sub-pixel correspondences with low-resolution (LR) observations, exploiting space-redundant information, formulating space homogeneity constraints, and leveraging cross-image similarities in structures. To address these challenges, this paper proposes a unified optimization framework to estimate both the super-resolved RGB image and the super-resolved depth image from the multi-observed LR RGB-D images using their correlations. Using depth-assisted cross-image correspondences, the RGB image SR problem is formulated as an effective regularization function by incorporating the normalized bilateral total variation regularizer, and it is efficiently solved by a first-order primal-dual algorithm. The depth image SR estimate can be obtained by minimizing a nonlocal regression-based energy, which integrates the structural cues of the super-resolved RGB image in a detail-preserving fashion. Essentially, our unified optimization framework uses the RGB image and depth image as a priori knowledge that the SR process uses for better accuracy. Our extensive experiments on public RGB-D benchmarks and real data and our quantitative comparison with several state-of-the-art methods demonstrate the superiority of our method in terms of accuracy, versatility, and reliability of details and sharp feature preservation.
Qingzheng Wang, Shuai Li 0001, Hong Qin 0001, Aimin Hao
IEEE Trans. Image Process.4
2016 Robust Optimization-Based Coronary Artery Labeling From X-Ray Angiograms
abstract
In this paper, we present an efficient robust labeling method for coronary arteries from X-ray angiograms based on energy optimization. The fundamental goal of this research is to facilitate the analysis and diagnosis of interventional surgery in the most efficient way, and such effort could also improve the performance during doctor training, and surgery simulation and planning. Compared to the prior state-of-the-art, our method is much more robust to resist noises and is tolerant to even incomplete data because of the "built-in" nature of global optimization. We start with a fully parallelized algorithm based on Hessian matrix to extract the tubular structure from the X-ray angiograms as vessel candidates. Then, instead of using the candidates directly, we use the grow cut (Vezhnevets and V. Konouchine, Growcut: Interactive multi-label N-D image segmentation by cellular automata, in Proc. of Graphicon, 2005, pp. 150-156.) method, which is similar to graph cut (Boykov et al. , Fast approximate energy minimization via graph cuts, IEEE Trans. Pattern Anal. Mach. Intell. , vol. 23, no. 11, pp. 1222-1239, Nov. 2001.)but with better performance to extract the precise vessel structure from the images. Next, we use the fast marching method with second derivatives and cross neighbors to extract the accurate skeleton segments. After that, we propose an efficient method based on iterative closest point (Z. Zhang, Iterative point matching for registration of free-form curves and surfaces, Int J. Comput. Vis., vol. 13, no. 2, pp. 119-152, 1994.) to organize the skeleton segments by treating the continuity and similarity as extra constraints. Finally, we formulate the vessel labeling problem as an energy optimization problem and solve it using belief propagation. We also demonstrate several typical applications including flow velocity estimation, heart beat estimation, and vessel diameter estimation to show its practical uses in clinical diagnosis and treatment. Our experiments exhibit the correctness and robustness, as well as the high performance of our algorithm. We envision that our system would be of high utility for diagnosis and therapy to treat vessel-related diseases in a clinical setting in the near future.
Xinglong Liu, Fei Hou 0001, Hong Qin 0001, Aimin Hao
IEEE J. Biomed. Health Informatics4
2015 Novel, Robust, and Efficient Guidewire Modeling for PCI Surgery Simulator Based on Heterogeneous and Integrated Chain-Mails
abstract
Despite the long R&D history of interactive minimally-invasive surgery and therapy simulations, the guide wire/catheter behavior modeling remains challenging in Percutaneous Coronary Intervention (PCI) surgery simulators. This is primarily due to the heterogeneous heart physiological structures and complex intravascular inter-dynamic procedures. To ameliorate, this paper advocates a novel, robust, and efficient guide wire/catheter modeling method based on heterogeneous and integrated chain-mails, that can afford medical practitioners and trainees the unique opportunity to experience the entire guide wire-dominant PCI procedures in virtual environments as our model aims to mimic what occurs in clinical settings. Our approach's originality is primarily founded upon this new method's unconditional stability, real time performance, flexibility, and high-fidelity realism for guide wire/catheter simulation. Considering the front end of the guide wire has different stiffness with its conjunctive slender body and the guide wire length is adaptive to the surrounding environment, we propose to model the spatially-varying six-degree of freedom behaviors by solely resorting to the generalized 3D chain-mails. Meanwhile, to effectively accommodate the motion constraints caused by the beating vessels and flowing blood, we integrate heterogeneous volumetric chain mails to streamline guide wire modeling and its interaction with surrounding substances. By dynamically coupling guide wire chain-mails with the surrounding media via virtual links, we are capable of efficiently simulating the collision-involved interdynamic behaviors of the guide wire. Finally, we showcase a PCI prototype simulator equipped with hap tic feedback for mimicing the guide wire intervention therapy, including pushing, pulling, and twisting operations, where the built-in high-fidelity, real-time efficiency, and stableness show great promise for its practical applications in clinical training and surgery rehearsal fields.
Shuai Li 0001, Hong Qin 0001, Aimin Hao
CAD/Graphics4
2015 Interactive volumetric segmentation through least-squares optimization of local hessian-constrained implicits
abstract
A great number of volumetric datasets have been routinely acquired everyday and their qualities are varying tremendously, without proper processing they could not be directly utilized. Specifically, volumetric segmentation plays a vital role in many downstream applications, including geometric modeling, scientific visualization, and medical diagnosis. So far, many volume segmentation methods have been proposed, Top et al. [2011] designed an interactive segmentation tool by interactively contouring on some sparse slices and Ijiri et al. [2013] developed a system to extract contours and evaluate the scalar field in spatial domain.
Shuai Li 0001, Aimin Hao, Hong Qin 0001
VRST3
2015 A novel integrated analysis-and-simulation approach for detail enhancement in FLIP fluid interaction
abstract
This paper advocates a novel integrated method to tightly couple simulation with analysis for the effective modeling and enhancement of scale-aware fluid details. It brings forth a suite of innovations in a unified framework, including depth-image-based space analysis for multi-scale detail detection, time-space analysis based on the logistic regression model that integrates both geometry and physics criteria, and depth-image-based sampling for quality-efficiency tradeoff. Our method contains an intertwined two-level processing architecture at its core. At the analysis level, we propose a rigorous time-space analysis model to pinpoint complex interacting regions, which can take into account multiple detail-relevant factors based on the depth-image sequence captured from FLIP-driven simulation sequence. At the simulation level, details are enhanced by animating extra diffuse materials, and augmenting the air-fluid mixing phenomenon. Directly benefitting from the flexibility of image-space-dominant processing, our unified framework can be entirely implemented on GPU, hence interactive performance could be guaranteed. Comprehensive experiments and evaluations on various diffuse phenomena (e.g., spray, foam, and bubble) have demonstrated its superiority in high-fidelity detail enhancement during fluid simulation and its interaction with surrounding environment for VR applications.
Lipeng Yang, Shuai Li 0001, Qing Xia 0002, Hong Qin 0001, Aimin Hao
VRST5
2015 Trivariate Biharmonic B-Splines
abstract
Abstract In this paper, we formulate a novel trivariate biharmonic B‐spline defined over bounded volumetric domain. The properties of bi‐Laplacian have been well investigated, but the straightforward generalization from bivariate case to trivariate one gives rise to unsatisfactory discretization, due to the dramatically uneven distribution of neighbouring knots in 3D. To ameliorate, our original idea is to extend the bivariate biharmonic B‐spline to the trivariate one with novel formulations based on quadratic programming, approximating the properties of localization and partition of unity. And we design a novel discrete biharmonic operator which is optimized more robustly for a specific set of functions for unevenly sampled knots compared with previous methods. Our experiments demonstrate that our 3D discrete biharmonic operators are robust for unevenly distributed knots and illustrate that our algorithm is superior to previous algorithms .
Fei Hou 0001, Hong Qin 0001, Aimin Hao
Comput. Graph. Forum3
2015 Real-time haptic manipulation and cutting of hybrid soft tissue models by extended position-based dynamics
abstract
Abstract This paper systematically describes an interactive dissection approach for hybrid soft tissue models governed by extended position‐based dynamics. Our framework makes use of a hybrid geometric model comprising both surface and volumetric meshes. The fine surface triangular mesh with high‐precision geometric structure and texture at the detailed level is employed to represent the exterior structure of soft tissue models. Meanwhile, the interior structure of soft tissues is constructed by coarser tetrahedral mesh, which is also employed as physical model participating in dynamic simulation. The less details of interior structure can effectively reduce the computational cost during simulation. For physical deformation, we design and implement an extended position‐based dynamics approach that supports topology modification and material heterogeneities of soft tissue. Besides stretching and volume conservation constraints, it enforces the energy preserving constraints, which take the different spring stiffness of material into account and improve the visual performance of soft tissue deformation. Furthermore, we develop mechanical modeling of dissection behavior and analyze the system stability. The experimental results have shown that our approach affords real‐time and robust cutting without sacrificing realistic visual performance. Our novel dissection technique has already been integrated into a virtual reality‐based laparoscopic surgery simulator. Copyright © 2015 John Wiley & Sons, Ltd.
JunJun Pan, Junxuan Bai, Xin Zhao 0025, Aimin Hao, Hong Qin 0001
Comput. Animat. Virtual Worlds4
2015 Real-time and robust object tracking in video via low-rank coherency analysis in feature space
Chenglizhao Chen, Shuai Li 0001, Hong Qin 0001, Aimin Hao
Pattern Recognit.4
2015 Structure-Sensitive Saliency Detection via Multilevel Rank Analysis in Intrinsic Feature Space
abstract
This paper advocates a novel multiscale, structure-sensitive saliency detection method, which can distinguish multilevel, reliable saliency from various natural pictures in a robust and versatile way. One key challenge for saliency detection is to guarantee the entire salient object being characterized differently from nonsalient background. To tackle this, our strategy is to design a structure-aware descriptor based on the intrinsic biharmonic distance metric. One benefit of introducing this descriptor is its ability to simultaneously integrate local and global structure information, which is extremely valuable for separating the salient object from nonsalient background in a multiscale sense. Upon devising such powerful shape descriptor, the remaining challenge is to capture the saliency to make sure that salient subparts actually stand out among all possible candidates. Toward this goal, we conduct multilevel low-rank and sparse analysis in the intrinsic feature space spanned by the shape descriptors defined on over-segmented super-pixels. Since the low-rank property emphasizes much more on stronger similarities among super-pixels, we naturally obtain a scale space along the rank dimension in this way. Multiscale saliency can be obtained by simply computing differences among the low-rank components across the rank scale. We conduct extensive experiments on some public benchmarks, and make comprehensive, quantitative evaluation between our method and existing state-of-the-art techniques. All the results demonstrate the superiority of our method in accuracy, reliability, robustness, and versatility.
Chenglizhao Chen, Shuai Li 0001, Hong Qin 0001, Aimin Hao
IEEE Trans. Image Process.4
2015 A parallelized 4D reconstruction algorithm for vascular structures and motions based on energy optimization
Xinglong Liu, Fei Hou 0001, Aimin Hao, Hong Qin 0001
Vis. Comput.3
2015 Metaballs-based physical modeling and deformation of organs for virtual surgery
JunJun Pan, Chengkai Zhao, Xin Zhao 0025, Aimin Hao, Hong Qin 0001
Vis. Comput.4
2014 Real-time path planning in emergency using non-uniform safety fields
abstract
We present a novel approach to direct virtual evacuees in emergency using non-uniform safety fields. The safety fields uses safety level to describe the danger degrees corresponding to the current situation of their positions. Our method also uses spread potential to generate smooth escape routes. We also combine our method with the most current agent-based collision avoidance algorithm to prevent oscillation. In practice, our algorithm can perform real-time navigation for dozens of evacuees in emergency scenarios.
Bangrui Liu, Aimin Hao
VR2
2014 Dissection of hybrid soft tissue models using position-based dynamics
abstract
This paper describes an interactive dissection approach for hybrid soft tissue models governed by position-based dynamics. Our framework makes use of a hybrid geometric model comprising both surface and volumetric meshes. The fine surface triangular mesh is used to represent the exterior structure of soft tissue models. Meanwhile, the interior structure of soft tissues is constructed by coarser tetrahedral meshes, which are also employed as physical models participating in dynamic simulation. The less details of interior structure can effectively reduce the computational cost of deformation and geometric subdivision during dissection. For physical deformation, we design and implement a position-based dynamics approach that supports topology modification and enforces the volume-preserving constraint. Experimental results have shown that, this hybrid dissection method affords real-time and robust cutting simulation without sacrificing realistic visual performance.
JunJun Pan, Junxuan Bai, Xin Zhao 0025, Aimin Hao, Hong Qin 0001
VRST4
2014 Hybrid Particle-grid Modeling for Multi-scale Droplet/Spray Simulation
abstract
Abstract This paper presents a novel hybrid particle‐grid method that tightly couples Lagrangian particle approach with Eulerian grid approach to simulate multi‐scale diffuse materials varying from disperse droplets to dissipating spray and their natural mixture and transition, originated from a violent (high‐speed) liquid stream. Despite the fact that Lagrangian particles are widely employed for representing individual droplets and Eulerian grid‐based method is ideal for volumetric spray modeling, using either one alone has encountered tremendous difficulties when effectively simulating droplet/spray mixture phenomena with high fidelity. To ameliorate, we propose a new hybrid model to tackle such challenges with many novel technical elements. At the geometric level, we employ the particle and density field to represent droplet and spray respectively, modeling their creation from liquid as well as their seamless transition. At the physical level, we introduce a drag force model to couple droplets and spray, and specifically, we employ Eulerian method to model the interaction among droplets and marry it with the widely‐used Lagrangian model. Moreover, we implement our entire hybrid model on CUDA to guarantee the interactive performance for high‐effective physics‐based graphics applications. The comprehensive experiments have shown that our hybrid approach takes advantages of both particle and grid methods, with convincing graphics effects for disperse droplets and spray simulation.
Lipeng Yang, Shuai Li 0001, Aimin Hao, Hong Qin 0001
Comput. Graph. Forum3
2014 Interactive texture design and synthesis from mesh sketches
Lili Wang 0006, Qinglin Qi, Wei Ke 0001, Aimin Hao
Frontiers Comput. Sci.5
2014 Interactive deformation and cutting simulation directly using patient-specific volumetric images
abstract
ABSTRACT This paper systematically advocates an interactive volumetric image manipulation framework, which can enable the rapid deployment and instant utility of patient‐specific medical images in virtual surgery simulation while requiring little user involvement. We seamlessly integrate multiple technical elements to synchronously accommodate physics‐plausible simulation and high‐fidelity anatomical structures visualization. Given a volumetric image, in a user‐transparent way, we build a proxy to represent the geometrical structure and encode its physical state without the need of explicit 3‐D reconstruction. On the basis of the dynamic update of the proxy, we simulate large‐scale deformation, arbitrary cutting, and accompanying collision response driven by a non‐linear finite element method. By resorting to the upsampling of the sparse displacement field resulted from non‐linear finite element simulation, the cut/deformed volumetric image can evolve naturally and serves as a time‐varying 3‐D texture to expedite direct volume rendering. Moreover, our entire framework is built upon CUDA (Beihang University, Beijing, China) and thus can achieve interactive performance even on a commodity laptop. The implementation details, timing statistics, and physical behavior measurements have shown its practicality, efficiency, and robustness. Copyright © 2013 John Wiley & Sons, Ltd.
Shuai Li 0001, Qinping Zhao, Shengfa Wang, Aimin Hao, Hong Qin 0001
Comput. Animat. Virtual Worlds4
2014 Real-time physical deformation and cutting of heterogeneous objects via hybrid coupling of meshless approach and finite element method
abstract
ABSTRACT This paper advocates a method for real‐time physical deformation and arbitrary cutting simulation of heterogeneous objects with multi‐material distribution, whose originality centers on the tight coupling of domain‐specific finite element method (FEM) and material distance‐aware meshless approach in a CUDA‐centric parallel simulation framework. We employ hierarchical hexahedron serving as basic building blocks for accurate material‐aware FEM simulation. Meanwhile, local meshless systems are designed to support cross‐FEM‐domain coupling and material‐sensitive propagation while respecting the regularity of finite elements. Directly benefiting from the structural regularity and uniformity of finite elements, our hybrid solution enables the local stiffness matrix pre‐computation and dynamic assembling, adaptive topological updating and precise cutting reconstruction. Moreover, our mathematically‐rigorous solver guarantees unconditional stableness. Experiments demonstrate the superiorities of our system. Copyright © 2014 John Wiley & Sons, Ltd.
Chen Yang 0002, Shuai Li 0001, Lili Wang 0006, Aimin Hao, Hong Qin 0001
Comput. Animat. Virtual Worlds4
2013 Efficient 3D Reconstruction of Vessels from Multi-views of X-Ray Angiography
abstract
In this paper, we present an efficient 3D vessels reconstruction algorithm based on multi-views of X-ray Angiography assisting interventional surgery. First, we extract the vascular-like structures from the image sequences using a geometrical analysis of multi-scale Hessian matrix eigen-system and use the fast marching method to extract the skeleton of the structure, from which we derive the vascular topological configurations. Second, we regard the 3D space as a Markov Random Field and formulate the reconstruction problem as an energy minimization problem with consistent, continuous and topological constraints to coarsely register and reconstruct the 3D vessels. Third, we refine the reconstructed vessels to register and reconstruct the 3D vessels accurately. We demonstrate our system in coronary arteries reconstruction for percutaneous coronary intervention surgery to help doctors learn about the configurations of the coronary arteries of specific patient during operation. We envision that our system will be used for clinic treatment to advance vessel reconstruction for diagnosis and therapy in the near future.
Xinglong Liu, Fei Hou 0001, Shuai Li 0001, Aimin Hao, Hong Qin 0001
CAD/Graphics4
2013 Direct Extraction of Feature Curves from Volume Image for Illustration and Vectorization Based on 2D/3D Curve Mapping
abstract
This paper proposes a parallel and direct semantic feature curve extraction method from 3D volume image for vectorization and illustration. Our approach is motivated by reconstructing 3D geometric information from multiple rendered images under multi-view in computer vision. The 2D rendered images are rich in the visual sense by color and opacity that convey the structure of volume data, so it is significant for the user to understand the structure of 3D volume data better if we can recover feature curves from those 2D images. Compared with conventional line extraction methods, which mainly focus on extracting feature curves from iso-surfaces in object space, we extract feature curves directly from volume images. Most of the computation can be computed in parallel on GPU with CUDA acceleration.
Lili Wang 0006, Fei Hou 0001, Aimin Hao, Hong Qin 0001
CAD/Graphics4
2013 ROI-Emphasized Volume Visualization Guided by Anisotropic Structure Tensor
abstract
Most of Focus Context visualization methods differentiate the magnification unit only by simply assigning each voxel/cell with an importance value while ignoring the shape content embedded in the volume data. In this paper, we take the volumetric structure information as important cue to facilitate Focus Context visualization, which can homogeneously or non-homogeneously scale the volume data in a structure-sensitive way.
Fei Hou 0001, Shuai Li 0001, Aimin Hao, Hong Qin 0001
CAD/Graphics4
2013 Multi-scale, multi-level, heterogeneous features extraction and classification of volumetric medical images
abstract
This paper articulates a novel method for the heterogeneous feature extraction and classification directly on volumetric images, which covers multi-scale point feature, multi-scale surface feature, multi-level curve feature, and blob feature. To tackle the challenge of complex volumetric inner structure and diverse feature forms, our technical solution hinges upon the integrated approach of locally-defined diffusion tensor (DT), DT-based anisotropic convolution kernel (DACK), DACK-based multi-scale analysis, and DT-governed curve feature growing. The extracted structural features can be further semantically classified. At the computational fronts, we design CUDA-based algorithm to conduct parallel computation for time consuming tasks. Various experiments and timing tests demonstrate the effectiveness, robustness, and high performance of our method.
Shuai Li 0001, Qinping Zhao, Shengfa Wang, Aimin Hao, Hong Qin 0001
ICIP4
2013 Unsupervised Co-segmentation of Complex Image Set via Bi-harmonic Distance Governed Multi-level Deformable Graph Clustering
abstract
Despite the recent success of extensive co-segmentation studies, they still suffer from limitations in accommodating multiple-foreground, large-scale, high-variability image set, as well as their underlying capability for parallel implementation. To improve, this paper proposes a bi-harmonic distance governed flexible method for the robust coherent segmentation of the overlapping/similar contents co-existing in image group, which is independent of supervised learning and any other user-specified prior. The central idea is the novel integration of bi-harmonic distance metric design and multi-level deformable graph generation for multi-level clustering, which gives rise to a host of unique advantages: accommodating multiple-foreground images, respecting both local structures and global semantics of images, being more robust and accurate, and being convenient for parallel acceleration. Critical pipeline of our method involves intrinsic content-coherent measuring, super-pixel assisted bottom-up clustering, and multi-level deformable graph clustering based cross-image optimization. We conduct extensive experiments on the iCoseg benchmark and Oxford flower datasets, and make comprehensive evaluations to demonstrate the superiority of our method via comparison with state-of-the-art methods collected in the MSRC database.
Jizhou Ma, Shuai Li 0001, Aimin Hao, Hong Qin 0001
ISM3
2013 Robust and high-fidelity guidewire simulation with applications in percutaneous coronary intervention system
abstract
Real-time and realistic physics-based simulation of deformable objects is of great value to medical intervention, training, and planning in virtual environments. This paper advocates a virtual-reality (VR) approach to minimally-invasive surgery/therapy (e.g., percutaneous coronary intervention) in medical procedures. In particular, we devise a robust and accurate physics-based modeling and simulation algorithm for the guidewire interaction with blood vessels. We also showcase a VR-based prototype system for simulating percutaneous coronary intervention and mimicing the intervention therapy, which affords the utility of flexible, slender guidewires to advance diagnostic or therapeutic catheters into a patient's vascular anatomy, supporting various real-world interaction tasks. The slender body of guidewires are modeled using the famous Cosserat theory of elastic rods. We derive the equations of motion for guidewires with continuous energies and integrate them with the implicit Euler solver, that guarantees robustness and stability. Our approach's originality is primarily founded upon its power, flexibility, and versatility when interacting with the surrounding environment, including novel strategies in the hybrid of geometry and physics, material variability, dynamic sampling, constraint handling and energy-driven physical responses. Our experimental results have shown that this prototype system is both stable and efficient with real-time performance. In the long run, our algorithm and system are expected to contribute to interactive VR-based procedure training and treatment planning.
Yurun Mao, Fei Hou 0001, Shuai Li 0001, Aimin Hao, Mingjing Ai, Hong Qin 0001
VRST4
2013 Multi-scale local features based on anisotropic heat diffusion and global eigen-structure
Shuai Li 0001, Hong Qin 0001, Aimin Hao
Sci. China Inf. Sci.3
2012 A Novel Material-Aware Feature Descriptor for Volumetric Image Registration in Diffusion Tensor Space
Shuai Li 0001, Qinping Zhao, Shengfa Wang, Tingbo Hou, Aimin Hao, Hong Qin 0001
ECCV (4)5
2012 Realtime Two-Way Coupling of Meshless Fluids and Nonlinear FEM
abstract
Abstract In this paper, we present a novel method to couple Smoothed Particle Hydrodynamics (SPH) and nonlinear FEM to animate the interaction of fluids and deformable solids in real time. To accurately model the coupling, we generate proxy particles over the boundary of deformable solids to facilitate the interaction with fluid particles, and develop an efficient method to distribute the coupling forces of proxy particles to FEM nodal points. Specifically, we employ the Total Lagrangian Explicit Dynamics (TLED) finite element algorithm for nonlinear FEM because of many of its attractive properties such as supporting massive parallelism, avoiding dynamic update of stiffness matrix computation, and efficient solver. Based on a predictor‐corrector scheme for both velocity and position, different normal and tangential conditions can be realized even for shell‐like thin solids. Our coupling method is entirely implemented on modern GPUs using CUDA. We demonstrate the advantage of our two‐way coupling method in computer animation via various virtual scenarios.
Lipeng Yang, Shuai Li 0001, Aimin Hao, Hong Qin 0001
Comput. Graph. Forum3
2010 Geodesic Model of Human Body
abstract
Anthropometry is widely applied to the research in skeleton extraction from surface meshes of human body. Especially the anatomical proportion can be employed as a benchmark in model segmentation and joint extraction. Unfortunately, the anatomical proportion is usually measured with the Euclidean distance, which makes it difficult to correlate it with the surface mesh. To bridge this gap, we take advantage of the property of the geodesic metrics that is invariance to rotation, translation, scaling and model pose, and propose an original geodesic model in which the length of each part of human body is measured by geodesic metrics, by which the anatomic proportions can be directly mapped to the contours of the mesh surface of human body in arbitrary pose. Combining the geodesic model with automatic extraction of feature points, we can determine the candidate scopes of joint positions and boundaries between the parts on meshes, and then refine the joint positions in the scopes using existing methods. And finally, we illustrate the utility of the geodesic model with an application to joint extraction.
Weihe Wu, Aimin Hao, Yongtao Zhao
CW2
2010 Automatic construction of 3D animatable facial avatars
abstract
Abstract Rigging for facial animation is an important but time‐consuming task, which generally requires experienced artists with knowledge of facial anatomy. In this paper, we investigate whether it is possible to produce a good animatable avatar automatically, given only a 3D static triangle mesh of the head. An automatic mechanism is devised for constructing multi‐layer animatable facial avatars for unseen faces. We evaluate our technique with a variety of models, and give a quantitative analysis of the constructed results. We also designed and conducted a user study for evaluating the perceived quality of the generated expressive animations. The results demonstrate that our method is an appropriate tool for naïve users to customize their personal 3D avatars. Copyright © 2010 John Wiley & Sons, Ltd.
Yujian Gao, Qinping Zhao, Aimin Hao, Tevfik Metin Sezgin, Neil A. Dodgson
Comput. Animat. Virtual Worlds3
2007 Best Fit Decreasing for Fully Automatic Compact Texture Atlas Generation
abstract
Texture atlas is widely used in many applications such as texture mapping, 3D Paint and illumination map texture synthesis etc. Several methods have been proposed for generating texture atlas, but most of them are not compact enough for some applications or require too much manual intervention. This paper presents a fully automatic BFD approach to generate compact texture atlas for triangular mesh models. Our main contribution is making some modifications to a piecewise mesh parameterization method to control the shape of charts by introducing a new local criterion, and developing a novel rectangle-packing algorithm.
Aimin Hao, Qinping Zhao
CAD/Graphics1
2007 A Method for Terrain Rendering without Seams Based on Image
abstract
This is a new method to avoid crack in terrain rendering. It does not use the traditional technique to deal with crack in geometry space. It draws terrain to a texture using FBO (frame buffer object) instead of to screen directly without operation to dispose seam. Then, it uses "inpainting", a technique in image processing, to treat the texture in image space, so that the seamless image rendering can be realized. As for rendering terrain, it uses a quad tree to manage terrain data, and preserves the values of height in a texture. Then it can fetch them from video memory. In rendering a node, it uses geomorphing operation to avoid popping.
Shaopeng Tang, Lili Wang 0006, Aimin Hao
CAD/Graphics3