VLDB 2026 Research / reviewers in the wild / expert
Shuai Li 0001
dblp:57/2281-1
· DBLP profile ↗
139ranked-venue papers
22as first author
72since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 89 · 9 first-author · 38 since 2021Artificial intelligence and machine learning · 34 · 9 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 6 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DECON: Reconstruction of Clothed-Geometric Multiple Humans from a Single Image via Geometry-Guided Decouplingabstract3D multi-human reconstruction from single images holds significant potential for advancing AR/VR applications. While remarkable progress has been made in single-human reconstruction, existing methods face challenges when reconstructing multiple humans. These challenges include: (1) severe inter-occlusion that disrupts individual body structures, and (2) the absence of physically plausible relative positioning among subjects. We present DECON, a novel DEcouple-and-reCONstruct framework that systematically addresses these limitations through two technical innovations: (1) a decouple-and-reconstruct framework with multi-view synthesis. It separates individuals and reconstructs detailed 3D bodies from a single image. (2) a Perspective-Aware Position Optimization (PAPO) approach. It ensures realistic positioning by fixing overlaps and gaps between subjects. Extensive experiments demonstrate our method's capability to reconstruct fully separated, anatomically complete 3D humans with clothed-geometric details and plausible interactions. Quantitative evaluations show a 54% reduction in Chamfer Distance and 35% in Point-to-Surface Distance compared to state-of-the-art methods. Yiming Jiang 0018, Wenfeng Song, Shuai Li 0001, Aimin Hao |
AAAI | 3 |
| 2026 | IntentMotion: Learning Intent-Aware Human Motion from Language in 3D ScenesabstractGenerating human motion in complex 3D scenes from text is a challenging task with broad applications. However, existing methods often overlook realistic physical contact, resulting in visually plausible but physically unrealistic motion, e.g., penetration. To alleviate this, we propose IntentMotion, a novel framework that generates human motion in 3D scenes from natural language instructions by explicitly modeling intent. We first introduce the Intention-Guided Contact Field (IGCF). This differentiable voxel-based contact region representation explicitly aligns parsed language roles with spatial contact regions through a hierarchical attention mechanism. IGCF is jointly trained with a diffusion-based motion generator, allowing contact predictions to adapt dynamically through gradient feedback. To improve the controllability and physics-aware motion, we further propose an Intention-Aware Diffusion Model (IADM), which decouples the high-level semantic planning from the low-level contact refinement in a coarse-to-fine process. The optimized contact cues are utilized to guide the synthesis of a coarse trajectory, followed by refining detailed pose sequences under IGCF supervision. Experiments on the HUMANISE and LINGO datasets demonstrate that our IntentMotion outperforms recent baselines in contact accuracy, semantic alignment, and generalization to unseen scenes. Wenfeng Song, Shi Zheng, Xingliang Jin, Aimin Hao, Fei Hou 0001, Xia Hou, Shuai Li 0001 |
AAAI | 8 |
| 2026 | Energy-based haptic rendering for real-time surgical simulation
Mingbo Hu, Wenli Xiu, Siming Zheng, Shuai Li 0001, Aimin Hao |
Comput. Graph. | 6 |
| 2026 | Deep-Saliency Foveated Ray Tracing For Real-time VR RenderingabstractImmersive VR applications demand high resolutions and refresh rates, posing significant challenges for real-time rendering. Foveated rendering mitigates this cost by exploiting properties of the Human Visual System (HVS), but conventional approaches often rely on oversimplified heuristic models that neglect high-level attentional cues, resulting in artifacts in peripheral regions. To this end, we present a neural saliency-driven foveated ray tracing framework that overcomes these limitations. Our method introduces a motion-aware foveation model to capture temporal dynamics and employs a lightweight convolutional neural network to predict saliency maps that reflect complex attentional patterns derived from eye-gaze data. The combination of these guides adaptive path tracing and filtering, enabling perceptually optimized rendering with minimal artifacts. Experimental results show that our approach improves perceptual quality over prior methods while sustaining real-time performance. Yang Gao 0032, Wencan Li, Shiyu Liang, Weizichuan Feng, Qing Xia 0002, Shuai Li 0001, Aimin Hao |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2026 | HFHuman: High-Fidelity Human Reconstruction From Single Image With Multi-Modality FusionabstractAccurately reconstructing high-fidelity human models from single images is critical for virtual reality applications. Existing methods often rely on 3D features from the estimated parametric human model to provide geometric priors. This approach addresses challenges such as missing limbs or deformations, which often arise due to viewpoint limitations and self-occlusion. However, accurately predicting 3D features from monocular images remains a significant challenge. This limitation poses difficulties for the fusion of 2D and 3D information. In this paper, we introduce HFHuman, a novel approach for high-fidelity human reconstruction from a single image using multi-modality fusion. HFHuman effectively fuses multiple modalities, including geometric and depth, directly from images. Our method introduces three key innovations: (1) a depth and geometric parallel reconstruction framework that simultaneously handles whole-body geometry and detailed depth reconstruction, refining a parameterized 3D human model under progressive depth guidance; (2) a pixel-voxel feature fusion strategy that combines pixel-aligned features with voxel-aligned features using a multi-modality adaptor; and (3) a depth-refined technique that integrates RGB imagery with surface normals and depth mapping. By addressing the challenge of blending 2D and 3D modalities, HFHuman results in more accurate and realistic human reconstructions. Experimental results demonstrate that HFHuman outperforms state-of-the-art methods, setting a new standard for realistic 3D human body reconstruction. Yiming Jiang 0018, Wenfeng Song, Shuai Li 0001, Aimin Hao |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2026 | FCMD: Fine-Grained Text-Driven Cohesive Motion Generation With Diffusion ModelabstractGenerating continuous and expressive human motion from textual descriptions is a critical challenge in applications such as gaming and filmmaking. Existing methods often struggle to maintain global coherence, realistic frame continuity, and smooth transitions. To address these limitations, we propose FCMD, a novel diffusion-based model for generating cohesive motion sequences from fine-grained textual descriptions. FCMD introduces three key innovations: (1) Fine-grained Text Fusion, which integrates detailed textual cues with transitional narratives to enhance semantic consistency; (2) History Motion Guidance, ensuring motion accuracy and consistency across consecutive frames; and (3) Smooth Stitching Sampling, which leverages preceding and current motion information to achieve seamless transitions. Additionally, FCMD employs a large language model (LLM) to refine motion datasets by extracting fine-grained textual descriptions. Extensive experiments demonstrate that FCMD outperforms state-of-the-art methods in generating coherent, natural, and highly controllable motion sequences. Shuai Li 0001, Wenfeng Song, Aimin Hao |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2026 | DynAvatar: Dynamic 3D Head Avatar Deformation With Expression Guided Gaussian SplattingabstractGenerating high-fidelity, expressive, and realistic 3D head avatars remains a fundamental challenge for immersive applications such as virtual reality, gaming, and telepresence. This task requires not only precise modeling of non-rigid facial deformations but also semantically controllable expression synthesis under diverse viewpoints and motion contexts. We present DynAvatar, a novel framework that integrates expression-guided deformation into the 3D Gaussian splatting pipeline to produce photorealistic and emotionally resonant head avatars. Our method introduces two key innovations: (1) an expression-guided Gaussian deformation module that tightly couples geometric displacement with high-level semantic cues, enabling fine-grained and anatomically meaningful facial animation; and (2) a spatial context embedding mechanism that encodes the canonical position of each Gaussian to preserve semantic coherence and spatial consistency during expression generation. Extensive experiments on both controlled and in-the-wild datasets demonstrate that DynAvatar significantly outperforms state-of-the-art methods in terms of visual realism, expression fidelity, and rendering quality. Wenfeng Song, Zhongyong Ye, Shuai Li 0001, Xia Hou, Aimin Hao |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | CtrlAvatar: Controllable Avatars Generation via Disentangled Invertible NetworksabstractAs virtual experiences grow in popularity, the demand for realistic, personalized, and animatable human avatars increases. Traditional methods, relying on fixed templates, often produce costly avatars that lack expressiveness and realism. To overcome these challenges, we introduce Controllable Avatars generation via disentangled invertible networks (CtrlAvatar), a real-time framework for generating lifelike and customizable avatars. CtrlAvatar uses disentangled invertible networks to separate the deformation process into implicit body geometry and explicit texture components. This approach eliminates the need for repeated occupancy reconstruction, enabling detailed and coherent animations. The body geometry component ensures anatomical accuracy, while the texture component allows for complex, artifact-free clothing customization. This architecture ensures smooth integration between body movements and surface details. By optimizing transformations with position-varying offsets from the avatar’s initial Linear Blend Skinning vertices, CtrlAvatar achieves flexible, natural deformations that adapt to various scenarios. Extensive experiments show that CtrlAvatar outperforms other methods in quality, diversity, controllability, and cost-efficiency, marking a significant advancement in avatar generation. Wenfeng Song, Fei Hou 0001, Shuai Li 0001, Aimin Hao, Xia Hou |
AAAI | 4 |
| 2025 | Genomics-Aware Multimodal Self-Supervised Learning for Cancer Survival PredictionabstractSurvival prediction in cancer diagnosis is a critical research task. Current methods often employ the multimodal feature fusion of pathological images and genomics data within a weakly-supervised learning paradigm. However, these approaches fail to efficiently learn the intrinsic features of large amount of unlabeled WSIs and neglect the strong associations between genomics data and pathological images, resulting in reduced prognostic accuracy. To address these challenges, we propose a novel Genomics-Aware Multimodal Self-Supervised Learning model that designs a multimodal pretext task, improving learning of intra-modal features and inter-modal correlations without additional annotations. Specifically, we randomly mask pathological patch features and fuse unmasked pathology representations with genomics representations via a cross-modal attention module. Then we add mask tokens to the genomicsguided pathology representation and reconstruct the missing parts via a reconstruction decoder. Experimental results on four TCGA datasets demonstrate the superior performance of our method compared to state-of-the-art methods, highlighting its potential for advancing survival prediction. Our code is available at https://github.com/sunkevin101/GMSL. Yuanbo He, Zining Liu, Jiahao Cui 0001, Shuai Li 0001 |
BIBM | 6 |
| 2025 | Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained ControlabstractSpeech-driven 3D talking face method should offer both accurate lip synchronization and controllable expressions. Previous methods solely adopt discrete emotion labels to globally control expressions throughout sequences while limiting flexible fine-grained facial control within the spatiotemporal domain. We propose a diffusion-transformer-based 3D talking face generation model, Cafe-Talk, which simultaneously incorporates coarse- and fine-grained multimodal control conditions. Nevertheless, the entanglement of multiple conditions challenges achieving satisfying performance. To disentangle speech audio and fine-grained conditions, we employ a two-stage training pipeline. Specifically, Cafe-Talk is initially trained using only speech audio and coarse-grained conditions. Then, a proposed fine-grained control adapter gradually adds fine-grained instructions represented by action units (AUs), preventing unfavorable speech-lip synchronization. To disentangle coarse- and fine-grained conditions, we design a swap-label training mechanism, which enables the dominance of the fine-grained conditions. We also devise a mask-based CFG technique to regulate the occurrence and intensity of fine-grained control. In addition, a text-based detector is introduced with text-AU alignment to enable natural language user input and further support multimodal control. Extensive experimental results prove that Cafe-Talk achieves state-of-the-art lip synchronization and expressiveness performance and receives wide acceptance in fine-grained control in user studies. Hejia Chen, Haoxian Zhang, Shoulong Zhang, Sisi Zhuang, Yuan Zhang 0020, Pengfei Wan 0001, Di Zhang 0026, Shuai Li 0001 |
ICLR | 9 |
| 2025 | ViMoGen: A Novel Motion Generator for Virtual Standard Patient
Xuehan Wang, Wenfeng Song, Shuai Li 0001, Xian'e Wang, Xia Hou |
ICXR | 4 |
| 2025 | The Effect of Unexpected Visual Stimuli on Short-Term Memory in Immersive Experience
Shoulong Zhang, Yutian Xiao, Xuejing Lu, Shuai Li 0001 |
ICXR | 6 |
| 2025 | Phys4DRT: Physics-based 4D Generation for Real-Time Interaction with Time-Frequency Supervision
Yuntian Xiao, Shoulong Zhang, Jiahao Cui 0001, Shuai Li 0001 |
ACM Multimedia | 6 |
| 2025 | Reactffusion: Physical Contact-guided Diffusion Model for Reaction Generation
Shoulong Zhang, Shuai Li 0001 |
ACM Multimedia | 4 |
| 2025 | Effects of interaction modalities and emotional states on user's perceived empathy with an LLM-based embodied conversational agent
Yang Gao 0032, Yangbin Dai, Guangtao Zhang, Aimin Hao, Shuai Li 0001 |
Int. J. Hum. Comput. Stud. | 6 |
| 2025 | Advancing MRI segmentation with CLIP-driven semi-supervised learning and semantic alignment
Kexuan Li, Jingjuan Liu, Xuehao Wang, Yuanbo He, Huadan Xue, Aimin Hao, Shuai Li 0001 |
Neurocomputing | 10 |
| 2025 | Saliency-Free and Aesthetic-Aware Panoramic Video NavigationabstractMost of the existing panoramic video navigation approaches are saliency-driven, whereby off-the-shelf saliency detection tools are directly employed to aid the navigation approaches in localizing video content that should be incorporated into the navigation path. In view of the dilemma faced by our research community, we rethink if the "saliency clues" are really appropriate to serve the panoramic video navigation task. According to our in-depth investigation, we argue that using "saliency clues" cannot generate a satisfying navigation path, failing to well represent the given panoramic video, and the views in the navigation path are also low aesthetics. In this paper, we present a brand-new navigation paradigm. Although our model is still trained on eye-fixations, our methodology can additionally enable the trained model to perceive the "meaningful" degree of the given panoramic video content. Outwardly, the proposed new approach is saliency-free, but inwardly, it is developed from saliency but biasing more to be "meaningful-driven"; thus, it can generate a navigation path with more appropriate content coverage. Besides, this paper is the first attempt to devise an unsupervised learning scheme to ensure all localized meaningful views in the navigation path have high aesthetics. Thus, the navigation path generated by our approach can also bring users an enjoyable watching experience. As a new topic in its infancy, we have devised a series of quantitative evaluation schemes, including objective verifications and subjective user studies. All these innovative attempts would have great potential to inspire and promote this research field in the near future. Chenglizhao Chen, Guangxiao Ma, Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | AttriDiffuser: Adversarially enhanced diffusion model for text-to-facial attribute image synthesis
Wenfeng Song, Zhongyong Ye, Xia Hou, Shuai Li 0001, Aimin Hao |
Pattern Recognit. | 5 |
| 2025 | TalkingStyle: Personalized Speech-Driven 3D Facial Animation With Style PreservationabstractIt is a challenging task to create realistic 3D avatars that accurately replicate individuals' speech and unique talking styles for speech-driven facial animation. Existing techniques have made remarkable progress but still struggle to achieve lifelike mimicry. This article proposes "TalkingStyle", a novel method to generate personalized talking avatars while retaining the talking style of the person. Our approach uses a set of audio and animation samples from an individual to create new facial animations that closely resemble their specific talking style, synchronized with speech. We disentangle the style codes from the motion patterns, allowing our method to associate a distinct identifier with each person. To manage each aspect effectively, we employ three separate encoders for style, speech, and motion, ensuring the preservation of the original style while maintaining consistent motion in our stylized talking avatars. Additionally, we propose a new style-conditioned transformer decoder, offering greater flexibility and control over the facial avatar styles. We comprehensively evaluate TalkingStyle through qualitative and quantitative assessments, as well as user studies demonstrating its superior realism and lip synchronization accuracy compared to current state-of-the-art methods. Wenfeng Song, Xuan Wang 0024, Shi Zheng, Shuai Li 0001, Aimin Hao, Xia Hou |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Frequency-Guided Network for Low-contrast Staining-free Dental Plaque SegmentationabstractTraditional dental plaque detection relies on medical staining reagents and professional intervention. Deep learning-based automatic staining-free dental plaque segmentation provides an alternative for patients to perform plaque detection at home without staining reagents. However, existing methods still struggle with low-contrast visual features between unstained plaque and healthy teeth. To address this, we propose a Frequency-Guided Network (FGN) for low-contrast staining-free dental plaque segmentation. We observe that dental plaque tends to concentrate specifically near the junction between the teeth and the gingiva. This junction demonstrates abrupt changes in pixel values, indicating high-frequency regions in the image. In other words, dental plaque tends to appear near the high-frequency regions of oral endoscope images. Exploiting this characteristic, we employ a frequency-guided decoupling module to separate the image into high-frequency and low-frequency regions automatically and expand the high-frequency region to encompass nearby potential dental plaque. Then we supervise two regions individually to specifically focus on the expended high-frequency region for localizing nearby dental plaque. Additionally, we propose a high-to-low frequency multiple tasks framework. In the first phase, the network segments the teeth region, and then we input the teeth mask into the second phase. In the second stage, the teeth mask allows us to have a higher frequency at the junction between the teeth and gums, thereby enhancing the effectiveness of frequency-guided decoupling. Furthermore, FGN integrates a frequency-driven refinement module to enhance the guidance quality of the teeth mask for the second phase. Extensive evaluations of the oral endoscope dataset demonstrate that our method outperforms existing high-performance segmentation methods. User studies also confirm that our approach achieves superior results to experienced dentists. https://frequency-guided-network.github.io/ Yiming Jiang 0018, Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
BIBM | 3 |
| 2024 | SL-SFGR: Segmentation Learning Coupling Spatial-Frequency Structure Information Enhancement for Guiding RegistrationabstractThe core of medical image registration lies in the alignment of corresponding structures. Hence, the effective construction of structure information in images is crucial for guiding registration. Some methods directly introduce explicit structure priors for assisting registration training, but obtaining high-precision structure priors is intrinsically challenging. As a segmentation model for obtaining structure priors, its capability is built upon the effective construction of structure information, thus its model learning can also be used to guide registration. However, existing registration-segmentation joint methods mostly only use segmentation results at the output level to constrain registration, neglecting the guidance of segmentation learning at the feature level for registration. Moreover, most existing methods only extract features in the image spatial domain for registration, overlooking the structure information in the frequency domain that is more easily captured to guide registration. To this end, this paper proposes an innovative registration method, namely Segmentation Learning Coupling Spatial-Frequency Structure Information Enhancement for Guiding Registration (SL-SFGR). Specifically, first, a semi-supervised segmentation learning network is constructed based on the registration deformation field to introduce structure features suitable for guiding registration. Second, an adaptive feature enhancement module is built in the spatial-frequency dual domain to further strengthen the inherent structure features. Finally, Dynamic Weight Average (DWA) is utilized for joint optimization of the model. The effectiveness of the proposed method has been verified on different brain MRI datasets. The related code is available at: https://github.com/goghfan/SL-SFGR/. Yuanbo He, Shuai Li 0001, Aimin Hao, Desen Cao |
BIBM | 3 |
| 2024 | mvEchoSeg: One-shot In-context Learning for Multi-view Echocardiography SegmentationabstractEchocardiography is the clinical standard for evaluation of cardiac morphology, and function, and providing hemodynamic parameters in patients with known or suspected heart disease. Due to the diversity and complexity of diagnostic tasks, comprehensive interpretation of echocardiography often requires multi-view imaging for integrated metric analysis. Deep learning methods have become the mainstream approach for echocardiography segmentation. However, achieving segmentation of multi-view echocardiography still requires a substantial amount of annotation. To remedy this, we propose a one-shot in-context learning network mvEchoSeg for multi-view echocardiography segmentation. This network requires only one annotated image for each view. Specifically, we propose a Task Prompt Identifier (TPI) module to identify the task type of the image and allocate the most precise task prompt for it, with minimal adaptation of CLIP and few-shot strategies. Additionally, we leverage a unified In-Context Model (ICM) capable of performing a diverse set of echocardiography segmentation tasks automatically. Using fine-tuning and low-rank adapters improved the performance of the pre-trained model, achieving significant results with minimal training cost. Furthermore, we collect a multi-view echocardiography dataset (MVECD) with 8 views to evaluate our method. The results show an improvement of more than 10% in the DICE score compared to the SOTA foundational medical image segmentation models. To our knowledge, this is the first exploration of a one-shot model for multi-view echocardiography segmentation. Our codes and models are available at https://github.com/stellating/mvEchoSeg. Ya Duan, Wenfeng Song, Nannan Li 0002, Aili Li, Shuai Li 0001 |
BIBM | 6 |
| 2024 | Radiographic Reports Generation via Retrieval Enhanced Cross-modal FusionabstractAccurate radiographic reports are crucial for effective clinical decision-making and patient safety, as they directly influence diagnosis and treatment plans. Existing models for radiographic report generation often struggle with integrating medical image and textual report features and addressing data imbalances. To overcome these limitations, we propose an enhanced cross-modal aligned retrieval-driven network (EARnet). Our model incorporates two key innovations: Enhanced Cross-modal Alignment (ECA) module and Case-based Retrieval Augmenter (CRA) module. ECA ensures the effective integration of medical visual and textual data by aligning these modalities. CRA helps mitigate data imbalance by improving the representation of medical information and ensuring a more balanced and comprehensive coverage of both normal and abnormal cases. This dual-module approach significantly improves the coherence and accuracy of the generated medical reports. Evaluation results demonstrate that our EARnet model substantially outperforms existing methods in terms of report quality and accuracy across multiple metrics. Our codes and models are available at https://github.com/lyf616/EARnet. Xia Hou, Wenfeng Song, Wenzhe You, Shuai Li 0001 |
BIBM | 6 |
| 2024 | Semi-Supervised Medical Image Segmentation with Cross-View Consistency and Contrastive LearningabstractMedical image segmentation plays a crucial role in many clinical applications. To alleviate the dependency on massive annotations, semi-supervised learning has attracted increasing attention. However, these methods face significant intra-class and inter-class variation and do not fully utilize the critical multi-view information inherent in medical images. This study proposes a novel network, CV-Net, which integrates multi-view information for semi-supervised medical image segmentation. Concretely, the network is based on Mean-Teacher architecture which largely narrows the empirical distribution gap between labeled and unlabeled data. The proposed cross-view consistency regularization module incorporates a dual-branch attention architecture to integrate consistent semantics while focusing on details, enhancing feature extraction capabilities. The proposed bi-semantic contrastive learning module leverages limited labels and explore pseudo-labels to define semantically similar regions, enhancing the representation capacity. Experiments conducted on two datasets demonstrated the effectiveness of the proposed network. CV-Net showed significant improvements across four metrics, evident with both 5% and 10% labeled data. Specifically, with 5% labeled data, the mean Dice increased by 1.37%. Compared with previous state-of-the-art methods, CV-Net achieved the best results, notably reducing both intra-class and inter-class errors. Kexuan Li, Jingjuan Liu, Xuehao Wang, Huadan Xue, Aimin Hao, Shuai Li 0001 |
BIBM | 8 |
| 2024 | Arbitrary Motion Style Transfer with Multi-Condition Motion Latent Diffusion ModelabstractComputer animation's quest to bridge content and style has historically been a challenging venture, with previous efforts often leaning toward one at the expense of the other. This paper tackles the inherent challenge of content-style duality, ensuring a harmonious fusion where the core narrative of the content is both preserved and elevated through stylistic enhancements. We propose a novel Multi-condition Motion Latent Diffusion Model (MCM-LDM) for Arbitrary Motion Style Transfer (AMST). Our MCM-LDM significantly emphasizes preserving trajectories, recognizing their fundamental role in defining the essence and fluidity of motion content. Our MCM-LDM's cornerstone lies in its ability first to disentangle and then intricately weave together motion's tripartite components: motion trajectory, motion content, and motion style. The critical insight of MCM-LDM is to embed multiple conditions with distinct priorities. The content channel serves as the primary flow, guiding the overall structure and movement, while the trajectory and style channels act as auxiliary components and synchronize with the primary one dynamically. This mechanism ensures that multi-conditions can seamlessly integrate into the main flow, enhancing the overall animation without overshadowing the core content. Empirical evaluations underscore the model's proficiency in achieving fluid and authentic motion style transfers, setting a new benchmark in the realm of computer animation. The source code and model are available at https://github.com/XingliangJin/MCM-LDM.git. Wenfeng Song, Xingliang Jin, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Xia Hou, Hong Qin 0001 |
CVPR | 3 |
| 2024 | HOIAnimator: Generating Text-Prompt Human-Object Animations Using Novel Perceptive Diffusion ModelsabstractTo date, the quest to rapidly and effectively produce human-object interaction (HOI) animations directly from textual descriptions stands at the forefront of computer vision research. The underlying challenge demands both a discriminating interpretation of language and a comprehen-sive physics-centric model supporting real-world dynamics. To ameliorate, this paper advocates HOIAnimator, a novel and interactive diffusion model with perception ability and also ingeniously crafted to revolutionize the animation of complex interactions from linguistic narratives. The effectiveness of our model is anchored in two ground-breaking innovations: (1) Our Perceptive Diffusion Models (PDM) brings together two types of models: one focused on hu-man movements and the other on objects. This combination allows for animations where humans and objects move in concert with each other, making the overall motion more realistic. Additionally, we propose a Perceptive Message Passing (PMP) mechanism to enhance the communication bridging the two models, ensuring that the animations are smooth and unified; (2) We devise an Interaction Contact Field (ICF), a sophisticated model that implicitly captures the essence of HOls. Beyond mere predictive contact points, the ICF assesses the proximity of human and object to their respective environment, informed by a probabilistic distribution of interactions learned throughout the denoising phase. Our comprehensive evaluation showcases HOlani-mator's superior ability to produce dynamic, context-aware animations that surpass existing benchmarks in text-driven animation synthesis. Wenfeng Song, Shuai Li 0001, Yang Gao 0032, Aimin Hao, Xia Hau, Chenglizhao Chen, Hong Qin 0001 |
CVPR | 3 |
| 2024 | SIE-DepthNet: Semantic-Guided Monocular Depth Estimation for Dynamic Environment
Zilong Song, Yang Gao 0032, Sijia Dai, Shuai Li 0001, Aimin Hao, Shoulong Zhang |
ICXR | 4 |
| 2024 | A Coupling Physics Model for Real-Time 4D Simulation of Cardiac Electromechanics
Jiahao Cui 0001, Shuai Li 0001, Aimin Hao |
Comput. Aided Des. | 3 |
| 2024 | Conditional room layout generation based on graph neural networks
Zhihan Yao, Jiahao Cui 0001, Shoulong Zhang, Shuai Li 0001, Aimin Hao |
Comput. Graph. | 5 |
| 2024 | CoupNeRF: Property-aware Neural Radiance Fields for Multi-Material Coupled Scenario ReconstructionabstractAbstract Neural Radiance Fields (NeRFs) have achieved significant recognition for their proficiency in scene reconstruction and rendering by utilizing neural networks to depict intricate volumetric environments. Despite considerable research dedicated to reconstructing physical scenes, rare works succeed in challenging scenarios involving dynamic, multi‐material objects. To alleviate, we introduce CoupNeRF, an efficient neural network architecture that is aware of multiple material properties. This architecture combines physically grounded continuum mechanics with NeRF, facilitating the identification of motion systems across a wide range of physical coupling scenarios. We first reconstruct specific‐material of objects within 3D physical fields to learn material parameters. Then, we develop a method to model the neighbouring particles, enhancing the learning process specifically in regions where material transitions occur. The effectiveness of CoupNeRF is demonstrated through extensive experiments, showcasing its proficiency in accurately coupling and identifying the behavior of complex physical scenes that span multiple physics domains. Jin Li 0068, Yang Gao 0032, Wenfeng Song, Yacong Li, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
Comput. Graph. Forum | 5 |
| 2024 | Correction: Automatic Generation of 3D Scene Animation Based on Dynamic Knowledge Graphs and Contextual Encoding
Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
Int. J. Comput. Vis. | 4 |
| 2024 | Dynamic attention augmented graph network for video accident anticipation
Wenfeng Song, Shuai Li 0001, Tao Chang, Ke Xie 0005, Aimin Hao, Hong Qin 0001 |
Pattern Recognit. | 2 |
| 2024 | Joints-Centered Spatial-Temporal Features Fused Skeleton Convolution Network for Action RecognitionabstractSkeleton-based action recognition is crucial for natural human-computer interaction, dynamic behavior analysis, and behavior surveillance. The key challenge is to effectively capture the intrinsic local-global clues of the activity. However, it remains challenging to efficiently leverage multidimensional information related to joints' local visual appearances, global spatial relationships, and coherent temporal cues. To address this challenge, we propose a joints-centered spatial-temporal feature-fused framework for action recognition, which exploits skeleton-based graph diffusion and convolution. Specifically, we employ Partial Differential Equation (PDE) based skeleton graph diffusion to automatically activate and diffuse the salient appearance features of joints. This approach simultaneously integrates the joints' appearance clues and their hierarchical relationships at both the super-pixel level and structure level. The diffused appearance-related features of the joints are further fused with skeleton-related spatial-temporal features, and the resulting fused features are fed into a skeleton convolution network for action recognition. Our method was extensively evaluated on two public datasets (NTU-RGBD and UWA3D), and the results demonstrate the improved accuracy and effectiveness of our approach. Our code will be public. Wenfeng Song, Tangli Chu, Shuai Li 0001, Nannan Li 0002, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | CenterFormer: A Novel Cluster Center Enhanced Transformer for Unconstrained Dental Plaque SegmentationabstractDental plaque segmentation is crucial for maintaining oral health. However, accurately segmenting dental plaque in unconstrained environments can be challenging due to its low contrast and high variability in appearance. While existing transformer-based networks rely on attention mechanisms for each pixel, they do not take into account the relationships between neighboring pixels. Consequently, feature extraction is limited, making it difficult to achieve accurate segmentation of low-contrast images. To address this issue, we propose a simple yet efficient cluster center transformer that improves dental plaque segmentation by clustering image pixels based on multiple levels of feature maps' intensity and texture information. By grouping similar pixels into regions, the proposed method enables the transformers to focus on the local contour and edge around the teeth regions, adapting to the low contrast and high variability of plaque appearance, leading to more accurate and efficient segmentation of dental plaque in dental images. Additionally, we designed Multiple Granularity Perceptions using a pyramid fusion mechanism to capture multiple scales of vision features, thereby enhancing the low-contrast vision features. The proposed method can benefit the dental diagnosis and treatment planning process by improving the accuracy and efficiency of dental plaque segmentation. Our proposed method achieved state-of-the-art results on the dental plaque dataset (Li et al., 2020), with intersection over union (IoU) of 60.91% and pixel accuracy (PA) of 76.81%, all of which were the highest among all methods, demonstrating its effectiveness in plaque segmentation in unconstrained environments. Wenfeng Song, Xuan Wang 0024, Shuai Li 0001, Aimin Hao |
IEEE Trans. Multim. | 4 |
| 2024 | MPMNet: A Data-Driven MPM Framework for Dynamic Fluid-Solid InteractionabstractHigh-accuracy, high-efficiency physics-based fluid-solid interaction is essential for reality modeling and computer animation in online games or real-time Virtual Reality (VR) systems. However, the large-scale simulation of incompressible fluid and its interaction with the surrounding solid environment is either time-consuming or suffering from the reduced time/space resolution due to the complicated iterative nature pertinent to numerical computations of involved Partial Differential Equations (PDEs). In recent years, we have witnessed significant growth in exploring a different, alternative data-driven approach to addressing some of the existing technical challenges in conventional model-centric graphics and animation methods. This article showcases some of our exploratory efforts in this direction. One technical concern of our research is to address the central key challenge of how to best construct the numerical solver effectively and how to best integrate spatiotemporal/dimensional neural networks with the available MPM's pressure solvers. In particular, we devise the MPMNet, a hybrid data-driven framework supporting the popular and powerful MPM, to combine the comprehensive properties of MPM in numerically handling physical behaviors ranging from fluid to deformable solids and the high efficiency of data-driven models. At the architectural level, our MPMNet comprises three primary components: A data processing module to describe the physical properties by way of the input fields; A deep neural network group to learn the spatiotemporal features; And an iterative refinement process to continue to reduce possible numerical errors. The goal of these special technical developments is to aim at involved numerical acceleration while preserving physical accuracy, realizing efficient and accurate fluid-solid interactions in a data-driven fashion. The extensive experimental results verify that our MPMNet can tremendously speed up the computation compared with the popular numerical methods as the complexity of interaction scenes increases while better retaining the numerical accuracy. Jin Li 0068, Yang Gao 0032, Ju Dai, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | A Unified Particle-Based Solver for Non-Newtonian Behaviors SimulationabstractIn this article, we present a unified framework to simulate non-Newtonian behaviors. We combine viscous and elasto-plastic stress into a unified particle solver to achieve various non-Newtonian behaviors ranging from fluid-like to solid-like. Our constitutive model is based on a Generalized Maxwell model, which incorporates viscosity, elasticity and plasticity in one non-linear framework by a unified way. On the one hand, taking advantage of the viscous term, we construct a series of strain-rate dependent models for classical non-Newtonian behaviors such as shear-thickening, shear-thinning, Bingham plastic, etc. On the other hand, benefiting from the elasto-plastic model, we empower our framework with the ability to simulate solid-like non-Newtonian behaviors, i.e., visco-elasticity/plasticity. In addition, we enrich our method with a heat diffusion model to make our method flexible in simulating phase change. Through sufficient experiments, we demonstrate a wide range of non-Newtonian behaviors ranging from viscous fluid to deformable objects. We believe this non-Newtonian model will enhance the realism of physically-based animation, which has great potential for computer graphics. Yang Gao 0032, Tianwei Cheng, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Expressive 3D Facial Animation Generation Based on Local-to-Global Latent Diffusionabstract3D Facial animations, crucial to augmented and mixed reality digital media, have evolved from mere aesthetic elements to potent storytelling media. Despite considerable progress in facial animation of neutral emotions, existing methods still struggle to capture the authenticity of emotions. This paper introduces a novel approach to capture fine facial expressions and generate facial animations using audio synchronization. Our method consists of two key components: First, the Local-to-global Latent Diffusion Model (LG-LDM) tailored for authentic facial expressions, which can integrate audio, time step, facial expressions, and other conditions towards possible encoding of emotionally rich yet latent features in response to possibly noisy raw audio signals. The core of LG-LDM is our carefully designed Facial Denoiser Model (FDM) for aligning the local-to-global animation feature with audio. Second, we redesign an Emotion-centric Vector Quantized-Variational AutoEncoder framework (EVQ-VAE) to finely decode the subtle differences under different emotions and reconstruct the final 3D facial geometry. Our work significantly contributes to the key challenges of emotionally realistic 3D facial animation for audio synchronization and enhances the immersive experience and emotional depth in augmented and mixed reality applications. We provide a reproducibility kit including our code, dataset, and detailed instructions for running the experiments. This kit is available at https://github.com/wangxuanx/Face-Diffusion-Model. Wenfeng Song, Xuan Wang 0024, Yiming Jiang 0018, Shuai Li 0001, Aimin Hao, Xia Hou, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Diffusing Coupling High-Frequency-Purifying Structure Feature Extraction for Brain Multimodal RegistrationabstractThe core of medical image registration is the alignment of corresponding structures. However, in multimodal image registration, substantial differences in appearance (intensity distribution) of the images often compel the registration model to prioritize intensity information over structure information, resulting in low accuracy of registration. Therefore, the disentangling structure information from intensity information is vital to improve the registration effectiveness. To this end, we propose a diffusing coupling high-frequency-purifying structure feature extraction for brain multimodal registration. Specifically, the denoising diffusion probabilistic models (DDPM) is firstly utilized to extract complete feature information from images. Then, the discrete cosine transform (DCT) is applied to purify high-frequency structure information from the complete feature information for registration. Furthermore, structure consistency constraint (SCC) is introduced based on purified structure information to emphasize the core position of the structure in registration. Through comprehensive comparisons with traditional and learning-based methods on the multimodal brain MRI dataset, our method demonstrates superior accuracy and stability in brain multimodal registration. Our code is available at https://github.com/goghfan/DDNet. Yuanbo He, Shuai Li 0001, Aimin Hao, Desen Cao |
BIBM | 3 |
| 2023 | Propose-and-Complete: Auto-regressive Semantic Group Generation for Personalized Scene Synthesis
Shoulong Zhang, Shuai Li 0001, Xinwei Huang, Wenchong Xu, Aimin Hao, Hong Qin 0001 |
BMVC | 2 |
| 2023 | Sequential Texts Driven Cohesive Motions Synthesis with Natural TransitionsabstractThe intelligent synthesis/generation of daily-life motion sequences is fundamental and urgently needed for many VR/metaverse-related applications. However, existing approaches commonly focus on monotonic motion generation (e.g., walking, jumping, etc.) based on single instruction-like text, which is still not intelligent enough and can’t meet practical demands. To this end, we propose a cohesive human motion sequence synthesis framework based on free-form sequential texts while ensuring semantic connection and natural transitions between adjacent motions. At the technical level, we explore the local-to-global semantic features of previous and current texts to extract relevant information. This information is used to guide the framework in understanding the semantics of the current moment. Moreover, we propose learnable tokens to adaptively learn the influence range of the previous motions towards natural transitions. These tokens can be trained to encode the relevant information into well-designed transition loss. To demonstrate the efficacy of our method, we conduct extensive experiments and comprehensive evaluations on the public dataset as well as a new dataset produced by us. All the experiments confirm that our method outperforms the state-of-the-art methods in terms of semantic matching, realism, and transition fluency. Our project is public available. https://druthrie.github.io/sequential-texts-to-motion/ Shuai Li 0001, Sisi Zhuang, Wenfeng Song, Hejia Chen, Aimin Hao |
ICCV | 1 |
| 2023 | Modality Profile - A New Critical Aspect to be Considered When Generating RGB-D Salient Object Detection Training SetabstractIt is widely acknowledged that selecting appropriate training data is crucial for obtaining good results in real-world testing, more so than utilizing complex network architectures. However, in the field of RGB-D SOD research, researchers have primarily focused on enhancing network architectures and have given less consideration to the choice of training and testing datasets, which may not translate well in practical applications. This paper aims to address an existing issue - how can we automatically generate a data-driven RGB-D SOD training dataset? We propose that in addition to scene similarity, the concept of "modality profile'' should be taken into account. The term "modality profile'' refers to the complementary status of modalities within a given dataset. A training dataset with a modality profile similar to the test dataset can significantly improve performance. To address this, we present a viable solution for automatically generating a training dataset with any desired modality profile in a weakly supervised manner. Our method also provides high-quality pseudo-GTs for all RGB-D images obtained from the web, making it suitable for training RGB-D SOD models. Extensive quantitative evaluations demonstrate the significance of the proposed "modality profile'' and confirm the superiority of the newly constructed training set guided by our "modality profile''. All codes, datasets, and results are available at this link. Xuehao Wang, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
ACM Multimedia | 2 |
| 2023 | ZetaDesign: an end-to-end deep learning method for protein sequence design and side-chain packingabstractComputational protein design has been demonstrated to be the most powerful tool in the last few years among protein designing and repacking tasks. In practice, these two tasks are strongly related but often treated separately. Besides, state-of-the-art deep-learning-based methods cannot provide interpretability from an energy perspective, affecting the accuracy of the design. Here we propose a new systematic approach, including both a posterior probability and a joint probability parts, to solve the two essential questions once for all. This approach takes the physicochemical property of amino acids into consideration and uses the joint probability model to ensure the convergence between structure and amino acid type. Our results demonstrated that this method could generate feasible, high-confidence sequences with low-energy side conformations. The designed sequences can fold into target structures with high confidence and maintain relatively stable biochemical properties. The side chain conformation has a significantly lower energy landscape without delegating to a rotamer library or performing the expensive conformational searches. Overall, we propose an end-to-end method that combines the advantages of both deep learning and energy-based methods. The design results of this model demonstrate high efficiency, and precision, as well as a low energy state and good interpretability. Junyu Yan, Shuai Li 0001, Aimin Hao, Qinping Zhao |
Briefings Bioinform. | 2 |
| 2023 | Analyzing part functionality via multi-modal latent space embedding and interweaving
Jiahao Cui 0001, Shuai Li 0001, Fei Hou 0001, Aimin Hao, Hong Qin 0001 |
Comput. Graph. | 2 |
| 2023 | Automatic Generation of 3D Scene Animation Based on Dynamic Knowledge Graphs and Contextual Encoding
Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
Int. J. Comput. Vis. | 4 |
| 2023 | Graph Diffusion Convolutional Network for Skeleton Based Semantic Recognition of Two-Person ActionsabstractGraph Convolutional Networks (GCNs) have successfully boosted skeleton-based human action recognition. However, existing GCN-based methods mostly cast the problem as separated person's action recognition while ignoring the interaction between the action initiator and the action responder, especially for the fundamental two-person interactive action recognition. It is still challenging to effectively take into account the intrinsic local-global clues of the two-person activity. Additionally, message passing in GCN depends on adjacency matrix, but skeleton-based human action recognition methods tend to calculate the adjacency matrix with the fixed natural skeleton connectivity. It means that messages can only travel along a fixed path at different layers of the network or in different actions, which greatly reduces the flexibility of the network. To this end, we propose a novel graph diffusion convolutional network for skeleton based semantic recognition of two-person actions by embedding the graph diffusion into GCNs. At technical fronts, we dynamically construct the adjacency matrix based on practical action information, so that we can guide the message propagation in a more meaningful way. Simultaneously, we introduce the frame importance calculation module to conduct dynamic convolution, so that we can avoid the negative effect caused by the traditional convolution, wherein the shared weights may fail to capture key frames or be affected by noisy frames. Besides, we comprehensively leverage the multidimensional features related to joints' local visual appearances, global spatial relationship and temporal coherency, and for different features, different metrics are designed to measure the similarity underlying the corresponding real physical law of the motions. Moreover, extensive experiments and comprehensive evaluations on four public large-scale datasets (NTU-RGB+D 60, NTU-RGB+D 120, Kinetics-Skeleton 400, and SBU-Interaction) demonstrate that our method outperforms the state-of-the-art methods. Shuai Li 0001, Xinxue He, Wenfeng Song, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | SC-GAN: Subspace Clustering based GAN for Automatic Expression Manipulation
Shuai Li 0001, Wenfeng Song, Aimin Hao, Hong Qin 0001 |
Pattern Recognit. | 1 |
| 2023 | An Intelligent Virtual Standard Patient for Medical Students Training Based on Oral Knowledge GraphabstractVirtual standard patient (VSP) is in high demand for medical students' diagnosis ability training in an efficient manner. Different from the traditional conversation system in medical dialogue generation, VSP needs a novel conversation paradigm to act as the patient instead of the doctor. However, existing conversation techniques still have limited ability in terms of generation of symptoms exhibited by patients with the personalized and knowledge-centered expressions. To alleviate these problems, we propose to construct a novel oral knowledge graph, which sufficiently provides medical clues of the certain disease. Accordingly, the VSP could accurately interact with the dentists for their underlying intention and express the symptoms characters in a natural style. To efficiently retrieve the related disease clues, the symptoms descriptions of the oral diseases are encoded into the oral knowledge graph, which could well organize the disease-centered symptom entities and speaking styles. Moreover, to transfer the common sense knowledge from existing large scale of medical knowledge graph to the specific oral knowledge graph, a coupled pre-trained Bert models is further designed to learn the related medical knowledge from coarse-level to fine-level hierarchically. Finally, a series of well-designed personalized templates are proposed to generate plausible and realistic answers in condition of the certain disease. We also conduct extensive user studies to demonstrate that the VSP satisfies the medical students' diagnosis practice requirement in terms of naturalness, realism, and topic relevance. Wenfeng Song, Xia Hou, Shuai Li 0001, Chenglizhao Chen, Danyang Gao, Xian'e Wang, Yuzhe Sun, Jianxia Hou, Aimin Hao |
IEEE Trans. Multim. | 3 |
| 2023 | FineStyle: Semantic-Aware Fine-Grained Motion Style Transfer with Dual Interactive-Flow FusionabstractWe present FineStyle, a novel framework for motion style transfer that generates expressive human animations with specific styles for virtual reality and vision fields. It incorporates semantic awareness, which improves motion representation and allows for precise and stylish animation generation. Existing methods for motion style transfer have all failed to consider the semantic meaning behind the motion, resulting in limited controls over the generated human animations. To improve, FineStyle introduces a new cross-modality fusion module called Dual Interactive-Flow Fusion (DIFF). As the first attempt, DIFF integrates motion style features and semantic flows, producing semantic-aware style codes for fine-grained motion style transfer. FineStyle uses an innovative two-stage semantic guidance approach that leverages semantic clues to enhance the discriminative power of both semantic and style features. At an early stage, a semantic-guided encoder introduces distinct semantic clues into the style flow. Then, at a fine stage, both flows are further fused interactively, selecting the matched and critical clues from both flows. Extensive experiments demonstrate that FineStyle outperforms state-of-the-art methods in visual quality and controllability. By considering the semantic meaning behind motion style patterns, FineStyle allows for more precise control over motion styles. Source code and model are available on https://github.com/XingliangJin/Fine-Style.git. Wenfeng Song, Xingliang Jin, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Xia Hou |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | Distribution-motivated 3D Style Characterization Based on Latent Feature Decomposition
Xinwei Huang, Shuai Li 0001, Shoulong Zhang, Aimin Hao, Hong Qin 0001 |
Comput. Aided Des. | 2 |
| 2022 | Erratum to: Self-adjustable hyper-graphs for video pose estimation based on spatial-temporal subspace construction
Jizhou Ma, Shuai Li 0001, Hong Qin 0001, Aimin Hao, Qinping Zhao |
Sci. China Inf. Sci. | 2 |
| 2022 | Self-adjustable hyper-graphs for video pose estimation based on spatial-temporal subspace construction
Jizhou Ma, Shuai Li 0001, Hong Qin 0001, Aimin Hao, Qinping Zhao |
Sci. China Inf. Sci. | 2 |
| 2022 | Automatic image matting and fusing for portrait synthesis
Zhike Yi, Wenfeng Song, Shuai Li 0001, Aimin Hao |
Sci. China Inf. Sci. | 3 |
| 2022 | Multi-scale and multi-level shape descriptor learning via a hybrid fusion network
Xinwei Huang, Nannan Li 0002, Qing Xia 0002, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
Graph. Model. | 4 |
| 2022 | Novel cross LSTM for predicting the changes of complementary pelvic angles between standing and sittingabstractSagittal spino-pelvic balance has been increasingly emphasized in hip surgery. The conversion between standing and sitting, characterized by complementary pelvic angles (pelvic tilt, pt and sacral slope, ss), involves a congruent sagittal spino-pelvic relationship. Hence, the changes of complementary pelvic angles pt, ss between standing and sitting could reflect the mechanism of sagittal spino-pelvic balance, and should be analyzed in evidence-based hip surgery planning. To this end, we propose a novel cross LSTM (C-LSTM) framework embedding the conversion between standing and sitting by cross-mapping, to predict the changes of complementary pelvic pt, ss between standing and sitting. Furthermore, to introduce the prior knowledge of the invariance of pelvic incidence, pi, two dual C-LSTMs are integrated to construct a much more powerful Fused C-LSTM. We have conducted extensive experiments on the sagittal standing-sitting dataset for the comprehensive evaluation of the proposed framework. Even in a small samples, Fused C-LSTM can achieve low prediction errors and high correlation between predicted and actual values. Notably, just based on static standing or sitting X-ray, Fused C-LSTM can obtain the change of complementary pt, ss between standing and sitting to assist in formulating a surgical hip plan that conforms to the sagittal spino-pelvic balance. Yuanbo He, Minwei Zhao, Tianfan Xu, Shuai Li 0001 |
J. Biomed. Informatics | 4 |
| 2022 | Recursive multi-model complementary deep fusion for robust salient object detection via parallel sub-networks
Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
Pattern Recognit. | 2 |
| 2022 | Salient Object Detection via Dynamic Scale RoutingabstractRecent research advances in salient object detection (SOD) could largely be attributed to ever-stronger multi-scale feature representation empowered by the deep learning technologies. The existing SOD deep models extract multi-scale features via the off-the-shelf encoders and combine them smartly via various delicate decoders. However, the kernel sizes in this commonly-used thread are usually "fixed". In our new experiments, we have observed that kernels of small size are preferable in scenarios containing tiny salient objects. In contrast, large kernel sizes could perform better for images with large salient objects. Inspired by this observation, we advocate the "dynamic" scale routing (as a brand-new idea) in this paper. It will result in a generic plug-in that could directly fit the existing feature backbone. This paper's key technical innovations are two-fold. First, instead of using the vanilla convolution with fixed kernel sizes for the encoder design, we propose the dynamic pyramid convolution (DPConv), which dynamically selects the best-suited kernel sizes w.r.t. the given input. Second, we provide a self-adaptive bidirectional decoder design to accommodate the DPConv-based encoder best. The most significant highlight is its capability of routing between feature scales and their dynamic collection, making the inference process scale-aware. As a result, this paper continues to enhance the current SOTA performance. Both the code and dataset are publicly available at https://github.com/wuzhenyubuaa/DPNet. Shuai Li 0001, Chenglizhao Chen, Hong Qin 0001, Aimin Hao |
IEEE Trans. Image Process. | 2 |
| 2022 | Automatic Dental Plaque Segmentation Based on Local-to-Global Features Fused Self-Attention NetworkabstractThe accurate detection of dental plaque at an early stage will definitely prevent periodontal diseases and dental caries. However, it remains difficult for the current dental examination to accurately recognize dental plaque without using medical dyeing reagent due to the low contrast between dental plaque and healthy teeth. To combat this problem, this paper proposes a novel network enhanced by a self-attention module for intelligent dental plaque segmentation. The key motivation is to directly utilize oral endoscope images (bypassing the need for dyeing reagent) and get accurate pixel-level dental plaque segmentation results. The algorithm needs to conduct self-attention at the super-pixel level and fuse the super-pixels' local-to-global features. Our newly-designed network architecture will afford the simultaneous fusion of multiple-scale complementary information guided by the powerful deep learning paradigm. The critical fused information includes the statistical distribution of the plaques color, the heat kernel signature (HKS) based local-to-global structure relationship, and the circle-LBP based local texture pattern in the nearby regions centering around the plaque area. To further refine the fuzed multiple-scale features, we devise an attention module based on CNN, which could focalize the regions of interest in plaque more easily, especially for many challenging cases. Extensive experiments and comprehensive evaluations confirm that, for a small-scale training dataset, our method could outperform the state-of-the-art methods. Meanwhile, the user studies verify the claim that our method is more accurate than conventional dental practice conducted by experienced dentists. Shuai Li 0001, Zhennan Pang, Wenfeng Song, Aimin Hao, Hong Qin 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | Deeper Look at Image Salient Object Detection: Bi-Stream Network With a Small Training DatasetabstractCompared with the conventional hand-crafted approaches, the deep learning based ISOD (image salient object detection) models have achieved tremendous performance improvements by training exquisitely crafted fancy networks over large-scale training sets. However, do we really need large-scale training set for ISOD? In this article, we provide a deeper insight into the interrelationship between the ISOD performance and the training data. To alleviate the conventional demands for large-scale training data, we provide a feasible way to construct a novel small-scale training set, which only contains 4 K images. To take full advantage of this new set, we propose a novel bi-stream network consisting of two different feature backbones. Benefit from the proposed gate control unit, this bi-stream network is able to achieve complementary fusion status for its subbranches. To our best knowledge, this is the first attempt to use a small-scale training set to compete with other large-scale ones; nevertheless, our method can still achieve the leading SOTA performance on all tested benchmark datasets. Both the code and dataset are publicly available athttps://github.com/wuzhenyubuaa/TSNet. Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | Point Cloud Semantic Scene Completion from RGB-D ImagesabstractIn this paper, we devise a novel semantic completion network, called point cloud semantic scene completion network (PCSSC-Net), for indoor scenes solely based on point clouds. Existing point cloud completion networks still suffer from their inability of fully recovering complex structures and contents from global geometric descriptions neglecting semantic hints. To extract and infer comprehensive information from partial input, we design a patch-based contextual encoder to hierarchically learn point-level, patch-level, and scene-level geometric and contextual semantic information with a divide-and-conquer strategy. Consider that the scene semantics afford a high-level clue of constituting geometry for an indoor scene environment, we articulate a semantics-guided completion decoder where semantics could help cluster isolated points in the latent space and infer complicated scene geometry. Given the fact that real-world scans tend to be incomplete as ground truth, we choose to synthesize scene dataset with RGB-D images and annotate complete point clouds as ground truth for the supervised training purpose. Extensive experiments validate that our new method achieves the state-of-the-art performance, in contrast with the current methods applied to our dataset. Shoulong Zhang, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
AAAI | 2 |
| 2021 | Global Correlation and Local Geometric Information Coupled Channel Contrast Learning for Thyroid Nodule Risk StratificationabstractThyroid nodule risk stratification based on ultra-sound images is vital for follow-up clinical treatment. Due to the complexity of the inter-risk stratification difference, experienced physicians are required to comprehensively analyze all ultrasound signs of the thyroid nodule to diagnose the corresponding risk stratification the thyroid nodule should belong to, however, which is often labor-intensive, subjective, and unstable. To this end, we propose a global correlation and local geometric information coupled channel contrast learning network for thyroid nodule risk stratification based on ultrasound images. Specifically, a channel contrast learning module by combining the contrastive learning with a novel cross-class interaction strategy is proposed to obtain the discriminative feature for different risk stratification levels. Furthermore, multiple ultrasound signs may be observed simultaneously in the same risk stratification level. To introduce the correlation among different ultrasound signs into the feature, a global correlation learning module is proposed by the spectral decomposition of the channel correlation matrix. Additionally, some ultrasound signs used as the basis for judging risk stratification are local characteristics. Therefore, a local geometry learning module is proposed using gradient matrix to model the local information of ultrasound signs to further strengthen the feature. Extensive experiments and comprehensive evaluations confirm that the proposed method achieves superior performance and presents to be promising in thyroid nodule risk stratification. Yuanbo He, Shuai Li 0001, Luying Gao |
BIBM | 3 |
| 2021 | Knowledge-inspired 3D Scene Graph Prediction in Point CloudabstractPrior knowledge integration helps identify semantic entities and their relationships in a graphical representation, however, its meaningful abstraction and intervention remain elusive. This paper advocates a knowledge-inspired 3D scene graph prediction method solely based on point clouds. At the mathematical modeling level, we formulate the task as two sub-problems: knowledge learning and scene graph prediction with learned prior knowledge. Unlike conventional methods that learn knowledge embedding and regular patterns from encoded visual information, we propose to suppress the misunderstandings caused by appearance similarities and other perceptual confusion. At the network design level, we devise a graph auto-encoder to automatically extract class-dependent representations and topological patterns from the one-hot class labels and their intrinsic graphical structures, so that the prior knowledge can avoid perceptual errors and noises. We further devise a scene graph prediction model to predict credible relationship triplets by incorporating the related prototype knowledge with perceptual information. Comprehensive experiments confirm that, our method can successfully learn representative knowledge embedding, and the obtained prior knowledge can effectively enhance the accuracy of relationship predictions. Our thorough evaluations indicate the new method can achieve the state-of-the-art performance compared with other scene graph prediction methods. Shoulong Zhang, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
NeurIPS | 2 |
| 2021 | Correction to: Long-Short Temporal-Spatial Clues Excited Network for Robust Person Re-identification
Shuai Li 0001, Wenfeng Song, Zheng Fang 0008, Jiaying Shi, Aimin Hao, Qinping Zhao, Hong Qin 0001 |
Int. J. Comput. Vis. | 1 |
| 2021 | Depth quality-aware selective saliency fusion for RGB-D image salient object detection
Xuehao Wang, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
Neurocomputing | 2 |
| 2021 | Improving video anomaly detection performance by mining useful data from unseen video frames
Renzhi Wu, Shuai Li 0001, Chenglizhao Chen, Aimin Hao |
Neurocomputing | 2 |
| 2021 | S-CUDA: Self-cleansing unsupervised domain adaptation for medical image segmentationabstractMedical image segmentation tasks hitherto have achieved excellent progresses with large-scale datasets, which empowers us to train potent deep convolutional neural networks (DCNNs). However, labeling such large-scale datasets is laborious and error-prone, which leads the noisy (or incorrect) labels to be an ubiquitous problem in the real-world scenarios. In addition, data collected from different sites usually exhibit significant data distribution shift (or domain shift). As a result, noisy label and domain shift become two common problems in medical imaging application scenarios, especially in medical image segmentation, which degrade the performance of deep learning models significantly. In this paper, we identify a novel problem hidden in medical image segmentation, which is unsupervised domain adaptation on noisy labeled data, and propose a novel algorithm named "Self-Cleansing Unsupervised Domain Adaptation" (S-CDUA) to address such issue. S-CUDA sets up a realistic scenario to solve the above problems simultaneously where training data (i.e., source domain) not only shows domain shift w.r.t. unsupervised test data (i.e., target domain) but also contains noisy labels. The key idea of S-CUDA is to learn noise-excluding and domain invariant knowledge from noisy supervised data, which will be applied on the highly corrupted data for label cleansing and further data-recycling, as well as on the test data with domain shift for supervised propagation. To this end, we propose a novel framework leveraging noisy-label learning and domain adaptation techniques to cleanse the noisy labels and learn from trustable clean samples, thus enabling robust adaptation and prediction on the target domain. Specifically, we train two peer adversarial networks to identify high-confidence clean data and exchange them in companions to eliminate the error accumulation problem and narrow the domain gap simultaneously. In the meantime, the high-confidence noisy data are detected and cleansed in order to reuse the contaminated training data. Therefore, our proposed method can not only cleanse the noisy labels in the training set but also take full advantage of the existing noisy data to update the parameters of the network. For evaluation, we conduct experiments on two popular datasets (REFUGE and Drishti-GS) for optic disc (OD) and optic cup (OC) segmentation, and on another public multi-vendor dataset for spinal cord gray matter (SCGM) segmentation. Experimental results show that our proposed method can cleanse noisy labels efficiently and obtain a model with better generalization performance at the same time, which outperforms previous state-of-the-art methods by large margin. Our code can be found at https://github.com/zzdxjtu/S-cuda. Luyan Liu, Shuai Li 0001, Kai Ma 0002, Yefeng Zheng 0001 |
Medical Image Anal. | 3 |
| 2021 | Hierarchical Object Relationship Constrained Monocular Depth Estimation
Shuai Li 0001, Jiaying Shi, Wenfeng Song, Aimin Hao, Hong Qin 0001 |
Pattern Recognit. | 1 |
| 2021 | A Plug-and-Play Scheme to Adapt Image Saliency Deep Model for Video DataabstractWith the rapid development of deep learning techniques, image saliency deep models trained solely by spatial information have occasionally achieved detection performance for video data comparable to that of the models trained by both spatial and temporal information. However, due to the lesser consideration of temporal information, the image saliency deep models may become fragile in the video sequences dominated by temporal information. Thus, the most recent video saliency detection approaches have adopted the network architecture starting with a spatial deep model that is followed by an elaborately designed temporal deep model. However, such methods easily encounter the performance bottleneck arising from the single stream learning methodology, so the overall detection performance is largely determined by the spatial deep model. In sharp contrast to the current mainstream methods, this paper proposes a novel plug-and-play scheme to weakly retrain a pretrained image saliency deep model for video data by using the newly sensed and coded temporal information. Thus, the retrained image saliency deep model will be able to maintain temporal saliency awareness, achieving much improved detection performance. Moreover, our method is simple yet effective for adapting any off-the-shelf pre-trained image saliency deep model to obtain high-quality video saliency detection. Additionally, both the data and source code of our method are publicly available. Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | A Global-Local Self-Adaptive Network for Drone-View Object DetectionabstractDirectly benefiting from the deep learning methods, object detection has witnessed a great performance boost in recent years. However, drone-view object detection remains challenging for two main reasons: (1) Objects of tiny-scale with more blurs w.r.t. ground-view objects offer less valuable information towards accurate and robust detection; (2) The unevenly distributed objects make the detection inefficient, especially for regions occupied by crowded objects. Confronting such challenges, we propose an end-to-end global-local self-adaptive network (GLSAN) in this paper. The key components in our GLSAN include a global-local detection network (GLDN), a simple yet efficient self-adaptive region selecting algorithm (SARSA), and a local super-resolution network (LSRN). We integrate a global-local fusion strategy into a progressive scale-varying network to perform more precise detection, where the local fine detector can adaptively refine the target's bounding boxes detected by the global coarse detector via cropping the original images for higher-resolution detection. The SARSA can dynamically crop the crowded regions in the input images, which is unsupervised and can be easily plugged into the networks. Additionally, we train the LSRN to enlarge the cropped images, providing more detailed information for finer-scale feature extraction, helping the detector distinguish foreground and background more easily. The SARSA and LSRN also contribute to data augmentation towards network training, which makes the detector more robust. Extensive experiments and comprehensive evaluations on the VisDrone2019-DET benchmark dataset and UAVDT dataset demonstrate the effectiveness and adaptivity of our method. Towards an industrial application, our network is also applied to a DroneBolts dataset with proven advantages. Our source codes have been available at https://github.com/dengsutao/glsan. Sutao Deng, Shuai Li 0001, Ke Xie 0005, Wenfeng Song, Xiao Liao, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Rethinking Image Salient Object Detection: Object-Level Semantic Saliency Reranking First, Pixelwise Saliency Refinement LaterabstractHuman attention is an interactive activity between our visual system and our brain, using both low-level visual stimulus and high-level semantic information. Previous image salient object detection (SOD) studies conduct their saliency predictions via a multitask methodology in which pixelwise saliency regression and segmentation-like saliency refinement are conducted simultaneously. However, this multitask methodology has one critical limitation: the semantic information embedded in feature backbones might be degenerated during the training process. Our visual attention is determined mainly by semantic information, which is evidenced by our tendency to pay more attention to semantically salient regions even if these regions are not the most perceptually salient at first glance. This fact clearly contradicts the widely used multitask methodology mentioned above. To address this issue, this paper divides the SOD problem into two sequential steps. First, we devise a lightweight, weakly supervised deep network to coarsely locate the semantically salient regions. Next, as a postprocessing refinement, we selectively fuse multiple off-the-shelf deep models on the semantically salient regions identified by the previous step to formulate a pixelwise saliency map. Compared with the state-of-the-art (SOTA) models that focus on learning the pixelwise saliency in single images using only perceptual clues, our method aims at investigating the object-level semantic ranks between multiple images, of which the methodology is more consistent with the human attention mechanism. Our method is simple yet effective, and it is the first attempt to consider salient object detection as mainly an object-level semantic reranking problem. Guangxiao Ma, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Data-Level Recombination and Lightweight Fusion Scheme for RGB-D Salient Object DetectionabstractExisting RGB-D salient object detection methods treat depth information as an independent component to complement RGB and widely follow the bistream parallel network architecture. To selectively fuse the CNN features extracted from both RGB and depth as a final result, the state-of-the-art (SOTA) bistream networks usually consist of two independent subbranches: one subbranch is used for RGB saliency, and the other aims for depth saliency. However, depth saliency is persistently inferior to the RGB saliency because the RGB component is intrinsically more informative than the depth component. The bistream architecture easily biases its subsequent fusion procedure to the RGB subbranch, leading to a performance bottleneck. In this paper, we propose a novel data-level recombination strategy to fuse RGB with D (depth) before deep feature extraction, where we cyclically convert the original 4-dimensional RGB-D into DGB, RDB and RGD. Then, a newly lightweight designed triple-stream network is applied over these novel formulated data to achieve an optimal channel-wise complementary fusion status between the RGB and D, achieving a new SOTA performance. Xuehao Wang, Shuai Li 0001, Chenglizhao Chen, Yuming Fang 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Simulating Multi-Scale, Granular Materials and Their Transitions With a Hybrid Euler-Lagrange SolverabstractMulti-scale granular materials, such as powdered materials and mudslides, are pretty common in nature. Modeling such materials and their phase transitions remains challenging since this task involves the delicate representations of various ranges of particles with multiple scales that cause their property variations among liquid, granular solid (i.e., particles), and smoke-like materials. To effectively animate the complicated yet intriguing natural phenomena involving multi-scale granular materials and their phase transitions in graphics with high fidelity, this article advocates a hybrid Euler-Lagrange solver to handle the behaviors of involved discontinuous fluid-like materials faithfully. At the algorithmic level, we present a unified framework that tightly couples the affine particle-in-cell (APIC) solver with density field to achieve the transformation spanning across granular particles, dust cloud, powders, and their natural mixtures. For example, a part of the granular particles could be transformed into dust cloud while interacting with air and being represented by density field. Meanwhile, the velocity decrease of the involved materials could also result in the transit from the density-field-driven dust to powder particles. Besides, to further enhance our modeling and simulation power to broaden the range of multi-scale materials, we introduce a moisture property for granular particles to control the transitions between particles and viscous liquid. At the geometric level, we devise an additional surface-tracking procedure to simulate the viscous liquid phase. We can arrive at delicate viscous behaviors by controlling the corresponding yield conditions. Through various experiments with the different scenes design being conducted in our unified framework, we can validate the mixed multi-scale materials' mutual transformation processes. Our unified framework furnished with a hybrid solver can significantly enhance the modeling flexibility and the animation potential of the particle-grid hybrid materials in graphics. Yang Gao 0032, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Design and Evaluation of Personalized Percutaneous Coronary Intervention Surgery Simulation SystemabstractIn recent years, medical simulators have been widely applied to a broad range of surgery training tasks. However, most of the existing surgery simulators can only provide limited immersive environments with a few pre-processed organ models, while ignoring the instant modeling of various personalized clinical cases, which brings substantive differences between training experiences and real surgery situations. To this end, we present a virtual reality (VR) based surgery simulation system for personalized percutaneous coronary intervention (PCI). The simulation system can directly take patient-specific clinical data as input and generate virtual 3D intervention scenarios. Specially, we introduce a fiber-based patient-specific cardiac dynamic model to simulate the nonlinear deformation among the multiple layers of the cardiac structure, which can well respect and correlate the atriums, ventricles and vessels, and thus gives rise to more effective visualization and interaction. Meanwhile, we design a tracking and haptic feedback hardware, which can enable users to manipulate physical intervention instruments and interact with virtual scenarios. We conduct quantitative analysis on deformation precision and modeling efficiency, and evaluate the simulation system based on the user studies from 16 cardiologists and 20 intervention trainees, comparing it to traditional desktop intervention simulators. The results confirm that our simulation system can provide a better user experience, and is a suitable platform for PCI surgery training and rehearsal. Shuai Li 0001, Jiahao Cui 0001, Aimin Hao, Qinping Zhao |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | Meta-RetinaNet for Few-shot Object Detection
Shaoqi Li, Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
BMVC | 3 |
| 2020 | Meta Transfer Learning for Adaptive Vehicle Tracking in UAV Videos
Wenfeng Song, Shuai Li 0001, Shaoqi Li, Aimin Hao, Hong Qin 0001, Qinping Zhao |
MMM (1) | 2 |
| 2020 | Cross-View Contextual Relation Transferred Network for Unsupervised Vehicle Tracking in Drone VideosabstractRecently CNN-centric object tracking methods have been gaining tremendous success in ground-view videos, however, it remains hard to cope with vehicle tracking in unmanned aerial vehicle (UAV) videos. The key difficulties mainly stem from lacking large-scale well-labeled training datasets and view-invariant appearance model for fast-moving drone-view vehicles. We enhance the vehicle's cross-view feature by exploring relations between the pivotal context and the target to facilitate unsupervised vehicle tracking. The relation is modeled as the relevance of the target and its contextual regions in the tracking task. Specifically, we propose a contextual relation actor-critic (CRAC) framework integrates an actor-critic agent with a dual GAN learning mechanism, which aims to dynamically search the related contextual regions and transfer the relations from ground-view to drone-view videos while retaining the discriminative features. We demonstrate that CRAC could be applied to several state-of-the-art trackers by extensive experiments and ablation studies on four public benchmarks. All the experiments confirm that, our CRAC can improve the performance of state-of-the-art methods in terms of accuracy, robustness, and versatility. Wenfeng Song, Shuai Li 0001, Tao Chang, Aimin Hao, Qinping Zhao, Hong Qin 0001 |
WACV | 2 |
| 2020 | Attention-based relation and context modeling for point cloud semantic segmentation
Zhiyu Hu, Dongbo Zhang 0004, Shuai Li 0001, Hong Qin 0001 |
Comput. Graph. | 3 |
| 2020 | Accelerating Liquid Simulation With an Improved Data-Driven MethodabstractAbstract In physics‐based liquid simulation for graphics applications, pressure projection consumes a significant amount of computational time and is frequently the bottleneck of the computational efficiency. How to rapidly apply the pressure projection and at the same time how to accurately capture the liquid geometry are always among the most popular topics in the current research trend in liquid simulations. In this paper, we incorporate an artificial neural network into the simulation pipeline for handling the tricky projection step for liquid animation. Compared with the previous neural‐network‐based works for gas flows, this paper advocates new advances in the composition of representative features as well as the loss functions in order to facilitate fluid simulation with free‐surface boundary. Specifically, we choose both the velocity and the level‐set function as the additional representation of the fluid states, which allows not only the motion but also the boundary position to be considered in the neural network solver. Meanwhile, we use the divergence error in the loss function to further emulate the lifelike behaviours of liquid. With these arrangements, our method could greatly accelerate the pressure projection step in liquid simulation, while maintaining fairly convincing visual results. Additionally, our neutral network performs well when being applied to new scene synthesis even with varied boundaries or scales. Yang Gao 0032, Quancheng Zhang, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
Comput. Graph. Forum | 3 |
| 2020 | Spatiotemporal consistency-based adaptive hand-held video stabilization
Shuai Li 0001, Hong Qin 0001, Aimin Hao |
Sci. China Inf. Sci. | 2 |
| 2020 | Dynamic particle partitioning SPH model for high-speed fluids simulation
Yang Gao 0032, Jin Li 0068, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
Graph. Model. | 4 |
| 2020 | Long-Short Temporal-Spatial Clues Excited Network for Robust Person Re-identification
Shuai Li 0001, Wenfeng Song, Zheng Fang 0008, Jiaying Shi, Aimin Hao, Qinping Zhao, Hong Qin 0001 |
Int. J. Comput. Vis. | 1 |
| 2020 | Adaptive appearance modeling via hierarchical entropy analysis over multi-type features
Jizhou Ma, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
Pattern Recognit. | 2 |
| 2020 | Multi-Cue Semi-Supervised Color Constancy With Limited Training SamplesabstractColor constancy is one of the fundamental tasks in computer vision. Many supervised methods, including recently proposed Convolutional Neural Networks (CNN)-based methods, have been proved to work well on this problem, but they often require a sufficient number of labeled data. However, it is expensive and time-consuming to collect a large number of labeled training images with accurately measured illumination. In order to reduce the dependence on labeled images and leverage unlabeled ones without measured illumination, we propose a novel semi-supervised framework with limited training samples for illumination estimation. Our key insight is that the images with similar features from different cues will share similar lighting conditions. Consequently, three graphs based on three visual cues, low-level RGB color distribution, mid-level initial illuminant estimates and high-level scene content, are constructed to represent the relationship among different images. Then a multi-cue semi-supervised color constancy method (MSCC) is proposed after integrating these three graphs into a unified model. Extensive experiments on benchmark datasets demonstrate that our proposed MSCC method outperforms nearly all the existing supervised methods with limited labeled samples. Even with no unlabeled samples, MSCC still obtains better performance and stableness than most supervised methods. Xinwei Huang, Bing Li 0001, Shuai Li 0001, Weihua Xiong, Xuanwu Yin, Weiming Hu 0004, Hong Qin 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Context-Interactive CNN for Person Re-IdentificationabstractDespite growing progresses in recent years, cross-scenario person re-identification remains challenging, mainly due to the pedestrians commonly surrounded by highly-complex environment contexts. In reality, the human perception mechanism could adaptively find proper contextualized spatial-temporal clues towards pedestrian recognition. However, conventional methods fall short in adaptively leveraging the long-term spatial-temporal information due to ever-increasing computational cost. Moreover, CNN-based deep learning methods are hard to conduct optimization due to the non-differentiable property of the built-in context search operation. To ameliorate, this paper proposes a novel Context-Interactive CNN (CI-CNN) to dynamically find both spatial and temporal contexts by embedding multi-task Reinforcement Learning (MTRL). The CI-CNN streamlines the multi-task reinforcement learning by using an actor-critic agent to capture the temporal-spatial context simultaneously, which comprises a context-policy network and a context-critic network. The former network learns policies to determine the optimal spatial context region and temporal sequence range. Based on the inferred temporal-spatial cues, the latter one focuses on the identification task and provides feedback for the policy network. Thus, CI-CNN can simultaneously zoom in/out the perception field in spatial and temporal domain for the context interaction with the environment. By fostering the collaborative interaction between the person and context, our method could achieve outstanding performance on various public benchmarks, which confirms the rationality of our hypothesis, and verifies the effectiveness of our CI-CNN framework. Wenfeng Song, Shuai Li 0001, Tao Chang, Aimin Hao, Qinping Zhao, Hong Qin 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Accurate and Robust Video Saliency Detection via Self-Paced DiffusionabstractConventional video saliency detection methods frequently follow the common bottom-up thread to estimate video saliency within the short-term fashion. As a result, such methods can not avoid the obstinate accumulation of errors when the collected low-level clues are constantly ill-detected. Also, being noticed that a portion of video frames, which are not nearby the current video frame over the time axis, may potentially benefit the saliency detection in the current video frame. Thus, we propose to solve the aforementioned problem using our newly-designed key frame strategy (KFS), whose core rationale is to utilize both the spatial-temporal coherency of the salient foregrounds and the objectness prior (i.e., how likely it is for an object proposal to contain an object of any class) to reveal the valuable long-term information. We could utilize all this newly-revealed long-term information to guide our subsequent “self-paced” saliency diffusion, which enables each key frame itself to determine its diffusion range and diffusion strength to correct those ill-detected video frames. At the algorithmic level, we first divide a video sequence into short-term frame batches, and the object proposals are obtained in a frame-wise manner. Then, for each object proposal, we utilize a pre-trained deep saliency model to obtain high-dimensional features in order to represent the spatial contrast. Since the contrast computation within multiple neighbored video frames (i.e., the non-local manner) is relatively insensitive to the appearance variation, those object proposals with high-quality low-level saliency estimation frequently exhibit strong similarity over the temporal scale. Next, the long-term common consistency (e.g., appearance models/movement patterns) of the salient foregrounds could be explicitly revealed via similarity analysis accordingly. We further boost the detection accuracy via long-term information guided saliency diffusion in a self-paced manner. We have conducted extensive experiments to compare our method with 16 state-of-the-art methods over 4 largest public available benchmarks, and all results demonstrate the superiority of our method in terms of both accuracy and robustness. Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | Salient Object Detection via Multiple Instance Joint Re-LearningabstractIn recent years deep neural networks have been widely applied to visual saliency detection tasks with remarkable detection performance improvements. As for the salient object detection in single image, the automatically computed convolutional features frequently demonstrate high discriminative power to distinguish salient foregrounds from its non-salient surroundings in most cases. Yet, the obstinate feature conflicts still persist, which naturally gives rise to the learning ambiguity, arriving at massive failure detections. To solve such problem, we propose to jointly re-learn common consistency of inter-image saliency and then use it to boost the detection performance. Its core rationale is to utilize the easy-to-detect cases to re-boost much harder ones. Compared with the conventional methods, which focus on their problem domain within the single image scope, our method attempts to utilize those beyond-scope information to facilitate the current salient object detection. To validate our new approach, we have conducted a comprehensive quantitative comparisons between our approach and 13 state-of-the-art methods over 5 publicly available benchmarks, and all the results suggest the advantage of our approach in terms of accuracy, reliability, and versatility. Guangxiao Ma, Chenglizhao Chen, Shuai Li 0001, Chong Peng 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Multim. | 3 |
| 2020 | Contextualized CNN for Scene-Aware Depth Estimation From Single RGB ImageabstractDirectly benefited from deep learning techniques, depth estimation from single image has gained great momentum in recent years. However, most of the existing approaches treat depth prediction as an isolated problem without taking into consideration high-level semantic context information, which results in inefficient utilization of training dataset and unavoidably requires a large number of captured depth data during the training phase. To ameliorate, this paper develops a novel scene-aware contextualized convolution neural network (CCNN), which characterizes the semantic context relationship at the class-level and refines depth at the pixel-level. Our newly-proposed CCNN is built upon the intrinsic exploitation of context-dependent depth association, including inner-object continuous depth and inter-object depth change priors nearby. Specifically, rather than conducting regression on depth in single CNN, we make the first attempt to integrate both class-level and pixel-level conditional random fields (CRFs) based probabilistic graphical model into the powerful CNN framework to simultaneously learn different-level features within the same CNN layer. With our CCNN, the former model will guide the latter one to learn the contextualized RGB-Depth mapping. Hence, CCNN has desirable properties in both class-level integrity and pixel-level discrimination, which makes it ideal to share such two-level convolutional features in parallel during the end-to-end training with the commonly-used back-propagation algorithm. We conduct extensive experiments and comprehensive evaluations on public benchmarks involving various indoor and outdoor scenes, and all the experiments confirm that, our method outperforms the state-of-the-art depth estimation methods, especially for the cases where only small-scale training data are readily available. Wenfeng Song, Shuai Li 0001, Aimin Hao, Qinping Zhao, Hong Qin 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | Stage-wise Salient Object Detection in 360° Omnidirectional Image via Object-level Semantical Saliency RankingabstractThe 2D image based salient object detection (SOD) has been extensively explored, while the 360° omnidirectional image based SOD has received less research attention and there exist three major bottlenecks that are limiting its performance. Firstly, the currently available training data is insufficient for the training of 360° SOD deep model. Secondly, the visual distortions in 360° omnidirectional images usually result in large feature gap between 360° images and 2D images; consequently, the widely used stage-wise training-a widely-used solution to alleviate the training data shortage problem, becomes infeasible when conducing SOD in 360° omnidirectional images. Thirdly, the existing 360° SOD approach has followed a multi-task methodology that performs salient object localization and segmentation-like saliency refinement at the same time, being faced with extremely large problem domain, making the training data shortage dilemma even worse. To tackle all these issues, this paper divides the 360° SOD into a multi-staqe task, the key rationale of which is to decompose the original complex problem domain into sequential easy sub problems that only demand for small-scale training data. Meanwhile, we learn how to rank the "object-level semantical saliency", aiming to locate salient viewpoints and objects accurately. Specifically, to alleviate the training data shortage problem, we have released a novel dataset named 360-SSOD, containing 1,105 360° omnidirectional images with manually annotated object-level saliency ground truth, whose semantical distribution is more balanced than that of the existing dataset. Also, we have compared the proposed method with 13 SOTA methods, and all quantitative results have demonstrated the performance superiority. Guangxiao Ma, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2020 | Compressing animated meshes with fine details using local spectral analysis and deformation transfer
Chengju Chen, Qing Xia 0002, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
Vis. Comput. | 3 |
| 2020 | Personalized cardiovascular intervention simulation systemabstractBackground This study proposes a series of geometry and physics modeling methods for personalized cardiovascular intervention procedures, which can be applied to a virtual endovascular simulator. Methods Based on personalized clinical computed tomography angiography (CTA) data, mesh models of the cardiovascular system were constructed semi-automatically. By coupling 4D magnetic resonance imaging (MRI) sequences corresponding to a complete cardiac cycle with related physics models, a hybrid kinetic model of the cardiovascular system was built to drive kinematics and dynamics simulation. On that basis, the surgical procedures related to intervention instruments were simulated using specially-designed physics models. These models can be solved in real-time; therefore, the complex interactions between blood vessels and instruments can be well simulated. Additionally, X-ray imaging simulation algorithms and realistic rendering algorithms for virtual intervention scenes are also proposed. In particular, instrument tracking hardware with haptic feedback was developed to serve as the interaction interface of real instruments and the virtual intervention system. Finally, a personalized cardiovascular intervention simulation system was developed by integrating the techniques mentioned above. Results This system supported instant modeling and simulation of personalized clinical data and significantly improved the visual and haptic immersions of vascular intervention simulation. Conclusions It can be used in teaching basic cardiology and effectively satisfying the demands of intervention training, personalized intervention planning, and rehearsing. Aimin Hao, Jiahao Cui 0001, Shuai Li 0001, Qinping Zhao |
Virtual Real. Intell. Hardw. | 3 |
| 2019 | Fine-Grained Thyroid Nodule Classification via Multi-Semantic Attention NetworkabstractThyroid nodule classification in ultrasound images has gained great momentum based on deep convolutional neural networks in recent years. Nevertheless, it is still challenging to intelligently classify the fine-grained thyroid nodules, which is significant for the subsequent clinical treatments. The difficulties mainly stem from four aspects: few fine-grained training dataset, highly-variable appearances of intra-class nodules, overall-similar characteristics of inter-class nodules, and the low resolution and contrast degree of the ultrasonic images as well as the influence of intrinsic speckle noises. In this paper, we propose a multi-semantic attention networks (MSAN) for fine-grained thyroid nodule classification in ultrasound images. Specifically, we employ a main network branch for coarse granularity feature extraction, which only focuses on the benign and malignant characteristics, and simultaneously employ multi-semantic network branches to extract discriminative features from the fine-grained pathological categories. Meanwhile, we introduce an self-attention scheme together with global average pooling (GAP) in our network, which facilitates to learn from the dynamically-selected nodule regions ranging from local to global. Extensive experiments demonstrate that, our MSAN gives rise to significant improvement of classification accuracy and outperforms the state-of-the-art methods. Shuai Li 0001, Wenfeng Song, Zhennan Pang, Aimin Hao, Hong Qin 0001 |
BIBM | 1 |
| 2019 | Few-Shot Learning for Monocular Depth Estimation Based on Local Object RelationshipabstractMonocular depth estimation has gained great momentum and achieved growing success recently. Nonetheless, due to the intrinsic difficulty associated with large-scale RGB-D data capture for training purpose and the inefficient utilization of existing training datasets, it is still challenging to accommodate flexibly-changing scenarios. To ameliorate, we propose a fewshot learning method for monocular depth estimation augmented by local object-object relationship. Our method is based on the insight that the depth changing between neighboring objects is relatively stable across diverse but similar scenarios. At the technical front, we first learn the object relationship based on the relative distance between single objects. Towards this goal, we design a CNN architecture to simultaneously encode the object spatial context into object-object relationship features and encode the original image into global context features. Hence we can complementally leverage few-shot dataset with only a few samples for depth estimation while preserving the global depth changing range and respecting the local object-object depth details. As a result, our novel approach could estimate depth from various indoor RGB images, which greatly alleviates the training dataset dependency in monocular depth estimation. Finally, we conduct extensive experiments and comprehensive evaluations on the widely-used public benchmarks, and all the experiments confirm that, our method outperforms the state-of-the-art depth estimation methods, especially for the cases where only smallscale training samples are available. Shuai Li 0001, Jiaying Shi, Wenfeng Song, Aimin Hao, Hong Qin 0001 |
ICTAI | 1 |
| 2019 | Context-Aware Network for 3D Human Pose Estimation from Monocular RGB ImageabstractConvolutional Neural Network (CNN) has brought tremendous improvements in estimating 3D human pose from a monocular RGB image. However, the task of 3D human pose estimation still remains extremely challenging, especially when the task is geared towards estimating the depth of human body parts. Different from 2D human pose estimation, which focuses on the fusion of spatial information and context information, depth estimation demands more context information. Inspired by this, we build a Context-Aware Network (CAN) which can fully explore the context information to discover the underlying relationships among different body parts. The key ingredient of our network is High-Level Depth Estimation Module (HLDEM) designed to extract context information effectively. Additionally, multi-scale supervision is introduced in our network to extract context information at different scales. Experimental results show that our network achieves competitive performance compared with state-of-the-art methods on Human3.6M dataset. Binyi Yin, Dongbo Zhang 0004, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
IJCNN | 3 |
| 2019 | A Hybrid Method for Powdered Materials ModelingabstractPowdered materials, such as sand and flour, are quite common in nature, whose properties always range from granular particles to smog materials under the air friction while throwing. This paper presents a hybrid method that tightly couples APIC solver with density field to accomplish the transformation of continuous powdered materials varying among granular particles, smog, powders and their natural mixtures. In our method, a part of the granular particles will be transformed to dust smog while interacting with air and represented by density field, then, as velocity decreases the density-based dust will deposit to powder particles. We construct a unified framework to imitate the mutual transformation process for the powdered materials of different scales, which greatly enhance the details of particle-based materials modeling. We have conducted extensive experiments to verify the performance of our model, and get satisfactory results in terms of stability, efficiency and visual authenticity as expected. Yang Gao 0032, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
VRST | 3 |
| 2019 | Quantitative and flexible 3D shape dataset augmentation via latent space embedding and deformation learning
Jiarui Liu 0003, Qing Xia 0002, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
Comput. Aided Geom. Des. | 3 |
| 2019 | Learning multi-view manifold for single image based modeling
Jiahao Cui 0001, Shuai Li 0001, Qing Xia 0002, Aimin Hao, Hong Qin 0001 |
Comput. Graph. | 2 |
| 2019 | Efficient 4D shape completion from sparse samples via cubic spline fitting in linear rotation-invariant space
Qing Xia 0002, Chengju Chen, Jiarui Liu 0003, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
Comput. Graph. | 4 |
| 2019 | Hybrid 4D cardiovascular modeling based on patient-specific clinical images for real-time PCI surgery simulation
Shuai Li 0001, Zhijun Xie, Qing Xia 0002, Aimin Hao, Hong Qin 0001 |
Graph. Model. | 1 |
| 2019 | Bidirectional Optimization Coupled Lightweight Networks for Efficient and Robust Multi-Person 2D Pose Estimation
Shuai Li 0001, Zheng Fang 0008, Wenfeng Song, Aimin Hao, Hong Qin 0001 |
J. Comput. Sci. Technol. | 1 |
| 2019 | Multitask Cascade Convolution Neural Networks for Automatic Thyroid Nodule Detection and RecognitionabstractThyroid ultrasonography is a widely used clinical technique for nodule diagnosis in thyroid regions. However, it remains difficult to detect and recognize the nodules due to low contrast, high noise, and diverse appearance of nodules. In today's clinical practice, senior doctors could pinpoint nodules by analyzing global context features, local geometry structure, and intensity changes, which would require rich clinical experience accumulated from hundreds and thousands of nodule case studies. To alleviate doctors' tremendous labor in the diagnosis procedure, we advocate a machine learning approach to the detection and recognition tasks in this paper. In particular, we develop a multitask cascade convolution neural network (MC-CNN) framework to exploit the context information of thyroid nodules. It may be noted that our framework is built upon a large number of clinically confirmed thyroid ultrasound images with accurate and detailed ground truth labels. Other key advantages of our framework result from a multitask cascade architecture, two stages of carefully designed deep convolution networks in order to detect and recognize thyroid nodules in a pyramidal fashion, and capturing various intrinsic features in a global-to-local way. Within our framework, the potential regions of interest after initial detection are further fed to the spatial pyramid augmented CNNs to embed multiscale discriminative information for fine-grained thyroid recognition. Experimental results on 4309 clinical ultrasound images have indicated that our MC-CNN is accurate and effective for both thyroid nodules detection and recognition. For the correct diagnosis rate of malignant and benign thyroid nodules, its mean Average Precision (mAP) performance can achieve up to [Formula: see text] accuracy, which outperforms the common CNNs by [Formula: see text] on average. In addition, we conduct rigorous user studies to confirm that our MC-CNN outperforms experienced doctors, yet only consuming roughly [Formula: see text] ( 1/48) of doctors' examination time on average. Therefore, the accuracy and efficiency of our new method exhibit its great potential in clinical applications. Wenfeng Song, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | An efficient FLIP and shape matching coupled method for fluid-solid and two-phase fluid simulations
Yang Gao 0032, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
Vis. Comput. | 2 |
| 2018 | A Novel Radiogenomics Framework for Genomic and Image Feature Correlation using Deep Learning
Shuai Li 0001, Hongze Han, Dong Sui, Aimin Hao, Hong Qin 0001 |
BIBM | 1 |
| 2018 | High-fidelity Compression of Dynamic Meshes with Fine Details using Piece-wise Manifold Harmonic BasesabstractMesh-based animation, usually represented as dynamic meshes with fixed connectivity, is becoming more and more prevalent in movies, games and other graphics applications nowadays, and there is a growing need to compactly store and rapidly transmit these meshes for practical use, especially for those with high-quality geometric details. In this paper, we explore a novel key-frame based dynamic mesh compression method, wherein we apply pose-similarity with spectral techniques to define piece-wise manifold harmonic bases to reduce spatial-temporal redundancy. We first partition the sequence into several clusters with similar poses, and then decompose the meshes in each cluster into primary poses and geometric details using the manifold harmonic bases derived from the extracted key-frame in that cluster. The primary poses can be characterized as linear combinations of manifold harmonic bases, and the geometric details can be recovered by deformation transfer technique. Thus, we only need a small number of key-frames and a few coefficients for compressing dynamic meshes, which saves a significant amount of storage comparing to traditional methods in which bases are stored explicitly. Furthermore, we apply a second-order linear prediction coding to the harmonic coefficients to further reduce the temporal redundancy. Our extensive experiments and evaluations on various datasets have manifested that our novel method could obtain a high compression ratio while preserving high-fidelity geometry details and guaranteeing limited human perceived distortion rate simultaneously. Chengju Chen, Qing Xia 0002, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
CGI | 3 |
| 2018 | Automatic Beautification for Group-Photo Facial Expressions Using Novel Bayesian GANs
Shuai Li 0001, Wenfeng Song, Hong Qin 0001, Aimin Hao |
ICANN (1) | 2 |
| 2018 | Learning from Weakly-Labeled Clinical Data for Automatic Thyroid Nodule Classification in Ultrasound ImagesabstractThis paper proposes a semi-supervised learning method based on weakly-labeled data to automatically classify ultrasound (US) thyroid nodules. Key to our new approach is the unification of multi-instance learning (MIL) with deep learning. Benefiting from that, our method can directly use off-the-shelf clinical data, which involves no labels to indicate nodule classes. To this end, we take the US images of a patient as a bag, and take the corresponding pathology report as the bag label. Specifically, we first propose a bag generating method, wherein the detected thyroid nodules are considered as instances corresponding to certain bag. After that, we design an effective EM algorithm to train a convolutional neural network (CNN) for nodule classification. We conduct extensive experiments and comprehensive evaluations on different datasets, and all the experiments confirm that, our method significantly outperforms state-of-the-art MIL algorithms, which exhibits great potential in clinical applications. Jianxiong Wang, Shuai Li 0001, Wenfeng Song, Hong Qin 0001, Aimin Hao |
ICIP | 2 |
| 2018 | Deep variance network: An iterative, improved CNN framework for unbalanced training datasets
Shuai Li 0001, Wenfeng Song, Hong Qin 0001, Aimin Hao |
Pattern Recognit. | 1 |
| 2018 | A Novel Bottom-Up Saliency Detection Method for Video With Dynamic BackgroundabstractAfter years of extensive studies, the salient motion detection problem has gained plausible performance improvement that was primarily propelled by the rapid development of self-adaptive top-down modeling techniques. Nevertheless, almost all the conventional solutions are still not robust enough to handle video sequences captured by hand-hold cameras. This is mainly due to the absence of the position alignment information that is indispensable for top-down background modeling. In contrast, the bottom-up video saliency detection methods, though achieving excellent salient motion detection in either stationary or nonstationary videos, still have rather poor detection performance in scenarios with massive dynamic background. In this letter, we explore a bottom-up saliency framework by introducing a novel spatial-temporal regional filter method to handle the dynamic background problem. Our key rationale is to assign large saliency value to those regions with stable spatial-temporal coherency while eliminating irregular, repeating dynamic background. As far as we know, this is the first work to address the dynamic background problem from the perspective of the bottom-up video saliency. We conduct massive quantitative evaluations over public available benchmarks to validate the effectiveness and robustness of our method. Chenglizhao Chen, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
IEEE Signal Process. Lett. | 3 |
| 2018 | Bilevel Feature Learning for Video Saliency DetectionabstractThis paper advocates a novel learning solution to the modeling of long-term spatial-temporal saliency consistency in order to boost the accuracy for video saliency detection. Conventional methods typically utilize the “slack” spatial-temporal model to locally ensure the smoothness of the computed video saliency, yet they could easily encounter the performance tradeoff dilemma (i.e., detection' accuracy and integrity). In contrast, our novel approach proposes the bilevel learning strategy to globally exploit the saliency consistency while overcoming the aforementioned difficulty. Our method first starts with the contrast computation of low-level saliency clues in a frame-wise manner. Then, based on such obtained saliency clues, we devise a novel bilevel Markov Random Field (bMRF) solution to conduct semantic labelling, which can explicitly indicates both the salient salient foregrounds and nonsalient nearby surroundings with high confidence while shrinking the low confidence remains. In such a way, the spatial-temporal consistency constraint is embedded intrinsically into the above explicit semantic labels, and we prevent the performance tradeoff problem from occurring. Next, based on those semantic labels made by our bMRF method, we further propose learning multiple nonlinear feature transformations to enlarge the feature margin between the salient foregrounds and the non-salient nearby surroundings, whose key rationale is to resort to long-term common consistencies to enforce the spatial-temporal smoothness. Thus, we can utilize these learned non-linear feature transformations to simultaneously suppress those short-term false-alarms and correct those hollow effects. To validate our new approach, we conduct extensive experiments on five publicly available benchmarks, and make comprehensive, quantitative evaluations between our method and 17 state-of-the-art techniques. All of the results demonstrate our method's advantages in terms of accuracy, reliability, robustness, and versatility. Chenglizhao Chen, Shuai Li 0001, Hong Qin 0001, Zhenkuan Pan 0001, Guowei Yang 0002 |
IEEE Trans. Multim. | 2 |
| 2017 | A novel fluid-solid coupling framework integrating FLIP and shape matching methodsabstractPhysically-based fluid animation and solid deformation driven by numerical simulation have manifested their significance for many graphics applications during the past two decades. For example, the fluid implicit particle (FLIP) method and shape matching technique based on position based dynamics (PBD) have demonstrated their unique graphics strength in fluid and solid animation, respectively. We propose a novel integrated approach supporting the seamless unification of FLIP and shape matching. We devise new algorithms to tackle existing difficulties when handling new phenomena such as high-fidelity fluid-solid interaction and solid melting. The key innovation of this paper is a unified Lagrangian framework that seamlessly blends FLIP and PBD based shape matching constraint towards the natural yet strong coupling between fluid and deformable solid. Within our integrated framework, it enables many complicated fluid-solid phenomena with ease. We conduct various kinds of experiments. All the results demonstrate the advantages of our unified hybrid approach towards visual fidelity, efficiency, stability, and versatility. Yang Gao 0032, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
CGI | 2 |
| 2017 | Hessian-constrained detail-preserving 3D implicit reconstruction from raw volumetric dataset
Shuai Li 0001, Dehui Yan, Aimin Hao, Hong Qin 0001 |
Comput. Graph. | 1 |
| 2017 | An efficient heat-based model for solid-liquid-gas phase transition and dynamic interaction
Yang Gao 0032, Shuai Li 0001, Lipeng Yang, Hong Qin 0001, Aimin Hao |
Graph. Model. | 2 |
| 2017 | Novel fluid detail enhancement based on multi-layer depth regression analysis and FLIP fluid simulationabstractAbstract In this paper, we propose a novel integrated method for effective modeling and realistic enhancement of scale‐sensitive fluid simulation details. The core of our method is the organic of multi‐layer depth image regression analysis and fluid implicit particle fluid simulation of which the regression analysis induces the criterion where the fluid details should be produced. First, we capture the depth buffer of the fluid surface dynamically from the top of scene. Second, we employ depth peeling technique to decompose the target fluid volume into multiple depth layers and conduct time‐space analysis over surface layers. Third, we propose a logistic regression‐based model to rigorously pinpoint the complex interacting regions, wherein multiple detail‐relevant factors are taken into account based on the captured multiple depth layers. Finally, details are enhanced by animating extra diffuse materials and augmenting the air‐fluid mixing phenomenon. It is evident that, with depth peeling technology, we can afford rigorous analysis not only across surface layers at different fluid depth but along the depth direction as well. After integrating the analysis results from these two sources, we are capable of performing detail enhancement both on the fluid surface and inside the fluid to obtain a great visual effect, even when large occlusion exists. Directly benefiting from the flexibility of image‐space‐dominant processing, our unified framework can be entirely implemented on graphics processing units and thus achieves interactive performance. For various fluid phenomena with different diffuse materials (e.g., spray, foam, and bubble), comprehensive experiments and evaluations have demonstrated its superiority in high‐fidelity fluid detail enhancement and its interaction with surrounding environment. Yuxing Qiu, Lipeng Yang, Shuai Li 0001, Qing Xia 0002, Hong Qin 0001, Aimin Hao |
Comput. Animat. Virtual Worlds | 3 |
| 2017 | Video Saliency Detection via Spatial-Temporal Fusion and Low-Rank Coherency DiffusionabstractThis paper advocates a novel video saliency detection method based on the spatial-temporal saliency fusion and low-rank coherency guided saliency diffusion. In sharp contrast to the conventional methods, which conduct saliency detection locally in a frame-by-frame way and could easily give rise to incorrect low-level saliency map, in order to overcome the existing difficulties, this paper proposes to fuse the color saliency based on global motion clues in a batch-wise fashion. And we also propose low-rank coherency guided spatial-temporal saliency diffusion to guarantee the temporal smoothness of saliency maps. Meanwhile, a series of saliency boosting strategies are designed to further improve the saliency accuracy. First, the original long-term video sequence is equally segmented into many short-term frame batches, and the motion clues of the individual video batch are integrated and diffused temporally to facilitate the computation of color saliency. Then, based on the obtained saliency clues, inter-batch saliency priors are modeled to guide the low-level saliency fusion. After that, both the raw color information and the fused low-level saliency are regarded as the low-rank coherency clues, which are employed to guide the spatial-temporal saliency diffusion with the help of an additional permutation matrix serving as the alternative rank selection strategy. Thus, it could guarantee the robustness of the saliency map's temporal consistence, and further boost the accuracy of the computed saliency map. Moreover, we conduct extensive experiments on five public available benchmarks, and make comprehensive, quantitative evaluations between our method and 16 state-of-the-art techniques. All the results demonstrate the superiority of our method in accuracy, reliability, robustness, and versatility. Chenglizhao Chen, Shuai Li 0001, Yongguang Wang, Hong Qin 0001, Aimin Hao |
IEEE Trans. Image Process. | 2 |
| 2017 | Unsupervised Multi-Class Co-Segmentation via Joint-Cut Over L1 -Manifold Hyper-Graph of Discriminative Image RegionsabstractThis paper systematically advocates a robust and efficient unsupervised multi-class co-segmentation approach by leveraging underlying subspace manifold propagation to exploit the cross-image coherency. It can combat certain image co-segmentation difficulties due to viewpoint change, partial occlusion, complex background, transient illumination, and cluttering texture patterns. Our key idea is to construct a powerful hyper-graph joint-cut framework, which incorporates mid-level image regions-based intra-image feature representation and L1-manifold graph-based inter-image coherency exploration. For local image region generation, we propose a bi-harmonic distance distribution difference metric to govern the super-pixel clustering in a bottom-up way. It not only affords drastic data reduction but also gives rise to discriminative and structure meaningful feature representation. As for the inter-image coherency, we leverage multi-type features involved L1-graph to detect the underlying local manifold from cross-image regions. As a result, the implicit supervising information could be encoded into the unsupervised hyper-graph joint-cut framework. We conduct extensive experiments and make comprehensive evaluations with other state-of-the-art methods over various benchmarks, including iCoseg, MSRC, and Oxford flower. All the results demonstrate the superiorities of our method in terms of accuracy, robustness, efficiency, and versatility. Jizhou Ma, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
IEEE Trans. Image Process. | 2 |
| 2016 | Detail-Preserving 3D Shape Modeling from Raw Volumetric Dataset via Hessian-Constrained Local Implicit Surfaces OptimizationabstractMassive routinely-acquired raw volumetric datasets are hard to be deeply exploited by cyber worlds related downstream applications due to the challenges in accurate and efficient shape modeling. This paper systematically advocates an interactive 3D shape modeling framework for raw volumetric datasets by iteratively optimizing Hessian-constrained local implicit surfaces. The key idea is to incorporate contour based interactive segmentation into the generalized local implicit surface reconstruction. Our framework allows a user to flexibly define derivative constraints up to the second order via intuitively placing contours on the cross sections of volumetric images and fine-tuning the eigenvector frame of Hessian matrix. It enables detail-preserving local implicit representation while combating certain difficulties due to ambiguous image regions, low-quality irregular data, close sheets, and massive coefficients involved extra computing burden. Moreover, we conduct extensive experiments on some volumetric images with blurry object boundaries, and make comprehensive, quantitative performance evaluation between our method and the state-of-the-art radial basis function based techniques. All the results demonstrate our method's advantages in the accuracy, detail-preserving, efficiency, and versatility of shape modeling. Shuai Li 0001, Dehui Yan, Aimin Hao, Hong Qin 0001 |
CW | 1 |
| 2016 | Automatic extraction of generic focal features on 3D shapes via random forest regression analysis of geodesics-in-heat
Qing Xia 0002, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
Comput. Aided Geom. Des. | 2 |
| 2016 | Coupling time-varying modal analysis and FEM for real-time cutting simulation of objects with multi-material sub-domains
Chen Yang 0002, Shuai Li 0001, Lili Wang 0006, Aimin Hao, Hong Qin 0001 |
Comput. Aided Geom. Des. | 2 |
| 2016 | Haptics-equiped interactive PCI simulation for patient-specific surgery training and rehearsing
Shuai Li 0001, Qing Xia 0002, Aimin Hao, Hong Qin 0001, Qinping Zhao |
Sci. China Inf. Sci. | 1 |
| 2016 | Automatic non-parametric image parsing via hierarchical semantic voting based on sparse-dense reconstruction and spatial-contextual cues
Xinyi An, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
Neurocomputing | 2 |
| 2016 | Robust salient motion detection in non-stationary videos via novel integrated strategies of spatio-temporal coherency clues and low-rank analysis
Chenglizhao Chen, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
Pattern Recognit. | 2 |
| 2016 | Super-Resolution of Multi-Observed RGB-D Images Based on Nonlocal Regression and Total VariationabstractThere is growing demand for accuracy in image processing and visualization, and the super-resolution (SR) technique for multi-observed RGB-D images has become popular, because it provides space-redundant information and produces a detailed reconstruction even with a large magnification factor. This technique has been thoroughly investigated in recent years. Nevertheless, technical challenges remain, such as finding sub-pixel correspondences with low-resolution (LR) observations, exploiting space-redundant information, formulating space homogeneity constraints, and leveraging cross-image similarities in structures. To address these challenges, this paper proposes a unified optimization framework to estimate both the super-resolved RGB image and the super-resolved depth image from the multi-observed LR RGB-D images using their correlations. Using depth-assisted cross-image correspondences, the RGB image SR problem is formulated as an effective regularization function by incorporating the normalized bilateral total variation regularizer, and it is efficiently solved by a first-order primal-dual algorithm. The depth image SR estimate can be obtained by minimizing a nonlocal regression-based energy, which integrates the structural cues of the super-resolved RGB image in a detail-preserving fashion. Essentially, our unified optimization framework uses the RGB image and depth image as a priori knowledge that the SR process uses for better accuracy. Our extensive experiments on public RGB-D benchmarks and real data and our quantitative comparison with several state-of-the-art methods demonstrate the superiority of our method in terms of accuracy, versatility, and reliability of details and sharp feature preservation. Qingzheng Wang, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
IEEE Trans. Image Process. | 2 |
| 2015 | Novel, Robust, and Efficient Guidewire Modeling for PCI Surgery Simulator Based on Heterogeneous and Integrated Chain-MailsabstractDespite the long R&D history of interactive minimally-invasive surgery and therapy simulations, the guide wire/catheter behavior modeling remains challenging in Percutaneous Coronary Intervention (PCI) surgery simulators. This is primarily due to the heterogeneous heart physiological structures and complex intravascular inter-dynamic procedures. To ameliorate, this paper advocates a novel, robust, and efficient guide wire/catheter modeling method based on heterogeneous and integrated chain-mails, that can afford medical practitioners and trainees the unique opportunity to experience the entire guide wire-dominant PCI procedures in virtual environments as our model aims to mimic what occurs in clinical settings. Our approach's originality is primarily founded upon this new method's unconditional stability, real time performance, flexibility, and high-fidelity realism for guide wire/catheter simulation. Considering the front end of the guide wire has different stiffness with its conjunctive slender body and the guide wire length is adaptive to the surrounding environment, we propose to model the spatially-varying six-degree of freedom behaviors by solely resorting to the generalized 3D chain-mails. Meanwhile, to effectively accommodate the motion constraints caused by the beating vessels and flowing blood, we integrate heterogeneous volumetric chain mails to streamline guide wire modeling and its interaction with surrounding substances. By dynamically coupling guide wire chain-mails with the surrounding media via virtual links, we are capable of efficiently simulating the collision-involved interdynamic behaviors of the guide wire. Finally, we showcase a PCI prototype simulator equipped with hap tic feedback for mimicing the guide wire intervention therapy, including pushing, pulling, and twisting operations, where the built-in high-fidelity, real-time efficiency, and stableness show great promise for its practical applications in clinical training and surgery rehearsal fields. Shuai Li 0001, Hong Qin 0001, Aimin Hao |
CAD/Graphics | 2 |
| 2015 | Interactive volumetric segmentation through least-squares optimization of local hessian-constrained implicitsabstractA great number of volumetric datasets have been routinely acquired everyday and their qualities are varying tremendously, without proper processing they could not be directly utilized. Specifically, volumetric segmentation plays a vital role in many downstream applications, including geometric modeling, scientific visualization, and medical diagnosis. So far, many volume segmentation methods have been proposed, Top et al. [2011] designed an interactive segmentation tool by interactively contouring on some sparse slices and Ijiri et al. [2013] developed a system to extract contours and evaluate the scalar field in spatial domain. Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
VRST | 2 |
| 2015 | A novel integrated analysis-and-simulation approach for detail enhancement in FLIP fluid interactionabstractThis paper advocates a novel integrated method to tightly couple simulation with analysis for the effective modeling and enhancement of scale-aware fluid details. It brings forth a suite of innovations in a unified framework, including depth-image-based space analysis for multi-scale detail detection, time-space analysis based on the logistic regression model that integrates both geometry and physics criteria, and depth-image-based sampling for quality-efficiency tradeoff. Our method contains an intertwined two-level processing architecture at its core. At the analysis level, we propose a rigorous time-space analysis model to pinpoint complex interacting regions, which can take into account multiple detail-relevant factors based on the depth-image sequence captured from FLIP-driven simulation sequence. At the simulation level, details are enhanced by animating extra diffuse materials, and augmenting the air-fluid mixing phenomenon. Directly benefitting from the flexibility of image-space-dominant processing, our unified framework can be entirely implemented on GPU, hence interactive performance could be guaranteed. Comprehensive experiments and evaluations on various diffuse phenomena (e.g., spray, foam, and bubble) have demonstrated its superiority in high-fidelity detail enhancement during fluid simulation and its interaction with surrounding environment for VR applications. Lipeng Yang, Shuai Li 0001, Qing Xia 0002, Hong Qin 0001, Aimin Hao |
VRST | 2 |
| 2015 | Multi-scale mesh saliency based on low-rank and sparse analysis in shape feature space
Shengfa Wang, Nannan Li 0002, Shuai Li 0001, Zhongxuan Luo, Zhixun Su, Hong Qin 0001 |
Comput. Aided Geom. Des. | 3 |
| 2015 | Real-time and robust object tracking in video via low-rank coherency analysis in feature space
Chenglizhao Chen, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
Pattern Recognit. | 2 |
| 2015 | Structure-Sensitive Saliency Detection via Multilevel Rank Analysis in Intrinsic Feature SpaceabstractThis paper advocates a novel multiscale, structure-sensitive saliency detection method, which can distinguish multilevel, reliable saliency from various natural pictures in a robust and versatile way. One key challenge for saliency detection is to guarantee the entire salient object being characterized differently from nonsalient background. To tackle this, our strategy is to design a structure-aware descriptor based on the intrinsic biharmonic distance metric. One benefit of introducing this descriptor is its ability to simultaneously integrate local and global structure information, which is extremely valuable for separating the salient object from nonsalient background in a multiscale sense. Upon devising such powerful shape descriptor, the remaining challenge is to capture the saliency to make sure that salient subparts actually stand out among all possible candidates. Toward this goal, we conduct multilevel low-rank and sparse analysis in the intrinsic feature space spanned by the shape descriptors defined on over-segmented super-pixels. Since the low-rank property emphasizes much more on stronger similarities among super-pixels, we naturally obtain a scale space along the rank dimension in this way. Multiscale saliency can be obtained by simply computing differences among the low-rank components across the rank scale. We conduct extensive experiments on some public benchmarks, and make comprehensive, quantitative evaluation between our method and existing state-of-the-art techniques. All the results demonstrate the superiority of our method in accuracy, reliability, robustness, and versatility. Chenglizhao Chen, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
IEEE Trans. Image Process. | 2 |
| 2014 | Hybrid Particle-grid Modeling for Multi-scale Droplet/Spray SimulationabstractAbstract This paper presents a novel hybrid particle‐grid method that tightly couples Lagrangian particle approach with Eulerian grid approach to simulate multi‐scale diffuse materials varying from disperse droplets to dissipating spray and their natural mixture and transition, originated from a violent (high‐speed) liquid stream. Despite the fact that Lagrangian particles are widely employed for representing individual droplets and Eulerian grid‐based method is ideal for volumetric spray modeling, using either one alone has encountered tremendous difficulties when effectively simulating droplet/spray mixture phenomena with high fidelity. To ameliorate, we propose a new hybrid model to tackle such challenges with many novel technical elements. At the geometric level, we employ the particle and density field to represent droplet and spray respectively, modeling their creation from liquid as well as their seamless transition. At the physical level, we introduce a drag force model to couple droplets and spray, and specifically, we employ Eulerian method to model the interaction among droplets and marry it with the widely‐used Lagrangian model. Moreover, we implement our entire hybrid model on CUDA to guarantee the interactive performance for high‐effective physics‐based graphics applications. The comprehensive experiments have shown that our hybrid approach takes advantages of both particle and grid methods, with convincing graphics effects for disperse droplets and spray simulation. Lipeng Yang, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
Comput. Graph. Forum | 2 |
| 2014 | Interactive deformation and cutting simulation directly using patient-specific volumetric imagesabstractABSTRACT This paper systematically advocates an interactive volumetric image manipulation framework, which can enable the rapid deployment and instant utility of patient‐specific medical images in virtual surgery simulation while requiring little user involvement. We seamlessly integrate multiple technical elements to synchronously accommodate physics‐plausible simulation and high‐fidelity anatomical structures visualization. Given a volumetric image, in a user‐transparent way, we build a proxy to represent the geometrical structure and encode its physical state without the need of explicit 3‐D reconstruction. On the basis of the dynamic update of the proxy, we simulate large‐scale deformation, arbitrary cutting, and accompanying collision response driven by a non‐linear finite element method. By resorting to the upsampling of the sparse displacement field resulted from non‐linear finite element simulation, the cut/deformed volumetric image can evolve naturally and serves as a time‐varying 3‐D texture to expedite direct volume rendering. Moreover, our entire framework is built upon CUDA (Beihang University, Beijing, China) and thus can achieve interactive performance even on a commodity laptop. The implementation details, timing statistics, and physical behavior measurements have shown its practicality, efficiency, and robustness. Copyright © 2013 John Wiley & Sons, Ltd. Shuai Li 0001, Qinping Zhao, Shengfa Wang, Aimin Hao, Hong Qin 0001 |
Comput. Animat. Virtual Worlds | 1 |
| 2014 | Real-time physical deformation and cutting of heterogeneous objects via hybrid coupling of meshless approach and finite element methodabstractABSTRACT This paper advocates a method for real‐time physical deformation and arbitrary cutting simulation of heterogeneous objects with multi‐material distribution, whose originality centers on the tight coupling of domain‐specific finite element method (FEM) and material distance‐aware meshless approach in a CUDA‐centric parallel simulation framework. We employ hierarchical hexahedron serving as basic building blocks for accurate material‐aware FEM simulation. Meanwhile, local meshless systems are designed to support cross‐FEM‐domain coupling and material‐sensitive propagation while respecting the regularity of finite elements. Directly benefiting from the structural regularity and uniformity of finite elements, our hybrid solution enables the local stiffness matrix pre‐computation and dynamic assembling, adaptive topological updating and precise cutting reconstruction. Moreover, our mathematically‐rigorous solver guarantees unconditional stableness. Experiments demonstrate the superiorities of our system. Copyright © 2014 John Wiley & Sons, Ltd. Chen Yang 0002, Shuai Li 0001, Lili Wang 0006, Aimin Hao, Hong Qin 0001 |
Comput. Animat. Virtual Worlds | 2 |
| 2013 | Efficient 3D Reconstruction of Vessels from Multi-views of X-Ray AngiographyabstractIn this paper, we present an efficient 3D vessels reconstruction algorithm based on multi-views of X-ray Angiography assisting interventional surgery. First, we extract the vascular-like structures from the image sequences using a geometrical analysis of multi-scale Hessian matrix eigen-system and use the fast marching method to extract the skeleton of the structure, from which we derive the vascular topological configurations. Second, we regard the 3D space as a Markov Random Field and formulate the reconstruction problem as an energy minimization problem with consistent, continuous and topological constraints to coarsely register and reconstruct the 3D vessels. Third, we refine the reconstructed vessels to register and reconstruct the 3D vessels accurately. We demonstrate our system in coronary arteries reconstruction for percutaneous coronary intervention surgery to help doctors learn about the configurations of the coronary arteries of specific patient during operation. We envision that our system will be used for clinic treatment to advance vessel reconstruction for diagnosis and therapy in the near future. Xinglong Liu, Fei Hou 0001, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
CAD/Graphics | 3 |
| 2013 | ROI-Emphasized Volume Visualization Guided by Anisotropic Structure TensorabstractMost of Focus Context visualization methods differentiate the magnification unit only by simply assigning each voxel/cell with an importance value while ignoring the shape content embedded in the volume data. In this paper, we take the volumetric structure information as important cue to facilitate Focus Context visualization, which can homogeneously or non-homogeneously scale the volume data in a structure-sensitive way. Fei Hou 0001, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
CAD/Graphics | 3 |
| 2013 | Multi-scale, multi-level, heterogeneous features extraction and classification of volumetric medical imagesabstractThis paper articulates a novel method for the heterogeneous feature extraction and classification directly on volumetric images, which covers multi-scale point feature, multi-scale surface feature, multi-level curve feature, and blob feature. To tackle the challenge of complex volumetric inner structure and diverse feature forms, our technical solution hinges upon the integrated approach of locally-defined diffusion tensor (DT), DT-based anisotropic convolution kernel (DACK), DACK-based multi-scale analysis, and DT-governed curve feature growing. The extracted structural features can be further semantically classified. At the computational fronts, we design CUDA-based algorithm to conduct parallel computation for time consuming tasks. Various experiments and timing tests demonstrate the effectiveness, robustness, and high performance of our method. Shuai Li 0001, Qinping Zhao, Shengfa Wang, Aimin Hao, Hong Qin 0001 |
ICIP | 1 |
| 2013 | Unsupervised Co-segmentation of Complex Image Set via Bi-harmonic Distance Governed Multi-level Deformable Graph ClusteringabstractDespite the recent success of extensive co-segmentation studies, they still suffer from limitations in accommodating multiple-foreground, large-scale, high-variability image set, as well as their underlying capability for parallel implementation. To improve, this paper proposes a bi-harmonic distance governed flexible method for the robust coherent segmentation of the overlapping/similar contents co-existing in image group, which is independent of supervised learning and any other user-specified prior. The central idea is the novel integration of bi-harmonic distance metric design and multi-level deformable graph generation for multi-level clustering, which gives rise to a host of unique advantages: accommodating multiple-foreground images, respecting both local structures and global semantics of images, being more robust and accurate, and being convenient for parallel acceleration. Critical pipeline of our method involves intrinsic content-coherent measuring, super-pixel assisted bottom-up clustering, and multi-level deformable graph clustering based cross-image optimization. We conduct extensive experiments on the iCoseg benchmark and Oxford flower datasets, and make comprehensive evaluations to demonstrate the superiority of our method via comparison with state-of-the-art methods collected in the MSRC database. Jizhou Ma, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
ISM | 2 |
| 2013 | Robust and high-fidelity guidewire simulation with applications in percutaneous coronary intervention systemabstractReal-time and realistic physics-based simulation of deformable objects is of great value to medical intervention, training, and planning in virtual environments. This paper advocates a virtual-reality (VR) approach to minimally-invasive surgery/therapy (e.g., percutaneous coronary intervention) in medical procedures. In particular, we devise a robust and accurate physics-based modeling and simulation algorithm for the guidewire interaction with blood vessels. We also showcase a VR-based prototype system for simulating percutaneous coronary intervention and mimicing the intervention therapy, which affords the utility of flexible, slender guidewires to advance diagnostic or therapeutic catheters into a patient's vascular anatomy, supporting various real-world interaction tasks. The slender body of guidewires are modeled using the famous Cosserat theory of elastic rods. We derive the equations of motion for guidewires with continuous energies and integrate them with the implicit Euler solver, that guarantees robustness and stability. Our approach's originality is primarily founded upon its power, flexibility, and versatility when interacting with the surrounding environment, including novel strategies in the hybrid of geometry and physics, material variability, dynamic sampling, constraint handling and energy-driven physical responses. Our experimental results have shown that this prototype system is both stable and efficient with real-time performance. In the long run, our algorithm and system are expected to contribute to interactive VR-based procedure training and treatment planning. Yurun Mao, Fei Hou 0001, Shuai Li 0001, Aimin Hao, Mingjing Ai, Hong Qin 0001 |
VRST | 3 |
| 2013 | Hierarchical feature subspace for structure-preserving deformation
Shengfa Wang, Tingbo Hou, Shuai Li 0001, Zhixun Su, Hong Qin 0001 |
Comput. Aided Des. | 3 |
| 2013 | Multi-scale local features based on anisotropic heat diffusion and global eigen-structure
Shuai Li 0001, Hong Qin 0001, Aimin Hao |
Sci. China Inf. Sci. | 1 |
| 2013 | Anisotropic Elliptic PDEs for Feature ClassificationabstractThe extraction and classification of multitype (point, curve, patch) features on manifolds are extremely challenging, due to the lack of rigorous definition for diverse feature forms. This paper seeks a novel solution of multitype features in a mathematically rigorous way and proposes an efficient method for feature classification on manifolds. We tackle this challenge by exploring a quasi-harmonic field (QHF) generated by elliptic PDEs, which is the stable state of heat diffusion governed by anisotropic diffusion tensor. Diffusion tensor locally encodes shape geometry and controls velocity and direction of the diffusion process. The global QHF weaves points into smooth regions separated by ridges and has superior performance in combating noise/holes. Our method's originality is highlighted by the integration of locally defined diffusion tensor and globally defined elliptic PDEs in an anisotropic manner. At the computational front, the heat diffusion PDE becomes a linear system with Dirichlet condition at heat sources (called seeds). Our new algorithms afford automatic seed selection, enhanced by a fast update procedure in a high-dimensional space. By employing diffusion probability, our method can handle both manufactured parts and organic objects. Various experiments demonstrate the flexibility and high performance of our method. Tingbo Hou, Shuai Li 0001, Zhixun Su, Hong Qin 0001, Shengfa Wang |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2012 | A Novel Material-Aware Feature Descriptor for Volumetric Image Registration in Diffusion Tensor Space
Shuai Li 0001, Qinping Zhao, Shengfa Wang, Tingbo Hou, Aimin Hao, Hong Qin 0001 |
ECCV (4) | 1 |
| 2012 | Realtime Two-Way Coupling of Meshless Fluids and Nonlinear FEMabstractAbstract In this paper, we present a novel method to couple Smoothed Particle Hydrodynamics (SPH) and nonlinear FEM to animate the interaction of fluids and deformable solids in real time. To accurately model the coupling, we generate proxy particles over the boundary of deformable solids to facilitate the interaction with fluid particles, and develop an efficient method to distribute the coupling forces of proxy particles to FEM nodal points. Specifically, we employ the Total Lagrangian Explicit Dynamics (TLED) finite element algorithm for nonlinear FEM because of many of its attractive properties such as supporting massive parallelism, avoiding dynamic update of stiffness matrix computation, and efficient solver. Based on a predictor‐corrector scheme for both velocity and position, different normal and tangential conditions can be realized even for shell‐like thin solids. Our coupling method is entirely implemented on modern GPUs using CUDA. We demonstrate the advantage of our two‐way coupling method in computer animation via various virtual scenarios. Lipeng Yang, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
Comput. Graph. Forum | 2 |