Weixin Si

dblp:49/9062 · also Wei-Xin Si · DBLP profile ↗
← Back
43ranked-venue papers
5as first author
34since 2021 · last 2026
0000-0002-3289-9714ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 5 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 16 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TFKAN: Time-frequency KAN for long-term time series forecasting
Xiaoyan Kui, Canwei Liu, Qinsong Li, Zhipeng Hu, Yangyang Shi, Weixin Si, Beiji Zou 0001
Neurocomputing6
2026 Tri-HGNet: A feature-driven dynamic hypergraph framework for medical image segmentation
Xiaoyan Kui, Lingxiao Liu, Qinsong Li, Haonan Yan, Weixin Si, Zuheng Ming, Beiji Zou 0001
Neurocomputing5
2026 A comprehensive survey on magnetic resonance image reconstruction
Xiaoyan Kui, Zijie Fan, Zexin Ji, Qinsong Li, Chengtao Liu, Weixin Si, Beiji Zou 0001
Image Vis. Comput.6
2026 Surgical Data Science in Time-Critical Contexts: A Roadmap Toward Brain-Inspired Computing
Yi Pan 0001, Shihao Zou, Jia-Wen Yang, Weixin Si
J. Comput. Sci. Technol.4
2026 Depth-induced prompt learning for laparoscopic liver landmark detection
abstract
• A new liver landmark detection dataset, L3D-2K, comprising 2,000 keyframes sourced from surgical videos with professional annotations. • A novel deep learning framework D2GPLand+ that utilizes RGB-D information for laparoscopic liver landmark detections. • Proposing the DPE module, which incorporates learnable prompts with contrastive learning to discriminate the geometric features of different landmark categories from depth clues. • Introducing the CUMamba block that concurrently conducts cross-modal interactions on spatial dimension and feature reparameterization on channel dimension for effective RGB-D fusion. • Introducing the AFA scheme to highlight anatomical structures by implicit and explicit edge emphasis and controlling detail levels. Laparoscopic liver surgery presents a highly intricate intraoperative environment with significant liver deformation, posing challenges for surgeons in locating critical liver structures. Anatomical liver landmarks can greatly assist surgeons in spatial perception in laparoscopic scenarios and facilitate preoperative-to-intraoperative registration. To advance research in liver landmark detection, we develop a new dataset called L3D-2K , comprising 2,000 keyframes with expert landmark annotations from surgical videos of 47 patients. Accordingly, we propose a baseline, D 2 GPLand+, which effectively leverages depth modality to boost landmark detection performance. Concretely, we introduce a Depth-aware Prompt Embedding (DPE) scheme, which dynamically extracts class-related global geometric cues with the guidance of self-supervised prompts from the SAM encoder. Further, a Cross-dimension Unified Mamba (CUMamba) block is designed to comprehensively incorporate RGB and depth features with the concurrent spatial and channel scanning mechanism. Besides, we bring out an Anatomical Feature Augmentation (AFA) module that captures anatomical cues and emphasizes key structures by optimizing feature granularity. For benchmarking purposes, we evaluate our method and 17 mainstream detection models on L3D, L3D-2K, and P2ILF datasets. Experimental results demonstrate that D 2 GPLand+ obtains superior performance on all three datasets. Our approach provides surgeons with guiding clues that facilitate surgical operations and decision-making in complex laparoscopic surgery. Our code and dataset are available at https://github.com/cuiruize/D2GPLand-Plus .
Ruize Cui, Weixin Si, Zhixi Li, Kai Wang 0092, Jialun Pei, Pheng-Ann Heng, Harry Qin
Medical Image Anal.2
2026 Integrating frequency-aware mamba with diffusion for 4D volumetric image synthesis
Yangyang Shi, Beiji Zou 0001, Xiaonian Deng, Yucong Zhang, Zehua Liu, Xiaoyan Kui, Weixin Si
Pattern Recognit.8
2026 Spatio-Temporal Representation Decoupling and Enhancement for Federated Instrument Segmentation in Surgical Videos
abstract
Surgical instrument segmentation under Federated Learning (FL) is a promising direction, which enables multiple surgical sites to collaboratively train the model without centralizing datasets. However, there exist very limited FL works in surgical data science, and FL methods for other modalities do not consider inherent characteristics in surgical domain: i) different scenarios show diverse anatomical backgrounds while highly similar instrument representation; ii) there exist surgical simulators which promote large-scale synthetic data generation with minimal efforts. In this paper, we propose a novel Personalized FL scheme, Spatio-Temporal Representation Decoupling and Enhancement (FedST), which wisely leverages surgical domain knowledge during both local-site and global-server training to boost segmentation. Concretely, our model embraces a Representation Separation and Cooperation (RSC) mechanism in local-site training, which decouples the query embedding layer to be trained privately, to encode respective backgrounds. Meanwhile, other parameters are optimized globally to capture the consistent representations of instruments, including the temporal layer to capture similar motion patterns. A textual-guided channel selection is further designed to highlight site-specific features, facilitating model adaptation to each site. Moreover, in global-server training, we propose Synthesis-based Explicit Representation Quantification (SERQ), which defines an explicit representation target based on synthetic data to synchronize the model convergence during fusion for improving model generalization. We construct a new PFL benchmark comprising five surgical sites from public datasets covering four types, with one out-of-federation site. FedST outperforms other state-of-the-art methods on federated sites (1.84% on IoU) and achieves a remarkable improvement on the out-of-federation site (45.29% on IoU). Our source code can be made available at: https://github.com/Meaw0415/FedST.
Xiaoming Qi, Chun-Mei Feng 0001, Jialun Pei, Weixin Si, Yueming Jin
IEEE Trans. Medical Imaging5
2026 Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space
abstract
Motion retrieval is crucial for motion acquisition, offering superior precision, realism, controllability, and editability compared to motion generation. Existing approaches leverage contrastive learning to construct a unified embedding space for motion retrieval from text or visual modality. However, these methods lack a more intuitive and user-friendly interaction mode and often overlook the sequential representation of most modalities for improved retrieval performance. To address these limitations, we propose a framework that aligns four modalities—text, audio, video, and motion—within a fine-grained joint embedding space, incorporating audio for the first time in motion retrieval to enhance user immersion and convenience. This fine-grained space is achieved through a sequence-level contrastive learning approach, which captures critical details across modalities for better alignment. To evaluate our framework, we augment existing text-motion datasets with synthetic but diverse audio recordings, creating two multi-modal motion retrieval datasets. Experimental results demonstrate superior performance over state-of-the-art methods across multiple sub-tasks, including an 10.16% improvement in R@10 for text-to-motion retrieval and a 25.43% improvement in R@1 for video-to-motion retrieval on the HumanML3D dataset. Furthermore, our results show that our 4-modal framework significantly outperforms its 3-modal counterpart, underscoring the potential of multi-modal motion retrieval for advancing motion acquisition.
Shiyao Yu, Zi-An Wang, Kangning Yin, Zheng Tian 0002, Weixin Si, Shihao Zou
IEEE Trans. Multim.6
2025 Surgical Workflow Recognition and Blocking Effectiveness Detection in Laparoscopic Liver Resection with Pringle Maneuver
abstract
Pringle maneuver (PM) in laparoscopic liver resection aims to reduce blood loss and provide a clear surgical view by intermittently blocking blood inflow of the liver, whereas prolonged PM may cause ischemic injury. To comprehensively monitor this surgical procedure and provide timely warnings of ineffective and prolonged blocking, we suggest two complementary AI-assisted surgical monitoring tasks: workflow recognition and blocking effectiveness detection in liver resections. The former presents challenges in real-time capturing of short-term PM, while the latter involves the intraoperative discrimination of long-term liver ischemia states. To address these challenges, we meticulously collect a novel dataset, called PmLR50, consisting of 25,037 video frames covering various surgical phases from 50 laparoscopic liver resection procedures. Additionally, we develop an online baseline for PmLR50, termed PmNet. This model embraces Masked Temporal Encoding (MTE) and Compressed Sequence Modeling (CSM) for efficient short-term and long-term temporal information modeling, and embeds Contrastive Prototype Separation (CPS) to enhance action discrimination between similar intraoperative operations. Experimental results demonstrate that PmNet outperforms existing state-of-the-art surgical workflow recognition methods on the PmLR50 benchmark. Our research offers potential clinical applications for the laparoscopic liver surgery community.
Diandian Guo, Weixin Si, Zhixi Li, Jialun Pei, Pheng-Ann Heng
AAAI2
2025 Versatile and Efficient Medical Image Super-Resolution Via Frequency-Gated Mamba
abstract
Medical image super-resolution (SR) is essential for enhancing diagnostic accuracy while reducing acquisition cost and scanning time. However, modeling both long-range anatomical structures and fine-grained frequency details with low computational overhead remains challenging. We propose FGMamba, a novel frequency-aware gated state-space model that unifies global dependency modeling and fine-detail enhancement into a lightweight architecture. Our method introduces two key innovations: a Gated Attention-enhanced State-Space Module (GASM) that integrates efficient state-space modeling with dualbranch spatial and channel attention, and a Pyramid Frequency Fusion Module (PFFM) that captures high-frequency details across multiple resolutions via FFT-guided fusion. Extensive evaluations across five medical imaging modalities (Ultrasound, OCT, MRI, CT, and Endoscopic) demonstrate that FGMamba achieves superior PSNR/SSIM while maintaining a compact parameter footprint (<0.75M), outperforming CNN-based and Transformerbased SOTAs. Our results validate the effectiveness of frequencyaware state-space modeling for scalable and accurate medical image enhancement. Source code and dataset will be made publicly available.
Wenfeng Huang, Xiangyun Liao, Wei Cao 0008, Wenjing Jia, Weixin Si
BIBM5
2025 Fast Intra-Operative Angiographic Parametric Imaging for Surgical Outcome Prediction of Interventional Cerebral Aneurysms Treatment
abstract
Endovascular intervention of cerebral aneurysms remains challenging due to the lack of intra-operative indicators that reflect hemodynamic alterations following stent implantation. This limitation makes it difficult to predict surgical outcomes and increases dependence on surgeon experience. In this study, we propose a fast and practical approach for converting intraoperative digital subtraction angiography (DSA) images into angiographic parametric imaging (API), enabling timely visualization of cerebral blood flow during surgery. We develop a stylusbased regions of interest (ROI) analysis interface. The proposed method has been validated against multiple reference standards, including CT perfusion (CTP), computational fluid dynamics (CFD), and 4D flow MRI, demonstrating strong consistency in hemodynamic quantification. Notably, unlike traditional methods that often take several hours, the proposed method has an average processing time of only 4.47 minutes, highlighting its suitability for intra-operative use. Furthermore, we designed a qualitative assessment strategy based on pre- and intra-operative ROI comparison to predict complication risk. The proposed method achieved an 87.5 % accuracy in predicting post-operative complications across 40 clinical cases, demonstrating particularly high sensitivity in detecting ischemia risk. These results demonstrate that the proposed method provides a reliable, efficient, and interpretable solution for intra-operative hemodynamic assessment, with strong potential for integration into clinical workflows. Code and data are available at: https://github.com/LZH970328/API.git.
Zehua Liu, Jianping Lv, Weixin Si
BIBM6
2025 TempDiffReg: Temporal Diffusion Model for Non-Rigid 2D-3D Vascular Registration
abstract
Transarterial chemoembolization (TACE) is a preferred treatment option for hepatocellular carcinoma and other liver malignancies, yet it remains a highly challenging procedure due to complex intra-operative vascular navigation and anatomical variability. Accurate and robust 2D-3D vessel registration is essential to guide microcatheter and instruments during TACE, enabling precise localization of vascular structures and optimal therapeutic targeting. To tackle this issue, we develop a coarse-to-fine registration strategy. First, we introduce a global alignment module, structure-aware perspective n-point (SA-PnP), to establish correspondence between 2D and 3D vessel structures. Second, we propose TempDiffReg, a temporal diffusion model that performs vessel deformation iteratively by leveraging temporal context to capture complex anatomical variations and local structural changes. We collected data from 23 patients and constructed 626 paired multi-frame samples for comprehensive evaluation. Experimental results demonstrate that the proposed method consistently outperforms state-of-the-art (SOTA) methods in both accuracy and anatomical plausibility. Specifically, our method achieves a mean squared error (MSE) of 0.63 mm and a mean absolute error (MAE) of 0.51 mm in registration accuracy, representing$66.7\%$lower MSE and$17.7\%$lower MAE compared to the most competitive existing approaches. It has the potential to assist less-experienced clinicians in safely and efficiently performing complex TACE procedures, ultimately enhancing both surgical outcomes and patient care. Code and data are available at: https://github.com/LZH970328/TempDiffReg.git
Zehua Liu, Shihao Zou, Jincai Huang 0003, Weixin Si
BIBM6
2025 Distance-Aware and Knowledge-Driven Vision Mamba U-Net for Radiotherapy Dose Prediction
abstract
Dose planning is essential in radiotherapy for cancer patients, yet current practice relies on iterative manual optimization, underscoring the need for automated prediction. Existing deep learning approaches remain limited because they often ignore the 3D spatial relationships between tumors and surrounding organs at risk (OARs), and clinical priors on safe dose thresholds. To overcome these limitations, we propose DKVMU-Net, a distance-aware and knowledge-driven Vision Mamba U-Net for automated dose prediction. Our framework incorporates Vision Mamba blocks to capture global, long-range dependencies from CT scans and OAR signed distance field (SDF) maps, which naturally encode spatial information. Additionally, we introduce a deformable dynamic feature enhancement module (DDFEM) for texture refinement, followed by a linear crossattention fusion module to improve cross-modality integration. A customized loss function is also designed to incorporate prior knowledge of OAR dose constraints, ensuring optimal target coverage and OAR protection. To alleviate the scarcity of doseplanning datasets, we collect an in-house radiotherapy lung cancer dataset (RLCD), consisting of CT volumes, OAR masks, and corresponding SDF maps from 116 patients. We evaluate our DKVMU-Net on both the in-house dataset and public available OpenKBP dataset. Compared with the sate-of-the-art method, our approach achieves an 11.6 % improvement in dose score (1.641 vs. 1.857) and 26.3 % in DVH score (6.481 vs. 8.799) on RLCD, and a 7.8 % improvement in dose score (2.421 vs. 2.626) and 13.9 % in DVH score (1.057 vs. 1.227) on OpenKBP. These results demonstrate the robustness and effectiveness of our approach.
Yangyang Shi, Xiaoyan Kui, Yucong Zhang, Shihao Zou, Zuheng Ming, Weixin Si, Azeddine Beghdadi, Beiji Zou 0001
BIBM6
2025 From Global to Local: Mamba-Based Hierarchical Registration for Respiratory Lung Deformation
abstract
Deformable image registration is essential in medical applications, as accurately estimating organ displacements across respiratory phases enables precise radiation dose planning in dynamic environments, mitigates damage to organs at risk (OARs), and thus improves patients' health-related quality of life. Although current learning-based methods have achieved impressive performance in small deformation registration, challenges remain due to their limited ability to capture large deformations occurring during respiration. To address this issue, we propose a novel Mamba-based hierarchical registration framework that effectively extracts both global and local features for accurate deformation prediction. Specifically, given a pair of source and target 3DCT volumes, we incorporate a foundation model pretrained on medical image registration tasks to enhance alignment accuracy. We further propose a directional-deformable Mamba scheme to facilitate global context extraction and local motion awareness. The directional Mamba component scans input features from multiple orientations to achieve broad contextual perception, while the deformable Mamba module employs adaptive directional scanning strategies to capture dynamic local variations. To overcome the scarcity of annotated respiratory data, we also collect a new respiratory lung cancer dataset comprising 100 annotated phases from 20 patients. Experimental results on our in-house dataset demonstrate that our method outperforms state-of-the-art approaches, achieving a 1.3 % improvement in overall Dice accuracy and a 1.6 dB increase in PSNR, underscoring its strong potential for clinical deployment. Code and test data are available at: https://github.com/yangyangshi806/Mamba_based_Registration.
Yangyang Shi, Yucong Zhang, Beiji Zou 0001, Xiaoyan Kui, Zexin Ji, Zuheng Ming, Azeddine Beghdadi, Weixin Si
BIBM8
2025 Medical Open Set Recognition via Intra-Class Clustering
abstract
In computational medical imaging, model's ability to identify whether a sample is from an unseen semantic category is critical in clinical deployments. However, conventional medical image recognition usually assumes a closed-set setting where all queries in testing are from pre-defined training categories and overlooks the fact that in practice it is possible to have queries from unknown categories such as unknown or unseen tissue. In this study, we particularly tackle this thorny challenge, namely Medical Open Set Recognition (MOSR), and explore it on medical image classification and diagnosis. The biggest challenge with this issue lies in deep model's overconfidence due to relatively large intra-class variance, which leads to incorrectly assigning an unknown sample to a known class with a high confidence level. To address this problem, we introduce intra-class clustering, which divides the samples assigned to each class into several low-variance sub-clusters. In addition, we propose to divide the samples uniformly to each cluster by optimal transport to achieve online clustering. Extensive experiments on 6 public medical imaging datasets demonstrate that a classification model trained with the proposed intra-class clustering dramatically alleviate the overconfidence problem with competitive accuracy and thus effective for improving MOSR performance. Our benchmarks and code will be publicly released when published.
Hanqiu Deng, Shihao Zou, Xiangyun Liao, Weixin Si
CW4
2025 Cerebrovascular Diseases Screening from Color Fundus Photography via Cross-View Fusion and Graph-Based Discrimination
Congyu Tian, Shihao Zou, Xiangyun Liao, Chubin Ou, Jianping Lv, Shanshan Wang 0002, Weixin Si
MICCAI (12)8
2025 Cost-effective Tangible Rehearsal Interface for Microsurgical Clipping of Intracranial Aneurysm
abstract
Microsurgical clipping (MC) is widely used for the treatment of intracranial aneurysms (IA). However, it is a high risk for neurosurgeons to perform this operation due to the complex intracranial anatomy, limited microscopic view, and confined operational space. To tackle the above issues, meticulous preoperative rehearsal is in urgent need to improve the neurosurgeons’ perception of patient-specific anatomy while designing the optimal surgical plan, which can greatly enhance the patients’ safety. Most existing MC simulators only support limited interaction methods, such as haptic devices, leading to substantial differences between the training experience and actual surgical situations. To this end, we present a mixed reality (MR) interface for IA MC rehearsal, which can provide neurosurgeons with a more immersive experience. Firstly, considering the labor-intensive labeling cost of reconstructing a 3D full-brain vascular model from CTA images, the simulator employs a cost-effective specific-to-general vessel stitching technique to generate 3D lesion-specific IA geometric models, which replaces normal vascular segments in a standard brain with a personalized-specific operating region by a scaling-constrained iterative closest point (ICP) algorithm. Besides, we design a marker-based tracking method allowing accessible natural human-computer interaction using real surgical instruments which can fuse virtual anatomy and real operation environments, enhancing the users’ spatial perception and tangible stimuli. Additionally, to ensure simulation stability while providing visually plausible clipping operations, we develop a collision distance constrained position-based dynamics (PBD) method with low-resolution sampling particles to simulate the deformation of aneurysm vessels. Quantitative experiments demonstrate the accuracy of our surgical instruments tracking and vessel deformation simulation, which can also fulfill the real-time performance of MC rehearsal. User study indicates that virtual rehearsals significantly improve spatial awareness and dexterity in handling aneurysms, and have the great potential to be applied in practical applications.
Wei Cao 0008, Zehua Liu, Jianping Lv, Weixin Si
VR7
2025 Boundary-aware dynamic re-weighting for semi-supervised medial image segmentation
Weili Jiang, Xifei Wei, Gwenolé Quellec, Weixin Si, Chubin Ou
Expert Syst. Appl.7
2025 Self-prompt contextual learning with AxialMamba for multi-label segmentation in carotid ultrasound
Congyu Tian, Xiangyun Liao, Jianping Lv, Weixin Si
Expert Syst. Appl.6
2025 WinGraphUNet: Advanced windowed graph modeling with remixed contextual learning for efficient medical image segmentation
Xiaoyan Kui, Haonan Yan, Qinsong Li, Lingxiao Liu, Weixin Si, Wei Liang 0005, Beiji Zou 0001
Knowl. Based Syst.5
2025 A new dataset and versatile multi-task surgical workflow analysis framework for thoracoscopic mitral valvuloplasty
Meng Lan, Weixin Si, Xinjian Yan, Xiaomeng Li 0001
Medical Image Anal.2
2025 Highly Efficient 3D Human Pose Tracking From Events With Spiking Spatiotemporal Transformer
abstract
Event camera, as an asynchronous vision sensor capturing scene dynamics, presents new opportunities for highly efficient 3D human pose tracking. Existing approaches typically adopt modern-day Artificial Neural Networks (ANNs), such as CNNs or Transformer, where sparse events are converted into dense images or paired with additional gray-scale images as input. Such practices, however, ignore the inherent sparsity of events, resulting in redundant computations, increased energy consumption, and potentially degraded performance. Motivated by these observations, we introduce the first sparse Spiking Neural Networks (SNNs) framework for 3D human pose tracking based solely on events. Our approach eliminates the need to convert sparse data to dense formats or incorporate additional images, thereby fully exploiting the innate sparsity of input events. Central to our framework is a novel Spiking Spatio-temporal Transformer, which enables bi-directional spatio-temporal fusion of spike pose features and provides a guaranteed similarity measurement between binary spike features in spiking attention. Moreover, we have constructed a largescale synthetic dataset, SynEventHPD, that features a broad and diverse set of 3D human motions, as well as much longer hours of event streams. Empirical experiments demonstrate the superiority of our approach over existing state-of-the-art (SOTA) ANN-based methods, requiring only 19.1% FLOPs and 3.6% energy cost. Furthermore, our approach outperforms existing SNN-based benchmarks in this task, highlighting the effectiveness of our proposed SNN framework. The dataset will be released upon acceptance, and code can be found at https://github.com/JimmyZou/HumanPoseTracking_SNN.
Shihao Zou, Yuxuan Mu, Wei Ji 0011, Zi-An Wang, Xinxin Zuo, Sen Wang 0003, Weixin Si, Li Cheng 0001
IEEE Trans. Circuits Syst. Video Technol.7
2025 Delving Into Quaternion Wavelet Transformer for Facial Expression Recognition in the Wild
abstract
The Facial Expression Recognition (FER) technique has increasingly matured over time. However, recognizing facial expressions in wild environments poses great challenges in achieving promising performance. The main obstacles arise from various factors, such as illumination changes, head pose variations, and occlusions. To overcome interferences from external environments and improve recognition accuracy, we propose a novel Quaternion Wavelet TRansformer (QWTR) model for FER in the wild. Specifically, we present a Quaternion Value Transformer (QVT) network that combines quaternion multi-head attention with quaternion CNN to capture emotional cues from global and local perception. To preserve the color structure while enhancing image contrast and brightness, we introduce a Quaternion Histogram Equalization (QHE) representation to transform color images into quaternion matrices representation. After that, to alleviate the impact of head pose and occlusion together with feature redundancy, a Quaternion Wavelet Feature Selection (QWFS) scheme is designed to decompose quaternion features and select the most correlated signals. Extensive experiments have been conducted on four in-the-wild FER datasets and several specific FER benchmarks under various conditions. The qualitative and quantitative results demonstrate thatQWTRoutperforms other state-of-the-art methods in FER benchmarks, e.g., 68.37% vs. 66.31% accuracy on the AffectNet dataset.
Yu Zhou 0049, Jialun Pei, Weixin Si, Harry Qin, Pheng-Ann Heng
IEEE Trans. Multim.3
2025 Generating High-Fidelity Clothed Human Dynamics with Temporal Diffusion
abstract
Clothed human modeling plays a crucial role in multimedia research, with applications spanning virtual reality, gaming, and fashion design. The goal is to learn clothed human dynamics from observations and then generate humans with high-fidelity clothing details for motion animation. Despite tremendous advancements in clothing shape analysis by existing approaches, the community still faces challenges in generating convincing visual effects of cloth dynamics, maintaining temporally smooth clothing details, and handling diverse clothing patterns. To address these challenges, we introduce ClothDiffuse, a temporal diffusion model that seamlessly integrates three key components into this task—temporal dynamics modeling, iterative refinement, and diversified generation. Our approach begins by using an encoder to extract high-level temporal features from input human body motions. These features are combined with a learnable pixel-aligned garment feature, serving as prior conditions for the shape decoder. The decoder then iteratively denoise Gaussian noise to produce clothing deformations over time on the input unclothed human bodies. To ensure that the results align with observations and adhere to physical plausibility for clothing shape inference, we propose two physics-inspired loss functions that preserve the intra-frame distances and inter-frame forces of clothing points. Additionally, the stochastic nature of the denoising process allows for the generation of diverse and plausible clothing shapes. Experiments show that our approach outperforms state-of-the-art methods in chamfer distance and visual effects, particularly for loose clothing such as dresses and skirts. Furthermore, our approach effectively adapts to out-of-domain clothing types and generates realistic clothes dynamics.
Shihao Zou, Yuanlu Xu, Nikolaos Sarafianos, Federica Bogo, Tony Tung, Weixin Si, Li Cheng 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2025 Semi-supervised intracranial aneurysm segmentation via reliable weight selection
Wei Cao 0008, Jianping Lv, Weixin Si
Vis. Comput.5
2025 MF-SAM: enhancing multi-modal fusion with Mamba in SAM-Med3D for GPi segmentation
Doudou Zhang, Junchi Ma, Linxia Xiao, Xiangyun Liao, Weixin Si
Vis. Comput.7
2024 Depth-Driven Geometric Prompt Learning for Laparoscopic Liver Landmark Detection
Jialun Pei, Ruize Cui, Yaoqian Li, Weixin Si, Harry Qin, Pheng-Ann Heng
MICCAI (6)4
2024 Versatile latent distribution-preserving tabular data synthesis-based endovascular treatment selection for intracranial aneurysm
Qian Yang 0005, Chubin Ou, Kang Li 0007, Yucong Zhang, Xiangyun Liao, Jianping Lv, Weixin Si
Expert Syst. Appl.8
2024 Synthesizing Feature-Aligned and Category-Aware Electronic Medical Records for Intracranial Aneurysm Rupture Prediction
abstract
Rupture prediction is crucial for precise treatment and follow-up management of patients with intracranial aneurysms (IAs). Considerable machine learning (ML) methods have been proposed to improve rupture prediction by leveraging electronic medical records (EMRs), however, data scarcity and category imbalance strongly influence performance. Thus, we propose a novel data synthesis method i.e., Transformer-based conditional GAN (TransCGAN), to synthesize highly authentic and category-aware EMRs to address above challenges. Specifically, we first align feature-wise context relationship and distribution between synthetic and original data to enhance synthetic data quality. To achieve this, we first integrate the Transformer structure into GAN to match the contextual relationship by processing the long-range dependencies among clinical factors and introduce a statistical loss to maintain distributional consistency by constraining the mean and variance of the synthesis features. Additionally, a conditional module is designed to assign the category of the synthesis data, thereby addressing the challenge of category imbalance. Subsequently, the synthetic data are merged with the original data to form a large-scale and category-balanced training dataset for IAs rupture prediction. Experimental results show that using TransCGAN's synthetic data enhances classifier performance, achieving AUC of 0.89 and outperforming state-of-the-art resampling methods by 5-33 in F1 score.
Qian Yang 0005, Caizi Li, Chubin Ou, Kang Li 0007, Xiangyun Liao, Chuanzhi Duan, Lequan Yu, Weixin Si
IEEE J. Biomed. Health Informatics8
2023 Semi-Supervised Intracranial Aneurysm Segmentation from CTA Images via Weight-Perceptual Self-Ensembling Model
Caizi Li, Ruiqiang Liu, Huan-Xin Zhong, Jun-Ming Fan, Weixin Si, Pheng-Ann Heng
J. Comput. Sci. Technol.5
2022 Amplitude-frequency-aware deep fusion network for optimal contact selection on STN-DBS electrodes
Linxia Xiao, Caizi Li, Yanjiang Wang 0001, Weixin Si, Doudou Zhang, Xiaodong Cai, Pheng-Ann Heng
Sci. China Inf. Sci.4
2022 Synergistic Digital Twin and Holographic Augmented-Reality-Guided Percutaneous Puncture of Respiratory Liver Tumor
abstract
Thermal ablation is an exciting new minimally invasive treatment that destroys liver tumors without removing them. It uses image guidance to place a needle through the skin into a liver tumor, which is highly dependent on surgeons’ experience. With the development of digital medicine, augmented reality (AR) has become a more intuitive and safer way to achieve real-time navigation. However, the technology is still in its infancy due to its limited accuracy and real-time performance. To address these problems, we syncretized the holographic AR with the digital twin technique to track the dynamic surgical scene and provide the 3-D navigation of heterogeneous target regions via internal motion prediction. To tackle the dilemma of real-time performance and precise internal motion estimation, a dynamic adaptation scheme is proposed to compensate for the time cost induced by the external/internal correlation model and data transmission. We carried out a series of experiments to validate our methods. With the proposed external/internal correlation model, the average estimation errors of the tumor and vessels are 2.18 and 2.79 mm, respectively. Besides, we performedin vivoexperiments on two beagle dogs with an artificial lesion in their liver, respectively, and the puncture accuracy of our method are 2.5 and 2.17 mm. The results show that on one hand, our method can fulfill the real-time requirement of AR via using the intraoperative data, which is also more precise than that with preoperative data. On the other hand, our method can provide more 3-D information for surgeons, such as vessels, which can well ensure the safety of operation.
Yangyang Shi, Xuesong Deng, Yuqi Tong, Ruotong Li, Lijie Ren, Weixin Si
IEEE Trans. Hum. Mach. Syst.7
2021 A global benchmark of algorithms for segmenting the left atrium from late gadolinium-enhanced cardiac magnetic resonance imaging
Zhaohan Xiong, Qing Xia 0002, Cheng Bian, Yefeng Zheng 0001, Sulaiman Vesal, Nishant Ravikumar, Andreas K. Maier, Xin Yang 0009, Pheng-Ann Heng, Dong Ni 0001, Caizi Li, Qianqian Tong 0001, Weixin Si, Élodie Puybareau, Younes Khoudli, Thierry Géraud, Jichao Zhao
Medical Image Anal.15
2021 Self-Ensembling Co-Training Framework for Semi-Supervised COVID-19 CT Segmentation
abstract
The coronavirus disease 2019 (COVID-19) has become a severe worldwide health emergency and is spreading at a rapid rate. Segmentation of COVID lesions from computed tomography (CT) scans is of great importance for supervising disease progression and further clinical treatment. As labeling COVID-19 CT scans is labor-intensive and time-consuming, it is essential to develop a segmentation method based on limited labeled data to conduct this task. In this paper, we propose a self-ensembled co-training framework, which is trained by limited labeled data and large-scale unlabeled data, to automatically extract COVID lesions from CT scans. Specifically, to enrich the diversity of unsupervised information, we build a co-training framework consisting of two collaborative models, in which the two models teach each other during training by using their respective predicted pseudo-labels of unlabeled data. Moreover, to alleviate the adverse impacts of noisy pseudo-labels for each model, we propose a self-ensembling strategy to perform consistency regularization for the up-to-date predictions of unlabeled data, in which the predictions of unlabeled data are gradually ensembled via moving average at the end of every training epoch. We evaluate our framework on a COVID-19 dataset containing 103 CT scans. Experimental results show that our proposed method achieves better performance in the case of only 4 labeled CT scans compared to the state-of-the-art semi-supervised segmentation networks.
Caizi Li, Qi Dou 0001, Fan Lin, Kebao Zhang, Zuxin Feng, Weixin Si, Xuesong Deng, Pheng-Ann Heng
IEEE J. Biomed. Health Informatics7
2019 Versatile numerical fractures removal for SPH-based free surface liquids
Weixin Si, Xiangyun Liao, Yinling Qian, Qiong Wang 0001, Pheng-Ann Heng
Comput. Graph.1
2019 Mixed reality based respiratory liver tumor puncture navigation
abstract
This paper presents a novel mixed reality based navigation system for accurate respiratory liver tumor punctures in radiofrequency ablation (RFA). Our system contains an optical see-through head-mounted display device (OST-HMD), Microsoft HoloLens for perfectly overlaying the virtual information on the patient, and a optical tracking system NDI Polaris for calibrating the surgical utilities in the surgical scene. Compared with traditional navigation method with CT, our system aligns the virtual guidance information and real patient and real-timely updates the view of virtual guidance via a position tracking system. In addition, to alleviate the difficulty during needle placement induced by respiratory motion, we reconstruct the patient-specific respiratory liver motion through statistical motion model to assist doctors precisely puncture liver tumors. The proposed system has been experimentally validated on vivo pigs with an accurate real-time registration approximately 5-mm mean FRE and TRE, which has the potential to be applied in clinical RFA guidance.
Ruotong Li, Weixin Si, Xiangyun Liao, Qiong Wang 0001, Reinhard Klein, Pheng-Ann Heng
Comput. Vis. Media2
2018 Augmented Reality-Based Personalized Virtual Operative Anatomy for Neurosurgical Guidance and Training
abstract
This paper presents a novel augmented reality (AR) interactive environment for neurosurgical training. Comparing with traditional virtual reality based neurosurgical simulator, our system provides a more natural and intuitive fashion for surgeons. To achieve holographic visualization of virtual brain on 3D-printed skull (workspace), the first step is to reconstruct the personalized anatomy structure from segmented MR imaging. Then, tailored to the computational power of HoloLens, we employ the mass-spring method to model the mechanical response of brain. After that, a precise registration method is employed to map the virtual-real spatial information, which can overlay the virtual operative brain on workspace. In addition, bimanual haptic interface is also integrated into our simulator, which is more similar with real neurosurgery. In experiments, we conduct accuracy validation on our registration method, as well as the validity test on the developed simulators. The results demonstrate that our simulator can provide high-accuracy augmented visualization effects and deep immersion for novice surgeons.
Weixin Si, Xianavun Liao, Qiong Wang 0001, Pheng-Ann Heng
VR1
2018 Thin-Feature-Aware Transport-Velocity Formulation for SPH-Based Liquid Animation
abstract
Realistic liquid animations with thin sheets or streams are crucial for creating fluid effects in digital media. However, it is challenging to simulate these appealing thin sheets or streams in the framework of smoothed particle hydrodynamics (SPH). The underlying reason for this challenge mainly lies in the inherent numerical instability of SPH due to inconsistent kernel interpolation, which is caused by the incomplete kernel support on the free surface and the particles' disorder dispersion within the simulation domain. To address this challenge, we propose a novel and effective approach to ensure the consistency of kernel interpolation at both internal flow and the free surface during the simulation such that these thin features can always be well maintained. First, we introduce a transport-velocity formulation to alleviate the disorder dispersion in the liquid domain. However, this formulation can only work in the internal flow, and it fails at the free surface because it cannot accurately estimate the density of particles there. To this end, we propose adaptively correcting the underestimated density caused by the incomplete kernel support of free-surface particles, which are identified by a geometry-aware anisotropic kernel, to counteract the inconsistent interpolation on the free surface. Then, we propose a novel scheme to further filter the background pressure to enhance the interactions between the internal flow and the free surface, as well as liquid and solid, such that the thin features generated from such interactions can be realistically simulated. The proposed approach can also achieve anticlumping and regularization effects in the entire simulation domain and, hence, further enhance the thin features in liquids. We evaluate our method on a variety of benchmark examples, and the results demonstrate that our method can achieve more appealing visual effects than state-of-the-art methods by realistically simulating more vivid thin features.
Weixin Si, Harry Qin, Zhuchao Chen, Xiangyun Liao, Qiong Wang 0001, Pheng-Ann Heng
IEEE Trans. Multim.1
2018 Animating Wall-Bounded Turbulent Smoke via Filament-Mesh Particle-Particle Method
abstract
Turbulent vortices in smoke flows are crucial for a visually interesting appearance. Unfortunately, it is challenging to efficiently simulate these appealing effects in the framework of vortex filament methods. The vortex filaments in grids scheme allows to efficiently generate turbulent smoke with macroscopic vortical structures, but suffers from the projection-related dissipation, and thus the small-scale vortical structures under grid resolution are hard to capture. In addition, this scheme cannot be applied in wall-bounded turbulent smoke simulation, which requires efficiently handling smoke-obstacle interaction and creating vorticity at the obstacle boundary. To tackle above issues, we propose an effective filament-mesh particle-particle (FMPP) method for fast wall-bounded turbulent smoke simulation with ample details. The Filament-Mesh component approximates the smooth long-range interactions by splatting vortex filaments on grid, solving the Poisson problem with a fast solver, and then interpolating back to smoke particles. The Particle-Particle component introduces smoothed particle hydrodynamics (SPH) turbulence model for particles in the same grid, where interactions between particles cannot be properly captured under grid resolution. Then, we sample the surface of obstacles with boundary particles, allowing the interaction between smoke and obstacle being treated as pressure forces in SPH. Besides, the vortex formation region is defined at the back of obstacles, providing smoke particles flowing by the separation particles with a vorticity force to simulate the subsequent vortex shedding phenomenon. The proposed approach can synthesize the lost small-scale vortical structures and also achieve the smoke-obstacle interaction with vortex shedding at obstacle boundaries in a lightweight manner. The experimental results demonstrate that our FMPP method can achieve more appealing visual effects than vortex filaments in grids scheme by efficiently simulating more vivid thin turbulent features.
Xiangyun Liao, Weixin Si, Hanqiu Sun, Harry Qin, Qiong Wang 0001, Pheng-Ann Heng
IEEE Trans. Vis. Comput. Graph.2
2017 Patch green coordinates based interactive embedded deformable model
abstract
Virtual surgery is a serious game which provides an opportunity to acquire cognitive and technical surgical skills via virtual surgical training and planning. However, interactively and realistically manipulating the human organ and simulating its motion under interaction is still a challenging task in this field. The underlying reason for this issue is the conflict requirements for physical constraints with high fidelity and real-time performance. To achieve realistic simulation of human organ motion with volume conservation, smooth interpolation under large deformation and precise frictional contact mechanics of global behavior in surgical scenario. This paper presents a novel and effective patch Green coordinates based interpolation for embedded deformable model to achieve the volume-preserving and smooth interpolation effects. Besides, we resolve the frictional contact mechanics for embedded deformable model, and further provide the precise boundary conditions for mechanical solver. In addition, our embedded deformable model is based on the total lagrangian explicit dynamics (TLED) finite element method (FEM) solver, which can well handle the large biological tissue deformation with both nonlinear geometric and material properties. In real compression experiments, our method can achieve liver deformation with average accuracy of 3.02 mm. Besides, the experimental results demonstrate that our method can also achieve smoother interpolation and volume-preserving effects than original embedded deformable model, and allows complex and accurate organ motion with mechanical interactions in virtual surgery.
Weixin Si, Xiangyun Liao, Qiong Wang 0001, Harry Qin, Pheng-Ann Heng
MIG1
2017 Filament-based realistic turbulent wake synthesis
abstract
Abstract Turbulent wake is crucial for the visually appealing effects of liquid. Unfortunately, it is challenging to realistically simulate this phenomenon with ring‐shaped vortical structures. To tackle this issue, we propose a filament‐based turbulent wake synthesis method for realistically simulating the turbulent wake with ring‐shaped vortical structures. The filaments are sampled at the separation points on the obstacle surface and emitted into the liquid flow to generate structured turbulent wake. Besides, the surface tension model is incorporated to generate natural turbulent wake diffusion visual effects in liquid by the anticurvature effects. The proposed approach can realistically and effectively synthesize the turbulent wake with ring‐shaped vortical structures and make it diffuse naturally. The experimental results demonstrate that our method outperforms than the vortex particle‐based method in synthesizing appealing turbulent wake.
Xiangyun Liao, Weixin Si, Qiong Wang 0001, Pheng-Ann Heng
Comput. Animat. Virtual Worlds2
2012 Parallel computing of 3D smoking simulation based on OpenCL heterogeneous platform
Weixin Si, Xiangyun Liao, Zhaoliang Duan, Yihua Ding, Jianhui Zhao 0001
J. Supercomput.2
2011 3D soft tissue warping dynamics simulation based on force asynchronous diffusion model
abstract
Abstract Soft tissue warping is one of the key technologies of medical dynamics simulation, such as surgical simulation, image guided surgery. In this paper, we present a novel simulation method which is stable and fast like linear models for soft tissue warping simulation. This method performs on the irregular mesh models, and it is able to represent the visual properties of physical processes with low computational complexity using the Force Asynchronous Diffusion Model (FADM) proposed in this paper. It contains three parts: model preprocessing, collision detection and simulation model solution. In model preprocessing, we establish three models based on the triangular mesh: the geometrical model, the physical model and the transitional model. A two‐level collision detection algorithm is presented based on the three models. At every time step of the simulation model solution, to more accurately reflect the internal physical properties of the soft tissue, we divide the springs in physical model into three kinds: tissue springs, connection springs and virtual springs; and we propose the asynchronous regions and active regions to simplify the computing process according to the realistic physical warping. Experimental results show the FAMD can achieve good warping effects on speed and realism. Copyright © 2011 John Wiley & Sons, Ltd.
Weixin Si, Xiangyun Liao, Zhaoliang Duan, Yihua Ding, Jianhui Zhao 0001
Comput. Animat. Virtual Worlds1