Yankun Cao

dblp:192/4192 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0003-0636-5417ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 GraphSP: Graph-Based Learning for Ultrasound Sequential Image Classification of Single-Patient with Limited Samples and Sparse Annotations
Aijing Feng, Baoning Liu, Anyu Li, Lan Ye, Abir Aal Issa, Tamer Abukhalil, Zhi Liu 0004, Yankun Cao
ICIC (21)9
2025 TF-Fusion: Time-Frequency Feature Fusion for High-Quality Single-Angle Ultrasound Imaging
abstract
Single-angle plane wave (SAPW) ultrasound imaging has gained significant attention in ultrafast and wearable ultrasound systems due to its high frame rate and hardware efficiency. However, the lack of transmit diversity leads to severe degradation in image quality, limiting its clinical utility. In contrast, Coherent Plane Wave Compounding (CPWC) improves spatial resolution and contrast by aggregating data from multiple transmission angles, but at the expense of a substantially reduced frame rate. To overcome this trade-off, we propose TF-Fusion, a novel Time-Frequency Domain Feature Fusion framework that enhances SAPW imaging by jointly exploiting both time- and frequency-domain features extracted from raw in-phase/quadrature (IQ) data. Specifically, our model utilizes a dual-branch encoder to learn compact and complementary representations from each domain, followed by a learnable fusion module that adaptively integrates the multi-domain information. This design facilitates effective noise and artifact suppression while preserving fine anatomical structures. We evaluate TF-Fusion on two public benchmarks—PICMUS and CUBDL—and demonstrate that our method significantly improves image resolution and contrast, achieving performance comparable to multi-angle CPWC methods. Notably, TF-Fusion maintains the high temporal resolution of SAPW, making it well-suited for real-time and resource-constrained ultrasound imaging scenarios.
Yankun Cao, Baolin Sun, Guangtao Zhai, Li-Zhen Cui 0001, Zhi Liu 0004
BIBM1
2025 DiffSegMem: A Novel Conditional Diffusion and Dynamic Memory Propagation Strategy for Left Ventricular Segmentation
abstract
Left ventricular segmentation in echocardiographic videos plays a crucial role in assessing various heart functions and disease diagnoses. However, due to the dynamic nature of echocardiography, maintaining consistent target segmentation across context shifts between frames poses a significant challenge. This task becomes even more difficult in the presence of noise interference in ultrasound images. In this paper, we propose DiffSegMem, which leverages the generative advantages of conditional diffusion model and a contextual memory storage architecture to enhance the accuracy of cross-frame segmentation in dynamic echocardiographic videos. Specifically, we introduce a noise-frequency domain aware dual-branch conditional encoding that establishes noise-resistant conditions at each sampling step, providing a reliable mask for the first frame of the echocardiographic sequence. As a result, our method does not require any prompting for the first video frame. Additionally, we propose a dynamic memory propagation strategy that utilizes memory transfer switches to extract memory from deeply linked video frames using a positional switch, with memory blocks carrying contextual clues prompting segmentation frame by frame. We evaluate our method on publicly available echocardiographic segmentation datasets and demonstrate state-of-the-art performance compared to existing models, outperforming current supervised methods for prompt-based segmentation.
Xifeng Hu, Jingchuan Wang, Yankun Cao, Zhi Liu 0004
BIBM3
2025 IVUS-Guided Rationality Evaluation Model for Clinical Stent Implantation Strategy
abstract
In the treatment of coronary artery disease, the stent implantation strategy (SIS) plays a decisive role in both the procedural success rate and the long-term prognosis of patients. However, SIS determination is primarily based on the experience of the doctors, visual evaluation based on coronary angiography, and the real-time interpretation of intravascular ultrasound (IVUS) during the procedure. To improve the rationality of stent implantation and reduce postoperative adverse events, we propose an objective, automated, and efficient method to evaluate the rationality of SIS. Specifically, a time-series chunk encoder is constructed to extract temporal features from IVUS video, enabling the model to handle long-sequence feature extraction while accommodating input videos of varying lengths. Furthermore, a frequency domain feature encoder is employed to extract spectral characteristics from IVUS video. Meanwhile, stent, balloon, and placement information is input into the model in text form, with feature extraction performed using two pretrained BERT models in medical Chinese and medical English. In order to ensure the discriminability of text features, a text memory bank and text matching loss function are designed to participate in model training in conjunction with other commonly used loss functions. Finally, two Transformer Layers integrate the temporal, spectral, and textual modalities, and a fully connected layer is used to determine the rationality of the SIS. Validation in a private rationality evaluation dataset for SIS demonstrates that our method effectively performs rationality evaluation tasks, providing valuable auxiliary recommendations for clinical decision-making.
Yankun Cao, Mengkang Fan, Zhi Liu 0004
BIBM2
2025 Cooperative metric learning-based hybrid transformer for automatic recognition of standard echocardiographic multi-views
Yankun Cao, Xiaoxiao Cui, Xifeng Hu, Yuezhong Zhang, Zhi Liu 0004, Li-Zhen Cui 0001, Shuo Li 0001
Future Gener. Comput. Syst.2
2024 SRMAR: Spatiotemporal Representation for Motion Artifact Removal in Intravascular Ultrasound
abstract
Intravascular ultrasound (IVUS) not only reveals changes within the vascular lumen but also illustrates the cross-sectional structure, encompassing aspects such as plaques, vessel wall thickness, morphology, and composition. However, during the image acquisition process, ultrasound imaging of vessels can lead to intraluminal misalignment due to motion, resulting in inaccurate measurement outcomes. Current methods for motion artifact removal in IVUS face the following challenges: (a) Gating, which extracts key gating frames to form a new artifact-free sequence, but often results in the loss of substantial useful information; and (b) Direct artifact removal, which requires lumen segmentation followed by registration, where the accuracy of registration is highly dependent on segmentation accuracy. To address these challenges, this paper proposes a robust direct artifact removal method based on spatiotemporal representations. Specifically, to address the issue of information loss in gating methods, a spatiotemporal representation network is introduced, which primarily relies on temporal granularity normalization and spatial position compensation. To tackle the problem of direct artifact removal methods heavily relying on segmentation accuracy, this paper integrates the segmentation network with the artifact removal network, allowing for mutual supervision. This approach ensures effective motion artifact removal even when segmentation accuracy is not optimal. Experimental results show that our method not only achieves state-of-the-art (SOTA) performance in both quantitative and qualitative evaluations but also maintains robust motion artifact removal even when the segmentation network is replaced with one of lower performance.
Yankun Cao, Guanjie Sun, Xiaoxiao Cui, Li-Zhen Cui 0001, Wenmiao Wang, Zhi Liu 0004, Yuezhong Zhang
BIBM1
2024 Multilevel Causality Learning for Multi-label Gastric Atrophy Diagnosis
Xiaoxiao Cui, Shanzhi Jiang, Baolin Sun, Yankun Cao, Zhen Li 0049, Chaoyang Lv, Zhi Liu 0004, Li-Zhen Cui 0001, Shuo Li 0001
MICCAI (3)5
2024 CausCLIP: Causality-Adapting Visual Scoring of Visual Language Models for Few-Shot Learning in Portable Echocardiography Quality Assessment
Xiaoxiao Cui, Yankun Cao, Yuezhong Zhang, Li-Zhen Cui 0001, Zhi Liu 0004, Shuo Li 0001
MICCAI (1)3
2022 MVFStain: Multiple virtual functional stain histopathology images generation based on specific domain mapping
Yankun Cao, Zhi Liu 0004, Jianye Wang, Xiaoyu Sui, Pengfei Zhang 0017, Li-Zhen Cui 0001, Shuo Li 0001
Medical Image Anal.2
2022 TRSA-Net: Task Relation Spatial Co-Attention for Joint Segmentation, Quantification and Uncertainty Estimation on Paired 2D Echocardiography
abstract
Clinical workflow of cardiac assessment on 2D echocardiography requires both accurate segmentation and quantification of the Left Ventricle (LV) from paired apical 4-chamber and 2-chamber. Moreover, uncertainty estimation is significant in clinically understanding the performance of a model. However, current research on 2D echocardiography ignores this vital task while joint segmentation with quantification, hence motivating the need for a unified optimization method. In this paper, we propose a multitask model with Task Relation Spatial co-Attention (referred as TRSA-Net) for joint segmentation, quantification, and uncertainty estimation on paired 2D echo. TRSA-Net achieves multitask joint learning by novelly exploring the spatial correlation between tasks. The task relation spatial co-attention learns the spatial mapping among task-specific features by non-local and co-excitation, which forcibly joints embedded spatial information in the segmentation and quantification. The Boundary-aware Structure Consistency (BSC) and Joint Indices Constraint (JIC) are integrated into the multitask learning optimization objective to guide the learning of segmentation and quantification paths. The BSC creatively promotes structural similarity of predictions, and JIC explores the internal relationship between three quantitative indices. We validate the efficacy of our TRSA-Net on the public CAMUS dataset. Extensive comparison and ablation experiments show that our approach can achieve competitive segmentation performance and highly accurate results on quantification.
Xiaoxiao Cui, Yankun Cao, Zhi Liu 0004, Xiaoyu Sui, Yuezhong Zhang, Li-Zhen Cui 0001, Shuo Li 0001
IEEE J. Biomed. Health Informatics2
2020 Unsupervised Graph Domain Adaptation for Neurodevelopmental Disorders Diagnosis
Bomin Wang, Zhi Liu 0004, Xiaoyan Xiao, Yankun Cao, Li-Zhen Cui 0001, Pengfei Zhang 0017
MICCAI (2)6
2020 Feature data processing: Making medical data fit deep neural networks
Zhi Liu 0004, Haixia Hou, Yankun Cao, Yuefeng Zhao, Wei Guo 0017, Li-Zhen Cui 0001
Future Gener. Comput. Syst.5