Yuanbo He

dblp:210/3759 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0001-8656-4496ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
YearPublicationVenuePosition
2026 MedSAM-guided geometry-aware 2D-3D feature fusion for medical image registration
Yuanbo He, Aimin Hao
Neural Networks2
2025 ArbitraryFlow: Towards One Step Generative Biomedical Image Segmentation
abstract
Biomedical image segmentation has witnessed significant advancements through deep learning, wherein diffusionbased generative models have emerged as compelling alternatives to traditional discriminative methodologies by reconceptualizing segmentation through an image-guided noise-to-mask generation paradigm. Yet, their clinical deployment remains constrained by intensive iterative denoising procedures necessitating hundreds to thousands of model function evaluations, compounded by inherent stochastic sampling mechanisms that invariably produce non-deterministic outputs violating the consistency imperatives of clinical decision-making requirements. To overcome these limitations, we propose ArbitraryFlow, a deterministic flowmatching framework that reformulates segmentation as a direct image-to-mask transport problem through an ordinary differential equation (ODE) modeling a continuous-time velocity field that governs the distribution transport dynamics. Crucially, our framework incorporates a two-time flow map mechanism that learns an average velocity field between arbitrary temporal endpoints, consequently enabling few(one)-step mask synthesis without recourse to ODE integration procedures. The velocity field is instantiated via a Diffusion Transformer (DiT) backbone equipped with adaptive layer normalization for effective dual time-conditioned feature modulation while preserving parametric and computational efficiency. Comprehensive experiments across three multimodal datasets (MSD-BraTS, GlaS, BUSI) demonstrate ArbitraryFlow achieves superior segmentation performance relative to established prior arts and consistently outperforms alternative sampling acceleration techniques under various inference budgets. Its single-step instantiation manifests a Dice decline of less than 1.5 % compared to its optimal 1000 step configuration, preserving parity with exhaustive diffusion baselines while delivering computational acceleration exceeding two orders of magnitude. ArbitraryFlow redefines the efficiency- accuracy trade-off in generative segmentation, enabling realtime clinical applications previously unattainable with diffusion paradigms, opening new possibilities for real-time biomedical image analysis. Our code is available at https://github.com/lyupengju/shortcut_flows.
Pengju Lyu, Haopeng Jing, Yuanbo He
BIBM4
2025 Genomics-Aware Multimodal Self-Supervised Learning for Cancer Survival Prediction
abstract
Survival prediction in cancer diagnosis is a critical research task. Current methods often employ the multimodal feature fusion of pathological images and genomics data within a weakly-supervised learning paradigm. However, these approaches fail to efficiently learn the intrinsic features of large amount of unlabeled WSIs and neglect the strong associations between genomics data and pathological images, resulting in reduced prognostic accuracy. To address these challenges, we propose a novel Genomics-Aware Multimodal Self-Supervised Learning model that designs a multimodal pretext task, improving learning of intra-modal features and inter-modal correlations without additional annotations. Specifically, we randomly mask pathological patch features and fuse unmasked pathology representations with genomics representations via a cross-modal attention module. Then we add mask tokens to the genomicsguided pathology representation and reconstruct the missing parts via a reconstruction decoder. Experimental results on four TCGA datasets demonstrate the superior performance of our method compared to state-of-the-art methods, highlighting its potential for advancing survival prediction. Our code is available at https://github.com/sunkevin101/GMSL.
Yuanbo He, Zining Liu, Jiahao Cui 0001, Shuai Li 0001
BIBM3
2025 Advancing MRI segmentation with CLIP-driven semi-supervised learning and semantic alignment
Kexuan Li, Jingjuan Liu, Xuehao Wang, Yuanbo He, Huadan Xue, Aimin Hao, Shuai Li 0001
Neurocomputing6
2024 SL-SFGR: Segmentation Learning Coupling Spatial-Frequency Structure Information Enhancement for Guiding Registration
abstract
The core of medical image registration lies in the alignment of corresponding structures. Hence, the effective construction of structure information in images is crucial for guiding registration. Some methods directly introduce explicit structure priors for assisting registration training, but obtaining high-precision structure priors is intrinsically challenging. As a segmentation model for obtaining structure priors, its capability is built upon the effective construction of structure information, thus its model learning can also be used to guide registration. However, existing registration-segmentation joint methods mostly only use segmentation results at the output level to constrain registration, neglecting the guidance of segmentation learning at the feature level for registration. Moreover, most existing methods only extract features in the image spatial domain for registration, overlooking the structure information in the frequency domain that is more easily captured to guide registration. To this end, this paper proposes an innovative registration method, namely Segmentation Learning Coupling Spatial-Frequency Structure Information Enhancement for Guiding Registration (SL-SFGR). Specifically, first, a semi-supervised segmentation learning network is constructed based on the registration deformation field to introduce structure features suitable for guiding registration. Second, an adaptive feature enhancement module is built in the spatial-frequency dual domain to further strengthen the inherent structure features. Finally, Dynamic Weight Average (DWA) is utilized for joint optimization of the model. The effectiveness of the proposed method has been verified on different brain MRI datasets. The related code is available at: https://github.com/goghfan/SL-SFGR/.
Yuanbo He, Shuai Li 0001, Aimin Hao, Desen Cao
BIBM2
2023 Diffusing Coupling High-Frequency-Purifying Structure Feature Extraction for Brain Multimodal Registration
abstract
The core of medical image registration is the alignment of corresponding structures. However, in multimodal image registration, substantial differences in appearance (intensity distribution) of the images often compel the registration model to prioritize intensity information over structure information, resulting in low accuracy of registration. Therefore, the disentangling structure information from intensity information is vital to improve the registration effectiveness. To this end, we propose a diffusing coupling high-frequency-purifying structure feature extraction for brain multimodal registration. Specifically, the denoising diffusion probabilistic models (DDPM) is firstly utilized to extract complete feature information from images. Then, the discrete cosine transform (DCT) is applied to purify high-frequency structure information from the complete feature information for registration. Furthermore, structure consistency constraint (SCC) is introduced based on purified structure information to emphasize the core position of the structure in registration. Through comprehensive comparisons with traditional and learning-based methods on the multimodal brain MRI dataset, our method demonstrates superior accuracy and stability in brain multimodal registration. Our code is available at https://github.com/goghfan/DDNet.
Yuanbo He, Shuai Li 0001, Aimin Hao, Desen Cao
BIBM2
2023 RGB and LUT based Cross Attention Network for Image Enhancement
Tengfei Shi, Chenglizhao Chen, Yuanbo He, Wenfeng Song, Aimin Hao
BMVC3
2023 Integrating Transient and Long-term Physical States for Depression Intelligent Diagnosis
Li Kuang, Huaiqian Ye, Yuanbo He
BMVC6
2023 Joint Probability Distribution Regression for Image Cropping
abstract
Image cropping aims at locating a candidate (rectangle region) with the highest aesthetic quality in professional photography. One solution of the previous methods is to generate a large number of candidates and then filter them, which leads to low efficiency. Another idea directly regresses the candidate coordinates to speed up but ignores the aesthetic subjectivity of the candidate’s evaluation, limiting the model’s performance. In this paper, we present an Aesthetic and Composition joint Probability Distribution regression Network (ACPD-Net) to explicitly investigate the process of generating the candidate with a joint probability distribution paradigm to improve the performance of cropping results in an efficient way. The joint probability distribution paradigm between location and size branch can identify the subjective aesthetic region and satisfy the objective composition rules in an end-to-end manner. Our method has been tested on the FCDB and FLMS datasets, which shows the superiority of ACPD-Net. The code is available at https://github.com/flyingbird93/ACPD-Net.
Tengfei Shi, Chenglizhao Chen, Yuanbo He, Wenfeng Song, Aimin Hao
ICIP3
2022 Novel cross LSTM for predicting the changes of complementary pelvic angles between standing and sitting
abstract
Sagittal spino-pelvic balance has been increasingly emphasized in hip surgery. The conversion between standing and sitting, characterized by complementary pelvic angles (pelvic tilt, pt and sacral slope, ss), involves a congruent sagittal spino-pelvic relationship. Hence, the changes of complementary pelvic angles pt, ss between standing and sitting could reflect the mechanism of sagittal spino-pelvic balance, and should be analyzed in evidence-based hip surgery planning. To this end, we propose a novel cross LSTM (C-LSTM) framework embedding the conversion between standing and sitting by cross-mapping, to predict the changes of complementary pelvic pt, ss between standing and sitting. Furthermore, to introduce the prior knowledge of the invariance of pelvic incidence, pi, two dual C-LSTMs are integrated to construct a much more powerful Fused C-LSTM. We have conducted extensive experiments on the sagittal standing-sitting dataset for the comprehensive evaluation of the proposed framework. Even in a small samples, Fused C-LSTM can achieve low prediction errors and high correlation between predicted and actual values. Notably, just based on static standing or sitting X-ray, Fused C-LSTM can obtain the change of complementary pt, ss between standing and sitting to assist in formulating a surgical hip plan that conforms to the sagittal spino-pelvic balance.
Yuanbo He, Minwei Zhao, Tianfan Xu, Shuai Li 0001
J. Biomed. Informatics1
2021 Global Correlation and Local Geometric Information Coupled Channel Contrast Learning for Thyroid Nodule Risk Stratification
abstract
Thyroid nodule risk stratification based on ultra-sound images is vital for follow-up clinical treatment. Due to the complexity of the inter-risk stratification difference, experienced physicians are required to comprehensively analyze all ultrasound signs of the thyroid nodule to diagnose the corresponding risk stratification the thyroid nodule should belong to, however, which is often labor-intensive, subjective, and unstable. To this end, we propose a global correlation and local geometric information coupled channel contrast learning network for thyroid nodule risk stratification based on ultrasound images. Specifically, a channel contrast learning module by combining the contrastive learning with a novel cross-class interaction strategy is proposed to obtain the discriminative feature for different risk stratification levels. Furthermore, multiple ultrasound signs may be observed simultaneously in the same risk stratification level. To introduce the correlation among different ultrasound signs into the feature, a global correlation learning module is proposed by the spectral decomposition of the channel correlation matrix. Additionally, some ultrasound signs used as the basis for judging risk stratification are local characteristics. Therefore, a local geometry learning module is proposed using gradient matrix to model the local information of ultrasound signs to further strengthen the feature. Extensive experiments and comprehensive evaluations confirm that the proposed method achieves superior performance and presents to be promising in thyroid nodule risk stratification.
Yuanbo He, Shuai Li 0001, Luying Gao
BIBM2