Yue Sun 0001

dblp:69/4248-1 · DBLP profile ↗
← Back
53ranked-venue papers
5as first author
44since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 39 · 4 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 1 first-author · 21 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 HiFi-Mesh: High-Fidelity Efficient 3D Mesh Generation via Compact Autoregressive Dependence
abstract
High-fidelity 3D meshes can be tokenized into one-dimension (1D) sequences and directly modeled using autoregressive approaches for faces and vertices. However, existing methods suffer from insufficient resource utilization, resulting in slow inference and the ability to handle only small-scale sequences, which severely constrains the expressible structural details. We introduce the Latent Autoregressive Network (LANE), which incorporates compact autoregressive dependencies in the generation process, achieving a 6× improvement in maximum generatable sequence length compared to existing methods. To further accelerate inference, we propose the Adaptive Computation Graph Reconfiguration (AdaGraph) strategy, which effectively overcomes the efficiency bottleneck of traditional serial inference through spatiotemporal decoupling in the generation process. Experimental validation demonstrates that LANE achieves superior performance across generation speed, structural detail, and geometric consistency, providing an effective solution for high-quality 3D mesh generation.
Tao Tan 0002, Qinquan Gao, Zhiwen Cao, Xiaohong Liu 0001, Yue Sun 0001
AAAI6
2026 SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition
abstract
Large Language Models (LLMs) hold rich implicit knowledge and powerful transferability. In this paper, we explore the combination of LLMs with the human skeleton to perform action classification and description. However, when treating LLM as a recognizer, two questions arise: 1) How can LLMs understand the skeleton? 2) How can LLMs distinguish among actions? To address these problems, we introduce a novel paradigm named learning Skeleton representation with visual-motion knowledge for Action Recognition (SUGAR). In our pipeline, we first utilize off-the-shelf large-scale video models as a knowledge base to generate visual, motion information related to actions. Then, we propose to supervise skeleton learning through this prior knowledge to yield discrete representations. Finally, we use the LLM with untouched pre-training weights to understand these representations and generate the desired action targets and descriptions. Notably, we present a Temporal Query Projection (TQP) module to continuously model the skeleton signals with long sequences. Experiments on several skeleton-based action classification benchmarks demonstrate the efficacy of our SUGAR. Moreover, experiments on zero-shot scenarios show that SUGAR is more versatile than linear-based methods.
Qilang Ye, Yu Zhou 0015, Jie Zhang 0081, Xuanming Guo, Mingkui Tan, Weicheng Xie 0001, Yue Sun 0001, Tao Tan 0002, Xiaochen Yuan, Ghada Khoriba, Zitong Yu
AAAI9
2026 Anatomy-guided prompting with cross-modal self-alignment for whole-body PET-CT breast cancer segmentation
Jiaju Huang, Xinglong Liang, Shaobin Chen, Yue Sun 0001, Greta S. P. Mok, Shuo Li 0001, Tao Tan 0002
Medical Image Anal.5
2026 Leveraging modality-guided pre-training for dual-prompt-driven multi-cancer PET-CT segmentation
Xinglong Liang, Jiaju Huang, Tianyu Zhang 0006, Luyi Han, Xin Wang 0121, Chunyao Lu, Yue Sun 0001, Jonas Teuwen, Tao Tan 0002, Ritse Mann
Medical Image Anal.8
2026 S2DENet: Shallow suppression and deep enhancement network for general ultrasound image segmentation
Xintao Pang, Jinlin Yang, Zhifan Gao, Chuan Lin 0003, Yue Sun 0001, Shuo Li 0001, Peter H. N. de With, Tao Tan 0002
Medical Image Anal.5
2026 Incorporating global-local tissue changes to predict future breast cancer from longitudinal screening mammograms
Xin Wang 0121, Tao Tan 0002, Eric Marcus, Chunyao Lu, Luyi Han, Antonio Portaluri, Ruisheng Su, Tianyu Zhang 0006, Xinglong Liang, Regina Beets-Tan, Katja Pinker-Domenig, Yue Sun 0001, Ritse Mann, Jonas Teuwen
Medical Image Anal.14
2026 TF-LLM: Enhanced time series analysis with time-frequency large language models
Yuhang Zhang 0034, Zitong Yu, Mingtong Dai, Yue Sun 0001, Tao Tan 0002
Neural Networks4
2026 3D-IRMM-Net: An Interpretable Radiologist-Mimicking 3D Multimodal Network for Lesion Detection in ABUS With LLM-Based Guidance and Uncertainty Quantification
abstract
Automated breast ultrasound (ABUS) has emerged as a promising tool for breast lesion detection, but most existing deep learning models for ABUS lack transparency and fail to align with clinical reasoning processes. This limits their interpretability and hinders clinical adoption. Therefore, we propose 3D-IRMM-Net, a radiologist-mimicking 3D multimodal network. It integrates visual, textual, and semantic cues to emulate hierarchical diagnostic reasoning. At its core, the network integrates two novel components: the 3D Scale-Expert Convolution (3D-SEC) block, which enables parameter-efficient, topology-aware feature routing via a shared graph convolutional layer; and the Adaptive Spatial-Gaussian (ASG) block, which enhances lesion localization in ABUS by modeling spatial dependencies through multi-scale Gaussian attention. In addition, we propose LLM-Vision activator, a natural-language-driven mechanism that uses language to guide 3D visual attention activation and enables 3D spatial-perceptual feature learning via a pretrained 2D VLM, boosting training efficiency. By aligning AI inference with clinical workflows, 3D-IRMM-Net enhances both detection rates and interpretability. Extensive experiments on ABUS2025, TDSC-ABUS2023, and FPHXD-ABUS datasets confirm its state-of-the-art performance across multiple metrics. At an FPPI of 0.5, our method achieved a patient-level detection rates of 89.7±1.1% (internal) and 85.1±2.0% (external); at an FPPI of 4, they rose to 96.1±1.6% and 91.5 ±1.9%, respectively. Furthermore, the model provides principled uncertainty quantification, improving trustworthiness in real-world deployment. These results confirm the effectiveness of 3D-IRMM-Net in bridging the gap between AI-driven detection and clinical practice.
Xin Qian 0001, Wei Ke 0001, Luoxi Zhu, Yue Sun 0001, Zhang Xiong 0001, Lingyun Bao, Tao Tan 0002
IEEE Trans. Circuits Syst. Video Technol.4
2026 Anatomy-Aware MR-Imaging-Only Radiotherapy
abstract
The synthesis of computed tomography images can supplement electron density information and eliminate MR-CT image registration errors. Consequently, an increasing number of MR-to-CT image translation approaches are being proposed for MR-only radiotherapy planning. However, due to substantial anatomical differences between various regions, traditional approaches often require each model to undergo independent development and use. In this paper, we propose a unified model driven by prompts that dynamically adapt to the different anatomical regions and generates CT images with high structural consistency. Specifically, it utilizes a region-specific attention mechanism, including a region-aware vector and a dynamic gating factor, to achieve MRI-to-CT image translation for multiple anatomical regions. Qualitative and quantitative results on three datasets of anatomical parts demonstrate that our models generate clearer and more anatomically detailed CT images than other state-of-the-art translation models. The results of the dosimetric analysis also indicate that our proposed model generates images with dose distributions more closely aligned to those of the real CT images. Thus, the proposed model demonstrates promising potential for enabling MR-only radiotherapy across multiple anatomical regions. we have released the source code for our RSAM model. The repository is accessible to the public at: https://github.com/yhyumi123/RSAM.
Hao Yang 0026, Yue Sun 0001, Chi Kin Lam, Qiang Zhao 0005, Xiangyu Xiong, Kunyan Cai, Behdad Dashtbozorg, Chenggang Yan 0001, Tao Tan 0002
IEEE Trans. Image Process.2
2026 BRPDNet: A BioRegion Prompt Distillation Network for Physiological Monitoring
abstract
Physiological signal extraction from video data is challenging in dynamic and occluded environments, requiring both accuracy and real-time performance. Existing methods struggle to balance accuracy with model efficiency, particularly under partial facial occlusion or redundant signals. We propose BRPDNet, a novel framework for efficient physiological signal extraction which includes a BioRegion Prompt module for adaptive convolution and a Hyper Distillation module to reduce signal redundancy, ensuring high accuracy and robustness, especially in dynamic and occluded environments. Additionally, the teacher-student network structure enhances the model's adaptability to occlusions and reduces computational complexity without relying on explicit segmentation. Experimental results show that BRPDNet outperforms state-of-the-art models in accuracy, robustness, and efficiency across multiple datasets. For instance, BRPDNet achieves an Mean Absolute Error (MAE) of 1.55 beats per minute (bpm) and a Pearson Correlation Coefficient (PCC) of 0.76 on PURE and UBFC-rPPG datasets with fewer parameters than existing models, ensuring efficient real-time performance.
Zhengxuan Chen, Bin Huang 0014, Kangyang Cao, Tao Tan 0002, Bingsheng Huang, Chan-Tong Lam, Yue Sun 0001
IEEE J. Biomed. Health Informatics7
2026 HRMamba: Fusing Luminance Information for Remote Physiological Measurement in Varied Lighting Conditions
abstract
Camera-based photoplethysmography (cbPPG) represents a non-invasive technique for capturing physiological parameters through facial videos, enabling the extraction of vital signs such as heart rate, respiration rate, and blood oxygen saturation without direct physical contact. Existing deep learning methods face two core challenges when dealing with cbPPG: firstly, extracting weak PPG signals from video segments with large spatial and temporal redundancy and understanding their periodic patterns in long contexts; secondly, accurately extracting PPG signals in complex lighting environments, especially in low-light conditions. To address these issues, this paper proposes an end-to-end method based on Mamba, named HRMamba. This method employs temporal difference mamba to process temporal signals and combines bidirectional state space to enable Mamba to robustly understand the scene and learn the periodic patterns of PPG. Furthermore, a luminance post-processing module is designed to extract luminance information from the video without enhancing lighting or altering the original video data, and embed it into the PPG signal. Experimental results demonstrate that HRMamba achieves state-of-the-art performance, and the designed luminance post-processing module can be applied in various lighting environments, significantly enhancing the performance in dark environments without degrading the performance in normal light scenes.
Nuoer Long, Wei Ke 0001, Chan-Tong Lam, Tao Tan 0002, Zitong Yu, Yue Sun 0001
IEEE J. Biomed. Health Informatics7
2026 SABPI-Net: A Structure-Aware Bidirectional Proxy Interaction Network for Infantile Retinal Disease Diagnosis
abstract
Delayed treatment of infantile retinal disease can reduce its effectiveness and may cause severe and irreversible damage. Automated diagnosis of infant retinal diseases faces challenges including subtle early lesions, diverse clinical phenotypes, imaging variations, and imbalanced data. To address these, which cannot be well addressed by existing general foundation models, we propose structure-aware bidirectional proxy interaction network (SABPI-Net) in a universal learning framework. SABPI-Net incorporates a high-frequency mapping branch, and employs a proposed proxy interaction attention module to enable effective interaction between its trunk feature encoding branch and the high-frequency mapping branch, thereby facilitating enhanced perception of retinal detail structures. Domain-agnostic embedding space self-matching, guided by a memory-bank low-frequency component replacement strategy, promotes domain-invariant learning and consistent model performance under diverse image styles. Finally, the tail-aware feature fusion strategy for fine-tuning further enhances the model's diagnostic sensitivity to tailed diseases. In this study, three classification tasks related to infant retinal diseases are implemented on the largest clinical infant retina dataset to date, covering 19 infant retinal diseases or normal conditions. SABPI-Net achieves superior performance compared to 13 SOTA methods, with 95.32% accuracy on mainstream clinical tasks, 73.58% on ROP five-stage classification, and 84.25% on multi-disease classification, representing improvements of 1.57%, 1.88%, and 4.71% respectively over the best competing methods. Extensive experiments demonstrate the effectiveness and superiority of SABPI-Net in diagnosing infant retinal diseases.
Shaobin Chen, Huazhu Fu, Jiaju Huang, Zhenquan Wu, Behdad Dashtbozorg, Bai Ying Lei, Yue Sun 0001
IEEE Trans. Medical Imaging9
2026 Multi-Granularity Query Network With Adaptive Category Feature Embedding for Behavior Recognition
abstract
Behavior recognition is a highly challenging task, particularly in scenarios requiring unified recognition across both human and animal subjects. Most existing approaches primarily focus on single-species datasets or rely heavily on prior information such as species labels, positional annotations, or skeletal keypoints, which limits their applicability in real-world scenarios where species labels may be ambiguous or annotations are insufficient. To address these limitations, we propose a query-based Multi-Granularity Behavior Recognition Network that directly mines cross-species shared spatiotemporal behavior patterns from raw video inputs. Specifically, we design a Multi-Granularity Query module to effectively fuse fine-grained and coarse-grained features, thereby enhancing the model's capability in capturing spatiotemporal dynamics at different granularities. Additionally, we introduce a Category Query Decoder that leverages learnable category query vectors to achieve explicit behavior category modeling and mapping. Without relying on any extra annotations, the proposed method achieves unified recognition of multi-species and multi-category behaviors, setting a new state-of-the-art on the Animal Kingdom dataset and demonstrating strong generalization ability on the Charades dataset.
Nuoer Long, Yonghao Dang, Chengpeng Xiong, Shaobin Chen, Tao Tan 0002, Wei Ke 0001, Chan-Tong Lam, Jianqin Yin, Peter H. N. de With, Yue Sun 0001
IEEE Trans. Multim.11
2025 SHIELDNet: Multi-Region Fusion and Denoising for Enhanced rPPG Signal Extraction in Healthcare Monitoring
abstract
Accurate extraction of Remote Photoplethysmography (rPPG) signals from video data is critical for medical applications such as remote patient monitoring. However, the process is hindered by significant challenges, including noise interference, occlusions, and multi bio-region signal processing. To address these, we propose SHIELDNet, an efficient and robust framework for real-time extraction of rPPG signals from multiple anatomical regions, incorporating advanced noise reduction mechanisms. SHIELDNet integrates a novel Differential Attention (DA) module, which adaptively focuses on multiple anatomical regions, enabling the model to effectively handle dynamic real-world conditions. Additionally, the network leverages an advanced Efficient Space Attention Module (ESAM) to enhance spatial feature extraction and multi bio-region signal fusion. BioRegion Prompt Module (BRPM) is further introduced to prioritize region-specific features, reducing the model's dependence on facial features alone. Futhermore, we introduce M-rPPG dataset, a comprehensive multi bio-region reference for BioRegion-based studies with full-body details at higher resolution than existing datasets. Extensive evaluations on multiple public datasets demonstrate significant improvements in Mean Absolute Error (MAE$=\mathbf{5. 2 8} \mathbf{~ b p m}$) and Pearson Correlation Coefficient ($\mathbf{P C C} \boldsymbol{=} \mathbf{0. 8 0}$), outperforming current state-of-the-art models. SHIELDNet provides an effective solution for noncontact, multi bio-region rPPG monitoring. We will release our code upon acceptance.
Zhengxuan Chen, Tao Tan 0002, Chan-Tong Lam, Yue Sun 0001
BIBM6
2025 Advancing Clinical Generalization in Remote Photoplethysmography via Age-Specific Physiological Features
abstract
Camera-based remote photoplethysmography (rPPG) enables contactless monitoring of important physiological signals such as heart rate (HR) and respiratory rate (RR), with transformative clinical potential. Despite recent progress under laboratory conditions, most deep-learning approaches overlook age-specific physiological variations and therefore perform poorly in complex clinical scenarios. In this paper, we propose a novel rPPG framework that integrates both explicit and implicit age-specific physiological feature adaptation mechanisms. First, a dynamic channel weighting (DCW) module that enriches the rPPG channel features according to age-related skin optical properties. Second, we propose a prompt-embedded hyper-convolutional (P-HC) layer dynamically adjusts the age-aware convolution kernels using explicit domain priors. Third, a factorized spatio-temporal attention (FAST) mechanism implicitly decomposes spatio-temporal features into age-invariant physiological bases and age-dependent modulation coefficients, thereby enhancing cross-age generalizability. Extensive experiments on five public datasets and clinical validation on 25 preterm infants in neonatal intensive care units demonstrate that the proposed method outperforms state-of-the-art approaches. These findings validate the method's clinical potential for demoaraphically inclusive monitoring.
Ieong Weng San, Xiangmin Luo, Zhengxuan Chen, Tao Tan 0002, Yue Sun 0001
BIBM6
2025 Grad-MTSeg: Mitigating Multi-Task Gradient Conflicts via Hierarchical Gradient Optimization for NPC Radiotherapy Delineation
abstract
The precise delineation of nasopharyngeal carcinoma (NPC) is a critical prerequisite for radiation therapy, but manual methods are inefficient and inconsistent. Current automated segmentation techniques are challenged by the complexity of multi-modal inputs and multi-target outputs, including Organs at Risk (OARs), Gross Tumor Volume of the Primary Tumor (GTVp), Gross Tumor Volume of the Nodal Metastases (GTVn), clinical target volume prescribed with 70 Gy (CTV70), and clinical target volume prescribed with 63 Gy (CTV63). This process is frequently hindered by gradient conflicts during multitask optimization. To resolve these issues, we propose Grad-MTSeg, a novel deep learning framework. Our approach introduces two core innovations: Unidirectional Anatomic Guidance (UAG) to leverage CT structural priors for improved MRI-based segmentation, and Hierarchical Gradient Optimization (HGO) to alleviate destructive gradient interference among tasks. Our framework improves segmentation accuracy for relevant NPC tasks by effectively resolving conflicts across OARs, GTVs, and CTVs. Validation on three external datasets confirms that Grad-MTSeg provides an efficient and precise solution for complex multimodal segmentation, advancing the automation of NPC radiotherapy planning.
Junqiang Ma, Luyi Han, Dengqiang Jia, Tao Tan 0002, Henry H. Y. Tong, Anne W. M. Lee, Sung Inda Soong, Yue Sun 0001
BIBM9
2025 UA-MAE: An Uncertainty-Aware Masked Autoencoder for Breast Lesion Segmentation in Ultrasound Images
abstract
Accurate segmentation of breast lesions is vital for diagnosing breast diseases. Masked image modeling (MIM) with random masking performs well in self-supervised learning but struggles in breast ultrasound segmentation due to (1) ambiguous representations from similar intensities near lesion boundaries and (2) a bias toward irrelevant regions. We propose UA-MAE, an uncertainty-aware masked autoencoder that uses pixel-wise uncertainty maps to dynamically select masking patches, prioritizing boundaries and morphologically relevant lesion areas. Experiments on two public datasets for pre-training and three for fine-tuning show UA-MAE outperforming four state-of-theart SSL methods and two supervised approaches in segmentation accuracy across diverse breast ultrasound images. The code is available at https://github.com/yXiangXiong/UA-MAE.
Xiangyu Xiong, Yue Sun 0001, Jiaju Huang, Da Huang 0004, Shaobin Chen, Zhuoneng Zhang, Tao Tan 0002
BIBM2
2025 MoEdit: On Learning Quantity Perception for Multi-object Image Editing
abstract
Multi-object images are prevalent in various real-world scenarios, including augmented reality, advertisement design, and medical imaging. Efficient and precise editing of these images is critical for these applications. With the advent of Stable Diffusion (SD), high-quality image generation and editing have entered a new era. However, existing methods often struggle to consider each object both individually and part of the whole image editing, both of which are crucial for ensuring consistent quantity perception, resulting in suboptimal perceptual performance. To address these challenges, we propose MoEdit, an auxiliaryfree multi-object image editing framework. MoEdit facilitates high-quality multi-object image editing in terms of style transfer, object reinvention, and background regeneration, while ensuring consistent quantity perception between inputs and outputs, even with a large number of objects. To achieve this, we introduce the Feature Compensation (FeCom) module, which ensures the distinction and separability of each object attribute by minimizing the in-between interlacing. Additionally, we present the Quantity Attention (QTTN) module, which perceives and preserves quantity consistency by effective control in editing, without relying on auxiliary tools. By leveraging the SD model, MoEdit enables customized preservation and modification of specific concepts in inputs with high quality. Experimental results demonstrate that our MoEdit achieves State-Of-The-Art (SOTA) performance in multi-object image editing. Data and codes are available at https://github.com/Tear-kitty/MoEdit.
Ka-Hou Chan, Yue Sun 0001, Chan-Tong Lam, Tong Tong 0001, Zitong Yu, Keren Fu, Xiaohong Liu 0001, Tao Tan 0002
CVPR3
2025 Wasserstein-Regularized Conformal Prediction under General Distribution Shift
abstract
Conformal prediction yields a prediction set with guaranteed $1-\alpha$ coverage of the true target under the i.i.d. assumption, which can fail and lead to a gap between $1-\alpha$ and the actual coverage. Prior studies bound the gap using total variation distance, which cannot identify the gap changes under distribution shift at different $\alpha$, thus serving as a weak indicator of prediction set validity. Besides, existing methods are mostly limited to covariate shifts, while general joint distribution shifts are more common in practice but less researched. In response, we first propose a Wasserstein distance-based upper bound of the coverage gap and analyze the bound using probability measure pushforwards between the shifted joint data and conformal score distributions, enabling a separation of the effect of covariate and concept shifts over the coverage gap. We exploit the separation to design algorithms based on importance weighting and regularized representation learning (WR-CP) to reduce the Wasserstein bound with a finite-sample error bound. WR-CP achieves a controllable balance between conformal prediction accuracy and efficiency. Experiments on six datasets prove that WR-CP can reduce coverage gaps to 3.2% across different confidence levels and outputs prediction sets 37% smaller than the worst-case approach on average.
Yue Sun 0001, Parv Venkitasubramaniam, Sihong Xie
ICLR3
2025 SABPI-Net: A Novel Structure-Aware Network for Accurate and Domain-Invariant Retinopathy of Prematurity Diagnosis
Shaobin Chen, Huazhu Fu, Tao Tan 0002, Jiaju Huang, Xiangyu Xiong, Zhenquan Wu, Behdad Dashtbozorg, Bai Ying Lei, Yue Sun 0001
MICCAI (10)11
2025 C2MAOT: Cross-modal Complementary Masked Autoencoder with Optimal Transport for Cancer Segmentation in PET-CT Images
Jiaju Huang, Shaobin Chen, Xinglong Liang, Zhuoneng Zhang, Yue Sun 0001, Tao Tan 0002
MICCAI (1)6
2025 BiMSRec: A Progressive Image Reconstruction Framework for Medical Image Fusion Guided by Multi-scale Deformation Fields
Nuoer Long, Xinyu Xie, Zitong Yu, Tao Tan 0002, Yue Sun 0001
MICCAI (2)6
2025 TRRG: Towards Truthful Radiology Report Generation With Cross-Modal Disease Clue Enhanced Large Language Models
Yue Sun 0001, Tao Tan 0002, Chao Hao, Yawen Cui, Xinqi Su, Weicheng Xie 0001, LinLin Shen, Zitong Yu
MICCAI (7)2
2025 Tumor Segmentation with Heterogeneity Clustering in Non-Contrast Breast MRI
Xinyu Xie, Luyi Han, Yonghao Li, Yaofei Duan, Yue Sun 0001, Muzhen He, Tao Tan 0002, Dinggang Shen
MICCAI (2)5
2025 FDF-VQVAE: A Frequency Disentanglement and Fusion Learning Framework for Multi-sequence MRI Enhancement
Xinghe Xie, Luyi Han, Yue Sun 0001, Chi Kin Lam, Jian Zheng 0001, Tong Tong 0001, Wei Ke 0001, Chan-Tong Lam, Tao Tan 0002
MICCAI (3)3
2025 SAMASK-CLTR: A Spatial-Aware Mask Guided Learning Model for Benign and Malignant Tumor Classification in ABUS
Peirong Xu, Luoqian Zhu, Jingkun Chen, Xin Qian 0001, Yue Sun 0001, Lingyun Bao, Tao Tan 0002
MICCAI (1)5
2025 MedIQA: A Scalable Foundation Model for Prompt-Driven Medical Image Quality Assessment
Siyi Xun, Yue Sun 0001, Jingkun Chen, Zitong Yu, Tong Tong 0001, Xiaohong Liu 0001, Mingxiang Wu, Tao Tan 0002
MICCAI (13)2
2025 RefineNet: Elevating Medical Foundation Models Through Quality-Centric Data Curation by MLLM-Annotated Proxy Distillation
Ningyi Zhang, Xin Wang 0121, Ka-Hou Chan, Jian Wu 0033, Chan-Tong Lam, Shanshan Wang 0010, Yue Sun 0001, Sio Kei Im, Tao Tan 0002
MICCAI (11)8
2025 An Efficient and Compact Network for Simultaneous Multi-Object Tracking and Behavior Monitoring in Pigeon Farming
abstract
Animal monitoring plays a crucial role in agriculture, especially in improving animal welfare and farming efficiency. However, existing methods cannot simultaneously achieve video-based tracking for each individual while monitoring their behaviors. To overcome this limitation and meet the demand for automated surveillance in large-scale pigeon farming, this study proposes an efficient and compact network for simultaneous multi-object tracking and behavior monitoring in pigeon farming. The network is capable of tracking each pigeon while recognizing their behaviors in a video sequence. The network combines the DETR detector of the lightweight RepViT backbone with the BoT-SORT motion tracker, which tracks accurately while saving costs. A text-guided multimodal action recognition model is introduced in the action recognition stage, which combines spatiotemporal video features with semantic text embedding to enhance classification accuracy. Experimental results show that the proposed method achieves the optimal tracking performance (IDF1: 96.58%) and action recognition accuracy (Top-1: 94.89%), while reducing the network complexity (13.6M parameters) and computational cost (45.67 GFLOPs). This method provides effective technical support to promote accurate management of poultry farming.
Jiefeng Xie, Tao Tan 0002, Yaoji Liu, Dachun Feng, Yue Sun 0001
SMC7
2025 MambaControl: Anatomy Graph-Enhanced Mamba ControlNet with Fourier Refinement for Diffusion-Based Disease Trajectory Prediction
abstract
Modelling disease progression in precision medicine requires capturing complex spatio-temporal dynamics while preserving anatomical integrity. Existing methods often struggle with longitudinal dependencies and structural consistency in progressive disorders. To address these limitations, we introduce MambaControl, a novel framework that integrates selective state-space modelling with diffusion processes for high-fidelity prediction of medical image trajectories. To better capture subtle structural changes over time while maintaining anatomical consistency, MambaControl combines Mamba-Based long-range modelling with graph- guided anatomical control to more effectively represent anatomical correlations. Furthermore, we introduce Fourier-Enhanced spectral graph representations to capture spatial coherence and multiscale detail, enabling MambaControl to achieve state-of-the-art performance in Alzheimer’s disease prediction. Quantitative and regional evaluations demonstrate improved progression prediction quality and anatomical fidelity, highlighting its potential for personalised prognosis and clinical decision support.
Hao Yang 0026, Tao Tan 0002, Weiqin Yang 0003, Kunyan Cai, Calvin Chen, Yue Sun 0001
SMC7
2025 Bi-branch bidirectional coupled interaction fusion network for multi-retinal diseases diagnosis
Shaobin Chen, Tao Tan 0002, Wei Ke 0001, Xiayu Xu, Yanwu Xu 0004, Chan-Tong Lam, Yue Sun 0001
Knowl. Based Syst.9
2025 An orchestration learning framework for ultrasound imaging: Prompt-Guided Hyper-Perception and Attention-Matching Downstream Synchronization
Shuo Li 0001, Shanshan Wang 0010, Zhifan Gao, Yue Sun 0001, Chan-Tong Lam, Xindi Hu, Xin Yang 0009, Dong Ni 0001, Tao Tan 0002
Medical Image Anal.5
2025 BLENet: A Bio-Inspired Lightweight and Efficient Network for Left Ventricle Segmentation in Echocardiography
Xintao Pang, Fengjuan Yao, Yue Sun 0001, Edmundo Patricio Lopes Lao, Chuan Lin 0003, Patrick Pang 0001, Wei Wang 0181, Zhifan Gao, Tao Tan 0002
IEEE Trans. Circuits Syst. Video Technol.4
2025 3MT-Net: A Multi-Modal Multi-Task Model for Breast Cancer and Pathological Subtype Classification Based on a Multicenter Study
abstract
Breast cancer poses a significant threat to women's health, and ultrasound plays a critical role in the assessment of breast lesions. This study introduces a prospective deep learning architecture, termed the "Multi-modal Multi-task Network" (3MT-Net), which integrates clinical data with B-mode and color Doppler ultrasound images. Specifically, an AM-CapsNet is employed to extract key features from ultrasound images, while a cascaded cross-attention mechanism is utilized to fuse clinical data. Moreover, an ensemble learning approach with an optimization algorithm is adopted to dynamically assign weights to different modalities, accommodating both high-dimensional and low-dimensional data. The 3MT-Net performs binary classification of benign versus malignant lesions and further classifies the pathological subtypes. Data were retrospectively collected from nine medical centers to ensure the broad applicability of the 3MT-Net. Two separate testsets were created and extensive experiments were conducted. Comparative analyses demonstrated that the AUC of the 3MT-Net outperforms the industry-standard computer-aided detection product, S-Detect, by 1.4% to 3.8%.
Yaofei Duan, Patrick Pang 0001, Rongsheng Wang 0004, Yue Sun 0001, Chuntao Liu, Xirong Yuan, Pengjie Song, Chan-Tong Lam, Ligang Cui, Tao Tan 0002
IEEE J. Biomed. Health Informatics5
2025 Guest Editorial: Multi-Modal Joint Learning in Healthcare Imaging
Tao Tan 0002, Yue Sun 0001, Shandong Wu
IEEE J. Biomed. Health Informatics3
2025 UniMRISegNet: Universal 3D Network for Various Organs and Cancers Segmentation on Multi-Sequence MRI
abstract
Three-dimensional organ and cancer segmentation based on multi-sequence MRI is crucial for assisting clinical diagnosis. However, current automated segmentation methods often focus on specific sequences, specific organs, and specific cancers, i.e., lack of generality. To address this issue, we propose a universal segmentation network for multi-sequence MRI (UniMRISegNet) that can segment multiple organs and cancers. UniMRISegNet features a shared encoder-decoder architecture equipped with contextual prompt generation (CPG) and prompt-conditioned dynamic convolution (PCDC) modules. The CPG module encodes sequence-specific, position-specific, and organ/cancer-specific text prompts as prior information to inform UniMRISegNet about the specific task to be executed. The PCDC module can adaptively generate model weights based on the assigned prompts, enhancing the segmentation capabilities of the UniMRISegNet for specific tasks. To mitigate discrepancies between different sequences of the same organ and capture similarities between related sequences, we design a novel loss function called Semantic-Aware Cosine Similarity Loss (SACSL), which integrates the cosine similarity of text embeddings to reconcile discrepancies and similarities between MRI sequences of the same organ. We created a large-scale annotated multi-sequence, multi-organ, and multi-cancer segmentation workflow (MSOCS), and demonstrated that our UniMRISegNet outperforms other universal networks and single-task networks on MSOCS. Furthermore, the universal weights from MSOCS can be transferred to never-before-seen downstream tasks, achieving superior performance compared to training from scratch.
Zhuoneng Zhang, Luyi Han, Tianyu Zhang 0006, Qinquan Gao, Tong Tong 0001, Yue Sun 0001, Tao Tan 0002
IEEE J. Biomed. Health Informatics7
2024 UniUSNet: A Promptable Framework for Universal Ultrasound Disease Prediction and Tissue Segmentation
abstract
Ultrasound is widely used in clinical practice due to its affordability, portability, and safety. However, current AI research often overlooks combined disease prediction and tissue segmentation. We propose UniUSNet, a universal framework for ultrasound image classification and segmentation. This model handles various ultrasound types, anatomical positions, and input formats, excelling in both segmentation and classification tasks. Trained on a comprehensive dataset with over 9.7K annotations from 7 distinct anatomical positions, our model matches state-of-the-art performance and surpasses single-dataset and ablated models. Zero-shot and fine-tuning experiments show strong generalization and adaptability with minimal fine-tuning. We plan to expand our dataset and refine the prompting mechanism, with model weights and code available at (https://github.com/Zehui-Lin/UniUSNet).
Zhuoneng Zhang, Xindi Hu, Zhifan Gao, Xin Yang 0009, Yue Sun 0001, Dong Ni 0001, Tao Tan 0002
BIBM6
2024 A Parameterized Generative Adversarial Network Using Cyclic Projection for Explainable Medical Image Classifications
abstract
Although current data augmentation methods are successful to alleviate the data insufficiency, conventional augmentation are primarily intra-domain while advanced generative adversarial networks (GANs) generate images remaining uncertain, particularly in small-scale datasets. In this paper, we propose a parameterized GAN (ParaGAN) that effectively controls the changes of synthetic samples among domains and highlights the attention regions for downstream classification. Specifically, ParaGAN incorporates projection distance parameters in cyclic projection and projects the source images to the decision boundary to obtain the class-difference maps. Our experiments show that ParaGAN can consistently outperform the existing augmentation methods with explainable classification on two small-scale medical datasets.
Xiangyu Xiong, Yue Sun 0001, Xiaohong Liu 0001, Chan-Tong Lam, Tong Tong 0001, Hao Chen 0037, Qinquan Gao, Wei Ke 0001, Tao Tan 0002
ICASSP2
2024 Automatic detection of breast lesions in automated 3D breast ultrasound with cross-organ transfer learning
abstract
Deep convolutional neural networks have garnered considerable attention in numerous machine learning applications, particularly in visual recognition tasks such as image and video analyses. There is a growing interest in applying this technology to diverse applications in medical image analysis. Automated three-dimensional Breast Ultrasound is a vital tool for detecting breast cancer, and computer-assisted diagnosis software, developed based on deep learning, can effectively assist radiologists in diagnosis. However, the network model is prone to overfitting during training, owing to challenges such as insufficient training data. This study attempts to solve the problem caused by small datasets and improve model detection performance. We propose a breast cancer detection framework based on deep learning (a transfer learning method based on cross-organ cancer detection) and a contrastive learning method based on breast imaging reporting and data systems (BI-RADS). When using cross organ transfer learning and BIRADS based contrastive learning, the average sensitivity of the model increased by a maximum of 16.05%. Our experiments have demonstrated that the parameters and experiences of cross-organ cancer detection can be mutually referenced, and contrastive learning method based on BI-RADS can improve the detection performance of the model.
B. A. O. Lingyun, Zhengrui Huang, Yue Sun 0001, Hui Chen 0020, Xiaochen Yuan, Tao Tan 0002
Virtual Real. Intell. Hardw.4
2023 A Hybrid Supervised Fusion Deep Learning Framework for Microscope Multi-Focus Images
Qiuhui Yang, Hao Chen 0037, Mingfeng Jiang, Jiong Zhang 0004, Yue Sun 0001, Tao Tan 0002
CGI (4)6
2023 Weakly Supervised Cerebellar Cortical Surface Parcellation with Self-Visual Representation Learning
Zhengwang Wu, Fenqiang Zhao, Yue Sun 0001, Dajiang Zhu, Tianming Liu 0001, Valerie Jewells, Weili Lin, Li Wang 0026, Gang Li 0001
MICCAI (8)5
2021 Construction of Longitudinally Consistent 4D Infant Cerebellum Atlases Based on Deep Learning
Liangjun Chen, Zhengwang Wu, Dan Hu 0004, Yuchen Pei, Fenqiang Zhao, Yue Sun 0001, Weili Lin, Li Wang 0026, Gang Li 0001
MICCAI (4)6
2021 Using BI-RADS Stratifications as Auxiliary Information for Breast Masses Classification in Ultrasound Images
abstract
Breast Ultrasound (BUS) imaging has been recognized as an essential imaging modality for breast masses classification in China. Current deep learning (DL) based solutions for BUS classification seek to feed ultrasound (US) images into deep convolutional neural networks (CNNs), to learn a hierarchical combination of features for discriminating malignant and benign masses. One existing problem in current DL-based BUS classification was the lack of spatial and channel-wise features weighting, which inevitably allow interference from redundant features and low sensitivity. In this study, we aim to incorporate the instructive information provided by breast imaging reporting and data system (BI-RADS) within DL-based classification. A novel DL-based BI-RADS Vector-Attention Network (BVA Net) that trains with both texture information and decoded information from BI-RADS stratifications was proposed for the task. Three baseline models, pre-trained DenseNet-121, ResNet-50 and Residual-Attention Network (RA Net) were included for comparison. Experiments were conducted on a large scale private main dataset and two public datasets, UDIAT and BUSI. On the main dataset, BVA Net outperformed other models, in terms of AUC (area under the receiver operating curve, 0.908), ACC (accuracy, 0.865), sensitivity (0.812) and precision (0.795). BVA Net also achieved the high AUC (0.87 and 0.882) and ACC (0.859 and 0.843), on UDIAT and BUSI. Moreover, we proposed a method that integrates both BVA Net binary classification and BI-RADS stratification estimation, called integrated classification. The introduction of integrated classification helped improving the overall sensitivity while maintaining a high specificity.
Qinyang Lu, Aijun Yu, Yi Xu 0001, Xiaoling Xia, Yue Sun 0001, Jing Xiao 0006, Lingyun Huang
IEEE J. Biomed. Health Informatics8
2021 Multi-Site Infant Brain Segmentation Algorithms: The iSeg-2019 Challenge
abstract
To better understand early brain development in health and disorder, it is critical to accurately segment infant brain magnetic resonance (MR) images into white matter (WM), gray matter (GM), and cerebrospinal fluid (CSF). Deep learning-based methods have achieved state-of-the-art performance; h owever, one of the major limitations is that the learning-based methods may suffer from the multi-site issue, that is, the models trained on a dataset from one site may not be applicable to the datasets acquired from other sites with different imaging protocols/scanners. To promote methodological development in the community, the iSeg-2019 challenge (http://iseg2019.web.unc.edu) provides a set of 6-month infant subjects from multiple sites with different protocols/scanners for the participating methods. T raining/validation subjects are from UNC (MAP) and testing subjects are from UNC/UMN (BCP), Stanford University, and Emory University. By the time of writing, there are 30 automatic segmentation methods participated in the iSeg-2019. In this article, 8 top-ranked methods were reviewed by detailing their pipelines/implementations, presenting experimental results, and evaluating performance across different sites in terms of whole brain, regions of interest, and gyral landmark curves. We further pointed out their limitations and possible directions for addressing the multi-site issue. We find that multi-site consistency is still an open issue. We hope that the multi-site dataset in the iSeg-2019 and this review article will attract more researchers to address the challenging and critical multi-site issue in practice.
Yue Sun 0001, Kun Gao 0002, Zhengwang Wu, Xiaopeng Zong, Zhihao Lei, Ying Wei 0007, Jun Ma 0016, Xiaoping Yang 0001, Xue Feng 0001, Li Zhao 0001, Trung Le Phan, Jitae Shin, Tao Zhong 0002, Yu Zhang 0064, Lequan Yu, Caizi Li, Ramesh Basnet, M. Omair Ahmad, M. N. S. Swamy 0001, Wenao Ma, Qi Dou 0001, Toan Duc Bui, Camilo Bermudez, Bennett A. Landman, Ian H. Gotlib, Kathryn L. Humphreys, Sarah Shultz, Longchuan Li, Sijie Niu, Weili Lin, Valerie Jewells, Dinggang Shen, Gang Li 0001, Li Wang 0026
IEEE Trans. Medical Imaging1
2020 Adaptive-Guided-Coupling-Probability Level Set for Retinal Layer Segmentation
abstract
Quantitative assessment of retinal layer thickness in spectral domain-optical coherence tomography (SD-OCT) images is vital for clinicians to determine the degree of ophthalmic lesions. However, due to the complex retinal tissues, high-level speckle noises and low intensity constraint, how to accurately recognize the retinal layer structure still remains a challenge. To overcome this problem, this paper proposes an adaptive-guided-coupling-probability level set method for retinal layer segmentation in SD-OCT images. Specifically, based on Bayes's theorem, each voxel probability representation is composed of two probability terms in our method. The first term is constructed as neighborhood Gaussian fitting distribution to characterize intensity information for each intra-retinal layer. The second one is boundary probability map generated by combining anatomical priors and adaptive thickness information to ensure surfaces evolve within a proper range. Then, the voxel probability representation is introduced into the proposed segmentation framework based on coupling probability level set to detect layer boundaries. A total of 1792 retinal B-scan images from 4 SD-OCT cubes in healthy eyes, 5 cubes in abnormal eyes with central serous chorioretinaopathy and 5 SD-OCT cubes in abnormal eyes with age-related macular disease are used to evaluate the proposed method. The experiment demonstrates that the segmentation results obtained by the proposed method have a good consistency with ground truth, and the proposed method outperforms six methods in the layer segmentation of uneven retinal SD-OCT images.
Yue Sun 0001, Sijie Niu, Xizhan Gao, Jie Su 0010, Jiwen Dong, Yuehui Chen, Li Wang 0026
IEEE J. Biomed. Health Informatics1
2019 Video-based discomfort detection for infants
Yue Sun 0001, Caifeng Shan, Tao Tan 0002, Xi Long 0001, Arash Pourtaherian, Svitlana Zinger, Peter H. N. de With
Mach. Vis. Appl.1
2015 Biomechanically Constrained Surface Registration: Application to MR-TRUS Fusion for Prostate Interventions
abstract
In surface-based registration for image-guided interventions, the presence of missing data can be a significant issue. This often arises with real-time imaging modalities such as ultrasound, where poor contrast can make tissue boundaries difficult to distinguish from surrounding tissue. Missing data poses two challenges: ambiguity in establishing correspondences; and extrapolation of the deformation field to those missing regions. To address these, we present a novel non-rigid registration method. For establishing correspondences, we use a probabilistic framework based on a Gaussian mixture model (GMM) that treats one surface as a potentially partial observation. To extrapolate and constrain the deformation field, we incorporate biomechanical prior knowledge in the form of a finite element model (FEM). We validate the algorithm, referred to as GMM-FEM, in the context of prostate interventions. Our method leads to a significant reduction in target registration error (TRE) compared to similar state-of-the-art registration algorithms in the case of missing data up to 30%, with a mean TRE of 2.6 mm. The method also performs well when full segmentations are available, leading to TREs that are comparable to or better than other surface-based techniques. We also analyze robustness of our approach, showing that GMM-FEM is a practical and reliable solution for surface-based registration.
Siavash Khallaghi, C. Antonio Sánchez, Abtin Rasoulian, Yue Sun 0001, Farhad Imani, Amir Khojaste, Orcun Goksel, Cesare Romagnoli, Hamidreza Abdi, Silvia D. Chang, Parvin Mousavi, Aaron Fenster, Aaron D. Ward, Sidney S. Fels, Purang Abolmaesumi
IEEE Trans. Medical Imaging4
2015 Three-Dimensional Nonrigid MR-TRUS Registration Using Dual Optimization
abstract
In this study, we proposed an efficient nonrigid magnetic resonance (MR) to transrectal ultrasound (TRUS) deformable registration method in order to improve the accuracy of targeting suspicious regions during a three dimensional (3-D) TRUS guided prostate biopsy. The proposed deformable registration approach employs the multi-channel modality independent neighborhood descriptor (MIND) as the local similarity feature across the two modalities of MR and TRUS, and a novel and efficient duality-based convex optimization-based algorithmic scheme was introduced to extract the deformations and align the two MIND descriptors. The registration accuracy was evaluated using 20 patient images by calculating the TRE using manually identified corresponding intrinsic fiducials in the whole gland and peripheral zone. Additional performance metrics [Dice similarity coefficient (DSC), mean absolute surface distance (MAD), and maximum absolute surface distance (MAXD)] were also calculated by comparing the MR and TRUS manually segmented prostate surfaces in the registered images. Experimental results showed that the proposed method yielded an overall median TRE of 1.76 mm. The results obtained in terms of DSC showed an average of 80.8±7.8% for the apex of the prostate, 92.0±3.4% for the mid-gland, 81.7±6.4% for the base and 85.7±4.7% for the whole gland. The surface distance calculations showed an overall average of 1.84±0.52 mm for MAD and 6.90±2.07 mm for MAXD.
Yue Sun 0001, Jing Yuan 0001, Wu Qiu, Martin Rajchl, Cesare Romagnoli, Aaron Fenster
IEEE Trans. Medical Imaging1
2014 3D Prostate TRUS Segmentation Using Globally Optimized Volume-Preserving Prior
Wu Qiu, Martin Rajchl, Fumin Guo, Yue Sun 0001, Eranga Ukwatta, Aaron Fenster, Jing Yuan 0001
MICCAI (1)4
2014 Dual optimization based prostate zonal segmentation in 3D MR images
Wu Qiu, Jing Yuan 0001, Eranga Ukwatta, Yue Sun 0001, Martin Rajchl, Aaron Fenster
Medical Image Anal.4
2014 Prostate Segmentation: An Efficient Convex Optimization Approach With Axial Symmetry Using 3-D TRUS and MR Images
abstract
We propose a novel global optimization-based approach to segmentation of 3-D prostate transrectal ultrasound (TRUS) and T2 weighted magnetic resonance (MR) images, enforcing inherent axial symmetry of prostate shapes to simultaneously adjust a series of 2-D slice-wise segmentations in a "global" 3-D sense. We show that the introduced challenging combinatorial optimization problem can be solved globally and exactly by means of convex relaxation. In this regard, we propose a novel coherent continuous max-flow model (CCMFM), which derives a new and efficient duality-based algorithm, leading to a GPU-based implementation to achieve high computational speeds. Experiments with 25 3-D TRUS images and 30 3-D T2w MR images from our dataset, and 50 3-D T2w MR images from a public dataset, demonstrate that the proposed approach can segment a 3-D prostate TRUS/MR image within 5-6 s including 4-5 s for initialization, yielding a mean Dice similarity coefficient of 93.2%±2.0% for 3-D TRUS images and 88.5%±3.5% for 3-D MR images. The proposed method also yields relatively low intra- and inter-observer variability introduced by user manual initialization, suggesting a high reproducibility, independent of observers.
Wu Qiu, Jing Yuan 0001, Eranga Ukwatta, Yue Sun 0001, Martin Rajchl, Aaron Fenster
IEEE Trans. Medical Imaging4
2013 Fast Globally Optimal Segmentation of 3D Prostate MRI with Axial Symmetry Prior
Wu Qiu, Jing Yuan 0001, Eranga Ukwatta, Yue Sun 0001, Martin Rajchl, Aaron Fenster
MICCAI (2)4
2013 Efficient Convex Optimization Approach to 3D Non-rigid MR-TRUS Registration
Yue Sun 0001, Jing Yuan 0001, Martin Rajchl, Wu Qiu, Cesare Romagnoli, Aaron Fenster
MICCAI (1)1