VLDB 2026 Research / reviewers in the wild / expert
Tao Tan 0002
dblp:06/7832-2
· DBLP profile ↗
110ranked-venue papers
11as first author
99since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 63 · 6 first-author · 57 since 2021Graphics, computer vision, multimedia, augmented reality and games · 49 · 3 first-author · 47 since 2021Artificial intelligence and machine learning · 29 · 4 first-author · 24 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HiFi-Mesh: High-Fidelity Efficient 3D Mesh Generation via Compact Autoregressive DependenceabstractHigh-fidelity 3D meshes can be tokenized into one-dimension (1D) sequences and directly modeled using autoregressive approaches for faces and vertices. However, existing methods suffer from insufficient resource utilization, resulting in slow inference and the ability to handle only small-scale sequences, which severely constrains the expressible structural details. We introduce the Latent Autoregressive Network (LANE), which incorporates compact autoregressive dependencies in the generation process, achieving a 6× improvement in maximum generatable sequence length compared to existing methods. To further accelerate inference, we propose the Adaptive Computation Graph Reconfiguration (AdaGraph) strategy, which effectively overcomes the efficiency bottleneck of traditional serial inference through spatiotemporal decoupling in the generation process. Experimental validation demonstrates that LANE achieves superior performance across generation speed, structural detail, and geometric consistency, providing an effective solution for high-quality 3D mesh generation. Tao Tan 0002, Qinquan Gao, Zhiwen Cao, Xiaohong Liu 0001, Yue Sun 0001 |
AAAI | 2 |
| 2026 | LUMIN: A Longitudinal Multi-modal Knowledge Decomposition Network for Predicting Breast Cancer RecurrenceabstractAccurate prediction of breast cancer recurrence after treatment is essential for improving long-term outcomes. However, existing models are limited by three key challenges: (1) they typically rely on single-modal data, missing cross-modal interactions; (2) they analyze static snapshots, failing to capture disease progression over time; and (3) they often perform coarse feature fusion, lacking semantic disentanglement and interpretability. To address these issues, we propose LUMIN (Longitudinal Multi-modal Knowledge Decomposition Network), a novel framework that integrates longitudinal mammograms and electronic health records (EHRs) for recurrence prediction. LUMIN leverages a vision-language contrastive pretraining backbone to align multi-modal representations and introduces two knowledge extraction modules: (1) a Cross-Modal Disentangled Knowledge Extractor (CM-DKE) that separates shared, complementary, and modality-specific information across imaging and text; and (2) a Temporal Evolution Disentangled Knowledge Extractor (TE-DKE) that captures time-invariant, time-varying, and time-specific features to model disease dynamics. Experiments on a large-scale dataset of 3,924 patients and 19,684 exams show that LUMIN significantly outperforms state-of-the-art baselines, demonstrating its effectiveness in capturing both multi-modal semantics and temporal heterogeneity for recurrence prediction. Chunyao Lu, Tianyu Zhang 0006, Xinglong Liang, Luyi Han, Xin Wang 0121, Nika Rasoolzadeh, Tao Tan 0002, Ritse Mann |
AAAI | 8 |
| 2026 | SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action RecognitionabstractLarge Language Models (LLMs) hold rich implicit knowledge and powerful transferability. In this paper, we explore the combination of LLMs with the human skeleton to perform action classification and description. However, when treating LLM as a recognizer, two questions arise: 1) How can LLMs understand the skeleton? 2) How can LLMs distinguish among actions? To address these problems, we introduce a novel paradigm named learning Skeleton representation with visual-motion knowledge for Action Recognition (SUGAR). In our pipeline, we first utilize off-the-shelf large-scale video models as a knowledge base to generate visual, motion information related to actions. Then, we propose to supervise skeleton learning through this prior knowledge to yield discrete representations. Finally, we use the LLM with untouched pre-training weights to understand these representations and generate the desired action targets and descriptions. Notably, we present a Temporal Query Projection (TQP) module to continuously model the skeleton signals with long sequences. Experiments on several skeleton-based action classification benchmarks demonstrate the efficacy of our SUGAR. Moreover, experiments on zero-shot scenarios show that SUGAR is more versatile than linear-based methods. Qilang Ye, Yu Zhou 0015, Jie Zhang 0081, Xuanming Guo, Mingkui Tan, Weicheng Xie 0001, Yue Sun 0001, Tao Tan 0002, Xiaochen Yuan, Ghada Khoriba, Zitong Yu |
AAAI | 10 |
| 2026 | BPEFNet: a bit plane enhanced fusion network for echocardiographic segmentation
Dang Li, Chi Kin Lam, Xintao Pang, Dashun Zheng, Penny Wong-On Chao, Patrick Pang 0001, Tao Tan 0002 |
Appl. Intell. | 8 |
| 2026 | DeepDesc: integrating retrieval-augmented generation with large language models for smart contract vulnerability detection
Tao Tan 0002, Xiao Chen 0002 |
Empir. Softw. Eng. | 1 |
| 2026 | Multi-task specialized expert model for hierarchical aspect-based sentiment analysis in consumer healthcare
Jiaxuan Li 0003, Jielong Guo, Patrick Pang 0001, Hugo Gonçalo Oliveira, Benjamin K. Ng, Tao Tan 0002 |
Expert Syst. Appl. | 6 |
| 2026 | BDCNet: Feature-decoupling and cross-task collaboration network with biological priors for cell segmentation and classification
Jinlin Yang, Xintao Pang, Chuan Lin 0003, Tao Tan 0002 |
Expert Syst. Appl. | 4 |
| 2026 | Feature distribution learning based on variance transfer and center shift for long-tailed classification
Chenxi Hong, Qiang Zhao 0005, Tao Tan 0002, Chenggang Yan 0001 |
Neurocomputing | 4 |
| 2026 | Anatomy-guided prompting with cross-modal self-alignment for whole-body PET-CT breast cancer segmentation
Jiaju Huang, Xinglong Liang, Shaobin Chen, Yue Sun 0001, Greta S. P. Mok, Shuo Li 0001, Tao Tan 0002 |
Medical Image Anal. | 9 |
| 2026 | Leveraging modality-guided pre-training for dual-prompt-driven multi-cancer PET-CT segmentation
Xinglong Liang, Jiaju Huang, Tianyu Zhang 0006, Luyi Han, Xin Wang 0121, Chunyao Lu, Yue Sun 0001, Jonas Teuwen, Tao Tan 0002, Ritse Mann |
Medical Image Anal. | 10 |
| 2026 | S2DENet: Shallow suppression and deep enhancement network for general ultrasound image segmentation
Xintao Pang, Jinlin Yang, Zhifan Gao, Chuan Lin 0003, Yue Sun 0001, Shuo Li 0001, Peter H. N. de With, Tao Tan 0002 |
Medical Image Anal. | 8 |
| 2026 | A hypergraph-based model for tumor prognosis using local and global information fusion on H&E-stained histology images
Yanfen Cui, Zhenhui Li, Xiuming Zhang, Su Yao, Dacheng Yang, Zhishun Liu, Shiwei Luo, Guangjun Yang, Lixu Yan, Xiangtian Zhao, Yingqiu Huo, Jiahui Ma, Wenfeng He, Tao Tan 0002, Anant Madabhushi, Jinglei Tang, Zaiyi Liu, Cheng Lu 0001 |
Medical Image Anal. | 23 |
| 2026 | Incorporating global-local tissue changes to predict future breast cancer from longitudinal screening mammograms
Xin Wang 0121, Tao Tan 0002, Eric Marcus, Chunyao Lu, Luyi Han, Antonio Portaluri, Ruisheng Su, Tianyu Zhang 0006, Xinglong Liang, Regina Beets-Tan, Katja Pinker-Domenig, Yue Sun 0001, Ritse Mann, Jonas Teuwen |
Medical Image Anal. | 2 |
| 2026 | From noisy labels to intrinsic structure: A geometric-structural dual-guided framework for noise-robust medical image segmentation
Tao Wang 0085, Zhenxuan Zhang, Yuanbo Zhou, Xinlin Zhang, Yuanbin Chen, Tao Tan 0002, Guang Yang 0006, Tong Tong 0001 |
Medical Image Anal. | 6 |
| 2026 | MOTDNet: Multi organ task decoupling network for cell segmentation
Jinlin Yang, Xintao Pang, Chuan Lin 0003, Tao Tan 0002 |
Medical Image Anal. | 4 |
| 2026 | TF-LLM: Enhanced time series analysis with time-frequency large language models
Yuhang Zhang 0034, Zitong Yu, Mingtong Dai, Yue Sun 0001, Tao Tan 0002 |
Neural Networks | 5 |
| 2026 | Event-aware temporal modeling and semantic alignment for long-form video question answering
Xichun Sheng, Haibo Gong, Liang Li 0003, Chenggang Yan 0001, Tao Tan 0002 |
Pattern Recognit. | 7 |
| 2026 | 3D-IRMM-Net: An Interpretable Radiologist-Mimicking 3D Multimodal Network for Lesion Detection in ABUS With LLM-Based Guidance and Uncertainty QuantificationabstractAutomated breast ultrasound (ABUS) has emerged as a promising tool for breast lesion detection, but most existing deep learning models for ABUS lack transparency and fail to align with clinical reasoning processes. This limits their interpretability and hinders clinical adoption. Therefore, we propose 3D-IRMM-Net, a radiologist-mimicking 3D multimodal network. It integrates visual, textual, and semantic cues to emulate hierarchical diagnostic reasoning. At its core, the network integrates two novel components: the 3D Scale-Expert Convolution (3D-SEC) block, which enables parameter-efficient, topology-aware feature routing via a shared graph convolutional layer; and the Adaptive Spatial-Gaussian (ASG) block, which enhances lesion localization in ABUS by modeling spatial dependencies through multi-scale Gaussian attention. In addition, we propose LLM-Vision activator, a natural-language-driven mechanism that uses language to guide 3D visual attention activation and enables 3D spatial-perceptual feature learning via a pretrained 2D VLM, boosting training efficiency. By aligning AI inference with clinical workflows, 3D-IRMM-Net enhances both detection rates and interpretability. Extensive experiments on ABUS2025, TDSC-ABUS2023, and FPHXD-ABUS datasets confirm its state-of-the-art performance across multiple metrics. At an FPPI of 0.5, our method achieved a patient-level detection rates of 89.7±1.1% (internal) and 85.1±2.0% (external); at an FPPI of 4, they rose to 96.1±1.6% and 91.5 ±1.9%, respectively. Furthermore, the model provides principled uncertainty quantification, improving trustworthiness in real-world deployment. These results confirm the effectiveness of 3D-IRMM-Net in bridging the gap between AI-driven detection and clinical practice. Xin Qian 0001, Wei Ke 0001, Luoxi Zhu, Yue Sun 0001, Zhang Xiong 0001, Lingyun Bao, Tao Tan 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | Anatomy-Aware MR-Imaging-Only RadiotherapyabstractThe synthesis of computed tomography images can supplement electron density information and eliminate MR-CT image registration errors. Consequently, an increasing number of MR-to-CT image translation approaches are being proposed for MR-only radiotherapy planning. However, due to substantial anatomical differences between various regions, traditional approaches often require each model to undergo independent development and use. In this paper, we propose a unified model driven by prompts that dynamically adapt to the different anatomical regions and generates CT images with high structural consistency. Specifically, it utilizes a region-specific attention mechanism, including a region-aware vector and a dynamic gating factor, to achieve MRI-to-CT image translation for multiple anatomical regions. Qualitative and quantitative results on three datasets of anatomical parts demonstrate that our models generate clearer and more anatomically detailed CT images than other state-of-the-art translation models. The results of the dosimetric analysis also indicate that our proposed model generates images with dose distributions more closely aligned to those of the real CT images. Thus, the proposed model demonstrates promising potential for enabling MR-only radiotherapy across multiple anatomical regions. we have released the source code for our RSAM model. The repository is accessible to the public at: https://github.com/yhyumi123/RSAM. Hao Yang 0026, Yue Sun 0001, Chi Kin Lam, Qiang Zhao 0005, Xiangyu Xiong, Kunyan Cai, Behdad Dashtbozorg, Chenggang Yan 0001, Tao Tan 0002 |
IEEE Trans. Image Process. | 11 |
| 2026 | UniEmo: Unifying Emotional Understanding and Generation With Learnable Expert QueriesabstractEmotional understanding and generation are often treated as separate tasks, yet they are inherently complementary and can mutually enhance each other. In this paper, we propose the UniEmo, a unified framework that seamlessly integrates these two tasks. The key challenge lies in the abstract nature of emotions, necessitating the extraction of visual representations beneficial for both tasks. To address this, we propose a hierarchical emotional understanding chain with learnable expert queries that progressively extracts multi-scale emotional features, thereby serving as a foundational step for unification. Simultaneously, we fuse these expert queries and emotional representations to guide the diffusion model in generating emotion-evoking images. To enhance the diversity and fidelity of the generated emotional images, we further introduce the emotional correlation coefficient and emotional condition loss into the fusion process. This step facilitates fusion and alignment for emotional generation guided by the understanding. In turn, we demonstrate that joint training allows the generation component to provide implicit feedback to the understanding part. Furthermore, we propose a novel data filtering algorithm to select high-quality and diverse emotional images generated by the well-trained model, which explicitly feedback into the understanding part. Together, these generation-driven dual feedback processes enhance the model's understanding capacity. Extensive experiments show that UniEmo significantly outperforms state-of-the-art methods in both emotional understanding and generation tasks. The code for the proposed method is available at https://github.com/JiuTian-VL/UniEmo. Lingsen Zhang, Zitong Yu, Rui Shao 0001, Tao Tan 0002, Liqiang Nie |
IEEE Trans. Image Process. | 5 |
| 2026 | BRPDNet: A BioRegion Prompt Distillation Network for Physiological MonitoringabstractPhysiological signal extraction from video data is challenging in dynamic and occluded environments, requiring both accuracy and real-time performance. Existing methods struggle to balance accuracy with model efficiency, particularly under partial facial occlusion or redundant signals. We propose BRPDNet, a novel framework for efficient physiological signal extraction which includes a BioRegion Prompt module for adaptive convolution and a Hyper Distillation module to reduce signal redundancy, ensuring high accuracy and robustness, especially in dynamic and occluded environments. Additionally, the teacher-student network structure enhances the model's adaptability to occlusions and reduces computational complexity without relying on explicit segmentation. Experimental results show that BRPDNet outperforms state-of-the-art models in accuracy, robustness, and efficiency across multiple datasets. For instance, BRPDNet achieves an Mean Absolute Error (MAE) of 1.55 beats per minute (bpm) and a Pearson Correlation Coefficient (PCC) of 0.76 on PURE and UBFC-rPPG datasets with fewer parameters than existing models, ensuring efficient real-time performance. Zhengxuan Chen, Bin Huang 0014, Kangyang Cao, Tao Tan 0002, Bingsheng Huang, Chan-Tong Lam, Yue Sun 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | HRMamba: Fusing Luminance Information for Remote Physiological Measurement in Varied Lighting ConditionsabstractCamera-based photoplethysmography (cbPPG) represents a non-invasive technique for capturing physiological parameters through facial videos, enabling the extraction of vital signs such as heart rate, respiration rate, and blood oxygen saturation without direct physical contact. Existing deep learning methods face two core challenges when dealing with cbPPG: firstly, extracting weak PPG signals from video segments with large spatial and temporal redundancy and understanding their periodic patterns in long contexts; secondly, accurately extracting PPG signals in complex lighting environments, especially in low-light conditions. To address these issues, this paper proposes an end-to-end method based on Mamba, named HRMamba. This method employs temporal difference mamba to process temporal signals and combines bidirectional state space to enable Mamba to robustly understand the scene and learn the periodic patterns of PPG. Furthermore, a luminance post-processing module is designed to extract luminance information from the video without enhancing lighting or altering the original video data, and embed it into the PPG signal. Experimental results demonstrate that HRMamba achieves state-of-the-art performance, and the designed luminance post-processing module can be applied in various lighting environments, significantly enhancing the performance in dark environments without degrading the performance in normal light scenes. Nuoer Long, Wei Ke 0001, Chan-Tong Lam, Tao Tan 0002, Zitong Yu, Yue Sun 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2026 | Multi-Granularity Query Network With Adaptive Category Feature Embedding for Behavior RecognitionabstractBehavior recognition is a highly challenging task, particularly in scenarios requiring unified recognition across both human and animal subjects. Most existing approaches primarily focus on single-species datasets or rely heavily on prior information such as species labels, positional annotations, or skeletal keypoints, which limits their applicability in real-world scenarios where species labels may be ambiguous or annotations are insufficient. To address these limitations, we propose a query-based Multi-Granularity Behavior Recognition Network that directly mines cross-species shared spatiotemporal behavior patterns from raw video inputs. Specifically, we design a Multi-Granularity Query module to effectively fuse fine-grained and coarse-grained features, thereby enhancing the model's capability in capturing spatiotemporal dynamics at different granularities. Additionally, we introduce a Category Query Decoder that leverages learnable category query vectors to achieve explicit behavior category modeling and mapping. Without relying on any extra annotations, the proposed method achieves unified recognition of multi-species and multi-category behaviors, setting a new state-of-the-art on the Animal Kingdom dataset and demonstrating strong generalization ability on the Charades dataset. Nuoer Long, Yonghao Dang, Chengpeng Xiong, Shaobin Chen, Tao Tan 0002, Wei Ke 0001, Chan-Tong Lam, Jianqin Yin, Peter H. N. de With, Yue Sun 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | SHIELDNet: Multi-Region Fusion and Denoising for Enhanced rPPG Signal Extraction in Healthcare MonitoringabstractAccurate extraction of Remote Photoplethysmography (rPPG) signals from video data is critical for medical applications such as remote patient monitoring. However, the process is hindered by significant challenges, including noise interference, occlusions, and multi bio-region signal processing. To address these, we propose SHIELDNet, an efficient and robust framework for real-time extraction of rPPG signals from multiple anatomical regions, incorporating advanced noise reduction mechanisms. SHIELDNet integrates a novel Differential Attention (DA) module, which adaptively focuses on multiple anatomical regions, enabling the model to effectively handle dynamic real-world conditions. Additionally, the network leverages an advanced Efficient Space Attention Module (ESAM) to enhance spatial feature extraction and multi bio-region signal fusion. BioRegion Prompt Module (BRPM) is further introduced to prioritize region-specific features, reducing the model's dependence on facial features alone. Futhermore, we introduce M-rPPG dataset, a comprehensive multi bio-region reference for BioRegion-based studies with full-body details at higher resolution than existing datasets. Extensive evaluations on multiple public datasets demonstrate significant improvements in Mean Absolute Error (MAE$=\mathbf{5. 2 8} \mathbf{~ b p m}$) and Pearson Correlation Coefficient ($\mathbf{P C C} \boldsymbol{=} \mathbf{0. 8 0}$), outperforming current state-of-the-art models. SHIELDNet provides an effective solution for noncontact, multi bio-region rPPG monitoring. We will release our code upon acceptance. Zhengxuan Chen, Tao Tan 0002, Chan-Tong Lam, Yue Sun 0001 |
BIBM | 4 |
| 2025 | Advancing Clinical Generalization in Remote Photoplethysmography via Age-Specific Physiological FeaturesabstractCamera-based remote photoplethysmography (rPPG) enables contactless monitoring of important physiological signals such as heart rate (HR) and respiratory rate (RR), with transformative clinical potential. Despite recent progress under laboratory conditions, most deep-learning approaches overlook age-specific physiological variations and therefore perform poorly in complex clinical scenarios. In this paper, we propose a novel rPPG framework that integrates both explicit and implicit age-specific physiological feature adaptation mechanisms. First, a dynamic channel weighting (DCW) module that enriches the rPPG channel features according to age-related skin optical properties. Second, we propose a prompt-embedded hyper-convolutional (P-HC) layer dynamically adjusts the age-aware convolution kernels using explicit domain priors. Third, a factorized spatio-temporal attention (FAST) mechanism implicitly decomposes spatio-temporal features into age-invariant physiological bases and age-dependent modulation coefficients, thereby enhancing cross-age generalizability. Extensive experiments on five public datasets and clinical validation on 25 preterm infants in neonatal intensive care units demonstrate that the proposed method outperforms state-of-the-art approaches. These findings validate the method's clinical potential for demoaraphically inclusive monitoring. Ieong Weng San, Xiangmin Luo, Zhengxuan Chen, Tao Tan 0002, Yue Sun 0001 |
BIBM | 5 |
| 2025 | Tree-Diffusion: Octree-Based Conditional Diffusion Model for Small Bowel Skeleton Generation with Geometric Direction ModelingabstractAccurate 3D reconstruction of the small bowel skeleton is vital for understanding intestinal morphology, de-tecting structural abnormalities, and supporting diagnosis, yet limited resolution, organ adhesion, complex anatomy, and scarce annotations make continuous skeleton extraction from masks challenging. Voxel-based methods often struggle with the sparse topology and geometric directionality inherent in the small bowel skeleton, leading to inefficiency and high memory cost. To address these limitations, we propose a novel octree-based conditional diffusion model (i.e., Tree-Diffusion) that generates anatomically consistent small bowel skeletons guided by 3D segmentation masks. Specifically, we introduce two modules that captures structural priors from masks and topology characteristics from skeletons, ensuring cross-domain alignment and high-quality skeleton generation. Besides, we design a synthesis strategy to generate anatomically plausible skeleton-mask pairs, serving as topological priors to guide the diffusion model toward realis-tic structure predictions. To efficiently represent the elongated skeleton, we adopt an octree- based spatial encoding of hierarchical geometric features. Compared with baselines, our model achieves superior performance in anatomical fidelity, directional consistency, and inference efficiency. The code is available at: https://github.com/Small-Bowel-Skeleton-GenerationlCode Zhichao Liang, Dengqiang Jia, Yaofei Duan, Xinyu Xie, Kaicong Sun, Zhiming Cui 0001, Tao Tan 0002, Dinggang Shen |
BIBM | 8 |
| 2025 | Grad-MTSeg: Mitigating Multi-Task Gradient Conflicts via Hierarchical Gradient Optimization for NPC Radiotherapy DelineationabstractThe precise delineation of nasopharyngeal carcinoma (NPC) is a critical prerequisite for radiation therapy, but manual methods are inefficient and inconsistent. Current automated segmentation techniques are challenged by the complexity of multi-modal inputs and multi-target outputs, including Organs at Risk (OARs), Gross Tumor Volume of the Primary Tumor (GTVp), Gross Tumor Volume of the Nodal Metastases (GTVn), clinical target volume prescribed with 70 Gy (CTV70), and clinical target volume prescribed with 63 Gy (CTV63). This process is frequently hindered by gradient conflicts during multitask optimization. To resolve these issues, we propose Grad-MTSeg, a novel deep learning framework. Our approach introduces two core innovations: Unidirectional Anatomic Guidance (UAG) to leverage CT structural priors for improved MRI-based segmentation, and Hierarchical Gradient Optimization (HGO) to alleviate destructive gradient interference among tasks. Our framework improves segmentation accuracy for relevant NPC tasks by effectively resolving conflicts across OARs, GTVs, and CTVs. Validation on three external datasets confirms that Grad-MTSeg provides an efficient and precise solution for complex multimodal segmentation, advancing the automation of NPC radiotherapy planning. Junqiang Ma, Luyi Han, Dengqiang Jia, Tao Tan 0002, Henry H. Y. Tong, Anne W. M. Lee, Sung Inda Soong, Yue Sun 0001 |
BIBM | 5 |
| 2025 | Spatiotemporal Uncertainty-Aware Mamba-Transformer Synergy: Breast Cancer Detection in ABUSabstractBreast cancer significantly affects women's health, making early and accurate diagnosis through ultrasound examination essential. However, automated breast ultrasound(ABUS) lesion detection models encounter challenges, including uncertain noise and difficulties in locating small lesions. This paper proposes SUA-MT, a multi-video object detection network using uncertainty-aware transformers and temporal-spatial mamba. To address the performance degradation caused by low-quality frames and noise interference in input data, we propose the uncertainty-gated transformer decoder (UGTD), which dynamically adjusts attention weights to focus on high-confidence regions while suppressing attention to redundant areas. The spatialtemporal(ST) mamba module is designed to model the long-term dependencies and 3D spatial features of different plane video frames, making full use of temporal and spatial information. In general, SUA-MT not only supports dynamic video length input, but also combines multidimensional video modeling of lesions (transverse, sagittal and coronal plane), making full use of the complementarity of multi-view information and spatiotemporal information to enhance the ability to locate small lesions. Experimental results on a combined set of internal and publicly datasets demonstrate that the SUA-MT method achieves state-of-the-art performance compared to existing video detection approaches. Specifically, SUA-MT attains a mean precision of 80.1 % and mean recall of 84.2 %, providing an efficient and robust solution for lesion detection in ABUS. Xin Qian 0001, Luoxi Zhu, Zhang Xiong 0001, Wei Ke 0001, Lingyun Bao, Tao Tan 0002 |
BIBM | 6 |
| 2025 | UA-MAE: An Uncertainty-Aware Masked Autoencoder for Breast Lesion Segmentation in Ultrasound ImagesabstractAccurate segmentation of breast lesions is vital for diagnosing breast diseases. Masked image modeling (MIM) with random masking performs well in self-supervised learning but struggles in breast ultrasound segmentation due to (1) ambiguous representations from similar intensities near lesion boundaries and (2) a bias toward irrelevant regions. We propose UA-MAE, an uncertainty-aware masked autoencoder that uses pixel-wise uncertainty maps to dynamically select masking patches, prioritizing boundaries and morphologically relevant lesion areas. Experiments on two public datasets for pre-training and three for fine-tuning show UA-MAE outperforming four state-of-theart SSL methods and two supervised approaches in segmentation accuracy across diverse breast ultrasound images. The code is available at https://github.com/yXiangXiong/UA-MAE. Xiangyu Xiong, Yue Sun 0001, Jiaju Huang, Da Huang 0004, Shaobin Chen, Zhuoneng Zhang, Tao Tan 0002 |
BIBM | 7 |
| 2025 | Vision-Language Semantic Guidance for Ejection Fraction Assessment in EchocardiographyabstractEjection fraction (EF) is a key indicator of cardiac function, crucial for diagnosing heart failure and guiding treatment. Its estimation from echocardiography is challenged by morphological changes across cardiac phases and low-quality, noisy boundaries. We propose EFusionNet, a multimodal segmentation framework that integrates echocardiographic images with structured diagnostic text to enhance segmentation and EF assessment. Clinical phrases (e.g., “irregular boundary”) are embedded via a domain-specific language model into both input fusion and UNet skip connections, enabling phase-aware feature calibration. A feature fusion enhancement module (FFEM) refines spatial localization, while a multi-objective loss enforces uncertainty learning and semantic consistency. Evaluated on CAMUS and EchoNet-Dynamic datasets, EFusionNet achieves Dice scores of 91.2%/90.1% and EFMAEof 4.8/5.0, outperforming baselines and improving reliable, interpretable EF estimation. Dashun Zheng, Patrick Pang 0001, Jiaxuan Li 0003, Edmundo Patricio Lopes Lao, Yapeng Wang 0001, Zhifan Gao, Tao Tan 0002 |
BIBM | 9 |
| 2025 | MoEdit: On Learning Quantity Perception for Multi-object Image EditingabstractMulti-object images are prevalent in various real-world scenarios, including augmented reality, advertisement design, and medical imaging. Efficient and precise editing of these images is critical for these applications. With the advent of Stable Diffusion (SD), high-quality image generation and editing have entered a new era. However, existing methods often struggle to consider each object both individually and part of the whole image editing, both of which are crucial for ensuring consistent quantity perception, resulting in suboptimal perceptual performance. To address these challenges, we propose MoEdit, an auxiliaryfree multi-object image editing framework. MoEdit facilitates high-quality multi-object image editing in terms of style transfer, object reinvention, and background regeneration, while ensuring consistent quantity perception between inputs and outputs, even with a large number of objects. To achieve this, we introduce the Feature Compensation (FeCom) module, which ensures the distinction and separability of each object attribute by minimizing the in-between interlacing. Additionally, we present the Quantity Attention (QTTN) module, which perceives and preserves quantity consistency by effective control in editing, without relying on auxiliary tools. By leveraging the SD model, MoEdit enables customized preservation and modification of specific concepts in inputs with high quality. Experimental results demonstrate that our MoEdit achieves State-Of-The-Art (SOTA) performance in multi-object image editing. Data and codes are available at https://github.com/Tear-kitty/MoEdit. Ka-Hou Chan, Yue Sun 0001, Chan-Tong Lam, Tong Tong 0001, Zitong Yu, Keren Fu, Xiaohong Liu 0001, Tao Tan 0002 |
CVPR | 9 |
| 2025 | ONDA-Pose: Occlusion-Aware Neural Domain Adaptation for Self-Supervised 6D Object Pose EstimationabstractSelf-supervised 6D object pose estimation has received increasing attention in computer vision recently. Some typical works in literature attempt to translate the synthetic images with object pose labels generated by object CAD models into the real domain, and then use the translated data for training. However, their performance is generally limited, since (i) there still exists a domain gap between the translated images and the real images and (ii) the translated images can not sufficiently reflect occlusions that exist in many real images. To address these problems, we propose an Occlusion-Aware Neural Domain Adaptation method for self-supervised 6D object Pose estimation, called ONDA-Pose. The proposed method comprises three main steps. Firstly, by utilizing both the training real images without pose labels and a CAD model, we explore a CAD-like radiance field for rendering corresponding synthetic images that have similar textures to those generated by the CAD model. Then, a backbone pose estimator trained on the synthetic data is employed to provide initial pose estimations for the synthetic images rendered from the CAD-like radiance field, and the initial object poses are refined by a global object pose refiner to generate pseudo object pose labels. Finally, the backbone pose estimator is further self-supervised as the final pose estimator by jointly utilizing the real images with pseudo object pose labels and the synthetic images rendered from the CAD-like radiance field. Experimental results on three public datasets demonstrate that ONDA-Pose significantly outperforms the comparative state-of-the-art methods in most cases. Tao Tan 0002, Qiulei Dong |
CVPR | 1 |
| 2025 | Contrastive Learning via Randomly Generated Deep SupervisionabstractUnsupervised visual representation learning has gained significant attention in the computer vision community, driven by recent advancements in contrastive learning. Most existing contrastive learning frameworks rely on instance discrimination as a pretext task, treating each instance as a distinct category. However, this often leads to intra-class collision in a large latent space, compromising the quality of learned representations. To address this issue, we propose a novel contrastive learning method that utilizes randomly generated supervision signals. Our framework incorporates two projection heads: one handles conventional classification tasks, while the other employs a random algorithm to generate fixed-length vectors representing different classes. The second head executes a supervised contrastive learning task based on these vectors, effectively clustering instances of the same class and increasing the separation between different classes. Our method, Contrastive Learning via Randomly Generated Supervision(CLRGS), significantly improves the quality of feature representations across various datasets and achieves state-of-the-art performance in contrastive learning tasks. Zili Ma, Ka-Hou Chan, Yue Liu 0001, Tong Tong 0001, Qinquan Gao, Guangtao Zhai, Xiaohong Liu 0001, Tao Tan 0002 |
ICASSP | 9 |
| 2025 | AdaMHF: Adaptive Multimodal Hierarchical Fusion for Survival PredictionabstractThe integration of pathologic images and genomic data for survival analysis has gained increasing attention with advances in multimodal learning. However, current methods often ignore biological characteristics, such as heterogeneity and sparsity, both within and across modalities, ultimately limiting their adaptability to clinical practice. To address these challenges, we propose AdaMHF: Adaptive Multimodal Hierarchical Fusion, a framework designed for efficient, comprehensive, and tailored feature extraction and fusion. AdaMHF is specifically adapted to the uniqueness of medical data, enabling accurate predictions with minimal resource consumption, even under challenging scenarios with missing modalities. Initially, AdaMHF employs an experts expansion and residual structure to activate specialized experts for extracting heterogeneous and sparse features. Extracted tokens undergo refinement via selection and aggregation, reducing the weight of non-dominant features while preserving comprehensive information. Subsequently, the encoded features are hierarchically fused, allowing multi-grained interactions across modalities to be captured. Furthermore, we introduce a survival prediction benchmark designed to resolve scenarios with missing modalities, mirroring real-world clinical conditions. Extensive experiments on TCGA datasets demonstrate that AdaMHF surpasses current state-of-the-art (SOTA) methods, showcasing exceptional performance in both complete and incomplete modality settings. Code is available in AdaMHF. Shuaiyu Zhang, Xun Lin, Rongxiang Zhang, Yong Xu 0001, Tao Tan 0002, Xubin Zheng, Zitong Yu |
ICME | 6 |
| 2025 | Region-Based Text-Consistent Augmentation for Multimodal Medical Segmentation
Kunyan Cai, Chenggang Yan 0001, Liangqiong Qu, Shuai Wang 0003, Tao Tan 0002 |
MICCAI (3) | 6 |
| 2025 | SABPI-Net: A Novel Structure-Aware Network for Accurate and Domain-Invariant Retinopathy of Prematurity Diagnosis
Shaobin Chen, Huazhu Fu, Tao Tan 0002, Jiaju Huang, Xiangyu Xiong, Zhenquan Wu, Behdad Dashtbozorg, Bai Ying Lei, Yue Sun 0001 |
MICCAI (10) | 4 |
| 2025 | ADAptation: Reconstruction-Based Unsupervised Active Learning for Breast Ultrasound Diagnosis
Yaofei Duan, Yuhao Huang 0001, Xin Yang 0009, Luyi Han, Xinyu Xie, Ka-Hou Chan, Ligang Cui, Sio Kei Im, Dong Ni 0001, Tao Tan 0002 |
MICCAI (16) | 12 |
| 2025 | C2MAOT: Cross-modal Complementary Masked Autoencoder with Optimal Transport for Cancer Segmentation in PET-CT Images
Jiaju Huang, Shaobin Chen, Xinglong Liang, Zhuoneng Zhang, Yue Sun 0001, Tao Tan 0002 |
MICCAI (1) | 8 |
| 2025 | DpDNet: An Dual-Prompt-Driven Network for Universal PET-CT Segmentation
Xinglong Liang, Jiaju Huang, Luyi Han, Tianyu Zhang 0006, Xin Wang 0121, Chunyao Lu, Lishan Cai, Tao Tan 0002, Ritse Mann |
MICCAI (6) | 9 |
| 2025 | BiMSRec: A Progressive Image Reconstruction Framework for Medical Image Fusion Guided by Multi-scale Deformation Fields
Nuoer Long, Xinyu Xie, Zitong Yu, Tao Tan 0002, Yue Sun 0001 |
MICCAI (2) | 5 |
| 2025 | TRRG: Towards Truthful Radiology Report Generation With Cross-Modal Disease Clue Enhanced Large Language Models
Yue Sun 0001, Tao Tan 0002, Chao Hao, Yawen Cui, Xinqi Su, Weicheng Xie 0001, LinLin Shen, Zitong Yu |
MICCAI (7) | 3 |
| 2025 | Synergy-Guided Regional Supervision of Pseudo Labels for Semi-supervised Medical Image Segmentation
Tao Wang 0085, Xinlin Zhang, Yuanbin Chen, Yuanbo Zhou, Longxuan Zhao, Tao Tan 0002, Tong Tong 0001 |
MICCAI (8) | 6 |
| 2025 | Tumor Segmentation with Heterogeneity Clustering in Non-Contrast Breast MRI
Xinyu Xie, Luyi Han, Yonghao Li, Yaofei Duan, Yue Sun 0001, Muzhen He, Tao Tan 0002, Dinggang Shen |
MICCAI (2) | 7 |
| 2025 | FDF-VQVAE: A Frequency Disentanglement and Fusion Learning Framework for Multi-sequence MRI Enhancement
Xinghe Xie, Luyi Han, Yue Sun 0001, Chi Kin Lam, Jian Zheng 0001, Tong Tong 0001, Wei Ke 0001, Chan-Tong Lam, Tao Tan 0002 |
MICCAI (3) | 9 |
| 2025 | SAMASK-CLTR: A Spatial-Aware Mask Guided Learning Model for Benign and Malignant Tumor Classification in ABUS
Peirong Xu, Luoqian Zhu, Jingkun Chen, Xin Qian 0001, Yue Sun 0001, Lingyun Bao, Tao Tan 0002 |
MICCAI (1) | 7 |
| 2025 | MedIQA: A Scalable Foundation Model for Prompt-Driven Medical Image Quality Assessment
Siyi Xun, Yue Sun 0001, Jingkun Chen, Zitong Yu, Tong Tong 0001, Xiaohong Liu 0001, Mingxiang Wu, Tao Tan 0002 |
MICCAI (13) | 8 |
| 2025 | UltraTwin: Towards Cardiac Anatomical Twin Generation from Multi-view 2D Ultrasound
Junxuan Yu, Yaofei Duan, Yuhao Huang 0001, Rongbo Ling, Weihao Luo, Jingxian Xu, Qiongying Ni, Yongsong Zhou, Binghan Li, Haoran Dou, Yanfen Chu, Feng Geng, Zhe Sheng, Zhifeng Ding, Yuhang Zhang 0034, Tao Tan 0002, Dong Ni 0001, Zhongshan Gou, Xin Yang 0009 |
MICCAI (16) | 22 |
| 2025 | Paired Image Generation with Diffusion-Guided Diffusion Models
Haoxuan Zhang, Wenju Cui, Yuzhu Cao, Tao Tan 0002, Yunsong Peng, Jian Zheng 0001 |
MICCAI (4) | 4 |
| 2025 | RefineNet: Elevating Medical Foundation Models Through Quality-Centric Data Curation by MLLM-Annotated Proxy Distillation
Ningyi Zhang, Xin Wang 0121, Ka-Hou Chan, Jian Wu 0033, Chan-Tong Lam, Shanshan Wang 0010, Yue Sun 0001, Sio Kei Im, Tao Tan 0002 |
MICCAI (11) | 10 |
| 2025 | A Diffusion-Driven Temporal Super-Resolution and Spatial Consistency Enhancement Framework for 4D MRI imaging
Xuanru Zhou, Jiarun Liu, Shoujun Yu, Hao Yang 0026, Cheng Li 0008, Tao Tan 0002, Shanshan Wang 0002 |
MICCAI (10) | 6 |
| 2025 | An Efficient and Compact Network for Simultaneous Multi-Object Tracking and Behavior Monitoring in Pigeon FarmingabstractAnimal monitoring plays a crucial role in agriculture, especially in improving animal welfare and farming efficiency. However, existing methods cannot simultaneously achieve video-based tracking for each individual while monitoring their behaviors. To overcome this limitation and meet the demand for automated surveillance in large-scale pigeon farming, this study proposes an efficient and compact network for simultaneous multi-object tracking and behavior monitoring in pigeon farming. The network is capable of tracking each pigeon while recognizing their behaviors in a video sequence. The network combines the DETR detector of the lightweight RepViT backbone with the BoT-SORT motion tracker, which tracks accurately while saving costs. A text-guided multimodal action recognition model is introduced in the action recognition stage, which combines spatiotemporal video features with semantic text embedding to enhance classification accuracy. Experimental results show that the proposed method achieves the optimal tracking performance (IDF1: 96.58%) and action recognition accuracy (Top-1: 94.89%), while reducing the network complexity (13.6M parameters) and computational cost (45.67 GFLOPs). This method provides effective technical support to promote accurate management of poultry farming. Jiefeng Xie, Tao Tan 0002, Yaoji Liu, Dachun Feng, Yue Sun 0001 |
SMC | 2 |
| 2025 | MambaControl: Anatomy Graph-Enhanced Mamba ControlNet with Fourier Refinement for Diffusion-Based Disease Trajectory PredictionabstractModelling disease progression in precision medicine requires capturing complex spatio-temporal dynamics while preserving anatomical integrity. Existing methods often struggle with longitudinal dependencies and structural consistency in progressive disorders. To address these limitations, we introduce MambaControl, a novel framework that integrates selective state-space modelling with diffusion processes for high-fidelity prediction of medical image trajectories. To better capture subtle structural changes over time while maintaining anatomical consistency, MambaControl combines Mamba-Based long-range modelling with graph- guided anatomical control to more effectively represent anatomical correlations. Furthermore, we introduce Fourier-Enhanced spectral graph representations to capture spatial coherence and multiscale detail, enabling MambaControl to achieve state-of-the-art performance in Alzheimer’s disease prediction. Quantitative and regional evaluations demonstrate improved progression prediction quality and anatomical fidelity, highlighting its potential for personalised prognosis and clinical decision support. Hao Yang 0026, Tao Tan 0002, Weiqin Yang 0003, Kunyan Cai, Calvin Chen, Yue Sun 0001 |
SMC | 2 |
| 2025 | Joint segmentation of retinal layers and fluid lesions in optical coherence tomography with cross-dataset learning
Xiayu Xu, Hualin Wang, Yulei Lu, Hanze Zhang, Tao Tan 0002, Jianqin Lei |
Artif. Intell. Medicine | 5 |
| 2025 | A universal parameter-efficient fine-tuning approach for stereo image super-resolution
Yuanbo Zhou, Yuyang Xue, Xinlin Zhang, Tao Wang 0085, Tao Tan 0002, Qinquan Gao, Tong Tong 0001 |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | FedHNR: Federated hierarchical resilient learning for echocardiogram segmentation with annotation noise
Wanli Ding, Weiyuan Lin, Tao Tan 0002, Zhifan Gao |
Expert Syst. Appl. | 4 |
| 2025 | Dynamic mask stitching-guided region consistency for semi-supervised 3D medical image segmentation
Dongsheng Ruan, Yang Li 0097, Tao Tan 0002, Lianming Wu, Guang Yang 0006, Mingfeng Jiang |
Expert Syst. Appl. | 4 |
| 2025 | DiffSteISR: Harnessing diffusion prior for superior real-world stereo image super-resolution
Yuanbo Zhou, Xinlin Zhang, Tao Wang 0085, Tao Tan 0002, Qinquan Gao, Tong Tong 0001 |
Neurocomputing | 5 |
| 2025 | An Optic Nerve Segmentation Model Based on Fully-Convolutional-Based Masked Autoencoders and Direction FieldabstractUltrasound measurement of optic nerve sheath diameter (ONSD) is considered a noninvasive method for estimating elevated intracranial pressure (ICP) in patients. Clinical trials have demonstrated a strong correlation between changes in ONSD and changes in ICP. Therefore, accurate segmentation of the ONSD is crucial for noninvasive ICP assessment. In this paper, we propose a two-stage self-supervised semantic segmentation method to enhance optic nerve segmentation. In the pre-training phase, we use a fully convolutional-based masked autoencoder (FCMAE) to reconstruct full images from partially masked inputs. The encoder of FCMAE aggregates contextual information to infer the masked image regions, and this pretrained encoder is then migrated to the segmentation task for parameter initialization. In the fine-tuning phase, we perform the optic nerve segmentation task. After obtaining the initial segmentation results through the UPerNet network, we use a direction field (DF) module to compute a vector of DFs pointing to the nearest edge of the optic nerve for each pixel. This DF information is then used to refine the initial segmentation results via the feature correction module. The model was trained on a dataset of optic nerve sheath images collected from hospital patients and achieved a Dice score of 98.03%. Our proposed method exhibits superior performance across all metrics compared to other segmentation models. Mingfeng Jiang, Q. Huang, Xin Huang 0030, Jucheng Zhang, Chunshuang Wu, T. Huang, Ling Xia 0001, Tao Tan 0002, Y. Chu |
Int. J. Pattern Recognit. Artif. Intell. | 8 |
| 2025 | Bi-branch bidirectional coupled interaction fusion network for multi-retinal diseases diagnosis
Shaobin Chen, Tao Tan 0002, Wei Ke 0001, Xiayu Xu, Yanwu Xu 0004, Chan-Tong Lam, Yue Sun 0001 |
Knowl. Based Syst. | 2 |
| 2025 | FetalFlex: Anatomy-guided diffusion model for flexible control on fetal ultrasound image synthesis
Yaofei Duan, Tao Tan 0002, Yuhao Huang 0001, Yuanji Zhang, Patrick Pang 0001, Xinru Gao, Guowei Tao, Xiang Cong, Lianying Liang, Guangzhi He, Linliang Yin, Xuedong Deng, Xin Yang 0009, Dong Ni 0001 |
Medical Image Anal. | 2 |
| 2025 | An orchestration learning framework for ultrasound imaging: Prompt-Guided Hyper-Perception and Attention-Matching Downstream Synchronization
Shuo Li 0001, Shanshan Wang 0010, Zhifan Gao, Yue Sun 0001, Chan-Tong Lam, Xindi Hu, Xin Yang 0009, Dong Ni 0001, Tao Tan 0002 |
Medical Image Anal. | 10 |
| 2025 | CVFSNet: A Cross View Fusion Scoring Network for end-to-end mTICI scoring
Weijin Xu, Tao Tan 0002, Wentao Liu 0004, Xipeng Pan, Yiming Deng, Theo van Walsum, Matthijs van der Sluijs, Ruisheng Su |
Medical Image Anal. | 2 |
| 2025 | Data augmentation strategies for semi-supervised medical image segmentation
Dongsheng Ruan, Yang Li 0097, Yongquan Wu, Tao Tan 0002, Guang Yang 0006, Mingfeng Jiang |
Pattern Recognit. | 6 |
| 2025 | BLENet: A Bio-Inspired Lightweight and Efficient Network for Left Ventricle Segmentation in Echocardiography
Xintao Pang, Fengjuan Yao, Yue Sun 0001, Edmundo Patricio Lopes Lao, Chuan Lin 0003, Patrick Pang 0001, Wei Wang 0181, Zhifan Gao, Tao Tan 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 11 |
| 2025 | 3MT-Net: A Multi-Modal Multi-Task Model for Breast Cancer and Pathological Subtype Classification Based on a Multicenter StudyabstractBreast cancer poses a significant threat to women's health, and ultrasound plays a critical role in the assessment of breast lesions. This study introduces a prospective deep learning architecture, termed the "Multi-modal Multi-task Network" (3MT-Net), which integrates clinical data with B-mode and color Doppler ultrasound images. Specifically, an AM-CapsNet is employed to extract key features from ultrasound images, while a cascaded cross-attention mechanism is utilized to fuse clinical data. Moreover, an ensemble learning approach with an optimization algorithm is adopted to dynamically assign weights to different modalities, accommodating both high-dimensional and low-dimensional data. The 3MT-Net performs binary classification of benign versus malignant lesions and further classifies the pathological subtypes. Data were retrospectively collected from nine medical centers to ensure the broad applicability of the 3MT-Net. Two separate testsets were created and extensive experiments were conducted. Comparative analyses demonstrated that the AUC of the 3MT-Net outperforms the industry-standard computer-aided detection product, S-Detect, by 1.4% to 3.8%. Yaofei Duan, Patrick Pang 0001, Rongsheng Wang 0004, Yue Sun 0001, Chuntao Liu, Xirong Yuan, Pengjie Song, Chan-Tong Lam, Ligang Cui, Tao Tan 0002 |
IEEE J. Biomed. Health Informatics | 12 |
| 2025 | Multi-Modal Longitudinal Representation Learning for Predicting Neoadjuvant Therapy Response in Breast Cancer TreatmentabstractLongitudinal medical imaging is crucial for monitoring neoadjuvant therapy (NAT) response in clinical practice. However, mainstream artificial intelligence (AI) methods for disease monitoring commonly rely on extensive segmentation labels to evaluate lesion progression. While self-supervised vision-language (VL) learning efficiently captures medical knowledge from radiology reports, existing methods focus on single time points, missing opportunities to leverage temporal self-supervision for disease progression tracking. In addition, extracting dynamic progression from longitudinal unannotated images with corresponding textual data poses challenges. In this work, we explicitly account for longitudinal NAT examinations and accompanying reports, encompassing scans before NAT and follow-up scans during mid-/post-NAT. We introduce the multi-modal longitudinal representation learning pipeline (MLRL), a temporal foundation model, that employs multi-scale self-supervision scheme, including single-time scale vision-text alignment (VTA) learning and multi-time scale visual/textual progress (TVP/TTP) learning to extract temporal representations from each modality, thereby facilitates the downstream evaluation of tumor progress. Our method is evaluated against several state-of-the-art self-supervised longitudinal learning and multi-modal VL methods. Results from internal and external datasets demonstrate that our approach not only enhances label efficiency across the zero-, few- and full-shot regime experiments but also significantly improves tumor response prediction in diverse treatment scenarios. Furthermore, MLRL enables interpretable visual tracking of progressive areas in temporal examinations, offering insights into longitudinal VL foundation tools and potentially facilitating the temporal clinical decision-making process. Tao Tan 0002, Xin Wang 0121, Regina Beets-Tan, Tianyu Zhang 0006, Luyi Han, Antonio Portaluri, Chunyao Lu, Xinglong Liang, Jonas Teuwen, Ritse Mann |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Guest Editorial: Multi-Modal Joint Learning in Healthcare Imaging
Tao Tan 0002, Yue Sun 0001, Shandong Wu |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | UniMRISegNet: Universal 3D Network for Various Organs and Cancers Segmentation on Multi-Sequence MRIabstractThree-dimensional organ and cancer segmentation based on multi-sequence MRI is crucial for assisting clinical diagnosis. However, current automated segmentation methods often focus on specific sequences, specific organs, and specific cancers, i.e., lack of generality. To address this issue, we propose a universal segmentation network for multi-sequence MRI (UniMRISegNet) that can segment multiple organs and cancers. UniMRISegNet features a shared encoder-decoder architecture equipped with contextual prompt generation (CPG) and prompt-conditioned dynamic convolution (PCDC) modules. The CPG module encodes sequence-specific, position-specific, and organ/cancer-specific text prompts as prior information to inform UniMRISegNet about the specific task to be executed. The PCDC module can adaptively generate model weights based on the assigned prompts, enhancing the segmentation capabilities of the UniMRISegNet for specific tasks. To mitigate discrepancies between different sequences of the same organ and capture similarities between related sequences, we design a novel loss function called Semantic-Aware Cosine Similarity Loss (SACSL), which integrates the cosine similarity of text embeddings to reconcile discrepancies and similarities between MRI sequences of the same organ. We created a large-scale annotated multi-sequence, multi-organ, and multi-cancer segmentation workflow (MSOCS), and demonstrated that our UniMRISegNet outperforms other universal networks and single-task networks on MSOCS. Furthermore, the universal weights from MSOCS can be transferred to never-before-seen downstream tasks, achieving superior performance compared to training from scratch. Zhuoneng Zhang, Luyi Han, Tianyu Zhang 0006, Qinquan Gao, Tong Tong 0001, Yue Sun 0001, Tao Tan 0002 |
IEEE J. Biomed. Health Informatics | 8 |
| 2024 | UniUSNet: A Promptable Framework for Universal Ultrasound Disease Prediction and Tissue SegmentationabstractUltrasound is widely used in clinical practice due to its affordability, portability, and safety. However, current AI research often overlooks combined disease prediction and tissue segmentation. We propose UniUSNet, a universal framework for ultrasound image classification and segmentation. This model handles various ultrasound types, anatomical positions, and input formats, excelling in both segmentation and classification tasks. Trained on a comprehensive dataset with over 9.7K annotations from 7 distinct anatomical positions, our model matches state-of-the-art performance and surpasses single-dataset and ablated models. Zero-shot and fine-tuning experiments show strong generalization and adaptability with minimal fine-tuning. We plan to expand our dataset and refine the prompting mechanism, with model weights and code available at (https://github.com/Zehui-Lin/UniUSNet). Zhuoneng Zhang, Xindi Hu, Zhifan Gao, Xin Yang 0009, Yue Sun 0001, Dong Ni 0001, Tao Tan 0002 |
BIBM | 8 |
| 2024 | A Parameterized Generative Adversarial Network Using Cyclic Projection for Explainable Medical Image ClassificationsabstractAlthough current data augmentation methods are successful to alleviate the data insufficiency, conventional augmentation are primarily intra-domain while advanced generative adversarial networks (GANs) generate images remaining uncertain, particularly in small-scale datasets. In this paper, we propose a parameterized GAN (ParaGAN) that effectively controls the changes of synthetic samples among domains and highlights the attention regions for downstream classification. Specifically, ParaGAN incorporates projection distance parameters in cyclic projection and projects the source images to the decision boundary to obtain the class-difference maps. Our experiments show that ParaGAN can consistently outperform the existing augmentation methods with explainable classification on two small-scale medical datasets. Xiangyu Xiong, Yue Sun 0001, Xiaohong Liu 0001, Chan-Tong Lam, Tong Tong 0001, Hao Chen 0037, Qinquan Gao, Wei Ke 0001, Tao Tan 0002 |
ICASSP | 9 |
| 2024 | Improving Neoadjuvant Therapy Response Prediction by Integrating Longitudinal Mammogram Generation with Cross-Modal Radiological Reports: A Vision-Language Alignment-Guided Model
Xin Wang 0121, Tianyu Zhang 0006, Luyi Han, Chunyao Lu, Xinglong Liang, Jonas Teuwen, Regina Beets-Tan, Tao Tan 0002, Ritse Mann |
MICCAI (1) | 10 |
| 2024 | Non-adversarial Learning: Vector-Quantized Common Latent Space for Multi-sequence MRI
Luyi Han, Tao Tan 0002, Tianyu Zhang 0006, Xin Wang 0121, Chunyao Lu, Xinglong Liang, Haoran Dou, Yunzhi Huang, Ritse Mann |
MICCAI (11) | 2 |
| 2024 | Ordinal Learning: Longitudinal Attention Alignment Model for Predicting Time to Future Breast Cancer Events from Mammograms
Xin Wang 0121, Tao Tan 0002, Eric Marcus, Luyi Han, Antonio Portaluri, Tianyu Zhang 0006, Chunyao Lu, Xinglong Liang, Regina Beets-Tan, Jonas Teuwen, Ritse Mann |
MICCAI (1) | 2 |
| 2024 | Variational Field Constraint Learning for Degree of Coronary Artery Ischemia Assessment
Qi Zhang 0078, Xiujian Liu, Heye Zhang, Chenchu Xu, Guang Yang 0006, Yixuan Yuan, Tao Tan 0002, Zhifan Gao |
MICCAI (3) | 7 |
| 2024 | MLC: Multi-level consistency learning for semi-supervised left atrium segmentation
Zhebin Shi, Mingfeng Jiang, Yang Li 0097, Bo Wei 0004, Yongquan Wu, Tao Tan 0002, Guang Yang 0006 |
Expert Syst. Appl. | 7 |
| 2024 | BSANet: Boundary-aware and scale-aggregation networks for CMR image segmentation
Dan Zhang 0026, Chenggang Lu, Tao Tan 0002, Behdad Dashtbozorg, Xi Long 0001, Xiayu Xu, Jiong Zhang 0004, Caifeng Shan |
Neurocomputing | 3 |
| 2024 | Synthesis-based imaging-differentiation representation learning for multi-sequence 3D/4D MRI
Luyi Han, Tao Tan 0002, Tianyu Zhang 0006, Yunzhi Huang, Xin Wang 0121, Jonas Teuwen, Ritse Mann |
Medical Image Anal. | 2 |
| 2024 | An end-to-end multi-task deep learning framework for bronchoscopy image classification
Rojin Setayeshi, Javad Vahidi, Ehsan Kozegar, Tao Tan 0002 |
Multim. Syst. | 4 |
| 2024 | LYSTO: The Lymphocyte Assessment Hackathon and Benchmark DatasetabstractWe introduce LYSTO, the Lymphocyte Assessment Hackathon, which was held in conjunction with the MICCAI 2019 Conference in Shenzhen (China). The competition required participants to automatically assess the number of lymphocytes, in particular T-cells, in images of colon, breast, and prostate cancer stained with CD3 and CD8 immunohistochemistry. Differently from other challenges setup in medical image analysis, LYSTO participants were solely given a few hours to address this problem. In this paper, we describe the goal and the multi-phase organization of the hackathon; we describe the proposed methods and the on-site results. Additionally, we present post-competition results where we show how the presented methods perform on an independent set of lung cancer slides, which was not part of the initial competition, as well as a comparison on lymphocyte assessment between presented methods and a panel of pathologists. We show that some of the participants were capable to achieve pathologist-level performance at lymphocyte assessment. After the hackathon, LYSTO was left as a lightweight plug-and-play benchmark dataset on grand-challenge website, together with an automatic evaluation platform. Yiping Jiao, Jeroen van der Laak, Shadi Albarqouni, Tao Tan 0002, Abhir Bhalerao, Shenghua Cheng, Jiabo Ma, John Pocock, Josien P. W. Pluim, Navid Alemi Koohbanani, Raja Muhammad Saad Bashir, Shan E Ahmed Raza, Sibo Liu, Simon Graham, Suzanne C. Wetstein, Syed Ali Khurram, Nasir M. Rajpoot, Mitko Veta, Francesco Ciompi |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | ERNet: Edge Regularization Network for Cerebral Vessel Segmentation in Digital Subtraction Angiography ImagesabstractStroke is a leading cause of disability and fatality in the world, with ischemic stroke being the most common type. Digital Subtraction Angiography images, the gold standard in the operation process, can accurately show the contours and blood flow of cerebral vessels. The segmentation of cerebral vessels in DSA images can effectively help physicians assess the lesions. However, due to the disturbances in imaging parameters and changes in imaging scale, accurate cerebral vessel segmentation in DSA images is still a challenging task. In this paper, we propose a novel Edge Regularization Network (ERNet) to segment cerebral vessels in DSA images. Specifically, ERNet employs the erosion and dilation processes on the original binary vessel annotation to generate pseudo-ground truths of False Negative and False Positive, which serve as constraints to refine the coarse predictions based on their mapping relationship with the original vessels. In addition, we exploit a Hybrid Fusion Module based on convolution and transformers to extract local features and build long-range dependencies. Moreover, to support and advance the open research in the field of ischemic stroke, we introduce FPDSA, the first pixel-level semantic segmentation dataset for cerebral vessels. Extensive experiments on FPDSA illustrate the leading performance of our ERNet. Weijin Xu, Yinghuan Shi, Tao Tan 0002, Wentao Liu 0004, Xipeng Pan, Yiming Deng, Ruisheng Su |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Automatic detection of breast lesions in automated 3D breast ultrasound with cross-organ transfer learningabstractDeep convolutional neural networks have garnered considerable attention in numerous machine learning applications, particularly in visual recognition tasks such as image and video analyses. There is a growing interest in applying this technology to diverse applications in medical image analysis. Automated three-dimensional Breast Ultrasound is a vital tool for detecting breast cancer, and computer-assisted diagnosis software, developed based on deep learning, can effectively assist radiologists in diagnosis. However, the network model is prone to overfitting during training, owing to challenges such as insufficient training data. This study attempts to solve the problem caused by small datasets and improve model detection performance. We propose a breast cancer detection framework based on deep learning (a transfer learning method based on cross-organ cancer detection) and a contrastive learning method based on breast imaging reporting and data systems (BI-RADS). When using cross organ transfer learning and BIRADS based contrastive learning, the average sensitivity of the model increased by a maximum of 16.05%. Our experiments have demonstrated that the parameters and experiences of cross-organ cancer detection can be mutually referenced, and contrastive learning method based on BI-RADS can improve the detection performance of the model. B. A. O. Lingyun, Zhengrui Huang, Yue Sun 0001, Hui Chen 0020, Xiaochen Yuan, Tao Tan 0002 |
Virtual Real. Intell. Hardw. | 10 |
| 2024 | Combining machine and deep transfer learning for mediastinal lymph node evaluation in patients with lung cancerabstractThe prognosis and survival of patients with lung cancer are likely to deteriorate with metastasis. Using deep-learning in the detection of lymph node metastasis can facilitate the noninvasive calculation of the likelihood of such metastasis, thereby providing clinicians with crucial information to enhance diagnostic precision and ultimately improve patient survival and prognosis In total, 623 eligible patients were recruited from two medical institutions. Seven deep learning models, namely Alex, GoogLeNet, Resnet18, Resnet101, Vgg16, Vgg19, and MobileNetv3 (small), were utilized to extract deep image histological features. The dimensionality of the extracted features was then reduced using the Spearman correlation coefficient (r ≥ 0.9) and Least Absolute Shrinkage and Selection Operator. Eleven machine learning methods, namely Support Vector Machine, K-nearest neighbor, Random Forest, Extra Trees, XGBoost, LightGBM, Naive Bayes, AdaBoost, Gradient Boosting Decision Tree, Linear Regression, and Multilayer Perceptron, were employed to construct classification prediction models for the filtered final features. The diagnostic performances of the models were assessed using various metrics, including accuracy, area under the receiver operating characteristic curve, sensitivity, specificity, positive predictive value, and negative predictive value. Calibration and decision-curve analyses were also performed. The present study demonstrated that using deep radiomic features extracted from Vgg16, in conjunction with a prediction model constructed via a linear regression algorithm, effectively distinguished the status of mediastinal lymph nodes in patients with lung cancer. The performance of the model was evaluated based on various metrics, including accuracy, area under the receiver operating characteristic curve, sensitivity, specificity, positive predictive value, and negative predictive value, which yielded values of 0.808, 0.834, 0.851, 0.745, 0.829, and 0.776, respectively. The validation set of the model was assessed using clinical decision curves, calibration curves, and confusion matrices, which collectively demonstrated the model's stability and accuracy In this study, information on the deep radiomics of Vgg16 was obtained from computed tomography images, and the linear regression method was able to accurately diagnose mediastinal lymph node metastases in patients with lung cancer. Jianfang Zhang, Lijuan Ding, Tao Tan 0002, Qing Li 0073 |
Virtual Real. Intell. Hardw. | 4 |
| 2024 | ARGA-Unet: Advanced U-net segmentation model using residual grouped convolution and attention mechanism for brain tumor MRI image segmentationabstractMagnetic resonance imaging (MRI) has played an important role in the rapid growth of medical imaging diagnostic technology, especially in the diagnosis and treatment of brain tumors owing to its non-invasive characteristics and superior soft tissue contrast. However, brain tumors are characterized by high non-uniformity and non-obvious boundaries in MRI images because of their invasive and highly heterogeneous nature. In addition, the labeling of tumor areas is time-consuming and laborious. To address these issues, this study uses a residual grouped convolution module, convolutional block attention module, and bilinear interpolation upsampling method to improve the classical segmentation network U-net. The influence of network normalization, loss function, and network depth on segmentation performance is further considered. In the experiments, the Dice score of the proposed segmentation model reached 97.581%, which is 12.438% higher than that of traditional U-net, demonstrating the effective segmentation of MRI brain tumor images. In conclusion, we use the improved U-net network to achieve a good segmentation effect of brain tumor MRI images. Siyi Xun, Sixu Duan, Tong Tong 0001, Qinquan Gao, Chan-Tong Lam, Menghan Hu, Tao Tan 0002 |
Virtual Real. Intell. Hardw. | 10 |
| 2023 | Improved YOLOX Framework for Automatic Large Intracranial Artery Stenosis Detection in Digital Subtraction Angiography ImagesabstractIschemic stroke has a very high mortality and disability rate, and intracranial artery stenosis is an important cause of ischemic stroke. At present, transvascular interventional surgery is an effective remedy to treat intracranial artery stenosis, and as the gold standard in surgery, Digital Subtraction Angiography (DSA) images can effectively display the outline of blood vessels and the flow of blood. Detecting and locating the stenosis from DSA images is a challenging problem due to the large variation in the thickness of the blood vessel and the complex shape of the blood vessel. In this paper, we collect a dataset with 2860 DSA sequence samples and annotate stenosis locations, constructing the first automatic detection and localization method for stenosis in DSA images. In addition, considering that the commonly used Intersection-over-Union (IoU) loss ignores the similarity indicators of the image patches in the prediction box and the ground-truth (GT) box, a plug-and-play loss function that considers the image similarity between the prediction box and the GT box is proposed to effectively improve network performance. Extensive experiments demonstrate the effectiveness of our approach, which outperforms classical detectors. Weijin Xu, Tao Tan 0002, Wentao Liu 0004, Yiming Deng, Xipeng Pan, Ruisheng Su |
BIBM | 3 |
| 2023 | Controllable Deep Learning Denoising Model for Ultrasound Images Using Synthetic Noisy Image
Mingfu Jiang, Chenzhi You, Heye Zhang, Zhifan Gao, Tao Tan 0002 |
CGI (1) | 7 |
| 2023 | A Hybrid Supervised Fusion Deep Learning Framework for Microscope Multi-Focus Images
Qiuhui Yang, Hao Chen 0037, Mingfeng Jiang, Jiong Zhang 0004, Yue Sun 0001, Tao Tan 0002 |
CGI (4) | 7 |
| 2023 | SMOC-Net: Leveraging Camera Pose for Self-Supervised Monocular Object Pose EstimationabstractRecently, self-supervised 6D object pose estimation, where synthetic images with object poses (sometimes jointly with un-annotated real images) are used for training, has attracted much attention in computer vision. Some typical works in literature employ a time-consuming differentiable renderer for object pose prediction at the training stage, so that (i) their performances on real images are generally limited due to the gap between their rendered images and real images and (ii) their training process is computationally expensive. To address the two problems, we propose a novel Network for Self-supervised Monocular Object pose estimation by utilizing the predicted Camera poses from unannotated real images, called SMOC-Net. The proposed network is explored under a knowledge distillation framework, consisting of a teacher model and a student model. The teacher model contains a backbone estimation module for initial object pose estimation, and an object pose refiner for refining the initial object poses using a geometric constraint (called relative-pose constraint) derived from relative camera poses. The student model gains knowledge for object pose estimation from the teacher model by imposing the relative-pose constraint. Thanks to the relative-pose constraint, SMOC-Net could not only narrow the domain gap between synthetic and real data but also reduce the training cost. Experimental results on two public datasets demonstrate that SMOC-Net outperforms several state-of-the-art methods by a large margin while requiring much less training time than the differentiable-renderer-based methods. Tao Tan 0002, Qiulei Dong |
CVPR | 1 |
| 2023 | An Explainable Deep Framework: Towards Task-Specific Fusion for Multi-to-One MRI Synthesis
Luyi Han, Tianyu Zhang 0006, Yunzhi Huang, Haoran Dou, Xin Wang 0121, Chunyao Lu, Tao Tan 0002, Ritse Mann |
MICCAI (10) | 8 |
| 2023 | DisAsymNet: Disentanglement of Asymmetrical Abnormality on Bilateral Mammograms Using Self-adversarial Learning
Xin Wang 0121, Tao Tan 0002, Luyi Han, Tianyu Zhang 0006, Chunyao Lu, Regina Beets-Tan, Ruisheng Su, Ritse Mann |
MICCAI (7) | 2 |
| 2023 | Synthesis of Contrast-Enhanced Breast MRI Using T1- and Multi-b-Value DWI-Based Hierarchical Fusion Network with Attention Mechanism
Tianyu Zhang 0006, Luyi Han, Anna D'Angelo, Xin Wang 0121, Chunyao Lu, Jonas Teuwen, Regina Beets-Tan, Tao Tan 0002, Ritse Mann |
MICCAI (7) | 9 |
| 2023 | Light-VQA: A Multi-Dimensional Quality Assessment Model for Low-Light Video EnhancementabstractRecently, Users Generated Content (UGC) videos becomes ubiquitous in our daily lives. However, due to the limitations of photographic equipments and techniques, UGC videos often contain various degradations, in which one of the most visually unfavorable effects is the underexposure. Therefore, corresponding video enhancement algorithms such as Low-Light Video Enhancement (LLVE) have been proposed to deal with the specific degradation. However, different from video enhancement algorithms, almost all existing Video Quality Assessment (VQA) models are built generally rather than specifically, which measure the quality of a video from a comprehensive perspective. To the best of our knowledge, there is no VQA model specially designed for videos enhanced by LLVE algorithms. To this end, we first construct a Low-Light Video Enhancement Quality Assessment (LLVE-QA) dataset in which 254 original low-light videos are collected and then enhanced by leveraging 8 LLVE algorithms to obtain 2,060 videos in total. Moreover, we propose a quality assessment model specialized in LLVE, named Light-VQA. More concretely, since the brightness and noise have the most impact on low-light enhanced VQA, we handcraft corresponding features and integrate them with deep-learning-based semantic features as the overall spatial information. As for temporal information, in addition to deep-learning-based motion features, we also investigate the handcrafted brightness consistency among video frames, and the overall temporal information is their concatenation. Subsequently, spatial and temporal information is fused to obtain the quality-aware representation of a video. Extensive experimental results show that our Light-VQA achieves the best performance against the current State-Of-The-Art (SOTA) on LLVE-QA and public dataset. Dataset and Codes can be found at https://github.com/wenzhouyidu/Light-VQA. Yunlong Dong, Xiaohong Liu 0001, Xunchu Zhou, Tao Tan 0002, Guangtao Zhai |
ACM Multimedia | 5 |
| 2023 | TransMRSR: transformer-based self-distilled generative prior for brain MRI super-resolution
Xiaohong Liu 0001, Tao Tan 0002, Menghan Hu, Xiaoer Wei, Tingli Chen, Bin Sheng 0001 |
Vis. Comput. | 3 |
| 2023 | PCTMF-Net: heart sound classification with parallel CNNs-transformer and second-order spectral analysisabstractHeart disease is a common condition worldwide and has become one of the leading causes of death worldwide. The electrocardiogram (PCG) is a safe, painless, and non-invasive test that captures bioacoustic information reflecting the function of the heart by capturing the acoustic signal of the patient’s heart. Nowadays, based on biosignal processing and artificial intelligence technologies, automated heart sound classification is playing an increasingly important role in clinical applications. In this paper, we propose a new parallel CNNs-transformer network with multi-scale feature context aggregation (PCTMF-Net). It combines the advantages of CNNs and transformer to achieve efficient heart sound classification. In PCTMF-Net, firstly, the heart tone signal features are extracted using the second-order spectral analysis, and a transformer-based MHTE-4 (multi-head transformer encoder with four attention heads) is designed to encode and aggregate the contextual information, and then, two CNNs feature extractors are designed in parallel with MHTE-4 to capture the hierarchical features. Finally, the feature vectors obtained from CNNs and MHTE-4 through feature fusion in PCTMF-Net will be fed into the fully connected layer for predicting the classification results of heart sounds. In addition, we perform validation based on two publicly available mutually exclusive heart sound datasets and conduct extensive experiments and comparisons of existing algorithms under different metrics. The experimental results show that our proposed method achieves 99.36% accuracy on the Yaseen dataset and 93% accuracy on the PhysioNet dataset. It surpasses current algorithms in terms of accuracy, recall and F 1-score metrics. The aim of this study is to apply these new techniques and methods to improve the diagnostic accuracy and validity of heart disease for clinical use. Rongsheng Wang 0004, Yaofei Duan, Dashun Zheng, Xiaohong Liu 0001, Chan-Tong Lam, Tao Tan 0002 |
Vis. Comput. | 7 |
| 2022 | Towards Idea Mining: Problem-Solution Phrase Extraction from Text
Haixia Liu 0001, Tim J. Brailsford, James Goulding, Tomas Maul, Tao Tan 0002, Debanjan Chaudhuri |
ADMA (2) | 5 |
| 2022 | Multi-modal trained artificial intelligence solution to triage chest X-ray for COVID-19 using pristine ground-truth, versus radiologists
Tao Tan 0002, Bipul Das, Ravi Soni, Mate Fejes, Hongxu Yang, Sohan Ranjan, Daniel Attila Szabo, Vikram Melapudi, K. S. Shriram, Utkarsh Agrawal, László Ruskó, Zita Herczeg, Barbara Darázs, Pal Tegzes, Lehel Ferenczi, Rakesh Mullick, Gopal Avinash |
Neurocomputing | 1 |
| 2022 | Guest Editorial Artificial Intelligence in Pre-DICOMabstractThe papers in this special section focus on artificial intelligence pre-DICOM medical imaging. AI for medical imaging is applied in three domains: pre-DICOM, pre-processing and clinical applications. Clinical applications mainly cover topics such as disease detection, classification, segmentation, registration. Pre-processing components are mainly designed for facilitating applications using image transformation such as image normalization, noise reduction, bias correction in MR. AI in the pre-DICOM domain is expected to improve imaging workflow, image protocol selection, imaging quality, imaging scanning time before images are converted into DICOM format for radiologists to review. The trends of AI publications in medical imaging have been gradually extended from clinical applications to pre-processing and, to pre-DICOM. The papers in this special section seek to present and highlight the latest development on applying advanced deep learning techniques in pre-DICOM space. The papers highlight the latest development on applying advanced deep learning techniques in pre-DICOM space. Tao Tan 0002, Ravi Soni, Jungong Han, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | Pristine Annotations-Based Multi-modal Trained Artificial Intelligence Solution to Triage Chest X-Ray for COVID-19
Tao Tan 0002, Bipul Das, Ravi Soni, Mate Fejes, Sohan Ranjan, Daniel Attila Szabo, Vikram Melapudi, K. S. Shriram, Utkarsh Agrawal, László Ruskó, Zita Herczeg, Barbara Darázs, Pal Tegzes, Lehel Ferenczi, Rakesh Mullick, Gopal Avinash |
MICCAI (7) | 1 |
| 2021 | Lesion Segmentation in Ultrasound Using Semi-Pixel-Wise Cycle Generative Adversarial NetsabstractBreast cancer is the most common invasive cancer with the highest cancer occurrence in females. Handheld ultrasound is one of the most efficient ways to identify and diagnose the breast cancer. The area and the shape information of a lesion is very helpful for clinicians to make diagnostic decisions. In this study we propose a new deep-learning scheme, semi-pixel-wise cycle generative adversarial net (SPCGAN) for segmenting the lesion in 2D ultrasound. The method takes the advantage of a fully convolutional neural network (FCN) and a generative adversarial net to segment a lesion by using prior knowledge. We compared the proposed method to a fully connected neural network and the level set segmentation method on a test dataset consisting of 32 malignant lesions and 109 benign lesions. Our proposed method achieved a Dice similarity coefficient (DSC) of 0.92 while FCN and the level set achieved 0.90 and 0.79 respectively. Particularly, for malignant lesions, our method increases the DSC (0.90) of the fully connected neural network to 0.93 significantly (p 0.001). The results show that our SPCGAN can obtain robust segmentation results. The framework of SPCGAN is particularly effective when sufficient training samples are not available compared to FCN. Our proposed method may be used to relieve the radiologists' burden for annotation. Zheren Li, Biyuan Wang, Yuji Qi, Bingbin Yu, Farhad G. Zanjani, Aiwen Zheng, Remco Duits, Tao Tan 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 9 |
| 2021 | Deep Learning Methods for Lung Cancer Segmentation in Whole-Slide Histopathology Images - The ACDC@LungHP Challenge 2019abstractAccurate segmentation of lung cancer in pathology slides is a critical step in improving patient care. We proposed the ACDC@LungHP (Automatic Cancer Detection and Classification in Whole-slide Lung Histopathology) challenge for evaluating different computer-aided diagnosis (CADs) methods on the automatic diagnosis of lung cancer. The ACDC@LungHP 2019 focused on segmentation (pixel-wise detection) of cancer tissue in whole slide imaging (WSI), using an annotated dataset of 150 training images and 50 test images from 200 patients. This paper reviews this challenge and summarizes the top 10 submitted methods for lung cancer segmentation. All methods were evaluated using metrics using the precision, accuracy, sensitivity, specificity, and DICE coefficient (DC). The DC ranged from 0.7354 ±0.1149 to 0.8372 ±0.0858. The DC of the best method was close to the inter-observer agreement (0.8398 ±0.0890). All methods were based on deep learning and categorized into two groups: multi-model method and single model method. In general, multi-model methods were significantly better (p 0.01) than single model methods, with mean DC of 0.7966 and 0.7544, respectively. Deep learning based methods could potentially help pathologists find suspicious regions for further analysis of lung cancer in WSI. Tao Tan 0002, Xichao Teng, Xiaoliang Sun, Lihong Liu, Byungjae Lee, Yilong Li 0002, Qianni Zhang, Shujiao Sun, Yushan Zheng, Junyu Yan, Yiyu Hong, Junsu Ko, Hyun Jung, Ching-Wei Wang, Vladimir Yurovskiy, Pavel Maevskikh, Vahid Khanagha, Daiqiang Li, Peter J. Schüffler, Hui Chen 0020, Yuling Tang, Geert Litjens 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Guest Editorial: Deep Learning in Ultrasound ImagingabstractAmong the different imaging modalities, ultrasound is the most widespread modality for visualizing human tissue due to it being low-cost, non-ionizing, real-time with immediate feedback to the sonographer, convenient to operate, widely available and well established, with a very large number of images generated in a single setting. On the other hand, ultrasound imaging suffers from the disadvantage of being user dependent and of variable quality,which makes the automated interpretation of ultrasound images often very difficult. In recent years, algorithms in medical imaging have been significantly improved thanks to the advent of deep learning methods (including convolutional neural networks, recurrent neural networks, autoencoders, or generative adversarial networks). To address the various challenges of automatically processing and interpreting ultrasound images, deep learning techniques have been gradually applied to various types of ultrasound data (such as B-mode ultrasound, Doppler ultrasound, or contrast-enhanced ultrasound), acquired with a range of different probes, with the aim of improving image quality, for organ segmentation, device localization and tracking, for tissue characterization, and ultimately to improve disease diagnosis and therapeutic outcome. The papers in this special section seek to present and highlight the latest development on applying advanced deep learning techniques in ultrasound imaging. Caifeng Shan, Tao Tan 0002, Shandong Wu, Julia A. Schnabel |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | Local to Global Learning: Gradually Adding Classes for Training Deep Neural NetworksabstractWe propose a new learning paradigm, Local to Global Learning (LGL), for Deep Neural Networks (DNNs) to improve the performance of classification problems. The core of LGL is to learn a DNN model from fewer categories (local) to more categories (global) gradually within the entire training set. LGL is most related to the Self-Paced Learning (SPL) algorithm but its formulation is different from SPL. SPL trains its data from simple to complex, while LGL from local to global. In this paper, we incorporate the idea of LGL into the learning objective of DNNs and explain why LGL works better from an information-theoretic perspective. Experiments on the toy data, CIFAR-10, CIFAR-100, and ImageNet dataset show that LGL outperforms the baseline and SPL-based algorithms. Hao Cheng 0005, Dongze Lian, Shenghua Gao, Tao Tan 0002, Yanlin Geng |
CVPR | 5 |
| 2019 | SLAOE-NN: A Deep Network with Structure Learning for Aspect and Opining co-Extraction for NLPabstractThe task of co-extracting aspects and opinion terms is intended to explicitly extract aspect terms that describe entity features and opinion terms that express emotions from user-generated text. An effective way to accomplish this task is to exploit the relationship between aspect terms and opinion terms by parsing the syntax structure of each sentence. However, this method requires a lot of effort to parse and is highly dependent on the quality of the parsing results. In this paper, we present a deep learning model called SLAOE-NN (Structural Learning and Aspect and Opining Extraction Neural Networks). The proposed model provides an end-to-end solution and does not require any other language resources for preprocessing. Particularly, we use ON-LSTM to generate hidden layer with language structure information which can generate constituency tree unsupervised and we serve it as an auxiliary task for aspect and opinion terms extraction. For aspect terms and opinion terms extract task, we propose different attention mechanism, which can exploit the indirect relationship between aspect term and corresponding opining term to achieve more accurate information extraction. The experimental results of SemEval's three benchmark datasets in 2014 and 2015 show that our model achieves the-state-of-art performance compared to several baselines. Kunling Liu, Shiqun Yin, Tao Tan 0002 |
ICTAI | 3 |
| 2019 | On the Convergence Speed of AMSGRAD and BeyondabstractIn ICLR's (2018) best paper "On the Convergence of Adam and Beyond", the author points out the shortcomings in Adam's convergence proof, proposes an AMSGRAD algorithm that can guarantee convergence as the number of iterations increases. However, through some comparative experiments, this paper finds that there are two problems in the convergence process of AMSGRAD algorithm. Firstly, the AMSGRAD algorithm is easy to oscillate; Secondly, the AMSGRAD algorithm converges slowly. After analysis, the above two problems can be solved by the following ways. When gt-1gt> 0, this paper adds the momentum term in Momentum algorithm to the AMSGRAD algorithm to accelerate convergence. When gt-1gt≤ 0, this paper use SGD algorithm instead of AMSGRAD algorithm to update the model weights. In order to eliminate some negative effects of the previous parameter gradient on the current parameter gradient and reduce the oscillation amplitude of the objective function, the first-order and second-order moment estimations of the parameter gradient are recalculated when gt-1gt≤ 0. Therefore, this paper proposes the ACADG algorithm, which not only can improve the convergence speed, suppress the oscillation amplitude of the objective function, but also can improve the accuracy of training and test data sets. Tao Tan 0002, Shiqun Yin, Kunling Liu, Man Wan |
ICTAI | 1 |
| 2019 | Transferring from ex-vivo to in-vivo: Instrument Localization in 3D Cardiac Ultrasound Using Pyramid-UNet with Hybrid Loss
Hongxu Yang, Caifeng Shan, Tao Tan 0002, Alexander F. Kolen, Peter H. N. de With |
MICCAI (5) | 3 |
| 2019 | Epileptic seizure detection in EEG signals using sparse multiscale radial basis function networks and the Fisher vector approach
Yang Li 0010, Wei-Gang Cui, Yuzhu Guo, Tao Tan 0002 |
Knowl. Based Syst. | 6 |
| 2019 | Video-based discomfort detection for infants
Yue Sun 0001, Caifeng Shan, Tao Tan 0002, Xi Long 0001, Arash Pourtaherian, Svitlana Zinger, Peter H. N. de With |
Mach. Vis. Appl. | 3 |
| 2018 | Mass Segmentation in Automated 3-D Breast Ultrasound Using Adaptive Region Growing and Supervised Edge-Based Deformable ModelabstractAutomated 3-D breast ultrasound has been proposed as a complementary modality to mammography for early detection of breast cancers. To facilitate the interpretation of these images, computer aided detection systems are being developed in which mass segmentation is an essential component for feature extraction and temporal comparisons. However, automated segmentation of masses is challenging because of the large variety in shape, size, and texture of these 3-D objects. In this paper, the authors aim to develop a computerized segmentation system, which uses a seed position as the only priori of the problem. A two-stage segmentation approach has been proposed incorporating shape information of training masses. At the first stage, a new adaptive region growing algorithm is used to give a rough estimation of the mass boundary. The similarity threshold of the proposed algorithm is determined using a Gaussian mixture model based on the volume and circularity of the training masses. In the second stage, a novel geometric edge-based deformable model is introduced using the result of the first stage as the initial contour. In a data set of 50 masses, including 38 malignant and 12 benign lesions, the proposed segmentation method achieved a mean Dice of 0.74 ± 0.19 which outperformed the adaptive region growing with a mean Dice of 0.65 ± 0.2 (p-value < 0.02). Moreover, the resulting mean Dice was significantly (p-value < 0.001) better than that of the distance regularized level set evolution method (0.52 ± 0.27). The supervised method presented in this paper achieved accurate mass segmentation results in terms of Dice measure. The suggested segmentation method can be utilized in two aspects: 1) to automatically measure the change in volume of breast lesions over time and 2) to extract features for a computer aided detection or diagnosis system. Ehsan Kozegar, Mohsen Soryani, Hamid Behnam, Masoumeh Salamati, Tao Tan 0002 |
IEEE Trans. Medical Imaging | 5 |
| 2013 | Chest wall segmentation in automated 3D breast ultrasound scans
Tao Tan 0002, Bram Platel, Ritse Mann, Henkjan J. Huisman, Nico Karssemeijer |
Medical Image Anal. | 1 |
| 2013 | Computer-Aided Detection of Cancer in Automated 3-D Breast UltrasoundabstractAutomated 3-D breast ultrasound (ABUS) has gained a lot of interest and may become widely used in screening of dense breasts, where sensitivity of mammography is poor. However, reading ABUS images is time consuming, and subtle abnormalities may be missed. Therefore, we are developing a computer aided detection (CAD) system to help reduce reading time and prevent errors. In the multi-stage system we propose, segmentations of the breast, the nipple and the chestwall are performed, providing landmarks for the detection algorithm. Subsequently, voxel features characterizing coronal spiculation patterns, blobness, contrast, and depth are extracted. Using an ensemble of neural-network classifiers, a likelihood map indicating potential abnormality is computed. Local maxima in the likelihood map are determined and form a set of candidates in each image. These candidates are further processed in a second detection stage, which includes region segmentation, feature extraction and a final classification. On region level, classification experiments were performed using different classifiers including an ensemble of neural networks, a support vector machine, a k-nearest neighbors, a linear discriminant, and a gentle boost classifier. Performance was determined using a dataset of 238 patients with 348 images (views), including 169 malignant and 154 benign lesions. Using free response receiver operating characteristic (FROC) analysis, the system obtains a view-based sensitivity of 64% at 1 false positives per image using an ensemble of neural-network classifiers. Tao Tan 0002, Bram Platel, Roel Mus, László K. Tabár, Ritse Mann, Nico Karssemeijer |
IEEE Trans. Medical Imaging | 1 |
| 2012 | Computer-Aided Lesion Diagnosis in Automated 3-D Breast Ultrasound Using Coronal SpiculationabstractA computer-aided diagnosis (CAD) system for the classification of lesions as malignant or benign in automated 3-D breast ultrasound (ABUS) images, is presented. Lesions are automatically segmented when a seed point is provided, using dynamic programming in combination with a spiral scanning technique. A novel aspect of ABUS imaging is the presence of spiculation patterns in coronal planes perpendicular to the transducer. Spiculation patterns are characteristic for malignant lesions. Therefore, we compute spiculation features and combine them with features related to echotexture, echogenicity, shape, posterior acoustic behavior and margins. Classification experiments were performed using a support vector machine classifier and evaluation was done with leave-one-patient-out cross-validation. Receiver operator characteristic (ROC) analysis was used to determine performance of the system on a dataset of 201 lesions. We found that spiculation was among the most discriminative features. Using all features, the area under the ROC curve (A(z)) was 0.93, which was significantly higher than the performance without spiculation features (A(z)=0.90, p=0.02). On a subset of 88 cases, classification performance of CAD (A(z)=0.90) was comparable to the average performance of 10 readers (A(z)=0.87). Tao Tan 0002, Bram Platel, Henkjan J. Huisman, Clara I. Sánchez, Roel Mus, Nico Karssemeijer |
IEEE Trans. Medical Imaging | 1 |