VLDB 2026 Research / reviewers in the wild / expert
Bin Pu
dblp:237/6362
· DBLP profile ↗
69ranked-venue papers
16as first author
68since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 10 first-author · 38 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 5 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 16 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Topology-Inspired Backward-Free Framework for Test-Time Adaptation in Medical DetectionabstractRecently, Test-Time Adaptation (TTA) has gained increasing attention in medical imaging due to its ability to improve model generalization under domain shifts without retraining. In particular, directly applying a well-trained model across various medical centers faces significant performance degradation caused by variations in equipment, operators, imaging conditions, and scanning skill levels of sonographers. Existing TTA methods either rely on parameter adaptation that increases computational cost or apply simple prediction fusion that ignores anatomical structure knowledge. To address these limitations, we propose a novel backward-free Topology-aware TTA framework named T^3 that integrates Structural Perception Modeling (SPM) and Box Regression Adaptation (BRA). SPM is implemented through an organ space heatmap generated via Gaussian kernel superposition. This heatmap encodes anatomical topology without requiring additional training or source data. BRA further improves localization and classification by fusing detection outputs based on the contribution of detected results to anatomically meaningful peak points from the heatmaps. Extensive experiments were conducted across six cross-domain scenarios, and the results demonstrate that our method achieves state-of-the-art cross-domain detection performance while maintaining high efficiency, offering a practical and robust solution for real-world medical diagnostic applications. Bin Pu, Xingguo Lv, Jiewen Yang, Lei Zhao 0013, Zuozhu Liu, Kenli Li 0001 |
AAAI | 1 |
| 2026 | Organ-Aware Routing Mixture-of-Retrieval Augmented Generation for Fetal Ultrasound ReportingabstractFetal ultrasound screening is a uniquely complex diagnostic task involving the simultaneous assessment of multiple fetal organs—each with its own anatomical and clinical context—within a single examination. Automating report generation for such cases poses a significant challenge: unlike existing methods that focus on single-organ radiology tasks (e.g., chest X-rays), fetal ultrasound requires reasoning over a structured, multiple-to-multiple setting, i.e., multi-organ images corresponding to a multi-section report. In this paper, we introduce FetusR, the first large-scale dataset for multi-organ fetal ultrasound reporting, containing 15,594 real-world cases with rich organ-wise annotations. To address the intrinsic image-report alignment, we propose Organ-Aware Routing Mixture-of-Retrieval Augmented Generation (ORM-RAG) inspired by the Mixture-of-Experts paradigm. Our method decomposes the complex alignment problem into multiple one-to-one sub-retrieval tasks. Specifically, ORM-RAG integrates (1) an organ-aware mixture-of-retrieval module that partitions the retrieval space into organ-specific corpora for independent retrieval, and (2) a dynamic routing mechanism that selectively aggregates high-confidence organ-specific reports while filtering uncertain ones. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art baselines across both textual similarity and clinical accuracy metrics. Our work opens a new direction for long-form, structured report generation in real-world, multi-organ medical imaging scenarios. Bin Pu, Rongbin Li, Xinpeng Ding, Lei Zhao 0013, Chaoqi Chen, Shengli Li 0001, Kenli Li 0001 |
AAAI | 1 |
| 2026 | Unified Mixture-of-Experts Framework for Joint Cardiac and Vascular Ultrasound Analysis and Report GenerationabstractEchocardiography and vascular ultrasound are essential for comprehensive cardiovascular assessment, yet manual evaluation and writing reports are labor-intensive, time-consuming, and require expertise from both cardiology and vascular surgery departments. Current automated report generation systems mainly focus on X-ray or CT, often neglecting echocardiographic modalities and critical quantitative parameters like aortic diameter and main pulmonary artery diameter, limiting their clinical utility. Moreover, the interdependence between cardiac and peripheral vascular health necessitates cross-departmental insights, which existing methods fail to incorporate. To address these limitations, we first propose the vision-language framework named the Echo-Cardiac-Vascular (ECV), for joint cardiac and vascular ultrasound report generation and parameter measurements. ECV introduces a Mixture-of-Experts vision encoder tailored for distinct ultrasound subtypes, a structured parameter measurement module for accurate quantification, and task-specific decoders that generate interpretable, multimodal diagnostic reports. Our framework, trained on 10K+ paired records, achieves high accuracy, improving diagnostic efficiency, consistency, and cross-disciplinary clinical applicability. Bin Pu, Jiewen Yang, Xingguo Lv, Kenli Li 0001 |
AAAI | 1 |
| 2026 | MPA: Multimodal Prototype Augmentation for Few-Shot LearningabstractRecently, Few-shot Learning (FSL) has become a popular task that aims to recognize new classes from only a few labeled examples and has been widely applied in fields such as natural science, remote sensing, and medical images. However, most existing methods focus only on the visual modality and compute prototypes directly from raw support images, which lack comprehensive and rich multimodal information. To address these limitations, we propose a novel Multimodal Prototype Augmentation FSL framework called MPA, including LLM-based Multi-Variant Semantic Enhancement (LMSE), Hierarchical Multi-View Augmentation (HMA), and an Adaptive Uncertain Class Absorber (AUCA). LMSE leverages large language models to generate diverse paraphrased category descriptions, enriching the support set with additional semantic cues. HMA exploits both natural and multi-view augmentations to enhance feature diversity (e.g., changes in viewing distance, camera angles, and lighting conditions). AUCA models uncertainty by introducing uncertain classes via interpolation and Gaussian sampling, effectively absorbing uncertain samples. Extensive experiments on four single-domain and six cross-domain FSL benchmarks demonstrate that MPA achieves superior performance compared to existing state-of-the-art methods across most settings. Notably, MPA surpasses the second-best method by 12.29% and 24.56% in the single-domain and cross-domain setting, respectively, in the 5-way 1-shot setting. Liwen Wu, Lei Zhao 0013, Qika Lin, Shaowen Yao 0001, Zuozhu Liu, Bin Pu |
AAAI | 8 |
| 2026 | MedForge: Interpretable Medical Deepfake Detection via Forgery-aware ReasoningabstractText-guided image editors can now manipulate authentic medical scans with high fidelity, enabling lesion implantation/removal that threatens clinical trust and safety. Existing defenses are inadequate for healthcare. Medical detectors are largely black-box, while MLLM-based explainers are typically post-hoc, lack medical expertise, and may hallucinate evidence on ambiguous cases. We present MedForge, a data-and-method solution for pre-hoc, evidence-grounded medical forgery detection. We introduce MedForge-90K, a large-scale benchmark of realistic lesion edits across 19 pathologies with expert-guided reasoning supervision via doctor inspection guidelines and gold edit locations. Building on it, MedForge-Reasoner performs localize-then-analyze reasoning, predicting suspicious regions before producing a verdict, and is further aligned with Forgery-aware GSPO to strengthen grounding and reduce hallucinations. Experiments demonstrate state-of-the-art detection accuracy and trustworthy, expert-aligned explanations. Kai He 0001, Qingyuan Lei, Bin Pu, Jian Zhang 0087, Yuling Xu, Mengling Feng |
ACL (1) | 4 |
| 2026 | Concept Relationship Embedding-Based Interactive Web Application for Explainable Medical DiagnosisabstractDeep learning has made remarkable progress in medical image analysis, yet its black-box nature still limits interpretability and clinician trust. Concept-based modeling offers a promising direction for explainable AI by integrating human-understandable concepts. However, existing approaches typically rely on global concept annotations and infer diagnosis based solely on the presence or absence of individual concepts. This oversimplified paradigm ignores the rich relationships among concepts and their causal influence on disease outcomes. To overcome these limitations, we propose the Concept Relationship Embedding Model (CREM) for interpretable medical diagnosis. CREM mirrors coarse-to-fine clinical reasoning by first extracting fine-grained subregional concepts, then explicitly encoding their relationships as a concept interaction graph, and finally performing causal inference between concepts and diagnoses to enable reliable and transparent diagnostic predictions. We evaluate CREM on four public medical imaging benchmarks, where it achieves state-of-the-art performance on both concept recognition and disease classification tasks, while exhibiting improved robustness, label efficiency, and interpretability. Furthermore, we deploy CREM as an interactive web-based demo that allows clinicians to visualize concept activations, trace diagnostic reasoning paths, and iteratively refine concept cues, facilitating effective human-in-the-loop decision-making. Lei Zhao 0013, Xingguo Lv, Qika Lin, Kaize Shi, Xiaoming Qi, Bin Pu, Kenli Li 0001 |
WWW | 6 |
| 2026 | Simple is what you need for efficient and accurate medical image segmentationabstractWhile modern segmentation models often prioritize performance over practicality, we advocate for a design philosophy that prioritizes simplicity and efficiency, and strive to design high-performance segmentation models. This paper presents SimpleUNet, a scalable, lightweight medical image segmentation framework. The key is that we proposed a simple yet effective partial feature selection mechanism for reducing information redundancy and thus facilitating compact model design. Additionally, we found that adjusting the model width is a straightforward yet easily overlooked tactic for lightweight model design, thereby preventing exponential parameter growth across network stages. By integrating an almost parameter-free channel attention module, the performance of the developed models can be improved with minimal overhead. Leveraging these techniques, our record-breaking model SimpleUNet with only 16 KB parameters surpasses LBUNet and other lightweight benchmarks across multiple public datasets. Impressively, the 0.67 MB variant achieves superior efficiency and accuracy, attaining a mean DSC/IoU of 85.76%/75.60% on a curated multi-center breast lesion dataset, surpassing both U-Net and TransUNet. Evaluations on skin lesion datasets (ISIC 2017/2018: mDice 84.86%/88.77%) and endoscopic polyp segmentation (KVASIR-SEG: 86.46%/76.48% mDice/mIoU) confirm consistent dominance over state-of-the-art models. Although our current SimpleUNet architecture does not rely on exotic or custom operators, it is fundamentally designed to embrace future innovations. The framework remains fully compatible with emerging operator-level advancements, allowing effortless integration and seamless upgrades without structural modifications. Codes can be found at https://github.com/Frankyu5666666/SimpleUNet . Yayan Chen, Guannan He, Qing Zeng 0005, Meiling Liang, Dandan Luo, Yimei Liao, Cheng Kang, Delong Yang, Bocheng Liang, Bin Pu, Shengli Li 0001 |
Expert Syst. Appl. | 13 |
| 2026 | A multilevel alignment and cross-fusion knowledge distillation framework for vision transformer-based medical image segmentation
Pengchen Liang, Jianguo Chen 0001, Renkai Wu, Zhuangzhuang Chen, Bin Pu, Qing Chang 0004, Guo Ran |
Future Gener. Comput. Syst. | 6 |
| 2026 | ToMo-UDA++: Unsupervised Domain Adaptation for Anatomical Structure Detection Using Enhanced Topology and Morphology Knowledge
Bin Pu, Jiewen Yang, Xingguo Lv, Xingbo Dong, Lei Zhao 0013, Shengli Li 0001, Kenli Li 0001, Xiaomeng Li 0001 |
Int. J. Comput. Vis. | 1 |
| 2026 | Dual dynamic graph attention network driven deep reinforcement learning for flexible job-Shop scheduling
Yan Kang 0003, Tianjing Li, Lei Zhao 0013, Zhuangzhuang Chen, Bin Pu |
Knowl. Based Syst. | 7 |
| 2026 | Task-specific knowledge distillation from the vision foundation model for enhanced medical image segmentation
Pengchen Liang, Haishan Huang, Bin Pu, Quanhong Zeng, Jianguo Chen 0001 |
Knowl. Based Syst. | 3 |
| 2026 | CFS-SMOTE: A cluster sample filtering-based synthetic minority oversampling technique for imbalanced clinical data
Zhaozhao Xu, Panzheng Xu, Fangyuan Yang, Junding Sun, Pengchen Liang, Yudong Zhang 0001, Chaosheng Tang, Deguang Li, Bin Pu |
Knowl. Based Syst. | 9 |
| 2026 | MTLQ-ViT: Multi-granularity Tail-enhanced Logarithmic Quantization for Vision Transformers
Yan Kang 0003, Shouhao Xu, Qika Lin, Kai He 0001, Zhuangzhuang Chen, Bin Pu |
Pattern Recognit. | 7 |
| 2026 | Topology-Preserving retinal vascular segmentation via sparse persistent homology and MoE convolution
Benteng Ma, Xiaomeng Li 0001, Bin Pu, Kwang-Ting Cheng |
Pattern Recognit. | 3 |
| 2026 | CED: CLIP-guided entropy dynamics for robust test-time adaptation in harsh visual conditions
Liwen Wang 0002, Xingbo Dong, Yen-Lung Lai, Bin Pu, Zhao Liu 0009, Qika Lin, Zhe Jin 0001 |
Pattern Recognit. | 4 |
| 2026 | Collaborative Coarse-to-Fine Disease Learning With Discharge Summary Awareness for EHR Event PredictionabstractDeep learning-based models have been widely used to predict electronic health record (EHR) events by exploiting diagnostic characteristics. Despite significant progress, three limitations remain: 1) effectively modeling dynamic relationships among diseases, 2) fully leveraging diagnosis code ontologies from multiple perspectives, and 3) incorporating unstructured discharge summaries. To address these challenges, we propose a coarse-to-fine disease learning framework with patient notes for EHR event prediction, tailored to capture both dynamic and static disease characteristics. First, we construct a fine-grained dynamic disease graph by removing disease weakly correlated disease pairs based on co-occurrence distributions. Second, disease embeddings are refined by integrating coarse and fine-grained information within the hierarchical structure of ICD-9-CM codes. In addition, discharge summaries are combined with auxiliary patient notes for collaborative disease learning. Finally, gated recurrent units, location-based attention, and soft attention mechanisms are utilized to further enhance embedding representations. Experiments on two real-world EHR datasets, MIMIC-III and MIMIC-IV, demonstrate that our model consistently outperforms nine baseline methods in EHR prediction. The source code can be found at https://github.com/YNU-L/CCDLD. Yan Kang 0003, Zhuolun Li, Bin Pu, Xingbo Dong, Jiewen Yang, Lei Zhao 0013, Benteng Ma, Ningshu Li, Jianguo Chen 0001, Philip S. Yu |
IEEE Trans. Cybern. | 3 |
| 2026 | Incorporating Large Vision Model Distillation and Fuzzy Perception for Improving Disease DiagnosisabstractEarly disease diagnosis is critical for timely clinical intervention and treatment. Intelligent models have shown significant potential in addressing the challenges of misdiagnosis, especially given the shortage of experienced experts. However, there exists complexity of ultrasound image information, and subtle differences between positive and negative samples, combined with limited disease data, pose challenges in feature extraction, class imbalance, and the lack of representative positive class prototypes. To this end, we propose a large vision model distillation framework with fuzzy perception to improve rare disease diagnosis. Specifically, we first fine-tune the pre-trained Medical Segment Anything Model (MedSAM) to adapt it to the target domain. Through image augmentation and latent feature relationship distillation, we enhance feature extraction robustness and reduce inter-class ambiguity, which helps mitigate the impact of class imbalance on the lightweight student model. Second, as imaging style differences in ultrasound images are often more pronounced than subtle variations between positive and negative samples, we construct image style subsets and introduce a fuzzy style matching strategy to perceive these differences. Finally, we combine features from both the differential perception and semantic enhancement branches to strengthen disease classification. Extensive experiments on an internal fetal spina bifida dataset and three widely used imbalanced benign-malignant medical datasets demonstrate the effectiveness of the proposed method. Qika Lin, Huaxuan Wen, Bin Pu, Mengling Feng, Kenli Li 0001 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2026 | ÆMMamba: An Efficient Medical Segmentation Model With Edge EnhancementabstractMedical image segmentation is critical for disease diagnosis, treatment planning, and prognosis assessment, yet the complexity and diversity of medical images pose significant challenges to accurate segmentation. While Convolutional Neural Networks capture local features and Vision Transformers excel in the global context, both struggle with efficient long-range dependency modeling. Inspired by Mamba's State Space Modeling efficiency, we propose ÆMMamba, a novel multi-scale feature extraction framework built on the Mamba backbone network. ÆMMamba integrates several innovative modules: the Efficient Fusion Bridge (EFB) module, which employs a bidirectional state-space model and attention mechanisms to fuse multi-scale features; the Edge-Aware Module (EAM), which enhances low-level edge representation using Sobel-based edge extraction; and the Boundary Sensitive Decoder (BSD), which leverages inverse attention and residual convolutional layers to handle cross-level complex boundaries. ÆMMamba achieves state-of-the-art performance across 8 medical segmentation datasets. On polyp segmentation datasets (Kvasir, ClinicDB, ColonDB, EndoScene, ETIS), it records the highest mDice and mIoU scores, outperforming methods like MADGNet and Swin-UMamba, with a standout mDice of 72.22 on ETIS, the most challenging dataset in this domain. For lung and breast segmentation, ÆMMamba surpasses competitors such as H2Former and SwinUnet, achieving Dice scores of 84.24 on BUSI and 79.83 on COVID-19 Lung. And on the LGG brain MRI dataset, ÆMMamba attains an mDice of 87.25 and an mIoU of 79.31, outperforming all compared methods. Xingbo Dong, Iman Yi Liao, Zhe Jin 0001, Zhaozhao Xu, Bin Pu |
IEEE J. Biomed. Health Informatics | 7 |
| 2026 | CMIS: A Class-Aware Multi-Structure Instance Segmentation Model for Fetal Brain Ultrasound Images With Fuzzy Region-Based ConstraintsabstractFetal anatomical structure segmentation in ultrasound images is essential for biometric measurement and disease diagnosis. However, current methods focus on a specific plane or a few structures, whereas obstetricians diagnose by considering multiple structures from different planes. In addition, existing methods struggle with segmenting fuzzy regions, which leads to performance degradation. We propose a real-time segmentation method called Class-aware Multi-structure Instance Segmentation (CMIS), designed to segment 19 key structures in 3 fetal brain planes to support brain-disease diagnosis. We extract instance information and generate class-aware attention for each class instead of dense instances to save computing resources and provide more informative details. Then we implement cross-layer and multi-scale fusion to obtain detailed prototypes. Finally, we fuse global attention with local prototypes cropped by boxes to generate masks and randomly perturb the boxes during training to enhance robustness. Moreover, we propose a new fuzzy region-based constraint loss to address the challenge of structures with varying scales and fuzzy boundaries. Extensive experiments on a fetal brain dataset demonstrate that CMIS outperforms 13 competing baselines, with an mDice of 83.41$\pm$0.03% at 37 FPS. CMIS also excels in external experiments on a fetal heart ultrasound dataset, achieving a mDice of 85.73$\pm$0.02% . These results demonstrate the effectiveness of CMIS in segmenting complex anatomical structures in ultrasound and its potential for real-time clinical applications. CMIS is limited to 2D normal standard planes ($\geq$19 weeks). Thus, its generalization to abnormal cases and broader datasets remains to be investigated. Mingxing Duan, Yuhuan Lu 0002, Bin Pu, Shuihua Wang, Kenli Li 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Leveraging Anatomical Consistency for Multi-Object Detection in Ultrasound Images via Source-free Unsupervised Domain AdaptationabstractSource-free unsupervised domain adaptation aims to eliminate domain shifts when data from the source domain and annotation from the target domain are not available. The multi-object detection tasks in medical image analysis are constrained by patient privacy and extremely huge annotation consumption. Hence, Source-free UDA is considered a more practical approach for eliminating the domain gap. However, relevant research that explores this topic is a dearth. In this paper, we design an Anatomy-aware Alignment Teacher-Student learning method using topological consistency based on a mean-teacher framework for Source-free UDA in multiple medical object detection named AATS, including Unsupervised Structure Refinement (USR) and Graph-aware Morphology Alignment (GMA). To match the student and teacher at the low-level and visual features, we propose the USR via an unsupervised clustering algorithm to group organs in ultrasound images. Based on USR, we obtain a graph with organ relations on the teacher branch. While in the student branch, we acquire visual features to construct graphical space and optimize the model with graph propagation. Finally, to match the student and teacher, GMA is designed to align the teacher and student based on both topology and morphology information that is derived from prior medical knowledge. Four groups of adaptation experiments were conducted on available medical datasets, and the outcomes demonstrate that our approach not only achieves state-of-the-art performance but also provides substantial advantages over existing methods. Bin Pu, Xingguo Lv, Jiewen Yang, Xingbo Dong, Yiqun Lin, Shengli Li 0001, Kenli Li 0001, Xiaomeng Li 0001 |
AAAI | 1 |
| 2025 | Anatomical Knowledge Mining and Matching for Semi-supervised Medical Multi-structure DetectionabstractIn medical image analysis, detecting multiple structures is crucial for evaluations and diagnosis but is often limited by the lack of high-quality annotations. Semi-supervised object detection emerges as a potent methodology to enhance model performance and generalization by leveraging a vast pool of unlabeled data alongside a minimal set of labeled data. A striking observation is that both unlabelled and labeled medical images contain a priori anatomical knowledge from human screening. In this work, we introduce a novel semi-supervised approach named Semi-akmm for mining and matching anatomical knowledge in ultrasound images. We develop an Adaptive Prior Knowledge Transfer (APKT) module to mine and explore the distribution and knowledge of potential proposal boxes by proposal proportion constraint. Furthermore, within a teacher-student learning framework, we put forward an Anatomical Structure Matching (ASM) module to facilitate co-learning consistent topological prior knowledge between the student and teacher models. To our knowledge, this marks the inception of an efficient semi-supervised medical multi-structure detection model. Our experiments across five publicly available ultrasound datasets demonstrate that Semi-akmm sets a new benchmark in performance with solid results that outperform existing methods. Bin Pu, Liwen Wang 0002, Jiewen Yang, Xingbo Dong, Benteng Ma, Zhuangzhuang Chen, Lei Zhao 0013, Shengli Li 0001, Kenli Li 0001 |
AAAI | 1 |
| 2025 | Direct Cardiovascular Disease Diagnosis From Multi-Modal Multi-View Ultrasound Via Unified Vision-Language ModelingabstractCardiovascular disease diagnosis via ultrasound screening relies on manually measured metrics and the experience level of human experts, which is time-consuming and may overlook subtle cross-anatomical pathological patterns. Recent vision-language models offer end-to-end diagnostic potential but lack mechanisms to handle heterogeneous multi-modal, multiview ultrasound data while preserving modality-specific semantics. To fill this gap, we propose an end-to-end framework called MMVL that directly fuses raw ultrasound sequences from diverse anatomical regions, bypassing intermediate measurements, and enabling direct diagnosis. We design lightweight adapters for domain-specific multi-modal feature fusion and refinement, a gating mechanism that dynamically reweights modality importance based on global context, and disease-aware prompt-guided classification. MMVL ensures robust performance across both common and rare conditions. The proposed multi-view, multimodal vision-language framework enables end-to-end cardiovascular disease diagnosis with a 10.9% accuracy gain, and opens a new avenue for automated and generalizable diagnostic solutions. Bin Pu, Jiewen Yang, Hangcheng Cao, Xingguo Lv, Lei Zhao 0013, Qika Lin, Yifan Zhu 0001, Kenli Li 0001 |
BIBM | 1 |
| 2025 | Test-Time Domain Generalization via Universe Learning: A Multi-Graph Matching Approach for Medical Image SegmentationabstractDespite domain generalization (DG) has significantly addressed the performance degradation of pre-trained models caused by domain shifts, it often falls short in real-world deployment. Test-time adaptation (TTA), which adjusts a learned model using unlabeled test data, presents a promising solution. However, most existing TTA methods struggle to deliver strong performance in medical image segmentation, primarily because they overlook the crucial prior knowledge inherent to medical images. To address this challenge, we incorporate morphological information and propose a framework based on multi-graph matching. Specifically, we introduce learnable universe embeddings that integrate morphological priors during multi-source training, along with novel unsupervised test-time paradigms for domain adaptation. This approach guarantees cycle-consistency in multi-matching while enabling the model to more effectively capture the invariant priors of unseen data, significantly mitigating the effects of domain shifts. Extensive experiments demonstrate that our method outperforms other state-of-the-art approaches on two medical image segmentation benchmarks for both multi-source and single-source domain generalization tasks. The source code is available at https://github.com/Yore0/TTDG-MGM. Xingguo Lv, Xingbo Dong, Liwen Wang 0002, Jiewen Yang, Lei Zhao 0013, Bin Pu, Zhe Jin 0001, Xuejun Li 0001 |
CVPR | 6 |
| 2025 | EA-KD: Entropy-Based Adaptive Knowledge DistillationabstractKnowledge distillation (KD) enables a smaller 'student' model to mimic a larger 'teacher' model by transferring knowledge from the teacher's output or features. However, most KD methods treat all samples uniformly, overlooking the varying learning value of each sample and thereby limiting effectiveness. In this paper, we propose Entropy- based Adaptive Knowledge Distillation (EA-KD), a simple yet effective plug-and-play KD method that prioritizes learning from valuable samples. EA-KD quantifies each sample's learning value by strategically combining the entropy of the teacher and student output, then dynamically reweights the distillation loss to place greater emphasis on high-entropy samples. Extensive experiments across diverse KD frameworks and tasks-including image classification, object detection, and large language model (LLM) distillation-demonstrate that EA-KD consistently enhances performance, achieving state-of-the-art results with negligible computational cost. Our code is available at https://github.com/cpsu00/EA-KD. Chi-Ping Su, Ching-Hsun Tseng, Bin Pu, Lei Zhao 0013, Jiewen Yang, Zhuangzhuang Chen, Shin-Jye Lee |
ICCV | 3 |
| 2025 | Personalized Federated Side-Tuning for Medical Image Classification
Jiayi Chen 0006, Benteng Ma, Yongsheng Pan, Bin Pu, Hengfei Cui, Yong Xia 0001 |
MICCAI (14) | 4 |
| 2025 | Concept-Induced Graph Perception Model for Interpretable Diagnosis
Lei Zhao 0013, Changjian Chen, Bin Pu, Xiaoming Qi, Fengfeng Peng, Chunlian Wang, Kenli Li 0001, Guanghua Tan |
MICCAI (12) | 3 |
| 2025 | Anatomical Structure Few-Shot Detection Utilizing Enhanced Human Anatomy Knowledge in Ultrasound Images
Bocheng Liang, Ningshu Li, Lei Zhao 0013, Hao Li 0021, Fengwei Yang, Bin Pu |
MICCAI (5) | 8 |
| 2025 | Learning to Zoom with Anatomical Relations for Medical Structure DetectionabstractAccurate anatomical structure detection is a critical preliminary step for diagnosing diseases characterized by structural abnormalities. In clinical practice, medical experts frequently adjust the zoom level of medical images to obtain comprehensive views for diagnosis. This common interaction results in significant variations in the apparent scale of anatomical structures across different images or fields of view. However, the information embedded in these zoom-induced scale changes is often overlooked by existing detection algorithms.
In addition, human organs possess a priori, fixed topological knowledge. To overcome this limitation, we propose ZR-DETR, a zoom-aware probabilistic framework tailored for medical object detection. ZR-DETR uniquely incorporates scale-sensitive zoom embeddings, anatomical relation constraints, and a Gaussian Process-based detection head. This architecture enables the framework to jointly model semantic context, enforce anatomical plausibility, and quantify detection uncertainty. Empirical validation across three diverse medical imaging benchmarks demonstrates that ZR-DETR consistently outperforms strong baselines in both single-domain and unsupervised domain adaptation scenarios. Bin Pu, Liwen Wang 0002, Xingbo Dong, Xingguo Lv, Zhe Jin 0001 |
NeurIPS | 1 |
| 2025 | CSP-SAM: CNN-Enhanced and Self-prompting SAM for Ultrasound Anatomical Structure Segmentation
Xingbo Dong, Bocheng Liang, Bin Pu, Zhe Jin 0001 |
PRCV (13) | 5 |
| 2025 | ThyFusion: A lightweight attribute enhancement module for thyroid nodule diagnosis using gradient and frequency-domain awareness
Guanyuan Chen, Ningbo Zhu, Bin Pu, Hongxia Luo, Kenli Li 0001 |
Neurocomputing | 4 |
| 2025 | Anatomical structures detection using topological constraint knowledge in fetal ultrasound
Juncheng Guo, Guanghua Tan, Bin Pu, Chunlian Wang, Shengli Li 0001, Kenli Li 0001 |
Neurocomputing | 4 |
| 2025 | A rule-guided interpretable lightweight framework for fetal standard ultrasound plane capture and biometric measurement
Jintang Li, Chunlian Wang, Bin Pu, Kenli Li 0001 |
Neurocomputing | 4 |
| 2025 | A key instance-guided frame-to-video information fusion network for thyroid ultrasound video instance segmentation
Guanyuan Chen, Ningbo Zhu, Bin Pu, Guanghua Tan, Hongxia Luo, Kenli Li 0001 |
Knowl. Based Syst. | 3 |
| 2025 | Low-light image enhancement with luminance duality
Xingguo Lv, Xingbo Dong, Jiewen Yang, Lei Zhao 0013, Bin Pu, Zhe Jin 0001 |
Knowl. Based Syst. | 5 |
| 2025 | DH-GAC: deep hierarchical context fusion network with modified geodesic active contour for multiple neurofibromatosis segmentation
Xiangqiong Wu, Guanghua Tan, Bin Pu, Mingxing Duan, Wenli Cai |
Neural Comput. Appl. | 3 |
| 2025 | TS-RePSO: A Three-Stage Feature Selection Method Combing ReliefF and PSO in BioinformaticsabstractThe inherent characteristics of high-dimensional feature redundancy of biomedical data lead to the "curse of dimensionality" in bioinformatics, which brings new challenges to feature selection problems. Recently, the two-stage approach combining the filter and wrapper methods has become popular for feature selection tasks. However, these two-stage or previous one-stage algorithms suffer from blindness in the setting of thresholds, and the search methods tend to fall into local optimum solutions. To this end, we propose a three-stage feature selection method that combines ReliefF and Particle swarm optimization as a specific case, called TS-RePSO, including the filter stage, grouping stage, and wrapper stage. Specifically, in the filter stage, ReliefF is utilized to compute the weights of the features and sort them in descending order. In the grouping stage, the ranked features are grouped based on the density equalization strategy so that the weight of groups in all groups is equal. In the wrapper stage, the proposed grouping PSO is employed to search for the grouped features and select them according to the in-group and out-group evaluation strategies. Extensive experiments are conducted on 5 benchmark datasets and 6 real-world datasets, and experiment results show that the proposed method achieves the best performance. Bin Pu, Haining Wang 0006, Zhaozhao Xu, Fangyuan Yang, Xiangqiong Wu, Jianguo Chen 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2025 | TKR-FSOD: Fetal Anatomical Structure Few-Shot Detection Utilizing Topological Knowledge ReasoningabstractFetal multi-anatomical structure detection in ultrasound (US) images can clearly present the relationship and influence between anatomical structures, providing more comprehensive information about fetal organ structures and assisting sonographers in making more accurate diagnoses, widely used in structure evaluation. Recently, deep learning methods have shown superior performance in detecting various anatomical structures in ultrasound images, but still have the potential for performance improvement in categories where it is difficult to obtain samples, such as rare diseases. Few-shot learning has attracted a lot of attention in medical image analysis due to its ability to solve the problem of data scarcity. However, existing few-shot learning research in medical image analysis focuses on classification and segmentation, and the research on object detection has been neglected. In this paper, we propose a novel fetal anatomical structure few-shot detection method in ultrasound images, TKR-FSOD, which learns topological knowledge through a Topological Knowledge Reasoning Module to help the model reason about and detect anatomical structures. Furthermore, we propose a Discriminate Ability Enhanced Feature Learning Module that extracts abundant discriminative features to enhance the model's discriminative ability. Experimental results demonstrate that our method outperforms the state-of-the-art baseline methods, exceeding the second-best method with a maximum margin of 4.8% on 5-shot of split 1 under four-chamber cardiac view. Bocheng Liang, Bin Pu, Jiewen Yang, Lei Zhao 0013, Yanqing Kong, Lixian Yang, Rentie Zhang, Hao Li 0021, Shengli Li 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | MambaSAM: A Visual Mamba-Adapted SAM Framework for Medical Image SegmentationabstractThe Segment Anything Model (SAM) has shown exceptional versatility in segmentation tasks across various natural image scenarios. However, its application to medical image segmentation poses significant challenges due to the intricate anatomical details and domain-specific characteristics inherent in medical images. To address these challenges, we propose a novel VMamba adapter framework that integrates a lightweight, trainable Visual Mamba (VMamba) branch with the pre-trained SAM ViT encoder. The VMamba adapter accurately captures multi-scale contextual correlations, integrates global and local information, and reduces ambiguities arising from local features only. Specifically, we propose a novel cross-branch attention (CBA) mechanism to facilitate effective interaction between the SAM and VMamba branches. This mechanism enables the model to learn and adapt more efficiently to the nuances of medical images, extracting rich, complementary features that enhance its representational capacity. Beyond architectural enhancements, we streamline the segmentation workflow by eliminating the need for prompt-driven input mechanisms. This results in an autonomous prediction model that reduces manual input requirements and improves operational efficiency. In addition, our method introduces only minimal additional trainable parameters, offering an efficient solution for medical image segmentation. Extensive evaluations of four medical image datasets demonstrate that our VMamba adapter framework achieves state-of-the-art performance. Specifically, on the ACDC dataset with limited training data, our method achieves an average Dice coefficient improvement of 0.18 and reduces the Hausdorff distance by 20.38 mm compared to the AutoSAM. Pengchen Liang, Leijun Shi, Bin Pu, Renkai Wu, Jianguo Chen 0001, Lite Xu, Zhuangzhuang Chen, Qing Chang 0004 |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Optical Flow-Enhanced Mamba U-Net for Cardiac Phase Detection in Ultrasound VideosabstractThe detection of cardiac phase in ultrasound videos, identifying end-systolic (ES) and end-diastolic (ED) frames, is a critical step in assessing cardiac function, monitoring structural changes, and diagnosing congenital heart disease. Current popular methods use recurrent neural networks to track dependencies over long sequences for cardiac phase detection, but often overlook the short-term motion of cardiac valves that sonographers rely on. In this paper, we propose a novel optical flow-enhanced Mamba U-net framework, designed to utilize both short-term motion and long-term dependencies to detect the cardiac phase in ultrasound videos. We utilize optical flow to capture the short-term motion of cardiac muscles and valves between adjacent frames, enhancing the input video. The Mamba layer is employed to track long-term dependencies across cardiac cycles. We then develop regression branches using the U-Net architecture, which integrates short-term and long-term information while extracting multi-scale features. Using this method, we can generate regression scores for each frame and identify keyframes (i.e., ES and ED frames). Additionally, we design a keyframe weighted loss function to guide the network to focus more on keyframes rather than intermediate period frames. Our method demonstrates superior performance compared to advanced baseline methods, achieving frame mismatches of 1.465 frames for ES and 0.842 frames for ED in the Fetal Echocardiogram dataset, where heart rates are higher and phase changes occur rapidly, and 2.444 frames and 2.072 frames in the publicly available adult Echonet-Dynamic dataset. Its accuracy and robustness in both fetal and adult datasets highlight its potential for clinical application. Yuhuan Lu 0002, Guanghua Tan, Bin Pu, Pak-Hei Yeung, Shengli Li 0001, Jagath C. Rajapakse, Kenli Li 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2024 | M3-UDA: A New Benchmark for Unsupervised Domain Adaptive Fetal Cardiac Structure DetectionabstractThe anatomical structure detection of fetal cardiac views is crucial for diagnosing fetal congenital heart disease. In practice, there is a large domain gap between different hospitals' data, such as the variable data quality due to differences in acquisition equipment. In addition, accurate annotation information provided by obstetrician experts is always very costly or even unavailable. This study explores the unsupervised domain adaptive fetal cardiac structure detection issue. Existing unsupervised domain adaptive object detection (UDAOD) approaches mainly focus on detecting objects in natural scenes, such as Foggy Cityscapes, where the structural relationships of natural scenes are uncertain. Unlike all previous UDAOD scenarios, we first collected a Fetal Cardiac Structure dataset from two hospital centers, called FCS, and proposed a multi-matching UDA approach (M3-UDA), including Histogram Matching (HM), Sub-structure Matching (SM), and Global-structure Matching (GM), to better transfer the topological knowledge of anatomical structure for UDA detection in medical scenarios. HM mitigates the domain gap between the source and target caused by pixel transformation. SM fuses the different angle information of the sub-structure to obtain the local topological knowledge for bridging the domain gap of the internal sub-structure. GM is designed to align the global topological knowledge of the whole organ from the source and target domain. Extensive experiments on our collected FCS and CardiacUDA, and experimental results show that M3-UDA outperforms existing UDAOD studies significantly. Datasets and source code are available at https://github.com/xmed-lab/M3-UDA. Bin Pu, Liwen Wang 0002, Jiewen Yang, Guannan He, Xingbo Dong, Shengli Li 0001, Zhe Jin 0001, Kenli Li 0001, Xiaomeng Li 0001 |
CVPR | 1 |
| 2024 | CardiacNet: Learning to Reconstruct Abnormalities for Cardiac Disease Assessment from Echocardiogram Videos
Jiewen Yang, Yiqun Lin, Bin Pu, Jiarong Guo, Xiaowei Xu 0004, Xiaomeng Li 0001 |
ECCV (23) | 3 |
| 2024 | Unsupervised Domain Adaptation for Anatomical Structure Detection in Ultrasound ImagesabstractModels trained on ultrasound images from one institution typically experience a decline in effectiveness when transferred directly to other institutions. Moreover, unlike natural images, dense and overlapped structures exist in fetus ultrasound images, making the detection of structures more challenging. Thus, to tackle this problem, we propose a new Unsupervised Domain Adaptation (UDA) method named ToMo-UDA for fetus structure detection, which consists of the Topology Knowledge Transfer (TKT) and the Morphology Knowledge Transfer (MKT) module. The TKT leverages prior knowledge of the medical anatomy of fetal as topological information, reconstructing and aligning anatomy features across source and target domains. Then, the MKT formulates a more consistent and independent morphological representation for each substructure of an organ. To evaluate the proposed ToMo-UDA for ultrasound fetal anatomical structure detection, we introduce FUSH$^2$, a new Fetal UltraSound benchmark, comprises Heart and Head images collected from Two health centers, with 16 annotated regions. Our experiments show that utilizing topological and morphological anatomy information in ToMo-UDA can greatly improve organ structure detection. This expands the potential for structure detection tasks in medical image analysis. Bin Pu, Xingguo Lv, Jiewen Yang, Guannan He, Xingbo Dong, Yiqun Lin, Shengli Li 0001, Tan Ying, Zhe Jin 0001, Kenli Li 0001, Xiaomeng Li 0001 |
ICML | 1 |
| 2024 | Boosting the Transferability of Adversarial Examples via Adaptive Attention and Gradient Purification MethodsabstractDeep neural networks are shown to be vulnerable to adversarial examples. Recently, various methods have been proposed to improve the transferability of adversarial examples. However, most of the existing methods add perturbations to the whole image without discrimination, causing the visual quality of the adversarial examples to degrade drastically. In addition, existing attack methods ignore the gradient information of secondary features, which affects the accuracy of generating adversarial perturbations. In this work, we propose Adaptive Attention and Gradient Purification Attack (AAGP) to address such issues. Specifically, we judge the mean and standard deviation of the gradient values to find out where the model is interested. Since different models share similar regions of attention, adding perturbations only to such areas can reduce the addition of adversarial perturbation and can also lead to better transferability of adversarial examples to other models. In addition, we disrupt the correlation of pixels at the distribution of secondary features by random discarding pixels in low-attention areas, generating more transferable perturbations through more accurate gradient information. Experimental results on ImageNet show that our method enhances the visibility of the adversarial examples and their transferability compared with several advanced baselines. Liwen Wu, Lei Zhao 0013, Bin Pu, Xin Jin 0005, Shaowen Yao 0001 |
IJCNN | 3 |
| 2024 | Bidirectional Recurrence for Cardiac Motion Tracking with Gaussian Process Latent CodingabstractQuantitative analysis of cardiac motion is crucial for assessing cardiac function. This analysis typically uses imaging modalities such as MRI and Echocardiograms that capture detailed image sequences throughout the heartbeat cycle. Previous methods predominantly focused on the analysis of image pairs lacking consideration of the motion dynamics and spatial variability. Consequently, these methods often overlook the long-term relationships and regional motion characteristic of cardiac. To overcome these limitations, we introduce the GPTrack, a novel unsupervised framework crafted to fully explore the temporal and spatial dynamics of cardiac motion. The GPTrack enhances motion tracking by employing the sequential Gaussian Process in the latent space and encoding statistics by spatial information at each time stamp, which robustly promotes temporal consistency and spatial variability of cardiac dynamics. Also, we innovatively aggregate sequential information in a bidirectional recursive manner, mimicking the behavior of diffeomorphic registration to better capture consistent long-term relationships of motions across cardiac regions such as the ventricles and atria. Our GPTrack significantly improves the precision of motion tracking in both 3D and 4D medical images while maintaining computational efficiency. The code is available at: https://github.com/xmed-lab/GPTrack. Jiewen Yang, Yiqun Lin, Bin Pu, Xiaomeng Li 0001 |
NeurIPS | 3 |
| 2024 | Learning Frequency and Structure in UDA for Medical Object Detection
Liwen Wang 0002, Guannan He, Shengli Li 0001, Bin Pu, Zhe Jin 0001, Wen Sha, Xingbo Dong |
PRCV (14) | 6 |
| 2024 | Graph-enhanced ensembles of multi-scale structure perception deep architecture for fetal ultrasound plane recognition
Guanghua Tan, Chunlian Wang, Bin Pu, Shengli Li 0001, Kenli Li 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | HICL: Hierarchical Intent Contrastive Learning for sequential recommendation
Yan Kang 0003, Yancong Yuan, Bin Pu, Yun Yang 0003, Lei Zhao 0013 |
Expert Syst. Appl. | 3 |
| 2024 | Boosting the Transferability of Ensemble Adversarial Attack via Stochastic Average Variance DescentabstractAdversarial examples have the property of transferring across models, which has created a great threat for deep learning models. To reveal the shortcomings in the existing deep learning models, the method of the ensemble has been introduced to the generating of transferable adversarial examples. However, most of the model ensemble attacks directly combine the different models’ output but ignore the large differences in optimization direction of them, which severely limits the transfer attack ability. In this work, we propose a new kind of ensemble attack method called stochastic average ensemble attack. Unlike the existing approach of averaging the outputs of each model as an integrated output, we continuously optimize the ensemble gradient in an internal loop using the model history gradient and the average gradient of different models. In this way, the adversarial examples can be updated in a more appropriate direction and make the crafted adversarial examples more transferable. Experimental results on ImageNet show that our method generates highly transferable adversarial examples and outperforms existing methods. Lei Zhao 0013, Zhizhi Liu, Sixing Wu, Liwen Wu, Bin Pu, Shaowen Yao 0001 |
IET Inf. Secur. | 6 |
| 2024 | A knowledge-interpretable multi-task learning framework for automated thyroid nodule diagnosis in ultrasound videos
Xiangqiong Wu, Guanghua Tan, Hongxia Luo, Zhilun Chen, Bin Pu, Shengli Li 0001, Kenli Li 0001 |
Medical Image Anal. | 5 |
| 2024 | A YOLOX-Based Deep Instance Segmentation Neural Network for Cardiac Anatomical Structures in Fetal Ultrasound ImagesabstractEchocardiography is an essential procedure for the prenatal examination of the fetus for congenital heart disease (CHD). Accurate segmentation of key anatomical structures in a four-chamber view is an essential step in measuring fetal growth parameters and diagnosing CHD. Currently, most obstetricians perform segmentation tasks manually, but the pixel-level operation is labor-intensive and requires extensive anatomical knowledge and clinical experience. As such, efficiently and accurately detecting structures from real-world fetal ultrasound images is a key challenge. In this paper, we propose a YOLOX-based deep instance segmentation neural network (i.e., IS-YOLOX) for cardiac anatomical structure location and segmentation in fetal ultrasound images. Specifically, we reconstruct a new instance segmentation branch based on a multi-task deep learning framework. We then design a new multi-level non-maximum suppression (NMS) mechanism to further improve the segmentation performance that consists of three levels of selection. Moreover, unlike two-stage instance segmentation approaches, our method does not rely on object detection results. To the best of our knowledge, this is the first study regarding instance segmentation on 13 types of anatomical structures in the fetal four-chamber view. Extensive experiments were carried out on clinical datasets, and the experimental results show that our method outperforms nine competitive baselines. Yuhuan Lu 0002, Kenli Li 0001, Bin Pu, Ningbo Zhu |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | MVSTT: A Multiview Spatial-Temporal Transformer Network for Traffic-Flow ForecastingabstractAccurate traffic-flow prediction remains a critical challenge due to complicated spatial dependencies, temporal factors, and unpredictable events. Most existing approaches focus on single- or dual-view learning and thus face limitations in systematically learning complex spatial-temporal features. In this work, we propose a novel multiview spatial-temporal transformer (MVSTT) network that can effectively learn complex spatial-temporal domain correlations and potential patterns from multiple views. First, we examine a temporal view and design a short-range gated convolution component from a short-term subview, and a long-range gated convolution component from a long-term subview. These two components effectively aggregate knowledge of the temporal domain at multiple granularities and mine patterns of node evolution across time steps. Meanwhile, in the spatial view, we design a dual-graph spatial learning module that captures fixed and dynamic spatial dependencies of nodes, as well as the evolution patterns of edges, from the static and dynamic graph subviews, respectively. In addition, we further design a spatial-temporal transformer to mine different levels of spatial-temporal features through multiview knowledge fusion. Extensive experiments on four real-world traffic datasets show that our method consistently outperforms the state-of-the-art baseline. The code of MVSTT is available at https://github.com/JianSoL/MVSTT. Bin Pu, Jiansong Liu, Yan Kang 0003, Jianguo Chen 0001, Philip S. Yu |
IEEE Trans. Cybern. | 1 |
| 2024 | MFISN: Modality Fuzzy Information Separation Network for Disease ClassificationabstractMost of the previous machine learning-based models for multi-modal medical diagnosis, primarily designed for unimodal images, usually do not fully leverage the potential of multimodal medical images, leading to limited classification accuracy. These conventional methods typically focus only on the intermodality common information, neglecting the intra-modality specific information and assuming that the common information is more effective in disease diagnosis. Moreover, they do not adequately address the impact of fuzzy information between different medical imaging modalities on diagnostic results. To this end, we propose a Modality Fuzzy Information Separation Network for disease classification, which extracts both common and specific information from fuzzy information to construct a comprehensive representation of multi-modal medical images. Specifically, we extract modality invariant features as common information by explicitly modeling and maximizing loss constraints on mutual information. For specific information extraction, a constraint on feature space independence between specific and common information is imposed on each modality. Above two steps, we concatenate common information and specific information to construct a comprehensive multi-modal representation for separating fuzzy information. Finally, we purposely design a decoder network to reconstruct medical images from uni-modal specific information and common information to demonstrate the effectiveness of the modality fuzzy information separation network. We conducted a validation of the proposed method's performance in classifying cardiomegaly, pneumothorax, edema, and skin disease. The experimental results substantiate the effectiveness of our proposed approach. Fengtao Nan, Bin Pu, Yingchun Fan, Jiewen Yang, Xingbo Dong, Zhaozhao Xu, Shuihua Wang |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | SKGC: A General Semantic-Level Knowledge Guided Classification Framework for Fetal Congenital Heart DiseaseabstractCongenital heart disease (CHD) is the most common congenital disability affecting healthy development and growth, even resulting in pregnancy termination or fetal death. Recently, deep learning techniques have made remarkable progress to assist in diagnosing CHD. One very popular method is directly classifying fetal ultrasound images, recognized as abnormal and normal, which tends to focus more on global features and neglects semantic knowledge of anatomical structures. The other approach is segmentation-based diagnosis, which requires a large number of pixel-level annotation masks for training. However, the detailed pixel-level segmentation annotation is costly or even unavailable. Based on the above analysis, we propose SKGC, a universal framework to identify normal or abnormal four-chamber heart (4CH) images, guided by a few annotation masks, while improving accuracy remarkably. SKGC consists of a semantic-level knowledge extraction module (SKEM), a multi-knowledge fusion module (MFM), and a classification module (CM). SKEM is responsible for obtaining high-level semantic knowledge, serving as an abstract representation of the anatomical structures that obstetricians focus on. MFM is a lightweight but efficient module that fuses semantic-level knowledge with the original specific knowledge in ultrasound images. CM classifies the fused knowledge and can be replaced by any advanced classifier. Moreover, we design a new loss function that enhances the constraint between the foreground and background predictions, improving the quality of the semantic-level knowledge. Experimental results on the collected real-world NA-4CH and the publicly FEST datasets show that SKGC achieves impressive performance with the best accuracy of 99.68% and 95.40%, respectively. Notably, the accuracy improves from 74.68% to 88.14% using only 10 labeled masks. Yuhuan Lu 0002, Guanghua Tan, Bin Pu, Bocheng Liang, Kenli Li 0001, Jagath C. Rajapakse |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | HFSCCD: A Hybrid Neural Network for Fetal Standard Cardiac Cycle Detection in Ultrasound VideosabstractIn the fetal cardiac ultrasound examination, standard cardiac cycle (SCC) recognition is the essential foundation for diagnosing congenital heart disease. Previous studies have mostly focused on the detection of adult CCs, which may not be applicable to the fetus. In clinical practice, localization of SCCs needs to recognize end-systole (ES) and end-diastole (ED) frames accurately, ensuring that every frame in the cycle is a standard view. Most existing methods are not based on the detection of key anatomical structures, which may not recognize irrelevant views and background frames, results containing non-standard frames, or even it does not work in clinical practice. We propose an end-to-end hybrid neural network based on an object detector to detect SCCs from fetal ultrasound videos efficiently, which consists of 3 modules, namely Anatomical Structure Detection (ASD), Cardiac Cycle Localization (CCL), and Standard Plane Recognition (SPR). Specifically, ASD uses an object detector to identify 9 key anatomical structures, 3 cardiac motion phases, and the corresponding confidence scores from fetal ultrasound videos. On this basis, we propose a joint probability method in the CCL to learn the cardiac motion cycle based on the 3 cardiac motion phases. In SPR, to reduce the impact of structure detection errors on the accuracy of the standard plane recognition, we use XGBoost algorithm to learn the relation knowledge of the detected anatomical structures. We evaluate our method on the test fetal ultrasound video datasets and clinical examination cases and achieve remarkable results. This study may pave the way for clinical practices. Bin Pu, Kenli Li 0001, Jianguo Chen 0001, Yuhuan Lu 0002, Qing Zeng 0005, Jiewen Yang, Shengli Li 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Real-Time Automatic M-Mode Echocardiography Measurement With Panel AttentionabstractMotion mode (M-mode) echocardiography is essential for measuring cardiac dimension and ejection fraction. However, the current diagnosis is time-consuming and suffers from diagnosis accuracy variance. This work resorts to building an automatic scheme through well-designed and well-trained deep learning to conquer the situation. That is, we proposed RAMEM, an automatic scheme of real-time M-mode echocardiography, which contributes three aspects to address the challenges: 1) provide MEIS, the first dataset of M-mode echocardiograms, to enable consistent results and support developing an automatic scheme; For detecting objects accurately in echocardiograms, it requires big receptive field for covering long-range diastole to systole cycle. However, the limited receptive field in the typical backbone of convolutional neural networks (CNN) and the losing information risk in non-local block (NL) equipped CNN risk the accuracy requirement. Therefore, we 2) propose panel attention embedding with updated UPANets V2, a convolutional backbone network, in a real-time instance segmentation (RIS) scheme for boosting big object detection performance; 3) introduce AMEM, an efficient algorithm of automatic M-mode echocardiography measurement, for automatic diagnosis; The experimental results show that RAMEM surpasses existing RIS schemes (CNNs with NL & Transformers as the backbone) in PASCAL 2012 SBD and human performances in MEIS. Ching-Hsun Tseng, Shao-Ju Chien, Po-Shen Wang, Shin-Jye Lee, Bin Pu, Xiaojun Zeng |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | TransFSM: Fetal Anatomy Segmentation and Biometric Measurement in Ultrasound Images Using a Hybrid TransformerabstractBiometric parameter measurements are powerful tools for evaluating a fetus's gestational age, growth pattern, and abnormalities in a 2D ultrasound. However, it is still challenging to measure fetal biometric parameters automatically due to the indiscriminate confusing factors, limited foreground-background contrast, variety of fetal anatomy shapes at different gestational ages, and blurry anatomical boundaries in ultrasound images. The performance of a standard CNN architecture is limited for these tasks due to the restricted receptive field. We propose a novel hybrid Transformer framework, TransFSM, to address fetal multi-anatomy segmentation and biometric measurement tasks. Unlike the vanilla Transformer based on a single-scale input, TransFSM has a deformable self-attention mechanism so it can effectively process multi-scale information to segment fetal anatomy with irregular shapes and different sizes. We devised a BAD to capture more intrinsic local details using boundary-wise prior knowledge, which compensates for the defects of the Transformer in extracting local features. In addition, a Transformer auxiliary segment head is designed to improve mask prediction by learning the semantic correspondence of the same pixel categories and feature discriminability among different pixel categories. Extensive experiments were conducted on clinical cases and benchmark datasets for anatomy segmentation and biometric measurement tasks. The experiment results indicate that our method achieves state-of-the-art performance in seven evaluation metrics compared with CNN-based, Transformer-based, and hybrid approaches. By Knowledge distillation, the proposed TransFSM can create a more compact and efficient model with high deploying potential in resource-constrained scenarios. Our study serves as a unified framework for biometric estimation across multiple anatomical regions to monitor fetal growth in clinical practice. Lei Zhao 0013, Guanghua Tan, Bin Pu, Qianghui Wu, Hongliang Ren 0001, Kenli Li 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | FARN: Fetal Anatomy Reasoning Network for Detection With Global Context Semantic and Local Topology RelationshipabstractAccurate recognition of fetal anatomical structure is a pivotal task in ultrasound (US) image analysis. Sonographers naturally apply anatomical knowledge and clinical expertise to recognizing key anatomical structures in complex US images. However, mainstream object detection approaches usually treat each structure recognition separately, overlooking anatomical correlations between different structures in fetal US planes. In this work, we propose a Fetal Anatomy Reasoning Network (FARN) that incorporates two kinds of relationship forms: a global context semantic block summarized with visual similarity and a local topology relationship block depicting structural pair constraints. Specifically, by designing the Adaptive Relation Graph Reasoning (ARGR) module, anatomical structures are treated as nodes, with two kinds of relationships between nodes modeled as edges. The flexibility of the model is enhanced by constructing the adaptive relationship graph in a data-driven way, enabling adaptation to various data samples without the need for predefined additional constraints. The feature representation is further refined by aggregating the outputs of the ARGR module. Comprehensive experimental results demonstrate that FARN achieves promising performance in detecting 37 anatomical structures across key US planes in tertiary obstetric screening. FARN effectively utilizes key relationships to improve detection performance, demonstrates robustness to small-scale, similar, and indistinct structures, and avoids some detection errors that deviate from anatomical norms. Overall, our study serves as a resource for developing efficient and concise approaches to model inter-anatomy relationships. Lei Zhao 0013, Guanghua Tan, Qianghui Wu, Bin Pu, Hongliang Ren 0001, Shengli Li 0001, Kenli Li 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | A Deep Graph Network with Multiple Similarity for User Clustering in Human-Computer InteractionabstractUser counterparts, such as user attributes in social networks or user interests, are the keys to more natural Human–Computer Interaction (HCI) . In addition, users’ attributes and social structures help us understand the complex interactions in HCI. Most previous studies have been based on supervised learning to improve the performance of HCI. However, in the real world, owing to signal malfunctions in user devices, large amounts of abnormal information, unlabeled data, and unsupervised approaches (e.g., the clustering method) based on mining user attributes are particularly crucial. This paper focuses on improving the clustering performance of users’ attributes in HCI and proposes a deep graph embedding network with feature and structure similarity (called DGENFS ) to cluster users’ attributes in HCI applications based on feature and structure similarity. The DGENFS model consists of a Feature Graph Autoencoder (FGA) module, a Structure Graph Attention Network (SGAT) module, and a Dual Self-supervision (DSS) module. First, we design an attributed graph clustering method to divide users into clusters by making full use of their attributes. To take full advantage of the information of human feature space, a k-neighbor graph is generated as a feature graph based on the similarity between human features. Then, the FGA and SGAT modules are utilized to extract the representations of human features and topological space, respectively. Next, an attention mechanism is further developed to learn the importance weights of different representations to effectively integrate human features and social structures. Finally, to learn cluster-friendly features, the DSS module unifies and integrates the features learned from the FGA and SGAT modules. DSS explores the high-confidence cluster assignment as a soft label to guide the optimization of the entire network. Extensive experiments are conducted on five real-world data sets on user attribute clustering. The experimental results demonstrate that the proposed DGENFS model achieves the most advanced performance compared with nine competitive baselines. Yan Kang 0003, Bin Pu, Yongqi Kou, Yun Yang 0003, Jianguo Chen 0001, Khan Muhammad 0001, Po Yang 0001, Mohammad Hijji |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | HN-PPISP: a hybrid network based on MLP-Mixer for protein-protein interaction site predictionabstractMOTIVATION: Biological experimental approaches to protein-protein interaction (PPI) site prediction are critical for understanding the mechanisms of biochemical processes but are time-consuming and laborious. With the development of Deep Learning (DL) techniques, the most popular Convolutional Neural Networks (CNN)-based methods have been proposed to address these problems. Although significant progress has been made, these methods still have limitations in encoding the characteristics of each amino acid in protein sequences. Current methods cannot efficiently explore the nature of Position Specific Scoring Matrix (PSSM), secondary structure and raw protein sequences by processing them all together. For PPI site prediction, how to effectively model the PPI context with attention to prediction remains an open problem. In addition, the long-distance dependencies of PPI features are important, which is very challenging for many CNN-based methods because the innate ability of CNN is difficult to outperform auto-regressive models like Transformers. RESULTS: To effectively mine the properties of PPI features, a novel hybrid neural network named HN-PPISP is proposed, which integrates a Multi-layer Perceptron Mixer (MLP-Mixer) module for local feature extraction and a two-stage multi-branch module for global feature capture. The model merits Transformer, TextCNN and Bi-LSTM as a powerful alternative for PPI site prediction. On the one hand, this is the first application of an advanced Transformer (i.e. MLP-Mixer) with a hybrid network for sequence-based PPI prediction. On the other hand, unlike existing methods that treat global features altogether, the proposed two-stage multi-branch hybrid module firstly assigns different attention scores to the input features and then encodes the feature through different branch modules. In the first stage, different improved attention modules are hybridized to extract features from the raw protein sequences, secondary structure and PSSM, respectively. In the second stage, a multi-branch network is designed to aggregate information from both branches in parallel. The two branches encode the features and extract dependencies through several operations such as TextCNN, Bi-LSTM and different activation functions. Experimental results on real-world public datasets show that our model consistently achieves state-of-the-art performance over seven remarkable baselines. AVAILABILITY: The source code of HN-PPISP model is available at https://github.com/ylxu05/HN-PPISP. Yan Kang 0003, Xinchao Wang, Bin Pu, Xuekun Yang, Yulong Rao, Jianguo Chen 0001 |
Briefings Bioinform. | 4 |
| 2023 | An end-to-end anti-shaking multi-focus image fusion approach
Jiayu Ji, Xuanyin Wang, Jixiang Tang, Ze'an Liu, Bin Pu |
Image Vis. Comput. | 6 |
| 2023 | TMHSCA: a novel hybrid two-stage mutation with a sine cosine algorithm for discounted {0-1} knapsack problems
Yan Kang 0003, Haining Wang 0006, Bin Pu, Jiansong Liu, Shin-Jye Lee, Xuekun Yang, Liu Tao |
Neural Comput. Appl. | 3 |
| 2023 | Correction to: TMHSCA: a novel hybrid two-stage mutation with a sine cosine algorithm for discounted {0-1} knapsack problems
Yan Kang 0003, Haining Wang 0006, Bin Pu, Jiansong Liu, Shin-Jye Lee, Xuekun Yang, Liu Tao |
Neural Comput. Appl. | 3 |
| 2023 | A Hybrid Two-Stage Teaching-Learning-Based Optimization Algorithm for Feature Selection in BioinformaticsabstractThe "curse of dimensionality" brings new challenges to the feature selection (FS) problem, especially in bioinformatics filed. In this paper, we propose a hybrid Two-Stage Teaching-Learning-Based Optimization (TS-TLBO) algorithm to improve the performance of bioinformatics data classification. In the selection reduction stage, potentially informative features, as well as noisy features, are selected to effectively reduce the search space. In the following comparative self-learning stage, the teacher and the worst student with self-learning evolve together based on the duality of the FS problems to enhance the exploitation capabilities. In addition, an opposition-based learning strategy is utilized to generate initial solutions to rapidly improve the quality of the solutions. We further develop a self-adaptive mutation mechanism to improve the search performance by dynamically adjusting the mutation rate according to the teacher's convergence ability. Moreover, we integrate a differential evolutionary method with TLBO to boost the exploration ability of our algorithm. We conduct comparative experiments on 31 public data sets with different data dimensions, including 7 bioinformatics datasets, and evaluate our TS-TLBO algorithm compared with 11 related methods. The experimental results show that the TS-TLBO algorithm obtains a good feature subset with better classification performance, and indicates its generality to the FS problems. Yan Kang 0003, Haining Wang 0006, Bin Pu, Liu Tao, Jianguo Chen 0001, Philip S. Yu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | An ultrasound standard plane detection model of fetal head based on multi-task learning and hybrid knowledge graph
Lei Zhao 0013, Kenli Li 0001, Bin Pu, Jianguo Chen 0001, Shengli Li 0001, Xiangke Liao |
Future Gener. Comput. Syst. | 3 |
| 2022 | MobileUNet-FPN: A Semantic Segmentation Model for Fetal Ultrasound Four-Chamber Segmentation in Edge Computing EnvironmentsabstractThe apical four-chamber (A4C) view in fetal echocardiography is a prenatal examination widely used for the early diagnosis of congenital heart disease (CHD). Accurate segmentation of A4C key anatomical structures is the basis for automatic measurement of growth parameters and necessary disease diagnosis. However, due to the ultrasound imaging arising from artefacts and scattered noise, the variability of anatomical structures in different gestational weeks, and the discontinuity of anatomical structure boundaries, accurately segmenting the fetal heart organ in the A4C view is a very challenging task. To this end, we propose to combine an explicit Feature Pyramid Network (FPN), MobileNet and UNet, i.e., MobileUNet-FPN, for the segmentation of 13 key heart structures. To our knowledge, this is the first AI-based method that can segment so many anatomical structures in fetal A4C view. We split the MobileNet backbone network into four stages and use the features of these four phases as the encoder and the upsampling operation as the decoder. We build an explicit FPN network to enhance multi-scale semantic information and ultimately generate segmentation masks of key anatomical structures. In addition, we design a multi-level edge computing system and deploy the distributed edge nodes in different hospitals and city servers, respectively. Then, we train the MobileUNet-FPN model in parallel at each edge node to effectively reduce the network communication overhead. Extensive experiments are conducted and the results show the superior performance of the proposed model on the fetal A4C and femoral-length images. Bin Pu, Yuhuan Lu 0002, Jianguo Chen 0001, Shengli Li 0001, Ningbo Zhu, Wei Wei 0006, Kenli Li 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | HWOA: an intelligent hybrid whale optimization algorithm for multi-objective task selection strategy in edge cloud computing system
Yan Kang 0003, Xuekun Yang, Bin Pu, Xiaokang Wang 0001, Haining Wang 0006, Puming Wang |
World Wide Web | 3 |
| 2021 | Fetal cardiac cycle detection in multi-resource echocardiograms using hybrid classification framework
Bin Pu, Ningbo Zhu, Kenli Li 0001, Shengli Li 0001 |
Future Gener. Comput. Syst. | 1 |
| 2021 | Automatic Fetal Ultrasound Standard Plane Recognition Based on Deep Learning and IIoTabstractIntelligent ultrasound imaging based on deep learning is one of the important applications in the field of intelligent medical care. In this article, we propose an automatic fetal ultrasound standard plane recognition (FUSPR) model based on deep learning in the Industrial Internet of Things (IIoT) environment. We build a distributed ultrasound data processing and predicting platform by using the IIoT and high-performance computing (HPC) technology. The FUSPR model deployed in the HPC center consists of a convolutional neural network (CNN) component and a recurrent neural network (RNN) component, which learns the spatial and temporal features of the ultrasound video stream by using multitask learning, respectively. The CNN component identifies fetal key anatomical structures from each video frame and accurately recognizes the potential four fetal standard planes. The RNN component obtains the temporal information between adjacent frames, and it realizes precise localization and tracking of fetal organs across frames. In addition, we introduce two feature fusion strategies into the FUSPR model, i.e., CNN fusion and RNN fusion, to fit the spatial sequence and motion representation in the video stream, thereby effectively improving the accuracy and robustness of the model. Extensive experiments conducted on more than 1000 ultrasound videos show that the FUSPR model is superior to the competing baselines in terms of accuracy and performance. Bin Pu, Kenli Li 0001, Shengli Li 0001, Ningbo Zhu |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | Energy-Efficient Target Tracking With UASNs: A Consensus-Based Bayesian ApproachabstractTarget tracking has been considered as one of the most important applications of underwater acoustic sensor networks. However, the long propagation delay, high-energy consumption, and strong noise properties of the underwater environment make target tracking more challenging as compared with terrestrial sensor networks. This article is concerned with an energy-efficient tracking issue for underwater targets, subject to an asynchronous clock, power restriction, and noise measurement constraints. The tracking process can be divided into two phases, i.e., position acquisition and persistent tracking. In the first phase, we establish the relationship between propagation delay and position, through which an asynchronous localization algorithm is developed for sensor nodes to estimate the position of target. Based on the estimated position, a consensus-based Bayesian filter is designed for sensor nodes in the second phase to enable persistent tracking. In particular, the consensus fusion strategy and duty-cycle mechanism are jointly adopted to improve the tracking accuracy and prolong the network lifetime. Moreover, the convergence analyses for the proposed approach are also presented. Finally, simulation and experimental results reveal that the proposed tracking approach can reduce the influence of malicious measurements, while the energy efficiency can be significantly improved as compared with the other works. Jing Yan 0001, Bin Pu, Xiaoyuan Luo, Cailian Chen, Xin-Ping Guan |
IEEE Trans Autom. Sci. Eng. | 3 |