VLDB 2026 Research / reviewers in the wild / expert
Lei Zhao 0013
dblp:87/734-13
· DBLP profile ↗
32ranked-venue papers
10as first author
30since 2021 · last 2026
0000-0002-2583-0081ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Topology-Inspired Backward-Free Framework for Test-Time Adaptation in Medical DetectionabstractRecently, Test-Time Adaptation (TTA) has gained increasing attention in medical imaging due to its ability to improve model generalization under domain shifts without retraining. In particular, directly applying a well-trained model across various medical centers faces significant performance degradation caused by variations in equipment, operators, imaging conditions, and scanning skill levels of sonographers. Existing TTA methods either rely on parameter adaptation that increases computational cost or apply simple prediction fusion that ignores anatomical structure knowledge. To address these limitations, we propose a novel backward-free Topology-aware TTA framework named T^3 that integrates Structural Perception Modeling (SPM) and Box Regression Adaptation (BRA). SPM is implemented through an organ space heatmap generated via Gaussian kernel superposition. This heatmap encodes anatomical topology without requiring additional training or source data. BRA further improves localization and classification by fusing detection outputs based on the contribution of detected results to anatomically meaningful peak points from the heatmaps. Extensive experiments were conducted across six cross-domain scenarios, and the results demonstrate that our method achieves state-of-the-art cross-domain detection performance while maintaining high efficiency, offering a practical and robust solution for real-world medical diagnostic applications. Bin Pu, Xingguo Lv, Jiewen Yang, Lei Zhao 0013, Zuozhu Liu, Kenli Li 0001 |
AAAI | 5 |
| 2026 | Organ-Aware Routing Mixture-of-Retrieval Augmented Generation for Fetal Ultrasound ReportingabstractFetal ultrasound screening is a uniquely complex diagnostic task involving the simultaneous assessment of multiple fetal organs—each with its own anatomical and clinical context—within a single examination. Automating report generation for such cases poses a significant challenge: unlike existing methods that focus on single-organ radiology tasks (e.g., chest X-rays), fetal ultrasound requires reasoning over a structured, multiple-to-multiple setting, i.e., multi-organ images corresponding to a multi-section report. In this paper, we introduce FetusR, the first large-scale dataset for multi-organ fetal ultrasound reporting, containing 15,594 real-world cases with rich organ-wise annotations. To address the intrinsic image-report alignment, we propose Organ-Aware Routing Mixture-of-Retrieval Augmented Generation (ORM-RAG) inspired by the Mixture-of-Experts paradigm. Our method decomposes the complex alignment problem into multiple one-to-one sub-retrieval tasks. Specifically, ORM-RAG integrates (1) an organ-aware mixture-of-retrieval module that partitions the retrieval space into organ-specific corpora for independent retrieval, and (2) a dynamic routing mechanism that selectively aggregates high-confidence organ-specific reports while filtering uncertain ones. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art baselines across both textual similarity and clinical accuracy metrics. Our work opens a new direction for long-form, structured report generation in real-world, multi-organ medical imaging scenarios. Bin Pu, Rongbin Li, Xinpeng Ding, Lei Zhao 0013, Chaoqi Chen, Shengli Li 0001, Kenli Li 0001 |
AAAI | 5 |
| 2026 | MPA: Multimodal Prototype Augmentation for Few-Shot LearningabstractRecently, Few-shot Learning (FSL) has become a popular task that aims to recognize new classes from only a few labeled examples and has been widely applied in fields such as natural science, remote sensing, and medical images. However, most existing methods focus only on the visual modality and compute prototypes directly from raw support images, which lack comprehensive and rich multimodal information. To address these limitations, we propose a novel Multimodal Prototype Augmentation FSL framework called MPA, including LLM-based Multi-Variant Semantic Enhancement (LMSE), Hierarchical Multi-View Augmentation (HMA), and an Adaptive Uncertain Class Absorber (AUCA). LMSE leverages large language models to generate diverse paraphrased category descriptions, enriching the support set with additional semantic cues. HMA exploits both natural and multi-view augmentations to enhance feature diversity (e.g., changes in viewing distance, camera angles, and lighting conditions). AUCA models uncertainty by introducing uncertain classes via interpolation and Gaussian sampling, effectively absorbing uncertain samples. Extensive experiments on four single-domain and six cross-domain FSL benchmarks demonstrate that MPA achieves superior performance compared to existing state-of-the-art methods across most settings. Notably, MPA surpasses the second-best method by 12.29% and 24.56% in the single-domain and cross-domain setting, respectively, in the 5-way 1-shot setting. Liwen Wu, Lei Zhao 0013, Qika Lin, Shaowen Yao 0001, Zuozhu Liu, Bin Pu |
AAAI | 3 |
| 2026 | Concept Relationship Embedding-Based Interactive Web Application for Explainable Medical DiagnosisabstractDeep learning has made remarkable progress in medical image analysis, yet its black-box nature still limits interpretability and clinician trust. Concept-based modeling offers a promising direction for explainable AI by integrating human-understandable concepts. However, existing approaches typically rely on global concept annotations and infer diagnosis based solely on the presence or absence of individual concepts. This oversimplified paradigm ignores the rich relationships among concepts and their causal influence on disease outcomes. To overcome these limitations, we propose the Concept Relationship Embedding Model (CREM) for interpretable medical diagnosis. CREM mirrors coarse-to-fine clinical reasoning by first extracting fine-grained subregional concepts, then explicitly encoding their relationships as a concept interaction graph, and finally performing causal inference between concepts and diagnoses to enable reliable and transparent diagnostic predictions. We evaluate CREM on four public medical imaging benchmarks, where it achieves state-of-the-art performance on both concept recognition and disease classification tasks, while exhibiting improved robustness, label efficiency, and interpretability. Furthermore, we deploy CREM as an interactive web-based demo that allows clinicians to visualize concept activations, trace diagnostic reasoning paths, and iteratively refine concept cues, facilitating effective human-in-the-loop decision-making. Lei Zhao 0013, Xingguo Lv, Qika Lin, Kaize Shi, Xiaoming Qi, Bin Pu, Kenli Li 0001 |
WWW | 1 |
| 2026 | ToMo-UDA++: Unsupervised Domain Adaptation for Anatomical Structure Detection Using Enhanced Topology and Morphology Knowledge
Bin Pu, Jiewen Yang, Xingguo Lv, Xingbo Dong, Lei Zhao 0013, Shengli Li 0001, Kenli Li 0001, Xiaomeng Li 0001 |
Int. J. Comput. Vis. | 5 |
| 2026 | Dual dynamic graph attention network driven deep reinforcement learning for flexible job-Shop scheduling
Yan Kang 0003, Tianjing Li, Lei Zhao 0013, Zhuangzhuang Chen, Bin Pu |
Knowl. Based Syst. | 3 |
| 2026 | Collaborative Coarse-to-Fine Disease Learning With Discharge Summary Awareness for EHR Event PredictionabstractDeep learning-based models have been widely used to predict electronic health record (EHR) events by exploiting diagnostic characteristics. Despite significant progress, three limitations remain: 1) effectively modeling dynamic relationships among diseases, 2) fully leveraging diagnosis code ontologies from multiple perspectives, and 3) incorporating unstructured discharge summaries. To address these challenges, we propose a coarse-to-fine disease learning framework with patient notes for EHR event prediction, tailored to capture both dynamic and static disease characteristics. First, we construct a fine-grained dynamic disease graph by removing disease weakly correlated disease pairs based on co-occurrence distributions. Second, disease embeddings are refined by integrating coarse and fine-grained information within the hierarchical structure of ICD-9-CM codes. In addition, discharge summaries are combined with auxiliary patient notes for collaborative disease learning. Finally, gated recurrent units, location-based attention, and soft attention mechanisms are utilized to further enhance embedding representations. Experiments on two real-world EHR datasets, MIMIC-III and MIMIC-IV, demonstrate that our model consistently outperforms nine baseline methods in EHR prediction. The source code can be found at https://github.com/YNU-L/CCDLD. Yan Kang 0003, Zhuolun Li, Bin Pu, Xingbo Dong, Jiewen Yang, Lei Zhao 0013, Benteng Ma, Ningshu Li, Jianguo Chen 0001, Philip S. Yu |
IEEE Trans. Cybern. | 6 |
| 2025 | Anatomical Knowledge Mining and Matching for Semi-supervised Medical Multi-structure DetectionabstractIn medical image analysis, detecting multiple structures is crucial for evaluations and diagnosis but is often limited by the lack of high-quality annotations. Semi-supervised object detection emerges as a potent methodology to enhance model performance and generalization by leveraging a vast pool of unlabeled data alongside a minimal set of labeled data. A striking observation is that both unlabelled and labeled medical images contain a priori anatomical knowledge from human screening. In this work, we introduce a novel semi-supervised approach named Semi-akmm for mining and matching anatomical knowledge in ultrasound images. We develop an Adaptive Prior Knowledge Transfer (APKT) module to mine and explore the distribution and knowledge of potential proposal boxes by proposal proportion constraint. Furthermore, within a teacher-student learning framework, we put forward an Anatomical Structure Matching (ASM) module to facilitate co-learning consistent topological prior knowledge between the student and teacher models. To our knowledge, this marks the inception of an efficient semi-supervised medical multi-structure detection model. Our experiments across five publicly available ultrasound datasets demonstrate that Semi-akmm sets a new benchmark in performance with solid results that outperform existing methods. Bin Pu, Liwen Wang 0002, Jiewen Yang, Xingbo Dong, Benteng Ma, Zhuangzhuang Chen, Lei Zhao 0013, Shengli Li 0001, Kenli Li 0001 |
AAAI | 7 |
| 2025 | Direct Cardiovascular Disease Diagnosis From Multi-Modal Multi-View Ultrasound Via Unified Vision-Language ModelingabstractCardiovascular disease diagnosis via ultrasound screening relies on manually measured metrics and the experience level of human experts, which is time-consuming and may overlook subtle cross-anatomical pathological patterns. Recent vision-language models offer end-to-end diagnostic potential but lack mechanisms to handle heterogeneous multi-modal, multiview ultrasound data while preserving modality-specific semantics. To fill this gap, we propose an end-to-end framework called MMVL that directly fuses raw ultrasound sequences from diverse anatomical regions, bypassing intermediate measurements, and enabling direct diagnosis. We design lightweight adapters for domain-specific multi-modal feature fusion and refinement, a gating mechanism that dynamically reweights modality importance based on global context, and disease-aware prompt-guided classification. MMVL ensures robust performance across both common and rare conditions. The proposed multi-view, multimodal vision-language framework enables end-to-end cardiovascular disease diagnosis with a 10.9% accuracy gain, and opens a new avenue for automated and generalizable diagnostic solutions. Bin Pu, Jiewen Yang, Hangcheng Cao, Xingguo Lv, Lei Zhao 0013, Qika Lin, Yifan Zhu 0001, Kenli Li 0001 |
BIBM | 5 |
| 2025 | Test-Time Domain Generalization via Universe Learning: A Multi-Graph Matching Approach for Medical Image SegmentationabstractDespite domain generalization (DG) has significantly addressed the performance degradation of pre-trained models caused by domain shifts, it often falls short in real-world deployment. Test-time adaptation (TTA), which adjusts a learned model using unlabeled test data, presents a promising solution. However, most existing TTA methods struggle to deliver strong performance in medical image segmentation, primarily because they overlook the crucial prior knowledge inherent to medical images. To address this challenge, we incorporate morphological information and propose a framework based on multi-graph matching. Specifically, we introduce learnable universe embeddings that integrate morphological priors during multi-source training, along with novel unsupervised test-time paradigms for domain adaptation. This approach guarantees cycle-consistency in multi-matching while enabling the model to more effectively capture the invariant priors of unseen data, significantly mitigating the effects of domain shifts. Extensive experiments demonstrate that our method outperforms other state-of-the-art approaches on two medical image segmentation benchmarks for both multi-source and single-source domain generalization tasks. The source code is available at https://github.com/Yore0/TTDG-MGM. Xingguo Lv, Xingbo Dong, Liwen Wang 0002, Jiewen Yang, Lei Zhao 0013, Bin Pu, Zhe Jin 0001, Xuejun Li 0001 |
CVPR | 5 |
| 2025 | EA-KD: Entropy-Based Adaptive Knowledge DistillationabstractKnowledge distillation (KD) enables a smaller 'student' model to mimic a larger 'teacher' model by transferring knowledge from the teacher's output or features. However, most KD methods treat all samples uniformly, overlooking the varying learning value of each sample and thereby limiting effectiveness. In this paper, we propose Entropy- based Adaptive Knowledge Distillation (EA-KD), a simple yet effective plug-and-play KD method that prioritizes learning from valuable samples. EA-KD quantifies each sample's learning value by strategically combining the entropy of the teacher and student output, then dynamically reweights the distillation loss to place greater emphasis on high-entropy samples. Extensive experiments across diverse KD frameworks and tasks-including image classification, object detection, and large language model (LLM) distillation-demonstrate that EA-KD consistently enhances performance, achieving state-of-the-art results with negligible computational cost. Our code is available at https://github.com/cpsu00/EA-KD. Chi-Ping Su, Ching-Hsun Tseng, Bin Pu, Lei Zhao 0013, Jiewen Yang, Zhuangzhuang Chen, Shin-Jye Lee |
ICCV | 4 |
| 2025 | Concept-Induced Graph Perception Model for Interpretable Diagnosis
Lei Zhao 0013, Changjian Chen, Bin Pu, Xiaoming Qi, Fengfeng Peng, Chunlian Wang, Kenli Li 0001, Guanghua Tan |
MICCAI (12) | 1 |
| 2025 | Anatomical Structure Few-Shot Detection Utilizing Enhanced Human Anatomy Knowledge in Ultrasound Images
Bocheng Liang, Ningshu Li, Lei Zhao 0013, Hao Li 0021, Fengwei Yang, Bin Pu |
MICCAI (5) | 4 |
| 2025 | Adaptive detection method for driver fatigue using facial multisource dynamic behavior fusion
Lei Zhao 0013, Chaoning Yu |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Low-light image enhancement with luminance duality
Xingguo Lv, Xingbo Dong, Jiewen Yang, Lei Zhao 0013, Bin Pu, Zhe Jin 0001 |
Knowl. Based Syst. | 4 |
| 2025 | PSFHS challenge report: Pubic symphysis and fetal head segmentation from intrapartum ultrasound images
Jieyun Bai, Zhanhong Ou, Gregor Köhler, Raphael Stock, Klaus H. Maier-Hein, Marawan Elbatel, Robert Martí, Xiaomeng Li 0001, Yaoyang Qiu, Panjie Gou, Gongping Chen, Lei Zhao 0013, Jianxun Zhang 0002, Yu Dai 0002, Fangyijie Wang, Guénolé C. M. Silvestre, Kathleen M. Curran, Hongkun Sun, Pengzhou Cai, Libin Lan, Dong Ni 0001, Mei Zhong, Gaowen Chen, Víctor M. Campello, Yaosheng Lu, Karim Lekadir |
Medical Image Anal. | 13 |
| 2025 | Corrigendum to "PSFHS challenge report: pubic symphysis and fetal head segmentation from intrapartum ultrasound images" [Medical Image Analysis 99 (2025),103353]
Jieyun Bai, Zhanhong Ou, Gregor Köhler, Raphael Stock, Klaus H. Maier-Hein, Marawan Elbatel, Robert Martí, Xiaomeng Li 0001, Yaoyang Qiu, Panjie Gou, Gongping Chen, Lei Zhao 0013, Jianxun Zhang 0002, Yu Dai 0002, Fangyijie Wang, Guénolé C. M. Silvestre, Kathleen M. Curran, Hongkun Sun, Pengzhou Cai, Libin Lan, Dong Ni 0001, Mei Zhong, Gaowen Chen, Víctor M. Campello, Yaosheng Lu, Karim Lekadir |
Medical Image Anal. | 13 |
| 2025 | TKR-FSOD: Fetal Anatomical Structure Few-Shot Detection Utilizing Topological Knowledge ReasoningabstractFetal multi-anatomical structure detection in ultrasound (US) images can clearly present the relationship and influence between anatomical structures, providing more comprehensive information about fetal organ structures and assisting sonographers in making more accurate diagnoses, widely used in structure evaluation. Recently, deep learning methods have shown superior performance in detecting various anatomical structures in ultrasound images, but still have the potential for performance improvement in categories where it is difficult to obtain samples, such as rare diseases. Few-shot learning has attracted a lot of attention in medical image analysis due to its ability to solve the problem of data scarcity. However, existing few-shot learning research in medical image analysis focuses on classification and segmentation, and the research on object detection has been neglected. In this paper, we propose a novel fetal anatomical structure few-shot detection method in ultrasound images, TKR-FSOD, which learns topological knowledge through a Topological Knowledge Reasoning Module to help the model reason about and detect anatomical structures. Furthermore, we propose a Discriminate Ability Enhanced Feature Learning Module that extracts abundant discriminative features to enhance the model's discriminative ability. Experimental results demonstrate that our method outperforms the state-of-the-art baseline methods, exceeding the second-best method with a maximum margin of 4.8% on 5-shot of split 1 under four-chamber cardiac view. Bocheng Liang, Bin Pu, Jiewen Yang, Lei Zhao 0013, Yanqing Kong, Lixian Yang, Rentie Zhang, Hao Li 0021, Shengli Li 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | ArmVR: Innovative Design Combining Virtual Reality Technology and Mechanical Equipment in Stroke Rehabilitation TherapyabstractThe rising incidence of stroke has created a significant global public health challenge. The immersive qualities of virtual reality (VR) technology, along with its distinct advantages, make it a promising tool for stroke rehabilitation. To address this challenge, developing VR-based upper limb rehabilitation systems has become a critical research focus. This study developed and evaluated an innovative ArmVR system that combines VR technology with rehabilitation hardware to improve recovery outcomes for stroke patients. Through comprehensive assessments, including neurofeedback, pressure feedback, and subjective feedback, the results suggest that VR technology has the potential to positively support the recovery of cognitive and motor functions. Different VR environments affect rehabilitation outcomes: forest scenarios aid emotional relaxation, while city scenarios better activate motor centers in stroke patients. The study also identified variations in responses among different user groups. Normal users showed significant changes in cognitive function, whereas stroke patients primarily experienced motor function recovery. These findings suggest that VR-integrated rehabilitation systems possess great potential, and personalized design can further enhance recovery outcomes, meet diverse patient needs, and ultimately improve quality of life. Jing Qu 0001, Lingguo Bu, Zhongxin Chen, Yalu Jin, Lei Zhao 0013, Shantong Zhu, Fenghe Guo |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Boosting the Transferability of Adversarial Examples via Adaptive Attention and Gradient Purification MethodsabstractDeep neural networks are shown to be vulnerable to adversarial examples. Recently, various methods have been proposed to improve the transferability of adversarial examples. However, most of the existing methods add perturbations to the whole image without discrimination, causing the visual quality of the adversarial examples to degrade drastically. In addition, existing attack methods ignore the gradient information of secondary features, which affects the accuracy of generating adversarial perturbations. In this work, we propose Adaptive Attention and Gradient Purification Attack (AAGP) to address such issues. Specifically, we judge the mean and standard deviation of the gradient values to find out where the model is interested. Since different models share similar regions of attention, adding perturbations only to such areas can reduce the addition of adversarial perturbation and can also lead to better transferability of adversarial examples to other models. In addition, we disrupt the correlation of pixels at the distribution of secondary features by random discarding pixels in low-attention areas, generating more transferable perturbations through more accurate gradient information. Experimental results on ImageNet show that our method enhances the visibility of the adversarial examples and their transferability compared with several advanced baselines. Liwen Wu, Lei Zhao 0013, Bin Pu, Xin Jin 0005, Shaowen Yao 0001 |
IJCNN | 2 |
| 2024 | A training and assessment system for human-computer interaction combining fNIRS and eye-tracking data
Jing Qu 0001, Lingguo Bu, Lei Zhao 0013 |
Adv. Eng. Informatics | 3 |
| 2024 | HICL: Hierarchical Intent Contrastive Learning for sequential recommendation
Yan Kang 0003, Yancong Yuan, Bin Pu, Yun Yang 0003, Lei Zhao 0013 |
Expert Syst. Appl. | 5 |
| 2024 | Boosting the Transferability of Ensemble Adversarial Attack via Stochastic Average Variance DescentabstractAdversarial examples have the property of transferring across models, which has created a great threat for deep learning models. To reveal the shortcomings in the existing deep learning models, the method of the ensemble has been introduced to the generating of transferable adversarial examples. However, most of the model ensemble attacks directly combine the different models’ output but ignore the large differences in optimization direction of them, which severely limits the transfer attack ability. In this work, we propose a new kind of ensemble attack method called stochastic average ensemble attack. Unlike the existing approach of averaging the outputs of each model as an integrated output, we continuously optimize the ensemble gradient in an internal loop using the model history gradient and the average gradient of different models. In this way, the adversarial examples can be updated in a more appropriate direction and make the crafted adversarial examples more transferable. Experimental results on ImageNet show that our method generates highly transferable adversarial examples and outperforms existing methods. Lei Zhao 0013, Zhizhi Liu, Sixing Wu, Liwen Wu, Bin Pu, Shaowen Yao 0001 |
IET Inf. Secur. | 1 |
| 2024 | The End-to-End Fetal Head Circumference Detection and Estimation in Ultrasound ImagesabstractIn prenatal examinations, the fetal head circumference (HC) measurement is essential for assessing fetal weight and health conditions. The sonographers obtain the fetal HC manually by fitting peripheral skull ellipse in clinical practice, which is highly subjective, time-consuming, and experience-dependent. Recently, many fetal HC automatic measurement algorithms have been proposed to improve workflow efficiency in prenatal examination. But most automatic measurement algorithms focus on using fetal head segmentation as an intermediate processing step, and HC estimation relies heavily on segmentation results, which causes the accumulation of errors in the above two stages. Independent of the segmentation method, we design a regression network to generate the oriented bounding box to detect the head contour, and directly obtain the fetal head parameters with a pixel-based ellipse regression (PER) loss. Moreover, an effective 3D attention mechanism is integrated into the network to estimate HC more precisely without adding parameters in complex ultrasound images. The extensive experimental results on the public HC18 and our clinical dataset show that the proposed network provides a feasible scheme for end-to-end estimating fetal HC, and avoids the mistake brought by the intermediary processes. Lei Zhao 0013, Ningshu Li, Guanghua Tan, Jianguo Chen 0001, Shengli Li 0001, Mingxing Duan |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2024 | TransFSM: Fetal Anatomy Segmentation and Biometric Measurement in Ultrasound Images Using a Hybrid TransformerabstractBiometric parameter measurements are powerful tools for evaluating a fetus's gestational age, growth pattern, and abnormalities in a 2D ultrasound. However, it is still challenging to measure fetal biometric parameters automatically due to the indiscriminate confusing factors, limited foreground-background contrast, variety of fetal anatomy shapes at different gestational ages, and blurry anatomical boundaries in ultrasound images. The performance of a standard CNN architecture is limited for these tasks due to the restricted receptive field. We propose a novel hybrid Transformer framework, TransFSM, to address fetal multi-anatomy segmentation and biometric measurement tasks. Unlike the vanilla Transformer based on a single-scale input, TransFSM has a deformable self-attention mechanism so it can effectively process multi-scale information to segment fetal anatomy with irregular shapes and different sizes. We devised a BAD to capture more intrinsic local details using boundary-wise prior knowledge, which compensates for the defects of the Transformer in extracting local features. In addition, a Transformer auxiliary segment head is designed to improve mask prediction by learning the semantic correspondence of the same pixel categories and feature discriminability among different pixel categories. Extensive experiments were conducted on clinical cases and benchmark datasets for anatomy segmentation and biometric measurement tasks. The experiment results indicate that our method achieves state-of-the-art performance in seven evaluation metrics compared with CNN-based, Transformer-based, and hybrid approaches. By Knowledge distillation, the proposed TransFSM can create a more compact and efficient model with high deploying potential in resource-constrained scenarios. Our study serves as a unified framework for biometric estimation across multiple anatomical regions to monitor fetal growth in clinical practice. Lei Zhao 0013, Guanghua Tan, Bin Pu, Qianghui Wu, Hongliang Ren 0001, Kenli Li 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | FARN: Fetal Anatomy Reasoning Network for Detection With Global Context Semantic and Local Topology RelationshipabstractAccurate recognition of fetal anatomical structure is a pivotal task in ultrasound (US) image analysis. Sonographers naturally apply anatomical knowledge and clinical expertise to recognizing key anatomical structures in complex US images. However, mainstream object detection approaches usually treat each structure recognition separately, overlooking anatomical correlations between different structures in fetal US planes. In this work, we propose a Fetal Anatomy Reasoning Network (FARN) that incorporates two kinds of relationship forms: a global context semantic block summarized with visual similarity and a local topology relationship block depicting structural pair constraints. Specifically, by designing the Adaptive Relation Graph Reasoning (ARGR) module, anatomical structures are treated as nodes, with two kinds of relationships between nodes modeled as edges. The flexibility of the model is enhanced by constructing the adaptive relationship graph in a data-driven way, enabling adaptation to various data samples without the need for predefined additional constraints. The feature representation is further refined by aggregating the outputs of the ARGR module. Comprehensive experimental results demonstrate that FARN achieves promising performance in detecting 37 anatomical structures across key US planes in tertiary obstetric screening. FARN effectively utilizes key relationships to improve detection performance, demonstrates robustness to small-scale, similar, and indistinct structures, and avoids some detection errors that deviate from anatomical norms. Overall, our study serves as a resource for developing efficient and concise approaches to model inter-anatomy relationships. Lei Zhao 0013, Guanghua Tan, Qianghui Wu, Bin Pu, Hongliang Ren 0001, Shengli Li 0001, Kenli Li 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Needle Trajectory Prediction for Percutaneous Kidney Biopsy in 5G-Powered Teleultrasound Navigation SystemabstractNeedle insertion is a critical component of many remote surgical procedures, including biopsies, injections, neurosurgery, and brachytherapy cancer treatments. However, precise visualization of the biopsy needle trajectory remains challenging due to specular reflection, speckle noise, and needle-like anatomical features. This paper proposes a visual feedback prediction framework for ultrasound-assisted percutaneous kidney biopsy in 5G remote surgery, aiming to enhance operator confidence, reduce procedure time, and minimize the risk of unintended bleeding. Building upon this framework, we design a Lightweight-Accuracy Needle Trajectory Prediction (LA-NTP) model by minimizing the backbone and optimizing the multi-module prediction process, incorporating innovative training strategies (i.e., angle-aware geometric and trajectory augmentation losses). The experimental results demonstrate that it achieves competitive performance with only 20.3% of the model size of the previous best real-time method and a 3.7-fold increase in inference speed. Even in challenging scenarios involving large insertion depths and steep angles, our method provides stable and precise navigation. Lei Zhao 0013, Guanghua Tan, Jiewen Lai, Chwee Ming Lim, Weng Kin Wong, Hongliang Ren 0001, Kenli Li 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2022 | A design method for an intelligent manufacturing and service system for rehabilitation assistive devices and special groups
Zilin Wang 0004, Li-Zhen Cui 0001, Wei Guo 0017, Lei Zhao 0013, Xiaosong Gu, Weizhong Tang, Lingguo Bu, Weiming Huang 0001 |
Adv. Eng. Informatics | 4 |
| 2022 | An ultrasound standard plane detection model of fetal head based on multi-task learning and hybrid knowledge graph
Lei Zhao 0013, Kenli Li 0001, Bin Pu, Jianguo Chen 0001, Shengli Li 0001, Xiangke Liao |
Future Gener. Comput. Syst. | 1 |
| 2021 | Driver behavior detection via adaptive spatial attention mechanism
Lei Zhao 0013, Lingguo Bu, Su Han |
Adv. Eng. Informatics | 1 |
| 2020 | Driver drowsiness recognition via transferred deep 3D convolutional network and state probability vector
Lei Zhao 0013, Zengcai Wang, Huanbing Gao |
Multim. Tools Appl. | 1 |
| 2011 | InfoNetOLAPer: Integrating InfoNetWarehouse and InfoNetCube with InfoNetOLAP
Chuan Li 0002, Philip S. Yu, Lei Zhao 0013, Wangqun Lin |
Proc. VLDB Endow. | 3 |