VLDB 2026 Research / reviewers in the wild / expert
Yaqi Wang 0002
dblp:98/7674-2
· DBLP profile ↗
26ranked-venue papers
5as first author
26since 2021 · last 2026
0000-0002-4627-3392ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-annotation agreement and prediction consistency networks: Improving semi-supervised segmentation of medical images with ambiguous boundaries
Shuai Wang 0003, Tengjin Weng, Yang Shen 0011, Zhidong Zhao, Yixiu Liu, Pengfei Jiao, Zhiming Cheng, Yaoqi Sun, Yaqi Wang 0002 |
Artif. Intell. Medicine | 11 |
| 2026 | Graph Neural Networks Based Analog Circuit Link Prediction
Guanyuan Pan, Tiansheng Zhou, Jianxiang Zhao, Yugui Lin, Bingtao Ma, Yaqi Wang 0002, Pietro Liò, Shuai Wang 0003 |
Eng. Appl. Artif. Intell. | 7 |
| 2026 | MICCAI STS 2024 challenge: Semi-supervised instance-level tooth segmentation in panoramic X-ray and CBCT imagesabstractOrthopantomogram (OPGs) and Cone-Beam Computed Tomography (CBCT) are vital for dentistry, but creating large datasets for automated tooth segmentation is hindered by the labor-intensive process of manual instance-level annotation. This research aimed to benchmark and advance semi-supervised learning (SSL) as a solution for this data scarcity problem. We organized the 2nd Semi-supervised Teeth Segmentation (STS 2024) Challenge at MICCAI 2024. We provided a large-scale dataset comprising over 90,000 2D images and 3D axial slices, which includes 2380 OPG images and 330 CBCT scans, all featuring detailed instance-level FDI annotations on part of the data. The challenge attracted 114 (OPG) and 106 (CBCT) registered teams. To ensure algorithmic excellence and full transparency, we rigorously evaluated the valid, open-source submissions from the top 10 (OPG) and top 5 (CBCT) teams, respectively. All successful submissions were deep learning-based SSL methods. The winning semi-supervised models demonstrated impressive performance gains over a fully-supervised nnU-Net baseline trained only on the labeled data. For the 2D OPG track, the top method improved the Instance Affinity (IA) score by over 44 percentage points. For the 3D CBCT track, the winning approach boosted the Instance Dice score by 61 percentage points. This challenge demonstrates the potential benefit benefit of SSL for complex, instance-level medical image segmentation tasks where labeled data is scarce. The most effective approaches consistently leveraged hybrid semi-supervised frameworks that combined knowledge from foundational models like SAM with multi-stage, coarse-to-fine refinement pipelines. Both the challenge dataset and the participants' submitted code have been made publicly available on GitHub (https://github.com/ricoleehduu/STS-Challenge-2024), ensuring transparency and reproducibility. Yaqi Wang 0002, Jun Liu 0027, Jiaxue Ni, Hongyuan Zhang 0002, Jin Liu 0025, Can Han, Kaiwen Fu, Changkai Ji, Xinxu Cai, Junqiang Chen, Qianni Zhang, Dahong Qian, Shuai Wang 0003, Huiyu Zhou 0001 |
Medical Image Anal. | 1 |
| 2026 | PolyS-Net: A joint learning framework for depth-aware and scale-aware polyp size estimation
Sijia Du, Yaqi Wang 0002, Chen Liu 0026, Jun Wang 0041, Ruilan Wang, Huiyu Zhou 0001, Qingwei Zhang, Dahong Qian |
Pattern Recognit. | 3 |
| 2026 | MICCAI 2023 STS Challenge: A retrospective study of semi-supervised approaches for teeth segmentationabstractComputer-aided diagnosis greatly enhances personalized treatment planning and diagnostic efficiency by providing accurate dental anatomy through teeth segmentation. However, it still constrained by the scarcity of high-quality annotated dental datasets. To address this issue, this paper presents a dataset combining both 2D panoramic X-rays with over 6,500 images and 3D CBCT with over 580 volumes (88,500+ slices) to support the Semi-supervised Teeth Segmentation (STS) Challenge, which includes partially meticulous annotations and covers all age groups. Moreover, multi-phase semi-supervised teeth segmentation algorithms and high-confidence pseudo-labels refinement strategies were proposed by competitors during this challenge. Algorithms were verified on this proposed dataset and good segmentation performance were achieved, over 93+ and 80+ Dice score were obtained for top three 2D and 3D participants, demonstrating the high quality of this proposed dataset. This paper also summarizes the diverse methods employed by the top-ranking teams in the MICCAI 2023 STS Challenge. Our dataset is publicly accessible through Zenodo ( https://zenodo.org/records/10597292 ), and the participants’ code is hosted on GitHub ( https://github.com/ricoleehduu/STS-Challenge ). Yaqi Wang 0002, Shuai Wang 0003, Dahong Qian, Hongyuan Zhang 0002, Ruilong Dan, Qianni Zhang, Xingru Huang, Jun Liu 0027, Zhean Ma, Weiwei Cui 0003, Shan Luo 0003, Chengkai Wang, Jiaxue Ni, Dongyun Liu, Zhouhao Lin, Chunshi Wang, Qiupu Chen, Mingqian Li, Huiyu Zhou 0001, Qun Jin |
Pattern Recognit. | 1 |
| 2026 | An Unsupervised Learning-Based Multidimensional Scaling Approach for Placement and Routing of Monolithic Microwave Integrated CircuitabstractThe layout method for Radio Frequency/Monolithic Microwave Integrated Circuit (RF/MMIC) is one of the key components in achieving MMIC design automation. Our work proposes a multi-stage progressive automated layout framework addressing MMIC layout challenges under 0.25 μmGaAs pHEMT technology. This framework employs dimensionality reduction from unsupervised learning, integrates practical layout design rules, and achieves automated layout through analytical optimization methods. Applied to multiple cases including filters, low-noise amplifiers (LNAs), and power amplifiers (PAs), the method successfully generated manufacturable layouts; the process achieves end-to-end automation from placement to routing, producing layouts that are verified to be DRC-clean. Simulation-verified results demonstrate comparable performance to manual designs while achieving ≥10× speedup over conventional methods. We believe this work contributes valuable attempts toward automated MMIC layout generation and advances progress in MMIC design automation. Yaqi Wang 0002, Bin You, Jun Liu 0027 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | End-to-End Echocardiogram Video Analysis for Automated Fetal Congenital Heart Disease DiagnosisabstractFetal congenital heart disease (FCHD) is a major cause of perinatal mortality, yet accurate prenatal echocardiographic screening remains challenged by complex fetal anatomy, variable fetal posture, and an abundance of non-diagnostic frames. Although deep learning has accelerated automated echocardiography, dedicated solutions for fetal imaging are scarce. We compiled 388 expertly annotated fetal 2-D echocardiogram videos (284 normal, 104 FCHD) and introduce an end-to-end framework that: (i) isolates diagnostic frames with a Monte Carlo Keyframe Selector, (ii) pinpoints salient anatomy with an Evidence Region Extractor, and (iii) integrates local and global cues via a Cross Retrieval Module. Extensive evaluations demonstrate substantial gains in diagnostic robustness and accuracy, underscoring the framework's potential to advance prenatal cardiac screening. Can Han, Chenyu Zhu, Tan Zhou, Yaqi Wang 0002, Shiya Yao, Baoying Ye, Dahong Qian |
BIBM | 5 |
| 2025 | Volumetric Axial Disentanglement Enabling Advancing in Medical Image SegmentationabstractInformation retrieved from three dimensions is treated uniformly in CNN-based volumetric segmentation methods. However, such neglect of axial disparities fails to capture true spatio-temporal variations. This paper introduces the volumetric axial disentanglement to address the disparities in spatial information along different axial dimensions. Building on this concept, we propose the Post-Axial Refiner (PaR) module to refine segmentation masks by implementing axial disentanglement on the specific axis of the volumetric medical sequences. As a plug-and-play enhancement to existing volumetric segmentation architecture, PaR further utilizes specialized attention approaches to learn disentangled post-decoding features, enhancing spatial representation and structural detail. Validation on various datasets demonstrates PaR's consistent elevation of segmentation precision and boundary clarity across 11 baselines and different imaging modalities, achieving state-of-the-art performance on multiple datasets. Experimental tests demonstrate the ability of volumetric axial disentanglement to refine the segmentation of volumetric medical images. Code is released at https://github.com/IMOP-lab/PaR-Pytorch. Xingru Huang, Jian Huang 0015, Tianyun Zhang, Yaqi Wang 0002, Ruipu Tang, Shaowei Jiang, Jin Liu 0025, Renjie Ruan, Xiaoshuai Zhang |
IJCAI | 6 |
| 2025 | Robust Real-Time Endoscopic Stereo Matching Under Fuzzy Tissue Boundaries
Can Han, Sijia Du, Yaqi Wang 0002, Dahong Qian |
PRCV (14) | 4 |
| 2025 | A spatial-spectral and temporal dual prototype network for motor imagery brain-computer interface
Can Han, Chen Liu 0026, Jun Wang 0072, Yaqi Wang 0002, Crystal Cai, Dahong Qian |
Knowl. Based Syst. | 4 |
| 2025 | A Novel Multi-Modal Population-Graph Based Framework for Patients of Esophageal Squamous Cell Cancer Prognostic Risk PredictionabstractPrognostic risk prediction is pivotal for clinicians to appraise the patient's esophageal squamous cell cancer (ESCC) progression status precisely and tailor individualized therapy treatment plans. Currently, CT-based multi-modal prognostic risk prediction methods have gradually attracted the attention of researchers for their universality, which is also able to be applied in scenarios of preoperative prognostic risk assessment in the early stages of cancer. However, much of the current work focuses only on CT images of the primary tumor, ignoring the important role that CT images of lymph nodes play in prognostic risk prediction. Additionally, it is important to consider and explore the inter-patient feature similarity in prognosis when developing models. To solve these problems, we proposed a novel multi-modal population-graph based framework leveraging CT images including primary tumor and lymph nodes combined with clinical, hematology, and radiomics data for ESCC prognostic risk prediction. A patient population graph was constructed to excavate the homogeneity and heterogeneity of inter-patient feature embedding. Moreover, a novel node-level multi-task joint loss was proposed for graph model optimization through a supervised-based task and an unsupervised-based task. Sufficient experimental results show that our model achieved state-of-the-art performance compared with other baseline models as well as the gold standard on discriminative ability, risk stratification, and clinical utility. Shuai Wang 0003, Yaqi Wang 0002, Chengkai Wang, Huiyu Zhou 0001, Yatao Zhang |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | MMFusion: Multi-modality Diffusion Model for Lymph Node Metastasis Diagnosis in Esophageal Cancer
Chengkai Wang, Huiyu Zhou 0001, Yatao Zhang, Yaqi Wang 0002, Shuai Wang 0003 |
MICCAI (5) | 6 |
| 2024 | PGKD-Net: Prior-guided and Knowledge Diffusive Network for Choroid SegmentationabstractThe thickness of the choroid is considered to be an important indicator of clinical diagnosis. Therefore, accurate choroid segmentation in retinal OCT images is crucial for monitoring various ophthalmic diseases. However, this is still challenging due to the blurry boundaries and interference from other lesions. To address these issues, we propose a novel prior-guided and knowledge diffusive network (PGKD-Net) to fully utilize retinal structural information to highlight choroidal region features and boost segmentation performance. Specifically, it is composed of two parts: a Prior-mask Guided Network (PG-Net) for coarse segmentation and a Knowledge Diffusive Network (KD-Net) for fine segmentation. In addition, we design two novel feature enhancement modules, Multi-Scale Context Aggregation (MSCA) and Multi-Level Feature Fusion (MLFF). The MSCA module captures the long-distance dependencies between features from different receptive fields and improves the model's ability to learn global context. The MLFF module integrates the cascaded context knowledge learned from PG-Net to benefit fine-level segmentation. Comprehensive experiments are conducted to evaluate the performance of the proposed PGKD-Net. Experimental results show that our proposed method achieves superior segmentation accuracy over other state-of-the-art methods. Our code is made up publicly available at: https://github.com/yzh-hdu/choroid-segmentation. Yaqi Wang 0002, Zehua Yang, Xindi Liu, Dechao Chen, Gangyong Jia, Juan Ye, Xingru Huang |
Artif. Intell. Medicine | 1 |
| 2024 | AMSC-Net: Anatomy and multi-label semantic consistency network for semi-supervised fluid segmentation in retinal OCTabstractAutomated segmentation of pathological fluid regions is crucial for digital diagnosis and individualized therapy under optical coherence tomography (OCT) images. Nonetheless, methods that rely on enormous annotations impede their clinical applications, as pixel-wise fine-grained labels are extraordinarily costly and require the expertise of ophthalmology. Though there exist works based on consistency-based and pseudo-label-based methods for reducing annotation reliance, the underutilization of consistency mechanisms and ignorance of fluid anatomical structures result in sub-optimal segmentation performance. In this work, we proposed a novel AMSC-Net devoted to semi-supervised fluid segmentation and achieves a 73.95% Dice score with 5% labeled data. Concretely, we developed a Heterogeneous Architecture Consistency (HAC) strategy based on our dual decoders enjoying different inductive biases. Moreover, we invented a Multi-label Semantic Consistency Loss (MSC-Loss) module in hierarchical semantic features and an Anatomy Contour Consistency Loss (ACC-Loss) module integrating anatomical constraint. These modules complementarily reinforce the quality of pseudo labels and boost semi-supervised training for robust segmentation results. To verify the superiority and bedside potential of our proposed AMSC-Net, we collected a large-scale fluid segmentation dataset composed of 22517 OCT images. Extensive quantitative and qualitative experiments validated the efficacy of our AMSC-Net with multiple novel techniques. Also, experimental results in a public fluid segmentation dataset demonstrated that our method achieves state-of-the-art performance. Code will be available at: https://github.com/ZeroOneGame/S4_Fluid_OCT. Yaqi Wang 0002, Ruilong Dan, Shan Luo 0003, Lingling Sun, Qicen Wu, Kangming Yan, Xin Ye 0006, Dingguo Yu |
Expert Syst. Appl. | 1 |
| 2024 | Mmy-net: a multimodal network exploiting image and patient metadata for simultaneous segmentation and diagnosis
Renshu Gu, Yueyu Zhang, Lisha Wang, Dechao Chen, Yaqi Wang 0002, Ruiquan Ge, Zicheng Jiao, Juan Ye, Gangyong Jia, Linyan Wang |
Multim. Syst. | 5 |
| 2024 | A Self-Supervised Learning Based Framework for Eyelid Malignant Melanoma Diagnosis in Whole Slide ImagesabstractEyelid malignant melanoma (MM) is a rare disease with high mortality. Accurate diagnosis of such disease is important but challenging. In clinical practice, the diagnosis of MM is currently performed manually by pathologists, which is subjective and biased. Since the heavy manual annotation workload, most pathological whole slide image (WSI) datasets are only partially labeled (without region annotations), which cannot be directly used in supervised deep learning. For these reasons, it is of great practical significance to design a laborsaving and high data utilization diagnosis method. In this paper, a self-supervised learning (SSL) based framework for automatically detecting eyelid MM is proposed. The framework consists of a self-supervised model for detecting MM areas at the patch-level and a second model for classifying lesion types at the slide level. A squeeze-excitation (SE) attention structure and a feature-projection (FP) structure are integrated to boost learning on details of pathological images and improve model performance. In addition, this framework also provides visual heatmaps with high quality and reliability to highlight the likely areas of the lesion to assist the evaluation and diagnosis of the eyelid MM. Extensive experimental results on different datasets show that our proposed method outperforms other state-of-the-art SSL and fully supervised methods at both patch and slide levels when only a subset of WSIs are annotated. It should be noted that our method is even comparable to supervised methods when all WSIs are fully annotated. To the best of our knowledge, our work is the first SSL method for automatic diagnosis of MM at the eyelid and has a great potential impact on reducing the workload of human annotations in clinical practice. Zijing Jiang, Linyan Wang, Yaqi Wang 0002, Gangyong Jia, Guodong Zeng, Jun Wang 0072, Dechao Chen, Guiping Qian, Qun Jin |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | Sketch-Supervised Histopathology Tumour Segmentation: Dual CNN-Transformer With Global Normalised CAMabstractDeep learning methods are frequently used in segmenting histopathology images with high-quality annotations nowadays. Compared with well-annotated data, coarse, scribbling-like labelling is more cost-effective and easier to obtain in clinical practice. The coarse annotations provide limited supervision, so employing them directly for segmentation network training remains challenging. We present a sketch-supervised method, called DCTGN-CAM, based on a dual CNN-Transformer network and a modified global normalised class activation map. By modelling global and local tumour features simultaneously, the dual CNN-Transformer network produces accurate patch-based tumour classification probabilities by training only on lightly annotated data. With the global normalised class activation map, more descriptive gradient-based representations of the histopathology images can be obtained, and inference of tumour segmentation can be performed with high accuracy. Additionally, we collect a private skin cancer dataset named BSS, which contains fine and coarse annotations for three types of cancer. To facilitate reproducible performance comparison, experts are also invited to label coarse annotations on the public liver cancer dataset PAIP2019. On the BSS dataset, our DCTGN-CAM segmentation outperforms the state-of-the-art methods and achieves 76.68 % IOU and 86.69 % Dice scores on the sketch-based tumour segmentation task. On the PAIP2019 dataset, our method achieves a Dice gain of 8.37 % compared with U-Net as the baseline network. Yilong Li 0002, Linyan Wang, Xingru Huang, Yaqi Wang 0002, Ruiquan Ge, Huiyu Zhou 0001, Juan Ye, Qianni Zhang |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | GOMPS: Global Attention-based Ophthalmic Image Measurement and Postoperative Appearance Prediction SystemabstractAccurate measurements of ophthalmic parameters and postoperative appearance prediction are essential for the diagnosis and treatment of many ophthalmic diseases. Nevertheless, it remains challenging due to (1) inconsistent ophthalmic image sampling standards, including ocular-camera distance, facial angle, and patient number, (2) complicated ocular morphology, such as subconjunctival hemorrhage, ocular movements, lighting effects, and morphological aging. It is difficult for a model to measure parameters and make predictions in variable sampling methods and morphology conditions. Therefore, the Global attention-based Ophthalmic Image Measurement and Postoperative Appearance Prediction System (GOMPS) is proposed, which quantifies ophthalmic image parameters to diagnose disease and simultaneously predict postoperative appearance of blepharoptosis. By perceiving the global structure of the ophthalmic image, GOMPS makes logical inference predictions of the sclera and cornea morphology, to overcome the above difficulties. Concretely, a global attention unit (GAU) and a novel global attention structure-aware network (GASA-Net) are designed to enhance GOMPS’s global structure awareness ability to perform logical reasoning. Extensive experimental results on our collected ophthalmic dataset for diagnosis & prediction (OD2P) demonstrate that GOMPS surpasses the state-of-the-art methods in segmentation accuracy and achieves the current optimal performance in measurement and postoperative prediction under many clinical scenes. Xingru Huang, Lixia Lou, Ruilong Dan, Lingxiao Chen, Guodong Zeng, Gangyong Jia, Qun Jin, Juan Ye, Yaqi Wang 0002 |
Expert Syst. Appl. | 11 |
| 2023 | ASCAM-Former: Blind image quality assessment based on adaptive spatial & channel attention merging transformer and image to patch weights sharingabstractBlind Image Quality Assessment (BIQA) is a challenging, unsolved research topic which is crucial for analyzing, understanding, and improving visual experience. Recently, transformer-based BIQA models are drawing increasing attention due to their powerful capacity in modeling global dependencies amongst tokens. However, existing works tend to apply self-attention mechanism for exploring the spatial dependencies whilst neglecting the impact of channel-wise self-attention. In this paper, we explore the feasibility of incorporating attention mechanism in a channel-wise manner for BIQA. By systematically studying the interactions between channel-wise and spatial-wise attention, an adaptive spatial and channel attention merging Transformer (ASCAM-Former) is then proposed for aggregating both the spatial-wise and channel-wise attention information. In addition, to accommodate IQA datasets containing both image and patch quality labels, an image to patch weights sharing (I2PWS) scheme is designed to take advantage of local quality learning tasks for reinforcing the learning of global quality, and vice versa. The experimental results indicate that channel-wise attention mechanism is as competitive as spatial-wise for IQA tasks, and the proposed ASCAM-Former yield accurate prediction on both authentically and synthetically distorted image quality datasets. Suiyu Zhang, Yaqi Wang 0002, Dingguo Yu |
Expert Syst. Appl. | 3 |
| 2023 | POST-IVUS: A perceptual organisation-aware selective transformer framework for intravascular ultrasound segmentationabstractIntravascular ultrasound (IVUS) is recommended in guiding coronary intervention. The segmentation of coronary lumen and external elastic membrane (EEM) borders in IVUS images is a key step, but the manual process is time-consuming and error-prone, and suffers from inter-observer variability. In this paper, we propose a novel perceptual oganisation-aware selective transformer framework that can achieve accurate and robust segmentation of the vessel walls in IVUS images. In this framework, temporal context-based feature encoders extract efficient motion features of vessels. Then, a perceptual oganisation-aware selective transformer module is proposed to extract accurate boundary information, supervised by a dedicated boundary loss. The obtained EEM and lumen segmentation results will be fused in a temporal constraining and fusion module, to determine the most likely correct boundaries with robustness to morphology. Our proposed methods are extensively evaluated in non-selected IVUS sequences, including normal, bifurcated, and calcified vessels with shadow artifacts. The results show that the proposed methods outperform the state-of-the-art, with a Jaccard measure of 0.92 for lumen and 0.94 for EEM on the IVUS 2011 open challenge dataset. This work has been integrated into a software QCU-CMS2 to automatically segment IVUS images in a user-friendly environment. Xingru Huang, Retesh Bajaj, Yilong Li 0002, Xin Ye 0006, Ji Lin 0004, Francesca Pugliese, Anantharaman Ramasamy, Yaqi Wang 0002, Ryo Torii, Jouke Dijkstra, Huiyu Zhou 0001, Christos V. Bourantas, Qianni Zhang |
Medical Image Anal. | 9 |
| 2023 | Stimulus-guided adaptive transformer network for retinal blood vessel segmentation in fundus imagesabstractAutomated retinal blood vessel segmentation in fundus images provides important evidence to ophthalmologists in coping with prevalent ocular diseases in an efficient and non-invasive way. However, segmenting blood vessels in fundus images is a challenging task, due to the high variety in scale and appearance of blood vessels and the high similarity in visual features between the lesions and retinal vascular. Inspired by the way that the visual cortex adaptively responds to the type of stimulus, we propose a Stimulus-Guided Adaptive Transformer Network (SGAT-Net) for accurate retinal blood vessel segmentation. It entails a Stimulus-Guided Adaptive Module (SGA-Module) that can extract local-global compound features based on inductive bias and self-attention mechanism. Alongside a light-weight residual encoder (ResEncoder) structure capturing the relevant details of appearance, a Stimulus-Guided Adaptive Pooling Transformer (SGAP-Former) is introduced to reweight the maximum and average pooling to enrich the contextual embedding representation while suppressing the redundant information. Moreover, a Stimulus-Guided Adaptive Feature Fusion (SGAFF) module is designed to adaptively emphasize the local details and global context and fuse them in the latent space to adjust the receptive field (RF) based on the task. The evaluation is implemented on the largest fundus image dataset (FIVES) and three popular retinal image datasets (DRIVE, STARE, CHASEDB1). Experimental results show that the proposed method achieves a competitive performance over the other existing method, with a clear advantage in avoiding errors that commonly happen in areas with highly similar visual features. The sourcecode is publicly available at: https://github.com/Gins-07/SGAT. Ji Lin 0004, Xingru Huang, Huiyu Zhou 0001, Yaqi Wang 0002, Qianni Zhang |
Medical Image Anal. | 4 |
| 2023 | MSCA-Net: Multi-scale contextual attention network for skin lesion segmentationabstractLesion segmentation algorithms automatically outline lesion areas in medical images, facilitating more effective identification and assessment of the clinically relevant features, and improving the efficacy and diagnosis accuracy. However, most fully convolutional network based segmentation methods suffer from spatial and contextual information loss when decreasing image resolution. To overcome this shortcoming, this paper proposes a skin lesion segmentation model , namely, the Multi-Scale Contextual Attention Network (MSCA-Net), which can exploit the multi-scale contextual information in images. Inspired by the skip connection of U-Net, we design a multi-scale bridge (MSB) module which interacts with multi-scale features to effectively fuse the multi-scale contextual information of the encoder and decoder path features. We further propose a global-local channel spatial attention module (GL-CSAM), aiming at capturing global contextual information. In addition, to take full advantage of the multi-scale features of the decoder, we propose a scale-aware deep supervision (SADS) module to achieve hierarchical iterative deep supervision. Comprehensive experimental results on the public dataset of ISIC 2017, ISIC 2018, and PH 2 show that our proposed method outperforms other state-of-the-art methods, demonstrating the efficacy of our method in skin lesion segmentation. Our code is available at https://github.com/YonghengSun1997/MSCA-Net . Yongheng Sun, Duwei Dai, Qianni Zhang, Yaqi Wang 0002, Songhua Xu, Chunfeng Lian |
Pattern Recognit. | 4 |
| 2023 | CDNet: Contrastive Disentangled Network for Fine-Grained Image Categorization of Ocular B-Scan UltrasoundabstractPrecise and rapid categorization of images in the B-scan ultrasound modality is vital for diagnosing ocular diseases. Nevertheless, distinguishing various diseases in ultrasound still challenges experienced ophthalmologists. Thus a novel contrastive disentangled network (CDNet) is developed in this work, aiming to tackle the fine-grained image categorization (FGIC) challenges of ocular abnormalities in ultrasound images, including intraocular tumor (IOT), retinal detachment (RD), posterior scleral staphyloma (PSS), and vitreous hemorrhage (VH). Three essential components of CDNet are the weakly-supervised lesion localization module (WSLL), contrastive multi-zoom (CMZ) strategy, and hyperspherical contrastive disentangled loss (HCD-Loss), respectively. These components facilitate feature disentanglement for fine-grained recognition in both the input and output aspects. The proposed CDNet is validated on our ZJU Ocular Ultrasound Dataset (ZJUOUSD), consisting of 5213 samples. Furthermore, the generalization ability of CDNet is validated on two public and widely-used chest X-ray FGIC benchmarks. Quantitative and qualitative results demonstrate the efficacy of our proposed CDNet, which achieves state-of-the-art performance in the FGIC task. Ruilong Dan, Gangyong Jia, Shuai Wang 0003, Ruiquan Ge, Guiping Qian, Qun Jin, Juan Ye, Yaqi Wang 0002 |
IEEE J. Biomed. Health Informatics | 11 |
| 2022 | AGMB-Transformer: Anatomy-Guided Multi-Branch Transformer Network for Automated Evaluation of Root Canal TherapyabstractAccurate evaluation of the treatment result on X-ray images is a significant and challenging step in root canal therapy since the incorrect interpretation of the therapy results will hamper timely follow-up which is crucial to the patients’ treatment outcome. Nowadays, the evaluation is performed in a manual manner, which is time-consuming, subjective, and error-prone. In this article, we aim to automate this process by leveraging the advances in computer vision and artificial intelligence, to provide an objective and accurate method for root canal therapy result assessment. A novel anatomy-guided multi-branch Transformer (AGMB-Transformer) network is proposed, which first extracts a set of anatomy features and then uses them to guide a multi-branch Transformer network for evaluation. Specifically, we design a polynomial curve fitting segmentation strategy with the help of landmark detection to extract the anatomy features. Moreover, a branch fusion module and a multi-branch structure including our progressive Transformer and Group Multi-Head Self-Attention (GMHSA) are designed to focus on both global and local features for an accurate diagnosis. To facilitate the research, we have collected a large-scale root canal therapy evaluation dataset with 245 root canal therapy X-ray images, and the experiment results show that our AGMB-Transformer can improve the diagnosis accuracy from 57.96% to 90.20% compared with the baseline network. The proposed AGMB-Transformer can achieve a highly accurate evaluation of root canal therapy. To our best knowledge, our work is the first to perform automatic root canal therapy evaluation and has important clinical value to reduce the workload of endodontists. Guodong Zeng, Jun Wang 0041, Qun Jin, Lingling Sun, Qianni Zhang, Qisi Lian, Guiping Qian, Neng Xia, Ruizi Peng, Shuai Wang 0003, Yaqi Wang 0002 |
IEEE J. Biomed. Health Informatics | 14 |
| 2021 | Enhanced Diagnosis of Pneumothorax with an Improved Real-Time Augmentation for Imbalanced Chest X-rays Data Based on DCNNabstractPneumothorax is a common pulmonary disease that can lead to dyspnea and can be life-threatening. X-ray examination is the main means to diagnose this disease. Computer-aided diagnosis of pneumothorax on chest X-ray, as a prerequisite for a timely cure, has been widely studied, but it is still not satisfactory to achieve highly accurate results. In this paper, an image classification algorithm based on the deep convolutional neural network (DCNN) is proposed for high-resolution medical image analysis of pneumothorax X-rays, which features a Network In Network (NIN) for cleaning the data, random histogram equalization data augmentation processing, and a DCNN. The experimental results indicate that the proposed method can effectively increase the correct diagnosis rate of pneumothorax, and the Area under Curve (AUC) of the test verified in the experiment is 0.9844 on ZJU-2 test data and 0.9906 on the ChestX-ray14, respectively. In addition, a large number of atmospheric pleura samples are visualized and analyzed based on the experimental results and in-depth learning characteristics of the algorithm. The analysis results verify the validity of feature extraction for the network. Combined with the results of these two aspects, the proposed X-ray image processing algorithm can effectively improve the classification accuracy of pneumothorax photographs. Yaqi Wang 0002, Lingling Sun, Qun Jin |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2021 | Multiscale Attention Guided Network for COVID-19 Diagnosis Using Chest X-Ray ImagesabstractCoronavirus disease 2019 (COVID-19) is one of the most destructive pandemic after millennium, forcing the world to tackle a health crisis. Automated lung infections classification using chest X-ray (CXR) images could strengthen diagnostic capability when handling COVID-19. However, classifying COVID-19 from pneumonia cases using CXR image is a difficult task because of shared spatial characteristics, high feature variation and contrast diversity between cases. Moreover, massive data collection is impractical for a newly emerged disease, which limited the performance of data thirsty deep learning models. To address these challenges, Multiscale Attention Guided deep network with Soft Distance regularization (MAG-SD) is proposed to automatically classify COVID-19 from pneumonia CXR images. In MAG-SD, MA-Net is used to produce prediction vector and attention from multiscale feature maps. To improve the robustness of trained model and relieve the shortage of training data, attention guided augmentations along with a soft distance regularization are posed, which aims at generating meaningful augmentations and reduce noise. Our multiscale attention model achieves better classification performance on our pneumonia CXR image dataset. Plentiful experiments are proposed for MAG-SD which demonstrates its unique advantage in pneumonia classification over cutting-edge models. The code is available at https://github.com/JasonLeeGHub/MAG-SD. Jingxiong Li, Yaqi Wang 0002, Shuai Wang 0003, Jun Wang 0041, Jun Liu 0027, Qun Jin, Lingling Sun |
IEEE J. Biomed. Health Informatics | 2 |